Compression of nucleotide sequencing data

By converting read depth values into ordinal class assignments and mapping them to reference sequences, large nucleotide datasets are efficiently processed and analyzed, reducing time and storage needs while enabling comprehensive genetic insights.

JP2026501259APending Publication Date: 2026-01-14KEYGENE NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025536431
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-22
Filing Date
2023-12-22
Publication Date
2026-01-14

AI Technical Summary

Technical Problem

Large-scale nucleotide sequencing datasets are expensive and time-consuming to analyze due to their size and require significant data storage, necessitating improved methods for processing and utilizing these datasets efficiently.

Method used

A method involving converting read depth values in nucleotide sequencing datasets into single-letter ordinal class assignments (RRD) and mapping these to reference nucleotide sequences at base pair resolution, which can be used for data compression and subsequent analysis, including machine learning applications.

Benefits of technology

This approach significantly reduces processing time and storage requirements while enabling detailed analysis of gene expression patterns and genome sequences, facilitating applications in evolutionary research and medical diagnostics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026501259000005
    Figure 2026501259000005
  • Figure 2026501259000006
    Figure 2026501259000006
  • Figure 2026501259000007
    Figure 2026501259000007
Patent Text Reader

Abstract

The present invention relates to the field of genetics, and more particularly to quantitative sequencing data. Methods for processing and / or compressing quantitative sequencing data, computer-readable storage media for containing such processed data, and computing devices including at least one processor configured to process such data are provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention is in the field of analyzing and processing biological datasets, preferably large-scale biological datasets. Such large-scale datasets are preferably nucleotide datasets obtained by high-throughput sequencing. These nucleotide datasets can be generated by high-throughput sequencing of the genome or transcriptome of one or more samples. [Background technology]

[0002] The genotype of an organism is determined by the genetic material inherited from parents to offspring, which constitutes the genome, which codes for all of the organism's life processes. In eukaryotes, this genetic material includes the double-stranded DNA strings that make up the chromosomes in the nucleus of the organism's cells. Phenotype is understood as the observable characteristics of an organism. In plant and animal biology, these observable characteristics are referred to as traits. The phenotype of an organism is largely determined by the RNAs and / or proteins encoded by the genome and their expression levels. These RNAs and / or proteins are encoded on genetic loci in the genome, referred to as genes.

[0003] A gene is a region of DNA and can be thought of as a unit of transcription. Although some genes are transcribed into RNA that does not code for proteins (referred to as non-coding RNA), a gene generally refers to a structure containing a coding sequence composed of one or more exons that can be translated into proteins together. Exons are preceded and followed by regulatory sequences and, optionally (in eukaryotic cells), are interrupted by introns. Gene expression involves the process of transcribing a gene to form a pre-messenger RNA (pre-mRNA) that contains a string of exons interrupted by introns and short regulatory sequences, referred to as 5'-UTR and 3'-UTR, respectively. A process called post-transcriptional modification excises potential introns and adds a 5'-cap and polyadenylation tail (abbreviated as polyA tail) to the 5' and 3' ends of the mRNA, respectively, to form a mature mRNA. This mature mRNA can then be translated into a protein.

[0004] To understand and possibly even predict an organism's phenotype from its genotype, it is important to understand the sequences that code for RNA and / or proteins, as well as the (non-expressed) sequences that control gene expression. Next-generation sequencing technologies have elucidated the genome sequences of many species. Due to the uniqueness of genome sequences, next-generation sequencing technologies have proven valuable in many applications, such as forensic science, diagnosis of genetic or infectious diseases, and biodiversity analysis. Comparative genomics between related species has been used to elucidate gene sequences based on the high degree of sequence conservation between related species. Quantitative sequencing methods have provided an additional dimension. For example, understanding the amount of genomic DNA of a species and / or the amount of a specific transcript of an individual present in a sample is important in various fields, such as forensic science, medical diagnostics, and agricultural science. Most of these techniques rely on the read depth of next-generation sequencing (see, e.g., Reinecke et al. BMC Bioinformatics. 2015;16(17)). The read depth of a particular nucleotide sequence reflects the abundance of the original nucleic acid present in the sample, which may reflect, for example, the abundance of a particular microorganism present in an environmental sample, the number of chromosome copies present in a sample for prenatal screening, or the number of tumor cells present in a biopsy sample. Apart from DNA sequencing, a read-depth-dependent quantitative sequencing method, designated "RNA-seq," is also available for quantifying RNA transcripts. RNA-seq is deep sequencing of cellular RNA. It is a quantitative method that can be used to determine the expression level of mature mRNA, and therefore the expression of genes in cells, without requiring prior sequence knowledge of the gene (Want et al. Nat Rev Genet. 2009;10(1):57-63).In addition to polyadenylated messenger RNA (mRNA) transcripts, RNA-Seq can be applied to investigate various populations of RNA, including total RNA, pre-mRNA, and noncoding RNAs (ncRNAs) such as ribosomal RNA (rRNA) and transfer RNA (tRNA), both of which are involved in RNA translation; small nuclear RNAs (snRNAs) involved in splicing; small nucleolar RNAs (snoRNAs) involved in ribosomal RNA modification; microRNAs (miRNAs); transcriptionally active siRNAs (tasiRNAs) and piwi-interacting RNAs (piRNAs), both of which regulate gene expression at the post-transcriptional level; and long noncoding RNAs (lnoRNAs), which are involved in chromatin remodeling, transcriptional regulation, and post-transcriptional processing (Kukurba and Montgomery, Cold Spring Harb Protoc 2015;11:951-969). In RNA-Seq, RNA is extracted and, depending on the protocol used, a subset of RNA molecules, such as mature mRNAs, is isolated by enrichment of polyadenylated transcripts. The isolated RNA molecules are reverse transcribed into cDNA fragments, optionally amplified, and then sequenced. These sequencing reads can then be realigned with a previously sequenced reference genome or reference transcript, or optionally assembled without a reference (Oshlack et al. Genome Biol 2010;11(12):220; Wang et al. Nat Rev Genet 2009;10(1):57-63 and Auer et al. Brief Funct Genomics 2012;11(1):57-62).

[0005] RNA-seq data can be used to study the transcriptome, also known as transcriptomics. Transcriptome analysis can provide detailed and quantitative information, particularly on gene expression, alternative splicing, and allele-specific expression. For example, mature mRNA levels, expressed as amounts expressed per cell, are the most frequently studied RNA species because they encode proteins. (Ciaran Evans, Johanna Hardin, and Daniel M. Stoebel, "Selecting between-sample RNA-seq normalization methods from the perspective of their assumptions," Briefings in Bioinformatics, 19(5), 2018, 776-79). Summary of the Invention

[0006] Recent advances in sequencing workflows, from sample preparation and sequencing platforms to bioinformatics data analysis, have enabled the sophisticated use of quantitative sequencing methods in a variety of fields, from evolutionary studies to medical applications. However, despite the availability of such tools, these datasets are expensive, and their large nature slows the analysis process and requires large data storage capacity. Therefore, improved methods for using available and / or acquired quantitative sequencing datasets are needed.

[0007] The present invention is summarized in the following embodiments.

[0008] Embodiment 1. A method for processing a nucleotide sequencing dataset, comprising: (a) providing a nucleotide sequencing dataset comprising read depth values ​​assigned to nucleotide sequences; (b) providing a reference nucleotide sequence; (c) converting the read depth values ​​present in the nucleotide sequencing dataset into a single letter indicating the relative read depth (RRD) by ordinal class assignment; (d) mapping the single letters to the reference nucleotide sequence of (b) at base pair resolution; and is preferably a computer-implemented method.

[0009] Embodiment 2. The method of embodiment 1, wherein the nucleotide sequencing dataset of (a) is a transcriptome dataset obtained by RNA-seq.

[0010] Embodiment 3. The method of embodiment 1, wherein the nucleotide sequencing dataset of (a) is a genomic dataset obtained by DNA sequencing.

[0011] Embodiment 4 The method of any one of the preceding embodiments, wherein the reference nucleotide sequence is a genomic sequence, preferably a whole genome sequence.

[0012] Embodiment 5. Step (a) comprises: (1) providing a nucleic acid sample; (2) optionally enriching and / or isolating a subset of nucleic acid molecules from at least a portion of said nucleic acid sample; (3) preparing a sequencing library of the nucleic acid molecules of (1) or the enriched and / or isolated nucleic acid molecules of (2); (4) (Sub)step of obtaining a nucleotide sequence by sequencing the sequencing library of (3); (5) assigning read depth values ​​to the nucleotide sequences obtained in (4), thereby obtaining a nucleotide sequencing dataset comprising read depth values ​​assigned to the nucleotide sequences of the sequenced nucleic acid molecules; 13. The method of any one of the preceding embodiments, comprising:

[0013] Embodiment 6. The method of embodiment 5, wherein the nucleic acid molecules of (1) and optionally the enriched and / or isolated nucleic acid molecules of (2) are mRNA molecules, and the method further comprises the (sub)step of converting mRNA to said cDNA prior to step (3) of preparing a sequencing library of cDNA.

[0014] Embodiment 7. The method of any one of the preceding embodiments, wherein the reference nucleotide sequence of (b) is obtained by (whole) genome sequencing.

[0015] Embodiment 8. In step (c), the read depth values ​​are sorted in N+1 consecutive RRDs, each RRD including a read depth upper limit, and the read depth value is assigned to the lowest RRD whose read depth value does not exceed the read depth upper limit of the RRD, the read depth upper limit of the lowest RRD is 0, and for the remaining RRDs, the read depth upper limits are: UB=f·x RV (Formula 2) (In the formula, - UB is the upper read depth limit for a given RRD, x is preferably a cardinal number having a value between 1 and Euler's number (e), f is a scaling factor preferably having a value of at least 1, - RV is a given ranking value selected from the group consisting of {1, 2, 3, 4, ...... N} of the RRD for which the upper bound is calculated, and - N is the ranking value of the top RRD) 13. The method of any one of the preceding embodiments, wherein

[0016] Embodiment 9. The value of the base number x is x=(MRD / f) 1 / N (Formula 4) (In the formula, - x is the base of equations 2 and 3, - MRD is the maximum read depth, and - N is the ranking value of the top RRD) 9. The method of embodiment 8, wherein

[0017] Embodiment 10. The value of the scaling factor f is f=MRD / e N (Formula 5) (In the formula, - f is the scaling factor for equations 2 and 4, - MRD is the maximum read depth, - e is Euler's number, and - N is the ranking value of the top RRD) 10. The method according to embodiment 8 or 9, wherein

[0018] Embodiment 11.f is a method for determining whether MRD≧e N is given by Equation 5, and MRD <e N When f=1, the read depth upper bound of the lowest RRD is 0, and for the remaining RRDs, the read depth upper bound is MRD≧e N in the case of, UB=(MRD·e RV ) / e N (Formula 6) is given by, and MRD <e N in the case of, UB=(MRD) RV / N (Formula 8) where: - UB is the upper read depth limit for a given RRD, - MRD is the maximum read depth, - e is Euler's number, - RV is a given ranking value selected from the group consisting of {1, 2, 3, 4, ...... N} of the RRD for which the upper bound is calculated, and - The method according to any one of embodiments 8 to 10, wherein N is the ranking value of the top RRD.

[0019] Embodiment 11. The method of any one of the preceding embodiments, performed on a plurality of samples.

[0020] Embodiment 12. Use of processed nucleotide sequencing data obtained or obtainable by any one of the methods described in embodiments 1 to 11, preferably in a computer-implemented method for predicting genome sequences and / or comparing gene expression patterns within or between samples.

[0021] Embodiment 13. The use of embodiment 12, wherein the computer-implemented method comprises machine learning.

[0022] Embodiment 14. A computer-readable storage medium comprising processed nucleotide sequencing data obtained by the method according to any one of embodiments 1 to 11.

[0023] Embodiment 15. A computing device including at least one processor, the processor configured to execute the method of any one of embodiments 1 to 11. [Brief explanation of the drawings]

[0024] [Figure 1] 1 illustrates a method for processing a nucleotide sequencing dataset. [Figure 2] 1 illustrates a method for obtaining a nucleotide sequencing dataset that includes read depth values ​​assigned to nucleotide sequences. [Figure 3] 1 illustrates a method for predicting genes and gene expression levels in a nucleic acid sample. [Figure 4] 1 illustrates a computing device for carrying out the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0025] definition Various terms relating to the methods, compositions, uses, and other aspects of the present invention are used throughout the specification and claims. Such terms have their ordinary meaning in the art to which the invention pertains, unless otherwise indicated. Other specifically defined terms are to be construed in a manner consistent with the definition provided herein. Although any methods and materials similar or equivalent to those described herein can be used in the practice of testing the present invention, the preferred materials and methods are described herein.

[0026] It will be apparent to one of ordinary skill in the art how to carry out the conventional techniques used in the methods of the present invention. Conventional practices in molecular biology, biochemistry, computational chemistry, cell culture, recombinant DNA, bioinformatics, genomics, sequencing, and related fields are well known to those of ordinary skill in the art and are described, for example, in the following references: Sambrook et al. Molecular Cloning. A Laboratory Manual, 2002; nd Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989; Ausubel et al. Current Protocols in Molecular Biology, John Wiley&Sons, New York, 1987 and periodic updates; and the series Methods in Enzymology, Academic Press, San Diego.

[0027] "A," "an," and "the": These singular terms include plural referents unless the content clearly indicates otherwise. Thus, for example, reference to a "cell" includes a combination of two or more cells, and the like.

[0028] As used herein, the term "about" is used to describe and account for small variations. For example, the term may refer to ±(+ or -) 10% or less, e.g., ±5% or less, ±4% or less, ±3% or less, ±2% or less, ±1% or less, ±0.5% or less, ±0.1% or less, or ±0.05% or less. Additionally, amounts, ratios, and other numerical values ​​may be presented in range format. Such range formats are used for convenience and brevity and should be interpreted flexibly to include numerical values ​​explicitly specified as range limits, but should also be understood to include all individual numerical values ​​or subranges subsumed within the range, as if each numerical value and subrange were expressly specified. For example, a ratio in the range of about 1 to about 200 should be understood to include not only the explicitly stated limits of about 1 and about 200, but also individual ratios such as about 2, about 3, and about 4, as well as subranges such as about 10 to about 50, about 20 to about 100, etc.

[0029] "Amplification," as used in reference to nucleic acids or nucleic acid reactions, refers to an in vitro method of generating copies of a specific nucleic acid, such as a target nucleic acid and / or a tagged nucleic acid. Numerous methods for amplifying nucleic acids are known in the art, including, but not limited to, polymerase chain reaction, ligase chain reaction, strand displacement amplification reaction, rolling circle amplification reaction, transcription-mediated amplification methods such as NASBA (e.g., U.S. Pat. No. 5,409,818), loop-mediated amplification (e.g., "LAMP" amplification using a loop-forming sequence as described in U.S. Pat. No. 6,410,278), and isothermal amplification reactions. The nucleic acid being amplified may be composed of or derived from DNA or RNA, or a mixture of DNA and RNA (including modified DNA and / or RNA). The product resulting from the amplification of a nucleic acid molecule or molecules (i.e., "amplification product") may be either DNA or RNA, or a mixture of both DNA and RNA nucleosides or nucleotides, or modified DNA or RNA nucleosides or nucleotides, regardless of whether the starting nucleic acid is DNA, RNA, or both.

[0030] As used herein, the term "nucleic acid sample," also referred to herein as "sample," refers to a sample containing one or more nucleic acids, preferably genomic DNA molecules and / or RNA molecules expressed from a genome, where the expressed RNA molecules are also referred to herein as a "transcriptome." A nucleic acid sample can be a biological sample derived from one or more biological sources and can be obtained from one or more of the same or different individuals, such as humans, plants, animals, bacteria, fungi, algae, insects, etc. A biological sample can be derived from cells, tissues, biopsies, or bodily fluids. Optionally, the nucleic acid sample is an environmental sample, such as a sample from soil or water, and contains a mixture of nucleic acids from different species, optionally in the form of, for example, whole organisms (e.g., bacteria or viruses) and / or blood, urine, feces, gametes, and / or skin of an organism. A sample can contain a mixture of substances containing one or more RNA molecules or RNA transcripts, typically, but not necessarily, in liquid form.

[0031] "Comprises" is to be interpreted as inclusive and open-ended, rather than exclusive. Specifically, this term and variations thereof mean that the specified features, steps or components are included. These terms should not be interpreted to exclude the presence of other features, steps or components.

[0032] As used herein, the terms "double stranded" and "duplex" refer to two complementary polynucleotides that base pair, i.e., hybridize together. Complementary nucleotide strands are also known in the art as reverse complements.

[0033] "Expression": This refers to the process by which a DNA region operably linked to appropriate regulatory regions, particularly a promoter, is transcribed into RNA. In the case of RNA that encodes a protein, the RNA can be translated into a protein or peptide.

[0034] The term "gene" refers to a DNA fragment containing a region (transcribed region) that is transcribed into an RNA molecule (e.g., a pre-mRNA or non-coding RNA) by a cellular RNA polymerase enzyme in a process called transcription. The transcribed region that can be translated into a protein is called an open reading frame (ORF) and begins with a three-letter code designated as a start codon and ends with one of three stop codons. An ORF may contain exons and one or more introns. Upstream of the ORF is the 5' UTR, and downstream is the 3' UTR, which constitutes the boundaries of the transcribed RNA. In the case of a pre-mRNA, during maturation, introns are spliced ​​out, and a 5' cap and polyadenylation tail (abbreviated as polyA tail) are added to the 5' and 3' ends of the RNA, respectively, to form the mature mRNA. Thus, the mature mRNA is composed of the following nucleotide sequence elements: 5' UTR, exons, 3' UTR, and polyA tail. Together, the exons constitute the coding sequence (CDS). The mature mRNA can then be translated into protein in a process called translation. An ORF is associated (or operably linked) with non-transcribed and / or non-translated regulatory sequences at its 5' and / or 3' end, such as promoter sequences that can bind transcription factors that promote the recruitment of RNA polymerase and the initiation of transcription. Apart from promoter sequences, regulatory sequences can function as enhancers and / or silencers of transcription, for example, by binding to specific enhancer or inhibitory elements and / or by influencing chromatin structure.

[0035] The term "nucleotide" includes, but is not limited to, naturally occurring nucleotides, including guanine, cytosine, adenine, thymine, and uracil (G, C, A, T, and U, respectively). The term "nucleotide" is further intended to include moieties containing not only the known purine and pyrimidine bases, but also modified and other heterocyclic bases. Such modifications include methylated purines or pyrimidines, acylated purines or pyrimidines, alkylated ribose, or other heterocycles. Additionally, the term "nucleotide" includes moieties containing haptens or fluorescent labels and may contain other sugars in addition to traditional ribose and deoxyribose sugars. Modified nucleosides or nucleotides include modifications to the sugar moiety, for example, in which one or more hydroxyl groups have been replaced with halogen atoms or aliphatic groups, or functionalized as ethers, amines, and the like.

[0036] The terms "nucleic acid," "polynucleotide," and "nucleic acid molecule" are used interchangeably herein to refer to polymers of any length, such as those containing nucleotides (e.g., deoxyribonucleotides or ribonucleotides) of any length, e.g., greater than about 2 bases, greater than about 10 bases, greater than about 100 bases, greater than about 500 bases, greater than 1000 bases, or up to about 10,000 bases or more (e.g., PNAs, as described in U.S. Pat. No. 5,948,902 and references cited therein). Nucleic acids can hybridize with naturally occurring nucleic acids in a sequence-specific manner similar to two naturally occurring nucleic acids, e.g., participate in Watson-Crick base pairing interactions. Furthermore, nucleic acids and polynucleotides can be isolated (and optionally subsequently fragmented) from cells, tissues, and / or bodily fluids. Nucleic acids can be, for example, genomic DNA (gDNA), mitochondria, cell-free DNA (cfDNA), DNA from a sequencing library, and / or RNA from a sequencing library. The nucleic acid is preferably a double-stranded molecule, unless the context makes it clear that a single-stranded molecule is intended.

[0037] As used herein, the term "oligonucleotide" refers to a single-stranded polymer of nucleotides, preferably about 2 to 200 nucleotides or up to 500 nucleotides in length. Oligonucleotides can be synthetic or enzymatically produced and, in some embodiments, are about 10 to 50 nucleotides in length. Oligonucleotides can contain ribonucleotide monomers (i.e., oligoribonucleotides) or deoxyribonucleotide monomers. Oligonucleotides can be, for example, about 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 100, 100 to 150, 150 to 200, or about 200 to 250 nucleotides in length.

[0038] "Sequence" or "Nucleotide Sequence": Refers to the order of nucleotides of or within a nucleic acid. In other words, any order of nucleotides in a nucleic acid can be referred to as a sequence or nucleic acid sequence. For example, a target sequence is the order of nucleotides contained in one strand of a DNA duplex.

[0039] As used herein, "sequence identity" is defined as the relationship between two or more amino acid (polypeptide or protein) sequences or two or more nucleotide (polynucleotide) sequences, as determined by comparing the sequences. In the art, "identity" also means the degree of sequence relatedness between amino acid or nucleic acid sequences, as the case may be, as determined by the match between strings of such sequences. "Similarity" between two amino acid sequences is determined by comparing the amino acid sequence of one polypeptide and its conserved amino acid substitutes with the sequence of a second polypeptide. "Identity" and "similarity" can be readily calculated using known methods. The percentage of sequence identity / similarity can be determined over the entire length of the sequences.

[0040] "Sequence identity" and "sequence similarity" can be determined by aligning two amino acid sequences or two nucleotide sequences using a global or local alignment algorithm, depending on the length of the two sequences. Sequences of similar length are preferably aligned using a global alignment algorithm (e.g., Needleman Wunsch) that optimally aligns sequences over their entire length, while sequences of significantly different lengths are preferably aligned using a local alignment algorithm (e.g., Smith Waterman). Thus, sequences can be referred to as "substantially identical" or "essentially similar" if they share at least a certain minimum percentage of sequence identity (defined below) (e.g., when optimally aligned using the GAP or BESTFIT programs with default parameters). GAP uses the Needleman and Wunsch global alignment algorithm to align two sequences over their entire length (full length), maximizing the number of matches and minimizing the number of gaps. When two sequences are similar in length, global alignment is preferably used to determine sequence identity. Generally, the GAP default parameters are used: gap creation penalty = 50 (nucleotides) / 8 (proteins), gap extension penalty = 3 (nucleotides) / 2 (proteins). For nucleotides, the default weight matrix is ​​nwsgapdna, and for proteins, the default weight matrix is ​​Blosum62 (Henikoff & Henikoff, 1992, PNAS 89, 915-919).Sequence alignments and percent sequence identity scores can be determined using computer programs such as the GCG Wisconsin Package, Version 10.3, available from Accelrys Inc., 9685 Scranton Road, San Diego, CA 92121-3752 USA, or open source software such as the EmbossWIN version 2.10.0 programs "needle" (which uses the global Needleman-Wunsch algorithm) or "water" (which uses the local Smith-Waterman algorithm), using the same parameters as GAP described above or default settings (for both "needle" and "water" and for both protein and DNA alignments, the default gap opening penalty is 10.0, the default gap extension penalty is 0.5, and the default scoring matrix is ​​Blossum62 for proteins and DNAFull for DNA). When the overall lengths of the sequences differ significantly, local alignments, such as those using the Smith-Waterman algorithm, are preferred.

[0041] Alternatively, the percentage of similarity or identity can be determined by searching public databases using algorithms such as FASTA and BLAST. Thus, the nucleic acid and protein sequences of the present invention can further be used as a "query sequence" to perform searches against public databases to identify, for example, other family members or related sequences. Such searches can be performed using the BLASTn and BLASTx programs (version 2.0) of Altschul, et al. (1990) J. Mol. Biol. 215:403-10. BLAST nucleotide searches can be performed with the NBLAST program, score = 100, word length = 12, to obtain nucleotide sequences homologous to the nucleic acid molecules of the present invention. BLAST protein searches can be performed with the BLASTx program, score = 50, word length = 3, to obtain amino acid sequences homologous to the protein molecules of the present invention. To obtain gapped alignments for comparison purposes, Gapped BLAST can be utilized as described in Altschul et al., (1997) Nucleic Acids Res. 25(17):3389-3402. When using BLAST and Gapped BLAST programs, the default parameters of the respective programs (e.g., BLASTx and BLASTn) can be used. See the National Center for Biotechnology Information homepage (http: / / www.ncbi.nlm.nih.gov / ).

[0042] As used herein, the term "sequencing" refers to a method by which the identity of at least about 10 consecutive nucleotides of a polynucleotide (e.g., the identity of at least about 20, at least about 50, at least about 100, or at least about 200 or more consecutive nucleotides) is obtained. The term "next-generation sequencing" refers to so-called parallelized sequencing-by-synthesis or sequencing-by-ligation platforms currently employed by, for example, Illumina, Life Technologies, PacBio, Roche, and Complete Genomics. Next-generation sequencing also includes electronic detection-based methods such as nanopore sequencing, as commercialized by Oxford Nanopore Technologies, or Ion Torrent technology, as commercialized by Life Technologies.

[0043] In a first aspect, there is provided a method for processing nucleotide sequencing data, comprising the steps of: (a) providing a nucleotide sequence analysis dataset comprising read depth values ​​assigned to nucleotide sequences; (b) providing a reference nucleotide sequence; (c) converting the read depth values ​​present in the nucleotide sequencing dataset into a single letter indicating the relative read depth (RRD) by ordinal class assignment; (d) mapping the single letters to the reference nucleotide sequence of (b) at base pair resolution; A method is provided that includes:

[0044] 1 depicts a flowchart showing the method, comprising step 101 providing a nucleotide sequencing dataset comprising read depth values ​​assigned to nucleotide sequences, step 102 providing a reference nucleotide sequence, step 103 converting the read depth values ​​present in the nucleotide sequencing dataset into single letters indicating relative read depth (RRD) via ordinal class assignment, and step 104 mapping these single letters to the reference nucleotide sequence in (b) at base pair resolution. Thus, in FIG. 1, step 101 corresponds to step (a) described herein, step 102 corresponds to step (b) described herein, step 103 corresponds to step (c) described herein, and step 104 corresponds to step (d) described herein.

[0045] The nucleotide sequencing dataset of (a) is preferably the output of a nucleotide sequencing method, preferably the output of sequencing nucleic acid molecules of a sample. Preferably, the nucleotide sequencing dataset of (a) is obtained by a quantitative sequencing method. Alternatively, the dataset is derived from a public database. As specified in (a), the sequencing dataset comprises read depth values ​​assigned to nucleotide sequences. A read depth value, also referred to herein as read depth, is herein understood as the number of times a given nucleotide or nucleotide sequence is read, for example, in a nucleic acid sequencing method that sequences nucleic acid molecules of a sample. The dataset of step (a) preferably comprises read depth values ​​assigned to nucleotide sequences of nucleic acid molecules present in the sample. Since each read depth value can indicate the respective amount of corresponding nucleic acid molecules present in the sample, a nucleotide sequencing dataset comprising read depth values ​​can also be referred to as a quantitative nucleotide sequencing dataset.

[0046] Provided is a method for optimally converting read depth values ​​of one or more nucleotide sequencing datasets into a single character that is itself a class assignment, but can also be subsequently used as ground truth for ordinal classification, preferably while preserving as much information as possible from the nucleotide sequencing dataset data. Here, "ground truth" is understood as the reality that is modeled by machine learning or the like. The conversion of read depth values ​​into (ordinal) RRDs can also be referred to as binning, and RRDs can be referred to as bins. Another name given to bins is ordinal variables or categories. Ordinal class assignment converts read depth values ​​into ordinal variables or categories that can be ranked from lowest to highest. In the method provided herein, these categories, bins, or RRDs are indicated by a single character, each representing a range of consecutive read depth values. Thus, read depth values ​​are grouped into RRDs, and each read depth value within a specific range of read depth values ​​is assigned to the same RRD. Each RRD has a specific read depth upper limit (or upper limit value), and read depths above the upper limit are assigned to a higher-ranked RRD. Similarly, each RRD has a specific read depth lower limit, and read depths having values ​​equal to or less than the read depth lower limit (or lower limit value) are assigned to lower RRDs.

[0047] If the number of different RRDs, and therefore the number of single characters, in step (c) is less than the number of different read depth values ​​present in the dataset provided in (a), step (c) of the method provided herein results in data compression, and therefore the above method can be expressed as a method of compressing nucleotide sequencing data of a nucleic acid sample.

[0048] The nucleotide sequencing dataset of step a) includes read depth values ​​assigned to nucleotide sequences. In step c), each of these read depth values ​​is assigned to a specific RRD. By aligning these nucleotide sequences or their complements to one or more reference nucleotide sequences of b), an RRD can be assigned to each nucleotide of the one or more reference nucleotide sequences in step c), a process referred to herein as mapping an RRD to a reference nucleotide sequence at base pair resolution. As used herein, "base pair resolution" is understood to mean assigning an RRD to each base pair of a double-stranded DNA sequence, or assigning an RRD to each base or nucleotide of a single-stranded DNA or RNA, which is optionally one strand of genomic DNA.

[0049] Optionally, the read depth values ​​of the nucleotide sequencing dataset are converted to RRDs before mapping a single character representing the RRD per nucleotide to the reference nucleotide sequence. Alternatively, the read depth values ​​of the nucleotide sequencing dataset are mapped nucleotide by nucleotide to the reference nucleotide sequence, thereby rendering a string of read depth values, and then the string values ​​are converted to RRDs. In either case, the result of step (d) is a nucleotide sequence represented by one character per nucleotide that represents that RRD in the reference nucleotide sequence, also referred to herein as a string of RRDs mapped to the reference nucleotide sequence or an RRD mapped to the reference nucleotide sequence at base pair resolution.

[0050] Optionally, the reference nucleotide sequence in b) is a genome sequence, the nucleotide sequence of the dataset is a transcript or RNA, preferably mRNA, more preferably mature mRNA, and the nucleotide sequencing dataset is an RNA-seq dataset or a transcriptome dataset. When the dataset is an RNA-seq dataset, the read depth value may correspond to the expression level value of the transcript. Since each RNA-nucleotide of the RNA sequence is transcribed from a specific DNA-nucleotide of the genome, such a DNA-nucleotide can be assigned a specific single letter representing the RRD of the transcript encoded by the DNA-nucleotide. The RRD, and thus the single letter representing the RRD, represents an expression category inferred from the transcriptome data and can be assigned to the genome at a base pair resolution.

[0051] Preferably, the dataset of step (a) comprises nucleotide sequences, the nucleotide sequences having been assigned read depth values. At least some of the sequences of step (a) may correspond to reference nucleotide sequences of step (b). Thus, preferably, at least some of the sequences of the dataset of step (a) have at least 50%, 60%, 70%, 80%, 90%, 95% or 100% identity to the reference sequences of step (b).

[0052] Preferably, the reference nucleotide sequences of (b) are related to the sequences of the dataset of (a). In particular, one or more, optionally all, of the nucleotide sequences of the dataset of (a) or their complementary sequences match at least a portion of at least the reference nucleotide sequences of (b). In other words, one or more nucleotide sequences of the dataset of (a) or their complementary sequences exhibit at least 50%, 60%, 70%, 80%, 90%, 95%, or 100% identity over their entire length with at least one (part of) said reference nucleotide sequences. Optionally, one or more of the nucleotide sequences of the dataset of (a) or their complementary sequences perfectly match at least a portion of at least one of the reference nucleotide sequences of (b).

[0053] Optionally, at least one (optionally all) of the reference nucleotide sequences of (b) is a whole or part of a genomic nucleotide sequence. The genomic sequence may be a genomic sequence of a cell or organelle, such as a genomic sequence of a nucleus, a mitochondrion, or a chloroplast. The part of the genomic sequence may be a sequence of particular interest, e.g., a sequence highly suggestive of the particular species from which the genomic sequence is derived, or a particular sequence containing a gene of interest, such as, but not limited to, a chromosomal region, a marker region, a conserved region, a (hyper)variable region, a quantitative trait locus (QTL), or a particular gene associated with a particular crop trait or a particular animal or human disease. Optionally, the reference genomic sequence of step (b) is of a particular species or individual. Optionally, in step (b), multiple reference sequences of multiple species or individuals are provided. The reference sequences may be from a public database or may be obtained by sequencing a nucleic acid sample.

[0054] Optionally, the one or more reference nucleotide sequences of (b) and the nucleotide sequencing dataset of (a) are obtained by sequencing the same nucleic acid sample or by sequencing nucleic acid samples of the same species. The nucleotide sequences of the dataset of (a) can be sequences of nucleic acid molecules that are fragments, amplicons, and / or transcription products of the (genomic, ribosomal, or mitochondrial) reference nucleotide sequence of (b). For example, if the reference nucleotide sequence is a nucleotide sequence of a genomic nucleic acid, the nucleotide sequencing dataset of step (a) can be obtained from RNA transcription products from said genomic nucleic acid.

[0055] Converting read depth values ​​of (quantitative) DNA and / or RNA sequencing datasets into single characters mapped to a reference sequence at base pair resolution allows for further processing and / or analysis of such large amounts of compressed data, e.g., in deep learning, machine learning and / or artificial intelligence processes and / or applications, thereby enabling a wide variety of (quantitative) nucleic acid-related problems to be addressed, ranging from evolutionary research problems to medical applications, while significantly reducing processing time and storage space.

[0056] The methods provided herein comprise: (1) providing a nucleic acid sample; (2) optionally enriching and / or isolating a subset of nucleic acid molecules from at least a portion of said sample; (3) preparing a sequencing library of the nucleic acid molecules of (1) or the enriched and / or isolated nucleic acid molecules of (2); (4) (Sub)step of obtaining a nucleotide sequence by sequencing the sequencing library of (3); (5) assigning read depth values ​​to the nucleotide sequences obtained in (4), thereby creating a nucleotide sequencing dataset that includes both the nucleotide sequences and the read depth values ​​assigned to the nucleotide sequences; may include:

[0057] 2 depicts a flowchart showing the method, comprising step 1011 of providing a nucleic acid sample containing nucleic acid molecules, step 1012 of optionally enriching and / or isolating a subset of nucleic acid molecules from at least a portion of the sample, step 1013 of preparing a sequencing library of the (1011) nucleic acid molecules or the (1012) enriched and / or isolated nucleic acid molecules, step 1014 of sequencing the sequencing library, and step 1015 of assigning read depth values ​​to the nucleotide sequences obtained in step 1014. Thus, in FIG. 2, step 1011 corresponds to step (a)(1) described herein, step 1012 corresponds to step (a)(2) described herein, step 1013 corresponds to step (a)(3) described herein, step 1014 corresponds to step (a)(4) described herein, and step 1015 corresponds to step (a)(5) described herein.

[0058] The nucleic acid sample of the methods provided herein may be derived from a single individual, such as, but not limited to, a human, animal, plant, fungus, insect, virus, or microorganism. The sample may be a cell sample, tissue sample, and / or liquid biopsy sample. The sample may be a sample of individuals within a population and / or a sample of a specific portion of individuals, such as, but not limited to, a sample of plant leaves, fruit, or roots, or animal or human blood, saliva, cancerous tissue, urine, or feces, and / or a sample from a plant exposed to biotic stress, or a sample from a human or animal undergoing medical treatment, taken from an individual under specific circumstances. Optionally, the nucleic acid sample is a pooled sample, optionally taken from different individuals or the same individual. Optionally, these multiple individuals are of different species. Optionally, the nucleic acid sample is an environmental sample, such as, but not limited to, a soil or wastewater sample. When the nucleic acid sample is a human or animal sample, the sample collection step is preferably non-invasive. Preferably, the sample collection step is not part of the methods described herein. Preferably, the methods provided herein are ex vivo and / or in vitro methods.

[0059] The nucleotide sequencing dataset of (a) may be a DNA sequencing dataset, preferably a quantitative DNA sequencing dataset. Optionally, the dataset is obtained from a nucleic acid sample comprising two or more DNA molecules. Additionally or alternatively, the nucleotide sequencing dataset of (a) may be an RNA sequencing dataset, preferably a quantitative RNA sequencing dataset. Optionally, the dataset is obtained from a nucleic acid sample as defined herein comprising two or more RNA molecules.

[0060] Optionally, the nucleotide sequencing dataset of (a) comprises nucleotide sequences of two or more (different) individuals, species, and / or cell types. Preferably, the read depth values ​​of said dataset reflect the amount of nucleic acid in a nucleic acid sample, e.g., the amount of nucleic acid of two or more (different) individuals, species, and / or cell types present in the sample. For example, said sample may be an environmental sample, or may be a plant, animal, and / or human sample comprising nucleic acid molecules of multiple different microorganisms. Non-limiting examples of environmental samples are soil, drinking water, surface water, sewage, etc. Non-limiting examples of plant, animal, and / or human samples are plant parts, blood, skin, saliva, intestine (gastric juice), urine, feces, etc. Preferably, in step (b), reference sequences of these different individuals, species, and / or cell types (optionally different microorganisms) are provided. Preferably, the reference nucleotide sequences are genomic sequences of different individuals, species, and / or cell types. Optionally, the reference sequences are RNA sequences of different individuals, species, and / or cell types. The nucleotide sequencing dataset may be a (quantitative) DNA or RNA sequencing dataset. In this embodiment, a single letter representing an RRD is mapped to one of multiple reference nucleotide sequences by aligning the nucleotide sequence or its complementary sequence to a reference genome sequence. Nucleotides in the reference sequence that align with these nucleotide sequences or their complementary sequences are assigned to each RRD, and the RRDs are assigned at (genomic) base pair resolution. When the sample is an environmental sample, the resulting processed data obtainable by the method described herein can be used, for example, to analyze the quantitative presence of various microorganisms in the environmental sample. The one or more reference nucleotide sequences in (b) are preferably of a specific portion of a genome sequence, such as a portion encoding ribosomal RNA (e.g., the nucleotide sequence of the 16S rRNA gene when screening for the presence and amount of various microorganisms in an environmental or biological sample), and the nucleotide sequencing dataset may be a quantitative DNA sequencing dataset or an RNA-seq dataset.Optionally, the reference sequence may be an RNA sequence (optionally a 16S rRNA gene), and the nucleotide sequencing dataset may be an RNA-seq dataset. Those skilled in the art will understand that the possible applications of this method are not limited to the detection of various microorganisms. For example, instead of the quantitative presence of different microorganisms, the quantitative presence of two or more (different) cell types may be analyzed. The different cell types may be cells from different individuals; for example, in the case of a sample from a pregnant animal or human, the nucleic acid sample of the methods provided herein may include cells from the mother and child, fetus, or embryo. Additionally or alternatively, the different cell types may be cells from different tissues, optionally cells from the same individual, for example, cells from different organs, or cells from different health states (e.g., healthy cells versus cancerous and / or metastatic cells). As another non-limiting example, the quantitative presence of a pathogen in, for example, a biological or environmental sample may be determined. Optionally, (parts of) parasites, bacteria or viruses may be quantitatively detected in biological or environmental samples, such as, but not limited to, quantitative detection of Covid-19 in biological or environmental samples.

[0061] The nucleotide sequencing dataset (a) may be a transcriptome dataset consisting of RNA read depth values ​​that reflect the amount of transcripts or RNA species present in a nucleic acid sample. Transcriptome data may be obtained by (massively parallel) RNA sequencing (RNA-seq) or other methods that provide quantitative expression level data. The transcriptome dataset may be obtained by RNA-seq. Preferably, to avoid amplification bias, RNA-seq is performed without an amplification step. Such amplification-free methods may be based on single-molecule-based platforms such as PacBio single-molecule real-time (SMRT) sequencing (Kukurba and Montgomery, Cold Spring Harb Protoc. 2015 Nov; 2015(11):951-969). Optionally, the transcriptome dataset provides the presence and quantity of total RNA-transcripts or (pre)mRNA of the provided sample, preferably obtained by RNA-seq, or a specific subset thereof, such as any one of other subsets, including but not limited to, rRNA, tRNA, snRNA, snoRNA, miRNA, piRNA, tasiRNA, lncRNA, and combinations thereof. Preferably, the provided sample has been processed to allow RNA sequencing, preferably to allow sequencing of either mRNA or other RNA transcripts. Preferably, (a part of) the genome of the provided sample may be used as a reference nucleotide sequence in step (b) of the method provided herein.

[0062] Preferably, step (a)(1) of providing a nucleic acid sample includes a sample collection step and, optionally, extraction, enrichment, and / or isolation steps of nucleic acids, DNA, and / or RNA from the sample, which are then subjected to nucleic acid, DNA, and / or RNA sequencing, thereby providing a quantitative nucleic acid, DNA, and / or RNA (RNA-seq) dataset for the sample. Those skilled in the art are aware of suitable techniques for extraction, enrichment, and / or isolation of total nucleic acid molecules (DNA and / or RNA) and / or specific types or subsets of the nucleic acid molecules. In the case of DNA, extraction can be, for example, extraction of cellular or cell-free DNA from body fluid samples such as blood or plasma or environmental samples. Those skilled in the art are well aware of methods for extracting DNA. Cellular DNA extraction can include steps of cell lysis using detergents and surfactants, protein and RNA digestion, ethanol precipitation, phenol-chloroform extraction, and / or (mini)column purification. Many DNA extraction kits are available for specific samples, such as plant cells, tissues, and soil. Several RNA extraction methods, such as phenol extraction and extraction using commercially available kits (e.g., Qiagen RNeasy Kit, Zymo Research Direct-zol), are described in Scholes and Lewis, BMC Genomics (2020) 21:249, incorporated herein by reference. Enrichment and / or isolation can be directed to a subset of nucleic acid molecules, the DNA or RNA fraction, e.g., the mRNA fraction in the case of RNA. Enrichment and / or isolation of mRNA from a nucleic acid sample can be performed, for example, by selectively capturing polyA-tailed RNA from the sample using magnetic beads conjugated with poly(dT) oligonucleotides. Additionally or alternatively, other DNA and / or RNA subsets can be of interest and can be enriched and / or isolated in step (a)(2) of the methods provided herein.Those skilled in the art are aware of suitable methods for enriching and / or isolating distinct RNA subsets, such as, for example, size fractionation by gel electrophoresis, silica spin columns for binding and elution of small RNAs, or methods that exploit the difference in solubility characteristics between small RNA molecules (preferably less than 100 nucleotides releasable) including mRNA and rRNA and large RNA molecules (no longer solubilized) in water (Choi et al. RNA Biol. 2018;15(6):763-772). Additionally or alternatively, the method may include a step (a)(2) of enriching DNA and / or RNA subsets consisting of and / or annealing to specific sequences, using techniques such as capture probe hybridization, e.g., using magnetic beads conjugated with capture probes or oligonucleotides consisting of sequences capable of annealing to the DNA and / or RNA subsets.

[0063] Optionally, the nucleotide sequencing dataset is an RNA dataset, preferably a quantitative RNA dataset consisting of read depth values ​​of one or more RNA sequences. Thus, preferably, the nucleotide sequence dataset is a transcriptome dataset, and the method provided herein can be expressed as a method for processing transcriptome data of a nucleic acid sample, the method comprising: (a) obtaining a transcriptome dataset comprising read depth values ​​assigned to the nucleotide sequences of transcripts present in a sample; (b) obtaining a reference genome nucleotide sequence; (c) converting the read depth values ​​present in the dataset into a single letter representing RRD by ordinal class assignment; (d) mapping the single letters to the reference nucleotide sequence of (b) at base pair resolution; Includes.

[0064] Preferably, the sequence of the reference genome is transcribed into transcripts of the transcriptome dataset defined herein. Optionally, in RNA-seq, the transcribed RNA (optionally enriched and / or isolated RNA) may be converted into complementary DNA (cDNA). The cDNA is then processed to form a sequencing library for (deep) sequencing. Those skilled in the art will recognize how to handle and process samples for RNA sequencing. The RNA sequencing is preferably mRNA (deep) sequencing. Thus, in the case of RNA, preferably, the RNA provided in step (a)(1) or (a)(2) is converted into cDNA, and a sequencing library is prepared from the cDNA in step (a)(3). Optionally, before sequencing, the cDNA is amplified using primers containing UMIs. The sequencing data obtained in step (4) of the cDNA corresponds to the original RNA. Optionally, step (b) of obtaining a reference nucleotide sequence is performed by sequencing. Optionally, the method provided herein comprises steps (a)(1) to (a)(3) above, wherein the sample provided in step (a)(1) is divided into at least two fractions, one fraction of the sample is subjected to steps (a)(3) to (a)(5) and optionally step (a)(2) defined herein, and another fraction of the sample is subjected to genome sequencing. The sequencing may be sequencing of a target genome sequence and / or a genome sequence of interest. Preferably, the genome sequencing is whole genome sequencing. The obtained genome sequence may serve as a reference sequence in step (b) of the method described herein. Optionally, (a part of) the sample is processed prior to genome sequencing, wherein the processing preferably comprises subjecting the sample to DNA enrichment and / or isolation and to preparation of a sequencing library, and the sequencing library is then subjected to (whole) genome sequencing.

[0065] Thus, in one embodiment, the method provided herein comprises steps (a)(1) to (a)(3) as defined herein, wherein step (1) further comprises dividing the provided sample into at least two equal parts (or fractions), the first part being subjected to genomic analysis to obtain the reference sequences provided in (b), and the second part being subjected to transcriptome analysis in steps (a)(3) and optionally step (a)(2) as defined herein to obtain the nucleotide sequencing database provided in (a). Preferably, the method comprises step (a)(2) of enriching a portion of the sample for an mRNA fraction. Optionally, the portion of the sample subjected to genomic analysis to obtain the reference nucleotide sequences in step (b) comprises a DNA enrichment and / or isolation step prior to the genomic sequencing step. Preferably, the DNA sequencing is (whole) genome DNA sequencing. Thus, the sample of the methods provided herein comprises nucleic acids, the nucleic acids preferably comprising DNA and RNA, the nucleic acids preferably comprising genomic DNA and mRNA. In an alternative embodiment, the reference (whole or partial) genome sequence used in step (b) of the method described herein can be obtained from a sequencing library and / or a public database. The reference genome sequence can be a publicly available genome sequence from the particular species from which the sample is derived. For example, if the sample is from a human, a publicly available human reference genome sequence can be used. As another non-limiting example, if the sample is a Heinz tomato, a publicly available reference genome sequence for Heinz tomatoes can be used.

[0066] The nucleotide sequencing dataset of step (a) comprises a read depth value. Preferably, the read depth per nucleic acid or per nucleotide (DNA or RNA) reflects the amount of said particular nucleic acid (DNA or RNA) present in the sample. In the case of RNA-seq, the read depth value may reflect the expression level of the transcript in the tissue and / or cell from which the sample is derived. The read depth is an arbitrary value that can range from 0 (i.e., not present in the nucleotide sequencing dataset) to a high multi-digit number, since the range of DNA and / or RNA present in the sample is very wide, e.g., in the case of RNA-seq, there is a wide range of RNA expression levels throughout the genome.

[0067] Converting read depth values ​​to a single character according to the methods provided herein allows for read depth information to be captured, optionally across the entire genome, at base pair resolution without the need for normalization. Using a single character for read depth values ​​minimizes data storage space and computational requirements when using the data for further processing, such as deep learning, while preserving as much nucleotide sequencing data as possible. When the nucleotide sequencing data is RNA-seq data and the reference nucleotide sequence is a genomic sequence, the RRD obtainable by the methods of the present invention may represent the relative expression of the genomic base pairs in the sample. When the sample is a sample from an individual, the methods provided herein may result in the representation of the genomic sequence by a character string (RRD) having expression information at base pair resolution. When the nucleotide sequencing data is obtained by quantitative genomic DNA sequencing and the reference nucleotide sequence is a genomic sequence, the RRD may represent the quantitative presence or absence of the genomic base pairs in the sample. When the sample is an environmental sample, the methods provided herein may result in the representation of one or more genomic sequences by a character string (RRD) having information on the quantitative presence or absence of genomic DNA at base pair resolution.

[0068] The obtained single strings (RRDs) can be used in analyses, preferably (computer-implemented) analyses of (large-scale) nucleic acid sequencing datasets, such as (large-scale) RNA-seq datasets and / or (large-scale) quantitative DNA sequencing datasets. The obtained single strings can also be used, for example, in artificial intelligence processes or methods for predicting genes, exons and / or introns in a genome sequence. Thus, the methods provided herein allow transcriptome data to be easily mapped to genome sequences, making them directly usable for further processing, such as deep learning, machine learning and / or artificial intelligence processes and / or applications.

[0069] Any single-character system may be suitable for use in step (c) of the method of the present invention.

[0070] Without being limited thereto, since the ranking from low to high is inherent and intuitive, it may be straightforward to use single digits of (Arabic) numerals, i.e., numbers selected from the group consisting of {0, 1, 2, 3, 4, 5, 6, 7, 8, 9} to represent RRDs (i.e., 0 is the least significant single digit and 9 is the most significant single digit). Optionally, the total number of categories represented by RRDs in the class assignment model is 10. In principle, any single letter may be used for each RRD, but if a total of 10 RRDs are used, single digits of (Arabic) numerals may be used, and without being limited to such a system, the RRDs are assigned 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 from lowest to highest, with the lowest range being assigned the number 0, the next range being assigned the number 1, and the highest range being assigned the number 9.

[0071] In addition to the 10 single-character symbols of the Arabic numeral system, there are many other suitable alternatives. For example, the 52 letters of the (Latin) alphabet, the 94 characters of the ASCII table, (part of) Unicode characters, etc. can be used. For example, if the reading depth values ​​are classified in a total of 52 RRDs, all 52 letters of the alphabet can be used to indicate the RRDs, and optionally the order of the letters of the alphabet can be respected (i.e., "a" for the lowest RRD and "z" for the highest RRD). Similarly, other single-character systems can be used, such as symbols of the ASCII code table and Unicode characters. Those skilled in the art will understand that these numbers per system (e.g., Arabic numerals, Latin alphabet, symbols of the ASCII code table, Unicode characters, etc.) can be easily increased, for example, by using symbols of different colors, different sizes, or, in the case of alphabetic characters, by using lowercase letters, uppercase letters, and / or different font types.

[0072] The number of RRDs or single characters is less than, preferably significantly less than, the maximum read depth of the obtained nucleotide sequencing data set. Preferably, the number of RRDs used in the methods provided herein is at least about 5, 10, 100, 500, 1000, 5000, 10000, 50000, or at least about 100000 times lower than the maximum read depth of the obtained data set.

[0073] It is understood by those skilled in the art that when it is desired to divide the read depth value of nucleotide sequencing data into less than the total number of characters of a specific system, for example, less than 52 ranges when the Latin alphabet is used, a part of the system can be used, for example, fewer letters of the alphabet can be used as single characters for RRD.Similarly, when it is desired to divide the read depth value of nucleotide sequencing data into a number that exceeds the total number of characters of a specific system, a combination of systems can be used.For example, when the read depth value is divided into 52 or more RRDs, the alphabet can be used in combination with, for example, the Arabic numeral system.It is understood by those skilled in the art that other (combined or partial) systems of distinct single characters can also be used.

[0074] The assignment of read depth values ​​to RRDs can be performed by first setting the boundaries of each RRD. Independent of the single-character system used in step (c) of the method of the present invention, each RRD ranked from the bottom has a ranking value. In this specification, for the purpose of calculating the boundaries, the ranking values ​​are each assigned a natural number (i.e., selected from the set {0, 1, 2, 3, 4.....N}), where N is the highest ranking value, and the total number of RRDs is given by N+1 (i.e., if the total number of RRDs is 10, the RRD with the highest ranking value will have a ranking value of 9).

[0075] Preferably, the read depth values ​​of the nucleotide sequencing dataset are converted into RRDs by class assignment, and the total number of RRDs is 10. Although the last single letter used for each RRD is irrelevant as long as each RRD receives a unique single letter, the inventors have developed an algorithm to calculate the boundaries of each RRD based on the ranking of these RRDs (referred to herein as the ranking value of the RRD), and low to high ranked RRDs are assigned consecutive ranking values ​​of 0, 1, 2, 3, 4, 5, 6, 7, 8, and 9, respectively, i.e., 0 is the lowest ranked RRD and 9 is the highest ranked RRD.

[0076] Preferably, the read depth values ​​of a nucleotide sequencing dataset are converted into RRDs by class assignment, and the total number of RRDs is 10. The final single letter used for each RRD is irrelevant as long as each RRD receives its own unique single letter. However, the inventors have developed an algorithm to calculate the boundaries of each RRD based on the ranking of these RRDs (referred to herein as the ranking value of the RRD), where low- to high-ranking RRDs are assigned consecutive ranking values ​​of 0, 1, 2, 3, 4, 5, 6, 7, 8, and 9, respectively, i.e., 0 is the lowest-ranking RRD, and 9 is the highest-ranking RRD. Since these ranking values ​​themselves are single digits, optionally, the single letters used to indicate RRDs in step (c) of the method for processing nucleotide sequencing data provided herein are identical to the ranking values ​​of these RRDs. In other words, the ranking value of an RRD can be identical to (the single letter of) the RRD. However, those skilled in the art will understand that these 10 ranking values ​​can be converted to any other 10-fold single-letter system. As a non-limiting example, an RRD with a ranking value of 0 may be designated with the letter "a," an RRD with a ranking value of 1 may be designated with the letter "b," and an RRD with a ranking value of 1 may be designated with the letter "c."

[0077] As mentioned above, each RRD has a specific read depth upper limit, and read depth values ​​exceeding the upper limit are assigned to higher-level RRDs. In other words, read depth values ​​are assigned to the lowest-level RRD that do not exceed the read depth upper limit.

[0078] In principle, any model can be used to fit the read count values ​​on the RRDs, including, but not limited to, linear, exponential, polynomial, logarithmic, power, etc. Preferably, the model selected depends on the nature of the data in the dataset, and the model is used to capture as much information as possible. For example, if the dataset is a quantitative DNA sequencing dataset reflecting the amount and type of genomic DNA of one or more species present in an environmental sample, a linear model can be used. In a linear model, the read count values ​​are divided evenly among the total number of RRDs. If the dataset is an RNA-seq dataset reflecting the amount and type of genes expressed from the genome, an exponential model can be used.

[0079] Preferably, a "zero dogma" is applied regardless of the model selected for class assignment. This can be explained as follows: Some nucleotides of a reference nucleotide sequence may not be represented or may not exist in a nucleotide sequencing dataset. For example, in a sample from which a nucleotide sequencing dataset, which may be an RNA-seq dataset, is obtained, some nucleotide sequences of the genome are not expressed or are expressed at levels below the detection level. Therefore, the "zero dogma" is understood herein as assigning nucleotides of a reference nucleotide sequence for which there is no data or zero transcripts in the nucleotide sequencing dataset to the lowest RRD. Without being limited thereto, this lowest RRD may be assigned the single character "0". Preferably, nucleotides of a reference nucleotide sequence for which a read count value greater than zero is found in the nucleotide sequencing dataset are assigned to a higher RRD based on the class assignment model.

[0080] In a linear regression model, the read count values ​​from zero to the maximum read count value of the dataset can be evenly divided among the remaining RRDs. The ranking values ​​of the RRDs can be used to define upper and / or lower bounds for the RRDs. For example, when ranking the RRDs from lowest to highest, the lowest RRD is assigned a ranking value of 0, and each subsequent RRD is assigned a ranking value one higher than the ranking value of the previous RRD. For example, for 10 RRDs, the RRDs ranked from lowest to highest are assigned ranking values ​​of 0, 1, 2, 3, 4, 5, 6, 7, 8, and 9, respectively. Using these ranking values ​​of the RRDs, the following formula can be used to calculate the upper bound of each RRD according to a linear class assignment model: UB=(RV / N)*MRD (Equation 1) During the ceremony, - UB is the upper read depth limit for a given RRD, - RV is a given ranking value selected from the group consisting of {0, 1, 2, 3, 4, ......N} of the RRD for which the upper bound is calculated; - N is the ranking value of the top RRD, and - MRD is the maximum read depth.

[0081] In a non-limiting example, if N is an upper limit of 9, the total number of RRDs is 10(N+1), MRD=9000, and the read depth value is assigned to the particular RRD if it is within the range of the RRD shown in Table 1.

[0082] [Table 1]

[0083] In the example of Table 1, the upper bound of an RRD is understood to be the maximum value that fits within the RRD. However, one skilled in the art will understand that the upper bound can also be determined as the lowest value of the RRD that is then ranked. In the latter case, for example, the range of an RRD with a ranking value of 1 should be displayed as [1000, 2000>. As will be apparent to one skilled in the art, it does not matter what system is used, as long as the same system is used throughout.

[0084] It should be understood that the reverse ranking method can be carried out herein and achieves the same result.Each RRD has a specific read depth lower limit, and the read depth value that is equal to or less than said lower limit is assigned to a lower RRD.In this ranking method, the read depth value is assigned to the top RRD whose read depth value is not equal to or less than the read depth lower limit.In ranking, it may be sufficient to only identify the upper limit of each RRD.Similarly, it may be sufficient to only identify the lower limit of each RRD.

[0085] Optionally, exponential class assignment model is used.This exponential model is particularly suitable for processing RNA-seq data.In this model, the lower and upper limits of the read depth of the lowest RRD can be 0 (zero dogma), the upper limit of the read depth of the highest RRD can be the maximum read depth of nucleotide sequencing data set, and any one of the RRDs in the middle position can have the upper limit of the read depth given by the following formula: UB=f·x RV (Formula 2) During the ceremony, - UB is the upper read depth limit for a given RRD, x is preferably a cardinal number having a value between 1 and Euler's number (e), f is a scaling factor preferably having a value of at least 1, - RV is a given ranking value selected from the group consisting of {1, 2, 3, 4, ...... N} of the RRD for which the upper bound is calculated, and - N is the ranking value of the top RRD.

[0086] Euler's number (e) is known as 2.718281... According to this exponential model, the lower and upper bounds of the lowest RRD are 0 (zero dogma), and then the lower bound of every ranked RRD is equal to the upper bound of the lower RRD, and the total number of RRDs is N1. Therefore, preferably, each RRD higher than the lowest RRD has a read depth lower bound given by the following formula: LB=(f·x RV-1 ) (Formula 3) During the ceremony, - LB is the lower bound of the read depth for a given RRD, x is a cardinal number having a value selected from the range from 1 to Euler's number (e), f is a scaling factor having a value of at least 1; - RV is a given ranking value selected from the group consisting of {1, 2, 3, 4, ...... N} of the RRD for which the upper bound is calculated, and - N is the ranking value of the top RRD.

[0087] Thus, in a preferred embodiment, a read depth value is assigned a particular RRD if it falls within the region of said RRD according to the exponential model of Equations 2 and 3 shown in Table 2.

[0088] [Table 2]

[0089] The base x of the exponential model in Equations 2 and 3 is x=(MRD / f) 1 / N (Formula 4) wherein: - x is the base of equations 2 and 3, - MRD is the maximum read depth, - f is the scaling factor for equations 2 and 3, and - N is the ranking value of the top RRD.

[0090] The scaling factor f of the exponential models of Equations 2, 3, and 4 may have a value that depends on the maximum read depth (MRD) of the dataset. More specifically, f is expressed by the formula: f=MRD / e N (Formula 5) wherein: - f is the scaling factor for equations 2, 3 and 4, - MRD is the maximum read depth, - e is Euler's number, and - N is the ranking value of the top RRD.

[0091] When the base number x and the scaling factor f are defined according to Equations 4 and 5, by integrating Equations 4 and 5 into Equation 2, the calculation formula for the UB of the RRD is: UB=(MRD·e RV ) / e N (Formula 6) In the formula, - UB is the upper read depth limit for a given RRD, - MRD is the maximum read depth, - RV is a given ranking value selected from the group consisting of {1, 2, 3, 4, ...... N} of the RRD for which the upper bound is calculated, and - N is the ranking value of the top RRD.

[0092] Similarly, when the base number x and the scaling factor f are defined according to Equation 4 and Equation 5, by integrating Equation 4 and Equation 5 into Equation 3, the calculation formula for LB of RRD is: LB = (MRD·e RV-1 ) / e N (Formula 7) In the formula, - UB is the upper read depth limit for a given RRD, - MRD is the maximum read depth, - RV is a given ranking value selected from the group consisting of {1, 2, 3, 4, ...... N} of the RRD for which the upper bound is calculated, and - N is the ranking value of the top RRD.

[0093] In one embodiment, the maximum read depth (MRD) of the dataset is set to a value e N f has the value of Equation 5 (i.e., MRD) unless it is less than <e N where e is Euler's number and N is the ranking value of the top RRD), in this case f has a value of 1. This may be the case for datasets from samples with low amounts of starting material, such as, but not limited to, single-cell RNA-seq datasets. For example, in this embodiment, when N=9, for datasets with a maximum read depth (MRD) of less than 8103, f may have a value of 1, and for datasets with a maximum read depth of at least 8103, f may have a value defined in Equation 5. Thus, in this embodiment, MRD≧e N , then f has the value of Equation 5, and the upper bound can be calculated using Equation 6, and the MRD <e N In this embodiment, f preferably has a value of 1. <e N In this case, by integrating f=1 in Equation 4 and Equation 2, the calculation formula for UB of RRD is: UB=(MRD) RV / N (Formula 8) In the formula, - UB is the upper read depth limit for a given RRD, - MRD is the maximum read depth, - RV is a given ranking value selected from the group consisting of {1, 2, 3, 4, ...... N} of the RRD for which the upper bound is calculated, and - N is the ranking value of the top RRD.

[0094] Similarly, in this embodiment, MRD <e N In this case, by integrating f=1 in Equation 4 and Equation 3, the calculation formula for LB of RRD is: LB=(MRD) RV-1 / N (Formula 9) In the formula, - UB is the upper read depth limit for a given RRD, - MRD is the maximum read depth, - RV is a given ranking value selected from the group consisting of {1, 2, 3, 4, ...... N} of the RRD for which the upper bound is calculated, and - N is the ranking value of the top RRD.

[0095] Therefore, in a preferred embodiment, the maximum read depth is at least e N For a dataset where e, the upper bound of RRD can be calculated using Equation 6, and the lower bound of RRD can be calculated using Equation 7 (e.g., when N=9, MRD≧8103). N If the maximum read depth is less than 1 / 3 (e.g., N=9, MRD<8103), the upper limit of RD can be calculated using Equation 8, and the lower limit of RRD can be calculated using Equation 9.

[0096] A read depth value is assigned to a particular RRD if the value is within the region of the RRD. In a preferred embodiment, N=9, and the regions of the RRD are shown in Table 3.

[0097] [Table 3]

[0098] Table 4 provides the preferred read depth value region for the RRD of the exponential model provided herein for N=9.

[0099] [Table 4]

[0100] In addition to a single letter indicating the RRD, further information, preferably at base pair resolution, may be added to the reference nucleotide sequence. For example, information about modifications, such as epigenetic modifications of nucleotides, may be added. This information may be captured by a single letter as defined herein. Optionally, these letters are selected from the same set of letters used to identify the RRD. Optionally, the order and / or specific position of this data may indicate whether it is for the letter representing the RRD or for the letter representing information about the nucleotide modification. As an example, the letter indicating the RRD may be in the first line, and the letter indicating the nucleotide modification may be in a second line at a specific position relative to the first line, e.g., below the first line. As another example, for each nucleotide in the reference sequence, all letters indicating different properties of each nucleotide (e.g., RRD and modification) are located adjacent to each other at fixed positions. In such cases, and in the case of two different properties, the reference sequence is represented by a string of single-letter pairs, where the first letter of the pair indicates one of the properties and the second letter of the pair indicates the other property. Optionally, the letters indicating the nucleotide modifications are selected from a different set of letters than those used for the RRDs. For example, optionally, the RRDs are designated by numbers (e.g., each RRD is selected from the group of 0, 1, 2, 3, 4, 5, 6, 7, 8, and 9), and the modifications are designated by letters (e.g., "A" for methylation, "B" for glycosylation, "C" for bond isomerization of uridine (formation of pseudouridine), etc.). Optionally, additional information can be added using additional letters and / or additional positions (e.g., when similar single letters are used). For example, DNA-seq (such as CNV-seq) data can be added to RNA-seq data.

[0101] It is also possible to include additional information in the same single character of the RRD. For example, if the total number of RRDs is 10 and there is one nucleotide modification, all combinations can be represented by a total of 20 different single characters. For example, if a nucleotide is unmodified, the nucleotide is initially assigned an RRD selected from the group consisting of 0, 1, 2, 3, 4, 5, 6, 7, 8, and 9 according to the RRD level. If a nucleotide is methylated, the nucleotide can be assigned an RRD selected from the group consisting of a, b, c, d, e, f, g, h, i, and j according to the RRD level. The obtained RRD string can contain a mixture of numbers and letters. Therefore, this method can further include a step of obtaining data on nucleotide modifications. This data is preferably obtained or can be obtained from the same sample as the sample that provides the nucleotide sequencing dataset including the read depth value.

[0102] Further provided is the use of processed nucleotide sequencing data obtainable by the methods provided herein for further processing and / or analysis. For example, in the case of transcriptome data, the processed data (preferably a string of RRDs mapped to a reference genome sequence) obtained by the methods of the present invention can be used for further processing and / or analysis, such as comparing expression profiling between different samples, screening for alternative splicing, analyzing allele-specific expression, and / or predicting one or more genome sequences. Preferably, the RRDs mapped to the reference genome sequence at base pair resolution, obtainable by the method for processing a transcriptome dataset obtainable herein, are used to compare and / or profile expression patterns between multiple samples. The methods provided herein can omit additional normalization steps required to correct for experimental variations, such as library fragment size, sequence composition bias, and read depth, in order to accurately estimate the expression level values ​​of genes and / or RNA transcripts in different samples.

[0103] For further processing and / or analysis, multiple nucleotide sequencing datasets can be combined. For example, nucleotide sequencing datasets from the same or similar individuals and / or nucleotide sequencing datasets from the same or similar cell or tissue samples can be combined. Preferably, these multiple nucleotide sequencing datasets are combined after processing by the method provided herein. In other words, preferably, multiple nucleotide sequencing datasets are each converted into RRDs, and these RRDs are combined. Alternatively, the nucleotide sequencing datasets are combined before processing by the method provided herein. In other words, in such an embodiment, the nucleotide sequencing datasets are preferably combined after normalization between libraries based on the number of reads per library, and then converted into RRDs. Methods for normalizing across multiple nucleotide sequencing datasets are known in the art, such as, but not limited to, Limma (Ritchie et al., Nucleic Acids Research, 2015, 43(7):e47), Combat (Johnson et al., Biostatistics, 2007, 8(1):118-127; Leek et al., Bioinformatics, 2012, 28(6), pp.882-883), and further methods described in Cole et al. (Cell Syst. 2019, Apr 24;8(4):315-328). Preferably, the processed nucleotide sequencing datasets to be combined were mapped to the same reference nucleotide sequence in step (b) of the method provided herein. Combining the processed nucleotide sequencing datasets may result in, for example, RRD complementarity for one or more nucleotide sequences of the reference nucleotide sequence for which data is present in only one of the combined datasets. Alternatively or additionally, such combination may result in multiple RRDs for a particular reference nucleotide sequence base pair that are present in more than one of the data sets for which the data was combined.Preferably, these multiple RRDs for a single reference nucleotide are reduced to a single RRD.If RRDs are numerical, it is desirable to calculate the average or median value, round it to a single digit, and then map it to the reference nucleotide sequence to reduce it to a single value.Preferably, the lowest RRD is excluded.More specifically, when there is a combination of processed nucleotide sequencing data sets, and a base pair of the reference sequence in one or more of these data sets represents any read (i.e., a read depth greater than 0), the base pair is assigned the penultimate lowest RRD of the reference sequence in the combined data set.

[0104] If the RRD is represented by other characters, such as letters of the alphabet, these characters are first converted to a numerical value corresponding to the rank of each character, the mean or median is calculated, and then rounded, if necessary, to reconcile with a particular character, which is then mapped to a reference nucleotide sequence.

[0105] Optionally, the nucleotide sequencing data of the methods provided herein is part of a method for comparing nucleic acid molecular profiles between different samples, where a profile is understood herein as information regarding the type of nucleic acid, the nucleotide sequence and / or the amount of nucleic acid molecules.

[0106] Optionally, the method for processing a nucleotide sequencing dataset is part of a method for comparing nucleic acid molecular profiles between different samples, and / or an RRD mapped to a reference nucleotide sequence at base pair resolution, obtainable by the method for processing a nucleotide sequencing dataset defined herein, is used to compare nucleic acid molecular profiles between different samples. Optionally, the different samples are obtained from the same or similar individuals. The different samples may be obtained and / or collected from the individuals at different time points, from different tissues, and / or before and after (different) treatments. In one embodiment, the method for processing a nucleotide sequencing dataset is part of a method for comparing nucleic acid molecular profiles of the same or similar individuals under different circumstances. There are an almost infinite variety of circumstances, such as, but not limited to, differences in exposure to biotic stress, abiotic stress, nutrient availability, water, toxins, sunlight, etc., or differences between individuals themselves, such as age, disease, location of sample collection, etc. (e.g., in the case of plants, one sample may be a fruit sample and another sample may be a root sample from the same plant). Alternatively, different samples are obtained from different individuals, preferably from the same or similar tissues, at the same or similar time points and / or after the same or similar treatments.

[0107] The method for comparing nucleic acid molecular profiles between different samples preferably includes providing multiple different samples, and steps (a), (c), and / or (d) can be performed in parallel or sequentially for each sample. Optionally, step (a) is performed in parallel for multiple samples up to the sequencing step of step (a)(3) defined herein, and optionally, the nucleic acid molecules of each sample are labeled or tagged with a sample identifier sequence, and the samples and / or (isolated and / or enriched) tagged nucleic acid molecules are then pooled before subjecting the pooled sample to nucleotide sequencing. After sequencing, the data, preferably RNA-seq data, can be demultiplexed based on the sample identifier sequence. The subsequent (demultiplexed) RRDs can be mapped to a reference nucleotide sequence, preferably a genomic sequence associated with the original sample. In the case of different samples from different individuals of the same species, the same reference nucleotide sequence, preferably a reference genomic sequence, can be used for each such different individual.

[0108] Preferably, the RRDs mapped to the reference nucleotide sequence at base pair resolution, which can be obtained by the method for processing nucleotide sequencing datasets provided herein, are used to compare nucleic acid profiles between multiple samples. In this method, preferably, the additional normalization step required to correct experimental variations such as sequencing library fragment size, sequence composition bias, and read depth is omitted to accurately compare the nucleic acid molecule levels of different samples.

[0109] In the case of transcriptome data, the methods provided herein can be part of a method for comparing expression profiles between different samples.

[0110] Furthermore, in the case of a transcriptome dataset, the method provided herein can be part of a method for predicting one or more genome sequences. A gene sequence is understood as a sequence containing an open reading frame. An open reading frame is a sequence between a start codon and a stop codon in genomic DNA, which can be transcribed and then processed to form mRNA. An open reading frame includes one or more exons and, optionally, one or more introns. The prediction of a gene sequence can be a prediction of at least one exon, intron, exon and / or intron boundaries, an open reading frame, a regulatory sequence, and the entire gene sequence, wherein the entire gene sequence includes an open reading frame and one or more transcriptional regulatory sequences, such as, but not limited to, a promoter sequence. Using RRDs mapped to a genome sequence at base pair resolution, which can be obtained by the method for processing a transcriptome dataset defined herein, unknown gene sequences can be predicted based on the expression pattern exhibited by the RRDs mapped to known reference gene sequences whose locations in the genome are also known.

[0111] Optionally, provided herein is a process for predicting genes and gene expression levels of a sample without obtaining a transcriptome dataset of the sample. Preferably, the prediction is performed by machine learning. Thus, (A) providing RRDs mapped to one or more reference genome sequences at base pair resolution, obtainable by the method for processing a transcriptome dataset provided herein; (B) providing gene annotation data for at least a portion of the one or more reference genome sequences; (C) assigning, by a training engine, the gene annotation data of (B) to RRDs mapped to one or more reference genome sequences of (A); (D) training a machine learning model using the gene annotation data assigned to the RRDs mapped in step (C), wherein during training the machine learning model learns to assign gene annotations to one or more additional genome sequences to which the RRDs are mapped; (E) obtaining a machine learning model; Also provided is a method for predicting one or more genomic sequences and optionally gene expression levels of said genomic sequences, comprising:

[0112] 3 depicts a flowchart showing the method, comprising step 301 of providing RRDs mapped to one or more reference genome sequences at base pair resolution, step 302 of providing gene annotation data for at least a portion of the one or more reference genome sequences, step 303 of assigning gene annotation data to RRDs by a training engine, step 304 of training a machine learning model using the gene annotation data assigned to the mapped RRDs in step 303, and step 305 of obtaining the machine learning model. Thus, in FIG. 3 , step 301 corresponds to step (A) described herein, step 302 corresponds to step (B) described herein, step 303 corresponds to step (C) described herein, step 304 corresponds to step (D) described herein, and step 305 corresponds to step (E) described herein.

[0113] Preferably, the assignment by the training engine in step (C) is automated and / or performed automatically. Preferably, the method is a computer-implemented method. Optionally, the one or more additional genomic sequences in step (D) are sequences related to the sequences for which genetic annotation data is provided in (B), where "related" as used herein may mean that the genomic sequences are from samples of the same individual, different individuals of the same species, or individuals of different species of the same genus, family, or order. Furthermore, "related" as used herein may mean that the transcriptome data mapped to the genomic sequence are obtained from individuals who have undergone a similar treatment.

[0114] Optionally, the methods provided herein are performed using multiple transcriptome datasets from multiple samples, and / or multiple strings of RRDs are obtained, optionally mapped to a genome sequence. Optionally, the multiple samples are highly similar, preferably from the same species, the same individual, the same tissue type, the same tissue, and / or the same cell type. Optionally, the multiple samples are from the same species, the same tissue type, but from different individuals. Optionally, the multiple samples are from the same individual, but from different tissue types. Optionally, the multiple samples are from the same individual, but from different cell types. Optionally, the multiple samples are from the same species and the same tissue, but the different individuals have been treated differently, e.g., mutagenized or not, exposed to biotic stress or not, exposed to abiotic stress or not, etc. Optionally, the multiple samples are from the same species and the same tissue, but the different individuals are at different developmental stages, e.g., younger or older.

[0115] Optionally, the method of predicting a gene sequence defined herein is preceded by the method of processing transcriptome data defined herein. Optionally, the method of predicting a gene sequence comprises predicting a gene sequence and / or comparing gene expression levels of a plurality of samples.

[0116] Optionally, the method of processing a nucleotide sequencing dataset, comparing nucleic acid molecular profiles and / or gene expression levels, and / or predicting a genome sequence is a computer-implemented method. Preferably, the computer is programmed to perform any one of the methods provided herein. Preferably, the computer comprises a computer-readable storage means comprising a (compressed) nucleotide sequencing dataset processed according to the methods provided herein. Also provided is a computer-readable storage means configured to perform the methods of processing a nucleotide sequencing dataset, predicting a genome sequence, and / or comparing nucleic acid molecular profiling and / or gene expression levels provided herein. Also provided is a computer program product comprising instructions, when executed by a processor system, causing the processor system to perform the methods of processing nucleotide sequencing data, predicting a genome sequence, and / or comparing nucleic acid molecular profiling and / or gene expression levels provided herein. Preferably, the methods provided herein are methods performed by an electronic device.

[0117] Further provided is a storage means for storing instructions, the instructions configured to cause the processor to perform the methods for processing nucleic acid sequencing data, predicting genomic sequences and / or performing nucleic acid molecular profiling and / or comparing gene expression levels provided herein.

[0118] Preferably, the processed nucleotide sequencing data obtained or obtainable by the methods provided herein is stored in a computer readable storage means. Also provided is a computer readable storage means comprising the processed nucleotide sequencing data obtained or obtainable by the methods provided herein.

[0119] Further provided is the use of a processed nucleotide sequencing dataset obtained or obtainable by the methods provided herein and / or a computer readable storage means comprising said processed nucleotide sequencing dataset obtained or obtainable by the methods provided herein for further processing and / or analysis, such as predicting one or more genomic sequences and / or comparing nucleic acid molecular profiling and / or gene expression profiling between different samples.

[0120] FIG. 4 illustrates a computing device 400, such as a mobile phone, tablet, laptop, desktop, television, or other computing device, for performing the methods of the present invention, for example, any of FIGS.

[0121] The device 400 may include at least one of a processor 401 , a display 402 , a communication unit 403 , a keyboard, a memory 405 , a camera 406 and other input / output units 407 .

[0122] The processor 401 is configured to execute programs / instructions stored in the memory 405 by controlling other components such as the display 402, the communication unit 403, the keyboard 404, the memory 405, the camera 406 and other input / output units 407.

[0123] Display 402 is controlled by processor 401 and may perform all display functions (and / or input functions, if it is a touchscreen) in the present invention, such as displaying any of the nucleotide sequencing dataset, the reference nucleotide sequence, a single letter indicating the RRD, and the mapping results, for example, in the method of Figure 1, displaying any of the subsets, sequencing libraries, and read depth values, for example, in the method of Figure 2, and displaying any of the RRDs, one or more reference genome sequences, and gene annotation data, for example, in the method of Figure 3. If display 402 is a touchscreen, it may be used to input all relevant information for carrying out the method, such as information obtained in any of the "providing" steps of either Figures 1 and 1. or information obtained in steps requiring user input (e.g., steps 303 and / or 304), and keyboard 404 may perform the same function.

[0124] The communication unit 403 is controlled by the processor 401 and may perform all communication functions in the present invention. For example, when an external device 410 (e.g., a server) is used to perform some of the functions of the steps in Figures 1 to 3 (e.g., step 101 in Figure 1 may be performed on the local device 400, and the final step 104 may be performed on the local device 400, but other steps in Figure 1 may be performed on the external device 410 (e.g., a server), or all of the steps may be performed on the local device 400, or some of the steps may be performed on the local device 400 and the remaining parts of the steps may be performed on at least one or more external devices 410), messages may be communicated via the communication unit 403. Optionally, datasets and other information used in the present invention, such as, for example, any of the nucleotide sequence dataset, reference nucleotide sequence, single letter indicating RRD, mapping results, subset, sequencing library, read depth value, RRD, one or more reference genome sequences, gene annotation data, and training engine / model, may be stored in memory of the external device 410 or the local device 400. If they are stored in the external device 410, they are notified to the local device 400 via the communication unit 403.

[0125] The memory 405 may be configured to store instructions for carrying out the methods of the present invention, such as the information described above.

[0126] The camera 406 may be configured to capture images or may be used as an input device to scan documents / images to obtain the above-mentioned information, although this is optional in the present invention.

[0127] Other input / output units 407 may be configured to perform other or the same input / output functions of the present invention, for example to receive user input for the allocation of training engines and training of models.

[0128] Device 400 may be configured for use in computer-implemented methods for predicting processed nucleotide sequence data, e.g., genomic sequences, obtained or obtainable by any of the methods of the present invention, and / or comparing gene expression patterns within or between samples.

Claims

1. 1. A method for processing a nucleotide sequencing dataset, comprising: (a) providing said nucleotide sequencing dataset comprising read depth values ​​assigned to nucleotide sequences; (b) providing a reference nucleotide sequence; (c) converting the read depth values ​​present in the nucleotide sequencing dataset into a single letter indicating relative read depth (RRD) by ordinal class assignment; (d) mapping said single characters to said reference nucleotide sequence of (b) at base pair resolution; wherein the method is a computer-implemented method.

2. 2. The method of claim 1, wherein the nucleotide sequencing dataset of (a) is a transcriptome dataset obtained by RNA-seq.

3. 2. The method of claim 1, wherein the nucleotide sequencing dataset of (a) is a genomic dataset obtained by DNA sequencing.

4. The method according to any one of claims 1 to 3, wherein the reference nucleotide sequence is a genomic sequence, preferably a whole genome sequence.

5. Step (a) (1) providing a nucleic acid sample; (2) optionally enriching and / or isolating a subset of nucleic acid molecules from at least a portion of said nucleic acid sample; (3) preparing a sequencing library of the nucleic acid molecules of (1) or the enriched and / or isolated nucleic acid molecules of (2); (4) (Sub) step of obtaining a nucleotide sequence by sequencing the sequencing library of (3); (5) assigning read depth values ​​to the nucleotide sequences obtained in (4), thereby obtaining a nucleotide sequencing dataset comprising read depth values ​​assigned to the nucleotide sequences of the sequenced nucleic acid molecules; The method according to any one of claims 1 to 4, comprising:

6. 6. The method of claim 5, wherein the nucleic acid molecules of (1) and optionally the enriched and / or isolated nucleic acid molecules of (2) are mRNA molecules, and the method further comprises the (sub)step of converting the mRNA to the cDNA prior to step (3) of preparing a cDNA sequencing library.

7. The method according to any one of claims 1 to 6, wherein the reference nucleotide sequence of (b) is obtained by (whole) genome sequencing.

8. In step (c), the read depth values ​​are sorted in N+1 consecutive RRDs, each RRD including a read depth upper limit, a read depth value is assigned to a lowest RRD whose read depth value does not exceed the read depth upper limit of the RRD, the read depth upper limit of the lowest RRD is 0, and for the remaining RRDs, the read depth upper limit is UB = f·x RV (Equation 2) (In the formula, - UB is the upper read depth limit for a given RRD; x is preferably a cardinal number having a value between 1 and Euler's number (e); f is a scaling factor preferably having a value of at least 1, RV is a given ranking value selected from the group consisting of {1, 2, 3, 4, ... N} of the RRD for which the upper bound is calculated; and - N is the ranking value of the top RRD) The method according to any one of claims 1 to 7, wherein

9. The value of the base x is x = (MRD / f) 1/N (Equation 4) (In the formula, x is the base number of formulas 2 and 3, - MRD is the maximum read depth, and N is the ranking value of the top RRD. The method of claim 8, wherein

10. The value of the scaling factor f is f = MRD / e N (Equation 5) (In the formula, f is the scaling factor for equations 2 and 4, MRD is the maximum read depth, e is the Euler number, and N is the ranking value of the top RRD.

10. The method of claim 8 or 9, wherein the

11. f is MRD≧e N , is given by Equation 5, and MRD<e N If f=1, the read depth upper limit of the lowest RRD is 0, and for the remaining RRDs, the read depth upper limit is MRD≧e N in the case of, UB = (MRD・e RV ) / e N (Equation 6) and MRD<e N in the case of, UB = (MRD) RV/N (Equation 8) where: - UB is the upper read depth limit for a given RRD; MRD is the maximum read depth, - e is the Euler number, RV is a given ranking value selected from the group consisting of {1, 2, 3, 4, ... N} of the RRD for which the upper bound is calculated; and The method according to any one of claims 8 to 10, wherein N is the ranking value of the top RRD.

12. The method of any one of claims 1 to 11, performed on a plurality of samples.

13. Preferably, use of processed nucleotide sequencing data obtained or obtainable by the method of any one of claims 1 to 12 in a computer-implemented method for predicting genome sequence and / or comparing gene expression patterns within or between samples.

14. 14. The use of embodiment 13, wherein the computer-implemented method comprises machine learning.

15. A computer readable storage medium comprising processed nucleotide sequencing data obtained by the method of any one of claims 1 to 12.

16. A computing device including at least one processor, said processor configured to perform the method of any one of claims 1 to 12.