Compositions, kits, uses of the compositions, uses of the kits, and methods

A composition of three artificial RNA molecules with specific sequences addresses the limitations of single spike-ins by ensuring consistent RNA extraction and detection, improving data quality and assay linearity in RNA analysis.

JP2025523920APending Publication Date: 2025-07-25HUMMINGBIRD DIAGNOSTICS GMBH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025502580
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-20
Filing Date
2023-07-03
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing RNA analysis methods face challenges in accurately quantifying small RNAs from biological samples due to biases in extraction and detection, with single spike-ins like cel-miR-39-3p failing to reflect primary sequence diversity and relative abundances, leading to inconsistent assay linearity and quality control issues.

Method used

A composition comprising at least three artificial RNA molecules with specific nucleotide sequences (SEQ ID NOs: 1 to 4) is used as a spike-in cocktail to mimic endogenous miRNAs, ensuring recovery at expected levels and linearity, and serving as a quality control and normalization tool for RNA extraction and detection.

Benefits of technology

The cocktail improves RNA data quality by maintaining linearity and consistency across assays, enabling robust evaluation of assay efficiency and correcting for batch effects, thereby enhancing the accuracy of RNA analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025523920000006
    Figure 2025523920000006
  • Figure 2025523920000007
    Figure 2025523920000007
  • Figure 2025523920000008
    Figure 2025523920000008
Patent Text Reader

Abstract

The present invention relates to a composition that can be used as a spike-in cocktail. Furthermore, the present invention relates to a kit containing the composition. Furthermore, the present invention relates to a method for inspecting a sample using the composition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a composition that can be used as a spike-in cocktail. Further, the present invention relates to a kit containing the composition. Further, the present invention relates to a method for examining a sample using the composition.

Background Art

[0002] RNAs such as small RNAs in biological samples such as body fluids play an important role as prognostic and diagnostic biomarkers for many human disease states. However, the accurate analysis of these RNAs in biological samples such as body fluids is very important regarding the significance of these data in the biomarker field.

[0003] The analysis of RNA expression profiles involves several complex steps. (1) It is necessary to extract RNA containing small RNAs from a biological source such as blood or saliva without distorting the original relative abundance of the RNA (it is necessary to maintain linearity). (2) It is necessary to reliably quantify the abundance of small RNAs. Since both steps can introduce biases inconsistently (e.g., due to the presence of inhibitors of cDNA synthesis), it is important to standardize and monitor the efficiency of RNA extraction and detection. These challenges can be addressed in several ways.

[0004] First, performing RNA integrity analysis of endogenous RNA can be used for quality control by comparing it with the electrophoretic RNA profile expected based on previously analyzed samples, for example, the trace of a small RNA bioanalyzer. However, such general analysis only provides a rough view of the integrity of the sample and lacks detailed information.

[0005] Second, it is possible to define reference values for endogenous small RNAs, such as housekeeping genes belonging to nuclear or splicing RNAs, which serve as a reference when comparing newly measured samples. Samples with housekeeping genes outside the expected range may thus be excluded from the analysis as outliers. However, using the intrinsic characteristics of samples for quality control makes it impossible to distinguish between samples that have undergone technically incomplete experimental processing and samples that exhibit true biological outliers.

[0006] To overcome this problem, academic and commercial laboratories have started using exogenous spike-in sequences that can mimic endogenous target RNA populations, thereby enabling the monitoring of assay inefficiencies. The most commonly used spike-in for small RNA analysis is cel-miR-39-3p from C. elegans worms. However, there are several drawbacks to using a single spike-in, such as cel-miR-39-3p. First, a single spike-in does not fully reflect the considerable primary sequence diversity of small RNAs as a whole. Effects specific to the primary sequence, such as RNA secondary structure, free energy, and GC content, can have a profound impact on extraction and detection efficiency, and a single species of spike-in is insufficient to faithfully reflect this. Furthermore, a single spike-in cannot assess the accurate preservation of relative abundances at different RNA levels. Linearity, the golden rule of nucleic acid isolation and detection, cannot be evaluated from a single data point. Rather, to robustly evaluate assay linearity, at least three different spike-ins need to be administered at different concentrations and the Pearson coefficient calculated.

[0007] To overcome the aforementioned drawbacks when using a single spike-in such as cel-miR-39-3p, the inventors designed and optimized a universal spike-in system of artificial small RNA molecules that broadly reflects the behavior of endogenous miRNAs during RNA extraction and detection. This mixture of artificial small RNA molecules is added to a biological sample such as a clinical sample at the start of the process and functions as an end-to-end quality control means. As a result of the analysis, only biological samples such as clinical samples in which the artificial small RNA molecules are recovered at the expected levels, order, and linearity are used for downstream analysis. Furthermore, the evaluation of the cocktail of artificial small RNA molecules can function as a bioinformatics-based normalization tool for removing batch effects in different assays.

[0008] Furthermore, the inventors have found that artificial small RNA molecules can be used to identify unwanted adducts to target RNA molecules. By excluding target RNA molecules having such adducts, it has been found that the quality of the RNA data is improved, and as a result, the target RNA analysis is improved. Artificial small RNA molecules are also referred to as spike-ins.

Summary of the Invention

Means for Solving the Problems

[0009] In a first aspect, the present invention relates to a composition comprising at least three RNA molecules, wherein the at least three RNA molecules are RNA molecules having nucleotide sequences according to SEQ ID NOs: 1 to 4, fragments thereof, and sequences having at least 80%, preferably at least 85%, more preferably at least 90%, still more preferably at least 95%, for example, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99% sequence identity, and are selected from the group consisting of.

[0010] In a second aspect, the present invention relates to a kit comprising the composition of the first aspect. In a third aspect, the present invention relates to the (standard) use of the composition of the first aspect or the kit of the second aspect for process control, sample testing, normalization and / or data processing control. Preferably, the data processing control includes raw data processing control or is metadata processing control. In a fourth aspect, the present invention relates to a method of inspecting a sample, comprising the step of evaluating the sample with respect to at least three RNA molecules contained in or derived from the composition of the first aspect. In a fifth aspect, the present invention relates to a method for optimizing a biological sample, comprising the step of performing the method of the fourth aspect. In a sixth aspect, the present invention relates to a method for optimized RNA preparation or RNA analysis from a biological sample, comprising the step of performing the method of the fourth aspect. In a seventh aspect, the present invention relates to a method for improving the quality of an (RNA) dataset, comprising the step of performing the method of the fourth aspect.

[0011] In an eighth aspect, the present invention relates to a method for improving the quality of an (RNA) dataset, (i) determining the sequences of RNA molecules in the sample, (ii) determining 5' and / or 3' end adducts to RNA molecules that are not part of the RNA molecules in their naturally occurring form, and, (iii) excluding RNA molecules having 5' and / or 3' end adducts from (RNA) dataset analysis / removing RNA molecules having 5' and / or 3' end adducts from the (RNA) dataset.

[0012] Preferably, the (RNA) dataset analysis is raw (RNA) data analysis or the (RNA) dataset is a raw (RNA) dataset. The summary of this invention does not necessarily describe all features of the present invention. Other embodiments will become apparent from consideration of the following detailed description.

Best Mode for Carrying Out the Invention

[0013] (Definition) Before the present invention is described in detail below, it should be understood that the present invention is not limited to the specific methodologies, protocols, and reagents described herein, and these may vary. It should also be understood that the terminology used herein is for the purpose of describing only particular embodiments and is not intended to limit the scope of the present invention, which will be limited only by the appended claims. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.

[0014] Preferably, the terminology used herein is defined as described in the "Multilingual Glossary of Biotechnology Terms: IUPAC Recommendations" (edited by Leuenberger, H.G.W, Nagel, B. and Kolbl, H., 1995, Helvetica Chimica Acta, Basel CH-4010, Switzerland).

[0015] Throughout the text of this specification, a number of documents are cited. Each document cited in this document (including all patents, patent applications, scientific papers, manufacturer's specifications, instructions, GenBank Accession Number sequence submissions, etc.) is incorporated by reference in its entirety, regardless of whether it is above or below this document. The content of this document should in no case be construed as an admission that the invention has no right to antedate such disclosure by virtue of a prior invention. In case of any conflict between the definitions or teachings of such reference documents and the definitions or teachings described herein, the description herein shall prevail.

[0016] As used herein, the term "comprise", or variations such as "comprises" or "comprising", means including a particular integer or group of integers and not excluding other integers or groups of integers. The term "consisting essentially of" means including a particular integer or group of integers while excluding modifications or other integers that have a substantial effect or change on that particular integer. The term "consisting of", or variations such as "consists of", means including a particular integer or group of integers while excluding other integers or groups of integers.

[0017] The terms "a", "an", "the", and similar reference terms used herein (especially in the claims) are to be construed as covering both the singular and the plural unless otherwise specified herein or the context clearly indicates otherwise. As used herein, the term "about" indicates a certain variation from the quantitative value immediately preceding it. In particular, the term "about" allows a variation of ±5% from the preceding quantitative value unless otherwise indicated or inferred. The use of the term "about" includes the particular quantitative value itself unless otherwise clearly indicated. For example, the expression "about 80°C" allows a variation of ±4°C and refers to the range of 76°C to 84°C.

[0018] The similarity of nucleotide and amino acid sequences, i.e., the ratio of sequence identity, can be determined by sequence alignment. Such alignments are performed using several well-known algorithms, preferably the mathematical algorithm of Karlin and Altschul (Karlin & Altschul (1993) Proc. Natl. Acad. Sci. USA 90:5873-5877), hmmalign (HMMER package) or the CLUSTAL algorithm (Thompson, J.D., Higgins, D.G. & Gibson, T.J. (1994) Nucleic Acids Res. 22, 4673-80) or the CLUSTALW2 algorithm (Larkin MA, Blackshields G, Brown NP, Chenna R, McGettigan PA, McWilliam H, Valentin F, Wallace IM, Wilm A, Lopez R, Thompson JD, Gibson TJ, Higgins DG. (2007). Clustal W and Clustal X version 2.0. Bioinformatics, 23, 2947-2948).

[0019] The degree of sequence identity (sequence matching) can be calculated using, for example, BLAST, BLAT or BlastZ (or BlastX). Similar algorithms are incorporated in the BLASTN and BLASTP programs of Altschul et al. (1990) J. Mol. Biol. 215:403-410. The BLAST protein search is performed using, for example, the BLASTP program available at the website: http: / / blast.ncbi.nlm.nih.gov / Blast.cgi?PROGRAM=blastp&BLAST_PROGRAMS=blastp&PAGE_TYPE=BlastSearch&SHOW_DEFAULTS=on&LINK_LOC=blasthome.

[0020] The parameters of the preferred algorithm used are the default parameters set at the specified website. Expected value threshold = 10, word size = 3, maximum number of matches within query range = 0, matrix = BLOSUM62, gap cost = existence: 11 extension: 1, composition adjustment = conditional composition score matrix adjustment with the database of non-redundant protein sequences (nr).

[0021] To obtain a gapped alignment for comparison purposes, gapped BLAST is utilized as described in Altschul et al. (1997) Nucleic Acids Res. 25:3389 - 3402. When using the BLAST and gapped BLAST programs, the default parameters of each program are used. The sequence matching analysis may be supplemented by established homology mapping techniques such as Shuffle-LAGAN (Brudno M., Bioinformatics 2003b, 19 Suppl 1:I54 - I62) or Markov random fields.

[0022] As used herein, the term "nucleotide" refers to an organic molecule composed of a nucleoside and a phosphate. In particular, a nucleotide is composed of three subunit molecules: a nucleobase, a five-carbon sugar (ribose or deoxyribose), and a phosphate group consisting of 1 - 3 phosphates. The four nucleobases of DNA are guanine, adenine, cytosine, and thymine, and in RNA, uracil is used instead of thymine. Nucleotides function as monomeric units of nucleic acid polymers such as deoxyribonucleic acid (DNA) and ribonucleic acid (RNA). Thus, nucleotides are the molecular building blocks of DNA and RNA.

[0023] The terms "nucleotide sequence" or "polynucleotide" are used interchangeably herein and refer to single-stranded and double-stranded polymers of nucleotide monomers, 2'-deoxyribonucleotides (DNA) and ribonucleotides (RNA) linked by phosphodiester bonds between nucleotides, or analogs between nucleotides, and related counterions (e.g., H + , NH4 +, trialkylammonium, Mg 2+ , Na + etc.) are also included, but not limited thereto. The nucleotide sequence or polynucleotide may be composed entirely of deoxyribonucleotides, entirely of ribonucleotides, or of a chimeric mixture thereof, and may include nucleotide analogs. The nucleotide monomer units can include any of the nucleotides described herein, including nucleotides and / or nucleotide analogs.

[0024] As used herein, the term "nucleic acid molecule" refers to single-stranded and double-stranded polymers of nucleotide monomers, 2'-deoxyribonucleotides (DNA) and ribonucleotides (RNA) joined by phosphodiester bonds between nucleotides, or analogs between nucleotides, and related counterions, e.g., H + , NH4 + , trialkylammonium, Mg 2+ , Na + etc. are included, but not limited thereto. The nucleic acid molecule may be composed of only deoxyribonucleotides, only ribonucleotides, or a chimeric mixture thereof, and may include nucleotide analogs.

[0025] As used herein, the term "RNA molecule" refers to a polymeric form of ribonucleotides of any length. Similar to DNA, RNA is assembled as a strand of nucleotides, but unlike DNA, RNA exists in nature as a single strand that folds back on itself rather than as a paired double strand.

[0026] Cell organisms use messenger (mRNA) that conveys genetic information (using the nitrogen bases guanine, uracil, adenine, and cytosine represented by the letters G, U, A, C) to direct the synthesis of specific proteins. Some RNA molecules play an active role by catalyzing biological reactions within cells, controlling gene expression, or sensing and transmitting responses to cell signals. One such active process is protein synthesis, which is a universal function where RNA molecules direct the synthesis of proteins on ribosomes. In this process, transfer RNA (tRNA) molecules are used to carry amino acids to the ribosome, and ribosomal RNA (rRNA) binds the amino acids to form the encoded protein.

[0027] In the context of the present invention, the RNA molecule is an artificial RNA molecule. Thus, it does not occur naturally. It is foreign to the cell. It is synthetically designed / produced. The RNA molecule is included in a composition that can be used as a general-purpose additive system or cocktail.

[0028] The inventors designed and optimized RNA molecules that widely reflect the behavior of endogenous miRNAs during RNA extraction and detection. Specifically, the inventors generated 20 types of random sequences with a length of 21 nucleotides (reflecting the typical length of miRNAs). Next, molecular properties such as melting temperature (Tm °C) and GC% were evaluated. Thereafter, based on several criteria, a shortlist of 10 artificial sequences for wet-lab verification was selected. Specifically, the aim was to minimize primer dimers and select RNAs with relatively weak secondary structures. They selected RNA molecules as spike-ins with a GC content of 38.1% to 61.9%, covering most of the GC content of endogenous miRNAs. Finally, the inventors selected four artificial RNA molecules with nucleotide sequences according to SEQ ID NO: 1 to SEQ ID NO: 4 as components of a composition that can be used as a universal spike-in system or cocktail. To efficiently monitor the inefficiency of the test, at least three of these four RNA molecules must be present in the composition. For example, to robustly evaluate the linearity of the test, at least three different RNA molecules (spike-ins) administered at different concentrations are required for statistical analysis, such as calculating the Pearson coefficient.

[0029] The RNA molecule has a secondary structure that minimizes primer dimer formation, a G / C content of 38.1 to 61.9%, thus encompassing most of the endogenous small RNAs such as miRNAs, and is characterized by having a 5'-phosphate group to reflect endogenous mature small RNAs such as miRNAs.

[0030] As used herein, the term "small RNA molecule" refers to a polymeric RNA molecule having a length of less than 200 ribonucleotides, preferably less than 50 ribonucleotides. Specifically, the small RNA molecule has a length of 10 to less than 200 ribonucleotides. More specifically, the small RNA molecule has a length of 10 to less than 50 ribonucleotides. Small RNA molecules are typically non-coding RNA molecules. RNA molecules having nucleotide sequences according to SEQ ID NOs: 1 to 4 are small RNA molecules.

[0031] As used herein, the term "composition" refers to a composition comprising at least three RNA molecules, at least three of which are selected from the group consisting of RNA molecules having nucleotide sequences according to SEQ ID NOs: 1 to 4. The RNA molecules contained in the composition are artificial exogenous small RNA molecules. They do not exist naturally. Furthermore, they widely reflect the behavior of endogenous miRNAs during RNA extraction and detection.

[0032] In one embodiment, at least three RNA molecules contained in the composition have a characteristic distribution, particularly with respect to their amounts. In particular, at least three RNA molecules are contained in the composition in different amounts. More specifically, at least three RNA molecules are contained in the composition in a defined gradient of amounts. Specifically, at least three RNA molecules are titrated in a linear range of different amounts.

[0033] The composition can be used as a spike-in cocktail or can be regarded as a spike-in cocktail. The spike-in cocktail is a general-purpose spike-in cocktail that is independent of downstream detection methods.

[0034] As used herein, the term "spike-in" means adding (spiking) a known amount / quantity of RNA molecules to a sample. The sample may contain biological material. Also, the sample may be a processed sample based on or derived from biological material. Next, a method for measuring the reaction (recovery) of the spiked sample is implemented. This method can be implemented in any processing step to determine the sample throughput or processing quality and / or the analysis of the sample. For example, this method can be performed after the lysis of biological material, RNA extraction, amplification of RNA or DNA (DNA derived from RNA) and / or sequencing of DNA (DNA derived from RNA).

[0035] As used herein, the term "spike-in cocktail" refers to a composition composed of RNA molecules (also called RNA spike-ins) of known sequences and amounts / quantities. The RNA molecules contained in this composition are artificial exogenous small RNA molecules. These do not exist in nature. Furthermore, they widely reflect the behavior of endogenous miRNAs during RNA extraction and detection.

[0036] The spike-in cocktail according to the present invention is a composition containing at least three RNA molecules, and at least three RNA molecules are selected from the group consisting of RNA molecules having nucleotide sequences according to SEQ ID NO: 1 to SEQ ID NO: 4.

[0037] The RNA molecules are added to a sample such as a biological sample or a processed biological sample to access the performance of nucleic acid quantification tests, such as molecular biology tests like qPCR, next-generation sequencing (NGS) and / or microarray tests. As described above, at least three RNA molecules that make up the composition have a characteristic distribution, particularly with respect to their amounts. In particular, at least three RNA molecules are included in the composition in different amounts. More specifically, at least three RNA molecules are included in the composition in a defined gradient of amounts. Thus, any RNA molecule included in the composition is present in a specific and known amount that is different from the amounts of the other RNA molecules. Specifically, at least three RNA molecules are titrated into different linear ranges of amounts.

[0038] In the method of the invention for the examination of a sample, the sample is evaluated with respect to at least three RNA molecules comprised in / derived from the composition described herein, and the at least three RNA molecules are selected from the group consisting of RNA molecules having a nucleotide sequence according to SEQ ID NO: 1 to SEQ ID NO: 4. The at least three RNA molecules are part of the sample. They are added to the sample.

[0039] In this method, it is evaluated whether at least three RNA molecules (during and / or after the processing and analysis of the sample) still have this characteristic distribution. Specifically, it is evaluated whether at least three RNA molecules are present at the expected levels, in the expected order and / or with the expected linearity. In this way, the quality or quantity of the processing and / or analysis of the sample can be controlled.

[0040] In this regard, it will be understood that the term "RNA molecule" may refer not only to the RNA molecule itself, but also to its alternatives, such as amplification products (e.g., cDNA derived therefrom).

[0041] As used herein, the term "expected level" means the level expected as a particular RNA molecular weight that is added to a sample after processing and / or analysis of the sample or that is included in a composition. As described above, the RNA molecule is added to the sample or is included in the composition in a specific amount. The amount of the RNA molecule may correspond / correlate to any expected level of the RNA molecule that can be measured during subsequent processing and / or analysis of the sample. Specifically, the amount of the RNA molecule may correspond / correlate to a specific number of reads (read count) (RPM) in a next-generation sequencing assay, and / or the amount of the RNA molecule may correspond / correlate to a specific cycle threshold (Ct) in a real-time PCR test. In particular, the cycle threshold (Ct) is the number of cycles in real-time PCR required to exceed a predefined threshold in the measured signal (e.g., fluorescence signal) of the amplified DNA. The more DNA (RNA) already present in the sample solution before PCR, the fewer the number of amplification cycles required to reach the corresponding threshold.

[0042] For example, if an RNA molecule having a nucleotide sequence according to SEQ ID NO: 1 is added to a sample in an amount of about 3400 amol or is included in the composition of the present invention, the predicted Ct value will be about 26 (the average Ct measured value is about 24.3), and if an RNA molecule having a nucleotide sequence according to SEQ ID NO: 2 is added to the sample in an amount of about 725 amol or is included in the composition of the present invention, the predicted Ct value will be about 29.3 (the average Ct measured value is about 27.1), and if an RNA molecule having a nucleotide sequence according to SEQ ID NO: 3 is added to the sample in an amount of about 20 amol or is included in the composition of the present invention, the predicted Ct value will be about 35.9 (the average Ct measured value is about 32), and / or if an RNA molecule having a nucleotide sequence according to SEQ ID NO: 4 is added to the sample in an amount of about 7 amol or is included in the composition of the present invention, the predicted Ct value will be about 39.2 (the average Ct measured value is about 33.6).

[0043] For example, when an RNA molecule having a nucleotide sequence according to SEQ ID NO: 1 is added to a sample in an amount of about 3400 amol or is included in the composition of the present invention, it results in a log2RPM of about 11.972 or a read count of about 30628; when an RNA molecule having a nucleotide sequence according to SEQ ID NO: 2 is added to a sample in an amount of about 725 amol or is included in the composition of the present invention, it results in a log2RPM of about 10.926 or a read count of about 5554; when an RNA molecule having a nucleotide sequence according to SEQ ID NO: 3 is added to a sample in an amount of about 20 amol or is included in the composition of the present invention, it results in a log2PRM of about 6.559 or a read count of about 491; and / or when an RNA molecule having a nucleotide sequence according to SEQ ID NO: 4 is added to a sample in an amount of about 7 amol or is included in the composition of the present invention, it results in a log2RPM of about 4.061 or a read count of about 94.

[0044] As used herein, the term "level" refers to the amount (measured in grams, moles, ion counts, etc.) or concentration (e.g., absolute or relative concentration, such as reads per million (RPM) or NGS count) of an RNA molecule contained in a composition or sample. The term "level" as used herein also includes scaled, normalized, or scaled and normalized amounts or values (e.g., RPM). In particular, the level of an RNA molecule is determined by sequencing, preferably next-generation sequencing (e.g., ABI SOLID, Illumina Genome Analyzer, Roche 454 GS FL, BGISEQ), nucleic acid hybridization (e.g., microarray or beads), nucleic acid amplification (e.g., PCR, RT-PCR, qRT-PCR, or high-throughput RT-PCR), polymerase extension, mass spectrometry, flow cytometry (e.g., LUMINEX), or any combination thereof. Specifically, the level of an RNA molecule is the expression level of the said RNA molecule.

[0045] As used herein, the term "expected order" means the order / rank expected for a particular RNA molecule added to a sample or included in a composition in subsequent sample processing and / or analysis. As described above, the RNA molecules added to the sample or included in the composition are in different amounts. Thus, if there are three RNA molecules added to the sample or included in the composition, the first RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 1) is present in the greatest amount, the second RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 2) is present in an amount less than the first RNA molecule, and the third RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 3) is present in an amount less than the first and second RNA molecules. If four RNA molecules are added to the sample or included in the composition, the first RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 1) is present in the greatest amount, the second RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 2) is present in an amount less than the first RNA molecule, the third RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 3) is present in an amount less than the first and second RNA molecules, and the fourth RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 4) is present in an amount less than the first, second, and third RNA molecules.

[0046] This expected order / rank represents an initial state / condition and must be restored during sample processing and / or analysis. In this regard, it should be noted that the amount of RNA molecules included in a particular order / rank in a sample or composition may correspond / correlate to the expected level of RNA molecules measured during subsequent sample processing and / or analysis. For example, the amount of RNA molecules may correspond to a particular number of reads in a next-generation sequencing device and / or the amount of RNA molecules may correspond / correlate to a particular cycle threshold (Ct) in a real-time PCR assay.

[0047] As used herein, the term "expected linearity" means the linearity expected for a particular RNA molecule added to a sample or included in a composition in subsequent sample processing and / or analysis.

[0048] This expected linearity represents an initial state / condition and must be restored during sample processing and / or analysis. As a result of the analysis, only samples in which the RNA molecule or its surrogate is recovered at the expected levels, order, and / or linearity are used for subsequent analysis and / or further processing.

[0049] To determine whether an RNA molecule is present in the expected order and / or linearity, statistical analysis is performed using data related to the RNA molecule (e.g., Ct value or RPM value). The statistical analysis includes, but is not limited to, Spearman (rank) correlation analysis and / or Pearson correlation analysis.

[0050] For example, to determine whether an RNA molecule is present in the expected order, preferably the Spearman rank correlation coefficient (Spearman ρ) is determined. Furthermore, to determine whether an RNA molecule is present in the expected linearity, preferably the Pearson correlation coefficient (Pearson r) is determined.

[0051] As used herein, the term "Spearman's rank correlation coefficient or Spearman ρ" is a nonparametric measure of rank correlation (the statistical dependence between the ranks of two variables). This assesses how accurately the relationship between two variables can be described using a monotonic function.

[0052] The Spearman correlation between two variables is equal to the Pearson correlation between the ranks of those two variables. The Pearson correlation assesses linear relationships, while the Spearman correlation assesses monotonic relationships (regardless of whether they are linear). When data values are not repeated, the Spearman correlation will be a perfect value of +1 or -1 when each variable is a perfect monotonic function of the other.

[0053] Intuitively, the Spearman correlation between two variables is high when the observations between the two variables have similar (or identical in the case of a correlation coefficient of 1) ranks (i.e., the relative position labels of the observations within the variable: 1st, 2nd, 3rd, etc.), and the Spearman correlation between two variables is low when the observations between the two variables have different (or completely opposite in the case of a correlation coefficient of -1) ranks.

[0054] In the context of the present invention, when the Spearman rank correlation coefficient (Spearman ρ) is 0.95 or greater, the RNA molecules are present in the expected order. In this case, the sample is further processed / analyzed. Thus, when the Spearman rank correlation coefficient (Spearman's ρ) is less than 0.95, the sample is discarded. In other words, such a sample is not further processed / analyzed.

[0055] The term "Pearson correlation coefficient or Pearson's r" as used herein is a measure of the linear correlation between two data sets. This is the ratio of the covariance between two variables to the product of their standard deviations. Thus, essentially, this is a normalized measure of covariance, and the result will always be a value between -1 and 1. Similar to covariance itself, this measure only reflects the linear correlation of variables and ignores many other types of relationships or correlations.

[0056] In connection with the present invention, an RNA molecule exhibits an expected linearity when the Pearson correlation coefficient (Pearson r) is 0.66 or greater. In this case, the sample is further processed / analyzed. Thus, if the Pearson correlation coefficient (Pearson r) is less than 0.66, the sample is discarded. In other words, such a sample is not further processed / analyzed.

[0057] Furthermore, the evaluation of a cocktail of artificial small RNA molecules functions as a bioinformatics normalization tool for removing batch effects in different assays. As used herein, the term "normalization" refers to the techniques necessary to compare RNA levels between different samples. For example, to compare small RNA levels in different samples, normalization of high-throughput small RNA sequencing data is required. Commonly used relative normalization approaches can lead to false conclusions due to small RNA populations that vary between tissues or body fluids. The inventors have developed a composition of RNA molecules (also referred to as an RNA spike-in) that enables absolute normalization of small RNA data in independent assays. Data from small RNA sequencing assays are typically normalized and reported as relative values such as reads per million genomic-matching reads (RPM). Normalization by relative values works well when it is assumed that small RNA subpopulations are present in equal proportions in different tissues or body fluid types being profiled. However, this assumption often fails because small RNA populations frequently vary between different tissue or body fluid types and various mutational backgrounds. Thus, the standard approach of comparing relatively normalized small RNA sequencing values can result in misleading results. In contrast, absolute normalization of small RNA sequencing data enables accurate comparison of small RNA levels across different cell types, mutant tissues, or disease states throughout the genome. The inventors have designed, for example, RNA molecules (also referred to as RNA spike-ins) that can be used for robust absolute normalization, such as small RNA sequencing data across independent assays.

[0058] The RNA molecule is a secondary structure that minimizes the formation of primer dimers, a G / C content of 38.1 to 61.9% (thus encompassing most endogenous small RNAs such as miRNAs), a 5' phosphate group that reflects endogenous mature small RNAs such as miRNAs, and is characterized by

[0059] As used herein, the term "quality" with respect to a sample (including biological material) refers to the level of degradation of a component in the sample relative to the time point when the component was included in a biological system such as a cell. For example, assessing the quality of RNA and / or DNA can include assessing the level of partial degradation of the RNA and / or DNA polymer. In some embodiments, the assessment of quality includes, for example, the assessment of partial degradation of RNA and / or DNA from a spike-in standard by assessing the presence of fragments (short molecules) of the spike-in standard in relation to a full-length spike-in standard.

[0060] As used herein, the term "quantity" in relation to a sample (including biological material) refers to the level of a component in the sample relative to the time point when the component was included in a biological system such as a cell. In some embodiments, assessing the quantity of RNA and / or DNA can include assessing the level of degradation or loss of the RNA and / or DNA polymer in the sample.

[0061] As used herein, the term "sample" refers to a mixture of biological materials having the composition of the present invention. Alternatively, the sample is based on / derived from a mixture of biological materials having the composition of the present invention. In this case, the sample is often a processed sample. A processed sample is obtained by mixing biological materials having the composition of the present invention and further processing the resulting mixture. For example, a processed sample can be a lysed sample, an extracted sample, an amplified sample, a sequenced sample, or a library-prepared sample.

[0062] As used herein, the term "biological material" refers to any material having a biological origin. The biological material is preferably a tissue material or a body fluid material.

[0063] As used herein, the term "body fluid material" refers to any liquid material obtained from an individual's body. The body fluid material may be urine, blood, saliva, breast milk, cerebrospinal fluid (CSF), earwax, gastric juice, mucus, lymph fluid, endolymph, perilymph, peritoneal fluid, pleural fluid, saliva, sebum (skin oil), semen, sweat, tears, oral mucosal specimen, vaginal secretion, liquid biopsy, or a vomit sample containing components or fractions thereof. The term "body fluid material" also encompasses body fluid fractions, such as blood fractions like blood cells, serum, or plasma, where blood cells represent the cellular fraction of blood and serum and plasma represent the cell-free fractions of blood.

[0064] As used herein, the term "blood material" encompasses whole blood or blood fractions. Preferably, the blood fraction is selected from the group consisting of a blood cell fraction, plasma, and serum. For example, the blood cell fraction includes red blood cells, white blood cells, and / or platelets. More preferably, the blood cell fraction is a white blood cell fraction or a mixture of red blood cells, white blood cells, and platelets.

[0065] The blood material may be provided by collecting blood from an individual, or may be provided using a previously isolated material. For example, the blood material may be collected from an individual by conventional blood collection techniques.

[0066] Whole blood samples can be collected using blood collection tubes. For example, they may be collected in PAXgene Blood RNA tubes, Tempus Blood RNA tubes, EDTA tubes, Na-citrate tubes, heparin tubes, or ACD (Acid citrate dextrose) tubes.

[0067] As used herein, the term "blood material", particularly whole blood material, may also be collected by blood spot techniques using, for example, a Mitra Microsampling Device. This technique typically requires a small sample of 45 - 60 μl or less in the case of humans. For example, whole blood can be extracted from an individual by fingertip puncture with a needle or lancet. Thus, the whole blood sample may be in the form of a blood drop. This blood drop is placed on a probe that is absorbent and can absorb whole blood, such as a hydrophilic polymer material like cellulose. Once sampling is complete, the blood drop is dried in air before being transferred or mailed to a laboratory for processing. Since the blood is dried, it is not considered harmful. Thus, no special precautions are required during handling or transportation. Once it arrives at the analysis facility, the desired component, such as RNA, is extracted from the dried blood spot into the supernatant and further analyzed. In this way, the RNA level is determined.

[0068] As used herein, the term "tissue material" refers to any tissue material obtained from an individual's body. The material may be tumor / cancerous tissue material or healthy tissue material from any organ of the individual. For example, the tissue material may be material from the lung, kidney, liver, brain, colon, breast, stomach, uterus, ovary, pancreas, or prostate. The material can be collected from an individual by conventional biopsy techniques.

[0069] As used herein, the term "target RNA" refers to an endogenous ribonucleotide sequence contained in a biological material for which detection is sought. Target RNA can be obtained from any source and may be composed of any number of different components. For example, target RNA is isolated from an organism, tissue, cell, or body fluid such as blood. Desirably, target RNA includes non-coding RNA. In particular, target RNA is a microRNA (miRNA) or miRNA isoform (isomiR) and may include variants, analogs, and mimics. Target RNA may, in some cases, also be designated as a target RNA molecule.

[0070] Furthermore, it will be understood that the term "target RNA" may refer not only to the target molecule itself, but also to its alternatives, such as amplification products (e.g., cDNA derived therefrom) and native sequences. In certain embodiments, the target RNA is a miRNA or miRNA isoform (isomiR) molecule. In certain embodiments, the target RNA is a mature small RNA, particularly a non-coding small RNA molecule (i.e., less than 200 ribonucleotides in length, e.g., less than 10 to 200 ribonucleotides in length). The target RNAs described herein can be obtained from any number of sources, including but not limited to humans and animals. These sources include, but are not limited to, whole blood, tissue biopsies, lymph, bone marrow, amniotic fluid, hair, skin, semen, biological warfare agents, anal secretions, vaginal secretions, sweat, saliva, oral mucosa specimens, etc. However, various environmental samples (e.g., agricultural, water, soil), generally research samples, generally purified samples, cultured cells, and lysed cells may also be used as samples. It will be understood that target RNA may be isolated from a sample using any of a variety of known procedures, such as, for example, Applied Biosystems ABI Prism™ 6100 Nucleic Acid PrepStation (Life Technologies, Foster City, CA) and ABI Prism™ 6700 Automated Nucleic Acid Workstation (Life Technologies, Foster City, CA), Ambion™ mirVana™ RNA isolation kit (Life Technologies, Austin, TX), PAXgene Blood RNA Kit (Qiagen, Hilden, Germany), etc., known to those of skill in the art.

[0071] Target RNA molecules having an origin inherent in biological material are different from the exogenous artificial RNA molecules also described herein. Exogenous artificial RNA molecules are added to biological material as part of a composition such as a spike-in cocktail, while target RNA molecules are part of the biological material.

[0072] As used herein, the term "miRNA" (which may also be denoted "microRNA") refers to a single-stranded RNA molecule. miRNAs are molecules that are 10 to 50 nucleotides in length, for example, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides in length and do not contain any labels and / or extension sequences (such as an extension of biotin).

[0073] miRNAs control gene expression and are encoded by genes transcribed from their DNA, but miRNAs are not translated into proteins (i.e., miRNAs are non-coding RNAs). Genes encoding miRNAs are longer than the processed mature miRNA molecules. miRNAs are first transcribed as longer precursor molecules (longer than 1000 nucleotides) called primary miRNA transcripts (pri-miRNAs). Pri-miRNAs have a hairpin structure and are processed by the Drosha enzyme (as part of the microprocessor complex). After processing by Drosha, the pri-miRNA becomes 60 - 100 nucleotides in length and is called precursor miRNA (pre-miRNA). At this point, the pre-miRNA is transported to the cytoplasm where it encounters the Dicer enzyme. Dicer cleaves the miRNA in two, generating a double-stranded miRNA strand. Conventionally, it was thought to be important in gene regulation to be incorporated into the RNA-induced silencing complex (RISC) and only one of these miRNA arms, which occur at higher concentrations intracellularly, was involved. This is called the "guide" strand and is designated miR. The other arm is called the "minor miRNA" or "passenger miRNA" and is often designated miR*. The passenger miRNA was thought to be completely degraded, but deep sequencing studies have revealed that some minor miRNAs persist and actually play a functional role in gene regulation. These developments have led to a change in the naming rules. Instead of the miR / miR* nomenclature, the miR-5p / miR-3p nomenclature has been adopted. In this new system, the 5’ arm of the miRNA is always designated miR-5p and the 3’ arm is designated miR-3p. The current nomenclature is as follows. A dash and a number follow the prefix "miR", and the latter often indicates the order of naming. For example, hsa-miR-16 was named before hsa-miR-342 and was probably discovered first.The capitalized "miR-" refers to mature miRNAs (e.g., hsa-miR-16-5p and hsa-miR-16-3p), the lower-case "mir-" refers to precursor miRNAs and precursor pri-miRNAs (e.g., hsa-mir-16), and "MIR" refers to the genes encoding them. However, since this is a recent change, the conventional miR / miR* nomenclature is still often used in the literature. After processing, the double-stranded miRNA strands bind to Argonaute (AGO) proteins to form a precursor of RISC. This complex unwinds the double strand, the passenger RNA strand is discarded, and the mature RISC carrying the mature single-stranded miRNA remains. Since miRNA suppresses the expression of target genes, it remains as part of RISC. This is the standard pathway in miRNA production, but various other pathways have also been discovered. These include the Drosha-independent pathways (mirtron pathway, snoRNA-derived pathway, shRNA-derived pathway, etc.) and the Dicer-independent pathways (pathways dependent on cleavage by AGO, other pathways dependent on tRNaseZ, etc.).

[0074] As used herein, the term "miRBase" refers to an established repository of verified miRNAs. miRBase (www.mirbase.org) is a searchable database of published miRNA sequences and annotations. Each entry in the miRBase sequence database represents the predicted hairpin portion of the miRNA transcript (referred to as mir in the database) and includes information on the location and sequence of the mature miRNA sequence (referred to as miR). The hairpin and mature sequences are searchable and browsable, and entries can also be searched by name, keyword, reference, and annotation. All sequence and annotation data can also be downloaded. In October 2018, miRbase version 22.1 was released. This is the current version.

[0075] As used herein, the term "isomiR" (or "miRNA isoform") refers to miRNAs with slightly different sequences due to changes in the cleavage site during miRNA biogenesis. In particular, incomplete cleavage by Drosha and Dicer or inversion of miRNAs may result in miRNAs with heterogeneous lengths and / or sequences. IsomiRs (miRNA isoforms) are mainly divided into three categories: 3′ isomiRs (one or more nucleotides are trimmed or added at the 3′ end), 5′ isomiRs (one or more nucleotides are trimmed or added at the 5′ end), and polymorphic isomiRs (some nucleotides within the sequence are different from the wild-type mature miRNA sequence). Increased expression of miRNA variants, or individual isomiRs, may lead to loss or reduction of function of the corresponding wild-type mature miRNA, or may result in the control of different transcriptomes. Recent studies have suggested that isomiRs probably play important roles in various cancers, tissues, and cell types. Therefore, the detection of miRNAs and isomiRs is absolutely necessary to accurately reflect the biological situation and make correct diagnostic and therapeutic decisions.

[0076] RNA can be isolated from samples using any of a variety of procedures well known in the art, such as the Applied Biosystems ABI Prism™ 6100 Nucleic Acid PrepStation (Life Technologies, Foster City, CA) and the ABI Prism™ 6700 Automated Nucleic Acid Workstation (Life Technologies, Foster City, CA), the Ambion® mirVana® RNA Isolation Kit (Life Technologies, Austin, TX), and the PAXgene Blood RNA Kit (Qiagen, Hilden, Germany).

[0077] The detection of RNA molecules requires the presence of a larger amount of RNA molecules. Through the process of reverse transcription, cDNA molecules are generated from RNA molecules. In various embodiments, these cDNA molecules are amplified. RNA molecules can also be directly amplified. As used herein, the term "amplification" refers to any means by which at least a portion of the nucleic acid molecules described herein are typically replicated in a template-dependent manner, including, but not limited to, a wide range of techniques for linearly or exponentially amplifying nucleic acid sequences. To amplify nucleic acid molecules, any of several methods can be used. Any in vitro means can be utilized to proliferate copies of the target sequence of the nucleic acid. These include linear, logarithmic, or other amplification methods.

[0078] Examples of amplifying techniques that can be used include, but are not limited to, PCR, quantitative PCR, quantitative fluorescence PCR (QF-PCR), multiplex fluorescence PCR (MF-PCR), real-time PCR (RT-PCR), single-cell PCR, restriction fragment length polymorphism PCR (PCR-RFLP), hot-start PCR, nested PCR, in situ colony PCR, in situ rolling circle amplification (RCA), bridge PCR, picotiter PCR, emulsion PCR, etc. Other suitable amplification methods include ligase chain reaction (LCR), transcription amplification, self-sustained sequence replication, selective amplification of target polynucleotide sequences, polymerase chain reaction using consensus sequence primers (CP-PCR), polymerase chain reaction using arbitrary primers (AP-PCR), PCR using degenerate oligonucleotide primers (DOP-PCR), nucleic acid-based sequence amplification (NABSA), etc.

[0079] In various embodiments, DNA molecules derived from RNA molecules are sequenced. As used herein, the term "sequencing" includes any method for determining the sequence of a nucleic acid molecule. Such methods include Maxam-Gilbert sequencing, chain termination methods, shotgun sequencing, PCR sequencing, bridge PCR, massively parallel sequencing (MPSS), Polony sequencing, pyrosequencing, Illumina (Solexa) sequencing, SOLiD sequencing, ion semiconductor sequencing, DNA nanoball sequencing, Heliscope single molecule sequencing, single molecule real time (SMRT) sequencing, nanopore DNA sequencing, sequencing by hybridization, sequencing using mass spectrometry, microfluidic Sanger sequencing, microscopy-based techniques, RNAP sequencing, (in vitro virus) high throughput sequencing (HTS).

[0080] As used herein, the term "next generation sequencing (NGS)" refers to new methods for rapidly and inexpensively sequencing nucleotide sequences. Next generation sequencing (NGS) is a high throughput method for rapidly sequencing the base pairs of a DNA or RNA sample. NGS, which supports a wide range of applications such as gene expression profiling, chromosome counting, detection of epigenetic changes, and molecular analysis, drives discovery and enables the future of personalized medicine. NGS is also referred to as second generation sequencing (SGS) or massively parallel sequencing (MPS).

[0081] As used herein, the term "kit of parts (hereinafter, kit)" refers to any combination of at least some of the components specified herein, which are combined into a functional unit, coexist spatially, and can further include components.

[0082] As used herein, the term "data processing control" refers to the initial stage of data analysis in which part of the data is deleted or its size is reduced in order to improve aspects of later analysis, such as improving the quality of the data and the reliability of the findings during analysis. Preferably, data processing control is raw data processing control.

[0083] As used herein, the term "raw data" refers to data collected during an inspection. These data have (not yet) been further processed. Raw data can be quantitative (numerical) and / or qualitative (descriptive). The size (generally large), file type, and structure of raw data vary depending on the technology used to create the data. Before obtaining ultimately analyzable data (e.g., detecting statistical differences between groups of biological samples with measured biological conditions / biological mappings, or patients with different disease states), the raw data needs to be processed using dedicated software.

[0084] In the process control of next-generation sequencing, the raw data generated after sequencing must be processed, and as a result, false positive and false negative results must be avoided. Therefore, for data analysis, data quality control and preprocessing are important. Data preprocessing not only evaluates each analysis step but also reduces the amount of low-quality sequencing reads. By removing such low-quality reads, the time and cost of computational analysis can be reduced, and highly reliable and high-quality results can be obtained.

[0085] In any downstream analysis of NGS data, false positive and false negative results are generated by the following. Inspection factors: such as sample contamination and PCR errors. Sequencing factors: This includes the quality of sequencing and data contamination caused by index hopping when dividing data. Parameter factors of analysis software: This includes alignment software and the exact type of parameter adjustment in downstream individual analyses.

[0086] The raw sequences generated after array determination include not only sequences of interest such as target RNA molecules, but also sequence biases (such as those due to systematic effects like Poisson sampling) and complex artifacts generated by the sequencing and inspection steps. These sequence biases and artifacts affect and interfere with accurate read alignment, thus affecting gene typing and variant calling. Therefore, preprocessing of raw sequence reads is necessary to enhance the reliability and quality of subsequent analyses and reduce the amount of computational resources required.

[0087] The inventors have discovered that 5' and / or 3' end adducts to (target) RNA molecules generated or occurring during a sequencing process, such as, for example, a next-generation sequencing process, can lead to misinterpretations and corrupt the results of the raw data. Specifically, it is necessary to exclude / remove adapter-contaminated sequence reads. The inventors have discovered that the quality of the raw data can be improved by excluding / removing adapter-contaminated sequence reads.

[0088] As used herein, the term "5' end adduct (also referred to as a prefix)" refers to nucleotides attached to the 5' end of an RNA molecule or a molecule derived therefrom, such as a cDNA molecule. These are in particular 5' sequencing adapters or remnants of other RNAs that are generated / present in the same physical mixture and have (in whole or in part) fused during the technical process of sequencing. In the case of sequencing adapters, the intended fusion may not have been followed by the intended cleavage for removing the adapter after completion of sequencing.

[0089] As used herein, the term "3'-end addition (also referred to as a suffix)" refers to nucleotides attached to the 3'-end of an RNA molecule or a molecule derived therefrom, such as a cDNA molecule. These are in particular remnants of other RNAs that are generated / present in the same physical mixture and that have (in whole or in part) fused during the technical process of sequencing. In this regard, it will be understood that the term "RNA molecule" may refer not only to the RNA molecule itself, but also to its alternatives, such as amplification products (e.g., cDNA derived therefrom).

[0090] Specifically, during next-generation sequencing (NGS), it has been confirmed that parts of other RNAs or adapter sequences are attached to one or both ends (5'-end and / or 3'-end) of a (target) RNA molecule. These parts are herein referred to as affixes and include both a prefix (5') and a suffix (3').

[0091] Preferably, the 5'-end and / or 3'-end additions have a length of 5 to 30 nucleotides, more preferably 5 to 20 nucleotides, even more preferably 7 to 15 nucleotides, such as 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides.

[0092] Specifically, the 5'-end and / or 3'-end additions are - for example, while using / using in combination a double-stranded RNA ligase such as T4 RNA ligase 2 (Rnl2) or Kod1 ligase to bind a 5'-adapter and / or a 3'-adapter to a (denatured) (target) RNA molecule, - for example, while using a reverse transcriptase (RT) such as MaximaH-RT or Tth polymerase to reverse transcribe a (target) RNA molecule to which a 5'-adapter and / or a 3'-adapter is bound (also referred to as a ligase product) into a cDNA molecule, - for example, while amplifying a (target) cDNA molecule by polymerase chain reaction (PCR), and / or - It is carried out / occurs during sequencing such as next-generation sequencing of the (said) cDNA molecule.

[0093] PCR may be, for example, real-time PCR (quantitative PCR or qPCR) such as TaqMan qPCR, multiplex PCR, nested PCR, high-fidelity PCR, fast PCR, hot-start PCR, or GC-rich PCR.

[0094] More specifically, the addition of the 5'-end and / or 3'-end adducts is carried out / occurs during the library preparation process for next-generation sequencing or during the next-generation sequencing process. Even more specifically, The process of next-generation sequencing preferably - Ligation of a 5'-adapter and / or 3'-adapter to a (denatured) (target) RNA molecule using, for example, a double-stranded RNA ligase such as T4 RNA ligase 2 (Rnl2) or Kod1 ligase, - Reverse transcription of a (target) RNA molecule (also called a ligase product) to which a 5'-adapter and / or 3'-adapter is ligated into a cDNA molecule using a reverse transcriptase (RT) such as Maxima H-RT or Tth polymerase, - Amplification of the (said) cDNA molecule using polymerase chain reaction (PCR), and / or - Sequencing of the (said) cDNA molecule is included.

[0095] In this specification, the terms "5'-end and / or 3'-end adduct" or "5'-end and / or 3'-end attachment" can be used interchangeably. As used herein, the term "adapter" refers to a non-biological RNA / DNA sequence that is intentionally added to the 5' or 3' end of a biological (target) RNA / cDNA molecule (derived from the sample to be sequenced) as part of the design of sequencing, such as NGS methods. This includes free adapter molecules that may fuse with biological RNA in an unintended manner. An adapter may be a combination / fusion of multiple individual adapters and index RNAs. Index RNAs are used as UMIs (unique molecular identifiers) to assign RNA molecules to their source samples during multiplexed sequencing, i.e., sequencing of multiple samples simultaneously. The UMI sequence is a short non-specific (random) sequence of a predetermined length, such as 12 nucleotides, but may vary in length and is incorporated among other adapter sequences.

[0096] As used herein, the term "RNA molecule in its native form" refers to an RNA molecule in the form in which it exists in nature or in a natural environment, such as a body fluid or tissue, e.g., whole blood. However, these RNA molecules may be further processed, e.g., processed into cDNA molecules. In this case, they are of natural origin.

[0097] As used herein, the term "RNA molecule in its naturally occurring form" further refers to an RNA molecule having an endogenous origin in a biological material / sample. The RNA molecule is part of the biological material / sample and may exist in its native functional form, or may be a fragment of a longer functional RNA sequence that has been degraded or processed as part of a specific biological process or by non-specific physical / chemical forces over time.

[0098] (Embodiments of the Invention) The present invention will be further described below. In the following description, different aspects of the present invention are defined in more detail. Each aspect thus defined may be combined with any one or more other aspects, unless the contrary is clearly indicated. In particular, any feature indicated as being preferred or advantageous may be combined with any one or more other features indicated as being preferred or advantageous, unless the contrary is clearly indicated.

[0099] Academic and commercial laboratories have started using exogenous spike-in sequences that can mimic endogenous target RNA populations, enabling the monitoring of assay inefficiencies. The most commonly used spike-in in small RNA analysis is cel-miR-39-3p from the nematode C. elegans. However, the use of a single spike-in, such as cel-miR-39-3p, has several drawbacks. First, a single spike-in does not fully reflect the considerable primary sequence diversity of the entire small RNA. The effects specific to the primary sequence, such as RNA secondary structure, free energy, and GC content, can have a profound impact on extraction and detection efficiency, and a single spike-in is insufficient to faithfully reflect this. Furthermore, a single spike-in cannot assess the true conservation status of the relative abundances of different RNA levels. The linearity, which is the golden rule in nucleic acid separation and detection, cannot be evaluated from a single data point. Rather, to reliably assess the linearity of the assay, at least three different spike-ins need to be added at different concentrations and the Pearson coefficient calculated.

[0100] To overcome the aforementioned drawbacks when using a single spike-in such as cel-miR-39-3p, the inventors designed and optimized a universal spike-in system of artificial small RNA molecules that broadly reflects the behavior of endogenous miRNAs during RNA extraction and detection. This mixture of artificial small RNA molecules is added to a sample consisting of a biological material such as a clinical material at the start of the process and functions as an end-to-end quality control means. As a result of the analysis, only samples in which the artificial small RNA molecules were recovered at the expected levels, order, and linearity are used for subsequent analysis. Furthermore, the evaluation of the cocktail of artificial small RNA molecules can function as a bioinformatics normalization tool for removing batch effects in different assays.

[0101] Furthermore, the inventors have found that artificial small RNA molecules can be used to identify unwanted attachments to target RNA molecules. By excluding target RNA molecules having such attachments, it has been found that the quality of the RNA data is improved, and as a result, the target RNA analysis is improved.

[0102] Artificial small RNA molecules are also referred to as spike-ins. Accordingly, in a first aspect, the present invention relates to a composition comprising at least three RNA molecules (e.g., three or four RNA molecules), wherein the at least three RNA molecules (e.g., three or four RNA molecules) are selected from the group consisting of nucleotide sequences of SEQ ID NOs: 1 to 4, fragments thereof, and sequences having at least 80%, preferably at least 85%, more preferably at least 90%, still more preferably at least 95%, e.g., 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% sequence identity thereto.

[0103] Specifically, the RNA molecules contained in the composition are (i) having a nucleotide sequence according to SEQ ID NO: 1, 2, 3, or 4, (ii) A nucleotide sequence that is a fragment of the nucleotide sequence according to (i), preferably 1 to 12 nucleotides, more preferably 1 to 8 nucleotides, most preferably 1 to 5 or 1 to 3 nucleotides, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides, a nucleotide sequence shorter than the nucleotide sequence according to (i), or (iii) A nucleotide sequence having at least 80%, preferably at least 85%, more preferably at least 90%, most preferably at least 95% or 99%, for example, at least 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99% sequence identity to the nucleotide sequence according to (i) or the nucleotide sequence fragment according to (ii).

[0104] The RNA molecule of SEQ ID NO: 1 has the following nucleotide sequence: GAUAGAUACGCCAGUACCGCC, The RNA molecule of SEQ ID NO: 2 has the following nucleotide sequence: AACGAAGCUCCACGAUGUAGG, The RNA molecule of SEQ ID NO: 3 has the following nucleotide sequence: UGUACGGAAAUAUUGGCUACC, and The RNA molecule of SEQ ID NO: 4 has the following nucleotide sequence: UUCAUACGUUGCCCAAUCCAG.

[0105] The RNA molecules contained in this composition are artificial exogenous small RNA molecules. These do not exist in nature. They are exogenous to the cell. Furthermore, they widely reflect the behavior of endogenous miRNAs during RNA extraction and detection. Furthermore, the RNA molecules contained in this composition Have a secondary structure that minimizes the formation of primer dimers, A G / C content of 38.1 to 61.9%, thus encompassing most of the endogenous small RNAs such as miRNAs, Have a 5' phosphate group that reflects endogenous mature small RNAs such as miRNAs.

[0106] Preferably, this composition contains four RNA molecules, and these four RNA molecules are RNA molecules having nucleotide sequences according to SEQ ID NO: 1 to SEQ ID NO: 4, fragments thereof, and those having at least 80%, preferably at least 85%, more preferably at least 90%, still more preferably at least 95%, for example 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99% sequence identity, and are selected from the group consisting of sequences.

[0107] Therefore, this composition (i) SEQ ID NO: 1, SEQ ID NO: 2 and SEQ ID NO: 3, (ii) SEQ ID NO: 2, SEQ ID NO: 3 and SEQ ID NO: 4, (iii) SEQ ID NO: 1, SEQ ID NO: 3 and SEQ ID NO: 4, (iv) SEQ ID NO: 1, SEQ ID NO: 2 and SEQ ID NO: 4, or (v) SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3 and SEQ ID NO: 4 can contain RNA molecules having nucleotide sequences according to.

[0108] Fragments of the RNA molecules listed in (i) to (v), or sequences having at least 80%, preferably at least 85%, more preferably at least 90%, still more preferably at least 95%, for example 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99% sequence identity to the RNA molecules listed in (i) to (v) are also included.

[0109] It should be noted that at least three RNA molecules, especially four RNA molecules, contained in the composition have a characteristic distribution. In particular, at least three RNA molecules, especially four RNA molecules, contained in the composition have a characteristic distribution with respect to their amounts. Specifically, at least three RNA molecules, particularly four RNA molecules, are included in the composition in different amounts. For example, any pair of the two RNA molecules included in the composition has different amounts.

[0110] Thus, when three RNA molecules are included in the composition, the first RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 1) is present in the largest amount, the second RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 2) is present in an amount less than that of the first RNA molecule, and the third RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 3) is present in an amount less than that of the first and second RNA molecules. When four RNA molecules are included in the composition, the first RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 1) is present in the largest amount, the second RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 2) is present in an amount less than that of the first RNA molecule, the third RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 3) is present in an amount less than that of the first and second RNA molecules, and the fourth RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 4) is present in an amount less than that of the first, second, and third RNA molecules.

[0111] More specifically, at least three RNA molecules, particularly four RNA molecules, are included in the composition in a defined amount gradient. In particular, at least three RNA molecules, particularly four RNA molecules, are titrated into different amounts in a linear range.

[0112] In one embodiment, at least three RNA molecules, particularly four RNA molecules, contained in the composition are present in an amount of 0.001 amol to 6000 amol, preferably 0.01 amol to 5000 amol, more preferably 0.1 amol to 4000 amol, still more preferably 1 amol to 3500 amol, for example 0.001, 0.002, 0.003, 0.004, 0.005, 0.006, 0.007, 0.008, 0.009, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 3000, 3500, 4000, 4500, 5000, 5500 or 6000 amol.

[0113] In one example, this composition contains at least three RNA molecules, the first RNA molecule is contained in an amount of about 3400 amol, the second RNA molecule is contained in an amount of about 725 amol, the third RNA molecule is contained in an amount of about 20 amol, or, the first RNA molecule is contained in an amount of about 1360 amol, the second RNA molecule is contained in an amount of about 290 amol, and the third RNA molecule is contained in an amount of about 80 amol. Specifically, the first RNA molecule has a nucleotide sequence according to SEQ ID NO: 1, the second RNA molecule has a nucleotide sequence according to SEQ ID NO: 2, and the third RNA molecule has a nucleotide sequence according to SEQ ID NO: 3.

[0114] In certain embodiments, the four RNA molecules comprised in the composition are present in an amount of from 0.001 amol to 6000 amol, preferably from 0.01 amol to 5000 amol, more preferably from 0.1 amol to 4000 amol, even more preferably from 1 amol to 3500 amol, for example in an amount of 0.001, 0.002, 0.003, 0.004, 0.005, 0.006, 0.007, 0.008, 0.009, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 3000, 3500, 4000, 4500, 5000, 5500 or 6000 amol.

[0115] In certain examples, the composition comprises four RNA molecules, the first RNA molecule comprises an amount of about 3400 amol, the second RNA molecule comprises an amount of about 725 amol, the third RNA molecule comprises an amount of about 20 amol, the fourth RNA molecule comprises an amount of about 7 amol, or, the first RNA molecule comprises an amount of about 1360 amol, the second RNA molecule comprises an amount of about 290 amol, the third RNA molecule comprises an amount of about 80 amol, the fourth RNA molecule comprises an amount of about 27 amol.

[0116] Specifically, the first RNA molecule has a nucleotide sequence according to SEQ ID NO: 1, the second RNA molecule has a nucleotide sequence according to SEQ ID NO: 2, the third RNA molecule has a nucleotide sequence according to SEQ ID NO: 3, and the fourth RNA molecule has a nucleotide sequence according to SEQ ID NO: 4.

[0117] Preferably, the composition is a solution. More preferably, the solution is an aqueous solution such as nuclease-free water. Even more preferably, no other components (other than the RNA molecules and water as the solvent) are comprised as part of the composition.

[0118] As described above, the analysis of RNA expression profiles involves several complex steps. (1) RNA, including small RNAs, must be extracted from biological samples such as blood or saliva without distorting the original relative abundance of the RNA (it is necessary to maintain linearity). (2) The abundance of small RNAs must be accurately quantified. Both steps are extremely important for standardizing and monitoring the efficiency of RNA extraction and detection because they can introduce inconsistent biases (such as the presence of inhibitors of cDNA synthesis).

[0119] The above composition is suitable as / is a spike-in cocktail as such. The spike-in cocktail is a general-purpose spike-in cocktail that is independent of subsequent detection methods. Furthermore, this composition is suitable as / is a standard for process control and / or normalization. For this purpose, the above composition is added to a sample containing biological material or a processed sample based on / derived from biological material.

[0120] In particular, this composition can be generally applied to any downstream processing of biological materials, samples containing biological materials, samples containing materials of biological origin, and / or samples based on / derived from biological materials.

[0121] The above biological material contains target RNA molecules required to be detected, for example, for diagnosing diseases such as cancer or neurodegenerative diseases such as Parkinson's disease (PD) or Alzheimer's disease (AD). The target RNA molecules are endogenous to the biological material and are different from exogenous artificial RNA molecules. Exogenous artificial RNA molecules are added to the biological material as part of a composition such as a spike-in cocktail, and the target RNA molecules are part of the biological material.

[0122] Process control specifically includes monitoring / detecting the inefficiency of inspections during the processing and analysis of target RNA, monitoring the efficiency of target RNA extraction and / or detection, standardizing target RNA extraction and detection, and monitoring the quantification and / or detection of target RNA. This composition can be used as an extraction control standard, a sequencing control standard, or a library preparation control standard.

[0123] Furthermore, this composition can also be used as a normalization indicator (compared with different samples together with the composition). In particular, this composition can be used as a normalization indicator (compared with different samples together with the composition) such that the measured spike-in expression value helps to control and correct inefficiencies in extraction, library preparation, or sequencing.

[0124] In a second aspect, the present invention relates to a kit comprising the composition of the first aspect. This composition is also called a spike-in cocktail. Preferably, this kit comprises means for determining the levels of at least three RNA molecules (e.g., three or four RNA molecules) contained in the composition. More preferably, this means a polynucleotide (probe) for detecting RNA molecules, a primer / primer pair for binding RNA molecules, an antibody capable of hybridizing and binding the polynucleotide probe and the RNA molecule, and / or means for performing next-generation sequencing (NGS).

[0125] The polynucleotide (probe) may be part of a microarray / bichip or attached to beads in a bead-based multiplex system. The primer / primer pair may be part of an RT-PCR system, a PCR system, or a next-generation sequencing system.

[0126] The means may further comprise a microarray, an RT-PCR system, a PCR system, a flow cytometer, a Luminex system, and / or a next-generation sequencing system.

[0127] The kit may include instructions on how to carry out the method of the present invention (see the fourth to sixth aspects). The kit is also useful for carrying out the method of the present invention (see the fourth to sixth aspects).

[0128] Furthermore, the kit may comprise a container and / or a data carrier.

[0129] The data carrier may be a non-electronic data carrier, such as a graphical data carrier like an information leaflet, an information sheet, a barcode, an access code, or an electronic data carrier such as a floppy disk, a compact disc (CD), a digital versatile disc (DVD), a microchip, or other semiconductor-based electronic data carriers. The access code may permit access to, for example, an Internet database, a centralized database, or a distributed database. The access code may permit access to application software that causes a computer to perform tasks for a computer user, or to a mobile application that is software designed to be executed on a smartphone and other mobile devices.

[0130] The data carrier may further include information regarding the expected levels, expected order / rank, and / or expected linearity of RNA molecules included in the composition in subsequent RNA processing / analysis assays. The data carrier may also include information or instructions regarding how to carry out the method of the present invention (see the fourth to sixth aspects).

[0131] The kit is preferably for in vitro use / is in vitro. In a third aspect, the present invention relates to the (in vitro) use of the composition of the first aspect or the kit (as a standard) of the second aspect for process control, sample testing, sample analysis, normalization, and / or data processing control.

[0132] The composition is also called a spike-in cocktail. Preferably, the data processing control includes raw data processing control. As described above, the analysis of RNA expression profiles involves several complex steps. (1) RNA containing small RNAs must be extracted from biological sources such as blood and saliva without distorting the original relative abundance of the RNA (linearity must be maintained). (2) The abundance of small RNAs must be accurately quantified. Both steps can introduce inconsistent biases (e.g., due to the presence of inhibitors of cDNA synthesis), so it is extremely important to standardize and monitor the efficiency of RNA extraction and detection.

[0133] The composition of the first aspect is used (as a standard) for process control, sample inspection, sample analysis, normalization, and / or data processing control. Preferably, the data processing control includes raw data processing control. For this purpose, the composition of the first aspect is added to a sample containing biological material or a processed sample based on / derived from biological material.

[0134] In particular, the composition is generally applicable to any downstream processing of biological materials, samples containing biological materials, samples containing materials of biological origin, and / or samples based on / derived from biological materials.

[0135] The biological material includes target RNA molecules that are required to be detected, for example, for diagnosing diseases such as cancer, neurodegenerative diseases such as Parkinson's disease (PD) and Alzheimer's disease (AD). The target RNA molecule has an endogenous origin in the biological material and is different from exogenous artificial RNA molecules. While exogenous artificial RNA molecules are added to the biological material as part of a composition such as a spike-in cocktail, the target RNA molecule is part of the biological material.

[0136] Process control specifically includes monitoring / detecting inefficiencies in inspections during the processing and analysis of target RNA, monitoring the efficiency of extraction and / or detection of target RNA, standardizing the extraction and detection of target RNA, and monitoring the quantification and / or detection of target RNA.

[0137] Preferably, process control includes quality control, quantity control, or end-to-end control. End-to-end control means specifically that only samples in which the recovery of artificial small RNA molecules is obtained at the levels, order, and linearity predicted by the analysis are used for downstream analysis.

[0138] The composition of the first aspect can be used particularly as an extraction control standard, a sequence control standard, or a library preparation control standard. Furthermore, the composition of the first aspect can be used particularly as a normalization index (for comparing samples with different compositions). This enables the removal of batch effects in different tests and / or the control of inefficiencies in the extraction, library preparation, and / or sequencing processes.

[0139] Process control, quality control, and preprocessing of data are important for data analysis because data generated after sequencing such as next-generation sequencing, especially raw data, needs to be processed so that false positives and false negatives do not appear in the results. Preprocessing of data not only evaluates each analysis step but also reduces the amount of low-quality sequence reads. By removing such low-quality reads, the time and cost of computational analysis can be reduced, and highly reliable and high-quality results can be obtained. In any downstream analysis of NGS data, false positive and false negative results Inspection factors: such as sample contamination and PCR errors, Sequencing factors: This includes the quality of sequencing and data contamination caused by index hopping when splitting data, Analysis software parameter factors: This includes the exact type of parameter adjustment in alignment software or downstream individual analysis, Generated by.

[0140] The raw sequences generated after sequencing contain not only sequences of interest such as target RNA molecules, but also sequence biases (e.g., due to systematic effects such as Poisson sampling), and complex artifacts generated by the sequencing and inspection procedures. These sequence biases and artifacts interfere with accurate read alignment that affects genotyping and variant calling. Therefore, preprocessing of raw sequence reads is essential to improve the reliability and quality of downstream analysis and reduce the amount of computational resources required.

[0141] The inventors have discovered that adducts at the 5'-end and / or 3'-end to the target RNA molecule, such as those generated or occurred during a sequencing process such as a next-generation sequencing process, can cause misunderstandings and distort the results of raw data. In particular, sequence reads contaminated with adapters need to be excluded. The inventors have found that by excluding sequence reads contaminated with adapters, the quality of raw data can be improved. Advantageously, 5'-end and / or 3'-end adducts to the target RNA molecule can be detected using at least three artificial small RNA molecules (spike-ins) contained in / derived from the composition of the first aspect of the present invention. The presence of 5'-end and / or 3'-end adducts identified for at least three artificial small RNA molecules (spike-ins) contained in / derived from the composition of the first aspect of the present invention indicates the presence of 5'-end and / or 3'-end adducts in the target RNA molecule. Thus, the processing of (raw) data can be controlled and particularly improved by excluding RNAs that are not of biological origin and should not be used for the biological interpretation of the data.

[0142] In a fourth aspect, the present invention relates to a method of examining a sample comprising the step of evaluating the sample with respect to at least three RNA molecules (e.g., three or four RNA molecules) contained in / derived from the composition of the first aspect.

[0143] In one embodiment, the sample is a mixture of biological materials having the composition of the first aspect. Preferably, the biological material is a tissue or a body fluid. More preferably, the body fluid is blood. Even more preferably, the blood is whole blood or a blood fraction. In particular, the blood fraction is selected from the group consisting of a blood cell fraction and plasma or serum. The blood cell fraction represents the cellular part of (whole) blood. Plasma and serum represent the cell-free part of (whole) blood. More specifically, the blood cell fraction contains red blood cells, white blood cells, or platelets, and the blood cell fraction is a fraction of red blood cells, white blood cells, or platelets, or the blood cell fraction is a mixture of red blood cells, white blood cells, and platelets.

[0144] In one alternative embodiment, the sample is based on / derived from a mixture of biological materials having the composition of the first aspect. Specifically, a sample based on / derived from a mixture of biological materials having the composition of the first aspect is a processed sample. More specifically, the processed sample is a lysed sample, an extracted sample, an amplified sample, a sequenced sample, or a library-prepared sample.

[0145] The processed sample may be obtained by adding the composition of the first aspect to the biological material, mixing the composition of the first aspect with the biological material, and further processing the sample. Preferably, the processed sample is obtained by mixing the biological material with the composition of the first aspect and further processing the resulting mixture.

[0146] In order to evaluate the sample with respect to at least three RNA molecules (e.g., three or four RNA molecules) contained in / derived from the composition of the first aspect, the levels of at least three of said RNA molecules (e.g., three or four RNA molecules) in the sample are preferably determined.

[0147] In a preferred embodiment, the evaluation comprises determining whether at least three RNA molecules (e.g., three or four RNA molecules) exhibit a characteristic distribution, or whether the characteristic distribution of the RNA molecules in the sample matches the expected characteristic distribution.

[0148] As described above, at least three RNA molecules (e.g., three or four RNA molecules) contained in the composition of the first aspect have a characteristic distribution, particularly with respect to their amounts. In particular, at least three RNA molecules (e.g., three or four RNA molecules) are contained in the composition of the first aspect in different amounts. More specifically, at least three RNA molecules (e.g., three or four RNA molecules) are contained in the composition of the first aspect in a defined gradient of amounts.

[0149] To determine whether at least three RNA molecules (e.g., three or four RNA molecules) exhibit a characteristic distribution, preferably, the levels of said at least three RNA molecules (e.g., three or four RNA molecules) in the sample are determined.

[0150] If a characteristic distribution is given, the sample is further processed and / or analyzed. In particular, a characteristic distribution is when at least three RNA molecules are present at expected levels, when at least three RNA molecules are present in an expected order / rank (in particular, defined by the levels of at least three RNA molecules, especially the quantitative relationship), and / or when at least three RNA molecules are present in an expected linearity.

[0151] If no characteristic distribution is given, the sample is not further processed and / or analyzed. In that case, it is discarded. In particular, a characteristic distribution is when at least three RNA molecules are not present at expected levels, when at least three RNA molecules are not present in an expected order / rank (in particular, defined by the levels of at least three RNA molecules, especially the quantitative relationship), and / or, when at least three RNA molecules are not present in an expected linearity.

[0152] The expected level means the level expected as a specific RNA molecular weight added to the sample in subsequent sample processing and / or analysis. As described above, only a specific amount of RNA molecules is added to the sample or contained in the sample. The amount of RNA molecules may correspond / correlate to the expected level of RNA molecules that can be measured during subsequent sample processing and / or analysis. Specifically, the amount of RNA molecules may correspond / correlate to a specific number of reads (reads per million (RPM)) in a next-generation sequencing assay, and / or the amount of RNA molecules may correspond / correlate to a specific cycle threshold (Ct) in a real-time PCR test. In particular, the cycle threshold (Ct) is the number of cycles in real-time PCR required to exceed a predefined threshold in the measurement signal (e.g., fluorescence signal) of the amplified DNA. The more DNA (RNA) present in the sample solution before PCR, the fewer amplification cycles are required to reach the corresponding threshold.

[0153] For example, when an RNA molecule having a nucleotide sequence according to SEQ ID NO: 1 is added to the sample in an amount of about 3400 amol or contained in the composition of the first aspect, the predicted Ct value is about 26 (the average Ct value is about 24.3), and when an RNA molecule having a nucleotide sequence according to SEQ ID NO: 2 is added to the sample in an amount of about 725 amol or contained in the composition of the first aspect, the predicted Ct value is about 29.3 (the average Ct value is about 27.1), and when an RNA molecule having a nucleotide sequence according to SEQ ID NO: 3 is added to the sample in an amount of about 20 amol or contained in the composition of the first aspect, the predicted Ct value is about 35.9 (the average Ct value is about 32), and / or when an RNA molecule having a nucleotide sequence according to SEQ ID NO: 4 is added to the sample in an amount of about 7 amol or contained in the composition of the first aspect, the predicted Ct value is about 39.2 (the average Ct value is about 33.6).

[0154] For example, when an RNA molecule having a nucleotide sequence according to SEQ ID NO: 1 is added to a sample in an amount of about 3400 amol or is included in the composition of the present invention, it results in a log2RPM of about 11.972 or a read count of about 30628; when an RNA molecule having a nucleotide sequence according to SEQ ID NO: 2 is added to a sample in an amount of about 725 amol or is included in the composition of the first aspect, it results in a log2RPM of about 10.926 or a read count of about 5554; when an RNA molecule having a nucleotide sequence according to SEQ ID NO: 3 is added to a sample in an amount of about 20 amol or is included in the composition of the first aspect, it results in a log2PRM of about 6.559 or a read count of about 491; and / or when an RNA molecule having a nucleotide sequence according to SEQ ID NO: 4 is added to a sample in an amount of about 7 amol or is included in the composition of the first aspect, it results in a log2RPM of about 4.061 or a read count of about 94.

[0155] The expected order / rank refers to the order / rank expected for specific RNA molecules added to a sample in subsequent sample processing and / or analysis. As described above, the RNA molecules added to a sample are each in different amounts. Thus, if three RNA molecules are added to a sample or are included in the composition of the first aspect, the first RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 1) is present in the largest amount, the second RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 2) is present in an amount less than the first RNA molecule, and the third RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 3) is present in an amount less than the first and second RNA molecules. If four RNA molecules are added to a sample or are included in the composition of the first aspect, the first RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 1) is present in the largest amount, the second RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 2) is present in an amount less than the first RNA molecule, the third RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 3) is present in an amount less than the first and second RNA molecules, and the fourth RNA molecule (e.g., an RNA molecule having a nucleotide sequence according to SEQ ID NO: 4) is present in an amount less than the first, second, and third RNA molecules.

[0156] This expected order / rank representing the initial state / condition must be recovered during sample processing and / or analysis. In this regard, it should be noted that the amount of RNA molecules included in a sample or composition in a specific order / rank may correspond / correlate to the expected levels of RNA molecules measured during subsequent sample processing and / or analysis. For example, the amount of RNA molecules may correspond to a specific number of reads in a next-generation sequencing assay, and the amount of RNA molecules may correspond / correlate to a specific cycle threshold (Ct) in a real-time PCR test.

[0157] The expected linearity means the linearity expected for specific RNA molecules added to the sample in subsequent sample processing and / or analysis. This expected linearity representing the initial state / condition must be restored during sample processing and / or analysis.

[0158] Therefore, only samples in which RNA molecules or their surrogates are recovered at the expected levels, order and / or linearity are used for subsequent analysis and / or further processing. To determine whether the RNA molecules are present in the expected order and / or linearity, statistical analysis is performed using data (e.g., Ct values or RPM) regarding the RNA molecules. The statistical analysis includes, but is not limited to, Spearman (rank) correlation analysis and / or Pearson correlation analysis.

[0159] For example, to determine whether the RNA molecules are present in the expected order, preferably the Spearman rank correlation coefficient (Spearman ρ) is determined. Further, to determine whether the RNA molecules are present in the expected linearity, preferably the Pearson correlation coefficient (Pearson r) is determined.

[0160] Preferably, when the Spearman rank correlation coefficient (Spearman ρ) is 0.95 or more, the RNA molecules are present in the expected order. In this case, the sample is further processed / analyzed. Therefore, when the Spearman rank correlation coefficient (Spearman ρ) is less than 0.95, the sample is discarded. In other words, such a sample is not further processed / analyzed.

[0161] Preferably, when the Pearson correlation coefficient (Pearson r) is 0.66 or more, the RNA molecules are present in the expected linearity. In this case, the sample is further processed / analyzed. Therefore, when the Pearson correlation coefficient (Pearson r) is less than 0.66, the sample is discarded. In other words, such a sample is not further processed / analyzed.

[0162] Further processing may include sample lysis, sample extraction, sample amplification, sample sequencing, and / or library preparation from the sample. Specifically, further processing includes cell lysis to release the nucleotide sequence contained in the sample, extraction of the nucleotide sequence contained in the sample, amplification of the nucleotide sequence contained in the sample, sequencing of the nucleotide sequence contained in the sample, and / or library preparation from the nucleotide sequence contained in the sample. In particular, the nucleotide sequence is a ribonucleic acid sequence. Preferably, the ribonucleic acid sequence belongs to a target RNA molecule. More preferably, the target RNA molecule is a small RNA molecule. Even more preferably, the small RNA molecule is a non-coding small RNA molecule. Most preferably, the non-coding small RNA molecule is a miRNA molecule and / or an isomiR molecule.

[0163] In this regard, it should be noted that artificial exogenous RNA molecules added to biological materials are treated in the same manner as endogenous target RNA molecules contained in the biological materials. RNAs such as small RNAs in biological samples such as body fluids play an important role as biomarkers for the prognosis and diagnosis of many human pathologies. However, the accurate analysis of these RNAs in biological samples such as body fluids is very important regarding the significance of these data in the biomarker field.

[0164] In the method according to the present invention for sample inspection, the sample is evaluated with respect to at least three RNA molecules contained in / derived from the composition described herein, and the at least three RNA molecules are selected from the group consisting of RNA molecules having nucleotide sequences according to SEQ ID NO: 1 to SEQ ID NO: 4. The at least three RNA molecules are part of the sample. They are added to the sample (as spike-ins).

[0165] In this method, it is alternatively or additionally verified / analyzed whether at least three RNA molecules having nucleotide sequences according to SEQ ID NO: 1 to SEQ ID NO: 4 contain 5'-end and / or 3'-end adducts. In this regard, it will be understood that the term "RNA molecule" may refer not only to the RNA molecule itself, but also to its alternatives, such as amplification products (e.g., cDNA derived therefrom).

[0166] Thus, in a preferred embodiment, the evaluation includes identifying 5'-end and / or 3'-end adducts of at least three RNA molecules (e.g., three or four RNA molecules) contained in / derived from the composition of the first aspect. Preferably, the 5'-end and / or 3'-end adducts have a length of at least 5 nucleotides. More preferably, the 5'-end and / or 3'-end adducts are 5 to 30 nucleotides, even more preferably 5 to 20 nucleotides, even even more preferably 7 to 15 nucleotides, for example, having a length of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides. The 5'-end adduct can also be called a prefix, the 3'-end adduct can also be called a suffix, and the 5'-end and 3'-end adducts can also be called an affix.

[0167] Specifically, at least five nucleotides extend beyond the original length of at least three RNA molecules. Prefixes, suffixes and affixes can extend the original length of the RNA molecule by only 1, 2, 3 or 4 nucleotides. However, in this context, a length of 5 nucleotides or more is required / recommended herein to reduce the risk of misidentifying the RNA molecule to which the 5'-end and / or 3'-end adduct is related.

[0168] The adducts to the 5'-end and / or 3'-end of at least three RNA molecules contained in / derived from the composition of the first aspect can be readily determined using sequencing methods and / or sequence alignment methods. Since the sequences of artificial small RNA molecules (spike-ins) are known, the identification / analysis of the added nucleotides is not a substantial problem and is a standard procedure for those skilled in the art. As described above, in this context, a minimum length of 5 nucleotides is required to designate as a 5'-end adduct and / or 3'-end adduct in order to reduce the risk of misidentifying RNA molecules having 5'-end and / or 3'-end adducts.

[0169] The 5'-end and / or 3'-end adducts may be the result of an RNA molecule / adapter fusion, an RNA molecule / RNA molecule fusion, or an adapter / adapter fusion. An adapter may be any non-biological RNA / DNA sequence that is intentionally added to the 5'-end or 3'-end of an RNA / cDNA molecule as part of sequencing, such as the NGS method. This includes free adapter molecules that may fuse with RNA in an unintended way. The adapter may be a combination / fusion of multiple individual adapters and index RNAs. Index RNAs are used as UMIs (Unique Molecular Identifiers) to assign samples from which RNA molecules are derived during multiplex sequencing, i.e., simultaneous sequencing of multiple samples. The UMI sequence is a short non-specific (random) sequence of a predetermined length, such as 12 nucleotides, although the length may vary.

[0170] The samples evaluated in the above embodiments preferably include target RNA molecules, such as those contained in biological materials such as whole blood, or those derived from biological materials such as whole blood. Therefore, in the next step, preferably, for at least three RNA molecules contained in / derived from the composition of the first aspect (artificial small RNA molecule / spike-in), it is determined whether the 5'-end and / or 3'-end adducts identified are also found in the target RNA molecule or its alternative (e.g., cDNA molecule) that is part of / included in the sample. In other words, the presence of the 5'-end and / or 3'-end adducts identified for at least three RNA molecules contained in / derived from the composition of the first aspect (artificial small RNA molecule / spike-in) preferably indicates the presence of 5'-end and / or 3'-end adducts in the target RNA molecule or its alternative (e.g., cDNA molecule).

[0171] Preferably, the 5'-end and / or 3'-end adducts are - for example, the ligation of a 5'-adapter and / or 3'-adapter to a (denatured) (target) RNA molecule using a double-stranded RNA ligase such as T4 RNA ligase 2 (Rnl2) or Kod1 ligase, - for example, the reverse transcription of a (target) RNA molecule (to which a 5'-adapter and / or 3'-adapter is ligated, also called the ligation product) to a cDNA molecule using a reverse transcriptase (RT) such as Maxima H-RT or Tth polymerase, - for example, the amplification of the (said) cDNA molecule by polymerase chain reaction (PCR), and / or - the sequencing of the (said) cDNA molecule such as next-generation sequencing, that occur during.

[0172] More preferably, the 5'-end and / or 3'-end adducts occur during the library preparation process or during the next-generation sequencing process for next-generation sequencing. Even more preferably, the next-generation sequencing process preferably - for example, the ligation of a 5'-adapter and / or 3'-adapter to a (denatured) (target) RNA molecule using a double-stranded RNA ligase such as T4 RNA ligase 2 (Rnl2) or Kod1 ligase, - Reverse transcription of a (target) RNA molecule (also called a binding product, having a 5' adapter and / or a 3' adapter bound thereto) into a cDNA molecule using a reverse transcriptase (RT) such as, for example, Maxima H-RT or Tth polymerase, - Amplification of the (said) cDNA molecule by, for example, polymerase chain reaction (PCR), and / or, - Sequencing of the (said) cDNA molecule, is included.

[0173] More preferably, the 5'-end and / or 3'-end adducts are CGATC (SEQ ID NO: 10), GGGGC (SEQ ID NO: 11), ACGATC (SEQ ID NO: 12), GGGCGT (SEQ ID NO: 13), CGGCGG (SEQ ID NO: 14), GGGGC G (SEQ ID NO: 15), GACGATC (SEQ ID NO: 16), GGGGC GT (SEQ ID NO: 17), GGGCGTG (SEQ ID NO: 18), GGGGGCG (SEQ ID NO: 19), GGGGG TG (SEQ ID NO: 20), GGGGGCGTG (SEQ ID NO: 21), CGGGGCGG (SEQ ID NO: 22), GGGAGGCC (SEQ ID NO: 23), GGAGGCGT (SEQ ID NO: 24), GGGCGTGG (SEQ ID NO: 25), TGGAGGCG (SEQ ID NO: 26), CGACGATC (SEQ ID NO: 27), GGGCGTT (SEQ ID NO: 28), GGGCGTGT (SEQ ID NO: 29), GGGGGCGT (SEQ ID NO: 30), GGGAGCCA (SEQ ID NO: 31), GGGGGGTGT (SEQ ID NO: 32), GGAGGCCC (SEQ ID NO: 33), CCGACGATC (SEQ ID NO: 34), GGGGGCGTG (SEQ ID NO: 35), TACCTGGTT (SEQ ID NO: 36), TGGAGGCGT (SEQ ID NO: 37), GGGCGTGGG (SEQ ID NO: 38), CGGCGGCGG (SEQ ID NO: 39), GGGGGGTGTA (SEQ ID NO: 40), GGGGGCGTT (SEQ ID NO: 41), GGCTGGGCG (SEQ ID NO: 42), TCGGGGGCGG (SEQ ID NO: 43), GGGGC GTGG (SEQ ID NO: 44), GGGGGAGCCA (SEQ ID NO: 45), GGGGGAGGCCC (SEQ ID NO: 46), CGGAGGGCGG (SEQ ID NO: 47), GTCCGCGATC (SEQ ID NO: 48), GTCGACGATC (SEQ ID NO: 49), CGGGCGGATC (SEQ ID NO: 50), TGGAGGCGTG (SEQ ID NO: 51), TCCGACGATC (SEQ ID NO: 52), GGGGGCGTGGG (SEQ ID NO: 53), AAGCGGGGCT (SEQ ID NO: 54), CGGGAGCCA (SEQ ID NO: 55), GTCCGACGATC (SEQ ID NO: 56), TCGGAGGGCGG (SEQ ID NO: 57), AGTCCGACGATC (SEQ ID NO: 58), AAGCGGGGCTGG (SEQ ID NO: 59), GTCCGACGGATC (SEQ ID NO: 60), TCGGGCTGGGGC (SEQ ID NO: 61), TACCTGGTTGAT (SEQ ID NO: 62), TCGGGGGCGGCGG (SEQ ID NO: 63).It is selected from the group consisting of adducts having nucleotide sequences according to CAGTCCGACGATC (SEQ ID NO: 64), TACCTGGTTGATC (SEQ ID NO: 65), TCGGGCTGGGGCG (SEQ ID NO: 66), TGGAGGCGTGGGT (SEQ ID NO: 67), ACAGTCCGACGATC (SEQ ID NO: 68), GGTCGGGCTGGGGC (SEQ ID NO: 69), CGGAAGCGTGCTGGG (SEQ ID NO: 70), GGTCGGGCTGGGGCG (SEQ ID NO: 71), TACAGTCCGACGATC (SEQ ID NO: 72), CGGAAGCGTGCTGGGC (SEQ ID NO: 73), CTACAGTCCGACGATC (SEQ ID NO: 74), TCTACAGTCCGACGATC (SEQ ID NO: 75), CGGAAGCGTGCTGGGCCC (SEQ ID NO: 76), TCGGGGGCGGCGGCGGCGG (SEQ ID NO: 77), TTCTACAGTCCGACGATC (SEQ ID NO: 78), TAGCAGCACATCATGGTT (SEQ ID NO: 79), GGATCATTA (SEQ ID NO: 80), GGGGC GTGGG (SEQ ID NO: 81) and TGGAGGCGTGGGT (SEQ ID NO: 82).

[0174] Specifically, determination of 5'-end and / or 3'-end adducts to a (target) RNA molecule includes analysis of whether the (target) RNA molecule has a sequence that is at least partially identical to an adapter sequence used in a sequencing, preferably a next-generation sequencing process.

[0175] More specifically, the adapter sequences used in the next-generation sequencing process are selected from the group consisting of TGGAATTCTCGGGTGCCAAGG (SEQ ID NO: 83), GTTCAGAGTTCTACAGTCCGACGATC (SEQ ID NO: 84), TGGAATTCTCGGGTGCCAAGG (SEQ ID NO: 85), GAATTCCACCACGTT CCCGTGG (SEQ ID NO: 86), AATGATACGGCGACCACCGAGATCTACACGTTCAGAGTTCTACAGTCCGA (SEQ ID NO: 87), CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 88), GTGACTGGAGTTCCTTGGCACCCGAGAATTCCA (SEQ ID NO: 89), CAAGCAGAAGACGGCATACGA (SEQ ID NO: 90), GTTCAGAGTTCTACAGTCCGACGATC (SEQ ID NO: 91), TCGTATGCCGTCTTCTGCTTGT (SEQ ID NO: 92), ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 93), CAAGCAGAAGACGGCATACGA (SEQ ID NO: 94), AATGATACGGCGACCACCGACAGGTTCAGAGTTCTACAGTCCGA (SEQ ID NO: 95), CGACAGGTTCAGAGTTCTACAGTCCGACGATC (SEQ ID NO: 96), AAAAAAAAAAAAAAAAAAAAAAAAAAAAA (SEQ ID NO: 97), AGATCGGAAGAGCACACGTCTGAACTCCAGTCA (SEQ ID NO: 98), AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT (SEQ ID NO: 99), GATCGGAAGAGCACACGTCTGAACTCCAGTCAC (SEQ ID NO: 100), ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 101), AATGATACGGCGACCACCGAGATCTACAC (SEQ ID NO: 102) and ACACTCTTTCCCTACACGACGCTCTTCCGATCT (SEQ ID NO: 103).

[0176] The raw data generated after array determination must be processed to avoid false positives and false negatives in the results. Therefore, process control, quality management, and preprocessing of data are important for data analysis. Preprocessing of data not only evaluates each analysis step but also reduces the amount of low-quality array reads. By removing such low-quality reads, the time and cost of computational analysis can be reduced, and highly reliable and high-quality results can be obtained.

[0177] In any downstream analysis of NGS data, false positive and false negative results are caused by the following factors. Inspection factors: such as sample contamination and PCR errors, Sequencing factors: This includes the quality of sequencing and data contamination caused by index hopping when dividing data, Analysis software factors: This includes alignment software or precise adjustment of parameters in downstream individual analyses.

[0178] The raw sequences generated after sequencing include not only sequences of interest such as target RNA molecules but also sequence biases (e.g., due to systematic effects such as Poisson sampling) and complex artifacts generated by the sequencing and inspection procedures. These sequence biases and artifacts affect and interfere with accurate read alignment, thus affecting gene typing and variant calling. Therefore, preprocessing of raw sequence reads is essential to enhance the reliability and quality of subsequent analyses and reduce the amount of computational resources required.

[0179] The inventors discovered that adducts added to the 5' end and / or 3' end of (target) RNA molecules generated or occurring during sequencing, such as next-generation sequencing, can lead to misunderstandings and falsify the results of raw data. Specifically, sequence reads contaminated with adapters need to be excluded. The inventors discovered that the quality of raw data can be improved by excluding sequence reads contaminated with adapters.

[0180] Thus, target RNA molecules having 5'-end and / or 3'-end adducts contained in the sample are excluded from further analysis / are not used further. Alternatively, or additionally, data from target RNA molecules having 5'-end and / or 3'-end adducts contained in the sample are excluded from further analysis / are not used further or do not form part of a data set, preferably a raw data set. For example, data regarding these target RNA molecules are simply excluded from the (raw) data set.

[0181] Preferably, the target RNA molecule is a small RNA molecule. More preferably, the small RNA molecule is a non-coding small RNA molecule. Even more preferably, the non-coding small RNA molecule is a miRNA and / or isomiR molecule.

[0182] The sample used in the above embodiments is preferably a lysed sample, an extracted sample, an amplified sample and / or a sequenced sample. Lysis of the sample involves lysis of the cells contained in the sample and is necessary to release the (target) RNA molecules contained therein. Extraction of the sample is necessary to extract the (target) RNA molecules contained therein. Amplification of the sample is necessary to amplify the (target) RNA molecules contained therein. Sequencing of the sample is necessary to sequence the (target) RNA molecules contained therein.

[0183] For example, a biological sample such as a whole blood sample is provided as the sample. In a first step, the sample, specifically the cells contained in the sample, are lysed and the (target) RNA molecules contained therein are released. Thereafter, cDNA is generated from the (target) RNA molecules using an adapter that binds to the (target) RNA molecules denatured by reverse transcription. The obtained cDNA molecules are then amplified, and finally sequencing such as next-generation sequencing of the cDNA molecules becomes possible.

[0184] In any of the above processes, an adduct at the 5'-end and / or 3'-end may be added to the (target) RNA molecule that needs to be identified, and the (target) RNA molecule having these 5'-end and / or 3'-end adducts needs to be excluded from subsequent analysis or raw data evaluation. This improves the quality of the data.

[0185] The analysis of RNA expression profiles involves several complex steps. (1) RNA containing small RNAs must be extracted from biological samples such as blood and saliva without distorting the original relative abundance of the RNA (it is necessary to maintain linearity). (2) The abundance of small RNAs must be accurately quantified. Since both steps can introduce inconsistent biases (e.g., the presence of inhibitors of cDNA synthesis), it is extremely important to standardize and monitor the efficiency of RNA extraction and detection. These issues can be addressed in several ways.

[0186] The inventors designed and optimized a general-purpose spike-in system of artificial small RNA molecules that widely reflects the behavior of endogenous miRNAs during RNA extraction and detection. This mixture of artificial small RNA molecules is added to a biological sample such as a clinical sample at the start of the treatment, functions as an end-to-end quality control means, and as a result of the analysis, only samples in which the artificial small RNA molecules are recovered at the expected levels, order, and linearity are used for downstream analysis. Accordingly, in a fifth aspect, the present invention relates to a method for optimizing the treatment of a biological sample, including the step of implementing the method of the fourth aspect.

[0187] In a sixth aspect, the present invention relates to a method for optimized RNA preparation from a biological sample or for RNA analysis of a biological sample, including the step of implementing the method of the fourth aspect. In a seventh aspect, the present invention relates to a method for improving the quality of an (RNA) dataset, including the step of implementing the method of the fourth aspect.

[0188] Data such as the raw data generated after array determination needs to be processed so that false positives and false negatives do not occur in the results. Therefore, process control, quality control, and preprocessing of data are important for data analysis. Preprocessing of data not only evaluates each analysis step but also reduces the amount of low-quality array reads. By removing such low-quality reads, the time and cost of computational analysis can be reduced, and highly reliable and high-quality results can be obtained. In any downstream analysis of NGS data, the causes of false positive and false negative results are the following factors. Inspection factors: such as sample contamination and PCR errors, Sequencing factors: This includes the quality of the sequence, data contamination caused by index hopping during data splitting, and / or Parameter factors of the analysis software: This includes the alignment software and the exact type of parameter adjustment in downstream individual analysis.

[0189] The raw sequences generated after sequencing include not only sequences of interest such as target RNA molecules but also sequence biases (e.g., due to systematic effects such as Poisson sampling) and complex artifacts generated by the sequencing and inspection procedures. These sequence biases and artifacts affect accurate read alignment, interfere with it, and affect gene typing and variant calling. Therefore, preprocessing of raw sequence reads is essential to improve the reliability and quality of downstream analysis and reduce the amount of computational resources required.

[0190] The inventors have found that, for example, the adducts added to the 5' end and / or 3' end of (target) RNA molecules generated or occurring during a sequencing process such as a next-generation sequencing process can cause misunderstandings and tamper with the results of raw data. In particular, sequence reads contaminated with adapters need to be excluded. The inventors have found that by excluding sequence reads contaminated with adapters, the quality of data such as raw data can be improved.

[0191] In an eighth aspect, the present invention is a method for improving the quality of an (RNA) dataset, comprising: (i) determining the sequence of (target) RNA molecules in a sample (in particular, alternative molecules such as cDNA molecules derived therefrom); (ii) determining 5' and / or 3' end adducts to the (target) RNA molecule that are not part of the (target) RNA molecule in its naturally occurring form; and (iii) excluding from the analysis of the (RNA) dataset / (deleting from the (RNA) dataset) (target) RNA molecules having 5' and / or 3' end adducts; The present invention relates to a method comprising the steps above.

[0192] In step (i) of the method above, the sequence of the RNA molecules (in particular, alternatives such as cDNA molecules derived therefrom) in the sample is determined. The sequence of the RNA molecules (in particular, alternatives such as cDNA molecules derived therefrom) in the sample may be determined by any method known to those skilled in the art. Known sequencing methods include, but are not limited to, Sanger sequencing, capillary electrophoresis and fragment analysis, or next-generation sequencing (NGS) methods. Preferably, the sequence of the RNA molecules (in particular, alternatives such as cDNA molecules derived therefrom) in the sample is determined by an NGS method.

[0193] Specifically, the determination of the sequence of the RNA molecules in the sample, in particular by next-generation sequencing methods, comprises: - denaturation of the (target) RNA molecule and ligation of a 5' adapter and / or a 3' adapter to the denatured (target) RNA molecule using a double-stranded RNA ligase such as, for example, T4 RNA ligase 2 (Rnl2) or Kod1 ligase; - reverse transcription of the RNA molecule (to which the 5' adapter and / or 3' adapter is ligated and which is also referred to as the ligation product) into a cDNA molecule using a reverse transcriptase (RT) such as, for example, Maxima H-RT or Tth polymerase; - Amplification of the (said) cDNA molecule, for example by polymerase chain reaction (PCR), and / or - Sequencing of the (said) cDNA molecule, particularly next-generation sequencing is included.

[0194] PCR is selected from the group consisting of real-time PCR (quantitative PCR or qPCR), preferably TaqMan qPCR, multiplex PCR, nested PCR, high-fidelity PCR, fast PCR, hot-start PCR, and GC-rich PCR.

[0195] In step (ii) of the above method, the 5'-end and / or 3'-end adducts to the (target) RNA molecule that are not part of the (target) RNA molecule in its naturally occurring form are determined. Naturally occurring RNA molecules are in the form that occurs in nature or the natural environment, for example, body fluids such as whole blood, or tissues. However, these RNA molecules may be further processed into their alternatives, such as cDNA molecules. In this case, they are of natural origin. Naturally derived RNA molecules have an endogenous origin in biological materials / samples. The said RNA molecules were originally part of the said biological materials / samples.

[0196] The 5'-end and / or 3'-end adducts preferably have a length of at least 5 nucleotides. More preferably, the 5'-end and / or 3'-end adducts are 5 nucleotides to 30 nucleotides, even more preferably 5 nucleotides to 20 nucleotides, even even more preferably 7 nucleotides to 15 nucleotides, for example, having a length of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. The 5'-end adduct is called a prefix, the 3'-end adduct is called a suffix, and the adducts at both the 5'-end and 3'-end can also be called an affix. Specifically, at least 5 nucleotides extend beyond the length of the natural RNA molecule.

[0197] Prefixes, suffixes, and affixes may be 1, 2, 3, or 4 nucleotides in length. However, a minimum length of 5 nucleotides is required / recommended herein to reduce the risk of misidentifying (target) RNA molecules that have adducts at the 5' end and / or 3' end.

[0198] Adducts at the 5' end and / or 3' end of the (target) RNA molecules contained in the sample can be determined by sequencing and comparing the 5' end / 3' end of the sequence with the sequence of the adduct. Adducts at the 5' end and / or 3' end are the result of RNA molecule / adapter fusions, RNA molecule / RNA molecule fusions, or incomplete adapter processing / removal.

[0199] An adapter may be a non-biological RNA / DNA sequence that is intentionally added to the 5' end or 3' end of a biological (target) RNA / cDNA molecule (derived from the sample to be sequenced) as part of the design of sequencing methods such as NGS methods. This also includes free adapter molecules that may fuse with biological RNA in an unintended way. An adapter may also be a combination / fusion of multiple individual adapters and index RNAs. Index RNAs are used as UMIs (unique molecular identifiers) to assign RNA molecules to their source samples during multiplexed sequencing, i.e., simultaneous sequencing of multiple samples. The UMI sequence may vary in length, for example, a short non-specific (random) sequence of a predetermined length of 12 nucleotides.

[0200] The adducts at the 5'-end and / or 3'-end are preferably CGATC (SEQ ID NO: 10), GGGGC (SEQ ID NO: 11), ACGATC (SEQ ID NO: 12), GGGCGT (SEQ ID NO: 13), CGGCGG (SEQ ID NO: 14), GGGGC G (SEQ ID NO: 15), GACGATC (SEQ ID NO: 16), GGGCGT (SEQ ID NO: 17), GGGCGTG (SEQ ID NO: 18), GGGGGCG (SEQ ID NO: 19), GGGGG TG (SEQ ID NO: 20), GGGGGCGTG (SEQ ID NO: 21), CGGGGCGG (SEQ ID NO: 22), GGGAGGCC (SEQ ID NO: 23), GGAGGCGT (SEQ ID NO: 24), GGGCGTGG (SEQ ID NO: 25), TGGAGGCG (SEQ ID NO: 26), CGACGATC (SEQ ID NO: 27), GGGGGCGTT (SEQ ID NO: 28), GGGCGTGT (SEQ ID NO: 29), GGGGGCGT (SEQ ID NO: 30), GGGAGCCA (SEQ ID NO: 31), GGGGGGTGT (SEQ ID NO: 32), GGAGGCCC (SEQ ID NO: 33), CCGACGATC (SEQ ID NO: 34), GGGGGCGTG (SEQ ID NO: 35), TACCTGGTT (SEQ ID NO: 36), TGGAGGCGT (SEQ ID NO: 37), GGGCGTGGG (SEQ ID NO: 38), CGGCGGCGG (SEQ ID NO: 39), GGGGGGTGTA (SEQ ID NO: 40), GGGGGCGTT (SEQ ID NO: 41), GGCTGGGCG (SEQ ID NO: 42), TCGGGGGCGG (SEQ ID NO: 43), GGGGGCGTGG (SEQ ID NO: 44), GGGGGAGCCA (SEQ ID NO: 45), GGGAGGCCC (SEQ ID NO: 46), CGGAGGGCGG (SEQ ID NO: 47), GTCCGCGATC (SEQ ID NO: 48), GTCGACGATC (SEQ ID NO: 49), CGGGCGGATC (SEQ ID NO: 50), TGGAGGCGTG (SEQ ID NO: 51), TCCGACGATC (SEQ ID NO: 52), GGGGGCGTGGG (SEQ ID NO: 53), AAGCGGGGGCT (SEQ ID NO: 54), CGGGGA GCCA (SEQ ID NO: 55), GTCCGACGATC (SEQ ID NO: 56), TCGGAGGGCGG (SEQ ID NO: 57), AGTCCGACGATC (SEQ ID NO: 58), AAGCGGGGGCTGG (SEQ ID NO: 59), GTCCGACGGATC (SEQ ID NO: 60), TCGGGGCTGGGGC (SEQ ID NO: 61), TACCTGGTTGAT (SEQ ID NO: 62), TCGGGGGCGGCGG (SEQ ID NO: 63)It is selected from the group consisting of additives having nucleotide sequences according to CAGTCCGACGATC (SEQ ID NO: 64), TACCTGGTTGATC (SEQ ID NO: 65), TCGGGCTGGGGCG (SEQ ID NO: 66), TGGAGGCGTGGGT (SEQ ID NO: 67), ACAGTCCGACGATC (SEQ ID NO: 68), GGTCGGGCTGGGGC (SEQ ID NO: 69), CGGAAGCGTGCTGGG (SEQ ID NO: 70), GGTCGGGCTGGGGCG (SEQ ID NO: 71), TACAGTCCGACGATC (SEQ ID NO: 72), CGGAAGCGTGCTGGGC (SEQ ID NO: 73), CTACAGTCCGACGATC (SEQ ID NO: 74), TCTACAGTCCGACGATC (SEQ ID NO: 75), CGGAAGCGTGCTGGGCCC (SEQ ID NO: 76), TCGGGGCGGCGGCGGCGG (SEQ ID NO: 77), TTCTACAGTCCGACGATC (SEQ ID NO: 78), TAGCAGCACATCATGGTT (SEQ ID NO: 79), GGATCATTA (SEQ ID NO: 80), GGGGC GTGGG (SEQ ID NO: 81) and TGGAGGCGTGGGT (SEQ ID NO: 82).

[0201] When the 5'-end and / or 3'-end adducts are detected or identified in the (target) RNA molecule or its substitute, the (target) RNA molecule or its substitute is excluded from the (RNA) dataset analysis in step (iii) of the above method, or the (target) RNA molecule or its substitute is removed from the (RNA) dataset in step (iii) of the above method.

[0202] Specifically, determination of 5'-end and / or 3'-end adducts to the (target) RNA molecule includes analysis of whether at least a partial sequence of the adapter sequence used in sequencing, preferably next-generation sequencing, is identical to the (target) RNA molecule.

[0203] More specifically, the adapter sequences used in the next-generation sequencing process are selected from the group consisting of TGGAATTCTCGGGTGCCAAGG (SEQ ID NO: 83), GTTCAGAGTTCTACAGTCCGACGATC (SEQ ID NO: 84), TGGAATTCTCGGGTGCCAAGG (SEQ ID NO: 85), GAATTCCACCACGTTCCCGTGG (SEQ ID NO: 86), AATGATACGGCGACCACCGAGATCTACACGTTCAGAGTTCTACAGTCCGA (SEQ ID NO: 87), CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 88), GTGACTGGAGTTCCTTGGCACCCGAGAATTCCA (SEQ ID NO: 89), CAAGCAGAAGACGGCATACGA (SEQ ID NO: 90), GTTCAGAGTTCTACAGTCCGACGATC (SEQ ID NO: 91), TCGTATGCCGTCTTCTGCTTGT (SEQ ID NO: 92), ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 93), CAAGCAGAAGACGGCATACGA (SEQ ID NO: 94), AATGATACGGCGACCACCGACAGGTTCAGAGTTCTACAGTCCGA (SEQ ID NO: 95), CGACAGGTTCAGAGTTCTACAGTCCGACGATC (SEQ ID NO: 96), AAAAAAAAAAAAAAAAAAAAAAAAAAAAA (SEQ ID NO: 97), AGATCGGAAGAGCACACGTCTGAACTCCAGTCA (SEQ ID NO: 98), AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT (SEQ ID NO: 99), GATCGGAAGAGCACACGTCTGAACTCCAGTCAC (SEQ ID NO: 100), ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 101), AATGATACGGCGACCACCGAGATCTACAC (SEQ ID NO: 102) and ACACTCTTTCCCTACACGACGCTCTTCCGATCT (SEQ ID NO: 103).

[0204] If the adapter array or a part thereof is detected or identified in the (target) RNA molecule or its substitute, in step (iii) of the above method, the (target) RNA molecule or its substitute is excluded from the (RNA) dataset analysis, or in step (iii) of the above method, the (target) RNA molecule or its substitute is removed from the (RNA) dataset. In particular, the RNA dataset analysis is RNA raw dataset analysis, and the RNA dataset is an RNA raw dataset.

[0205] Preferably, the (target) RNA molecule is a small RNA molecule. More preferably, the small RNA molecule is a non-coding small RNA molecule. Even more preferably, the non-coding small RNA molecule is a miRNA molecule and / or an isomiR molecule.

[0206] The sample used in the above method may be a processed sample. Preferably, the sample used in the above method is a lysed sample, an extracted sample, an amplified sample and / or a sequenced sample. Lysis of the sample involves lysis of the cells contained in the sample and is necessary to release the (target) RNA molecule contained therein. Extraction of the sample is necessary to extract the (target) RNA molecule contained in the sample. Amplification of the sample is necessary to amplify the (target) RNA molecule contained in the sample. Sequencing of the sample is necessary to sequence the (target) RNA molecule contained in the sample.

[0207] For example, a biological sample such as a whole blood sample is provided. In the first step, the sample, specifically the cells contained in the sample, are lysed and the RNA molecules contained in the cells are released. Then, cDNA is generated from the (target) RNA molecule using an adapter bound to the (target) RNA molecule denatured by reverse transcription. Then, the cDNA molecules are amplified, and finally, sequencing such as next-generation sequencing of the cDNA molecules becomes possible.

[0208] In any of the above processes, (target) RNA molecules may have adducts added to their 5'-ends and / or 3'-ends, and (target) RNA molecules that need to be identified must be excluded from the analysis of (RNA) datasets having these 5'-end and / or 3'-end adducts, or must be removed from the (RNA) dataset.

[0209] The sample contains or is derived from a biological material. Preferably, the biological material is a tissue or a body fluid. More preferably, the body fluid is blood. Even more preferably, the blood is whole blood or a blood fraction. In particular, the blood fraction is selected from the group consisting of a blood cell fraction and plasma or serum. The blood cell fraction represents the cellular part of (whole) blood. Plasma and serum represent the extracellular part of (whole) blood. More specifically, the blood cell fraction contains red blood cells, white blood cells or platelets, the blood cell fraction is a fraction of red blood cells, white blood cells or platelets, or the blood cell fraction is a mixture of red blood cells, white blood cells and platelets.

[0210] Those skilled in the art will be able to understand various modifications and variations without departing from the scope of the present invention. Although the present invention has been described in connection with specific preferred embodiments, it should be understood that the present invention should not be unduly limited to such specific embodiments as set forth in the claims. Indeed, all various modifications of the described embodiments for carrying out the present invention that are obvious to those skilled in the relevant art are intended to be within the scope of the present invention.

Brief Description of the Drawings

[0211] The following figures are for illustrative purposes only and should not be construed as limiting the scope of the present specification as set forth by the appended claims.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

[0212] <Example> The examples given below are for illustrative purposes and do not limit the invention described above. (Example 1) To overcome the aforementioned considerations when using a single spike-in such as cel-miR-39-3p, a universal spike-in system that widely reflects the behavior of endogenous miRNAs during RNA extraction and detection was designed and optimized.

[0213] Specifically, 20 random sequences of 21 nucleotide lengths (reflecting the typical length of microRNAs) were designed (not shown). Next, their molecular properties, such as melting temperature (Tm °C) and GC%, were evaluated. Then, based on several criteria, a shortlist of 10 artificial sequences was selected for wet-lab verification. The aim was to minimize primer dimer formation and select RNA molecules with relatively weak secondary structures. Furthermore, the RNA molecules must widely reflect the range of endogenous GC content. Therefore, spike-ins with a GC content of 38.1% - 61.9%, which cover most of the GC content of endogenous microRNAs, were selected.

[0214] The artificial sequences remaining as final candidates were used for the ordering of RNA synthesis containing 5'-phosphate groups (IDT, Newark, Virginia, USA). The miRcury LNA assay (QIAgen, Venlo, Netherlands) was ordered and used for the quantification of spike-ins using semi-quantitative reverse transcription PCR (qRT-PCR).

[0215] First, a 10-fold dilution series was prepared individually for each molecule for each spike-in, used in the miRucry assay, and measured on a Quantstudio Flex 6 (ThermoFisher Scientific, Waltham, Massachusetts, USA). The cycle threshold (Ct value) was plotted against the six dilution rates used. Figure 1 shows the data for four sequences (SEQ ID NO: 1 to SEQ ID NO: 4).

[0216] The spike-in of SEQ ID NO: 1 has the following nucleotide sequence: GAUAGAUACGCCAGUACCGCC, The spike-in of SEQ ID NO: 2 has the following nucleotide sequence: AACGAAGCUCCACGAUGUAGG, The spike-in of SEQ ID NO: 3 has the following nucleotide sequence: UGUACGGAAAUAUUGGCUACC, and, The spike-in of SEQ ID NO: 4 has the following nucleotide sequence: UUCAUACGUUGCCCAAUCCAG.

[0217] Data for other sequences are not shown. Based on the linear regression fitting to the curves of the plots shown in Figure 1, the gradient, R-squared, and primer efficiency were calculated. These are summarized in Table 1 below:

[0218]

Table 1

[0219] All four test spike-in arrays showed the expected gradient (close to 3), high R-squared (above 0.99), and very good primer efficiency (90 - 100%). Each specific miRcury LNA assay was used in qPCR tests with non-correlated spike-in RNA and biological PAXgene RNA samples. No amplification was observed (data not shown), suggesting that the miRcury LNA assay is specific only to the correlated RNA.

[0220] Next, each spike-in was prepared at a two-fold dilution and added to the PAXgene RNA sample, and the miRcury qPCR assay was performed. Figure 2 summarizes the data of this test. The coefficient of determination of each spike-in RNA was calculated from the data shown in Figure 2 and is shown in Table 2. Overall, a coefficient of determination of 0.99 or close to it was observed for all spike-in RNAs.

[0221]

Table 2

[0222] Next, test mixtures of the four spike-ins were prepared as RNA at the following concentrations with nuclease-free water. SEQ ID NO: 1 - 340 pM, SEQ ID NO: 2 - 72.5 pM, SEQ ID NO: 3 - 2 pM, SEQ ID NO: 4 - 0.7 pM.

[0223] 10 μl of this test mixture was added to the PAXgene lysis solution immediately before extraction. The QIAsymphony PAXgene extraction kit (QIAgen, Venlo, Netherlands) was used for PAXgene extraction and eluted with 200 μl of elution buffer. This RNA was used for expression measurement by qPCR. The results were summarized in Table 3, and a comparison of the Ct values of the predicted and measured values was performed. Overall, the Ct value of the measured value was lower than the predicted value, indicating good extraction efficiency.

[0224]

Table 3

[0225] In the next step, the PAXgene RNA samples described in Table 3 were used to prepare next-generation sequencing libraries using the miRNA QIAseq NGS Library Kit (Qiagen, Venlo, the Netherlands) according to the manufacturer's recommendations. Immediately after the cDNA purification step, the cDNA was diluted 1:10 and used in qPCR reactions using spike-in specific forward primers and a common reverse transcription primer. Specifically, the following primers were used. SpikeIn_1_forward: GATAGATACGCCAGTACCGCC (SEQ ID NO: 5), SpikeIn_2_forward: AACGAAGCTCCACGATGTAGG (SEQ ID NO: 6), SpikeIn_3_forward: TGTACGGAAATATTGGCTACC (SEQ ID NO: 7), SpikeIn_4_forward: TTCATACGTTGCCCAATCCAG (SEQ ID NO: 8), and, SpikeIn_reverse: ATGCATGCATGCATTGATGGTGCCTACAGTT (SEQ ID NO: 9).

[0226] The results are shown in Table 4, and the R-squared of the spike-in expression measured by qPCR in the NGS library is shown in Figure 3.

[0227]

Table 4

[0228] As a conclusion, acceptable linearity (0.95 or higher) measured with cDNA from the library preparation process was achieved. Finally, the cDNA samples from the previous inspection were used for library PCR and NGS. After preprocessing the data, the raw read counts for each spike-in (Table 5) were assigned. Even at the level of NGS read counts, it was observed that the spike-ins were expressed in the expected linear order (Figure 4).

[0229]

Table 5

[0230] (Example 2) The above artificial spike-in RNA (having a nucleotide sequence according to SEQ ID NOs: 1 to 4) was further used to identify artificial target RNA generated by next-generation sequencing (NGS) technology. Specifically, target RNAs with 5'-terminal and / or 3'-terminal adducts attached were identified. During the NGS process, it was found that parts of other RNAs were attached to both ends (5' and 3') of the spike-in RNA. These parts are herein referred to as affixes and include both a prefix (5') and a suffix (3'). It was also found that the adapter sequences of the adapters used in the NGS process were attached to the spike-in RNA.

[0231] The inventor's internal (HBDX) lung disease dataset (blood samples) and the external TCGA dataset (tissue samples of cancer patients) were used. The HBDX dataset contains spike-ins and was used for the identification of affixes. Even in TCGA, similar affixes were found (at a low rate) despite no spike-ins being added. Some affixes may bind to the 5'-adapter sequence GTTCAGAGTTCTACAGTCCGACGATC (SEQ ID NO: 84) commonly used during NGS of small RNAs. Refer to Figure 5 in this regard.

[0232] Identification of the affixes added to the spike-in RNA enables the identification of target RNAs that are likely to be artificial / modified, i.e., they were generated by the same "affix addition" mechanism as the spike-in RNA itself. Since these target RNAs can cause misunderstandings during data analysis, they should be removed to improve the quality of the dataset and subsequent analysis.

[0233] To prevent the incorrect removal of target RNAs, only target RNAs with a minimum affix length of 5 nucleotides or more were removed from the subsequent analysis / dataset. Figure 6 shows the distribution of spike-in RNAs with nucleotide sequences according to SEQ ID NO: 1 to SEQ ID NO: 4 and the affix length. The higher the expression of the spike-in, the more the number of discovered affixes.

[0234] Figure 7 shows the average expression of spike-in RNA (length: 21 nucleotides) and longer artifacts (spike-in and affixes with increasing length). Figure 8 shows the number of sequences removed from two datasets (HBDX, TCGA) when considering only the minimum length (or longer) affixes. The theoretical guideline is given as the proportion of accidental matches in all possible nucleotide sequences of the same length as the affix. The lower this theoretical error rate, the longer the minimum length must be selected. However, when the minimum length is 5, the error rate drops below 0.5% and may be acceptable in most settings.

Claims

**Claim 1** A composition comprising at least three RNA molecules, wherein at least three of said RNA molecules are selected from the group consisting of RNA molecules having a nucleotide sequence according to SEQ ID NO: 1 to SEQ ID NO: 4, fragments thereof, and sequences having at least 80% sequence identity thereto. **Claim 2** The composition according to claim 1, comprising four of said RNA molecules, wherein the four RNA molecules are selected from the group consisting of RNA molecules having a nucleotide sequence according to SEQ ID NO: 1 to SEQ ID NO: 4, fragments thereof, and sequences having at least 80% sequence identity thereto. **Claim 3** (i) SEQ ID NO: 1, SEQ ID NO: 2, and SEQ ID NO: 3, (ii) SEQ ID NO: 2, SEQ ID NO: 3, and SEQ ID NO: 4, (iii) SEQ ID NO: 1, SEQ ID NO: 3, and SEQ ID NO: 4, (iv) SEQ ID NO: 1, SEQ ID NO: 2, and SEQ ID NO: 4, or (v) SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, and SEQ ID NO: 4 The composition according to claim 1 or claim 2, comprising an RNA molecule having a nucleotide sequence according to. **Claim 4** The composition according to any one of claims 1 to 3, wherein at least three of said RNA molecules have a characteristic distribution. **Claim 5** The composition according to claim 4, wherein at least three of said RNA molecules have a characteristic distribution with respect to their amounts. **Claim 6** The composition according to claim 5, wherein at least three of said RNA molecules are included in different amounts. **Claim 7** The composition according to claim 6, wherein any pair of two of said RNA molecules comprised in said composition have different amounts. **Claim 8** The composition according to claim 6 or claim 7, wherein at least three of said RNA molecules are included in the composition in a defined gradient of amounts. **Claim 9** The composition according to any one of claims 1 to 8, wherein at least three of said RNA molecules are present in an amount of 0.001 amol to 6000 amol, preferably 0.01 amol to 5000 amol, more preferably 0.1 amol to 4000 amol, still more preferably 1 amol to 3500 amol. **Claim 10** The composition according to claim 9, wherein the four RNA molecules are present in an amount of 0.001 amol to 6000 amol, preferably 0.01 amol to 5000 amol, more preferably 0.1 amol to 4000 amol, still more preferably 1 amol to 3500 amol. **Claim 11** Comprising at least three of said RNA molecules, The first RNA molecule is included in an amount of about 3400 amol, the second RNA molecule is included in an amount of about 725 amol, the third RNA molecule is included in an amount of about 20 amol, or The composition according to any one of claims 1 to 10, wherein the first RNA molecule is included in an amount of about 1360 amol, the second RNA molecule is included in an amount of about 290 amol, and the third RNA molecule is included in an amount of about 80 amol. **Claim 12** comprising four of said RNA molecules, The first RNA molecule is included in an amount of about 3400 amol, the second RNA molecule is included in an amount of about 725 amol, the third RNA molecule is included in an amount of about 20 amol, and the fourth RNA molecule is included in an amount of about 7 amol, or The composition according to claim 11, wherein the first RNA molecule is included in an amount of about 1360 amol, the second RNA molecule is included in an amount of about 290 amol, the third RNA molecule is included in an amount of about 80 amol, and the fourth RNA molecule is included in an amount of about 27 amol. **Claim 13** The composition according to any one of claims 1 to 12, wherein the RNA molecule is an artificial RNA molecule (not existing in nature). **Claim 14** The composition according to any one of claims 1 to 13, which is suitable as a spike-in cocktail or is a spike-in cocktail. **Claim 15** The composition according to any one of claims 1 to 14, which is suitable as a standard for process control and / or normalization or is a standard. **Claim 16** The composition according to any one of claims 1 to 15, which is a solution. **Claim 17** The composition according to claim 16, wherein the solution is an aqueous solution. **Claim 18** The composition according to claim 17, wherein the aqueous solution is water. **Claim 19** The composition according to claim 18, wherein the water is nuclease-free water. **Claim 20** A kit comprising the composition according to any one of claims 1 to 19. **Claim 21** The kit according to claim 20, further comprising means for determining the levels of at least three of the RNA molecules comprised in the composition. **Claim 22** Use of the composition according to any one of claims 1 to 19, or the kit according to claim 20 or claim 21, as a (standard) for process control, sample inspection, normalization and / or data processing control. **Claim 23** Use of the kit according to claim 22, wherein the process control includes quality control or end-to-end control.

24. Use of the kit according to claim 22 or claim 23, wherein the data processing control includes raw data processing control.

25. A method for inspecting a sample, comprising: evaluating a sample with respect to at least three of the RNA molecules contained in or derived from the composition according to any one of claims 1 to 19.

26. The method according to claim 25, wherein the sample is a mixture of biological materials having the composition according to any one of claims 1 to 19.

27. The method according to claim 26, wherein the biological material is blood.

28. The method according to claim 27, wherein the blood is whole blood or a blood fraction.

29. The method according to claim 28, wherein the blood fraction is selected from the group consisting of a blood cell fraction and plasma or serum.

30. The method according to any one of claims 25 to 29, wherein the sample is based on or derived from a mixture of biological materials having the composition according to any one of claims 1 to 19.

31. The method according to any one of claims 25 to 30, wherein the sample is a processed sample.

32. The method according to claim 31, wherein the processed sample is a lysed sample, an extracted sample, an amplified sample, a sequenced sample, or a library-prepared sample.

33. The method according to claim 31 or claim 32, wherein the processed sample is obtained by mixing biological materials having the composition according to any one of claims 1 to 19 and subjecting the resulting mixture to further processing.

34. The method according to any one of claims 25 to 33, wherein the evaluation includes determining whether at least three of the RNA molecules exhibit a characteristic distribution.

35. The method according to claim 34, wherein when the characteristic distribution is given, the sample is further processed and / or analyzed.

36. The characteristic distribution is given when at least three of the RNA molecules are present at expected levels. when at least three of said RNA molecules are present in a predicted order / rank (defined by the relationship of the levels of at least three of said RNA molecules, particularly by amount), and / or The method according to claim 35, wherein at least three of said RNA molecules are present in a predicted linearity. **Claim 37** The method according to claim 36, wherein the predicted order has a Spearman's rank correlation coefficient (Spearman ρ) of 0.95 or more, and / or the predicted linearity has a Pearson's correlation coefficient (Pearson r) of 0.66 or more. **Claim 38** The method according to claim 34, wherein when the characteristic distribution is not given, the sample is not further processed and is discarded. **Claim 39** when at least three of said RNA molecules are not present at predicted levels, when at least three of said RNA molecules are not present in a predicted order / rank (defined by the relationship of the levels of at least three of said RNA molecules, particularly by amount), and / or The method according to claim 38, wherein when at least three of said RNA molecules are not present in a predicted linearity, the characteristic distribution is not given. **Claim 40** The method according to claim 39, wherein the predicted order has a Spearman's rank correlation coefficient (Spearman ρ) of 0.95 or more, and / or the predicted linearity has a Pearson's correlation coefficient (Pearson r) of 0.66 or more. **Claim 41** The method according to any one of claims 33 to 40, wherein the further processing includes lysing the sample, extracting the sample, amplifying the sample, sequencing the sample, and / or preparing a library from the sample. **Claim 42** The method according to claim 41, wherein the further processing includes lysing the cell to release the nucleotide sequence contained in the sample, extracting the nucleotide sequence contained in the sample, amplifying the nucleotide sequence contained in the sample, sequencing the nucleotide sequence contained in the sample, and / or preparing a library from the nucleotide sequence contained in the sample. **Claim 43** The method according to claim 42, wherein the nucleotide sequence is a ribonucleotide sequence. **Claim 44** The method according to claim 43, wherein the ribonucleotide sequence belongs to a target RNA molecule.

45. The method according to claim 44, wherein the target RNA molecule is a small RNA molecule.

46. The method according to claim 45, wherein the small RNA molecule is a non-coding small RNA molecule, preferably a miRNA molecule.

47. The method according to any one of claims 25 to 46, wherein the evaluation comprises identifying 5' and / or 3' end adducts of at least three of the RNA molecules comprised in or derived from the composition according to any one of claims 1 to 19.

48. The method according to claim 47, wherein the 5' and / or 3' end adduct has a length of at least 5 nucleotides.

49. The method according to claim 48, wherein at least the 5 nucleotides extend beyond the original length of at least three of the RNA molecules.

50. The method according to any one of claims 47 to 49, wherein the 5' and / or 3' end adduct is the result of an RNA molecule / adapter fusion, an RNA molecule / RNA molecule fusion, or an adapter / adapter fusion.

51. The method according to any one of claims 25 to 50, wherein the sample comprises a target RNA molecule.

52. The method according to claim 51, wherein the presence of the 5' and / or 3' end adducts identified for at least three of the RNA molecules comprised in or derived from the composition according to any one of claims 1 to 19 indicates the presence of 5' and / or 3' end adducts in the target RNA molecule.

53. The method according to claim 51 or 52, wherein the target RNA molecule having the 5' and / or 3' end adduct comprised in the sample is excluded from further analysis / not subjected to further use.

54. The method according to any one of claims 51 to 53, wherein data from the target RNA molecule having the 5' and / or 3' end adduct comprised in the sample is excluded from further analysis / not subjected to further use, or does not form part of a data set, preferably a raw data set.

55. The method according to any one of claims 51 to 54, wherein the target RNA molecule is a small RNA molecule.

56. The method according to claim 55, wherein the small RNA molecule is a non-coding small RNA molecule, preferably a miRNA molecule.

57. A method for optimizing a biological sample, comprising: The step of implementing the method according to any one of claims 25 to 56.

58. A method for optimized RNA preparation from a biological sample or for RNA analysis of the biological sample, comprising: The step of implementing the method according to any one of claims 25 to 56.

59. A method for improving the quality of an RNA dataset, comprising: The step of implementing the method according to any one of claims 25 to 56.

60. A method for improving the quality of an RNA dataset, comprising: (i) determining the sequence of RNA molecules in a sample; (ii) determining 5' and / or 3' end additions to the RNA molecules that are not part of the RNA molecules in their naturally occurring form; and (iii) excluding from the analysis of the RNA dataset RNA molecules having the 5' and / or 3' end additions / removing from the RNA dataset RNA molecules having the 5' and / or 3' end additions.

61. The determination of the sequence of the RNA molecules in the sample is denaturation of the RNA molecules and binding of a 5' adapter and / or a 3' adapter to the denatured RNA molecules; reverse transcription of the RNA molecules (to which the 5' adapter and / or the 3' adapter are bound) into cDNA molecules; amplification of the (said) cDNA molecules; and / or sequencing of the (said) cDNA molecules, preferably next-generation sequencing, according to claim 60.

62. The method according to claim 60 or claim 61, wherein the 5' and / or 3' end additions have a length of at least 5 nucleotides.

63. The method according to claim 62, wherein at least the 5 nucleotides extend beyond the length of the RNA molecules in their naturally occurring form.

64. The method according to any one of claims 60 to 63, wherein the 5'-end and / or the 3'-end adduct is the result of an RNA molecule / adapter fusion, an RNA molecule / RNA molecule fusion, or an adapter / adapter fusion.

65. The 5'-end and / or the 3'-end adduct is CGATC (SEQ ID NO: 10), GGGGC (SEQ ID NO: 11), ACGATC (SEQ ID NO: 12), GGGCGT (SEQ ID NO: 13), CGGCGG (SEQ ID NO: 14), GGGCGG (SEQ ID NO: 15), GACGATC (SEQ ID NO: 16), GGGCGT (SEQ ID NO: 17), GGGCGTG (SEQ ID NO: 18), GGGGGCG (SEQ ID NO: 19), GGGGGTG (SEQ ID NO: 20), GGGGCGTG (SEQ ID NO: 21), CGGGGGCGG (SEQ ID NO: 22), GGGA GCC (SEQ ID NO: 23), GGAGGC GT (SEQ ID NO: 24), GGGCGTGG (SEQ ID NO: 25), TGGAGGCG (SEQ ID NO: 26), CGACGATC (SEQ ID NO: 27), GGGC GTT (SEQ ID NO: 29), GGGGGCGT (SEQ ID NO: 30), GGGA GCC A (SEQ ID NO: 31), GGGGGTGT (SEQ ID NO: 32), GGAG GCC C (SEQ ID NO: 33), CCGACGATC (SEQ ID NO: 34), GGGGGCGTG (SEQ ID NO: 35), TACCTG GTT (SEQ ID NO: 36), TGGAGGCGT (SEQ ID NO: 37), GGGCGTGGG (SEQ ID NO: 38), CGGCGGCGG (SEQ ID NO: 39), GGGGGT GTA (SEQ ID NO: 40), GGGGGC GT T (SEQ ID NO: 41), GGCTGG GCG (SEQ ID NO: 42), TCGGGGGCGG (SEQ ID NO: 43), GGGGCGTGG (SEQ ID NO: 44), GGGGAGCC A (SEQ ID NO: 45), GGAG GCC C (SEQ ID NO: 46), CGGAGGGGCGG (SEQ ID NO: 47), GTCCGCGATC (SEQ ID NO: 48), GT CGACGATC (SEQ ID NO: 49), CGGGC GGA TC (SEQ ID NO: 50), TGGAGGCGTG (SEQ ID NO: 51), TCCGACGATC (SEQ ID NO: 52), GGGGGCGTGGG (SEQ ID NO: 53), AAGCGGGGGCT (SEQ ID NO: 54), CGGGGAGCC A (SEQ ID NO: 55), GTCCGACGATC (SEQ ID NO: 56), TCGGAGGGGCGG (SEQ ID NO: 57), AGTC CGACGATC (SEQ ID NO: 58), AAGCGGGGGCTGG (SEQ ID NO: 59), GTCCGACGGA TC (SEQ ID NO: 60), TCGGGCTGGGGGC (SEQ ID NO: 61), TACCTG GTTGAT (SEQ ID NO: 62), TCGGGGGCGGCGG (SEQ ID NO: 63), CAGTC CGACGATC (SEQ ID NO: 64), TACCTG GTTGATC (SEQ ID NO: 65)The method according to any one of claims 60 to 64, selected from the group consisting of adducts having a nucleotide sequence according to TCGGGCTGGGGCG (SEQ ID NO: 66), TGGAGGCGTGGT (SEQ ID NO: 67), ACAGTCCGACGATC (SEQ ID NO: 68), GGTCGGGCTGGGG (SEQ ID NO: 69), CGGAAGCGTGCTGGG (SEQ ID NO: 70), GGTCGGGCTGGGGCG (SEQ ID NO: 71), TACAGTCCGACGATC (SEQ ID NO: 72), CGGAAGCGTGCTGGGC (SEQ ID NO: 73), CTACAGTCCGACGATC (SEQ ID NO: 74), TCTACAGTCCGACGATC (SEQ ID NO: 75), CGGAAGCGTGCTGGGGCCC (SEQ ID NO: 76), TCGGGGGCGGCGGCGGCGGG (SEQ ID NO: 77), TTCTACAGTCCGACGATC (SEQ ID NO: 78), TAGCAGCACATCATGGTT (SEQ ID NO: 79), GGATCATTTA (SEQ ID NO: 80), GGGGCGTGGT (SEQ ID NO: 81), TGGAGGCGTGGT (SEQ ID NO: 82).

66. The determination of the 5'-end and / or the 3'-end adduct to the RNA molecule includes an analysis of whether the RNA molecule is at least partially sequence identical to the adapter sequence used in the sequencing, preferably the next-generation sequencing process, according to any one of claims 60 to 65.

67. The adapter sequence used in the next-generation sequencing process is The method according to claim 66, selected from the group consisting of: TGG AAT TCT CGG GTG CCA AGG (SEQ ID NO: 83), GTT CAG AGT TCT ACA GTC CGA CGA TC (SEQ ID NO: 84), TGG AAT TCT CGG GTG CCA AGG (SEQ ID NO: 85), GAA TTC CAC CAC GTT CCC GTG G (SEQ ID NO: 86), AAT GAT ACG GCG ACC ACC GAG ATC TAC ACG TTC AGA GTT CTA CAG TCC GA (SEQ ID NO: 87), CAA GCA GAA GAC GGC ATA CGA GAT (SEQ ID NO: 88), GTG ACT GGA GTT CCT TGG CAC CCG AGA ATT CA (SEQ ID NO: 89), CAA GCA GAA GAC GGC ATA CGA (SEQ ID NO: 90), GTT CAG AGT TCT ACA GTC CGA CGA TC (SEQ ID NO: 91), TCG TAT GCC GTC TTC TGC TTG T (SEQ ID NO: 92), ATC TCG TAT GCC GTC TTC TGC TTG (SEQ ID NO: 93), CAA GCA GAA GAC GGC ATA CGA (SEQ ID NO: 94), AAT GAT ACG GCG ACC ACC GAC AGG TTC AGA GTT CTA CAG TCC GA (SEQ ID NO: 95), CGA CAG GTT CAG AGT TCT ACA GTC CGA CGA TC (SEQ ID NO: 96), AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA AAAAA ​ ​

Citation Information

Patent Citations

  • Novel spike-in oligonucleotides for sequence data normalization

    JP2020519256A