Methods and products for biomarker identification

The challenge of RNA detection in blood samples is solved through long-read RNA sequencing and normalization processing techniques, achieving more uniform RNA distribution and higher detection sensitivity and specificity.

CN119998461APending Publication Date: 2025-05-13WOBAO GENOMICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380069963.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-21
Filing Date
2023-10-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art present challenges in the processing and detection of RNA in blood samples, including rapid degradation of RNA, active RNA degradation enzymes in the blood, and detection methods rely on highly expressed genes, resulting in a lack of detection breadth of the entire blood transcriptome.

Method used

Long read RNA sequencing technology is used, and normalized treatment is used to improve the uniformity of RNA in the sample, reducing dependence on highly expressed genes, thereby improving the sensitivity and specificity of sequencing.

Benefits of technology

It achieves a more uniform distribution of RNA in blood samples, improves the detection ability of disease biomarkers, reduces the requirements for data storage, and improves sequencing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998461A_ABST
    Figure CN119998461A_ABST
Patent Text Reader

Abstract

Methods and products for discovering disease biomarkers and diagnosing disease are provided. RNA or cDNA samples are typically dominated by sequences from highly expressed genes, which can negatively impact the analysis of the samples. The present invention provides methods and products for preparing a treated nucleic acid sample having a more uniform sequence distribution, which can then be analyzed to find and detect disease biomarkers. Methods of protecting RNA in a blood sample are also provided. These methods can be combined to further improve the ability to find and detect disease biomarkers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to methods, compositions, systems and kits for discovering disease biomarkers and diagnosing diseases. The present invention also relates to methods for processing blood samples for biomarker discovery and detection. Background Art

[0002] Disease screening can take many forms, and there has been a growing interest in molecular screening methods recently. For example, the detection of RNA specific to cancer in a simple blood test has the potential to provide an easily accepted and relatively non-invasive screening method, as well as prognostic biomarkers (Larson, MH et al. A comprehensive characterization of the cell-free transcriptome reveals tissue-and subtype-specific biomarkers for cancer detection. 1–11). This approach has also become affordable, allowing regular screening and aiding in early diagnosis. However, the handling and processing of samples for RNA sequencing is a challenge in the clinical setting. In addition, the development of sufficiently sensitive and specific methods has proven to be a challenge. As a result, there is currently no practical and accurate blood RNA detection platform for breast cancer or ovarian cancer diagnosis in the clinic.

[0003] With regard to blood processing, a major concern is maintaining the integrity of RNA. It is well known that RNA degrades rapidly. In addition, blood and other tissue types contain enzymes that actively degrade RNA. This results in a rapid decrease in both the quantity and quality of RNA in the collected blood. Therefore, during the process of collecting, storing and transporting blood, the RNA content is constantly degraded.

[0004] RNA expression is characterized by a wide range of expression levels in unique genes. There is a group of highly expressed genes, which are called housekeeping genes. These genes do not provide useful information, but they usually account for more than half of the total RNA in any given sample. In blood, this problem is even more serious, because globin RNA and ribosomal RNA usually account for more than 95% of the total RNA (Harrington, CA et al. RNA-Seq of human whole blood: Evaluation of globin RNA depletion on Ribo-Zero library method. Sci. Rep. 10, 1–12 (2020)). Therefore, a major challenge of current detection methods is that they rely on the depletion of these genes, or rely on a subset of targeted genes. These solutions lead to a lack of detection breadth for the entire blood transcriptome, and may explain why this method has not been successful.

[0005] Another question is how representative these data are of actual complexity and biological variation. Although there are several RNA detection assays, until recently, these methods only allow the detection of small fragments of each RNA molecule. This is problematic because each gene can be transcribed into many different RNA isoforms. This may be due to multiple combinations of transcription start sites, termination sites, and alternative splicing (Harrow, J. et al. GENCODE: producing a reference annotation for ENCODE. 7, 1–9 (2006)). These selective transcripts usually have different functions and are usually associated with specific cell and tissue types. As a result, detecting only small fragments of RNA does not tell the whole story. Summary of the invention

[0006] The inventors have developed a technique that works with long-read RNA sequencing in a manner that can provide better consistency, sensitivity, and specificity for RNA-based disease screening. RNA or cDNA samples are often dominated by sequences from highly expressed genes, which can negatively affect the analysis of the sample. The present invention provides methods and devices for preparing processed nucleic acid samples with a more uniform sequence distribution, which can then be analyzed to discover and / or detect disease biomarkers. Normalization improves the sequencing efficiency for transcript discovery and / or detection. First, genes and isoforms that are specific to the disease in question are more easily detected, and secondly, there is less redundancy in the data generated, reducing the requirements for data storage. Also provided is a method designed to protect RNA in blood samples. These methods can be combined to further improve the ability to discover and / or detect disease biomarkers.

[0007] It should be remembered that various aspects have been designed to be advantageously combined, and all such combinations are contemplated to be within the scope of the invention. It should also be understood that the selections described in relation to one area of ​​improvement will apply mutatis mutandis to other areas; e.g., sample type, disease, etc., as the case may be.

[0008] The present invention provides a method for discovering disease biomarkers, comprising: (i) providing a first cDNA sample and a second cDNA sample; (ii) normalizing the first cDNA sample and the second cDNA sample; (iii) sequencing the normalized first cDNA sample and the second cDNA sample; and (iv) comparing the sequencing outputs of the first cDNA sample and the second cDNA sample to discover disease biomarkers.

[0009] The first cDNA sample and the second cDNA sample may be from the same subject. In certain embodiments, the first cDNA sample is from the subject before disease treatment, and the second cDNA sample is from the same subject after disease treatment. In some specific embodiments, the first cDNA sample is from the subject before disease treatment, and the second cDNA sample is from the same subject after disease treatment begins and / or after disease treatment is completed.

[0010] The first cDNA sample and the second cDNA sample can be from different objects. In some specific embodiments, the first cDNA sample and the second cDNA sample are from objects suffering from the same disease of different grades or stages. In certain embodiments, the first cDNA sample and the second cDNA sample are from objects suffering from different diseases (e.g., different types of cancer). In certain embodiments, the first cDNA sample is from an object suffering from a disease, and the second cDNA sample is from an object not suffering from the disease.

[0011] Therefore, the present invention provides a method for discovering disease biomarkers, comprising: (i) providing a first cDNA sample from a subject suffering from a disease and a second cDNA sample from a subject not suffering from the disease; (ii) normalizing the first cDNA sample and the second cDNA sample; (iii) sequencing the normalized first cDNA sample and the second cDNA sample; and (iv) comparing the sequencing outputs of the first cDNA sample and the second cDNA sample to discover disease biomarkers.

[0012] Discovering disease biomarkers means identifying new biomarkers (indicators of biological states) for a particular disease, such as discovering previously unknown biomarkers for development into tests for the disease. Disease biomarkers may be useful for diagnosing a disease, characterizing a disease, predicting a response to treatment, detecting minimal residual disease, and / or prognosing a disease. By characterization, it is meant the classification and evaluation of a disease. Prognosis refers to predicting the possible outcome of a subject's disease. In certain embodiments, characterization and / or prognosis of a disease includes determining the grade and / or stage of a disease. In other embodiments, characterization of a disease includes determining a subtype of a disease. Disease biomarkers may be useful for indicating the likelihood that a subject with a particular disease will benefit from a particular treatment.

[0013] In some specific embodiments, the disease is cancer. Therefore, in certain embodiments, the characterization and / or prognosis of cancer includes determining the presence or absence of metastasis. Metastasis or metastatic disease is the spread of cancer from one organ or part to another non-adjacent organ or part. The resulting new occurrence of the disease is called metastasis. The characterization and / or prognosis of the disease may also include predicting biochemical recurrence and / or determining whether the cancer is invasive and / or determining whether the cancer has spread to lymph nodes. Aggressiveness refers to a cancer that grows rapidly, is more likely to spread, is more likely to recur, and / or is resistant to treatment.

[0014] According to a related aspect of the invention, there is provided a method for monitoring a subject, comprising: (i) providing a first cDNA sample from a subject at a first time point, and providing a second cDNA sample from the subject at a second time point; (ii) normalizing the first cDNA sample and the second cDNA sample; (iii) sequencing the normalized first cDNA sample and the second cDNA sample; and (iv) comparing the sequencing outputs of the first cDNA sample and the second cDNA sample.

[0015] Monitoring objects may include monitoring the response to disease treatment. The first time point may be before starting treatment, and the second time point may be during or after treatment. Comparing the sequencing output of the first cDNA sample and the second cDNA sample can provide an indication of whether the treatment is successful. For example, the presence or absence of a disease biomarker can indicate whether the treatment is successful. Comparing the sequencing output of the first cDNA sample and the second cDNA sample can include comparing each other and / or comparing with the sequencing output from a reference sample.

[0016] According to all aspects of the present invention, the disease biomarker can be a cDNA sequence. The cDNA sequence will correspond to an RNA sequence. The cDNA / RNA sequence can correspond to a protein or peptide. In some specific embodiments, the method also includes identifying RNA, transcriptional models, genes, proteins and / or peptides corresponding to the cDNA sequence. Therefore, the disease biomarker can be a cDNA molecule (of a specific sequence), a DNA molecule (of a specific sequence), an RNA molecule (of a specific sequence), a transcriptional model, a protein or a peptide.

[0017] In some embodiments, the method comprises finding more than one disease biomarker, optionally more than 10, 100, 1000, 10000, 100000, 1 million or 10 million disease biomarkers. In other embodiments, the method comprises finding 1 to 10, 1 to 100, 1 to 1000, 1 to 10000, 1 to 100000, 1 to 1 million or 10 to 10 million disease biomarkers.

[0018] Disease biomarkers can be suitable targets for therapeutic agents (e.g., vaccines, RNA treatments, and / or gene editing). The discovery of disease biomarkers specific to cancer cells can be an initial step in identifying cancer-specific antigens targeted by cancer vaccines. Therefore, in some embodiments, the method further includes identifying transcripts or proteins corresponding to disease biomarkers as cancer vaccine targets. The method may also include developing a cancer vaccine for a target, optionally an RNA vaccine.

[0019] According to a related aspect of the present invention, there is provided a method for discovering a cancer vaccine target, comprising: (i) providing a first cDNA sample from a subject having cancer and a second cDNA sample from a subject not having cancer; (ii) normalizing the first cDNA sample and the second cDNA sample; (iii) sequencing the normalized first cDNA sample and the second cDNA sample; and (iv) comparing the sequencing outputs of the first cDNA sample and the second cDNA sample to discover cancer vaccine targets.

[0020] In some specific embodiments, the method comprises providing 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40 or more cDNA samples from different subjects with the disease and / or providing 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40 or more cDNA samples from different subjects without the disease. Preferably, the method comprises providing 30 or more cDNA samples from different subjects with the disease (e.g., with cancer) and providing 30 or more cDNA samples from different subjects without the disease (e.g., without cancer).

[0021] In some specific embodiments, the first cDNA sample and the second cDNA sample are obtained from a biological fluid or a fluid or lysate produced by a biological material. The first cDNA sample and the second cDNA sample can be obtained from blood. The first cDNA sample can be obtained by extracting RNA from a biological sample (e.g., blood) obtained from an object suffering from a disease, and the second cDNA sample can be obtained by extracting RNA from a biological sample (e.g., blood) obtained from an object not suffering from a disease. Then cDNA is synthesized using RNA as a template (i.e., by reverse transcription). Therefore, in some specific embodiments, before step (i), the method may also include: (a) extracting RNA from a biological fluid or a fluid or lysate produced from a biological material from a subject suffering from a disease and a subject not suffering from a disease; and (b) Using RNA as a template to synthesize cDNA (i.e. converting RNA into cDNA).

[0022] In this manner, a first cDNA sample and a second cDNA sample can be generated.

[0023] A subject with a disease means that when the biological sample (biological fluid or biological material) from which the cDNA sample is derived is obtained from the subject, the subject has the disease. A subject without a disease means that when the biological sample (biological fluid or biological material) from which the cDNA sample is derived is obtained from the subject, the subject does not have the disease.

[0024] A subject not suffering from a disease may be a healthy subject.

[0025] A disease biomarker may be a cDNA molecule (of a specific sequence), an RNA molecule (of a specific sequence), a protein or a peptide that is detectable in a sample from a subject with the disease but not in a sample from a subject without the disease. Alternatively, a disease biomarker may be a cDNA molecule (of a specific sequence), an RNA molecule (of a specific sequence), a protein or a peptide that is not detectable in a sample from a subject with the disease but detectable in a sample from a subject without the disease.

[0026] A cancer vaccine target may be a cDNA molecule (of a specific sequence), an RNA molecule (of a specific sequence), a protein or a peptide that is detectable in a sample from a subject having cancer but not in a sample from a subject not having cancer.

[0027] In some embodiments, a disease biomarker is a cDNA sequence that is present in a first cDNA sample but not in a second cDNA sample, or a cDNA sequence that is present in a second cDNA sample but not in the first cDNA sample. In some embodiments, a disease biomarker is a transcript that is (uniquely) present in subjects with a particular disease. In other embodiments, a disease biomarker is a transcript that is (uniquely) not present in subjects with a particular disease.

[0028] In some embodiments, the cancer vaccine target is a cDNA sequence that is present in the first cDNA sample but not in the second cDNA sample. In some embodiments, the cancer vaccine target is a transcript that is uniquely present in subjects with a specific cancer. The transcript / cDNA sequence may correspond to a specific protein or peptide. At least a portion of the protein or peptide may form an antigen included in the cancer vaccine.

[0029] The present invention is capable of identifying transcripts found only in subjects with a disease, optionally cancer. Such transcript models can be identified by comparison with subjects not suffering from the disease, such as benign patients, and optionally a public transcriptome annotation database. The sequencing output and / or resulting transcriptome profile from subjects suffering from a particular disease, such as breast cancer, can be compared with the sequencing output and / or resulting transcriptome profile from subjects suffering from a different disease, such as ovarian cancer and / or colorectal cancer, to determine whether the biomarker is unique to a particular disease, such as breast cancer.

[0030] In another aspect, the present invention provides a method for diagnosing a disease in a subject, comprising: (i) providing a cDNA sample from a subject; (ii) normalizing the cDNA samples; and (iii) sequencing the normalized cDNA sample, wherein the sequencing output is used to identify whether the subject has the disease.

[0031] Diagnosis means determination that the subject has the disease at the time of testing.

[0032] In certain embodiments, the disease is cancer, and diagnosing the disease includes detecting minimal residual disease (cancer cells remaining in the subject during or after treatment).

[0033] In a related aspect, a method for characterizing and / or prognosing a disease in a subject is provided, comprising: (i) providing a cDNA sample from a subject; (ii) normalizing the cDNA samples; and (iii) sequencing the normalized cDNA sample, wherein the sequencing output is used to provide a characterization and / or prognosis of the disease.

[0034] In another aspect, a method for selecting a treatment for a disease in a subject is provided, comprising: (i) providing a cDNA sample from a subject; (ii) normalizing cDNA samples; (iii) sequencing the normalized cDNA sample, wherein the sequencing output is used to provide a diagnosis, characterization and / or prognosis of the disease; and (iv) selecting a treatment appropriate for the diagnosis, characterization and / or prognosis of the disease.

[0035] In another aspect, a method for predicting responsiveness of a subject suffering from a disease to a therapeutic agent is provided, comprising: (i) providing a cDNA sample from a subject; (ii) normalizing the cDNA samples; and (iii) sequencing the normalized cDNA sample, wherein the sequencing output is used to predict the subject's responsiveness to the therapeutic agent.

[0036] The methods described herein may also include treating the subject.

[0037] The method may include comparing the sequencing output of the normalized cDNA sample with one or more reference sequences or with the sequencing output of one or more control samples, optionally wherein the one or more control samples are from one or more subjects suffering from the disease and / or subjects not suffering from the disease. Preferably, the method includes comparing the sequencing output of the normalized cDNA sample with the sequencing output of one or more control samples from one or more subjects suffering from the disease.

[0038] Sequencing output means one or more sequences obtained from sequencing (normalized) cDNA. The sequence can be the original sequence or can be further processed. For example, low-quality reads can be filtered out and / or adapter sequences can be filtered out and removed. (Processed) sequence can be mapped to a human reference genome (for example, using Minimap2 to carry out) to prepare a transcriptome spectrum. One or more transcript models can be identified in a transcriptome spectrum (sequence mapped to genome). Transcript model represents a specific transcript, i.e., a specific RNA isoform or splice variant produced by a gene. Therefore, in some specific embodiments, the sequencing output used in the method defined herein (for example, it is compared to find disease biomarkers or is used to identify whether an object suffers from a disease) can be a transcriptome spectrum and / or a transcript model.

[0039] Using sequencing output to identify whether a subject has a disease may include detecting a disease biomarker. In some specific embodiments, using sequencing output to identify whether a subject has a disease includes detecting more than one disease biomarker, optionally more than 10, 100, 1000, 10000, 100000, 1 million or 10 million disease biomarkers. In other embodiments, using sequencing output to identify whether a subject has a disease includes detecting 1 to 10, 1 to 100, 1 to 1000, 1 to 10000, 1 to 100000, 1 to 1 million or 1 to 10 million disease biomarkers. Detecting a disease biomarker may include determining the presence or absence of a disease biomarker. Thus, in some embodiments, using the sequencing output to identify whether a subject has a disease comprises determining the presence or absence of more than one disease biomarker, optionally more than 10, 100, 1000, 10000, 100000, 1 million, or 10 million disease biomarkers. In other embodiments, using the sequencing output to identify whether a subject has a disease comprises determining the presence or absence of 1 to 10, 1 to 100, 1 to 1000, 1 to 10000, 1 to 100000, 1 to 1 million, or 1 to 10 million disease biomarkers.

[0040] The presence of a particular cDNA sequence in the sequencing output can indicate that the subject has a disease, for example, when a particular transcript (corresponding to a cDNA molecule) is uniquely present in a subject with a particular disease. Likewise, the presence of a particular cDNA sequence in the sequencing output can indicate a characterization and / or prognosis of a disease. The presence of a particular cDNA sequence in the sequencing output can allow prediction of the responsiveness of a subject with a disease to a therapeutic agent, for example, when a particular transcript is found to be correlated with the responsiveness of a subject with a disease to a particular therapeutic agent.

[0041] The absence of a particular cDNA sequence in the sequencing output may indicate that the subject suffers from a disease, for example, if a particular transcript (corresponding to a cDNA molecule) is not present in a subject suffering from a particular disease. Likewise, the absence of a particular cDNA sequence in the sequencing output may indicate a characterization and / or prognosis of a disease. The absence of a particular cDNA sequence in the sequencing output may allow prediction of the responsiveness of a subject suffering from a disease to a therapeutic agent, for example, if a particular transcript is found to be correlated with the responsiveness of a subject suffering from a disease to a particular therapeutic agent.

[0042] According to all aspects of the present invention, in some specific embodiments, the cDNA sample is obtained from a biological fluid or a fluid or lysate produced by a biological material. The cDNA sample can be obtained from blood. The cDNA sample can be obtained by extracting RNA from a biological sample (e.g., blood) obtained from a subject. The cDNA is then synthesized using RNA as a template (i.e., by reverse transcription). Therefore, in certain embodiments, before step (i), the method further comprises: (a) extracting RNA from a biological fluid from a subject or a fluid or lysate produced from biological material from a subject; and (b) Using RNA as a template to synthesize cDNA (i.e. converting RNA into cDNA).

[0043] In this way, a cDNA sample from a subject can be generated.

[0044] The method may also include reporting the outcome of the method to the subject. The result may be a diagnosis or prognosis of a disease. In some embodiments, the result is a specific level or stage of a disease (e.g., cancer).

[0045] Complementary DNA (cDNA) normalization (Alex S. Shcheglov, Pavel A. Zhulidov, Ekaterina A. Bogdanova, DAS Normalization of cDNA Libraries, Nucleic Acids Hybrid. CHAPTER 5, (2014)) solves the problem that high-abundance housekeeping genes reduce the sampling efficiency of target genes. Since RNA sequencing usually relies on the conversion of RNA to double-stranded cDNA, cDNA normalization utilizes the biochemical properties of cDNA to produce a uniform distribution of unique genes and isotypes in the cDNA library. In theory, if all unique RNA sequences are expressed in the same relative abundance, the maximum non-targeted sampling efficiency is generated. Therefore, the purpose of normalization is to redistribute the cDNA library (sample) to meet the standard as closely as possible.

[0046] Normalizing the cDNA samples results in the generation of normalized cDNA samples. In certain embodiments, normalization includes (selectively) increasing the relative abundance of less abundant sequences without targeting a specific sequence based on its nucleotide sequence (i.e., identity or homology to a known sequence).

[0047] "Normalization" means that the level of RNA or cDNA sequence in the sample is more equal.Therefore, the normalized cDNA sample can be a cDNA sample in which the amount of each unique cDNA sequence is more uniform than the amount in the same sample before normalization, i.e. the normalized cDNA sample is closer to the realization that each unique cDNA sequence has the same abundance (relative to other unique cDNA sequences in the normalized cDNA sample) than the same sample before normalization.In order to achieve this, the relative representativeness or level of less abundant sequence can be improved, and / or the relative representativeness or level of more abundant sequence can be reduced.In terms of the following meaning, the improvement of less abundant sequence / reduction of more abundant sequence is selective: if all sequences are improved / reduced to the same degree, relative abundance will remain unchanged.However, in the case of targeting specific sequence (e.g., using pre-defined probes) not based on the nucleotide composition of specific sequence (i.e., based on their identity or homology with known sequences), the relative representativeness or level of less abundant sequence can be improved and / or the relative representativeness or level of more abundant sequence can be reduced. According to all embodiments, less abundant sequence can be a unique sequence whose amount is lower than threshold value, for example, they exist in the cDNA sample before normalization with an amount lower than the average amount of unique sequence in sample. Less abundant sequence can exist in the cDNA sample before normalization with an amount lower than 0.1%, 1%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% than the average amount of unique sequence in sample. More abundant sequence can be a unique sequence whose amount is higher than threshold value, for example, they exist in the cDNA sample before normalization with an amount higher than the average amount of unique sequence in sample. More abundant sequence can exist in the cDNA sample before normalization with an amount higher than 0.1%, 1%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% than the average amount of unique sequence in sample. Relative abundance means the abundance relative to other unique sequences in sample.

[0048] In some specific embodiments, the normalized cDNA sample comprises cDNA sequences with substantially the same level. For example, wherein the level of the normalized cDNA sequence changes less than 50%, less than 40%, less than 30%, less than 20% or less than 10%. The normalized cDNA can be a normalized cDNA sample, wherein at least a portion of 10, 100, 1000 or 10000 kinds of the most abundant (unique) sequences in the cDNA sample has been removed or reduced (reduced by at least 1%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90%) in terms of copy number. Normalized cDNA can be a normalized cDNA sample in which the level of at least a portion of the 10, 100, 1000 or 10000 least abundant (unique) sequences in the cDNA sample has been increased in copy number, e.g., increased by at least 1%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90%. The method described herein for normalizing a cDNA sample can be a method for equalizing a cDNA sample, i.e., equalizing the relative abundance of each unique sequence.

[0049] According to all aspects of the invention, in some specific embodiments, normalizing cDNA samples reduces the variability of cDNA levels (e.g., reducing by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80% or at least 90%). Normalizing cDNA can achieve a more uniform distribution of cDNA sequences. The difference in abundance between the most abundant cDNA and the least abundant cDNA in a sample can be reduced (e.g., reducing by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80% or at least 90%). In certain embodiments, the cDNA sample is normalized to reduce the number of molecules (copy number) of the most abundant cDNA molecules (1, 10, 100, 1000 or 10000) by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80% or at least 90%. In some specific embodiments, the number of molecules (copy number) of the most abundant cDNA molecules in the (first and / or second) cDNA sample is reduced by at least 50% in the normalized cDNA. In other embodiments, the relative abundance of the least abundant cDNA molecules (1, 10, 100, 1000 or 10000) is increased by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80% or at least 90%. In some embodiments, the number of molecules (copy number) of the least abundant cDNA molecule in the (first and / or second) cDNA sample is increased by at least 50% in the normalized cDNA.

[0050] According to all aspects of the invention, in certain embodiments, normalizing the (first and / or second) cDNA sample increases the amount (copy number) of at least a portion of the low-abundance cDNA sequences in the (first and / or second) cDNA sample. The low-abundance cDNA sequence may be 50%, 40%, 30%, 20%, 10% or 1% of the (unique) sequence with the lowest copy number. Thus, normalization may include selectively increasing the amount of low-abundance cDNA in each cDNA sample.

[0051] In certain embodiments, the normalized cDNA is a cDNA that can be more easily analyzed. It can be sequenced more effectively because the relative representativeness of less abundant sequences has been improved. In some specific embodiments, normalization of (the first and / or second) cDNA sample does not include removing abundant (more abundant) cDNA molecules / sequences (e.g., with albumin, IgG, apolipoprotein AI, transferrin, apolipoprotein A-II, α1-proteinase inhibitor, α1-acid glycoprotein, thyroxine transporter, hepatoglobin (Hepatoglobin) and / or hemoglobin corresponding to those) from the sample (e.g., using duplex-specific nuclease or sequence targeting method). In other embodiments, normalization does not include targeting specific (unique) sequences (e.g., with albumin, IgG, apolipoprotein AI, transferrin, apolipoprotein A-II, α1-proteinase inhibitor, α1-acid glycoprotein, thyroxine transporter, hepatoglobin and / or hemoglobin corresponding to those). Thus, in some embodiments, normalization of cDNA samples is non-targeted, i.e., it does not involve targeting a specific sequence based on its nucleotide sequence (e.g., it does not involve targeting a specific sequence based on its identity or homology to a known sequence).

[0052] According to all aspects of the present invention, the term "sequence" can refer to all single nucleic acid (e.g., cDNA or RNA) molecules with 100% identical nucleotide sequences. Alternatively, the term "sequence" can refer to all single nucleic acid (e.g., cDNA or RNA) molecules with greater than 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55% or 50% homogeneity to each other. "% identity" between a query nucleic acid sequence and a subject nucleic acid sequence can be calculated over the entire length of the query sequence using a suitable algorithm (e.g., BLASTN, FASTA, Needleman-Wunsch, Smith-Waterman, LALIGN, or GenePAST / KERR) or software (e.g., DNASTAR Lasergene, GenomeQuest, EMBOSS needle, or EMBOSS infoalign) after performing a pairwise global sequence alignment using a suitable algorithm (e.g., Needleman-Wunsch or GenePAST / KERR) or software (e.g., DNASTAR Lasergene or GenePAST / KERR). The term "unique sequence" or "unique cDNA sequence" can refer to all individual nucleic acid (e.g., cDNA or RNA, as the case may be) molecules that meet or exceed a threshold % identity (e.g., 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% identity to each other). A "unique sequence" or "unique cDNA sequence" can be different from other sequences present in a sample (e.g., different from other sequences by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or 100 nucleotides, or different from other sequences by at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50%).

[0053] According to all aspects of the present invention, in some embodiments, the disease is not an infectious disease. In certain embodiments, the disease is cancer. The cancer can be an epithelial cancer. In some embodiments, the cancer is breast cancer, ovarian cancer and / or colorectal cancer. Preferably, the cancer is breast cancer. In other embodiments, the first cDNA sample is from a subject suffering from breast cancer, and the second cDNA sample is from a subject suffering from a benign breast condition.

[0054] In some embodiments, the methods can be used to diagnose more than one disease in a single procedure, for example, by detecting multiple cDNA molecules (derived from transcripts) that are each uniquely present in a subject with a particular disease.

[0055] In some embodiments, the method can be used to diagnose more than one cancer type. In certain embodiments, the method can be used to distinguish breast cancer from benign breast conditions.

[0056] The inventors have developed a selective amplification method (called "Level-Up") that can be used for cDNA normalization. Level-Up normalization allows cDNA libraries to be taken and the relative abundance of each unique transcript sequence to be equalized. This is done without depletion and without targeting. Essentially, this allows the creation of an ideal cDNA library for detecting all RNAs present in a sample.

[0057] Therefore, normalization of cDNA samples can be achieved by using a selective amplification method of single-stranded cDNA, the method comprising: (i) providing a cDNA sample comprising double-stranded cDNA templates, each template having a known 5' pre-attached adaptor and a known 3' pre-attached adaptor; (ii) denaturing the cDNA sample to generate a single-stranded cDNA template; (iii) re-associating the cDNA sample to generate a mixture of associated single-stranded cDNA templates and associated double-stranded cDNA templates; (iv) annealing a 5' adapter complex to a 5' pre-attached adapter of at least one post-association single-stranded cDNA template, and annealing a 3' adapter complex to a 3' pre-attached adapter of the same post-association single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide; (v) ligating the oligonucleotide from the 5' adapter complex to the 5' pre-attached adapter of the associated single-stranded cDNA template, and ligating the oligonucleotide from the 3' adapter complex to the 3' pre-attached adapter of the same associated single-stranded cDNA template; and (vi) Selectively amplifying the cDNA sample using primers specific for the ligated oligonucleotides.

[0058] Thus, in some embodiments, the first cDNA sample and the second cDNA sample comprise double-stranded cDNA templates, each template having a known 5' pre-attached adaptor and a known 3' pre-attached adaptor; And normalizing the first cDNA sample and the second cDNA sample comprises: (i) denaturing the cDNA sample to generate a single-stranded cDNA template; (ii) re-associating the cDNA sample to generate a mixture of associated single-stranded cDNA templates and associated double-stranded cDNA templates; (iii) annealing a 5' adapter complex to a 5' pre-attached adapter of at least one post-association single-stranded cDNA template, and annealing a 3' adapter complex to a 3' pre-attached adapter of the same post-association single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide; (iv) ligating the oligonucleotide from the 5' adapter complex to the 5' pre-attached adapter of the associated single-stranded cDNA template, and ligating the oligonucleotide from the 3' adapter complex to the 3' pre-attached adapter of the same associated single-stranded cDNA template; and (v) Selectively amplifying the cDNA sample using primers specific for the ligated oligonucleotides.

[0059] In other embodiments, the cDNA sample comprises double-stranded cDNA templates, each template having a known 5' pre-attached adaptor and a known 3' pre-attached adaptor; And normalization of cDNA samples includes: (i) denaturing the cDNA sample to generate a single-stranded cDNA template; (ii) re-associating the cDNA sample to generate a mixture of associated single-stranded cDNA templates and associated double-stranded cDNA templates; (iii) annealing a 5' adapter complex to a 5' pre-attached adapter of at least one post-association single-stranded cDNA template, and annealing a 3' adapter complex to a 3' pre-attached adapter of the same post-association single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide; (iv) ligating the oligonucleotide from the 5' adapter complex to the 5' pre-attached adapter of the associated single-stranded cDNA template, and ligating the oligonucleotide from the 3' adapter complex to the 3' pre-attached adapter of the same associated single-stranded cDNA template; and (v) Selectively amplifying the cDNA sample using primers specific for the ligated oligonucleotides.

[0060] The present invention also provides a method for discovering disease biomarkers, comprising: (i) providing a first cDNA sample from a subject having a disease and a second cDNA sample from a subject not having the disease, wherein the first cDNA sample and the second cDNA sample comprise double-stranded cDNA templates, each template having a known 5' pre-attached adaptor and a known 3' pre-attached adaptor; (ii) denaturing the first cDNA sample and the second cDNA sample to generate a single-stranded cDNA template; (iii) re-associating the first cDNA sample and the second cDNA sample to generate a mixture of associated single-stranded cDNA templates and associated double-stranded cDNA templates; (iv) annealing a 5' adapter complex to a 5' pre-attached adapter of at least one post-association single-stranded cDNA template and annealing a 3' adapter complex to a 3' pre-attached adapter of the same post-association single-stranded cDNA template in each of the first cDNA sample and the second cDNA sample, wherein each adapter complex comprises at least one oligonucleotide; (v) ligating the oligonucleotide from the 5' adapter complex to the 5' pre-attached adapter of the associated single-stranded cDNA template, and ligating the oligonucleotide from the 3' adapter complex to the 3' pre-attached adapter of the same associated single-stranded cDNA template; (vi) selectively amplifying the first cDNA sample and the second cDNA sample using primers specific for the ligated oligonucleotides; (vii) sequencing the selectively amplified first cDNA sample and the second cDNA sample; and (viii) comparing the sequencing outputs of the first cDNA sample and the second cDNA sample to discover disease biomarkers.

[0061] The first cDNA sample and the second cDNA sample remain separate and / or individually identifiable.

[0062] The present invention also provides a method for diagnosing a disease in a subject, the method comprising: (i) providing a cDNA sample from a subject comprising double-stranded cDNA templates, each template having a known 5' pre-attached adaptor and a known 3' pre-attached adaptor; (ii) denaturing the cDNA sample to generate a single-stranded cDNA template; (iii) re-associating the cDNA sample to generate a mixture of associated single-stranded cDNA templates and associated double-stranded cDNA templates; (iv) annealing a 5' adapter complex to a 5' pre-attached adapter of at least one post-association single-stranded cDNA template, and annealing a 3' adapter complex to a 3' pre-attached adapter of the same post-association single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide; (v) ligating the oligonucleotide from the 5' adapter complex to the 5' pre-attached adapter of the associated single-stranded cDNA template, and ligating the oligonucleotide from the 3' adapter complex to the 3' pre-attached adapter of the same associated single-stranded cDNA template; and (vi) selectively amplifying the cDNA sample using primers specific for the ligated oligonucleotides; and (vii) sequencing the selectively amplified cDNA sample, wherein the sequencing output is used to identify whether the subject has the disease.

[0063] In certain embodiments: (A) The 5' adaptor complex is a pre-oligonucleotide dimer comprising: (i) a pre-lig-oligonucleotide for ligation with a 5' pre-attached adaptor of the (after association) single-stranded cDNA template; and (ii) a frontlink-oligonucleotide for annealing to the 5' pre-attached adaptor and the front-lig-oligonucleotide, the frontlink-oligonucleotide comprising a region complementary to the 5' pre-attached adaptor and a region complementary to the front-lig-oligonucleotide, such that upon annealing, the end of the pre-lig-oligonucleotide is adjacent to the end of the 5' pre-attached adaptor to enable ligation of the pre-lig-oligonucleotide to the 5' pre-attached adaptor at the ligation site; and (B) The 3' adaptor complex is a post-oligonucleotide dimer comprising: (i) a post-lig-oligonucleotide for ligation with the 3' pre-attached adaptor of the (after association) single-stranded cDNA template; and (ii) a backlink-oligonucleotide for annealing to the 3' pre-attached adaptor and the back-lig-oligonucleotide, the back-link-oligonucleotide comprising a region complementary to the 3' pre-attached adaptor and a region complementary to the back-lig-oligonucleotide, Such that upon annealing, the end of the later lig-oligonucleotide is adjacent to the end of the 3' pre-attached adaptor to enable ligation of the later lig-oligonucleotide to the 3' pre-attached adaptor at the ligation site.

[0064] Suitably: (A) Front Linker-Oligonucleotide comprising: (i) a template overhang region at the end of the pre-ligation-oligonucleotide proximal to the region complementary to the 5' pre-attached adaptor, said template overhang region being non-complementary to the corresponding region of the (after association) single-stranded cDNA template; and / or (ii) a lig-oligonucleotide overhang region at the end of the pre-lig-oligonucleotide proximal to the region complementary to the pre-lig-oligonucleotide, said lig-oligonucleotide overhang region being non-complementary to the corresponding region of the pre-lig-oligonucleotide; and / or (B) Post-ligation - oligonucleotide comprising: (i) a template overhang region at the end of the post-ligation-oligonucleotide proximal to the region complementary to the 3' pre-attached adaptor, said template overhang region being non-complementary to a corresponding region of the (after association) single-stranded cDNA template; and / or (ii) a lig-oligonucleotide overhang region at the end of the second ligation-oligonucleotide proximal to the region complementary to the second lig-oligonucleotide, said lig-oligonucleotide overhang region being non-complementary to the corresponding region of the second lig-oligonucleotide.

[0065] Suitably, the length of the template overhang and / or lig-oligonucleotide overhang is about 1 bp to about 20 bp. The template overhang and / or lig-oligonucleotide overhang can be 2 bp to 19 bp, 3 bp to 18 bp, 2 bp to 17 bp, 3 bp to 16 bp, 2 bp to 15 bp, 3 bp to 14 bp, 2 bp to 13 bp, 3 bp to 12 bp, 2 bp to 11 bp, 3 bp to 10 bp, 2 bp to 9 bp, 3 bp to 8 bp, 2 bp to 7 bp, 3 bp to 6 bp, 2 bp to 5 bp, 3 bp to 5 bp, or 2 bp to 4 bp. Preferably, the template overhang and / or lig-oligonucleotide overhang is 3 bp.

[0066] The template overhang and / or the lig-oligonucleotide overhang may be at least 2 bp or at least 3 bp. Preferably, the template overhang and / or the lig-oligonucleotide overhang is at least 3 bp.

[0067] Suitably, the combined length of the front link-oligonucleotide and the front lig-oligonucleotide is less than about 300bp, and / or the combined length of the rear link-oligonucleotide and the rear lig-oligonucleotide is less than about 300bp. The combined length of the front link-oligonucleotide and the front lig-oligonucleotide can be at least about 200bp, and / or the combined length of the rear link-oligonucleotide and the rear lig-oligonucleotide can be at least about 100bp.

[0068] Suitably, the length of the front link-oligonucleotide and / or the back link-oligonucleotide is less than 200 bp.The length of the front link-oligonucleotide and / or the back link-oligonucleotide may be at least 50 bp.

[0069] Suitably, the pre-oligonucleotide dimer and / or the post-oligonucleotide dimer has at least one non-blunt end.

[0070] Suitably, the front ligation-oligonucleotide and / or the back ligation-oligonucleotide provide at least 5 bp of complementary binding on either side of the ligation site.

[0071] Suitably, the nucleotide sequence of the preceding oligonucleotide dimer is different and non-complementary to the nucleotide sequence of the following oligonucleotide dimer.

[0072] Suitably, at least one of the front oligonucleotide dimer and the rear oligonucleotide dimer may anneal to the (after association) single stranded cDNA template at a temperature exceeding 30°C.

[0073] Suitably, the concentration of pre-oligonucleotide dimers and / or the concentration of post-oligonucleotide dimers exceeds the predicted total single-stranded cDNA concentration or the total cDNA concentration in the cDNA sample.

[0074] Suitably, the duration of the step of re-association of the cDNA sample is from 0 to 24 hours, optionally from 0 to 8 hours, from 1 to 7 hours, from 1 to 24 hours or from 7 to 24 hours.

[0075] By using long-read sequencing, full-length RNA / cDNA can be detected, which will provide better information for identifying the tissue source and specific function of each RNA. Although long-read RNA sequencing is expensive compared to other assays, Level-Up technology makes it possible to reduce the amount of sequencing required, thereby reducing the overall cost.

[0076] Therefore, according to all aspects of the invention, in some embodiments, sequencing comprises the use of long read sequencing.Thus, sequencing can be performed by long read sequencing, for example long read sequencing that allows sequencing reads greater than 1000, 5000 or 10000 bp.

[0077] Level-Up makes it possible to analyze samples with RNA degradation by increasing the presence of low-abundance transcripts. However, it is preferred to minimize RNA degradation during processing of biological samples to generate cDNA samples. The inventors have developed methods for processing blood samples when RNA extraction is not performed on the day of blood collection. Specific steps for freezing, storing and thawing blood samples improve the condition of the samples, especially if they are to be sequenced for long reads.

[0078] Therefore, the present invention provides a method for processing a blood sample, comprising: (i) storing the blood sample at -15°C or below; (ii) thawing the blood sample at 5 to 30°C for at least 1 hour; and (iii) RNA was extracted from thawed blood samples.

[0079] The blood sample can be a liquid (i.e., non-dried) blood sample. The blood sample can be whole blood. In certain embodiments, the blood sample is in a sample tube. In some specific embodiments, the blood sample is not absorbed into a material (e.g., a sponge).

[0080] In certain embodiments, the blood samples are stored at -15°C to -80°C, -15°C to -70°C, -15°C to -60°C, -15°C to -50°C, -15°C to -40°C, -15°C to -30°C, or -15°C to -20°C.

[0081] The blood sample may be stored at or below -15°C within 12 hours, 8 hours, 5 hours, 2 hours, 1 hour, 30 minutes, 15 minutes, 5 minutes or 1 minute of collection. Preferably, the blood sample is stored at or below -15°C within 12 hours of collection (i.e., obtaining the blood sample from the subject).

[0082] In certain embodiments, the blood sample is stored at or below -20° C. The blood sample may be stored at or below -20° C. within 12 hours, 8 hours, 5 hours, 2 hours, 1 hour, 30 minutes, 15 minutes, 5 minutes, or 1 minute of collection. Preferably, the blood sample is stored at or below -20° C. within 12 hours of collection (i.e., obtaining the blood sample from the subject).

[0083] In certain embodiments, the blood sample is stored at -15°C or below (optionally -20°C or below) for at least 24 hours (and optionally no more than 72 hours, 1 week, 2 weeks, 4 weeks, 1 month or 2 months), and then stored at -70°C or below (optionally -80°C or below) for no more than 4, 5, 6, 7, 8, 9, 10, 11 or 12 months, or 2, 3, 4 or 5 years. Preferably, the storage at -70°C or below (optionally -80°C or below) is for no more than 5 years.

[0084] Preferably, thawing of the blood sample occurs on the same day as RNA extraction.

[0085] In some specific embodiments, the blood sample is thawed at 16 to 29° C., 17 to 28° C., 18 to 27° C., 18 to 26° C., or 18 to 25° C. Preferably, the blood sample is thawed at 18 to 25° C. The duration of the thawing step can be 1 to 5 hours, 2 to 5 hours, 1 to 4 hours, 2 to 4 hours, 1 to 3 hours, 2 to 3 hours, or 1 to 2 hours. In a preferred embodiment, the blood sample is thawed at 18 to 25° C. for 1 to 3 hours. More preferably, the blood sample is thawed at 18 to 25° C. for 3 hours.

[0086] In some embodiments, once the blood sample is completely thawed, the sample tube is inverted at least 5, 6, 7, 8, 9, or 10 times. Preferably, the sample tube is inverted 10 times. The blood sample can then be incubated at 18 to 25° C. for about 2 hours before RNA extraction.

[0087] RNA can be extracted from (thawed) blood samples using the Qiagen Paxgene Blood RNA Kit.

[0088] In some specific embodiments, prior to step (i), blood is received in a container (optionally a Paxgene blood RNA tube) at room temperature (5 to 30°C, preferably 18 to 25°C). The container may be inverted at least 5, 6, 7, 8, 9 or 10 times immediately after blood collection. Preferably, the container is inverted 10 times immediately after blood collection. The blood sample may be stored at or below -15°C (or at or below -20°C) immediately after inversion.

[0089] Therefore, the present invention provides a method for processing a blood sample, comprising: (i) receiving a blood sample in a container at a temperature between 18 and 25° C.; (ii) storing blood samples at or below -20°C within 12 hours of collection; (iii) thawing the blood sample at 18 to 25° C. for 1 to 3 hours; and (iv) RNA was extracted from thawed blood samples.

[0090] According to all aspects of the invention, in some embodiments, the blood sample is no more than 5ml, 3ml, 2.5ml or 2ml. Preferably, the blood sample is no more than 2.5ml. The blood sample can be 0.1ml to 5ml, 0.5ml to 5ml, 1ml to 4ml, or 2ml to 3ml.

[0091] The method for processing a blood sample can be combined with the above-mentioned method for using a cDNA sample. Therefore, in some specific embodiments, a cDNA sample is obtained from blood by following the step of the method for processing a blood sample described herein. In addition, the method may include the step of using the extracted RNA as a template to synthesize cDNA (i.e., converting the extracted RNA into cDNA, reverse transcription). In other embodiments, a first cDNA sample and a second cDNA sample are obtained from blood by following the step of the method for processing a blood sample described herein. In addition, the method may include the step of using the extracted RNA as a template to synthesize cDNA (i.e., converting the extracted RNA into cDNA, reverse transcription).

[0092] Therefore, the present invention provides a method for discovering disease biomarkers, comprising: (i) receiving a first blood sample from a subject having a disease in a first container and a second blood sample from a subject not having the disease in a second container at 18 to 25° C.; (ii) storing the first blood sample and the second blood sample at or below -20°C within 12 hours of collection from the subject; (iii) thawing the first blood sample and the second blood sample at 18 to 25° C. for 1 to 3 hours; (iv) extracting RNA from the thawed first blood sample and the second blood sample; (v) synthesizing cDNA using the extracted RNA as a template to form a first cDNA sample from a subject suffering from the disease and a second cDNA sample from a subject not suffering from the disease; (vi) normalizing the first cDNA sample and the second cDNA sample; (vii) sequencing the normalized first cDNA sample and the second cDNA sample; and (viii) comparing the sequencing outputs of the first cDNA sample and the second cDNA sample to discover disease biomarkers.

[0093] Additionally, the present invention provides a method for diagnosing a disease in a subject, comprising: (i) receiving a blood sample from a subject in a container at a temperature between 18 and 25° C.; (ii) storing blood samples at or below -20°C within 12 hours of collection from the subject; (iii) thawing the blood sample at 18 to 25° C. for 1 to 3 hours; (iv) extracting RNA from thawed blood samples; (v) synthesizing cDNA using the extracted RNA as a template to form a cDNA sample from the subject; (vi) normalizing the cDNA samples; and (vii) sequencing the normalized cDNA sample, wherein the sequencing output is used to identify whether the subject has the disease.

[0094] Another aspect of the present invention provides the use of an oligonucleotide dimer composition in a method for discovering cancer biomarkers and / or cancer vaccine targets, wherein the oligonucleotide dimer composition is used for selectively amplifying single-stranded cDNA by ligating oligonucleotides to the 5' end and 3' end of a post-association single-stranded cDNA template having a known 5' pre-attached adaptor and a 3' pre-attached adaptor, wherein the composition comprises: (A) a pre-oligonucleotide dimer comprising: (i) a pre-lig-oligonucleotide for ligation with a 5' pre-attached adaptor of the associated single-stranded cDNA template; and (ii) a pre-ligation-oligonucleotide for annealing to the 5' pre-attached adaptor and the pre-lig-oligonucleotide, the pre-ligation-oligonucleotide comprising a region complementary to the 5' pre-attached adaptor and a region complementary to the pre-lig-oligonucleotide, The ends of the pre-lig-oligonucleotide are adjacent to the ends of the 5' pre-attached adaptor upon annealing to enable the pre-lig-oligonucleotide to be attached to the 5' pre-attached adaptor at the ligation site. Connectivity; and (B) a rear oligonucleotide dimer comprising: (i) a post-lig-oligonucleotide for ligation with the 3' pre-attached adaptor of the associated post-single-stranded cDNA template; and (ii) a post-ligation-oligonucleotide for annealing to the 3' pre-attachment adaptor and the post-lig-oligonucleotide, the post-ligation oligonucleotide comprising a region complementary to the 3' pre-attachment adaptor and a region complementary to the post-lig-oligonucleotide, Such that upon annealing, the end of the later lig-oligonucleotide is adjacent to the end of the 3' pre-attached adaptor to enable ligation of the later lig-oligonucleotide to the 3' pre-attached adaptor at the ligation site.

[0095] The present invention also provides use of an oligonucleotide dimer composition in a method for diagnosing and / or prognosing cancer in a subject, wherein the oligonucleotide dimer composition is used for selectively amplifying single-stranded cDNA by ligating an oligonucleotide to the 5' end and 3' end of a post-association single-stranded cDNA template having a known 5' pre-attached adaptor and a 3' pre-attached adaptor, wherein the composition comprises: (A) a pre-oligonucleotide dimer comprising: (i) a pre-lig-oligonucleotide for ligation with a 5' pre-attached adaptor of the associated single-stranded cDNA template; and (ii) a pre-ligation-oligonucleotide for annealing to the 5' pre-attached adaptor and the pre-lig-oligonucleotide, the pre-ligation-oligonucleotide comprising a region complementary to the 5' pre-attached adaptor and a region complementary to the pre-lig-oligonucleotide, The ends of the pre-lig-oligonucleotide are adjacent to the ends of the 5' pre-attached adaptor upon annealing to enable the pre-lig-oligonucleotide to be attached to the 5' pre-attached adaptor at the ligation site. Connectivity; and (B) a rear oligonucleotide dimer comprising: (i) a post-lig-oligonucleotide for ligation with the 3' pre-attached adaptor of the associated post-single-stranded cDNA template; and (ii) a post-ligation-oligonucleotide for annealing to the 3' pre-attachment adaptor and the post-lig-oligonucleotide, the post-ligation oligonucleotide comprising a region complementary to the 3' pre-attachment adaptor and a region complementary to the post-lig-oligonucleotide, Such that upon annealing, the end of the later lig-oligonucleotide is adjacent to the end of the 3' pre-attached adaptor to enable ligation of the later lig-oligonucleotide to the 3' pre-attached adaptor at the ligation site.

[0096] The oligonucleotide dimer compositions defined herein may also be used in methods of characterizing a disease in a subject, in methods of selecting a treatment for a disease in a subject, and / or in methods of predicting the responsiveness of a subject suffering from a disease to a therapeutic agent.

[0097] In some specific embodiments: (A) Front Linker-Oligonucleotide comprising: (i) a template overhang region at the end of the pre-ligation-oligonucleotide proximal to the region complementary to the 5' pre-attached adaptor, said template overhang region being non-complementary to the corresponding region of the post-association single-stranded cDNA template; and / or (ii) a lig-oligonucleotide overhang region at the end of the pre-lig-oligonucleotide proximal to the region complementary to the pre-lig-oligonucleotide, said lig-oligonucleotide overhang region being non-complementary to the corresponding region of the pre-lig-oligonucleotide; and / or (B) Post-ligation - oligonucleotide comprising: (i) a template overhang region at the end of the post-ligation-oligonucleotide proximal to the region complementary to the 3' pre-attached adaptor, said template overhang region being non-complementary to a corresponding region of the post-association single-stranded cDNA template; and / or (ii) a lig-oligonucleotide overhang region at the end of the second ligation-oligonucleotide proximal to the region complementary to the second lig-oligonucleotide, said lig-oligonucleotide overhang region being non-complementary to the corresponding region of the second lig-oligonucleotide.

[0098] Suitably, the length of the template overhang and / or lig-oligonucleotide overhang is about 1 bp to about 20 bp. The template overhang and / or lig-oligonucleotide overhang can be 2 bp to 19 bp, 3 bp to 18 bp, 2 bp to 17 bp, 3 bp to 16 bp, 2 bp to 15 bp, 3 bp to 14 bp, 2 bp to 13 bp, 3 bp to 12 bp, 2 bp to 11 bp, 3 bp to 10 bp, 2 bp to 9 bp, 3 bp to 8 bp, 2 bp to 7 bp, 3 bp to 6 bp, 2 bp to 5 bp, 3 bp to 5 bp, or 2 bp to 4 bp. Preferably, the template overhang and / or lig-oligonucleotide overhang is 3 bp.

[0099] The template overhang and / or the lig-oligonucleotide overhang may be at least 2 bp, or at least 3 bp. Preferably, the template overhang and / or the lig-oligonucleotide overhang is at least 3 bp.

[0100] Suitably, the combined length of the front link-oligonucleotide and the front lig-oligonucleotide is less than about 300 bp, and / or the combined length of the rear link-oligonucleotide and the rear lig-oligonucleotide is less than about 300 bp.

[0101] Suitably, the length of the front link-oligonucleotide and / or the back link-oligonucleotide is less than 200 bp.

[0102] Suitably, the pre-oligonucleotide dimer and / or the post-oligonucleotide dimer has at least one non-blunt end.

[0103] Suitably, in use of the composition, the front ligation-oligonucleotide and / or the back ligation-oligonucleotide provide at least 5 bp of complementary binding on either side of the ligation site.

[0104] Suitably, the nucleotide sequence of the preceding oligonucleotide dimer is different and non-complementary to the nucleotide sequence of the following oligonucleotide dimer.

[0105] Suitably, the front oligonucleotide dimer and / or the rear oligonucleotide dimer may be annealed to the associated single stranded cDNA template at a temperature exceeding 30°C.

[0106] A selective amplification kit is provided for selectively amplifying low-abundance cDNA from a cDNA sample and / or for selectively amplifying cDNA containing a known adapter sequence, wherein the cDNA sample contains a cDNA template with a known 5' pre-attached adapter and a 3' pre-attached adapter, and the kit contains tools for preparing the oligonucleotide dimer composition as described above and tools for implementing the selective amplification method as described above.

[0107] Another aspect of the invention provides the use of a kit as described herein for use in a method for diagnosing and / or prognosing a disease in a subject. Also provided is the use of a kit as described herein for use in a method for characterizing a disease in a subject, for use in a method for selecting a treatment for a disease in a subject, and / or for use in a method for predicting the responsiveness of a subject with a disease to a therapeutic agent.

[0108] Also provided is a selective amplification kit for selectively amplifying low-abundance cDNA from a first cDNA sample from a subject suffering from a disease and a second cDNA sample from a subject not suffering from the disease, and / or for selectively amplifying cDNA containing a known adapter sequence, wherein the first cDNA sample and the second cDNA sample contain cDNA templates with known 5' pre-attached adapters and 3' pre-attached adapters, and the kit comprises tools for preparing the oligonucleotide dimer composition as described above and tools for implementing the selective amplification method as described above.

[0109] Another aspect of the invention provides the use of a kit as described herein in a method for discovering disease biomarkers and / or cancer vaccine targets.

[0110] In some specific embodiments, the tool for preparing the oligonucleotide dimer composition may include a front lig-oligonucleotide, a front connection-oligonucleotide, a rear lig-oligonucleotide and / or a rear connection-oligonucleotide as described herein. In other embodiments, the tool for preparing the oligonucleotide dimer composition may include a front oligonucleotide dimer and / or a rear oligonucleotide dimer as described herein.

[0111] The means for carrying out the selective amplification method may comprise primers specific for the front and / or rear lig-oligonucleotides.

[0112] In some specific embodiments, the kit may also include a hybridization buffer. The hybridization buffer may include HEPES1M (pH=7.5), NaCl 5M and H2O. The kit may also include a ligase and / or a ligase buffer. Any suitable ligase may be used. The ligase may be a gap repair ligase or a blunt end ligase. Optionally, the ligase may be a Taq DNA ligase. Suitable ligase buffers are also known and commercially available.

[0113] In other embodiments, the kit may further comprise primers for adding phosphate groups to the cDNA before using the cDNA as a cDNA sample. These primers are based on known 5' pre-attached adapter and known 3' pre-attached adapter sequences.

[0114] In some specific embodiments, the kit may also include suitable reagents for PCR, including polymerase, dinucleotide triphosphate (dNTP), MgCl2 and one or more, up to all of the buffer. Any suitable polymerase may be used. Typically, DNA polymerase is used to increase nucleic acid targets according to the present invention. Some examples include thermostable polymerases, such as Taq or Pfu polymerases and the various derivatives of these enzymes. Suitable buffer is also known and commercially available, and may be included in a PCR master mix that includes most of the components required for PCR amplification.

[0115] In other embodiments, the kit also includes suitable reagents for reverse transcription of RNA into cDNA, including reverse transcriptase. Any suitable reverse transcriptase can be used. Suitable buffer is also known and commercially available, and can be included in a reverse transcription master mix that includes most of the components required for reverse transcription.

[0116] In some embodiments, the kit further comprises suitable reagents for processing blood samples, such as containers (optionally PAXgene blood RNA tubes), RNA stabilization reagents, and / or blood cell lysis buffers. The reagents may be RNase-free. RNA stabilization reagents are commercially available and include (Sigma-Aldrich) and RNAprotect (Qiagen). Suitable RNA stabilization reagents may include EDTA, sodium citrate and / or ammonium sulfate, such as 70% (w / v) ammonium sulfate, 25 mM sodium citrate and / or 10 mM EDTA. The pH may be adjusted to 5.2 with sulfuric acid.

[0117] Another aspect of the invention provides a reagent kit for use in a method as described herein, the reagent kit comprising: a pre-lig-oligonucleotide, a pre-link-oligonucleotide, a post-lig-oligonucleotide and / or a post-link-oligonucleotide; and Primers specific for the pre-lig-oligonucleotide and / or the post-lig-oligonucleotide.

[0118] The pre-lig-oligonucleotide, pre-link-oligonucleotide, post-lig-oligonucleotide and / or post-link-oligonucleotide may be provided as an oligonucleotide dimer composition as described herein.

[0119] The reagent set may also include one or more, up to all, of the following: hybridization buffer (optionally including HEPES1M (pH=7.5), NaCl 5M and H2O), ligase, ligase buffer, primer pairs for adding phosphate groups to cDNA, DNA polymerase and / or dNTPs.

[0120] In other embodiments, the reagent set further comprises suitable reagents for processing blood samples, such as containers (optionally PAXgene blood RNA tubes), RNA stabilization reagents, and / or blood cell lysis buffers. The reagents may be RNase-free. RNA stabilization reagents are commercially available and include (Sigma-Aldrich) and RNAprotect (Qiagen). Suitable RNA stabilization reagents may include EDTA, sodium citrate and / or ammonium sulfate, such as 70% (w / v) ammonium sulfate, 25 mM sodium citrate and / or 10 mM EDTA. The pH may be adjusted to 5.2 with sulfuric acid. Incubation for cell lysis may be, for example, 1 minute to 3 hours.

[0121] In one embodiment, the reagent kit comprises: RNA stabilization reagent; (blood) cell lysis buffer; pre-lig-oligonucleotide, pre-link-oligonucleotide, post-lig-oligonucleotide and / or post-link-oligonucleotide; Primers specific for the pre-lig-oligonucleotide and / or the post-lig-oligonucleotide; Hybridization buffer; Ligase; Ligase buffer; and Primer pair used to add phosphate groups to cDNA.

[0122] Another important aspect of the oligonucleotide dimer composition for selective amplification of single-stranded cDNA described herein is the ability to select cDNA with a known adapter sequence. This aspect can be applied to single-cell sequencing, where adapters are necessary to assign reads to individual cells. In this application, cDNA sequences without cell identification barcodes / adapters appear in the cDNA library. These are called template switching oligonucleotides (TSO) artifacts, and are undesirable in single-cell sequencing projects because they cannot be assigned to source cells. The method of the present invention can be applied to select only cDNA sequences with the desired adapter sequence, thereby effectively limiting the sequencing of TSO artifacts. The method of the present invention can also be performed with a single-cell cDNA library to both remove TSO artifacts and increase the transcriptome coverage of each cell.

[0123] TSO cleanup can also be performed without normalization, in which case there is no need to include a step of reassociating the cDNA sample to produce a mixture of associated single-stranded cDNA templates and associated double-stranded cDNA templates.

[0124] Therefore, according to another aspect of the present invention, there is provided a method for diagnosing a disease in a subject, the method comprising: (i) providing a cDNA sample from a subject comprising a double-stranded cDNA template, a portion of the template having a known 5' pre-attached adaptor and a known 3' pre-attached adaptor; (ii) denaturing the cDNA sample to generate a single-stranded cDNA template; (iii) annealing a 5' adapter complex to a 5' pre-attached adapter of at least one single-stranded cDNA template, and annealing a 3' adapter complex to a 3' pre-attached adapter of the same single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide; (iv) ligating the oligonucleotide from the 5' adapter complex to the 5' pre-attached adapter of the single-stranded cDNA template, and ligating the oligonucleotide from the 3' adapter complex to the 3' pre-attached adapter of the same single-stranded cDNA template; (v) selectively amplifying the cDNA sample using primers specific for the ligated oligonucleotides; and (vi) sequencing the cDNA sample, wherein the sequencing output is used to determine whether the subject has the disease.

[0125] Another aspect of the present invention provides a method for discovering disease biomarkers, the method comprising: (i) providing a first cDNA sample from a subject having a disease and a second cDNA sample from a subject not having the disease, the samples comprising a double-stranded cDNA template, a portion of the template having a known 5' pre-attached adaptor and a known 3' pre-attached adaptor; (ii) denaturing each cDNA sample to generate a single-stranded cDNA template; (iii) annealing a 5' adapter complex to a 5' pre-attached adapter of at least one single-stranded cDNA template, and annealing a 3' adapter complex to a 3' pre-attached adapter of the same single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide; (iv) ligating the oligonucleotide from the 5' adapter complex to the 5' pre-attached adapter of the single-stranded cDNA template, and ligating the oligonucleotide from the 3' adapter complex to the 3' pre-attached adapter of the same single-stranded cDNA template; (v) selectively amplifying each cDNA sample using primers specific for the ligated oligonucleotides; and (vi) sequencing each cDNA sample and comparing the sequencing outputs of the first cDNA sample and the second cDNA sample to discover disease biomarkers.

[0126] According to all aspects of the invention, the cDNA sample may be derived from a single cell.

[0127] The above-mentioned embodiment of the method for selectively amplifying single-stranded cDNA has made necessary modifications to the method for selectively amplifying cDNA containing a known adapter sequence, and is not repeated for the sake of brevity. The oligonucleotide dimer composition is suitable for use in the method for selectively amplifying cDNA containing a known adapter sequence as defined herein. The method for selectively amplifying cDNA containing a known adapter sequence as defined herein can be used to discover disease biomarkers and / or cancer vaccine targets, or for diagnosis and / or prognosis of the disease. The kits and reagent sets defined herein are also suitable for use in the method for selectively amplifying cDNA containing a known adapter sequence.

[0128] In some embodiments, according to all aspects of the invention, the cDNA sample comprises no more than 800ng, 700ng, 500ng, 100ng, 20ng, 10ng, 5ng or 1ng of starting cDNA. The cDNA sample may comprise 1 to 800ng, 1 to 500ng, 5 to 100ng or 10 to 50ng of starting cDNA.

[0129] In some specific embodiments, according to all aspects of the present invention, RNA from a sample is first reverse transcribed into cDNA. Sample types include blood samples (particularly blood samples from plasma and serum), other body fluids, such as saliva, urine or lymph. Other sample types include solid tissues, including frozen tissues or formalin fixed, paraffin embedded (FFPE) materials. RNA can be messenger RNA (mRNA), microRNA (miRNA), etc. In such embodiments, RNA is typically reverse transcribed using a reverse transcriptase to form a complementary DNA (cDNA) molecule. Methods for reverse transcribing RNA into cDNA using a reverse transcriptase are well known in the art. Any suitable reverse transcriptase can be used, and some examples of suitable reverse transcriptases are widely available in the art. The initial cDNA molecule can be single-stranded until a DNA polymerase has been used to produce a complementary strand. Commercially available kits (e.g. The Single Cell / Low Input cDNA Synthesis & Amplification Module) can be used to convert RNA into double-stranded cDNA with 5' and 3' adapters. Primers based on 5' and 3' adapters can be used to add phosphate groups to the cDNA. Before the cDNA is used as a cDNA sample, a cDNA purification step (e.g., with ProNex or Ampure beads) can be performed.

[0130] Since the present invention requires only a small amount of starting cDNA, this can be generated from a small amount of RNA and / or no additional PCR cycles are required during the generation of cDNA. The RNA sample may contain no more than 3.5 μg, 3 μg, 2 μg, 1 μg, 500 ng, 100 ng, 10 ng or 1 ng of starting RNA. The RNA sample may contain 1 ng to 3 μg, 10 ng to 2 μg or 100 ng to 1 μg of starting RNA.

[0131] The present invention also provides a system or test kit for discovering disease biomarkers and / or cancer vaccine targets, comprising: (a) one or more test devices for normalizing a first cDNA sample from a subject suffering from a disease (cancer) and a second cDNA sample from a subject not suffering from the disease (cancer), and sequencing the normalized first cDNA sample and the second cDNA sample; (b) processor; and (c) a storage medium containing a computer application which, when executed by a processor, is configured to: (i) obtaining a first cDNA sample and a second cDNA sample determined on one or more test devices; 2. Sequence of cDNA sample; (ii) Calculate whether there are cDNAs present in the first cDNA sample but not in the second cDNA sample cDNA sequences in the sample, or whether there are cDNA sequences that are present in the second cDNA sample but not in the cDNA sequences in the first cDNA sample; and (iii) Outputting the result of step (ii) from the processor.

[0132] In a related aspect, a system or test kit is provided for diagnosing and / or prognosing a disease in a subject, comprising: (a) one or more test devices for normalizing a cDNA sample from a subject and sequencing the normalized cDNA sample (b) processor; and (c) a storage medium containing a computer application which, when executed by a processor, is configured to: (i) obtaining the sequence of the cDNA sample determined on one or more test devices (ii) Calculate whether the cDNA sequence exists, where the presence or absence of the cDNA sequence is related to Disease related; and (iii) Outputting from the processor whether the subject suffers from the disease and / or the prognosis of the disease.

[0133] The one or more test devices may use / comprise an oligonucleotide dimer composition as described herein.The one or more test devices may comprise a long read sequencer.

[0134] The system or test kit may also include a display for output from the processor.

[0135] Also provided is a computer application or a storage medium comprising a computer application as defined herein. DETAILED DESCRIPTION

[0136] The above and other aspects of the present invention will now be described in further detail, by way of example only, with reference to the following examples and accompanying drawings, in which:

[0137] Figure 1 is a schematic overview showing the addition of a linker sequence to the end of a single-stranded cDNA template;

[0138] Figure 2 is a schematic overview of one embodiment of a cDNA normalization process according to the present invention;

[0139] Figure 3 is a schematic overview of the annealing of the front and rear oligonucleotide dimer structures to the single-stranded cDNA template as shown, Figure 3A is its details, and Figure 3B A diagrammatic representation using some of the exemplary sequences described herein is provided;

[0140] Figure 4 is a graph showing the length distribution from gel electrophoresis of input cDNA;

[0141] Figure 5 It is shown that Figure 2 A graph of the length distribution from gel electrophoresis of normalized cDNA produced by the cDNA normalization process shown in ; and

[0142] Figure 6 The use of nanopore cDNA sequencing is shown in Figure 2 Saturation curves of normalized cDNA and input cDNA in normalization.

[0143] Figure 7is a 2D PCA plot from 13573 transcripts showing the clustering of cancer and control samples in the dataset.

[0144] In order to solve the problems in the current cDNA normalization technology, the inventors have developed an improved selective amplification method (called "Level-Up"). The method uses the same denaturation and re-hybridization process as the DSN method and column method described herein. However, the method of the present invention is different from other methods by using a non-depletion or addition mechanism. In other words, Level-Up provides methods and tools for increasing the amount of low-abundance cDNA in a sample. These methods and tools can be used for cDNA normalization in sequencing processes, or for other processes that will benefit from the amplification of low-abundance cDNA, such as the discovery, detection or identification of biomarkers.

[0145] In the context of the present invention, the following explanations of terms and methods are provided to better describe the disclosure and to provide guidance in the practice of the disclosure.

[0146] The phrase "selective amplification" is used to describe the method developed by the inventors to amplify a specific DNA template preferentially over other DNA templates (e.g., amplifying only single-stranded cDNA in a sample containing a mixture of single-stranded and double-stranded cDNA). The term is also used herein to describe the preferential amplification of a specific type of DNA, such as low-abundance DNA.

[0147] The term "adapter" is used to describe a short DNA sequence added to the end of a DNA template, such as those commonly used in RNA sequencing by ligating the adaptor to a cDNA template. A "3' pre-attached adaptor" refers to an adaptor with a known nucleotide sequence that has been added to the 3' end of a cDNA template. A "5' pre-attached adaptor" refers to an adaptor with a known nucleotide sequence that has been added to the 5' end of a cDNA template.

[0148] The terms "normalization" and "normalization score" refer to the process of leveling the abundance of different transcripts in a sample. This can be achieved by prior art methods of reducing the amount of high abundance transcripts, or by selectively amplifying low abundance transcripts using the methods described herein.

[0149] As used in the context of the present invention, the phrase "post-association single-stranded cDNA template" refers to a single-stranded cDNA template produced by de-association and re-association (i.e., denaturation and re-hybridization) of a double-stranded cDNA sample to form a mixture of single-stranded cDNA and double-stranded cDNA. The single-stranded cDNA that remains single-stranded after re-association is referred to as post-association single-stranded cDNA. If the re-association step is not performed, the phrases "post-association single-stranded cDNA template" and "single-stranded cDNA template" are interchangeable in the embodiments defined herein.

[0150] As used herein, the phrase "ligated adapter-cDNA template" refers to a cDNA template formed by ligating an adapter to a cDNA template.

[0151] The present invention encompasses an "adaptor complex" that is suitable for annealing to the end of a single-stranded cDNA template after association. The term "adaptor complex" refers to an adaptor comprising more than one component.

[0152] As used in the context of the present invention, the terms "pre-oligonucleotide dimer", "pre-k-linker" and "pre-dimer" refer to an adapter complex that can anneal to the 5' end of the single-stranded cDNA template after association. As used in the context of the present invention, the terms "post-oligonucleotide dimer", "post-k-linker" and "post-dimer" refer to an adapter complex that can anneal to the 3' end of the single-stranded cDNA template after association.

[0153] The terms "pre-lig-oligonucleotide" and "pre-lig" as used in the context of the present invention refer to the oligonucleotide component of the pre-dimer. The terms "post-lig-oligonucleotide" or "post-lig" as used in the context of the present invention refer to the oligonucleotide component of the post-dimer.

[0154] As used in the context of the present invention, the terms "front link-oligonucleotide" and "front linker" refer to the oligonucleotide component of the front dimer. The terms "back link-oligonucleotide" and "back linker" used in the context of the present invention refer to the oligonucleotide component of the back dimer.

[0155] The term "overhang" is used in the context of the present invention to describe the overhang region of the dimer sequence of the present invention, wherein the overhang region is non-complementary to the region it is paired with, so that once annealed with the associated single-stranded cDNA template, the overhang region does not bind to the region it is paired with.

[0156] Use the primer selective amplification single-stranded cDNA that is specific to the oligonucleotide connected.Because single-stranded DNA is template, the primer region of a primer in the specific primer pair is complementary to the single-stranded DNA molecule.Another primer in the specific primer pair comprises the complementary single-stranded DNA molecule formed during the amplification cycle and the primer region that therefore hybridizes with it.Therefore, a primer is complementary to an oligonucleotide connected, and another primer (at least partially) comprises the sequence of another oligonucleotide connected. Previously developed cDNA normalization method

[0157] Two forms of full-length cDNA normalization have been developed previously: the double-stranded specific nuclease (DuplexSpecific Nuclease, DSN) method (Zhulidov, PA et al. Simple cDNA normalization using kamchatka crab duplex-specific nuclease. Nucleic Acids Res. 32, e37 (2004)) and the hydroxyapatite column method (Andrews-Pfannkoch, C., Fadrosh, DW, Thorpe, J. & Williamson, SJ Hydroxyapatite-mediated separation of double-stranded DNA, single-strandedDNA, and RNA genomes from natural viral assemblages. Appl. Environ. Microbiol. 76, 5039–5045 (2010)). Both methods rely on the denaturation and re-hybridization of cDNA chains. When single-stranded cDNA moves around in the solution, the more abundant sequence has a greater probability of finding a complementary sequence that matches it to re-hybridize. Therefore, when rehybridization reaches its limit, the remaining single-stranded cDNA represents the normalized sequence library.

[0158] The difference between these two methods lies in the method used to separate the single-stranded cDNA library from the rehybridized double-stranded cDNA molecules.

[0159] In the DSN method, an enzyme that specifically cuts double-stranded DNA is used to break down all double-stranded cDNA in a solution. The solution is then purified and cDNA sequences that exceed a certain length are size-selected. These sequences are then amplified using the polymerase chain reaction (PCR).

[0160] In the column method, the denatured and rehybridized cDNA library is passed through a heated column filled with hydroxyapatite particles. Hydroxyapatite preferentially binds to larger DNA molecules. The size of the bound DNA is controlled by the concentration of the phosphate buffer in which the cDNA library is dissolved. Therefore, the concentration of the phosphate buffer must be adjusted specifically for cDNA molecules within a certain sequence length range. The cDNA is eluted through the column using an increased concentration of phosphate buffer to extract the increased size DNA molecules. Since single-stranded cDNA is approximately half the size of the rehybridized cDNA, if the average cDNA sequence length is known, the elution of the single-stranded portion can be controlled. The elution that occurs is intended to enrich the single-stranded cDNA, which is then amplified using PCR.

[0161] In both the DSN and column methods, known adapters must be ligated to the ends of the cDNA prior to normalization to facilitate PCR amplification (so that appropriate primers can be used).

[0162] Since both methods are subtractive in nature, in which a large portion of the cDNA is depleted, the amount of starting cDNA usually needs to be higher than 1 μg for the DSN method and higher than 4 μg for the column method.

[0163] Since the DSN method uses an enzyme that cuts all double-stranded cDNAs, it can theoretically deplete low-abundance sequences with the parts that match high-abundance sequences. This effect can also increase the possibility of forming PCR chimeras. When incomplete single-stranded cDNA sequences serve as primers for other sequences, thus combining the sequences in a manner that does not exist in nature, PCR chimeras are formed. PCR chimeras represent false positives of new isoforms and are difficult to distinguish from true alternative isoforms. Verifying PCR chimeras usually requires in-depth biochemical assays.

[0164] Since the column method only allows separation of high and low abundance fractions within a narrow size range, it has a significant bias towards longer cDNA sequences. The result is a loss of representation of longer RNA sequences. Column method

[0165] As mentioned above, the hydroxyapatite column method relies on the denaturation and re-hybridization of cDNA chains. When single-stranded cDNA moves around in the solution, the more abundant sequence has a greater probability of finding a complementary sequence that matches it for re-hybridization. In the column method, the denatured and re-hybridized cDNA library passes through a heated column filled with hydroxyapatite particles. Hydroxyapatite preferentially binds to larger DNA molecules. Since single-stranded cDNA is approximately half the size of the re-hybridized cDNA, if the average cDNA sequence length is known, the elution of the single-stranded portion can be controlled. The elution that occurs is intended to enrich the single-stranded cDNA, which is then amplified using PCR.

[0166] Using 4 μg of cDNA as starting material as described by Andrews-Pfannkoch et al. The hydroxyapatite column method was performed as described in (Andrews-Pfannkoch, C., Fadrosh, DW, Thorpe, J. & Williamson, SJ Hydroxyapatite-mediated separation of double-stranded DNA, single-stranded DNA, and RNA genomes from natural viral assemblages. Appl. Environ. Microbiol. 76, 5039–5045 (2010). The hydroxyapatite column method did not produce usable yields when 2 μg or less of cDNA was used.

[0167] Since the column method is based on separation by size, this approach results in a loss of representation of longer RNA sequences (longer than 4 kb), which is observable in the length distribution before and after normalization. DSN Method

[0168] As for the column method, the DSN method relies on the denaturation and re-hybridization of cDNA strands. Since single-stranded cDNA moves around in the solution, the more abundant sequence has a greater probability of finding a complementary sequence to match it for re-hybridization. In the DSN method, an enzyme that specifically cuts double-stranded DNA is used to decompose all double-stranded cDNA in the solution.

[0169] The commercially available Evrogen Trimmer-2 cDNA Normalization Kit uses the DSN method. The kit was used according to the manufacturer's instructions with 1 μg of cDNA starting sample. However, in order to generate enough material for long-read RNA sequencing, it was found necessary to use 2 μg of cDNA.

[0170] The DSN method was found to completely eliminate high abundance RNAs, and this is expected to be true for RNAs with sequence similarity to high abundance RNAs. Thus, over-depletion was observed, in which high abundance RNAs were not only reduced in number, but completely removed from the sample. Table 1 shows over-depletion for levels 1 to 20, and shows a selection of lower levels (55, 64, 77, 92, and 98) in which RNAs were significantly reduced but not completely depleted. Table 1 – Overdepletion of high-abundance RNA using the DSN approach

[0171] Additionally, the DSN approach creates conditions that can generate artificial chimeric sequences, which can appear as false positives in gene predictions. Selective Amplification Method (Level-up)

[0172] This method is Figures 1 to 3 In general, refer to Figure 2 , a cDNA sample containing cDNA with known 5' and 3' pre-attached adapters is denatured to provide a single-stranded cDNA sample; the sample is then re-hybridized or re-associated to provide a mixture of single-stranded cDNA and double-stranded cDNA. The single-stranded cDNA represents low-abundance cDNA. The single-stranded cDNA in the re-association sample is then modified by adding oligonucleotides to the 5' and 3' pre-attached adapters. The cDNA sample is amplified using primers for these oligonucleotides. This process selectively increases the content of only single-stranded cDNA in the re-association sample, thereby increasing the content of low-abundance cDNA in the sample, thereby providing a normalized sample.

[0173] The input cDNA library for the selective amplification method is double-stranded, and the double-stranded templates each contain a 5' pre-attached adapter of a known nucleotide sequence and a 3' pre-attached adapter of a known nucleotide sequence. Since the method of the present invention is an additive method, a lower amount of starting cDNA is required compared to the normalization method of the prior art. In the DSNase method, a minimum of 1 μg of input cDNA is required, while in the column method, a minimum of 4 μg of input cDNA is required. In testing, it was found that the method of the present invention can be applied with as little as 20 ng of starting cDNA.

[0174] The input cDNA is combined with hybridization buffer and the solution is heated to the denaturation temperature, about 98 degrees Celsius, to produce a denatured single-stranded cDNA template. After 5 to 10 minutes, the solution is then lowered to the re-hybridization temperature, about 68 degrees Celsius. Depending on the desired normalization amount, the solution is incubated at this temperature for 0 to 24 hours, such as 3 to 10 hours. 7 hours is a typical duration for the re-association step. This step produces a re-association or re-hybridization sample containing the re-association double-stranded cDNA template and the re-association single-stranded cDNA template.

[0175] After incubation, add oligonucleotide dimer of the present invention (the inventor is called K-joint). These oligonucleotide dimers are discussed in more detail below. At this time, the solution can be incubated at 68 degrees Celsius for 0 to 1 hour. 5 minutes is the typical duration of this incubation step. Then the solution is reduced to the annealing temperature of the K-joint, which is usually at 40 to 60 degrees Celsius, such as 44 degrees Celsius. The solution is incubated at this temperature for 10 minutes to 2 hours, such as 25 minutes. This step anneals the K-joint with the single-stranded cDNA template after association.

[0176] After this incubation period, DNA ligase is added together with the ligation mixture. The solution is incubated at this same temperature for 0.5 to 2 hours, for example 1 hour, and then cooled to room temperature. This step results in the formation of the adapter-cDNA template connected, wherein the oligonucleotide from the K-joint is connected to each end of the single-stranded cDNA template after association. At this point, the cDNA can be purified (e.g., using Pronex or Ampure beads) or can be directly used for PCR amplification using primers based on the K-joint sequence.

[0177] Following PCR amplification, the cDNA is then purified using any suitable means, and the resulting cDNA represents the normalized cDNA library.

[0178] Prior to PCR amplification, the post-association double-stranded cDNA template may be removed, but following testing, the inventors have shown that leaving the double-stranded cDNA in solution does not negatively impact the normalization process.

[0179] In fact, the double-stranded cDNA after association can also be analyzed and used, for example, to obtain an estimate of gene expression. This involves an additional (PCR) selective amplification step using primers for known 5' pre-attached adapters and known 3' pre-attached adapters, wherein the molecular barcode is contained in the primer. First, PCR amplification is performed using primers based on the K-joint sequence, followed by a single PCR cycle using primers for known 5' pre-attached adapters and known 3' pre-attached adapters. PCR can be paused to add other primers for the final cycle. Alternatively, cDNA can be purified using primers based on the K-joint sequence after PCR, and a new PCR for a single cycle is performed using primers for known 5' pre-attached adapters and known 3' pre-attached adapters. In both cases, the product will be a mixture of two distinguishable parts consisting of a sequence derived from the double-stranded cDNA template after association and a sequence derived from the single-stranded cDNA template after association. Molecular barcode addition allows source molecules to be identified during sequencing analysis. This aspect of the present invention can be used for multiple sequencing. Design of oligonucleotide dimer complexes

[0180] The oligonucleotide dimer composition of the present invention comprises a front oligonucleotide dimer and a rear oligonucleotide dimer (a front K-linker and a rear K-linker), both of which anneal to the same strand of the associated single-stranded cDNA.

[0181] Reference Figure 3 and 3A Each K-linker contains two oligonucleotide sequences; one is called the linker-oligonucleotide (also called a "linker"; Figure 3 and 3A in ), and the other is called a lig-oligonucleotide (also called a "lig"; in Figure 3 and 3A It is called "LU adapter" in the literature.

[0182] The linker sequence contains regions complementary to the known 3' / 5' pre-attachment adapter sequences (which were previously added to the cDNA) and a region complementary to the lig.

[0183] If still Figure 3 and 3A As depicted in , a linker can be designed to have an overhang region at one end that is non-complementary to a known pre-attached adapter sequence, which is referred to as a "template overhang". The opposite end of the linker can have a similar overhang that is non-complementary to the lig sequence, which is referred to as a "lig overhang" or "lig-oligonucleotide overhang".

[0184] The purpose of the connexon is to anneal the pre-attached adaptor and lig of the single-stranded cDNA template after the association, and in such a way, in use, one end of the cDNA template is adjacent to one end of the lig. By locating the cDNA template and lig in this way, DNA ligase can be used to connect the cDNA template to the lig, so that the lig sequence is added to the end of the cDNA template. The lig sequence is added to the 5' and 3' ends of the single-stranded cDNA template. The front K-joint and the back K-joint are used to add these ligs to the 5' and 3' ends of the cDNA template respectively.

[0185] In the above method, once lig has been added to each end of the single-stranded cDNA template, primers based on the added lig sequence are used to selectively amplify the single-stranded cDNA portion that has been successfully connected to both the front lig sequence and the back lig sequence. In this way, PCR can be used to amplify only the low-abundance post-association single-stranded cDNA portion.

[0186] The specific structure of the K-linker provides advantages for the overall normalized performance, where the front and back K-linkers have different functions provided by their specific structural features.

[0187] The pre-K-adapter that binds to the 5' end of the single-stranded cDNA template (shown in the figure as the reverse complement of the original RNA sequence) can be designed so that the K-adapter does not act as a primer during PCR amplification.

[0188] To provide additional advantages, the front K-adapter does not have a blunt end on the lig side to allow the use of a DNA ligase that can perform blunt-end ligation during selective amplification. Providing the lig side of the K-adapter complex with a non-blunt end also avoids ligation to other K-adapter complexes or to double-stranded cDNA in solution.

[0189] Since the front linker has a 5' to 3' directionality pointing away from the template, the linker itself cannot act as a primer for the template. However, in some cases, the front lig or PCR primer can potentially anneal to the linker during PCR amplification and undergo polymerase extension to present the template side sequence of the linker. Providing a linker with an overhang on the template side avoids the extended lig / primer from acting as a primer for a cDNA sequence (i.e., a high abundance sequence) that does not have a lig sequence added to its end. Therefore, a template overhang structure for the back linker can be provided.

[0190] The post-K-linker that binds to the 3' end of the single-stranded cDNA template (shown in the figure as the reverse complement of the original RNA sequence) can be designed so that the K-linker does not act as a primer during PCR amplification. This can be achieved by using a template overhang, which is described above and in Figure 3 and 3AShown in.

[0191] For the same reason that the front K-adapter can be designed to have an overhang on the lig side, the rear K-adapter can be provided without a blunt end on the lig side. If the rear K-adapter complex has a blunt end on the lig side, in some cases, the lig can potentially ligate with other K-adapter complexes or with double-stranded cDNA in the sample, which may result in ligation of the cDNA template.

[0192] The overhang is also used to reduce the annealing temperature of the K-joint complex, so that it is less likely to act as a primer for each other during the PCR amplification using a higher annealing temperature. The overhang can reduce undesired initiation. As a supplement or alternative, the overhang can provide an indicator that can measure the amount of undesired initiation. For example, if the K-joint complex with its overhang can trigger the template, the overhang sequence can be seen in the sequencing data, and it can be inferred that it is a product of undesired initiation.

[0193] All overhangs should ideally be about 1 bp to about 20 bp, preferably 3 bp. Longer overhangs can be used, but they will make the design more difficult because there are fewer sequence compositions that prevent undesired priming as the overhangs become longer.

[0194] The complementary regions between the cDNA template and the linker and between the linker and the lig should be long enough to anneal at the active temperature of the ligase to be used. Gap repair ligases are particularly preferred for this process, which typically require about five or more complementary bases on either side of the ligation site.

[0195] The combined length of the linker and the corresponding lig should ideally be less than about 300 bp, preferably less than about 200 bp, to reduce carry over during the purification process. Example sequence representation:

[0196] All example sequence structures are provided in a 5' to 3' orientation. Figure 3B Shown in. Pre-linker: OOOO0XXXXXXXXXXXXXXXFFFFFFFFFFFFFOOOOO Front lig: FFFFFFFFFFFFFFFFFFFF Post-linker: OOOOOBBBBBBBBBBBBBBBXXXXXXXXXXXXXXXOOOOO After lig: BBBBBBBBBBBBBBBBBBBB X - Nucleotide complementary to the 5' / 3' pre-attached adaptor O-overhang sequence F-sequence complementary to the front lig to be connected F-prelig sequence B-sequence complementary to the rear lig to be connected B-sequence of the rear lig Oligonucleotides used in the selective amplification method Examples of nucleotides and single-stranded cDNA templates:

[0197] Primer sequences from NEB / PacBio cDNA synthesis kit Iso-Seq Expression Fwd: Iso-Seq Expression Rev: Pre-linker: Pre-Lig and Primer: Post-linker: Rear Lig: Primers:

[0198] Overhangs are underlined. Regions of complementarity between oligonucleotides are shown in bold.

[0199] Single-stranded cDNA template: AAGCAGTGGTATCAACGCAGAGNNNNNNNNNNNNNNNNCAACCCTGCGACTTCATTGCC (i.e., the sequence of the 5'-3' - 5' pre-attached adaptor, the sequence of the cDNA represented by N, and the sequence of the 3' pre-attached adaptor; the regions complementary to the front-ligation oligonucleotide and the back-ligation oligonucleotide are shown in bold). Preparation of oligonucleotide dimers

[0200] Once designed, dimers can be prepared using standard techniques well known in the art. Amplification of oligonucleotide-cDNA templates

[0201] After ligating the ligase to the 3' and 5' ends of the cDNA template, the cDNA of the resulting solution can be purified using a suitable cDNA purification method. The purification step can be skipped, but skipping can result in reduced efficiency of PCR amplification.

[0202] After purification or after ligation, the resulting material can be PCR amplified using forward and reverse primers based on the lig sequence. Primer sequences can be selected that have a higher annealing temperature than the complementary region of the template / linker / lig to avoid unwanted priming.

[0203] If desired, the optimal number of PCR cycles can be determined by first performing a qPCR experiment to determine the inflection point of the amplification curve.

[0204] The cDNA resulting from PCR amplification can then be purified and used as input for any downstream process (eg, sequencing). Validation of Selectively Amplified Samples

[0205] The effect of selective amplification or normalization can be measured indirectly by measuring the length distribution of the cDNA library using gel electrophoresis, or directly using sequencing.

[0206] To identify the effect of normalization using gel electrophoresis, the length distribution profile of the input cDNA can be compared to the normalized cDNA. The input cDNA will typically show a peak along the length distribution, which corresponds to the high abundance transcript sequence ( Figure 4 ).

[0207] The normalized cDNA will have a length distribution similar to a normal distribution without a spike ( Figure 5 ). This represents a uniform distribution throughout the transcript sequence.

[0208] When sequencing is used to directly measure normalization, the preferred sequencing method is long read sequencing. This allows identification of different isoforms. The result of the direct measurement can be a graph of the number of reads per gene, or a saturation graph showing the number of new genes identified as the depth of sequencing increases.

[0209] To validate the method, the inventors performed nanopore cDNA sequencing on both the input cDNA library and the cDNA library generated by the method. They compared the saturation curves of each library ( Figure 6 ). The curve of the method of the present invention is many times higher than the curve of the input cDNA. This indicates that the cDNA library is more evenly distributed. Applications of Selective Amplification

[0210] The main application of the selective amplification method of the present invention is to improve the discovery and detection of low-abundance genes and isoforms. When the method of the present invention is combined with sequencing, the sampling efficiency of identifying all unique genes in a sample is improved.

[0211] The method of the present invention can be applied to any double-stranded cDNA library with known adapters at the ends and whose length can be amplified by PCR method. This means that it can be used for DNA sequencing.

[0212] Another important aspect of the method of the present invention is the ability to select cDNA with a known lig sequence. This aspect can be applied to single-cell sequencing, where adapters are necessary to assign reads to individual cells. In this application, cDNA sequences without cell identification barcodes / adapters appear in the cDNA library. These are called template switching oligonucleotide (TSO) artifacts and are undesirable in single-cell sequencing projects because they cannot be assigned to the cell of origin. The method of the present invention can be applied to a short re-hybridization step to select only cDNA sequences with the desired lig sequence, thereby effectively limiting the sequencing of TSO artifacts. The method of the present invention can also be performed with a single-cell cDNA library to both remove TSO artifacts and increase the transcriptome coverage of each cell. discuss

[0213] The selective cDNA amplification method of the present invention represents an innovative approach to achieve greater non-targeted discovery / detection of low-abundance nucleic acids. Because it is non-targeted, it is not necessary to know which sequences are low or high in abundance in a given sample.

[0214] The difference between it and the existing normalization method is that the addition method is used while the current method uses the depletion method. The addition method allows the method of the present invention to use a significantly smaller amount of starting cDNA. It helps in the detection of low-abundance nucleic acids and provides greater credibility for determining the absence of specific nucleic acids. The addition method also prevents excessive depletion and artificial chimerism, which are unique to the DSNase method. The method of the present invention also has a length deviation only to the extent that the PCR amplification has a length deviation. Therefore, it can be successfully operated with longer cDNA molecules.

[0215] The example systems, methods, and actions described in the previously presented embodiments are illustrative, and in alternative embodiments, certain actions may be performed in a different order, performed in parallel with each other, omitted entirely, and / or combined between different example embodiments, and / or certain additional actions may be performed without departing from the scope and spirit of the various embodiments. Therefore, such alternative embodiments are included in the embodiments described herein.

[0216] Although specific embodiments have been described in detail above, this description is for illustrative purposes only. Therefore, it should be understood that many of the above aspects are not intended to be required or essential elements unless otherwise expressly stated.

[0217] In addition to those described above, those skilled in the art who have the benefit of this disclosure may make modifications to the disclosed aspects of the example embodiments and may make equivalent components or acts corresponding to the disclosed aspects of the example embodiments without departing from the spirit and scope of the embodiments defined in the following claims, which scope is accorded the broadest interpretation so as to encompass such modifications and equivalent structures. Example

[0218] The present invention will be further understood by reference to the following experimental examples. Example 1 Using a novel transcriptome discovery platform to detect unique RNA signatures in epithelial cancers

[0219] The inventors have validated the use of Level-up using RNA from cancer cell lines and blood samples. In these preliminary studies, thousands of previously unannotated isoforms were discovered that have the potential to be used as cancer biomarkers. Figure 7 Clustering of cancer samples (blood samples taken from people who have been diagnosed with breast cancer) and control samples (blood samples taken from people who have not been diagnosed with breast cancer) based on the 13,573 transcripts identified in these initial studies is shown. The cancer samples can be easily distinguished from the control samples.

[0220] Example protocols for discovering unique RNA signatures in epithelial cancers are provided: Phase 1

[0221] Blood processing validation was performed using 5 samples from 5 individuals (25 samples total). Blood was collected and stored in PAXgene RNA blood tubes (blood capacity of each tube was 2.5 ml). After collection, the tubes were shipped on dry ice from the local collection center on the same day.

[0222] For each biological replicate, 5 collected tubes were treated according to different mock treatment conditions. These include: 1. Process the tubes on the day of collection. 2. Process the tubes 72 hours after collection. 3. Process the tubes 2 weeks after collection. 4. Process the tubes 4 weeks after collection. 5. Store the tubes long term at -80C and process within 4 to 12 months.

[0223] RNA was extracted from PAXgene tubes using the Qiagen PAXgene Blood RNA Kit (Cat. No. / ID: 762164). The single cell / low input cDNA synthesis & amplification module (E6421S) converts the resulting RNA into cDNA. The resulting cDNA is processed by normalization as described herein. The normalized cDNA library is then sequenced using an Oxford Nanopore Technologies Minion MK1C sequencer. Phase 2

[0224] Having optimized sample processing and ensured the quality of the data generated, the research will then progress to patient samples. Inclusion and Exclusion Criteria

[0225] Inclusion criteria: All premenopausal female patients aged 18 years and above but less than 50 years with invasive breast cancer or benign conditions (B2).

[0226] Exclusion criteria: women with any history of cancer or concurrently with other types of cancer, autoimmune diseases, and those unable or unwilling to give informed consent. Identification and recruitment of patients for prospective cohort studies

[0227] Eligible patients were identified from a multidisciplinary team (MDT) meeting based on the triple assessment results. When patients reviewed the results with their breast surgeon in the clinic, they were invited to participate in the study and given a patient information sheet (PIS).

[0228] The research team will follow up with patients who were given PIS after 72 hours. If patients choose to participate in the study, they will meet with the research nurse and sign the informed consent form, or give informed consent remotely. Anonymization

[0229] Once enrolled in the study, the patient is assigned a unique ID number by a research nurse, which is linked to background data via a secure database, where it is used to collect and identify the sample. Background data collection

[0230] The following data were collected for each patient: baseline demographics, menopausal status, past medical history, medication history, drug and alcohol intake, BMI, and family history.

[0231] Triple assessment including presentation, examination findings, imaging, and biopsy results were collated. Postoperative histology was also collected. Any additional prognostic information, such as the use of biomolecular assays, was also collected. Blood sample collection

[0232] On the morning of surgery, blood samples were obtained and stored at -20°C before being shipped to the processing center. For any patients not having surgery, such as those with fibroadenomas, they were invited to the day unit where blood tests were performed by the same team. The sample volume was 20 ml, with a maximum of 30 ml collected on the day. Blood sample processing

[0233] Blood samples from collection tubes were processed in a laminar flow class 2 safety cabinet using an RNA extraction kit provided by Qiagen. After RNA extraction, the remaining solid waste was autoclaved and discarded by a human medical waste removal service provider. Liquid waste will be decontaminated and disposed of in accordance with the University of Edinburgh guidelines for the use of human samples. Blood samples that were not processed on the day of arrival were stored in a secure -80°C freezer. Sealed containers were used to transport blood samples between the freezer and the safety cabinet. Data analysis

[0234] Transcriptomics data are stored and processed in secure Amazon Web Services cloud repositories and servers. No personal information is stored in the same computing locations.

[0235] Oxford Nanopore Technologies (ONT) Minion sequencing machine was used to sequence cDNA samples. This outputs the raw data as fast5 files. The latest high-precision basecaller (Bonito) from ONT was run to convert to fastq sequence files. Nanopore reads were quality filtered using seqkit and adapters were removed using pychopper. The trimmed reads were then mapped to the HG38 (or later) human reference genome assembly using Minimap2. Stages 3 to 4

[0236] The study was amended to initially increase the sample size and subsequently expand to colorectal and ovarian cancers. Statistical considerations and sample size

[0237] For Phase 1 of the study, five samples were processed. This was sufficient to generate technical and biological replicates and to perform an initial quality assessment.

[0238] Data handling and processing were performed as in Phase 2.

[0239] For Phase 2 of the study, 30 samples from benign patients and 30 samples from cancer patients were initially collected and processed. This data will inform subsequent sample sizes for Phases 3 and 4 of the overall study. Ethical approval

[0240] Ethical approval was sought from the National Research Ethics Service before the study began. Informed consent process

[0241] Written informed consent was obtained from all patients participating in this study for the collection of their demographic and clinical data and for the subsequent storage of blood samples and transcriptomic material and data. If the patient chose, they could give their consent by telephone. The patient could withdraw their consent at any time. Anyone deemed unable to give informed consent was excluded from this study. Quality Control

[0242] To assess that samples were correctly labeled and that there were no processing issues in each sample, computer quality control testing was performed. This testing consisted of checking for known genes that should be present in all samples and performing cluster analysis to assess for outliers that could represent sample contamination.

[0243] It also looks for genes present in the sample data that shouldn't be there - for example, genes from other species. Example 2 Blood sample collection procedure for full-length RNA extraction

[0244] cDNA is a DNA synthesized from an RNA template. Therefore, the quality of a cDNA sample is related to the RNA from which the cDNA sample is reverse transcribed. Prior art methods for treating blood samples before RNA extraction involve thawing frozen blood samples overnight. The inventors have found that this results in significant RNA degradation, which has a negative impact on long read sequencing. The inventors used the following protocol to treat blood samples before RNA extraction to minimize degradation and optimize RNA extraction for long read sequencing.

[0245] Blood sample collection procedure for full-length RNA extraction • At room temperature (18 to 25°C), draw 2.5 ml of blood into a PAXgene blood RNA tube. Immediately after blood collection, gently invert the blood tube 10 times. If RNA is to be extracted on the same day as sample collection, store the blood sample upright at room temperature (18 to 25°C) for 2 to 3 hours and then immediately proceed with RNA extraction using the PAXgene Blood RNA Kit.

[0246] If RNA extraction is not performed on the day of blood collection, follow the instructions below regarding sample freezing, storage, and thawing: Blood samples should be stored at -20°C or below immediately after collection. For long-term storage, freeze blood samples at -20°C for 24 hours before transferring to a -70°C or -80°C freezer. If the samples are to be transferred to a different location, ship them on dry ice to ensure that the samples remain frozen during transport. On the day of RNA extraction, thaw the blood sample by placing the sample tube upright on a stand and incubate at room temperature (18 to 25°C) for 1 to 3 hours. Once the blood is completely thawed, gently invert the sample tube 10 times, incubate at room temperature for another 2 hours, and then immediately perform RNA extraction using the PAXgene Blood RNA Kit.

[0247] As shown in Table 2, a shorter thawing time (3 hours) improved RNA integrity relative to overnight thawing. The RNA integrity number was calculated according to Schroeder et al. (The RIN: an RNA integrity number for assigning integrity values ​​to RNA measurements. BMC Molecular Biology 7, 3 (2006). https: / / doi.org / 10.1186 / 1471-2199-7-3 .) Calculated. Samples were placed at -20°C within 12 hours for the "-20°C 2 weeks; overnight thaw" and "-20°C 1 month; 3 hours thaw" tests (second and third tests). 5 individuals were used for each test. All thawing was performed at room temperature. Table 2 – Blood sample storage conditions and RNA integrity Blood sample storage conditions RNA integrity number Fresh 8.3±0.2 -20℃ for 2 weeks; thaw overnight 7.8±0.2 -20℃ for 1 month; 3 hours to defrost 8.4±0.3 4℃ for 3 days, then -20℃ for 2 months; 3 hours to thaw 8.0±0.3

[0248] The scope of the present invention is not limited by the specific embodiments described herein. In fact, according to the above description and accompanying drawings, various modifications of the present invention (except those described herein) will become obvious to those skilled in the art. Such modifications are intended to fall within the scope of the appended claims. In addition, all embodiments described herein are considered to be widely applicable and can be appropriately combined with any and all other consistent embodiments.

[0249] Various publications are cited herein, the disclosures of which are incorporated by reference in their entireties.

Claims

1. A method for discovering disease biomarkers, comprising: (i) providing a first cDNA sample from a subject suffering from a disease and a second cDNA sample from a subject not suffering from the disease; (ii) normalizing the first cDNA sample and the second cDNA sample; (iii) sequencing the normalized first cDNA sample and the second cDNA sample; as well as (iv) comparing the sequencing outputs of the first cDNA sample and the second cDNA sample to discover disease biomarkers.

2. A method for diagnosing a disease in a subject, comprising: (i) providing a cDNA sample from the subject; (ii) normalizing the cDNA samples; as well as (iii) sequencing the normalized cDNA sample, wherein the sequencing output is used to identify whether the subject has the disease.

3. The method of claim 2, comprising comparing the sequencing output of the normalized cDNA sample with one or more reference sequences or with the sequencing output of one or more control samples, optionally wherein the one or more control samples are from one or more subjects suffering from the disease and / or subjects not suffering from the disease.

4. The method of claim 2 or 3, wherein the cDNA sample is obtained from a biological fluid or a fluid or lysate produced from a biological material, optionally wherein the cDNA sample is obtained from blood.

5. The method of any one of claims 2 to 4, further comprising reporting the outcome of the method to the subject.

6. The method of any one of claims 1 to 5, wherein the disease is cancer.

7. The method of any preceding claim, wherein normalization comprises increasing the amount of low-abundance cDNA in each cDNA sample.

8. The method of claim 1, 6 or 7, wherein the first cDNA sample and the second cDNA sample comprise double-stranded cDNA templates, each template having a known 5' pre-attached adaptor and a known 3' pre-attached adaptor; And wherein normalizing the first cDNA sample and the second cDNA sample comprises: (i) denaturing the cDNA sample to generate a single-stranded cDNA template; (ii) re-associating the cDNA sample to generate a mixture of associated single-stranded cDNA templates and associated double-stranded cDNA templates; (iii) annealing a 5' adapter complex to the 5' pre-attached adapter of at least one post-association single-stranded cDNA template, and annealing a 3' adapter complex to the 3' pre-attached adapter of the same post-association single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide; (iv) ligating an oligonucleotide from the 5' adapter complex to the 5' pre-attached adapter of the associated single-stranded cDNA template, and ligating an oligonucleotide from the 3' adapter complex to the 3' pre-attached adapter of the same associated single-stranded cDNA template; and (v) selectively amplifying the cDNA sample using primers specific for the ligated oligonucleotides.

9. The method of any one of claims 2 to 7, wherein the cDNA sample comprises double-stranded cDNA templates, each template having a known 5' pre-attached adaptor and a known 3' pre-attached adaptor; And wherein normalizing the cDNA sample comprises: (i) denaturing the cDNA sample to generate a single-stranded cDNA template; (ii) re-associating the cDNA sample to generate a mixture of associated single-stranded cDNA templates and associated double-stranded cDNA templates; (iii) annealing a 5' adapter complex to the 5' pre-attached adapter of at least one post-association single-stranded cDNA template, and annealing a 3' adapter complex to the 3' pre-attached adapter of the same post-association single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide; (iv) ligating an oligonucleotide from the 5' adapter complex to the 5' pre-attached adapter of the associated single-stranded cDNA template, and ligating an oligonucleotide from the 3' adapter complex to the 3' pre-attached adapter of the same associated single-stranded cDNA template; and (v) selectively amplifying the cDNA sample using primers specific for the ligated oligonucleotides.

10. The method of claim 8 or 9, wherein: (A) The 5' adaptor complex is a pre-oligonucleotide dimer comprising: (i) a pre-lig-oligonucleotide for ligation with a 5' pre-attached adaptor of said (after association) single-stranded cDNA template; and (ii) a pre-attachment oligonucleotide for annealing to the 5' pre-attachment adaptor and the pre-lig-oligonucleotide, the pre-attachment oligonucleotide comprising a region complementary to the 5' pre-attachment adaptor and a region complementary to the pre-lig-oligonucleotide, such that upon annealing, the ends of the pre-lig-oligonucleotide are adjacent to the ends of the 5' pre-attached adaptor to enable ligation of the pre-lig-oligonucleotide to the 5' pre-attached adaptor at the ligation site; and (B) The 3' adaptor complex is a post-oligonucleotide dimer comprising: (i) a post-lig-oligonucleotide for ligation with the 3' pre-attached adaptor of the (after association) single-stranded cDNA template; and (ii) a post-ligation-oligonucleotide for annealing to the 3' pre-attachment adaptor and the post-lig-oligonucleotide, the post-ligation-oligonucleotide comprising a region complementary to the 3' pre-attachment adaptor and a region complementary to the post-lig-oligonucleotide, Such that upon annealing, the end of the post-lig-oligonucleotide is adjacent to the end of the 3' pre-attached adaptor to enable ligation of the post-lig-oligonucleotide to the 3' pre-attached adaptor at the ligation site.

11. The method of claim 10, wherein: (A) The pre-linker-oligonucleotide comprises: (i) a template overhang region at the end of the pre-ligation-oligonucleotide proximal to the region complementary to the 5' pre-attached adaptor, said template overhang region being non-complementary to a corresponding region of the (after association) single-stranded cDNA template; and / or (ii) a lig-oligonucleotide overhang region at the end of the pre-lig-oligonucleotide proximal to the region complementary to the pre-lig-oligonucleotide, said lig-oligonucleotide overhang region being non-complementary to the corresponding region of the pre-lig-oligonucleotide; and / or (B) The post-linking oligonucleotide comprises: (i) a template overhang region at the end of the post-ligation-oligonucleotide proximal to the region complementary to the 3' pre-attached adaptor, said template overhang region being non-complementary to a corresponding region of the (after association) single-stranded cDNA template; and / or (ii) a lig-oligonucleotide overhang region at the end of the second ligation-oligonucleotide proximal to the region complementary to the second lig-oligonucleotide, said lig-oligonucleotide overhang region being non-complementary to the corresponding region of the second lig-oligonucleotide.

12. The method of any preceding claim, wherein sequencing comprises using long read sequencing.

13. The method of any one of claims 1, 6 to 8, or 10 to 12, wherein the first cDNA sample and the second cDNA sample are obtained from a biological fluid or a fluid or lysate produced from a biological material, optionally wherein the first cDNA sample and the second cDNA sample are obtained from blood.

14. The method of any one of claims 1, 6 to 8 or 10 to 13, further comprising, prior to step (i), extracting RNA from a biological fluid or a fluid or lysate produced from a biological material and synthesizing cDNA using the RNA as a template, wherein the biological fluid or the biological material is from the subject suffering from the disease and the subject not suffering from the disease.

15. The method of any one of claims 2 to 7 or 9 to 12, further comprising, prior to step (i), extracting RNA from a biological fluid from the subject or a fluid or lysate produced from a biological material and synthesizing cDNA using the RNA as a template.

16. The method of any one of claims 1, 6 to 8, or 10 to 14, wherein the disease biomarker is a cDNA sequence that is present in the first cDNA sample but not in the second cDNA sample, or a cDNA sequence that is present in the second cDNA sample but not in the first cDNA sample.

17. A method for processing a blood sample, comprising: (i) storing the blood sample at -15°C or below; (ii) thawing the blood sample at 5 to 30° C. for at least 1 hour; as well as (iii) RNA was extracted from thawed blood samples.

18. The method of claim 17, wherein: (i) storing the blood sample at -20°C or below; (ii) thawing the blood sample at 18 to 25°C; and / or (iii) Thawing the blood sample for 1 to 4 hours.

19. The method of claim 17 or 18, wherein the blood sample is stored at or below -15°C or at or below -20°C within 12 hours of collection.

20. The method of any one of claims 2 to 7, 9 to 12 or 15, wherein the cDNA sample is obtained from blood by following the steps of the method of any one of claims 17 to 19 and synthesizing cDNA using the extracted RNA as a template.

21. The method of any one of claims 1, 6 to 8, 10 to 14 or 16, wherein the first cDNA sample and the second cDNA sample are obtained from blood by following the steps of the method of any one of claims 17 to 19 and synthesizing cDNA using the extracted RNA as a template.

22. Use of an oligonucleotide dimer composition in a method for discovering cancer biomarkers, the oligonucleotide dimer composition being used for selectively amplifying single-stranded cDNA by ligating oligonucleotides to the 5' end and 3' end of a post-association single-stranded cDNA template having a known 5' pre-attached adaptor and a 3' pre-attached adaptor, wherein the composition comprises: (A) a pre-oligonucleotide dimer comprising: (i) a pre-lig-oligonucleotide for ligation with a 5' pre-attached adaptor of the associated single-stranded cDNA template; and (ii) a pre-ligation-oligonucleotide for annealing to the 5' pre-attachment adaptor and the pre-lig-oligonucleotide, the pre-ligation-oligonucleotide comprising a region complementary to the 5' pre-attachment adaptor and a region complementary to the pre-lig-oligonucleotide, such that upon annealing, the ends of the pre-lig-oligonucleotide are adjacent to the ends of the 5' pre-attached adaptor to enable ligation of the pre-lig-oligonucleotide to the 5' pre-attached adaptor at the ligation site; and (B) a rear oligonucleotide dimer comprising: (i) a post-lig-oligonucleotide for ligation with the 3' pre-attached adaptor of the associated single-stranded cDNA template; and (ii) a post-ligation-oligonucleotide for annealing to the 3' pre-attachment adaptor and the post-lig-oligonucleotide, the post-ligation-oligonucleotide comprising a region complementary to the 3' pre-attachment adaptor and a region complementary to the post-lig-oligonucleotide, Such that upon annealing, the end of the post-lig-oligonucleotide is adjacent to the end of the 3' pre-attached adaptor to enable ligation of the post-lig-oligonucleotide to the 3' pre-attached adaptor at the ligation site.

23. The use according to claim 22, wherein: (A) The pre-linker-oligonucleotide comprises: (i) a template overhang region at the end of the pre-ligation-oligonucleotide proximal to the region complementary to the 5' pre-attached adaptor, said template overhang region being non-complementary to a corresponding region of the post-association single-stranded cDNA template; and / or (ii) a lig-oligonucleotide overhang region at the end of the pre-lig-oligonucleotide proximal to the region complementary to the pre-lig-oligonucleotide, said lig-oligonucleotide overhang region being non-complementary to the corresponding region of the pre-lig-oligonucleotide; and / or (B) The post-ligation oligonucleotide comprises: (i) a template overhang region at the end of the post-ligation-oligonucleotide proximal to the region complementary to the 3' pre-attached adaptor, said template overhang region being non-complementary to a corresponding region of the post-association single-stranded cDNA template; and / or (ii) a lig-oligonucleotide overhang region at the end of the second ligation-oligonucleotide proximal to the region complementary to the second lig-oligonucleotide, said lig-oligonucleotide overhang region being non-complementary to the corresponding region of the second lig-oligonucleotide.

24. A system or test kit for discovering disease biomarkers, comprising: (a) one or more test devices for normalizing a first cDNA sample from a subject suffering from a disease and a second cDNA sample from a subject not suffering from the disease, and sequencing the normalized first cDNA sample and the second cDNA sample; (b) processor; and (c) a storage medium comprising a computer application which, when executed by the processor, is configured to: (i) obtaining the sequences of the first cDNA sample and the second cDNA sample determined on the one or more test devices, (ii) calculating whether there is a cDNA sequence present in the first cDNA sample but not in the second cDNA sample, or whether there is a cDNA sequence present in the second cDNA sample but not in the first cDNA sample; as well as (iii) outputting the result of step (ii) from the processor.

25. A system or test kit for diagnosing a disease in a subject, comprising: (a) one or more test devices for normalizing a cDNA sample from a subject and sequencing the normalized cDNA sample, (b) processor; and (c) a storage medium comprising a computer application which, when executed by the processor, is configured to: (i) obtaining the sequence of the cDNA sample determined on the one or more test devices, (ii) calculating the presence or absence of a cDNA sequence, wherein the presence or absence of the cDNA sequence is associated with the disease; and (iii) outputting from the processor whether the subject suffers from the disease.