Methods and products for identifying biomarkers
The method enhances RNA-based disease screening by normalizing cDNA samples to improve sensitivity and specificity, addressing RNA degradation and housekeeping gene interference, enabling the detection of disease-specific biomarkers and isoforms.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- WOBBLE GENOMICS LTD
- Filing Date
- 2023-10-20
- Publication Date
- 2026-05-19
AI Technical Summary
Current methods for RNA-based disease screening, particularly in blood samples, face challenges such as rapid RNA degradation, high expression of housekeeping genes, and limited detection of RNA isoforms, leading to insufficient sensitivity and specificity for disease biomarker detection.
A method involving the preparation and normalization of cDNA samples to achieve a more uniform sequence distribution, followed by sequencing and comparison of these samples to discover disease biomarkers, including steps for amplification and analysis to enhance detection breadth and reduce redundancy.
The method improves the consistency, sensitivity, and specificity of RNA-based disease screening by facilitating the detection of disease-specific genes and isoforms, reducing data storage requirements, and enabling the identification of novel biomarkers for disease diagnosis and treatment.
Smart Images

Figure 2026515558000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to methods, compositions, systems and kits for discovering disease biomarkers and diagnosing diseases. The present invention also relates to methods for processing blood samples for use in biomarker discovery and detection.
Background Art
[0002] There are many forms of screening for diseases, and recently, increasing interest has been paid to molecular screening techniques. For example, the detection of cancer-specific RNA in a simple blood test has the potential to provide a relatively non-invasive screening format that is readily acceptable and prognostic biomarkers (Larson, M.H. et al. A comprehensive characterization of the cell-free transcriptome reveals tissue- and subtype-specific biomarkers for cancer detection. 1-11). In addition, this technique has become affordable, allowing for regular screening intervals and assisting in early diagnosis. However, the handling and processing of samples for RNA sequencing are difficult in a clinical setting. In addition, it has been found to be difficult to develop methods that are sufficiently sensitive and specific. As a result, there is currently no practical and accurate blood RNA detection platform for clinical use in breast cancer or ovarian cancer diagnosis.
[0003] Regarding the handling of blood, the greatest concern is to maintain the integrity of RNA. It is well known that RNA degrades rapidly. In addition, blood and other tissue types contain enzymes that actively degrade RNA. This results in a rapid decrease in both the amount and quality of RNA in blood after collection. Therefore, the RNA content always degrades through the processes of collecting, storing and transporting blood.
[0004] RNA expression is characterized by broad expression levels across specific genes. There is a set of highly expressed genes known as housekeeping genes. While these genes do not provide useful information, they typically constitute more than half of the total RNA in any given sample. This problem is even more extreme in blood, where globin RNA and ribosomal RNA typically account for more than 95% of the total RNA (Harrington, CA et al. RNA-Seq of human whole blood: Evaluation of globin RNA depletion on Ribo-Zero library method. Sci. Rep. 10, 1-12 (2020)). Therefore, a prominent problem with current detection methods is that they rely on either depletion of these genes or targeting subsets of genes. These solutions result in a lack of detection breadth across the entire blood transcriptome, which may explain why this approach has not been successful.
[0005] A further problem is how well the data represents the actual complexity and biological variability. While several RNA detection assays exist, until recently these methods only allowed for the detection of small fragments of individual RNA molecules. This is problematic because each gene can be transcribed into numerous different RNA isoforms. Due to multiple combinations of transcription start sites, stop sites, and alternative splicing, genes can be transcribed into numerous different RNA isoforms (Harrow, J. et al. GENCODE: producing a reference annotation for ENCODE. 7, 1-9 (2006)). These alternative transcripts often have distinct functions and are often associated with specific cell and tissue types. Consequently, detecting only small fragments of RNA does not tell the whole picture. [Overview of the project]
[0006] The inventors have developed a technique for handling long-read RNA sequencing in a manner that can provide better consistency, sensitivity, and specificity for RNA-based disease screening. RNA or cDNA samples are typically dominated by sequences of highly expressed genes, which can negatively impact the analysis of the sample. The present invention provides a method and device for preparing a processed nucleic acid sample with a more uniform sequence distribution, which can then be analyzed for the discovery and / or detection of disease biomarkers. Normalization increases the efficiency of sequencing for transcript discovery and / or detection. Firstly, genes and isoforms that are specific to the condition in question become easier to detect, and secondly, there is less redundancy in the generated data, reducing data storage requirements. A method designed to protect RNA in blood samples is also provided. The methods may be combined to further improve the ability to discover and / or detect disease biomarkers.
[0007] Various embodiments have been devised to allow for advantageous combinations, and it should be noted that all such combinations are envisioned within the scope of the present invention. Furthermore, it should be understood that options described in relation to one area of improvement, such as sample type or disease, may be applied mutatis mutandis to other areas as appropriate.
[0008] The present invention (i) A step of preparing a first cDNA sample and a second cDNA sample, (ii) A step of normalizing the first and second cDNA samples, (iii) the step of sequencing the normalized first and second cDNA samples, and (iv) A step to discover disease biomarkers by comparing the sequencing outputs of the first and second cDNA samples. This provides a method for discovering disease biomarkers, including those mentioned above.
[0009] The first cDNA sample and the second cDNA sample may be from the same subject. In a particular embodiment, the first cDNA sample is from the subject before disease treatment, and the second cDNA sample is from the same subject after disease treatment. In a particular embodiment, the first cDNA sample is from the subject before disease treatment, and the second cDNA sample is from the same subject after disease treatment has been initiated and / or after disease treatment has been completed.
[0010] The first cDNA sample and the second cDNA sample may be samples from different subjects. In certain embodiments, the first cDNA sample and the second cDNA sample are samples from subjects with the same disease but at different grades or stages. In certain embodiments, the first cDNA sample and the second cDNA sample are samples from subjects with different diseases, for example, different types of cancer. In certain embodiments, the first cDNA sample is a sample from a subject with the disease, and the second cDNA sample is a sample from a subject without the disease.
[0011] Therefore, the present invention is (i) A step of preparing a first cDNA sample from a subject with the disease and a second cDNA sample from a subject without the disease, (ii) A step of normalizing the first and second cDNA samples, (iii) the step of sequencing the normalized first and second cDNA samples, and (iv) A step to discover disease biomarkers by comparing the sequencing outputs of the first and second cDNA samples. This provides a method for discovering disease biomarkers, including those mentioned above.
[0012] Discovering disease biomarkers means identifying novel biomarkers (indicators of biological status) for a particular disease, for example, revealing previously unknown biomarkers and developing trials for that disease. Disease biomarkers may be suitable for use in diagnosing a disease, characterizing a disease, predicting a response to treatment, detecting low-level residual disease, and / or determining the prognosis of a disease. Characterization means classifying and evaluating a disease. Determining the prognosis refers to predicting possible outcomes of a disease for a subject. In certain embodiments, disease characterization and / or prognosis determination includes determining the grade and / or stage of the disease. In further embodiments, disease characterization includes determining the subtype of the disease. Disease biomarkers may be suitable for use in indicating that a subject with a particular disease is likely to benefit from a particular treatment.
[0013] In certain embodiments, the disease is cancer. Therefore, in certain embodiments, characterizing and / or determining the prognosis of cancer includes determining the presence or absence of metastasis. Metastatic or metastatic disease is the spread of cancer from one organ or part to another non-adjacent organ or part. New occurrences of disease thus produced are referred to as metastases. Characterizing and / or determining the prognosis of the disease may also include predicting biochemical recurrence and / or determining whether the cancer is high-grade and / or whether the cancer has spread to the lymph nodes. High-grade cancer refers to cancer that is rapidly growing, more likely to spread, more likely to recur, and / or resistant to treatment.
[0014] According to related embodiments of the present invention, (i) A step of preparing a first cDNA sample from the subject at a first time point and a second cDNA sample from the subject at a second time point. (ii) A step of normalizing the first and second cDNA samples, (iii) the step of sequencing the normalized first and second cDNA samples, and (iv) A step of comparing the sequencing outputs for the first and second cDNA samples, Methods for monitoring the subject are provided, including the following.
[0015] Monitoring subjects may include monitoring their response to disease treatment. The first time point may be before the treatment begins, and the second time point may be during or after the treatment. Comparing the sequencing outputs of the first and second cDNA samples may provide an indicator of whether the treatment was successful. For example, the presence or absence of disease biomarkers may indicate whether the treatment was successful. Comparing the sequencing outputs of the first and second cDNA samples may include comparing them to each other and / or comparing the sequencing outputs from a reference sample.
[0016] According to all aspects of the present invention, a disease biomarker may be a cDNA sequence. A cDNA sequence may correspond to an RNA sequence. A cDNA / RNA sequence may correspond to a protein or a peptide. In certain embodiments, the method further includes the step of identifying an RNA, transcript model, gene, protein and / or peptide corresponding to a cDNA sequence. Therefore, a disease biomarker may be a cDNA molecule (of a specific sequence), a DNA molecule (of a specific sequence), an RNA molecule (of a specific sequence), a transcript model, a protein or a peptide.
[0017] In certain embodiments, the method includes the step of discovering two or more disease biomarkers, optionally, more than 10, 100, 1,000, 10,000, 1,000,000, 1,000,000, or 10,000,000 disease biomarkers. In further embodiments, the method includes the step of discovering 1 to 10, 1 to 100, 1 to 1,000, 1 to 1,000, 1 to 1,000,000, 1 to 1,000,000, or 1 to 10,000,000 disease biomarkers.
[0018] Disease biomarkers may be suitable targets for therapeutic agents, such as vaccines, RNA therapies, and / or gene editing. The discovery of cancer cell-specific disease biomarkers may be an initial step in identifying cancer-specific antigens to be targeted by cancer vaccines. Therefore, in certain embodiments, the method further includes the step of identifying transcripts or proteins corresponding to disease biomarkers as cancer vaccine targets. The method may further include the step of developing cancer vaccines, optionally RNA vaccines, that target these biomarkers.
[0019] According to related embodiments of the present invention, (i) A step of preparing a first cDNA sample from a subject with cancer and a second cDNA sample from a subject without cancer, (ii) A step of normalizing the first and second cDNA samples, (iii) the step of sequencing the normalized first and second cDNA samples, and (iv) A step to discover cancer vaccine targets by comparing the sequencing outputs of the first and second cDNA samples, A method for discovering cancer vaccine targets is provided, including the method described above.
[0020] In certain embodiments, the method includes the steps of preparing 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40 or more cDNA samples from different subjects having a disease, and / or preparing 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40 or more cDNA samples from different subjects not having the disease. Preferably, the method includes the steps of preparing 30 or more cDNA samples from different subjects having a disease (e.g., having cancer), and preparing 30 or more cDNA samples from different subjects not having the disease (e.g., not having cancer).
[0021] In certain embodiments, the first cDNA sample and the second cDNA sample are obtained from a biological fluid, or a fluid or lysate generated from a biological material. The first cDNA sample and the second cDNA sample can be obtained from blood. The first cDNA sample may be obtained by extracting RNA from a biological sample (e.g., blood) obtained from a subject having a disease, and the second cDNA sample may be obtained by extracting RNA from a biological sample (e.g., blood) obtained from a subject not having a disease. Thereafter, cDNA is synthesized using RNA as a template (i.e., by reverse transcription). Thus, in certain embodiments, the method may further comprise, prior to step (i), (a) extracting RNA from a biological fluid, or a fluid or lysate generated from a biological material, from a subject having a disease and a subject not having a disease, and (b) synthesizing cDNA using RNA as a template (i.e., converting RNA to cDNA) may be further included. In this way, the first cDNA sample and the second cDNA sample can be generated.
[0022] A subject having a disease means that the subject has a disease when the biological sample (biological fluid or biological material) from which the cDNA sample is derived is collected from the subject. A subject not having a disease means that the subject does not have a disease when the biological sample (biological fluid or biological material) from which the cDNA sample is derived is collected from the subject.
[0023] A subject not having a disease can be a healthy subject.
[0024] A disease biomarker may be a cDNA molecule (of a specific sequence), RNA molecule (of a specific sequence), protein, or peptide that is detectable in samples from subjects with the disease but not in samples from subjects without the disease. Alternatively, a disease biomarker may be a cDNA molecule (of a specific sequence), RNA molecule (of a specific sequence), protein, or peptide that is not detectable in samples from subjects with the disease but is detectable in samples from subjects without the disease.
[0025] Cancer vaccine targets may be cDNA molecules (of a specific sequence), RNA molecules (of a specific sequence), proteins, or peptides that are detectable in samples from subjects with cancer but not in samples from subjects without cancer.
[0026] In certain embodiments, the disease biomarker is a cDNA sequence present in the first cDNA sample but not in the second cDNA sample, or a cDNA sequence present in the second cDNA sample but not in the first cDNA sample. In certain embodiments, the disease biomarker is a transcript that is present (intrinsically present) in subjects with a particular disease. In further embodiments, the disease biomarker is a transcript that is absent (intrinsically absent) in subjects with a particular disease.
[0027] In certain embodiments, the cancer vaccine target is a cDNA sequence present in the first cDNA sample but not in the second cDNA sample. In certain embodiments, the cancer vaccine target is a transcript uniquely present in subjects with a particular cancer. The transcript / cDNA sequence may correspond to a specific protein or peptide. At least a portion of the protein or peptide may form an antigen included in the cancer vaccine.
[0028] The present invention enables the identification of transcripts found only in subjects with a disease, and optionally, cancer. Such transcript models can be identified through comparison with subjects without the disease (e.g., benign patients), and optionally, with public transcriptome annotation databases. Sequencing outputs and / or obtained transcriptome profiles(s) from subjects with a specific disease (e.g., breast cancer) can be compared with sequencing outputs and / or obtained transcriptome profiles(s) from subjects with different diseases (e.g., ovarian cancer and / or colorectal cancer) to determine whether the biomarker is specific to a particular disease (e.g., breast cancer).
[0029] In a further embodiment, the present invention is (i) A step of preparing a cDNA sample from the subject, (ii) A step of normalizing the cDNA sample, and (iii) A step of sequencing a normalized cDNA sample, wherein the sequencing output is used to determine whether the subject has a disease. The present invention provides a method for diagnosing a disease in a subject, including [specific example].
[0030] To diagnose means to determine whether the subject has a disease at the time of the examination.
[0031] In a particular embodiment, the disease is cancer, and diagnosing the disease involves detecting a small amount of residual disease (cancer cells remaining in the subject during or after treatment).
[0032] In related embodiments, (i) A step of preparing a cDNA sample from the subject, (ii) A step of normalizing the cDNA sample, and (iii) A step of sequencing a normalized cDNA sample, wherein the sequencing output is used to characterize a disease and / or determine the prognosis. Methods are provided for characterizing a disease and / or determining the prognosis in a subject, including the above.
[0033] In further cases, (i) A step of preparing a cDNA sample from the subject, (ii) Step to normalize the cDNA sample, (iii) A step of sequencing a normalized cDNA sample, wherein the sequencing output is used to diagnose, characterize and / or determine the prognosis of a disease, and (iv) A step of selecting appropriate measures for the diagnosis, characterization and / or prognosis of the disease, A method is provided for selecting a treatment for a disease in a subject, including the following.
[0034] In further embodiments, (i) A step of preparing a cDNA sample from the subject, (ii) A step of normalizing the cDNA sample, and (iii) A step of sequencing a normalized cDNA sample, wherein the sequencing output is used to predict the responsiveness of the subject to a therapeutic agent. A method is provided for predicting the response of subjects with a disease to therapeutic agents.
[0035] The methods described herein may further include a step of treating the subject.
[0036] The method may include the step of comparing the sequencing output of a normalized cDNA sample with the sequencing output of one or more reference sequences or one or more control samples, optionally, one or more control samples being samples from one or more subjects with the disease and / or one or more subjects without the disease. Preferably, the method includes the step of comparing the sequencing output of a normalized cDNA sample with the sequencing output of one or more control samples from one or more subjects with the disease.
[0037] Sequencing output refers to one or more sequences obtained from sequencing cDNA (normalized). Sequences may be raw sequences or may be further processed. For example, low-quality reads may be filtered and / or adapter sequences may be filtered and removed. Processed sequences may be mapped to a human reference genome (e.g., using Minimap2) to prepare transcriptome profiles. One or more transcript models may be identified in the transcriptome profiles (sequences mapped to the genome). A transcript model represents a specific transcript, i.e., a specific RNA isoform or splice variant produced from a gene. Thus, in certain embodiments, the sequencing output used in the methods defined herein (e.g., used to compare for the discovery of disease biomarkers or to determine whether a subject has a disease) may be transcriptome profiles and / or transcript models.
[0038] Using sequencing output to determine whether a subject has a disease may include detecting disease biomarkers. In certain embodiments, using sequencing output to determine whether a subject has a disease may include detecting two or more disease biomarkers, optionally, more than 10, 100, 1,000, 10,000, 1,000,000, 1,000,000, 1,000,000, 1,000,000 or 10,000,000 disease biomarkers. In further embodiments, using sequencing output to determine whether a subject has a disease may include detecting 1 to 10, 1 to 100, 1 to 1,000, 1 to 1,000, 1 to 1,000,000, 1 to 1,000,000 or 1 to 10,000,000 disease biomarkers. Detecting disease biomarkers may include determining the presence or absence of disease biomarkers. Therefore, in certain embodiments, using the sequencing output to determine whether a subject has a disease includes determining the presence or absence of two or more disease biomarkers, optionally, more than 10, 100, 1,000, 10,000, 1,000,000, 1,000,000, or 10,000,000 disease biomarkers. In further embodiments, using the sequencing output to determine whether a subject has a disease includes determining the presence or absence of 1 to 10, 1 to 100, 1 to 1,000, 1 to 1,000, 1 to 1,000,000, 1 to 1,000,000, or 1 to 10,000,000 disease biomarkers.
[0039] The presence of a specific cDNA sequence in sequencing output may indicate that a subject has a disease, for example, if a specific transcript (corresponding to a cDNA molecule) is uniquely present in a subject with a particular disease. Similarly, the presence of a specific cDNA sequence in sequencing output may indicate disease characterization and / or prognosis. The presence of a specific cDNA sequence in sequencing output may enable prediction of a subject's responsiveness to a therapeutic agent, for example, if a specific transcript is found to correlate with a subject's responsiveness to a particular therapeutic agent.
[0040] The absence of a specific cDNA sequence in sequencing output may indicate that a subject has a disease, for example, if a specific transcript (corresponding to a cDNA molecule) is absent in a subject with a particular disease. Similarly, the absence of a specific cDNA sequence in sequencing output may indicate disease characterization and / or prognosis. The absence of a specific cDNA sequence in sequencing output may allow for prediction of a subject's responsiveness to a therapeutic agent, for example, if a specific transcript has been found to correlate with a subject's responsiveness to a particular therapeutic agent.
[0041] According to all aspects of the present invention, in certain embodiments, the cDNA sample is obtained from a biological fluid, or a fluid or lysate produced from a biological substance. The cDNA sample may be obtained from blood. The cDNA sample may be obtained by extracting RNA from a biological sample (e.g., blood) obtained from a subject. The cDNA is then synthesized using the RNA as a template (i.e., by reverse transcription). Therefore, in certain embodiments, the method is performed prior to step (i), (a) A step of extracting RNA from a biological fluid from a subject, or from a fluid or lysate produced from a biological substance, and (b) A step of synthesizing cDNA using RNA as a template (i.e., converting RNA to cDNA), This further includes the following. In this way, a cDNA sample can be generated from the subject.
[0042] The method may further include a step of reporting the outcome of the method to the subjects. The outcome may be a diagnosis of a disease or a determination of prognosis. In certain embodiments, the outcome may be a disease, for example, a specific grade or stage of cancer.
[0043] Complementary DNA (cDNA) normalization (Alex S. Shcheglov, Pavel A. Zhulidov, Ekaterina A. Bogdanova, DAS Normalization of cDNA Libraries, Nucleic Acids Hybrid, Chapter 5, (2014)) addresses the problem that high abundances of housekeeping genes reduce the sampling efficiency of target genes. Since RNA sequencing generally relies on converting RNA to double-stranded cDNA, cDNA normalization utilizes the biochemical properties of cDNA to generate a uniform distribution of unique genes and isoforms within a cDNA library. Theoretically, maximum non-target sampling efficiency is achieved when all unique RNA sequences are represented by the same relative abundance. Therefore, the goal of normalization is to redistribute the cDNA library (sample) to be as close as possible to this criterion.
[0044] Normalizing a cDNA sample results in the generation of a normalized cDNA sample. In certain embodiments, normalization involves increasing (selectively increasing) the relative abundance of less abundant sequences without targeting specific sequences based on the nucleotide sequence of the sequence (i.e., identity or homology with known sequences).
[0045] "Normalized" means that the levels of RNA or cDNA sequences in a sample are more uniform. Therefore, a normalized cDNA sample may be one in which the amount of each intrinsic cDNA sequence is more uniform than in the same sample before normalization; that is, a normalized cDNA sample is closer to achieving equal abundance of each intrinsic cDNA sequence (compared to other intrinsic cDNA sequences in the normalized cDNA sample) than in the same sample before normalization. To achieve this, it is possible to increase the relative representation or level of less abundant sequences and / or decrease the relative representation or level of more abundant sequences. The increase of less abundant sequences / decrease of more abundant sequences is selective in the sense that the relative abundance may remain the same if all sequences are increased / decreased to the same extent. However, it is possible to increase the relative representation or level of less abundant sequences and / or decrease the relative representation or level of more abundant sequences without targeting specific sequences based on the nucleotide composition of the sequences (i.e., based on their identity or homology with known sequences) (e.g., using predefined probes). According to all embodiments, less abundant sequences may be unique sequences having amounts below a threshold, for example, they are present in the unnormalized cDNA sample in amounts below the average amount of unique sequences in the sample. Less abundant sequences may be present in the unnormalized cDNA sample in amounts 0.1%, 1%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% lower than the average amount of unique sequences in the sample. More abundant sequences may be unique sequences having amounts above a threshold, for example, they are present in the unnormalized cDNA sample in amounts above the average amount of unique sequences in the sample. More abundant sequences may be present in the unnormalized cDNA sample in amounts 0.1%, 1%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% higher than the average amount of unique sequences in the sample.Relative abundance refers to the abundance of a sequence relative to other intrinsic sequences in the sample.
[0046] In certain embodiments, the normalized cDNA sample contains cDNA sequences having substantially the same level. For example, here the level of sequences in the normalized cDNA varies by less than 50%, less than 40%, less than 30%, less than 20%, or less than 10%. The normalized cDNA may be a normalized cDNA sample in which at least a fraction of the copies of 10, 100, 1,000, or 10,000 of the most abundant sequences (unique sequences) in the cDNA sample have been removed or reduced (reduced by at least 1%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%). A normalized cDNA can be a normalized cDNA sample in which the levels of at least some of the 10, 100, 1,000, or 10,000 least abundant sequences (unique sequences) in the cDNA sample have been increased in copy number by, for example, at least 1%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%. The methods for normalizing cDNA samples(s) described herein may be methods for equating cDNA samples(s), i.e., for equating the relative abundance of each unique sequence.
[0047] According to all aspects of the present invention, in certain embodiments, normalizing a cDNA sample reduces variability at the level of cDNA (e.g., by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%). Normalizing cDNA can achieve a more uniform distribution of cDNA sequences. The difference in abundance between the most abundant and least abundant cDNA in a sample can be reduced (e.g., by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%). In certain embodiments, normalizing a cDNA sample reduces the number of molecules (copy numbers) of the most abundant cDNA molecules (1, 10, 100, 1,000, or 10,000 of the most abundant cDNA molecules) by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%. In certain embodiments, the number of molecules (copy numbers) of the most abundant cDNA molecules in a cDNA sample (first and / or second cDNA sample) is reduced by at least 50% in the normalized cDNA. In further embodiments, the relative abundance of the least abundant cDNA molecule (1, 10, 100, 1,000, or 10,000 least abundant cDNA molecules) is increased by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%. In certain embodiments, the number of molecules (copy number) of the least abundant cDNA molecule in the cDNA sample (first and / or second cDNA sample) is increased by at least 50% in the normalized cDNA.
[0048] According to all aspects of the present invention, in certain embodiments, normalizing a cDNA sample (first and / or second cDNA sample) (or more) increases the amount (copy number) of at least some of the low-abundance cDNA sequences in the cDNA sample (first and / or second cDNA sample) (or more). Low-abundance cDNA sequences may be 50%, 40%, 30%, 20%, 10%, or 1% of the sequence with the fewest copies (unique sequence). Therefore, normalization may involve selectively increasing the amount of low-abundance cDNA in each cDNA sample.
[0049] In certain embodiments, normalized cDNA is cDNA that is easier to analyze. Because the relative representation of less abundant sequences is increased, it can be sequenced more efficiently. In certain embodiments, normalizing a cDNA sample (first and / or second cDNA sample)(or more) does not involve removing abundant (more abundant) cDNA molecules / sequences (e.g., cDNA molecules / sequences corresponding to albumin, IgG, apolipoprotein AI, transferrin, apolipoprotein A-II, α1-proteinase inhibitors, α1-glycoprotein, trans tyretin, hepatoglobin, and / or hemopexin) from the sample(or more) (e.g., using double-strand specific nucleases or sequence targeting methods). In further embodiments, normalization does not involve targeting specific sequences (unique sequences) (e.g., sequences corresponding to albumin, IgG, apolipoprotein AI, transferrin, apolipoprotein A-II, α1-proteinase inhibitors, α1-glycoprotein, trans tyretin, hepatoglobin, and / or hemopexin). Therefore, in certain embodiments, normalizing a cDNA sample is non-targeting, i.e., normalizing a cDNA sample does not involve targeting specific sequences based on the nucleotide sequence of the sequence (e.g., normalizing a cDNA sample does not involve targeting specific sequences based on sequence identity or homology with known sequences).
[0050] According to all aspects of the present invention, the term “sequence” may refer to all individual nucleic acid (e.g., cDNA or RNA) molecules having 100% identical nucleotide sequences. Alternatively, the term “sequence” may refer to all individual nucleic acid (e.g., cDNA or RNA) molecules having more than 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% identity with respect to one another. The "identity percentage" between the query nucleic acid sequence and the target nucleic acid sequence can be calculated over the entire length of the query sequence using a suitable algorithm (e.g., BLASTN, FASTA, Needleman-Wunsch, Smith-Waterman, LALIGN, or GenePAST / KERR) or software (e.g., DNASTAR Lasergene, GenomeQuest, EMBOSS needle, or EMBOSS infoalign), after the overall sequence alignment for each pair has been performed using a suitable algorithm (e.g., Needleman-Wunsch or GenePAST / KERR) or software (e.g., DNASTAR Lasergene, GenomeQuest, EMBOSS needle, or EMBOSS infoalign). The terms “unique sequence” or “unique cDNA sequence” may refer to all individual nucleic acid (e.g., cDNA or RNA, as appropriate) molecules that meet or exceed an identity % threshold (e.g., 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% identity to one another). A “unique sequence” or “unique cDNA sequence” may differ from other sequences present in the sample (e.g., by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or 100 nucleotides, or by at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of those sequences).
[0051] According to all aspects of the present invention, in certain embodiments, the disease is not an infectious disease. In certain embodiments, the disease is cancer. The cancer may be an epithelial carcinoma. In certain embodiments, the cancer is breast cancer, ovarian cancer, and / or colorectal cancer. Preferably, the cancer is breast cancer. In further embodiments, the first cDNA sample is a cDNA sample from a subject having breast cancer, and the second cDNA sample is a cDNA sample from a subject having a benign breast condition.
[0052] In certain embodiments, the method may be used to diagnose two or more diseases in a single process, for example, by detecting multiple cDNA molecules (derived from transcripts) that are uniquely present in each subject having a particular disease.
[0053] In certain embodiments, the method may be used to diagnose two or more types of cancer. In certain embodiments, the method may be used to distinguish between breast cancer and a benign breast condition.
[0054] The inventors have developed a selective amplification method (called "level-up") that can be used for cDNA normalization. Level-up normalization utilizes a cDNA library to equalize the relative abundance of each unique transcript sequence. Level-up normalization is performed without depletion or targeting. Essentially, level-up normalization enables the creation of an optimal cDNA library for detecting all RNA present in a sample.
[0055] Therefore, normalizing a cDNA sample is a method for selective amplification of single-stranded cDNA. (i) A step of preparing a cDNA sample containing a double-stranded cDNA template, each template having a known 5' pre-binding adapter and a known 3' pre-binding adapter, (ii) A step of denaturing the cDNA sample and preparing a single-stranded cDNA template, (iii) A step of reassociating the cDNA samples to prepare a mixture of the post-association single-stranded cDNA template and the post-association double-stranded cDNA template, (iv) Annealing a 5' adapter complex to the 5' pre-binding adapter of at least one post-association single-stranded cDNA template, and annealing a 3' adapter complex to the 3' pre-binding adapter of the same post-association single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide, (v) Ligating oligonucleotides from the 5' adapter complex to the 5' pre-binding adapter of the post-association single-stranded cDNA template, and ligating oligonucleotides from the 3' adapter complex to the 3' pre-binding adapter of the same post-association single-stranded cDNA template, (vi) A step of selectively amplifying the cDNA sample using primers specific to the ligated oligonucleotide, This can be achieved by using methods that include [specific methods].
[0056] Therefore, in certain embodiments, the first and second cDNA samples include a double-stranded cDNA template, each template having a known 5' pre-binding adapter and a known 3' pre-binding adapter. Normalizing the first and second cDNA samples is (i) Denaturing a cDNA sample to create a single-stranded cDNA template. (ii) Reassociate the cDNA samples to prepare a mixture of the resulting single-stranded cDNA template and the resulting double-stranded cDNA template. (iii) Annealing a 5' adapter complex to a 5' pre-binding adapter of at least one post-association single-stranded cDNA template, and annealing a 3' adapter complex to a 3' pre-binding adapter of the same post-association single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide. (iv) Ligating oligonucleotides from the 5' adapter complex to the 5' pre-binding adapter of the post-association single-stranded cDNA template, and ligating oligonucleotides from the 3' adapter complex to the 3' pre-binding adapter of the same post-association single-stranded cDNA template, (v) Selective amplification of the cDNA sample using primers specific to the ligated oligonucleotide. Includes.
[0057] In further embodiments, the cDNA sample comprises a double-stranded cDNA template, each template having a known 5' pre-binding adapter and a known 3' pre-binding adapter. Normalizing cDNA samples is (i) Denaturing a cDNA sample to create a single-stranded cDNA template. (ii) Reassociate the cDNA samples to prepare a mixture of the resulting single-stranded cDNA template and the resulting double-stranded cDNA template. (iii) Annealing a 5' adapter complex to a 5' pre-binding adapter of at least one post-association single-stranded cDNA template, and annealing a 3' adapter complex to a 3' pre-binding adapter of the same post-association single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide. (iv) Ligating oligonucleotides from the 5' adapter complex to the 5' pre-binding adapter of the post-association single-stranded cDNA template, and ligating oligonucleotides from the 3' adapter complex to the 3' pre-binding adapter of the same post-association single-stranded cDNA template, (v) Selective amplification of the cDNA sample using primers specific to the ligated oligonucleotide. Includes.
[0058] Furthermore, the present invention is (i) A step of preparing a first cDNA sample from a subject with the disease and a second cDNA sample from a subject without the disease, wherein the first and second cDNA samples include a double-stranded cDNA template, and each template has a known 5' pre-binding adapter and a known 3' pre-binding adapter. (ii) A step of denaturing the first and second cDNA samples to prepare a single-stranded cDNA template, (iii) Reassociating the first and second cDNA samples to prepare a mixture of the associated single-stranded cDNA template and the associated double-stranded cDNA template, (iv) In each of the first and second cDNA samples, the 5' adapter complex is annealed to the 5' pre-binding adapter of at least one post-associated single-stranded cDNA template, and the 3' adapter complex is annealed to the 3' pre-binding adapter of the same post-associated single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide. (v) Ligating oligonucleotides from the 5' adapter complex to the 5' pre-binding adapter of the post-association single-stranded cDNA template, and ligating oligonucleotides from the 3' adapter complex to the 3' pre-binding adapter of the same post-association single-stranded cDNA template, (vi) A step of selectively amplifying the first and second cDNA samples using primers specific to the ligated oligonucleotides. (vii) The step of sequencing the selectively amplified first and second cDNA samples, and (viii) A step to discover disease biomarkers by comparing the sequencing outputs of the first and second cDNA samples, This provides a method for discovering disease biomarkers, including those mentioned above.
[0059] The first and second cDNA samples are isolated from each other and / or can be identified separately. The present invention is a method for diagnosing a disease in a subject, (i) A step of preparing a cDNA sample from a subject, which includes a double-stranded cDNA template in which each template has a known 5' pre-binding adapter and a known 3' pre-binding adapter, (ii) A step of denaturing the cDNA sample and preparing a single-stranded cDNA template, (iii) A step of reassociating the cDNA samples to prepare a mixture of the post-association single-stranded cDNA template and the post-association double-stranded cDNA template, (iv) Annealing a 5' adapter complex to the 5' pre-binding adapter of at least one post-association single-stranded cDNA template, and annealing a 3' adapter complex to the 3' pre-binding adapter of the same post-association single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide, (v) Ligating oligonucleotides from the 5' adapter complex to the 5' pre-binding adapter of the post-association single-stranded cDNA template, and ligating oligonucleotides from the 3' adapter complex to the 3' pre-binding adapter of the same post-association single-stranded cDNA template, (vi) a step of selectively amplifying a cDNA sample using a primer specific to the ligated oligonucleotide, and (vii) A step of sequencing a selectively amplified cDNA sample, wherein the sequencing output is used to determine whether the subject has the disease. Further methods including this will be provided.
[0060] In a particular embodiment, (A) The 5' adapter complex is (i) A forward rig-oligonucleotide for ligating to the 5' pre-binding adapter of the (post-association) single-stranded cDNA template, and (ii) A forward link oligonucleotide for annealing to a 5' pre-bonding adapter and a forward rig oligonucleotide, comprising a region complementary to the 5' pre-bonding adapter and a region complementary to the forward rig oligonucleotide, Includes, During annealing, the terminal end of the forward lig-oligonucleotide is adjacent to the terminal end of the 5' pre-binding adapter, enabling ligation of the forward lig-oligonucleotide to the 5' pre-binding adapter at the ligation site; and, (B) The 3' adapter complex is (i) Post-association single-stranded cDNA template ligation to the 3' pre-binding adapter, and (ii) A backlink oligonucleotide for annealing to a 3' pre-binding adapter and a backlink oligonucleotide, comprising a region complementary to the 3' pre-binding adapter and a region complementary to the backlink oligonucleotide, Includes, This is a posterior oligonucleotide dimer in which, during annealing, the terminus of the posterior lig-oligonucleotide is adjacent to the terminus of the 3' pre-binding adapter, enabling ligation of the posterior lig-oligonucleotide to the 3' pre-binding adapter at the ligation site. Preferably, (A) The forward link oligonucleotide is (i) A template overhang region at the end of a forward link oligonucleotide adjacent to a region complementary to the 5' pre-binding adapter, which is incomparable to the corresponding region of the (post-association) single-stranded cDNA template, and / or (ii) A rig-oligonucleotide overhang region at the terminal of a forward link-oligonucleotide adjacent to a region complementary to the forward rig-oligonucleotide, which is a rig-oligonucleotide overhang region that is non-complementary to the corresponding region of the forward rig-oligonucleotide. including and / or (B) The back link oligonucleotide is (i) A template overhang region at the end of a backlink oligonucleotide adjacent to a region complementary to the 3' pre-binding adapter, which is incomparable to the corresponding region of the (post-association) single-stranded cDNA template, and / or (ii) A rig-oligonucleotide overhang region at the terminal of a posterior link-oligonucleotide adjacent to a region complementary to the posterior rig-oligonucleotide, which is a rig-oligonucleotide overhang region that is non-complementary to the corresponding region of the posterior rig-oligonucleotide. Includes.
[0061] Preferably, the template overhang and / or rig-oligonucleotide overhang is between approximately 1 bp and approximately 20 bp in length. The template overhang and / or rig-oligonucleotide overhang may be between 2 bp and 19 bp, between 3 bp and 18 bp, between 2 bp and 17 bp, between 3 bp and 16 bp, between 2 bp and 15 bp, between 3 bp and 14 bp, between 2 bp and 13 bp, between 3 bp and 12 bp, between 2 bp and 11 bp, between 3 bp and 10 bp, between 2 bp and 9 bp, between 3 bp and 8 bp, between 2 bp and 7 bp, between 3 bp and 6 bp, between 2 bp and 5 bp, between 3 bp and 5 bp, or between 2 bp and 4 bp. Preferably, the template overhang and / or rig-oligonucleotide overhang is 3 bp.
[0062] The template overhang and / or rig-oligonucleotide overhang may be at least 2 bp, or at least 3 bp. Preferably, the template overhang and / or rig-oligonucleotide overhang is at least 3 bp.
[0063] Preferably, the combined length of the forward link-oligonucleotide and the forward rig-oligonucleotide is less than approximately 300 bp, and / or the combined length of the backward link-oligonucleotide and the backward rig-oligonucleotide is less than approximately 300 bp. In some cases, the combined length of the forward link-oligonucleotide and the forward rig-oligonucleotide may be at least approximately 200 bp, and / or the combined length of the backward link-oligonucleotide and the backward rig-oligonucleotide may be at least approximately 100 bp.
[0064] Preferably, the forward and / or backward link oligonucleotides have a length of less than 200 bp. The forward and / or backward link oligonucleotides may have a length of at least 50 bp.
[0065] Preferably, the forward oligonucleotide dimer and / or the backward oligonucleotide dimer have at least one blunt end.
[0066] Preferably, the forward link oligonucleotide and / or backward link oligonucleotide provide a complementary bond of at least 5 bp on either side of the ligation site.
[0067] Preferably, the nucleotide sequence of the forward oligonucleotide dimer is different from and non-complementary to the nucleotide sequence of the backward oligonucleotide dimer.
[0068] Preferably, at least one of the forward oligonucleotide dimer and the backward oligonucleotide dimer can be annealed to a (post-association) single-stranded cDNA template at a temperature above 30°C.
[0069] Preferably, the concentrations of the forward oligonucleotide dimer and / or the backward oligonucleotide dimer exceed the predicted total single-stranded cDNA concentration or the total cDNA concentration in the cDNA sample.
[0070] Preferably, the step of reassociating the cDNA sample has a duration of 0 to 24 hours, optionally 0 to 8 hours, 1 to 7 hours, 1 to 24 hours, or 7 to 24 hours.
[0071] Long-read sequencing allows for the detection of full-length RNA / cDNA, providing better information for identifying the tissue origin and specific function of each RNA. While long-read RNA sequencing is more expensive than other assays, level-up techniques can reduce the amount of sequencing required, thus lowering the overall cost.
[0072] Accordingly, according to all aspects of the present invention, in certain embodiments, the sequencing includes the use of long-read sequencing. Thus, the sequencing may be sequencing by long-read sequencing, for example, enabling sequencing reads exceeding 1000, 5000, or 10000 bp.
[0073] Leveling up allows for the analysis of samples involving RNA degradation by increasing the presence of low-abundance transcripts. However, it is preferable if RNA degradation is minimized during processing of the biological sample to generate cDNA samples. The inventors have developed a method for processing blood samples when RNA extraction is not performed on the day of blood collection. Specific steps of freezing, storing, and thawing the blood samples improve the condition of the samples, especially if long-read sequencing is planned for them.
[0074] Therefore, the present invention is (i) A step of storing the blood sample at -15°C or below, (ii) The step of thawing the blood sample at 5-30°C for at least 1 hour, (iii) A step of extracting RNA from the thawed blood sample, The present invention provides a method for processing blood samples, including [specific components / methods].
[0075] The blood sample may be a liquid (i.e., non-dried) blood sample. The blood sample may be whole blood. In certain embodiments, the blood sample is contained within a sample tube. In certain embodiments, the blood sample is not absorbed by a material such as a sponge.
[0076] In certain embodiments, blood samples are stored at temperatures between -15°C and -80°C, -15°C and -70°C, -15°C and -60°C, -15°C and -50°C, -15°C and -40°C, -15°C and -30°C, or -15°C and -20°C.
[0077] Blood samples may be stored at -15°C or below within 12 hours, 8 hours, 5 hours, 2 hours, 1 hour, 30 minutes, 15 minutes, 5 minutes, or 1 minute of collection. Preferably, blood samples are stored at -15°C or below within 12 hours of collection (i.e., taking a blood sample from the subject).
[0078] In certain embodiments, blood samples are stored at -20°C or below. Blood samples may be stored at -20°C or below within 12 hours, 8 hours, 5 hours, 2 hours, 1 hour, 30 minutes, 15 minutes, 5 minutes, or 1 minute of collection. Preferably, blood samples are stored at -20°C or below within 12 hours of collection (i.e., taking a blood sample from the subject).
[0079] In a particular embodiment, the blood sample is stored at -15°C or below (optionally, -20°C or below) for at least 24 hours (and optionally, for 72 hours, 1 week, 2 weeks, 4 weeks, 1 month or less, or up to 2 months), and then stored at -70°C or below (optionally, -80°C or below) for 4, 5, 6, 7, 8, 9, 10, 11 or 12 months, or for 2, 3, 4 or 5 years or less. Preferably, storage at -70°C or below (optionally, -80°C or below) is for 5 years or less.
[0080] Preferably, the thawing of the blood sample is performed on the same day that the RNA is to be extracted. In certain embodiments, the blood sample is thawed at 16-29°C, 17-28°C, 18-27°C, 18-26°C, or 18-25°C. Preferably, the blood sample is thawed at 18-25°C. The duration of the thawing step may be 1-5 hours, 2-5 hours, 1-4 hours, 2-4 hours, 1-3 hours, 2-3 hours, or 1-2 hours. In preferred embodiments, the blood sample is thawed at 18-25°C for 1-3 hours. More preferably, the blood sample is thawed at 18-25°C for 3 hours.
[0081] In certain embodiments, once the blood sample is completely thawed, the sample tube is inverted at least 5, 6, 7, 8, 9, or 10 times. Preferably, the sample tube is inverted 10 times. The blood sample can then be incubated at 18–25°C for approximately 2 hours before RNA extraction.
[0082] RNA can be extracted from (thawed) blood samples using the Qiagen Paxgene Blood RNA Kit.
[0083] In certain embodiments, prior to step (i), blood is received in a container (optionally, a Paxgene Blood RNA Tube) at room temperature (5–30°C, preferably 18–25°C). The container may be inverted at least 5, 6, 7, 8, 9, or 10 times immediately after blood collection. Preferably, the container is inverted 10 times immediately after blood collection. The blood sample may be stored at -15°C or below (or -20°C or below) immediately after inversion.
[0084] Therefore, the present invention is (i) A step of receiving a blood sample into a container at 18-25°C, (ii) A step of storing the blood sample at -20°C or below within 12 hours of collection, (iii) A step of thawing the blood sample at 18-25°C for 1-3 hours, and (iv) Step of extracting RNA from the thawed blood sample, The present invention provides a method for processing blood samples, including [specific components / methods].
[0085] According to all aspects of the present invention, in certain embodiments, the blood sample is 5 ml, 3 ml, 2.5 ml, or 2 ml or less. Preferably, the blood sample is 2.5 ml or less. The blood sample may be 0.1 ml to 5 ml, 0.5 ml to 5 ml, 1 ml to 4 ml, or 2 ml to 3 ml.
[0086] The method for processing a blood sample may be combined with the method outlined above using a cDNA sample. Thus, in certain embodiments, a cDNA sample is obtained from blood by following the steps of the method for processing a blood sample described herein. In addition, the method may include the step of synthesizing cDNA using the extracted RNA as a template (i.e., reverse transcription, converting the extracted RNA into cDNA). In further embodiments, a first cDNA sample and a second cDNA sample are obtained from blood by following the steps of the method for processing a blood sample described herein. In addition, the method may include the step of synthesizing cDNA using the extracted RNA as a template (i.e., reverse transcription, converting the extracted RNA into cDNA).
[0087] Therefore, the present invention is (i) A step of receiving a first blood sample from a subject with a disease in a first container and a second blood sample from a subject without a disease in a second container at 18-25°C, (ii) The step of storing the first and second blood samples at -20°C or below within 12 hours of collection from the subject, (iii) A step of thawing the first and second blood samples at 18-25°C for 1-3 hours. (iv) A step of extracting RNA from the thawed first and second blood samples, (v) Synthesizing cDNA using the extracted RNA as a template to form a first cDNA sample from a subject with the disease and a second cDNA sample from a subject without the disease. (vi) Step of normalizing the first and second cDNA samples, (vii) The step of sequencing the normalized first and second cDNA samples, and (viii) A step to discover disease biomarkers by comparing the sequencing outputs of the first and second cDNA samples, This provides a method for discovering disease biomarkers, including those mentioned above. In addition, the present invention relates to a method for diagnosing a disease in a subject, (i) A step of receiving a blood sample from the subject into a container at 18-25°C, (ii) A step of storing the blood sample at -20°C or below within 12 hours of collection from the subject, (iii) A step of thawing the blood sample at 18-25°C for 1-3 hours. (iv) Step of extracting RNA from the thawed blood sample, (v) A step of synthesizing cDNA using the extracted RNA as a template to form a cDNA sample from the subject, (vi) A step to normalize the cDNA sample, and (vii) A step of sequencing a normalized cDNA sample, the sequencing output being used to determine whether the subject has the disease, This provides a method that includes [something].
[0088] A further aspect of the present invention provides the use of oligonucleotide dimer compositions for selective amplification of single-stranded cDNA by ligation of oligonucleotides to the 5' and 3' ends of a post-association single-stranded cDNA template having known 5' and 3' pre-binding adapters in a method for discovering cancer biomarkers and / or cancer vaccine targets, the composition is (A)(i) A forward rig-oligonucleotide for ligating to the 5' pre-binding adapter of the single-stranded cDNA template after association, and (ii) A forward link oligonucleotide for annealing to a 5' pre-bonding adapter and a forward rig oligonucleotide, comprising a region complementary to the 5' pre-bonding adapter and a region complementary to the forward rig oligonucleotide. Includes, A forward oligonucleotide dimer, in which, during annealing, the terminus of the forward lig-oligonucleotide is adjacent to the terminus of the 5' pre-binding adapter, enabling ligation of the forward lig-oligonucleotide to the 5' pre-binding adapter at the ligation site, and, (B)(i) A posterior rig-oligonucleotide for ligating the 3' pre-binding adapter of the single-stranded cDNA template after association, and (ii) A backlink oligonucleotide for annealing to a 3' pre-binding adapter and a backlink oligonucleotide, comprising a region complementary to the 3' pre-binding adapter and a region complementary to the backlink oligonucleotide. Includes, A back oligonucleotide dimer in which, during annealing, the end of the back lig-oligonucleotide is adjacent to the end of the 3' pre-binding adapter, enabling ligation of the back lig-oligonucleotide to the 3' pre-binding adapter at the ligation site. Includes.
[0089] Furthermore, the present invention provides for the use of oligonucleotide dimer compositions for selective amplification of single-stranded cDNA by ligation of oligonucleotides to the 5' and 3' ends of a post-association single-stranded cDNA template having known 5' and 3' pre-binding adapters in a method for diagnosing and / or determining the prognosis of cancer in a subject, the composition is (A)(i) A forward rig-oligonucleotide for ligating to the 5' pre-binding adapter of the single-stranded cDNA template after association, and (ii) A forward link oligonucleotide for annealing to a 5' pre-bonding adapter and a forward rig oligonucleotide, comprising a region complementary to the 5' pre-bonding adapter and a region complementary to the forward rig oligonucleotide. Includes, A forward oligonucleotide dimer, in which, during annealing, the terminus of the forward lig-oligonucleotide is adjacent to the terminus of the 5' pre-binding adapter, enabling ligation of the forward lig-oligonucleotide to the 5' pre-binding adapter at the ligation site, and, (B)(i) A posterior rig-oligonucleotide for ligating the 3' pre-binding adapter of the single-stranded cDNA template after association, and (ii) A backlink oligonucleotide for annealing to a 3' pre-binding adapter and a backlink oligonucleotide, comprising a region complementary to the 3' pre-binding adapter and a region complementary to the backlink oligonucleotide. Includes, A back oligonucleotide dimer in which, during annealing, the end of the back lig-oligonucleotide is adjacent to the end of the 3' pre-binding adapter, enabling ligation of the back lig-oligonucleotide to the 3' pre-binding adapter at the ligation site. Includes.
[0090] Furthermore, oligonucleotide dimer compositions as defined herein may be used in methods for characterizing a disease in a subject, methods for selecting a treatment for a disease in a subject, and / or methods for predicting the responsiveness of a subject having a disease to a therapeutic agent.
[0091] In a particular embodiment, (A) The forward link oligonucleotide is (i) A template overhang region at the end of a forward link oligonucleotide adjacent to a region complementary to the 5' pre-binding adapter, which is a template overhang region that is complementary to the corresponding region of the post-associated single-stranded cDNA template, and / or (ii) A rig-oligonucleotide overhang region at the terminal of a forward link-oligonucleotide adjacent to a region complementary to the forward rig-oligonucleotide, which is a rig-oligonucleotide overhang region that is non-complementary to the corresponding region of the forward rig-oligonucleotide. including and / or (B) The back link oligonucleotide is (i) A template overhang region at the end of a posterior link oligonucleotide adjacent to a region complementary to the 3' pre-binding adapter, which is a template overhang region that is complementary to the corresponding region of the post-associated single-stranded cDNA template, and / or (ii) A rig-oligonucleotide overhang region at the terminal of a posterior link-oligonucleotide adjacent to a region complementary to the posterior rig-oligonucleotide, which is a rig-oligonucleotide overhang region that is non-complementary to the corresponding region of the posterior rig-oligonucleotide. Includes.
[0092] Preferably, the template overhang and / or rig-oligonucleotide overhang is between approximately 1 bp and approximately 20 bp in length. The template overhang and / or rig-oligonucleotide overhang may be between 2 bp and 19 bp, between 3 bp and 18 bp, between 2 bp and 17 bp, between 3 bp and 16 bp, between 2 bp and 15 bp, between 3 bp and 14 bp, between 2 bp and 13 bp, between 3 bp and 12 bp, between 2 bp and 11 bp, between 3 bp and 10 bp, between 2 bp and 9 bp, between 3 bp and 8 bp, between 2 bp and 7 bp, between 3 bp and 6 bp, between 2 bp and 5 bp, between 3 bp and 5 bp, or between 2 bp and 4 bp. Preferably, the template overhang and / or rig-oligonucleotide overhang is 3 bp.
[0093] The template overhang and / or rig-oligonucleotide overhang may be at least 2 bp, or at least 3 bp. Preferably, the template overhang and / or rig-oligonucleotide overhang is at least 3 bp.
[0094] Preferably, the total length of the forward link oligonucleotide and the forward rig oligonucleotide is less than approximately 300 bp, and / or the total length of the backward link oligonucleotide and the backward rig oligonucleotide is less than approximately 300 bp.
[0095] Preferably, the forward and / or backward link oligonucleotides have a length of less than 200 bp.
[0096] Preferably, the forward oligonucleotide dimer and / or the backward oligonucleotide dimer have at least one blunt end.
[0097] Preferably, in the use of the composition, the forward-link oligonucleotide and / or backward-link oligonucleotide provide a complementary bond of at least 5 bp on either side of the ligation site.
[0098] Preferably, the nucleotide sequence of the forward oligonucleotide dimer is different from and non-complementary to the nucleotide sequence of the backward oligonucleotide dimer.
[0099] Preferably, the forward oligonucleotide dimer and / or backward oligonucleotide dimer can be annealed to the post-association single-stranded cDNA template at temperatures above 30°C.
[0100] A selective amplification kit is provided for selectively amplifying low-abundance cDNA from a cDNA sample and / or for selective amplification of cDNA containing a known adapter sequence, wherein the cDNA sample comprises a cDNA template having known 5' and 3' pre-binding adapters, and the kit comprises means for preparing the oligonucleotide dimer composition described above and means for carrying out the selective amplification method described above.
[0101] Further aspects of the present invention provide the use of the kit described herein in a method for diagnosing and / or determining the prognosis of a disease in a subject. Also provided are the use of the kit described herein in a method for characterizing a disease in a subject, a method for selecting a treatment for a disease in a subject, and / or a method for predicting the responsiveness of a subject with the disease to a therapeutic agent.
[0102] Furthermore, the selective amplification kit is provided for selectively amplifying low-abundance cDNA from a first cDNA sample from a diseased subject and a second cDNA sample from a non-disease subject, and / or for selective amplification of cDNA containing known adapter sequences, wherein the first and second cDNA samples comprise a cDNA template having known 5' and 3' pre-binding adapters, and the kit comprises means for preparing the above-mentioned oligonucleotide dimer composition and means for carrying out the above-mentioned selective amplification method.
[0103] A further aspect of the present invention provides the use of the kit described herein in a method for discovering disease biomarkers and / or cancer vaccine targets.
[0104] In certain embodiments, means for preparing oligonucleotide dimer compositions may include forward rig-oligonucleotides, forward link-oligonucleotides, backward rig-oligonucleotides and / or backward link-oligonucleotides as described herein. In further embodiments, means for preparing oligonucleotide dimer compositions may include forward oligonucleotide dimers and / or backward oligonucleotide dimers as described herein.
[0105] Means for carrying out a selective amplification method may include primers specific to the forward and / or backward rig-oligonucleotides.
[0106] In certain embodiments, the kit may further include a hybridization buffer. The hybridization buffer may include HEPES 1M (pH=7.5), NaCl 5M, and H2O. The kit may also include a ligase and / or a ligase buffer. Any suitable ligase may be used. The ligase may be a nick repair ligase or a blunt-end ligase. Optionally, the ligase may be a Taq DNA ligase. Suitable ligase buffers are also well known and commercially available.
[0107] In further embodiments, the kit may further include primers for adding phosphate groups to cDNA before use as a cDNA sample. These primers are based on known 5' pre-binding adapter and known 3' pre-binding adapter sequences.
[0108] In certain embodiments, the kit may further include suitable reagents for PCR, comprising one or more, up to all, of polymerase, dinucleotide triphosphate (dNTP), MgCl2, and buffer. Any suitable polymerase may be used. Generally, DNA polymerases are used to amplify the nucleic acid targets described in the present invention. Examples include heat-stable polymerases such as Taq or Pfu polymerase, and various derivatives of these enzymes. Suitable buffers are also well known and commercially available, and may be included in PCR master mixes that contain most of the components necessary for PCR amplification.
[0109] In further embodiments, the kit further comprises suitable reagents for reverse transcription of RNA to cDNA, including a reverse transcriptase. Any suitable reverse transcriptase may be used. Suitable buffers are also well known and commercially available, and may be included in a reverse transcription master mix that contains most of the components necessary for reverse transcription.
[0110] In certain embodiments, the kit further comprises suitable reagents for processing blood samples, such as a container (optionally, a PAXgene Blood RNA tube), an RNA stabilizing reagent, and / or a hematopoietic cell lysis buffer. The reagents may be RNase-free. RNA stabilizing reagents are commercially available and include RNAlater® (Sigma-Aldrich) and RNAprotect (Qiagen). Suitable RNA stabilizing reagents may include EDTA, sodium citrate, and / or ammonium sulfate, for example, 70% (w / v) ammonium sulfate, 25 mM sodium citrate, and / or 10 mM EDTA. The pH may be adjusted to 5.2 with sulfuric acid.
[0111] Further aspects of the present invention provide the use of a set of reagents in the method described herein, the set of reagents is Forward rig-oligonucleotides, forward link-oligonucleotides, backward rig-oligonucleotides and / or backward link-oligonucleotides, and Primers specific to forward and / or backward rig-oligonucleotides Includes.
[0112] Forward rig-oligonucleotides, forward link-oligonucleotides, backward rig-oligonucleotides and / or backward link-oligonucleotides may be provided as oligonucleotide dimer compositions described herein.
[0113] The reagent set may further include one or more, up to all, of the following: hybridization buffer (optionally containing HEPES 1M (pH=7.5), NaCl 5M, and H2O), ligase, ligase buffer, a primer pair for adding phosphate groups to cDNA, DNA polymerase, and / or dNTPs.
[0114] In further embodiments, the reagent set further comprises suitable reagents for processing blood samples, e.g., a container (optionally, a PAXgene Blood RNA tube), an RNA stabilizing reagent, and / or a blood cell lysis buffer. The reagents may be RNase-free. RNA stabilizing reagents are commercially available and include RNA Rater® (Sigma-Aldrich) and RNAprotect (Qiagen). Suitable RNA stabilizing reagents may include EDTA, sodium citrate, and / or ammonium sulfate, e.g., 70% (w / v) ammonium sulfate, 25 mM sodium citrate, and / or 10 mM EDTA. The pH may be adjusted to 5.2 with sulfuric acid. Incubation for cell lysis may be, for example, 1 minute to 3 hours.
[0115] In one embodiment, the reagent set is: RNA stabilization reagent, Cell (blood cell) lysis buffer, Forward rig-oligonucleotide, forward link-oligonucleotide, backward rig-oligonucleotide and / or backward link-oligonucleotide, Primers specific to forward and / or backward rig-oligonucleotides, Hybridization buffer, Ligauze, Ligase buffer, and Primer pair for adding a phosphate group to cDNA Includes.
[0116] Another important aspect of the oligonucleotide dimer compositions for selective amplification of single-stranded cDNA described herein is the ability to select cDNA having a known adapter sequence. This aspect can be applied to single-cell sequencing where an adapter is required to assign reads to individual cells. In this application, cDNA sequences without a barcode / adapter that identifies cells will occur in the cDNA library. These are known as template-switching oligo (TSO) artifacts and are undesirable in single-cell sequencing projects because they cannot be assigned to the cells from which they originate. This method can be applied solely to select cDNA sequences with a desired adapter sequence and therefore can effectively limit the sequencing of TSO artifacts. This method can also be performed on single-cell cDNA libraries to remove TSO artifacts and to improve the transcriptome coverage per cell.
[0117] TSO cleanup can also be performed without normalization, in which case it is not necessary to include the step of reassociating the cDNA samples to create a mixture of the post-association single-stranded cDNA template and the post-association double-stranded cDNA template.
[0118] Therefore, according to a further aspect of the present invention, a method for diagnosing a disease in a subject, (i) A step of preparing a cDNA sample from a subject containing a double-stranded cDNA template, wherein a portion of the template has a known 5' pre-binding adapter and a known 3' pre-binding adapter. (ii) A step of denaturing the cDNA sample and preparing a single-stranded cDNA template, (iii) Annealing a 5' adapter complex to a 5' pre-binding adapter of at least one single-stranded cDNA template, and annealing a 3' adapter complex to a 3' pre-binding adapter of the same single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide, (iv) Ligating oligonucleotides from the 5' adapter complex to the 5' pre-binding adapter of a single-stranded cDNA template, and ligating oligonucleotides from the 3' adapter complex to the 3' pre-binding adapter of the same single-stranded cDNA template, (v) A step of selectively amplifying a cDNA sample using a primer specific to the ligated oligonucleotide, and (vi) A step of sequencing a cDNA sample, the sequencing output being used to determine whether the subject has a disease, A method including this is provided.
[0119] A further aspect of the present invention is a method for discovering disease biomarkers, (i) A step of preparing a first cDNA sample from a subject with the disease and a second cDNA sample from a subject without the disease, wherein the sample contains a double-stranded cDNA template, and a portion of the template has a known 5' pre-binding adapter and a known 3' pre-binding adapter. (ii) A step of denaturing each cDNA sample and preparing a single-stranded cDNA template, (iii) Annealing a 5' adapter complex to a 5' pre-binding adapter of at least one single-stranded cDNA template, and annealing a 3' adapter complex to a 3' pre-binding adapter of the same single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide, (iv) Ligating oligonucleotides from the 5' adapter complex to the 5' pre-binding adapter of a single-stranded cDNA template, and ligating oligonucleotides from the 3' adapter complex to the 3' pre-binding adapter of the same single-stranded cDNA template, (v) A step of selectively amplifying each cDNA sample using primers specific to the ligated oligonucleotide, and (vi) A step to discover disease biomarkers by sequencing each cDNA sample and comparing the sequencing outputs for the first and second cDNA samples. This provides a method that includes [something].
[0120] According to all aspects of the present invention, the cDNA sample may be a cDNA sample from a single cell.
[0121] The embodiments of the selective amplification method for single-stranded cDNA described above are applicable mutatis mutandis to the selective amplification method for cDNA containing known adapter sequences and will not be repeated for the sake of brevity. Oligonucleotide dimer compositions are suitable for use in the methods defined herein for the selective amplification of cDNA containing known adapter sequences. The methods defined herein for the selective amplification of cDNA containing known adapter sequences may be used for the discovery of disease biomarkers and / or cancer vaccine targets, or in the process of diagnosing and / or determining the prognosis of diseases. The kits and reagent sets defined herein are also suitable for use in the methods for the selective amplification of cDNA containing known adapter sequences.
[0122] In some embodiments, according to all aspects of the present invention, the cDNA sample contains 800 ng, 700 ng, 500 ng, 100 ng, 20 ng, 10 ng, 5 ng, or 1 ng or less of starting cDNA. The cDNA sample may also contain 1 to 800 ng, 1 to 500 ng, 5 to 100 ng, or 10 to 50 ng of starting cDNA.
[0123] In certain embodiments, according to all aspects of the present invention, RNA from a sample is first reverse transcribed into cDNA. Sample types include blood samples (particularly plasma-derived, but also serum), saliva, urine, lymph, and other bodily fluids. Other sample types include solid tissues, including frozen tissue and formalin-fixed paraffin-embedded (FFPE) materials. RNA may be messenger RNA (mRNA), microRNA (miRNA), etc. In such embodiments, RNA is generally reverse transcribed using reverse transcriptase to form complementary DNA (cDNA) molecules. Methods for reverse transcribing RNA into cDNA using reverse transcriptase are well known in the art. Any suitable reverse transcriptase can be used, and examples of suitable reverse transcriptases are widely available in the art. The initial cDNA molecule may be single-stranded until a complementary strand is generated using DNA polymerase. RNA can be converted to double-stranded cDNA with 5' and 3' adapters using commercially available kits (such as the NEBNext® Single Cell / Low Input cDNA Synthesis & Amplification Module). Phosphates can be added to cDNA using primers based on 5' and 3' adapters. Before using cDNA as a cDNA sample, a cDNA purification step (e.g., using ProNex or Ampure beads) can also be performed.
[0124] Since the present invention requires only a small amount of starting cDNA, cDNA can be generated from a small amount of RNA and / or without requiring additional PCR cycles during cDNA generation. The RNA sample may contain 3.5 μg, 3 μg, 2 μg, 1 μg, 500 ng, 100 ng, 10 ng, or less than 1 ng of starting RNA. The RNA sample may contain 1 ng to 3 μg, 10 ng to 2 μg, or 100 ng to 1 μg of starting RNA.
[0125] Furthermore, the present invention is (a) One or more test devices for normalizing a first cDNA sample from a subject with the disease (cancer) and a second cDNA sample from a subject without the disease (cancer), and for sequencing the normalized first and second cDNA samples. (b) Processor, and (c) When executed by a processor, (i) Access the determined sequence(s) for the first and second cDNA samples of one or more test devices: (ii) Calculate whether there are any cDNA sequences that are present in the first cDNA sample but not in the second cDNA sample, or that are present in the second cDNA sample but not in the first cDNA sample, and (iii) Output the result of step (ii) from the processor. Storage media containing computer applications configured in such a way The present invention provides a system or test kit for discovering disease biomarkers and / or cancer vaccine targets, including those mentioned above.
[0126] In related embodiments, (a) One or more test devices for normalizing cDNA samples from subjects and sequencing the normalized cDNA samples, (b) Processor, and (c) When executed by a processor, (i) Access the determined sequence(s) for one or more cDNA samples from test devices, (ii) Calculate whether the cDNA sequence is present or absent, where the presence or absence of the cDNA sequence is associated with the disease, and (iii) Output from the processor whether the subject has a disease and / or a prognosis for the disease. Storage media containing computer applications configured in such a way A system or test kit is provided for diagnosing and / or determining the prognosis of a disease in a subject, including the above.
[0127] One or more test devices may use / include the oligonucleotide dimer compositions described herein. One or more test devices may include a long-read sequencer.
[0128] The system or test kit may further include a display for output from the processor.
[0129] Furthermore, a computer application as defined herein or a storage medium containing a computer application is provided. [Brief explanation of the drawing]
[0130] [Figure 1] This is a schematic diagram showing the addition of a ligation sequence to the end of a single-stranded cDNA template. [Figure 2] This is a schematic diagram of one embodiment of the cDNA normalization process according to the present invention. [Figure 3] Figure 3A is a schematic diagram of the forward and backward oligonucleotide dimer structures annealed to a single-stranded cDNA template, with Figure 3B being a diagram illustrating the sequence representation of the examples described herein. [Figure 3A] Figure 3A is a schematic diagram of the forward and backward oligonucleotide dimer structures annealed to a single-stranded cDNA template, with Figure 3B being a diagram illustrating the sequence representation of the examples described herein. [Figure 3B] Figure 3A is a schematic diagram of the forward and backward oligonucleotide dimer structures annealed to a single-stranded cDNA template, with Figure 3B being a diagram illustrating the sequence representation of the examples described herein. [Figure 4] This figure shows a graph illustrating the length distribution of input cDNA obtained by gel electrophoresis. [Figure 5] As shown in Figure 2, this is a graph showing the length distribution of normalized cDNA obtained from the cDNA normalization process, as determined by gel electrophoresis. [Figure 6]This figure shows the saturation curves of the input cDNA and the normalized cDNA, as shown in Figure 2, obtained using nanopore cDNA sequencing. [Figure 7] This figure shows a 2D PCA plot from 13,573 transcripts illustrating the clustering of cancer and control samples in the dataset.
[0131] [Detailed explanation] Next, the above and other aspects of the present invention will be described in further detail for illustrative purposes only, with reference to the following examples and accompanying figures.
[0132] To address the limitations of current cDNA normalization techniques, the inventors have developed an improved selective amplification method (referred to as "level-up"). This method utilizes the same denaturation and rehybridization steps as the DSN and column methods described herein. However, this method differs from other approaches by employing a non-depletion or additive mechanism. In other words, level-up provides methods and means for increasing the amount of low-abundance cDNA in a sample. These methods and means can be used in cDNA normalization in sequencing processes, or in other processes that benefit from the amplification of low-abundance cDNA, such as biomarker discovery, detection, or identification.
[0133] In the context of the present invention, the following explanations of terms and methods are provided to better describe the present disclosure and to provide guidance for the implementation of the present disclosure.
[0134] The term "selective amplification" is used to describe a method developed by the inventors that preferentially amplifies a particular DNA template over other DNA templates, for example, a method that amplifies only single-stranded cDNA in a sample containing a mixture of single-stranded and double-stranded cDNA. This term is also used herein to describe preferentially amplifying a particular category of DNA, such as low-abundance DNA.
[0135] The term "adapter" is used to describe short DNA sequences that are added to the ends of a DNA template, such as those commonly used in RNA sequencing, by ligating the adapter to the cDNA template. A "3' pre-bonding adapter" refers to an adapter with a known nucleotide sequence added to the 3' end of the cDNA template. A "5' pre-bonding adapter" refers to an adapter with a known nucleotide sequence added to the 5' end of the cDNA template.
[0136] The terms "normalization" and "normalized fraction" refer to the process of equalizing the abundance of different transcripts within a sample. This can be achieved by using conventional methods that reduce the amount of high-abundance transcripts, or by using the methods described herein that selectively amplify low-abundance transcripts.
[0137] As used in the context of the present invention, the term “post-association single-stranded cDNA template” refers to a single-stranded cDNA template produced by dissociating and reassociating (i.e., denaturing and rehybridizing) a sample of double-stranded cDNA to form a mixture of single-stranded and double-stranded cDNA. Single-stranded cDNA that remains single-stranded after reassociation is called post-association single-stranded cDNA. In embodiments defined herein where the reassociation step is not performed, the term “post-association single-stranded cDNA template” is interchangeable with “single-stranded cDNA template.”
[0138] As used herein, the term "ligation adapter cDNA template" refers to a cDNA template formed by ligating an adapter to a cDNA template.
[0139] The present invention encompasses an “adapter complex” suitable for annealing to the ends of a post-association single-stranded cDNA template. The term “adapter complex” refers to an adapter comprising multiple components.
[0140] As used in the context of this invention, the terms “forward oligonucleotide dimer,” “forward k-linker,” and “forward dimer” refer to an adapter complex that can anneal to the 5' end of a single-stranded cDNA template after association. As used in the context of this invention, the terms “backward oligonucleotide dimer,” “backward k-linker,” and “backward dimer” refer to an adapter complex that can anneal to the 3' end of a single-stranded cDNA template after association.
[0141] As used in the context of this invention, the terms "forward rig-oligonucleotide" and "forward rig" refer to the oligonucleotide component of the forward dimer. As used in the context of this invention, the terms "backward rig-oligonucleotide" or "backward rig" refer to the oligonucleotide component of the backward dimer.
[0142] As used in the context of this invention, the terms "forward-linked oligonucleotide" and "forward link" refer to the oligonucleotide component of a forward dimer. As used in the context of this invention, the terms "backward-linked oligonucleotide" and "backward link" refer to the oligonucleotide component of a backward dimer.
[0143] In the context of this invention, the term "overhang" is used to describe an overhang region of the dimer sequence of this invention, where the overhang region, once annealed to the post-associated single-stranded cDNA template, is complementary to its paired region and does not bind to that paired region.
[0144] Single-stranded cDNA is selectively amplified using primers specific to the ligated oligonucleotide. Since single-stranded DNA is the template, the primer region of one primer in the specific primer pair is complementary to the single-stranded DNA molecule. The other primer in the specific primer pair contains a primer region that is complementary to the complementary single-stranded DNA molecule formed during the amplification cycle and therefore hybridizes. Thus, one primer is complementary to one of the ligated oligonucleotides, and the other primer contains (at least a partial) sequence of the other ligated oligonucleotide.
[0145] Methods of cDNA normalization developed to date Two methods have been developed to date for the normalization of full-length cDNA: the duplex-specific nuclease (DSN) method (Zhulidov, PA et al., Simple cDNA normalization using kamchatka crab duplex-specific nuclease. Nucleic Acids Res. 32, e37 (2004)) and the hydroxyapatite column method (Andrews-Pfannkoch, C., Fadrosh, DW, Thorpe, J., and Williamson, SJ, Hydroxyapatite-mediated separation of double-stranded DNA, single-stranded DNA, and RNA genomes from natural viral assemblages. Appl. Environ. Microbiol. 76, 5039~5045 (2010)). Both methods rely on cDNA strand denaturation and rehybridization. As single-stranded cDNA moves through solution, sequences that are more abundant have a higher probability of finding complementary sequences that match for rehybridization. Therefore, when rehybridization reaches its limit, the remaining single-stranded cDNA represents a normalized sequence library.
[0146] The difference between the two methods lies in the approach taken to isolate a single-stranded cDNA library from the re-hybridized double-stranded cDNA molecule.
[0147] In the DSN method, all double-stranded cDNA in solution is degraded using an enzyme that specifically cleaves double-stranded DNA. The solution is then purified, and cDNA sequences longer than a certain length are size-selected. These sequences are then amplified using polymerase chain reaction (PCR).
[0148] In the column assay, a denatured and rehybridized cDNA library is passed through a heated column packed with hydroxyapatite granules. Hydroxyapatite preferentially binds to larger DNA molecules. The size of the bound DNA is controlled by the concentration of the phosphate buffer used to dissolve the cDNA library. Therefore, the phosphate buffer concentration must be specifically adjusted for cDNA molecules within a certain range of sequence lengths. The cDNA is eluted through the column while increasing the phosphate buffer concentration, and larger DNA molecules are extracted as the size increases. Since single-stranded cDNA is approximately half the size of rehybridized cDNA, the elution of the single-stranded fraction can be controlled if the average cDNA sequence length is known. The resulting eluate is intended to concentrate single-stranded cDNA and then amplify it using PCR.
[0149] In both DSN and column chromatography, known adapters must be attached to the ends of the cDNA before normalization to facilitate PCR amplification (so that appropriate primers can be used).
[0150] Both methods are essentially subtractive, depleting large fractions of cDNA. Therefore, the amount of starting cDNA generally required is 1 μg or more for the DSN method and 4 μg or more for the column method.
[0151] The DSN method uses an enzyme that cleaves all double-stranded cDNA, so theoretically, it can remove less abundant sequences in segments that match more abundant sequences. This effect can also increase the probability of PCR chimera formation. PCR chimeras are formed when an incomplete single-stranded cDNA sequence acts as a primer for other sequences, causing sequences to bind in ways that do not occur naturally. PCR chimeras are false positives for novel isoforms and are extremely difficult to distinguish from true alternative isoforms. Verification of PCR chimeras generally requires thorough biochemical assays.
[0152] Columnar chromatography can only separate high-abundance and low-abundance fractions within a narrow size range, resulting in a significant bias towards longer cDNA sequences. This effect leads to the loss of representation of longer RNA sequences.
[0153] Columnar method As described above, the hydroxyapatite column assay relies on the denaturation and rehybridization of cDNA strands. As single-stranded cDNA moves through solution, sequences that are more abundant have a higher probability of finding complementary sequences that match for rehybridization. In the column assay, the denatured and rehybridized cDNA library is passed through a heated column packed with hydroxyapatite granules. Hydroxyapatite preferentially binds to larger DNA molecules. Since single-stranded cDNA becomes approximately half the size of rehybridized cDNA, the elution of the single-stranded fraction can be controlled if the average cDNA sequence length is known. The resulting eluate is intended to concentrate the single-stranded cDNA and then amplify it using PCR.
[0154] The hydroxyapatite column assay was performed using 4 μg of cDNA as the starting sample, as described in Andrews-Pfannkoch et al. (Andrews-Pfannkoch, C., Fadrosh, DW, Thorpe, J., and Williamson, SJ.: Hydroxyapatite-mediated separation of double-stranded DNA, single-stranded DNA, and RNA genomes from natural viral assemblages. Appl. Environ. Microbiol. 76, 5039-5045 (2010)). The hydroxyapatite column assay did not produce usable yields when using cDNA of 2 μg or less.
[0155] Because the columnar method is based on size-based separation, it loses representation of long RNA sequences (4kb or longer), which can be observed in the length distribution before and after normalization.
[0156] DSN method Similar to column chromatography, the DSN method relies on the denaturation and rehybridization of cDNA strands. As single-stranded cDNA moves through solution, sequences that are more abundant have a higher probability of finding a matched complementary sequence for rehybridization. In the DSN method, an enzyme that specifically cleaves double-stranded DNA is used to degrade all double-stranded cDNA in solution.
[0157] The commercially available Evrogen Trimmer-2 cDNA normalization kit uses the DSN method. This kit was used according to the manufacturer's instructions, with 1 μg of cDNA as the starting sample. However, it was found that 2 μg of cDNA was necessary to generate sufficient material for long-read RNA sequencing.
[0158] The DSN method was found to completely remove high-abundance RNA, and this was expected to be true for RNA with sequences similar to high-abundance RNA as well. Therefore, over-depletion was observed, where high-abundance RNA was not simply reduced in quantity but completely removed from the sample. Table 1 shows the over-depletion for ranks 1-20 and the selection of lower ranks (55, 64, 77, 92, and 98) where RNA was significantly reduced but not completely removed.
[0159] [Table 1]
[0160] Furthermore, the DSN method can create conditions that lead to the generation of artificial chimeric sequences, which may appear as false positives in gene predictions.
[0161] Selective amplification method (level up) The method is shown in Figures 1 to 3. In summary, referring to Figure 2, a cDNA sample containing cDNA with known 5' and 3' pre-binding adapters is denatured to produce a single-stranded cDNA sample, and then this sample is re-hybridized or re-associated to produce a mixture of single-stranded and double-stranded cDNA. This single-stranded cDNA corresponds to low-abundance cDNA. The single-stranded cDNA in the re-associated sample is then modified by adding oligonucleotides to the 5' and 3' pre-binding adapters. The cDNA sample is amplified using primers for these oligonucleotides. This step selectively increases the content of only single-stranded cDNA in the re-associated sample, thereby increasing the content of low-abundance cDNA in the sample and producing a normalized sample.
[0162] The input cDNA library for selective amplification is double-stranded, and each double-stranded template contains a 5' pre-bonding adapter for a known nucleotide sequence and a 3' pre-bonding adapter for a known nucleotide sequence. Because this method is additive, it requires less starting cDNA compared to prior art normalization methods. The DSNase method requires a minimum of 1 μg of input cDNA, while the column method requires a minimum of 4 μg. Tests have shown that this method can be applied with as little as 20 ng of starting cDNA.
[0163] The input cDNA is combined with the hybridization buffer, and the solution is heated to the denaturation temperature of approximately 98 degrees Celsius to create a denatured single-stranded cDNA template. After 5-10 minutes, the solution is then cooled to the rehybridization temperature of approximately 68 degrees Celsius. Depending on the amount of normalization required, the solution is incubated at this temperature for 0-24 hours, for example, 3-10 hours. 7 hours is a typical time for the re-association step. This step generates a re-associated or re-hybridized sample containing the post-associated double-stranded cDNA template and the post-associated single-stranded cDNA template.
[0164] After incubation, the oligonucleotide dimers of the present invention (referred to by the inventors as K-linkers) are added. These oligonucleotide dimers are described in detail below. At this point, the solution can be incubated at 68°C for 0 to 1 hour. 5 minutes is a typical incubation time. The solution is then cooled to the annealing temperature of the K-linkers, which is typically 40 to 60°C, for example, 44°C. The solution is incubated at this temperature for 10 minutes to 2 hours, for example, 25 minutes. This step allows the K-linkers to anneal to the post-association single-stranded cDNA template.
[0165] After this incubation time, DNA ligase is added along with the ligation mix. This solution is incubated at the same temperature for 0.5–2 hours, for example, 1 hour, and then cooled to room temperature. This step forms a ligation adapter cDNA template in which the oligonucleotides of the K-linker are ligated to both ends of the post-association single-stranded cDNA template. At this point, the cDNA may be purified (e.g., using Pronex or Ampure beads) or used directly for PCR amplification with primers based on the K-linker sequence.
[0166] After PCR amplification, the cDNA is purified using an appropriate method, and the resulting cDNA corresponds to a normalized cDNA library.
[0167] Although the double-stranded cDNA template after association can be removed before PCR amplification, our tests have shown that leaving double-stranded cDNA in solution does not adversely affect the normalization process.
[0168] In fact, post-association double-stranded cDNA can also be analyzed and used, for example, to obtain estimates of gene expression. This involves an additional selective (PCR) amplification step using primers for known 5' pre-binding adapters and known 3' pre-binding adapters, with molecular barcoding included in the primers. PCR amplification is performed first with primers based on the K-linker sequence, followed by one PCR cycle using primers for known 5' pre-binding adapters and known 3' pre-binding adapters. PCR can be paused to add further primers for the final cycle. Alternatively, the cDNA can be purified after PCR with primers based on the K-linker sequence, and a new PCR cycle can be performed using primers for known 5' pre-binding adapters and known 3' pre-binding adapters. In either case, the product is a mixture of two distinguishable fractions containing sequences derived from the post-association double-stranded cDNA template and sequences derived from the post-association single-stranded cDNA template. The addition of molecular barcoding allows for the identification of source molecules during sequencing analysis. This aspect of the present invention can be used for multiplex sequencing.
[0169] Design of oligonucleotide dimer complexes The oligonucleotide dimer composition of the present invention comprises forward and backward oligonucleotide dimers (forward and backward K-linkers), both of which anneal to the same strand of the single-stranded cDNA after association.
[0170] Referring to Figures 3 and 3A, each K-linker contains two oligonucleotide sequences, one called the link-oligonucleotide (also known as the "link"; referred to as the "LU adapter linker" in Figures 3 and 3A) and the other called the rig-oligonucleotide (also known as the "rig"; referred to as the "LU adapter" in Figures 3 and 3A).
[0171] The link sequence includes regions complementary to known 3' / 5' pre-binding adapter sequences (pre-attached to cDNA) and regions complementary to the rig.
[0172] As shown in Figures 3 and 3A, the link can be designed to have an overhang region at one end that is non-complementary to a known pre-binding adapter sequence; this region is called the "template overhang." The other end of the link can have a similar overhang that is non-complementary to the rig sequence; this is called the "rig overhang" or "rig-oligonucleotide overhang."
[0173] The purpose of this link is to anneal to both the pre-binding adapter and the rig of the single-stranded cDNA template after association, so that when in use, one end of the cDNA template is adjacent to one end of the rig. By positioning the cDNA template and rig in this way, the cDNA template can be ligated to the rig using DNA ligase, and thus the rig sequence can be added to the end of the cDNA template. The rig sequence is added to both the 5' and 3' ends of the single-stranded cDNA template. Forward K-linkers and backward K-linkers are used to add these rigs to the 5' and 3' ends of the cDNA template, respectively.
[0174] In the method described above, when rigs are attached to both ends of a single-stranded cDNA template, primers based on the sequences of the attached rigs are used to selectively amplify the single-stranded cDNA fraction that has successfully ligated to both the anterior and posterior rig sequences. In this way, PCR can be used to amplify only the low-abundance post-association single-stranded cDNA fraction.
[0175] The specific structure of the K-linker, with forward and rear K-linkers having different functions provided by their unique structural characteristics, brings advantages to overall normalization performance.
[0176] A forward K-linker that binds to the 5' end of a single-stranded cDNA template (shown in the figure as the reverse complement of the original RNA sequence) can be designed so that the K-linker does not act as a primer during PCR amplification.
[0177] To offer further advantages, the forward K-linker does not have a blunt end on the rig side, allowing the use of DNA ligases that can perform blunt-end ligation during the selective amplification process. By providing a non-blunt end on the rig side of the K-linker complex, ligation with other K-linker complexes or double-stranded cDNA in solution is also avoided.
[0178] Because the forward linker's 5' to 3' directionality is detached from the template, the linker itself cannot act as a primer for the template. However, in some cases, the forward rig or PCR primer may anneal to the linker during PCR amplification and, upon polymerase extension, incorporate the template-side sequence of the link. By providing a link with an overhang on the template side, it is possible to avoid the extended rig / primer acting as a primer for cDNA sequences that do not have a rig sequence attached to their terminals, i.e., sequences that are abundant. Therefore, a template overhang structure for posterior links can be provided.
[0179] A posterior K-linker that binds to the 3' end of a single-stranded cDNA template (shown in the figure as the reverse complement of the original RNA sequence) can be designed so that the K-linker does not act as a primer during PCR amplification. This can be achieved using template overhang, which has been described above and is shown in Figures 3 and 3A.
[0180] Posterior K-linkers can be designed without blunt ends on the rig side for the same reasons that anterior K-linkers can be designed with an overhang on the rig side. If the posterior K-linker complex has blunt ends on the rig side, the rig may, in some cases, ligate to other K-linker complexes or double-stranded cDNA in the sample, which could lead to the ligation of the cDNA template.
[0181] Overhangs also play a role in lowering the annealing temperature of the K-linker complex, making it less likely for them to act as primers for each other during PCR amplification where higher annealing temperatures are used. Overhangs can reduce unintended priming. In addition or alternatively, overhangs can provide an indicator that allows for the measurement of the amount of unintended priming. For example, if a K-linker complex with an overhang can prime a template, the overhang sequence can be seen in the sequencing data and it can be concluded that it is a product of unintended priming.
[0182] All overhangs should ideally be between approximately 1 bp and 20 bp, preferably 3 bp. Longer overhangs can be used, but the design becomes more difficult because longer overhangs result in fewer sequence compositions that prevent unintended priming.
[0183] The complementary regions between the cDNA template and the links, and between the links and the rig, must be long enough to anneal at the activity temperature of the ligase used. Nick repair ligases are particularly preferred for this process, which generally requires about five or more complementary bases on either side of the ligation site.
[0184] To reduce carryover during the purification process, the combined length of the link and the corresponding rig should ideally be less than approximately 300 bp, and preferably less than approximately 200 bp.
[0185] Example representation of an array: Examples of all sequence structures are provided in the 5' to 3' direction, also shown in Figure 3B. Forward link: OOOOOXXXXXXXXXXXXXXXFFFFFFFFFFFFFFFOOOOO Forward rig: [ka] Back links: OOOOOBBBBBBBBBBBBBXXXXXXXXXXXXXXOOOOOO Rear rig: [ka] Nucleotides complementary to the X-5' / 3' pre-bonding adapter O-overhang arrangement A complementary arrangement to the forward rig that is F-ligated. F (italic) - Forward rig arrangement A complementary arrangement for the rear rig that is B-ligated. B (italic) - Rear rig arrangement
[0186] Examples of oligonucleotides and single-stranded cDNA templates used in selective amplification methods: NEB / PacBio cDNA synthesis kit primer sequences Iso-Seq Express Fwd: GGCAATGAAGTCGCAGGGTTG Iso-Seq Express Rev: AAGCAGTGGTATCAACGCAGAG Forward link: [ka] Forward rig and primer: [ka] Back links: [ka] Rear rig: [ka] Primer: [ka] Overhangs are indicated by underlines. Complementary regions between oligonucleotides are indicated in bold. Single-stranded cDNA template: [ka] (That is, the sequence of the 5'-3'-5' pre-binding adapter, the sequence of the cDNA represented by N, and the sequence of the 3' pre-binding adapter. Regions complementary to the forward and backward link oligonucleotides are shown in bold.)
[0187] Preparation of oligonucleotide dimers Once designed, the dimer can be prepared using standard techniques well known in this field.
[0188] Amplification of oligonucleotide-cDNA templates After ligating the rig to the 3' and 5' ends of the cDNA template, the resulting solution can be purified using an appropriate cDNA purification method. While the purification step can be omitted, doing so may reduce PCR amplification efficiency.
[0189] After purification or ligation, the obtained material can be PCR-amplified using forward and reverse primers based on the rig sequence. The primer sequences can be selected to have a higher annealing temperature than the complementary region of the template / link / rig to avoid unwanted priming.
[0190] If necessary, the optimal number of PCR cycles can be determined by first performing a qPCR experiment to identify the inflection point of the amplification curve.
[0191] Subsequently, the cDNA obtained from PCR amplification is purified and can be used as an input for downstream processes such as sequencing.
[0192] Validation of selectively amplified samples The effect of selective amplification or normalization can also be measured indirectly by measuring the length distribution of the cDNA library using gel electrophoresis, or directly using sequencing.
[0193] To confirm the effect of normalization using gel electrophoresis, the length distribution plot of the input cDNA can be compared with the normalized cDNA. The input cDNA generally shows peaks along the length distribution, which correspond to highly abundant transcript sequences (Figure 4).
[0194] The normalized cDNA has a length distribution similar to a normal distribution without sharp peaks (Figure 5). This represents a uniform distribution across the entire transcript sequence.
[0195] When using sequencing for direct measurement of normalization, the preferred sequencing method is long-read sequencing. This enables the identification of different isoforms. The results of direct measurement may be either a plot of the number of reads per gene or a saturation plot showing the number of newly identified genes as the sequencing depth increases.
[0196] To verify this method, the inventors performed Nanopore cDNA sequencing on both the input cDNA library and the cDNA library prepared by this method. The inventors compared the saturation curves of each library (Figure 6). The curve of this method is higher by multiple factors than the curve of the input cDNA. This indicates that the cDNA library is more evenly dispersed.
[0197] Applications of selective amplification The main use of the selective amplification of this method is to improve the discovery and detection of low-abundance genes and isoforms. Combining this method with sequencing improves the sampling efficiency for identifying all unique genes within a sample.
[0198] This method can be applied to any double-stranded cDNA library that has known adapters at its ends and is of a length that can be amplified by the PCR method. This means it can be used for DNA sequencing.
[0199] Another important aspect of this method is the ability to select cDNA with known lig sequences. This aspect can be applied to single-cell sequencing, which requires adapters to assign reads to individual cells. In this application, cDNA sequences without barcodes / adapters for cell identification occur within the cDNA library. These are known as template switching oligo (TSO) artifacts and are not desirable in single-cell sequencing projects as they cannot be assigned to their originating cells. By applying this method with a short re-hybridization step, only cDNA sequences with the desired lig sequences can be selected, and thus the sequencing of TSO artifacts can be effectively restricted. This method can also be implemented in single-cell cDNA libraries to remove TSO artifacts and improve the transcriptome coverage per cell.
[0200] Consideration The selective cDNA amplification of this method is an innovative way to achieve more non-target discovery / detection of low-abundance nucleic acids. Since it is not targeted, there is no need to know which sequences are low-abundance or high-abundance sequences in a given sample.
[0201] The difference from existing normalization approaches is that, while current approaches use depletion methods, this method uses additive methods. Additive methods allow the use of this method with very small amounts of starting cDNA. This assists in the detection of low-abundance nucleic acids and provides greater confidence in determining the absence of specific nucleic acids. Additive methods also prevent the excessive depletion and artificial chimerization inherent in the DSNase method. This method also exhibits length bias only to a certain extent, whereas PCR amplification does not exhibit length bias. Therefore, it can be successfully performed even with longer cDNA molecules.
[0202] The systems, methods, and actions described in the examples of previously shown embodiments are illustrative, and in alternative embodiments, certain actions may be performed in a different order, in parallel with one another, omitted entirely, and / or combined between embodiments of different examples, and / or certain additional actions may be performed without departing from the scope and spirit of the various embodiments. Accordingly, such alternative embodiments are included in the examples described herein.
[0203] Although specific embodiments are described in detail above, the description is for illustrative purposes only. Therefore, it should be understood that many of the above embodiments are not intended as necessary or essential elements unless explicitly indicated otherwise.
[0204] In addition to the embodiments described above, modifications to the embodiments disclosed in the examples, and equivalent components or actions corresponding to those embodiments, can be made by those skilled in the art in order to have the merits of the Disclosure without departing from the spirit and scope of the embodiments as defined in the following claims, and the claims should be construed in the broadest way to encompass such modifications and equivalent structures. [Examples]
[0205] The present invention can be further understood by referring to the following experimental examples.
[0206] Example 1 Using a novel transcriptome discovery platform to detect unique RNA signatures in epithelial cancers. The inventors investigated the use of RNA level-up from cancer cell lines and blood samples. These initial studies discovered thousands of previously unannotated isoforms with potential for use as cancer biomarkers. Figure 7 shows clustering of cancer samples (blood samples taken from individuals diagnosed with breast cancer) and control samples (blood samples taken from individuals not diagnosed with breast cancer) based on 13,573 transcripts identified in these initial studies. Cancer samples can be easily distinguished from control samples.
[0207] An example protocol for discovering unique RNA signatures in epithelial cancers is presented.
[0208] Stage 1 Blood handling validation will be performed using five samples from five individuals (25 samples in total). Blood will be collected and stored in PAXgene RNA blood tubes (each with a blood capacity of 2.5 ml). On the same day as collection, the tubes will be transported on dry ice from collection centers in various locations.
[0209] For each biological replicate, five collected tubes are processed according to various simulation handling conditions. These include: 1. Dispose of the pipes on the same day as collection. 2. Process the pipes 72 hours after collection. 3. Dispose of the pipes two weeks after collection. 4. Dispose of the pipes four weeks after collection. 5. Store the tubes at -80°C for an extended period and process them after 4 to 12 months.
[0210] Extract RNA from PAXgene tubes using the Qiagen PAXgene Blood RNA Kit (Catalog number / ID: 762164). Convert the resulting RNA into cDNA using the NEBNext® Single Cell / Low Input cDNA Synthesis & Amplification Module (E6421S). Process the resulting cDNA by the normalization described herein. Then, sequence the normalized cDNA library using an Oxford Nanopore Technologies Minion MK1C sequencer.
[0211] Stage 2 Once sample processing is optimized and the quality of the generated data is ensured, the study is then advanced to patient samples.
[0212] Selection criteria and exclusion criteria Selection criteria: All female patients aged 18 or older and under 50 years old, premenopausal, with invasive breast cancer or a benign condition (B2).
[0213] Exclusion criteria: Any history of cancer, or other types of cancer occurring concurrently, autoimmune diseases, and women who cannot or do not wish to give informed consent. Patient identification and recruitment into a prospective cohort study Eligible patients are identified based on the results of triple assessment from an interdisciplinary team (MDT) meeting. When patients receive the results from a breast surgeon at the clinic, they are invited to participate in the study and receive a patient information sheet (PIS). Patients who receive a PIS are followed up by the study team after 72 hours. If patients choose to participate in the study, they meet with a study nurse and sign a consent form or consent remotely.
[0214] Anonymization Once enrolled in the study, research nurses assign patients a unique ID number that is linked to their background data through a secure database and used to collect and identify samples within the database. Background data collection
[0215] The following data will be collected for each patient: basic demographics, menopausal status, past medical history, medication history, drug and alcohol use, BMI, and family history.
[0216] A triple evaluation including symptoms, laboratory findings, imaging, and biopsy results will be compared. Postoperative history will also be collected. Additionally, any additional prognostic information, such as the use of biomolecular assays, will be gathered.
[0217] Blood sample collection On the morning of surgery, blood samples are taken, stored at -20°C, and transported to the processing center. Any patients not undergoing surgery, such as those with fibroadenomas, are instructed to come in on the day when blood tests will be performed by the same team. The sample size is 20 ml, with a maximum collection volume of 30 ml per day.
[0218] Processing of blood samples Process blood samples from collection tubes in a layered flow class 2 safety cabinet using an RNA extraction kit provided by Qiagen. After RNA extraction, autoclave and dispose of any remaining solid waste using a human medical waste removal service provider. Liquid waste is purified and disposed of in accordance with the University of Edinburgh's guidance on the use of human samples. Blood samples not processed on the same day of arrival are stored in a stable -80°C freezer. Use sealed containers to transfer blood samples between the freezer and the safety cabinet.
[0219] Data Analysis Transcriptomics data is stored and processed in a secure Amazon Web Services cloud repository and servers. No personal information is stored in the same computer location.
[0220] cDNA sequencing is performed on the sample using an Oxford Nanopore Technologies (ONT) Minion sequencing machine. This outputs the raw data as a fast5 file. The latest high-precision basecaller (basecaller) from ONT (Bonito) is used to convert it to a fastq sequence file. Nanopore reads are filtered for quality using a seqkit, and adapters are removed using pychopper. The trimmed reads are then mapped to a human reference genome assembly (HG38 or a newer version) using Minimap2.
[0221] Stage 3-4 The study will be modified to initially increase sample size, and then expand to include colorectal cancer and ovarian cancer.
[0222] Statistical considerations and sample size For Stage 1 of the study, five samples will be processed. This is sufficient to produce technical and biological replicates and to ensure initial quality assessment.
[0223] Data handling and processing will be carried out in the same manner as in Stage 2.
[0224] For Stage 2 of the study, 30 samples from benign cells and 30 samples from cancer patients will be collected and processed initially. This data will inform the next sample sizes for Stages 3 and 4 of the overall study.
[0225] ethical approval Ethical approval from the National Research Ethics Service is required before commencing the research.
[0226] Informed Consent Process Written informed consent will be obtained from all patients participating in this study for the purpose of collecting patient demographic and clinical data, as well as for the storage of blood samples and subsequent transcriptome materials and data. If patients choose to do so, they will be asked for their consent by telephone. Patients can withdraw their consent at any time. Individuals deemed unable to give informed consent will be excluded from the study.
[0227] quality control In silico quality control testing is performed to assess that the samples are accurately labeled and that there are no handling issues with each sample. This testing consists of checking for known genes that should be present in all samples, and clustering analysis to identify outliers that may indicate sample contamination.
[0228] Furthermore, genes that shouldn't exist but are present in the sample data are searched for. For example, genes from other species.
[0229] Example 2 Blood sample collection technique for full-length RNA extraction cDNA is DNA synthesized from an RNA template. Therefore, the quality of a cDNA sample is related to the RNA it is reverse transcribed from. Conventional processing for handling blood samples before RNA extraction involves thawing frozen blood samples overnight. We have found that this causes significant RNA degradation, which negatively impacts long-read sequencing. We process blood samples before RNA extraction using the following protocol to minimize degradation and optimize RNA extraction for long-read sequencing.
[0230] Blood sample collection technique for full-length RNA extraction Collect 2.5 ml of blood at room temperature (18-25°C) into a PAXgene Blood RNA Tube. Gently invert the blood tube 10 times immediately after blood collection.
[0231] If RNA extraction is scheduled for the same day as sample collection, store the blood sample upright at room temperature (18-25°C) for 2-3 hours, then immediately perform RNA extraction using the PAXgene Blood RNA Kit. If RNA extraction is not performed on the day of blood collection, follow the instructions below regarding sample freezing, storage, and thawing: Blood samples should be stored at -20°C or below immediately after collection.
[0232] For long-term storage, freeze blood samples at -20°C for 24 hours, then transfer them to a freezer at -70°C or -80°C.
[0233] If the samples are to be transported to different locations, they should be transported on dry ice to ensure they remain frozen during transit.
[0234] On the day of RNA extraction, thaw the blood sample by placing the sample tube upright in a rack and incubating at room temperature (18-25°C) for 1-3 hours. Once the blood is completely thawed, gently invert the sample tube 10 times, incubate at room temperature for a further 2 hours, and then immediately perform RNA extraction using the PAXgene Blood RNA Kit.
[0235] As shown in Table 2, RNA integrity is improved with a shorter thawing period (3 hours) compared to overnight thawing. The RNA Integrity Number (RNA Integrity Number) was defined by Schroeder et al. (The RIN: an RNA integrity number for assigning integrity values to RNA measurements. BMC Molecular Biology 7, 3 (2006)). https: / / doi.org / 10.1186 / 1471-2199-7-3The calculations were performed as shown in [reference]. For the "-20°C for 2 weeks; thawed overnight" and "-20°C for 1 month; thawed for 3 hours" tests (tests 2 and 3), samples were placed at -20°C within 12 hours. Five individuals were used in each test. All thawing was performed at room temperature.
[0236] [Table 2]
[0237] The present invention is not limited in scope by the specific embodiments described herein. In fact, in addition to the embodiments described herein, various modifications of the present invention will become apparent to those skilled in the art from the above description and the accompanying drawings. Such modifications are intended to fall within the scope of the accompanying claims. Furthermore, all embodiments described herein are, as appropriate, broadly applicable and can be combined with any and all other consistent embodiments.
[0238] Various publications are cited herein, and their disclosures are incorporated herein by reference in their entirety.
Claims
1. A method for discovering disease biomarkers, (i) A step of preparing a first cDNA sample from a subject having the disease and a second cDNA sample from a subject not having the disease, (ii) A step of normalizing the first and second cDNA samples, (iii) The step of sequencing the normalized first and second cDNA samples, and (iv) A step of discovering disease biomarkers by comparing the sequencing outputs for the first and second cDNA samples, A method that includes this.
2. A method for diagnosing a disease in a subject, (i) A step of preparing a cDNA sample from the subject, (ii) The step of normalizing the cDNA sample, and (iii) A step of sequencing the normalized cDNA sample, wherein the sequencing output is used to determine whether the subject has a disease, A method that includes this.
3. The method according to claim 2, comprising the step of comparing the sequencing output of the normalized cDNA sample with one or more reference sequences or the sequencing output of one or more control samples, wherein the one or more control samples are optionally samples from one or more subjects having the disease and / or one or more subjects not having the disease.
4. The method according to claim 2 or 3, wherein the cDNA sample is obtained from a biological fluid or a fluid or solution produced from a biological substance, and optionally, the cDNA sample is obtained from blood.
5. The method according to any one of claims 2 to 4, further comprising the step of reporting the outcome of the method to the subject.
6. The method according to any one of claims 1 to 5, wherein the disease is cancer.
7. The method according to any one of the claims, wherein the normalization step includes increasing the amount of low-abundance cDNA in each cDNA sample.
8. The first and second cDNA samples each comprise a double-stranded cDNA template, each template having a known 5' pre-bonding adapter and a known 3' pre-bonding adapter. The step of normalizing the first and second cDNA samples is, (i) Denaturing the cDNA sample to prepare a single-stranded cDNA template, (ii) Reassociating the cDNA samples to produce a mixture of the associated single-stranded cDNA template and the associated double-stranded cDNA template. (iii) Annealing a 5' adapter complex to the 5' pre-binding adapter of at least one post-associated single-stranded cDNA template, and annealing a 3' adapter complex to the 3' pre-binding adapter of the same post-associated single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide. (iv) Ligating oligonucleotides from the 5' adapter complex to the 5' pre-binding adapter of the post-association single-stranded cDNA template, and ligating oligonucleotides from the 3' adapter complex to the 3' pre-binding adapter of the same post-association single-stranded cDNA template, (v) Selectively amplifying the cDNA sample using a primer specific to the ligated oligonucleotide. The method according to any one of claims 1, 6, or 7, including
9. The cDNA sample comprises a double-stranded cDNA template, each template having a known 5' pre-binding adapter and a known 3' pre-binding adapter. The step of normalizing the cDNA sample is (i) Denaturing the cDNA sample to prepare a single-stranded cDNA template, (ii) Reassociating the cDNA samples to produce a mixture of the associated single-stranded cDNA template and the associated double-stranded cDNA template. (iii) Annealing a 5' adapter complex to the 5' pre-binding adapter of at least one post-associated single-stranded cDNA template, and annealing a 3' adapter complex to the 3' pre-binding adapter of the same post-associated single-stranded cDNA template, wherein each adapter complex comprises at least one oligonucleotide. (iv) Ligating oligonucleotides from the 5' adapter complex to the 5' pre-binding adapter of the post-association single-stranded cDNA template, and ligating oligonucleotides from the 3' adapter complex to the 3' pre-binding adapter of the same post-association single-stranded cDNA template, (v) Selectively amplifying the cDNA sample using a primer specific to the ligated oligonucleotide. The method according to any one of claims 2 to 7, including
10. (A) The 5' adapter complex, (i) a forward rig oligonucleotide for ligating the 5' pre-binding adapter of the (post-association) single-stranded cDNA template, and (ii) A forward link oligonucleotide for annealing to the 5' pre-bonding adapter and the forward rig oligonucleotide, comprising a region complementary to the 5' pre-bonding adapter and a region complementary to the forward rig oligonucleotide, Includes, During annealing, the terminal end of the forward lig-oligonucleotide is adjacent to the terminal end of the 5' pre-binding adapter, and the forward oligonucleotide dimer enables ligation of the forward lig-oligonucleotide to the 5' pre-binding adapter at the ligation site, and (B) The 3' adapter complex, (i) a back-rig oligonucleotide for ligating the 3' pre-binding adapter of the (post-association) single-stranded cDNA template, and (ii) A backlink oligonucleotide for annealing to the 3' pre-bonding adapter and the backlink oligonucleotide, comprising a region complementary to the 3' pre-bonding adapter and a region complementary to the backlink oligonucleotide, Includes, During annealing, the end of the back lig-oligonucleotide is adjacent to the end of the 3' pre-binding adapter, enabling ligation of the back lig-oligonucleotide to the 3' pre-binding adapter at the ligation site, thus forming a back oligonucleotide dimer. The method according to claim 8 or 9.
11. (A) The forward link-oligonucleotide is (i) a template overhang region at the end of the forward link-oligonucleotide adjacent to the region complementary to the 5' pre-binding adapter, which is a template overhang region that is incomplementary to the corresponding region of the (post-association) single-stranded cDNA template, and / or (ii) A rig-oligonucleotide overhang region at the end of the forward link-oligonucleotide adjacent to the region complementary to the forward rig-oligonucleotide, the rig-oligonucleotide overhang region being incomplementary to the corresponding region of the forward rig-oligonucleotide, including and / or (B) The rear link-oligonucleotide is (i) a template overhang region at the end of the backlink-oligonucleotide adjacent to the region complementary to the 3' pre-binding adapter, which is incomparable to the corresponding region of the (post-associated) single-stranded cDNA template, and / or (ii) A rig-oligonucleotide overhang region at the end of the rear link-oligonucleotide adjacent to the region complementary to the rear rig-oligonucleotide, the rig-oligonucleotide overhang region being incomplementary to the corresponding region of the rear rig-oligonucleotide, The method according to claim 10, including the method described in claim 10.
12. The method according to any one of the claims, wherein the sequencing step includes the use of long-read sequencing.
13. The method according to any one of claims 1, 6-8, or 10-12, wherein the first cDNA sample and the second cDNA sample are obtained from biological fluid or a fluid or lysate produced from biological material, and optionally, the first cDNA sample and the second cDNA sample are obtained from blood.
14. The method according to any one of claims 1, 6 to 8, or 10 to 13, further comprising, prior to step (i), a step of extracting RNA from biological fluids, or fluids or lysates produced from biological substances, from subjects having the disease and subjects not having the disease, and a step of synthesizing cDNA using the RNA as a template.
15. The method according to any one of claims 2 to 7 or 9 to 12, further comprising, before step (i), a step of extracting RNA from a biological fluid from the subject, or from a fluid or lysate produced from a biological substance, and a step of synthesizing cDNA using the RNA as a template.
16. The method according to any one of claims 1, 6 to 8, or 10 to 14, wherein the disease biomarker is a cDNA sequence that is present in the first cDNA sample but not in the second cDNA sample, or that is present in the second cDNA sample but not in the first cDNA sample.
17. A method for processing blood samples, (i) A step of storing the blood sample at -15°C or below, (ii) The step of thawing the blood sample at 5 to 30°C for at least 1 hour, (iii) A step of extracting RNA from the thawed blood sample, A method that includes this.
18. (i) The blood sample is stored at -20°C or below. (ii) The blood sample is melted at 18-25°C and / or (iii The blood sample is allowed to thaw for 1 to 4 hours. The method according to claim 17.
19. The method according to claim 17 or 18, wherein the blood sample is stored at -15°C or below, or -20°C or below, within 12 hours of collection.
20. The method according to any one of claims 2 to 7, 9 to 12, or 15, wherein the cDNA sample is obtained from blood by synthesizing cDNA according to the steps of the method according to any one of claims 17 to 19 and using the extracted RNA as a template.
21. The method according to any one of claims 1, 6-8, 10-14, or 16, wherein the first cDNA sample and the second cDNA sample are obtained from blood by synthesizing cDNA according to the steps of the method according to any one of claims 17-19 and using the extracted RNA as a template.
22. Use of oligonucleotide dimer compositions for selective amplification of single-stranded cDNA by ligation of oligonucleotides to the 5' and 3' ends of a post-association single-stranded cDNA template having known 5' and 3' pre-binding adapters, in a method for discovering cancer biomarkers, wherein the composition is (A) (i) A forward rig oligonucleotide for ligating the 5' pre-binding adapter of the single-stranded cDNA template after association, and (ii) A forward link oligonucleotide for annealing to the 5' pre-bonding adapter and the forward rig oligonucleotide, comprising a region complementary to the 5' pre-bonding adapter and a region complementary to the forward rig oligonucleotide. Includes, A forward oligonucleotide dimer, wherein during annealing, the terminal end of the forward lig-oligonucleotide is adjacent to the terminal end of the 5' pre-binding adapter, enabling ligation of the forward lig-oligonucleotide to the 5' pre-binding adapter at the ligation site, and (B) (i) A back-rig oligonucleotide for ligating the 3' pre-binding adapter of the single-stranded cDNA template after association, and (ii) A backlink oligonucleotide for annealing to the 3' pre-bonding adapter and the backlink oligonucleotide, comprising a region complementary to the 3' pre-bonding adapter and a region complementary to the backlink oligonucleotide. Includes, A back oligonucleotide dimer in which, during annealing, the end of the back lig-oligonucleotide is adjacent to the end of the 3' pre-binding adapter, enabling ligation of the back lig-oligonucleotide to the 3' pre-binding adapter at the ligation site. Includes, use.
23. (A) The forward link-oligonucleotide is (i) a template overhang region at the end of the forward link-oligonucleotide adjacent to the region complementary to the 5' pre-binding adapter, which is incomparable to the corresponding region of the post-associated single-stranded cDNA template, and / or (ii) A rig-oligonucleotide overhang region at the end of the forward link-oligonucleotide adjacent to the region complementary to the forward rig-oligonucleotide, the rig-oligonucleotide overhang region being incomplementary to the corresponding region of the forward rig-oligonucleotide, including and / or (B) The rear link-oligonucleotide is (i) a template overhang region at the end of the backlink-oligonucleotide adjacent to the region complementary to the 3' pre-binding adapter, which is incomparable to the corresponding region of the post-associated single-stranded cDNA template, and / or (ii) A rig-oligonucleotide overhang region at the end of the rear link-oligonucleotide adjacent to the region complementary to the rear rig-oligonucleotide, the rig-oligonucleotide overhang region being incomplementary to the corresponding region of the rear rig-oligonucleotide, The use according to claim 22, including the use described in claim 22.
24. A system or test kit for discovering disease biomarkers, (a) One or more test devices for normalizing a first cDNA sample from a subject with the disease and a second cDNA sample from a subject without the disease, and for sequencing the normalized first and second cDNA samples. (b) Processor, and (c) When executed by the aforementioned processor, (i) Access the determined sequences for the first and second cDNA samples of the one or more test devices, (ii) Calculate whether there are any cDNA sequences that are present in the first cDNA sample but not in the second cDNA sample, or that are present in the second cDNA sample but not in the first cDNA sample, and (iii) Output the result of step (ii) from the processor. Storage media containing computer applications configured in such a way A system or test kit that includes this.
25. A system or test kit for diagnosing a disease in a subject, (a) One or more test devices for normalizing cDNA samples from a subject and sequencing the normalized cDNA samples, (b) Processor, and (c) When executed by the aforementioned processor, (i) Access the determined sequence for the cDNA sample of one or more test devices, (ii) Calculate whether the cDNA sequence is present or absent, where the presence or absence of the cDNA sequence is associated with the disease, and (iii) Output from the processor whether the subject has the disease. Storage media containing computer applications configured in such a way A system or test kit that includes this.