Ligation method, method for inhibiting formation of linker dimer, preparation method of NGS library, NGS library preparation kit and 3 'linker
By using a 3' linker with a specific base sequence and an independent ligase, the problem of linker dimer formation in NGS library preparation was solved, thereby improving sequencing efficiency and simplifying the process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ARKRAY INC
- Filing Date
- 2024-10-07
- Publication Date
- 2026-05-15
AI Technical Summary
During NGS library preparation, the formation of adapter dimers leads to a decrease in sequencing efficiency, and existing methods are complex and difficult to improve efficiency easily.
A 5'-terminal adenylated single-stranded DNA with a specific base sequence was used as a 3' adapter, and a separate ligase was used to ligate the 3' and 5' adapters to inhibit the formation of adapter dimers.
It improves sequencing efficiency, simplifies the NGS library preparation process, reduces the formation of adapter dimers, and enhances data acquisition efficiency.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This disclosure relates to a ligation method, a method for inhibiting the formation of adapter dimers, a method for preparing NGS libraries, an NGS library preparation kit, and a 3' adapter. Background Technology
[0002] In recent years, nucleic acid network quantification analysis, represented by next-generation sequencing (NGS), has begun to be applied to disease diagnosis and other applications. In library preparation for NGS, firstly, base sequences called 5' adapters and 3' adapters are attached to the 5' and 3' ends of the target base sequence, respectively. If the nucleic acid being analyzed is RNA, the target base sequence with adapters attached to both ends is then converted into cDNA via reverse transcription. Then, using primers containing the sequences required for sequence analysis, the DNA is amplified by polymerase chain reaction (PCR) (see, for example, Non-Patent Literature 1).
[0003] In library preparation, ligation requires strict condition setting. The type of ligase and the sequence of the adapter nucleic acid affect the success or failure of the ligation reaction. For example, it is known that ligase activity can vary from batch to batch (see Non-Patent Literature 2). Furthermore, regarding the influence of the adapter nucleic acid sequence, it is known that randomizing its terminal sequence can reduce quantitative bias related to the RNA type of the assay (see, for example, Patent Literature 1 and Non-Patent Literature 3). More specifically, in Non-Patent Literature 3, a 3' adapter with a randomized 5' terminal base sequence and RNA are ligated in the presence of a high concentration of polyethylene glycol (PEG). Then, after the ligation, the remaining unligated 3' adapters are removed using a polyacrylamide gel, purifying the RNA with the attached 3' adapter. Then, the purified RNA with the attached 3' adapter is further ligated to a 5' adapter with a randomized 3' terminal base sequence in the presence of a high concentration of PEG.
[0004] Patent Document 1: US Patent No. 9,487,825
[0005] Non-patent literature 1: Seda Eminaga et al. Quantification of microRNA expression with next-generation sequencing. Curr Protoc Mol Biol. 2013 July: Chapter 4: Unit 4.17
[0006] Non-patent literature 2: Andreas Mayer; L Stirling Churchman, Genome-wide profiling of RNA polymerase transcription at nucleotide resolution in human cells with native elongating transcript sequencing. Nat Protoc. 2016 Apr;11(4):813-33
[0007] Non-patent document 3: Haedong Kim et al. Bias-minimized quantification of microRNA reveals widespread alternative processing and 3' end modification. Nucleic Acids Res. 2019 Mar 18;47(5):2630-2640 Summary of the Invention
[0008] The problem that the invention aims to solve
[0009] During the ligation process described above, a side reaction can occur where the remaining 5' adapters, which were not ligated to the target base sequence, bind to the remaining 3' adapters, forming adapter dimers. Furthermore, PCR of adapter dimers also occurs during PCR of the target base sequence, resulting in reads from adapter dimers in subsequent sequencing. This reduces sequencing efficiency (i.e., the number of reads per unit of data). In Non-Patent Literature 3, the remaining 3' adapters are removed, and the RNA with the 3' adapters is purified before ligation to the 5' adapters; however, a method that can more easily improve sequencing efficiency is needed.
[0010] In view of this situation, this disclosure relates to a ligation method that can more easily improve sequencing efficiency, a method for inhibiting adapter dimer formation, a method for preparing NGS libraries, an NGS library preparation kit, and a 3' adapter.
[0011] Methods for solving problems
[0012] The means to solve the above problems include the following methods.
[0013] <1> A connection method comprising: A 5'-terminal adenylated single-stranded DNA, serving as a 3' adapter, is ligated to the 3' end of each of several different types of nucleic acids. The 5'-terminal base sequence of this adenylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table A below; and After the ligation of the 5'-terminal adenylated single-stranded DNA, single-stranded nucleic acids serving as 5' adapters are ligated to the 5' ends of each of the aforementioned types of nucleic acids.
[0014] [Table 1]
[0015] <2> according to <1> The ligation method wherein the ligase used in the ligation of the 3' connector and the ligase used in the ligation of the 5' connector are each independently selected from at least one of the group consisting of T4 RNA ligase 1, T4 RNA ligase 2, T4 DNA ligase and their variants.
[0016] <3> according to <1> or <2> In the aforementioned connection method, the base sequence at the 5' end of the 3' connector is any one of the base sequences shown in Table B below.
[0017] [Table 2]
[0018] <4> according to <1> ~ <3> In any one of the ligation methods, the ligase used in the ligation of the 3' connector is T4 RNA ligase 1.
[0019] <5> according to <1> ~ <4> The ligation method described in any one of the above-mentioned multiple types of nucleic acids are small RNAs.
[0020] <6> A method for inhibiting the formation of linker dimers, comprising: A 5'-terminal adenylated single-stranded DNA, serving as a 3' adapter, is ligated to the 3' end of each of several different types of nucleic acids. The 5'-terminal base sequence of this adenylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table A below; and After the ligation of the 5'-terminal adenylated single-stranded DNA, single-stranded nucleic acids serving as 5' adapters are ligated to the 5' ends of each of the aforementioned types of nucleic acids.
[0021] [Table 3]
[0022] <7> according to <6> The method for inhibiting adapter dimer formation, wherein the ligase used in the ligation of the 3' adapter and the ligase used in the ligation of the 5' adapter are each independently selected from at least one of the group consisting of T4 RNA ligase 1, T4 RNA ligase 2, T4 DNA ligase and their variants.
[0023] <8> A method for preparing an NGS library, comprising: through <1> ~ <5> The ligation method described in any one of the following steps involves ligating a 5'-terminal adenylated single-stranded DNA as a 3' adapter and a single-stranded nucleic acid as a 5' adapter to multiple types of nucleic acids with different sequences; synthesizing cDNA from the multiple types of nucleic acids ligated with the aforementioned 3' adapter and the aforementioned 5' adapter via reverse transcription; and amplifying the obtained cDNA via PCR.
[0024] <9> according to <8> The method for preparing the NGS library, wherein the multiple types of nucleic acids with different sequences are multiple types of RNA with different sequences.
[0025] <10> An NGS library preparation kit comprising 5'-terminal adenylated single-stranded DNA contained in a first container, wherein the 5'-terminal base sequence of the 5'-terminal adenylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table A below.
[0026] [Table 4]
[0027] <11> according to <10> The NGS library preparation kit wherein the 5' end sequence of the single-stranded DNA, excluding the adenosine nucleotide at the 5' end, is any one of the base sequences shown in Table B below.
[0028] [Table 5]
[0029] <12> according to <10> or <11> The NGS library preparation kit further comprises T4 RNA ligase 1 housed in a second container.
[0030] <13> like <10> ~ <12> The NGS library preparation kit according to any one of the following further comprises, in a manner that is housed in separate containers, at least one selected from the group consisting of a 5' adapter, a ligase, a reverse transcriptase, a reverse transcription primer, a PCR enzyme, and a PCR primer.
[0031] <14> A 3' adapter used in the preparation of a nucleic acid library is a single-stranded DNA with its 5' end nucleotide adenosine-modified, wherein the 5' end base sequence of the single-stranded DNA, excluding the 5' end adenosine nucleotide, is any of the base sequences shown in Table A below.
[0032] [Table 6]
[0033] <15> according to <14> The 3' adapter used in the preparation of the nucleic acid library, wherein the 5' end base sequence of the single-stranded DNA, excluding the adenosine nucleotide at the 5' end, is any of the base sequences shown in Table B below.
[0034] [Table 7]
[0035] <16> A 3' adapter used in ligation is a single-stranded DNA with its 5' end nucleotide adenosine-modified, wherein the 5' end base sequence of the single-stranded DNA, excluding the 5' end adenosine nucleotide, is any of the base sequences shown in Table B below.
[0036] [Table 8]
[0037] Invention Effects
[0038] According to this disclosure, a ligation method that can more easily improve sequencing efficiency, a method for inhibiting adapter dimer formation, a method for preparing NGS libraries, an NGS library preparation kit, and a 3' adapter are provided. Attached Figure Description
[0039] Figure 1 This is a graph showing the correlation between different 5' connector types and the effective segment rate of each 3' connector in Example 1.
[0040] Figure 2 This indicates the correlation between the quantitative values (average of n=4) of miRNA when the 3' and 5' adapters were ligated using two different T4 RNA ligases 1 in Example 2.
[0041] Figure 3 The correlation of the quantitative values (average of n=4) of miRNA when only the 3' adapter was ligated using two T4 RNA ligases 1 in Example 2 is shown.
[0042] Figure 4 The correlation of the quantitative values (average of n=4) of miRNA when only the 5' connector was ligated using two T4 RNA ligases 1 in Example 2 is shown.
[0043] Figure 5 The correlation between the quantitative values (average of n=2) of miRNAs when the 5' end base sequence (excluding the 5' end adenosine nucleotide) of Example 3 was the 3' linker of TCCA and two T4 RNA ligases 1 were used.
[0044] Figure 6 The correlation between the quantitative values (average of n=2) of miRNAs when the 5' end base sequence (excluding the 5' end adenosine nucleotide) of Example 3 was GACA 3' linker and two T4 RNA ligases 1 were used. Detailed Implementation
[0045] The embodiments for carrying out this disclosure will now be described in detail. However, the embodiments of this disclosure are not limited to the following embodiments. In the following embodiments, unless specifically stated otherwise, the constituent elements (including element steps, etc.) are not essential. The same applies to numerical values and their ranges, which are not intended to limit the embodiments of this disclosure.
[0046] The term "process" in this disclosure includes not only processes that are independent of other processes, but also processes that achieve their purpose, even when they cannot be clearly distinguished from other processes.
[0047] In this disclosure, the numerical range represented by “~” is respectively contained within the numerical values recorded before and after “~” as the minimum and maximum values.
[0048] In the numerical ranges described in this disclosure in stages, the upper or lower limit value described in one numerical range can be replaced with the upper or lower limit value of other numerical ranges described in stages. Furthermore, the upper or lower limit value of the numerical ranges described in this disclosure can be replaced with the values shown in the embodiments.
[0049] Each component in this disclosure may comprise multiple corresponding substances. In the case where multiple substances equivalent to each component are present in the composition, unless otherwise stated, the content or percentage of each component refers to the total content or percentage of the multiple substances present in the composition.
[0050] Connection Methods
[0051] The ligation method disclosed herein includes: ligating a 5'-terminal adenosylated single-stranded DNA as a 3' adapter to the 3' ends of each of a plurality of different types of nucleic acids, wherein the 5'-terminal base sequence of the 5'-terminal adenosylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table A below; and after ligating the aforementioned 5'-terminal adenosylated single-stranded DNA, ligating a single-stranded nucleic acid as a 5' adapter to the 5' ends of each of the plurality of nucleic acids.
[0052] [Table 9]
[0053] In this disclosure, "5'-terminal adenylated single-stranded DNA serving as a 3' adapter (the 5'-terminal base sequence of the 5'-terminal adenylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table A)" is also simply referred to as "the 3' adapter of this disclosure." "Single-stranded nucleic acid serving as a 5' adapter" is also simply referred to as "5' adapter." "Multiple types of nucleic acids with different sequences" refers to the nucleic acid targeted for sequencing, and is also simply referred to as "multiple types of nucleic acids" or "target nucleic acid." The 5'-terminal base sequence of the 5'-terminal adenylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, refers to the 5'-terminal base sequence assuming the single-stranded DNA is not adenylated.
[0054] The 3' adapter disclosed herein readily ligates to the 3' ends of multiple types of nucleic acids with different sequences. The reason for this ease of ligation is not yet clear, but it is believed that the 3' adapter of this disclosure has high ligation efficiency for the 3' ends of the target nucleic acids, thereby improving sequencing efficiency. One reason for this is that, in the 3' adapter of this disclosure, the first base at the 5' end (excluding the adenosine nucleotide) is adenine, thymine, or guanine.
[0055] Furthermore, in the ligation method disclosed herein, after the ligation of the 3' connector, a single-stranded nucleic acid serving as a 5' connector is ligated to the 5' end of each of the target nucleic acids.
[0056] Here, according to the existing ligation method, as mentioned above, the following side reaction is likely to occur: the remaining 5' adapters that are not ligated to the nucleic acid that is the sequencing target bind to the remaining 3' adapters to form adapter dimers.
[0057] On the other hand, according to the ligation method of this disclosure, the side reaction of the remaining 5' adapter not ligated to the nucleic acid being sequenced binding with the remaining 3' adapter to form adapter dimers is less likely to occur. One reason for the less likely formation of adapter dimers is that the 3' adapter of this disclosure is less likely to ligate and / or hybridize with the 5' adapter. More specifically, when the sequence of the four bases starting from the 5' end of the 3' adapter (excluding the adenosine nucleotide at the 5' end) is any of the specific base sequences shown in Table A above, the 3' adapter is less likely to ligate and / or hybridize with the 5' adapter. One reason for this is that, in the 3' adapter of this disclosure, the first base at the 5' end (excluding the adenosine nucleotide at the 5' end) is adenine, thymine, or guanine.
[0058] As described above, the 3' adapter of this disclosure is easy to ligate to the target nucleic acid, and furthermore, the 3' adapter of this disclosure is unlikely to form an adapter dimer with the 5' adapter. Based on the above reasons, it is believed that the ligation method according to this disclosure can more easily improve sequencing efficiency.
[0059] It should be noted that this disclosure is not subject to any of the above-mentioned speculative mechanisms.
[0060] <3' Connector Connection>
[0061] The ligation method disclosed herein includes: ligating a 5'-terminal adenosylated single-stranded DNA, which serves as a 3' adapter, to the 3' end of each of a plurality of types of nucleic acids, wherein the 5'-terminal base sequence of the 5'-terminal adenosine nucleotide is any of the base sequences shown in Table A below.
[0062] [Table 10]
[0063] [3' connector]
[0064] The 3' adapter of the present invention is a 5'-terminal adenylated single-stranded DNA, wherein the 5'-terminal base sequence other than the 5'-terminal adenosine nucleotide is any of the base sequences shown in Table A.
[0065] In the 3' connector of this disclosure, the 5' end base sequence, excluding the adenosine nucleotide at the 5' end, is the base sequence numbered 1 to 78 in Table A above.
[0066] In the 3' adapter of this disclosure, from the viewpoint of improving sequencing efficiency, the 5' end base sequence, excluding the adenosine nucleotide at the 5' end, is preferably selected from any one of the base sequences numbered 1 to 53 shown in Table A above, and more preferably from any one of the base sequences numbered 1 to 3, 5 to 7, 10 to 16, 18, 22, 24, 26 to 29, 31 to 32, 34 to 38, 41 to 48, and 51 to 52 shown in Table A above. More preferably, it is selected from any one of the groups consisting of the base sequences numbered 1, 3, 5, 11, 12, 16, 18, 22, 26, 29, 32, 35, 37, 42, 44, 47 and 51 shown in Table A above, and particularly preferably, it is selected from any one of the groups consisting of the base sequences numbered 5, 11, 16, 26, 29, 35, 37 and 42 shown in Table A above.
[0067] It should be noted that the sequence numbers 1 to 78 in Table A above have no special meaning.
[0068] In one approach, when the 5' end base sequence, excluding the 5' end adenosine nucleotide, is used as a 3' linker for sequencing, the effective read rate is greater than 1.0%. This is achieved by using 5' end adenosine-modified single-stranded DNA selected from any of the base sequences numbered 1 to 78 shown in Table A above as the 3' linker.
[0069] In one approach, when the 5' end base sequence, excluding the 5' end adenosine nucleotide, is used as the 3' linker for sequencing, the effective read rate is greater than 1.5%. This is achieved by using 5' end adenosine-modified single-stranded DNA selected from any of the base sequences numbered 1 to 53 shown in Table A above.
[0070] In one approach, when the 5' end base sequence, excluding the 5' end adenosine nucleotide, is used as the 3' linker for sequencing, the effective read rate is greater than 2.0%. This is achieved by using a 5' end adenosine-modified single-stranded DNA selected from any of the following sequences: numbers 1–3, 5–7, 10–16, 18, 22, 24, 26–29, 31–32, 34–38, 41–48, and 51–52, as the 5' end adenosine-modified single-stranded DNA.
[0071] In one approach, when the 5' end base sequence, excluding the 5' end adenosine nucleotide, is used as the 3' linker for sequencing, the effective read rate is greater than 3.0% when the 5' end adenosine-modified single-stranded DNA is selected from any of the following sequences: numbered 1, 3, 5, 11, 12, 16, 18, 22, 26, 29, 32, 35, 37, 42, 44, 47, and 51 as shown in Table A above.
[0072] In one approach, when the 5' end base sequence, excluding the 5' end adenosine nucleotide, is used as the 3' linker for sequencing, the effective read rate is greater than 5.0%. This is achieved by using 5' end adenosine-enhanced single-stranded DNA selected from any of the base sequences numbered 5, 11, 16, 26, 29, 35, 37, and 42 shown in Table A above.
[0073] It should be noted that, in this disclosure, the effective read rate (also known as the proportion of effective reads) represents the efficiency of acquiring target data during sequencing. The effective read rate is calculated as the ratio (%) of the number of effective reads to the total read count. For example, the effective read rate in microRNA refers to the proportion of reads obtained through sequencing that are identical or substantially identical to sequences in the miRBase database.
[0074] From the perspective of improving sequencing efficiency, in the 3' adapter disclosed herein, the first base at the 5' end, excluding the adenosine nucleotide at the 5' end, is adenine, thymine, or guanine, preferably adenine or thymine.
[0075] From the perspective of improving sequencing efficiency, in the 3' adapter disclosed herein, the second base at the 5' end, other than the adenosine nucleotide at the 5' end, is preferably adenine, guanine, or cytosine.
[0076] In the 3' linker of this disclosure, from the viewpoint of being able to appropriately suppress the formation of linker dimers and batch-to-batch variability of T4 RNA ligase 1, the 5' end base sequence, excluding the 5' end adenosine nucleotide, is preferably any of the base sequences shown in Table B below (i.e., numbers 5, 6, 7, 10, 16, 44, 59, 71, and 75 in Table A).
[0077] [Table 11]
[0078] In this disclosure, the following ligation method will be referred to as Embodiment B. This ligation method includes: ligating a 5'-terminal adenosylated single-stranded DNA, serving as a 3' adapter, to the 3' ends of each of a plurality of different types of nucleic acids, wherein the 5'-terminal base sequence of the 5'-terminal adenosylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table B above; and after ligating the 5'-terminal adenosylated single-stranded DNA, ligating a single-stranded nucleic acid, serving as a 5' adapter, to the 5' ends of each of the plurality of nucleic acids. Furthermore, "the 5'-terminal adenosylated single-stranded DNA serving as a 3' adapter (the 5'-terminal base sequence of the 5'-terminal adenosylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table B)" will be abbreviated as "the 3' adapter of Embodiment B".
[0079] It should be noted that, in this disclosure, the description of Embodiment B, except that the 5' end base sequence of the 3' connector, excluding the adenosine nucleotide at the 5' end, is any of the base sequences shown in Table B above, is the same as the description in the embodiments of this disclosure, including definitions, examples, and preferred methods, unless otherwise stated.
[0080] In the continuous acquisition of large-scale data such as next-generation sequencing, batch effects, caused by differences in laboratories and experimenters, often become a problem. Batch effects not only hinder the discovery of new biomarkers but also make the long-term application of existing diagnostic assays difficult. One major cause of batch effects in next-generation sequencing is the batch-to-batch variation of reagents used in library preparation. For example, the T4 RNA ligase used in the ligation reaction is known to exhibit batch-to-batch variation. There are two subtypes of T4 RNA ligase, T4 RNA ligase 1 and T4 RNA ligase 2, as well as variants. From the viewpoint of ease of acquisition, T4 RNA ligase 1 is advantageous. However, T4 RNA ligase 1 is prone to batch effects due to batch-to-batch variations. However, according to Embodiment B, a ligation method has been discovered that is less prone to the formation of adapter dimers and less susceptible to the batch-to-batch variations of T4 RNA ligase 1.
[0081] As mentioned above, in the ligation of adapters during nucleic acid library preparation, if the remaining 5' adapter binds to the remaining 3' adapter, adapter dimers are formed. Furthermore, T4 RNA ligase 1 exhibits batch-to-batch variability, which can easily lead to errors, i.e., batch effects, when continuously acquiring large-scale data through next-generation sequencing and other methods. The inventors have discovered that, firstly, ligation of the 3' adapter is more likely to cause batch-to-batch variability in T4 RNA ligase 1 compared to ligation of the 5' adapter.
[0082] Furthermore, the inventors focused on the possibility that the 5' end base sequence of the 3' adapter might affect the ease of adapter dimer formation, and investigated the relationship between the sequence of the four 5' end bases of the 3' adapter (excluding the 5' end adenosine nucleotide) and the ease of adapter dimer formation. This is because if the amount of adapter dimer produced is high, the adapter dimer will also be amplified by PCR during NGS library preparation, leading to a decrease in sequencing efficiency. As a result, it was found that the sequence of these four bases is related to the ease of adapter dimer formation, and the 5' end base sequence of the 3' adapter that is less likely to form adapter dimers can be determined. Theoretically, the implementation of Embodiment B is not limited, but a 3' adapter with the determined base sequence may be easy to ligate to the target nucleic acid and / or difficult to ligate to or hybridize with the 5' adapter.
[0083] Next, considering that the 5' end base sequence of the 3' linker, excluding the adenosine nucleotide at the 5' end, might also affect batch-to-batch variability of T4 RNA ligase 1, the relationship between the sequence of the four 5' end bases and batch-to-batch variability of T4 RNA ligase 1 was investigated in 3' linkers that do not readily form linker dimers. The results showed that the sequence of these four bases was associated with batch-to-batch variability of T4 RNA ligase 1, and the 5' end base sequence of the 3' linker that could reduce batch-to-batch variability could be identified. Theoretically, there are no limitations to implementation method B, but the identified base sequence may have the same affinity for T4 RNA ligase and be independent of the batch of T4 RNA ligase, thereby making the ligation efficiency more uniform.
[0084] In the 3' linker of Embodiment B, from the viewpoint of being able to appropriately suppress the formation of linker dimers and batch-to-batch variability of T4 RNA ligase 1, the 5' end base sequence other than the 5' end adenosine nucleotide is preferably ACAA, AACA, AACT, AGGA, and AACG (i.e., numbers 5, 6, 7, 10, and 16 in Table A), more preferably ACAA, AACA, AACT, and AACG (i.e., numbers 5, 6, 7, and 10 in Table A), and even more preferably AACA, AACT, and AACG (i.e., numbers 5, 6, and 7 in Table A).
[0085] In one embodiment, the effective number of reads during adapter ligation to multiple nucleic acids, PCR, and sequencing using the 3' adapter of embodiment B is preferably 1.0% or more, more preferably 1.5% or more, further preferably 2.0% or more, particularly preferably 3.0% or more, and extremely preferably 5.0% or more.
[0086] In the 3' linker of this disclosure, the bases after the fifth base of the 5'-terminal base sequence (excluding the 5'-terminal adenosine nucleotide) can be any base sequence. The bases after the fifth base of the 5'-terminal base sequence (excluding the 5'-terminal adenosine nucleotide) may or may not contain modifying bases. Examples of modifying bases include, but are not limited to, 4-acetylcytidine, dihydrouridine, inosine, and 1-methyladenosine.
[0087] From the viewpoint of facilitating the connection between the adapter and the target nucleic acid, the 3' adapter of this disclosure preferably has a base length of 6 to 1000 bases excluding the adenosine nucleotide at the 5' end, more preferably 6 to 100 bases, and even more preferably 10 to 30 bases.
[0088] The nucleotide at the 3' end of the 3' linker of this disclosure is preferably modified with a compound having a modified structure capable of inhibiting nucleic acid binding.
[0089] Examples of compounds with modified structures capable of inhibiting nucleic acid binding include dideoxynucleoside triphosphates (ddNTPs) and compounds containing amino groups. Examples of dideoxynucleoside triphosphates include dideoxycytidine triphosphate (ddCTP), dideoxyguanosine triphosphate (ddGTP), dideoxythymidine triphosphate (ddTTP), and dideoxyadenosine triphosphate (ddATP). Among compounds with modified structures capable of inhibiting nucleic acid binding, dideoxynucleoside triphosphates are preferred, and dideoxycytidine triphosphates are more preferred.
[0090] [Multiple types of nucleic acids with different sequences]
[0091] In this disclosure, "multiple types of nucleic acids with different sequences" means that there are multiple types of nucleic acids with different base sequences. The morphology of these multiple types of nucleic acids with different sequences is not particularly limited.
[0092] From the viewpoint of easy connection with the 3' connector of this disclosure, the aforementioned nucleic acid is preferably single-stranded. In this disclosure, the aforementioned nucleic acid can be, for example, DNA, RNA, analogs thereof, or fusions thereof. The nucleic acid can be small RNA (sRNA). Small RNA can be any of microRNA (miRNA), piRNA, and tsRNA. Furthermore, the nucleic acid can be composed only of A, T, G, C, and U bases, or it can be a nucleic acid not composed only of A, T, G, C, and U bases. Specifically, examples of nucleic acids not composed only of A, T, G, C, and U bases include, for example, nucleic acids that have undergone modifications such as DNA or RNA methylation or A-to-I RNA editing.
[0093] The source of the aforementioned nucleic acids is not particularly limited; they can be naturally derived nucleic acids (e.g., total RNA) or synthetic RNA. The source of nucleic acids can be, for example, biological samples, viral samples, environmental samples, or artificial samples.
[0094] As biologically derived samples, examples include those derived from animals, plants, fungi, or bacteria. Animals include humans and non-human animals. Non-human animals include non-human mammals (e.g., monkeys, dogs, cats, mice, rats, rabbits, cattle, horses, pigs, and sheep) and birds (e.g., chickens and quails).
[0095] Animal-derived samples, more specifically, can include serum, plasma, whole blood, urine, feces, saliva, bone marrow fluid, lymph, and tissues.
[0096] As environmental samples, examples include samples from soil or water.
[0097] Examples of artificial samples include artificial serum, artificial plasma, and artificial urine.
[0098] The aforementioned nucleic acids can be 10,000 to 300,000 types, 10,000 to 10,000 types, or 1,000 to 2,000 types.
[0099] The length of the nucleic acid is not particularly limited, but is preferably 1 to 100 bases, more preferably 10 to 50 bases, and even more preferably 16 to 40 bases.
[0100] [Connection Method]
[0101] Regarding the ligation method of this disclosure, in which the 3' adapter of this disclosure is ligated to the 3' ends of various types of nucleic acids, the detailed conditions can be performed according to conventional methods. For example, there are no particular limitations on the concentration, temperature, and time of each component when ligating the 3' adapter of this disclosure to the aforementioned nucleic acids.
[0102] The preferred reaction temperature for ligating the 3' adapter of this disclosure to the above-mentioned nucleic acid is 10°C to 40°C, more preferably 15°C to 30°C, further preferably 17°C to 25°C, and particularly preferably 20°C. The preferred reaction time is 10 minutes to 2 hours, more preferably 30 minutes to 90 minutes, further preferably 45 minutes to 75 minutes, and particularly preferably 1 hour.
[0103] One or more 3' adapters of this disclosure can be ligated to the 3' end of multiple types of nucleic acids.
[0104] (ligase)
[0105] In the ligation method of this disclosure, when ligating the 3' adapter of this disclosure to the 3' ends of various types of nucleic acids, the enzyme used for ligation (also called ligase) is not particularly limited, and known enzymes can be used. Only one type of ligase may be used, or multiple types may be used.
[0106] The ligase used for ligation of the 3' linker is preferably selected from at least one of the group consisting of T4 RNA ligase 1, T4 RNA ligase 2, T4 DNA ligase and their variants, and more preferably T4 RNA ligase 2 or its variants from the viewpoint of improving sequencing efficiency.
[0107] In embodiment B, the ligase used for ligation of the 3' linker is preferably at least one selected from the group consisting of T4 RNA ligase 1, T4 RNA ligase 2, T4 DNA ligase and their variants, and more preferably T4 RNA ligase 1 from the viewpoint of being able to appropriately suppress the formation of linker dimers and batch-to-batch variability of T4 RNA ligase 1.
[0108] Variants of T4 RNA ligase 1, T4 RNA ligase 2, or T4 DNA ligase refer to variants that possess a certain degree of sequence identity relative to the wild-type T4 RNA ligase 1, T4 RNA ligase 2, or T4 DNA ligase from which they originate—for example, more than 50%, 60%, 70%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity—and exhibit the same nucleic acid-adaptor ligation activity as the wild-type T4 RNA ligase 1, T4 RNA ligase 2, or T4 DNA ligase. It should be noted that the sequence identity here refers to the sequence identity when comparing the full lengths of the ligases.
[0109] Here, well-known methods can be cited as approaches to calculating the identity of amino acid sequences. For example, commercially available analytical tools or analytical tools accessible via electrical communication lines (the Internet) can be used. As an example, the identity of amino acid sequences can be calculated using the BLAST (Basic Local Alignment Search Tool) homology algorithm from the National Center for Biotechnology Information (NCBI) (http: / / www.ncbi.nlm.nih.gov / BLAST / ) using default (initial) parameters.
[0110] Additionally, variants of T4 RNA ligase 1, T4 RNA ligase 2, or T4 DNA ligase can also be proteins composed of amino acid sequences consisting of substitutions, deletions, insertions, and / or additions of one or more amino acids in the amino acid sequence of wild-type T4 RNA ligase 1, T4 RNA ligase 2, or T4 DNA ligase from which they originate, and possessing the same nucleic acid-linker ligation activity as wild-type T4 RNA ligase 1, T4 RNA ligase 2, or T4 DNA ligase. Here, one or more can be, for example, 1 to 100, 1 to 80, more preferably 1 to 40, further preferably 1 to 10, particularly preferably 1 to 5, extremely preferably 1 to 3, but not particularly limited.
[0111] Variants of T4 RNA ligase 2 may include those containing an amino acid sequence having 85% or more, 90% or more, 95% or more, 98% or more, or 99% or more identity with the amino acid sequence from the N-terminus of the wild-type T4 RNA ligase 2. Variants of T4 RNA ligase 2 may also include those containing the amino acid sequence from the N-terminus of the wild-type T4 RNA ligase 2, or those containing 1 to 10, 1 to 5, or 1 to 3 amino acids replaced, deleted, inserted, and / or added relative to the amino acid sequence from the N-terminus of the wild-type T4 RNA ligase 2. Examples of variants of T4 RNA ligase 2 include, for example, truncated T4 RNA ligase 2, truncated T4 RNA ligase 2 KQ, and truncated T4 RNA ligase 2 K227Q, with truncated T4 RNA ligase 2 or truncated T4 RNA ligase 2 KQ being preferred.
[0112] T4 RNA ligase 2 truncated variants are variants of T4 RNA ligase 2 with a C-terminal deletion, typically having the amino acid sequence from position 1 to position 249 of the N-terminus of T4 RNA ligase 2.
[0113] T4 RNA ligase 2 truncated variant K227Q is a variant in which the K (lysine) at position 227 of the amino acid sequence of T4 RNA ligase 2 is changed to Q (glutamine).
[0114] T4 RNA ligase 2 truncated KQ is a variant (two-point mutant) that further modifies the amino acid sequence of T4 RNA ligase 2 truncated K227Q by changing R (arginine) at position 55 to K (lysine).
[0115] Variants of T4 DNA ligase include, for example, Salt-T4 DNA ligase.
[0116] <5' Connector Connection>
[0117] The ligation method disclosed herein includes: after the ligation of the above-mentioned 3' adapter, ligating a single-stranded nucleic acid serving as a 5' adapter to the 5' end of each of the above-mentioned multiple types of nucleic acids.
[0118] [5' connector]
[0119] The 5' adapter is a single-stranded nucleic acid. The 5' adapter only needs to be a single-stranded nucleic acid; other characteristics are not particularly limited. The base length of the single-stranded nucleic acid used as the 5' adapter is also not particularly limited, but is preferably 5–100 bases long, more preferably 15–50 bases long, and even more preferably 20–40 bases long.
[0120] The nucleic acid in the single-stranded nucleic acid that serves as the 5' linker can be DNA, RNA, their analogues, or their fusions. That is, the 5' linker can be single-stranded DNA or single-stranded RNA. The nucleic acid in the single-stranded nucleic acid can be composed only of A, T, G, C, and U bases, or it can be composed of more than just A, T, G, C, and U bases. Specifically, examples of nucleic acids not composed of A, T, G, C, and U bases include, for example, nucleic acids that have undergone modifications such as DNA or RNA methylation or A-to-I RNA editing.
[0121] According to the ligation method disclosed herein, sequencing efficiency can be improved more easily regardless of the base sequence of the 5' adapter. Therefore, the base sequence of the 5' adapter is not particularly limited in this disclosure.
[0122] According to the ligation method of Embodiment B, it is believed that regardless of the base sequence of the 5' linker, the formation of linker dimers and batch-to-batch variability of T4 RNA ligase 1 can be suppressed. Therefore, in Embodiment B, the base sequence of the 5' linker is not particularly limited.
[0123] The 5' linker can be a base sequence that has more than 80% sequence identity with respect to the base sequences shown in Serial Numbers 1 to 6 below. The sequence identity can be more than 85%, more than 90%, more than 95%, or 100% (i.e., the base sequences are completely identical).
[0124] (Serial Number 1) 5'-GUUCAGAGUUCUACAGUCCGACGAUCGGAG-3'
[0125] (Serial Number 2) 5'-GUUCAGAGUUCUACAGUCCGACGAUCCACG-3'
[0126] (Serial Number 3) 5'-GUUCAGAGUUCUACAGUCCGACGAUCUUGU-3'
[0127] (Serial Number 4) 5'-GUUCAGAGUUCUACAGUCCGACGAUC-3'
[0128] (Serial Number 5) 5'-GUUCAGAGUUCUACAGUCCGACGAUCUCAU-3'
[0129] (Serial Number 6) 5'-GUUCAGAGUUCUACAGUCCGACGAUCUGCA-3'
[0130] Here, the method for calculating the identity of base sequences can be performed using well-known methods. For example, commercially available analytical tools or analytical tools accessible via electrical communication lines (the Internet) can be used. As an example, the identity of base sequences can be calculated using the BLAST (Basic Local Alignment Search Tool) homology algorithm from the National Center for Biotechnology Information (NCBI) (http: / / www.ncbi.nlm.nih.gov / BLAST / ) using default (initial settings) parameters.
[0131] The 5' adapter sequence can be a sequence consisting of one or more bases substituted, deleted, inserted, and / or added from the base sequences shown in Serial Numbers 1 to 6, and can be linked to nucleic acids in the same way as the 5' adapters of the base sequences shown in Serial Numbers 1 to 6. Here, one or more bases can be, for example, 1 to 80, preferably 1 to 40, more preferably 1 to 10, further preferably 1 to 5, even more preferably 1 to 3, particularly preferably 1 or 2, but there is no particular limitation.
[0132] [Multiple types of nucleic acids with different sequences]
[0133] The multiple types of nucleic acids used for ligating the 5' connector are nucleic acids for which a 3' connector is attached to the 3' end of each of the multiple types of nucleic acids that serve as the target nucleic acid. Primers may also be further attached to the multiple types of nucleic acids used for ligating the 5' connector. For example, if the multiple types of nucleic acids serving as the target nucleic acid are RNA, the multiple types of nucleic acids used for ligating the 5' connector may also be nucleic acids that have been annealed with reverse transcription primers attached to the multiple types of nucleic acids with 3' connectors.
[0134] The various types of nucleic acids obtained after the ligation of the 3' adapter described above can be purified after the ligation operation. Purification can be performed using organic solvents or solid-phase methods (e.g., silica membranes, ion exchange columns, or magnetic beads). It should be noted that the ligation method according to this disclosure can suppress the formation of adapter dimers, therefore purification may not be necessary.
[0135] [Connection Method]
[0136] Regarding the ligation method of this disclosure, the detailed conditions for ligating 5' adapters to the 5' ends of various types of nucleic acids can be performed according to conventional methods. For example, there are no particular limitations on the concentration, temperature, and time of the components when ligating 5' adapters to the aforementioned nucleic acids.
[0137] The preferred reaction temperature for ligating the 5' adapter to the aforementioned nucleic acid is 10°C to 40°C, more preferably 15°C to 30°C, further preferably 17°C to 25°C, and particularly preferably 20°C. The preferred reaction time is 10 minutes to 2 hours, more preferably 30 minutes to 90 minutes, further preferably 45 minutes to 75 minutes, and particularly preferably 1 hour.
[0138] One or more 5' adapters can be ligated to the 5' end of various types of nucleic acids.
[0139] (ligase)
[0140] In the ligation method disclosed herein, when ligating 5' adapters to the 5' ends of various types of nucleic acids, the enzyme used for ligation (also called ligase) is not particularly limited, and known enzymes can be used. Only one type of ligase may be used, or multiple types may be used.
[0141] The ligase used for ligation of the 5' linker is preferably at least one selected from the group consisting of T4 RNA ligase 1, T4 RNA ligase 2, T4 DNA ligase and their variants, more preferably T4 RNA ligase 1 or its variants.
[0142] The descriptions of the definitions, examples, and preferred methods of variants of T4 RNA ligase 1, T4 RNA ligase 2, and T4 DNA ligase are the same as those described in <3' Ligation of Connectors> above.
[0143] It should be noted that the ligases used in the ligation of the 3' and 5' connectors can be the same or different. In Embodiment B, from the viewpoint of simplicity, it is preferable to use the same ligase for both the 3' and 5' connectors. From the viewpoint that the ligation method of Embodiment B is particularly useful, it is more preferable to use T4 RNA ligase 1 for both.
[0144] Methods for inhibiting the formation of linker dimers
[0145] The method for inhibiting adapter dimer formation disclosed herein includes: ligating a 5'-terminal adenosylated single-stranded DNA as a 3' adapter to the 3' end of each of a plurality of nucleic acids with different sequences, wherein the 5'-terminal base sequence of the 5'-terminal adenosylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table A above; and after ligating the 5'-terminal adenosylated single-stranded DNA, ligating a single-stranded nucleic acid as a 5' adapter to the 5' end of each of the plurality of nucleic acids.
[0146] The definitions, examples, and preferred methods of the connection of the 3' connector and the 5' connector in the method for suppressing the formation of connector dimers disclosed herein are the same as those described in the above-mentioned <Connection of 3' Connectors> or <Connection of 5' Connectors>.
[0147] In the method for suppressing the formation of linker dimers disclosed herein, as described above, the base sequence at the 5' end of the 3' linker of the present disclosure is preferably selected from any one of the group consisting of base sequences numbered 1 to 53 shown in Table A.
[0148] In the method for inhibiting adapter dimer formation disclosed herein, as described above, the ligase used for ligation of the 3' adapter and the ligase used for ligation of the 5' adapter are preferably each independently selected from at least one of the group consisting of T4 RNA ligase 1, T4 RNA ligase 2, T4 DNA ligase and their variants.
[0149] Methods for preparing NGS libraries
[0150] The method for preparing the NGS library disclosed herein includes: ligating the 5'-terminal adenylated single-stranded DNA and the single-stranded nucleic acid as 5' adapters of the present disclosure to multiple types of nucleic acids with different sequences using the ligation method of the present disclosure; synthesizing cDNA from the multiple types of nucleic acids ligated with the above-mentioned 3' adapters and the above-mentioned 5' adapters by reverse transcription; and amplifying the obtained cDNA by PCR.
[0151] The method for preparing the NGS library of Embodiment B includes: ligating a 5'-terminal adenylated single-stranded DNA as the 3'-adapter of Embodiment B and a single-stranded nucleic acid as the 5'-adapter to multiple types of RNAs with different sequences by the ligation method of Embodiment B; synthesizing cDNA from the multiple types of RNAs ligated with the above 3'-adapter and the above 5'-adapter by reverse transcription reaction; and amplifying the obtained cDNA by PCR.
[0152] <Ligation>
[0153] The description of ligating the 5'-terminal adenylated single-stranded DNA as the 3'-adapter and the single-stranded nucleic acid as the 5'-adapter of the present disclosure to multiple types of nucleic acids by the ligation method of the present disclosure in the method for preparing the NGS library of the present disclosure is as described above.
[0154] <Synthesis of cDNA>
[0155] The method for preparing the NGS library of the present disclosure includes: synthesizing cDNA from multiple types of nucleic acids ligated with the above 3'-adapter and the above 5'-adapter by reverse transcription reaction.
[0156] In one mode, the reverse transcription reaction includes annealing the reverse transcription primer to the multiple types of nucleic acids ligated with the above 3'-adapter and the above 5'-adapter. Alternatively, the reverse transcription primer can also be annealed after the ligation of the 3'-adapter and before the ligation of the 5'-primer. The reverse transcription primer can have a sequence for further attaching an inherent sequence to each of the multiple types of nucleic acids. Examples of the inherent sequence include UMI (Unique Molecular Identifier) sequence, etc. By attaching the inherent sequence of each of the multiple types of nucleic acids to the cDNA, even if the cDNA is amplified by PCR as described later, the amount of each nucleic acid present before amplification can be determined.
[0157] The method for the reverse transcription reaction for synthesizing cDNA is not particularly limited. For example, the reverse transcriptase and the conditions of the reverse transcription reaction can also use known means. For example, the reverse transcription reaction can be carried out by reverse transcription of the multiple types of nucleic acids ligated with the 3'-adapter and 5'-adapter of the present disclosure in the presence of a reverse transcriptase, a primer, a dNTP mixture, and an appropriate buffer.
[0158] <Amplification of cDNA>
[0159] The method for preparing the NGS library of the present disclosure includes: amplifying the obtained cDNA by PCR. The method of PCR is not particularly limited. For example, the PCR enzyme and the conditions of PCR can also use known means.
[0160] <Order of each method>
[0161] The preferred method for preparing the NGS library disclosed herein includes, in sequence: ligating the 5'-terminal adenylated single-stranded DNA (as a 3' adapter) and the single-stranded nucleic acid (as a 5' adapter) of the present disclosure to multiple types of nucleic acids using the ligation method of the present disclosure; synthesizing cDNA from the multiple types of nucleic acids ligated with the above-mentioned 3' adapter and the above-mentioned 5' adapter by reverse transcription; and amplifying the obtained cDNA by PCR.
[0162] When ligating the 3' adapter of this disclosure to multiple types of nucleic acids and then ligating the 5' adapter, "ligating a single-stranded nucleic acid as a 5' adapter to the 5' end of each of the aforementioned multiple types of nucleic acids" means that a single-stranded nucleic acid as a 5' adapter is ligated to the 5' end of each of the multiple types of nucleic acids to which the 3' adapter of this disclosure is ligated.
[0163] On the other hand, when the 3' adapter of this disclosure is ligated after ligating the 5' adapter to multiple types of nucleic acids, "ligating the 5'-terminal adenosylated single-stranded DNA as a 3' adapter to the 3' end of each of the multiple types of nucleic acids" means ligating the 5'-terminal adenosylated single-stranded DNA as a 3' adapter of this disclosure to the 3' end of each of the multiple types of nucleic acids to which the 5' adapter is ligated.
[0164] In this disclosure, the description of the method for preparing the NGS library of Embodiment B is the same as the description of the method for preparing the NGS library of this disclosure, except that the ligation method of Embodiment B and the multiple types of nucleic acids are RNA, unless otherwise stated.
[0165] NGS Library Preparation Kit
[0166] The NGS library preparation kit disclosed herein contains 5'-terminal adenosine-coated single-stranded DNA housed in a first container, wherein the 5'-terminal base sequence of the 5'-terminal adenosine nucleotide, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table A below.
[0167] [Table 12]
[0168] In the NGS library preparation kit of the present invention, from the viewpoint of improving sequencing efficiency, the 5' end base sequence of the above-mentioned 5' end adenylated single-stranded DNA is preferably selected from any one of the group consisting of the base sequences numbered 1 to 53 shown in Table A above.
[0169] The NGS library preparation kit of embodiment B contains 5'-terminal adenylated single-stranded DNA housed in a first container, wherein the 5'-terminal base sequence of the 5'-terminal adenylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table B below.
[0170] [Table 13]
[0171] In the NGS library preparation kit of Embodiment B, from the viewpoint of being able to appropriately suppress the formation of adapter dimers and batch-to-batch variability of T4 RNA ligase 1, the 5' end base sequence of the above-mentioned 5' adenylated single-stranded DNA is preferably ACAA, AACA, AACT, AGGA and AACG (i.e., numbers 5, 6, 7, 10 and 16 in Table A), more preferably ACAA, AACA, AACT and AACG (i.e., numbers 5, 6, 7 and 10 in Table A), and even more preferably AACA, AACT and AACG (i.e., numbers 5, 6 and 7 in Table A).
[0172] The description of the 3' adapter in the NGS library preparation kit disclosed herein, including definitions, examples, and preferred methods, is the same as the description of the "3' adapter" described in the "Ligation Methods" section.
[0173] The container described above may also contain a buffer solution for stably preserving the 3' connector of this disclosure. Examples of components contained in the buffer solution include, for instance, Tris hydrochloric acid (Tris-HCl), Triton X (registered trademark), potassium chloride, sodium chloride, dithiothreitol (DTT), ethylenediaminetetraacetic acid (EDTA), bovine serum albumin (BSA), and glycerol.
[0174] <Other Ingredients>
[0175] The kit disclosed herein may further comprise at least one selected from the group consisting of a 5' adapter, a ligase, a reverse transcriptase, a reverse transcription primer, a PCR enzyme, and PCR primers, respectively, and housed in separate containers. From the viewpoint of suppressing batch-to-batch variability of T4 RNA ligase 1, the kit of embodiment B may further comprise T4 RNA ligase 1 housed in a second container. As a ligase, T4 RNA ligase 1 and a different ligase may also be included, respectively, in separate containers.
[0176] In addition, the kit disclosed herein may include an instruction manual.
[0177] The descriptions of the 5' adapter and ligase, including definitions, examples, and preferred methods, are identical to those for the "5' adapter" and "ligase" in the "Ligation Methods" section.
[0178] The reverse transcriptase, reverse transcription primers, PCR enzymes, and PCR primers can be well-known substances.
[0179] 3' adapters used in the preparation of nucleic acid libraries
[0180] The 3' adapter used in the preparation of the nucleic acid library disclosed herein is a single-stranded DNA with its 5' end nucleotide adenosine-modified. The 5' end base sequence of the single-stranded DNA, excluding the 5' end adenosine nucleotide, is any of the base sequences shown in Table A below.
[0181] [Table 14]
[0182] In the preparation of the nucleic acid library disclosed herein, the 3' adapter has a 5' end base sequence other than the 5' end adenosine nucleotide, which is the base sequence numbered 1 to 78 in Table A above.
[0183] In the 3' adapter of this disclosure, from the viewpoint of improving sequencing efficiency, the 5' end base sequence, excluding the adenosine nucleotide at the 5' end, is preferably selected from any one of the base sequences numbered 1 to 53 shown in Table A above, and more preferably from any one of the base sequences numbered 1 to 3, 5 to 7, 10 to 16, 18, 22, 24, 26 to 29, 31 to 32, 34 to 38, 41 to 48, and 51 to 52 shown in Table A above. More preferably, it is selected from any one of the groups consisting of the base sequences numbered 1, 3, 5, 11, 12, 16, 18, 22, 26, 29, 32, 35, 37, 42, 44, 47 and 51 shown in Table A above. Even more preferably, it is selected from any one of the groups consisting of the base sequences numbered 5, 11, 16, 26, 29, 35, 37 and 42 shown in Table A above.
[0184] The 3' adapter used in the preparation of the nucleic acid library in Implementation Method B is a single-stranded DNA with its 5' end nucleotide adenosine-modified. The 5' end base sequence of the single-stranded DNA, excluding the 5' end adenosine nucleotide, is any of the base sequences shown in Table B below.
[0185] [Table 15]
[0186] In the preparation of the nucleic acid library of Embodiment B, the 3' adapter used, from the viewpoint of being able to appropriately suppress the formation of adapter dimers and batch-to-batch variability of T4 RNA ligase 1, preferably has the following 5' base sequences other than the 5' adenosine nucleotide: ACAA, AACA, AACT, AGGA, and AACG (i.e., numbers 5, 6, 7, 10, and 16 in Table A), more preferably ACAA, AACA, AACT, and AACG (i.e., numbers 5, 6, 7, and 10 in Table A), and even more preferably AACA, AACT, and AACG (i.e., numbers 5, 6, and 7 in Table A).
[0187] In the preparation of the nucleic acid library in Embodiment B, the 3' adapter is detailed as described in the 3' adapter section of Embodiment B above.
[0188] 3' connector used in the connection
[0189] The 3' linker used in embodiment B is a single-stranded DNA with its 5' end nucleotides adenosine-modified. The 5' end base sequence of the single-stranded DNA, excluding the 5' end adenosine nucleotide, is any of the base sequences shown in Table B below.
[0190] [Table 16]
[0191] In the 3' linker used in embodiment B, from the viewpoint of being able to appropriately suppress the formation of linker dimers and batch-to-batch variability of T4 RNA ligase 1, the 5' end base sequence other than the 5' end adenosine nucleotide is preferably ACAA, AACA, AACT, AGGA, and AACG (i.e., numbers 5, 6, 7, 10, and 16 in Table A), more preferably ACAA, AACA, AACT, and AACG (i.e., numbers 5, 6, 7, and 10 in Table A), and even more preferably AACA, AACT, and AACG (i.e., numbers 5, 6, and 7 in Table A).
[0192] In the connection used in Embodiment B, the details of the 3' connector are as described in the 3' connector of Embodiment B above.
[0193] Example
[0194] Next, the embodiments of this disclosure will be described in detail through examples, but the embodiments of this disclosure are not limited to these examples.
[0195] Example 1: Study on the base sequence of the 3' linker used to suppress linker dimer formation and obtain sufficient effective reads.
[0196] In this embodiment, 3' adapters were screened to suppress adapter dimer formation and obtain a sufficient number of effective reads. This is because if a large amount of adapter dimers are produced, they will also be amplified by PCR during library preparation, leading to a decrease in sequencing efficiency.
[0197] <Preparation of the 3' Connector>
[0198] The 3' adapter disclosed herein is linked to multiple types of nucleic acids at the 5' end of the 3' adapter, therefore it is believed that the sequence at the 5' end of the 3' adapter affects the performance of the 3' adapter.
[0199] Therefore, 5'-terminal adenylated single-stranded DNA (hereinafter also referred to as 5'-terminal randomized DNA. The base sequence is represented by sequence number 7) was purchased, with the 5'-terminal base sequence randomized from positions 1 to 4, excluding the 5'-terminal adenosine nucleotide. It should be noted that the 3' end of the above-mentioned 5'-terminal randomized DNA was capped by dideoxycytidine modification.
[0200] (Serial Number 7) 5'-App / NNNNGTAGGCACCATCAAT / 3'ddC
[0201] In sequence number 7, N independently represents A, T, G, or C. That is, we get 256 (=4). 4 ) types of 5' end randomized DNA.
[0202] <Preparation of 5' Connector>
[0203] The 5' adapter connects to multiple types of nucleic acids at the 3' end. The following six base sequences (Sequence Nos. 1 to 6) were purchased as the 5' adapter.
[0204] (Serial Number 1) 5'-GUUCAGAGUUCUACAGUCCGACGAUCGGAG-3'
[0205] (Serial Number 2) 5'-GUUCAGAGUUCUACAGUCCGACGAUCCACG-3'
[0206] (Serial Number 3) 5'-GUUCAGAGUUCUACAGUCCGACGAUCUUGU-3'
[0207] (Serial Number 4) 5'-GUUCAGAGUUCUACAGUCCGACGAUC-3'
[0208] (Serial Number 5) 5'-GUUCAGAGUUCUACAGUCCGACGAUCUCAU-3'
[0209] (SEQ ID NO: 6) 5’-GUUCAGAGUUCUACAGUCCGACGAUCUGCA-3’
[0210] <Preparation of NGS Library>
[0211] Using the above 3’ adapter and 5’ adapter respectively, prepare the NGS library of small RNAs (in this example, human-derived serum RNA is used). The method for preparing the NGS library is carried out according to Quantification of microRNA Expression with Next-Generation Sequencing by Seda Eminaga et al., Curr Protoc Mol Biol. 2013 July; 04: Unit 4.17. That is, perform the operations of [1] ligation of 3’ adapter, [2] annealing of reverse transcription primer, [3] ligation of 5’ adapter, [4] reverse transcription, [5] PCR, and [6] purification of PCR products.
[0212] [1] Ligation of 3’ Adapter
[0213] Nuclease-free water 1.0 μL
[0214] T4 RNA Ligase Reaction Buffer (NEB / M0242) 1.0 μL
[0215] 100% DMSO (13445-74 / Nacalai Tesque) 1.0 μL
[0216] 50% PEG (NEB / B1004) 1.0 μL
[0217] Total RNA (human serum-derived RNA) 4.0 μL
[0218] Add the 3’ adapter of SEQ ID NO: 7 at 10 μM to the above mixture and incubate at 90 °C for 30 seconds.
[0219] T4 RNA Ligase 2 truncated form (NEB / M0242) 1.5 μL
[0220] SUPERase Inhibitor (Thermo Fisher / AM2694) 0.5 μL
[0221] Furthermore, add the above mixture and incubate at 22 °C for 60 minutes to ligate the total RNA with the 3’ adapter. / / There was a line break here in the original text, but it's not clear if it's intentional or a formatting error. I've kept it as it is for now.
[0222] [2] Annealing of RT Primer
[0223] For the solution obtained by the above [1] operation, add 1 μL of 10 μM RT primer, incubate at 90 °C for 30 seconds, and then incubate at 65 °C for 5 minutes.
[0224] The base sequence of the RT primer is represented by sequence number 8 below. The part represented by N is the UMI sequence. It should be noted that in the following sequences, N independently represents A, T, G, or C.
[0225] (Serial No. 8)5'-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNNNNATTGATGGTGCCTAC-3'
[0226] [3] Connection of 5' connector
[0227] 0.8 μL of nuclease-free water
[0228] ATP (NEB / 0204M) 1.4 μL
[0229] T4 RNA ligase 1 (NEB / 0204M) 1.5 μL
[0230] For the solution obtained by the above [2] operation, 10 μM of 5' adapter as any of the base sequences in sequence number 1 to sequence number 6 is added to the above mixture and incubated at 20°C for 60 minutes.
[0231] [4] Reverse transcription
[0232] 3.5 μL of nuclease-free water
[0233] 0.3 μL of 25 mM dNTP mixture (Nippon Gene / 312-07271)
[0234] 5×FS buffer (Thermo Fisher / 18080093) 1.5μL
[0235] 0.1M DTT (Thermo Fisher / 18080093) 1.5μL
[0236] RNase OUT(Thermo Fisher / 10777019) 0.375μL
[0237] SuperScript III(Thermo Fisher / 18080093) 1.5μL
[0238] Then, 7.5 μL of the solution obtained from the above incubation was added to the above mixture, incubated at 48°C for 30 seconds, and further incubated at 85°C for 5 minutes.
[0239] It should be noted that the base sequence of the RT primer is shown in Serial No. 8 above.
[0240] [5]PCR
[0241] 24.1 μL of nuclease-free water
[0242] Phusion HF buffer (NEB / M0530L) 10μL
[0243] 0.4 μL of 25 mM dNTP mixture (Nippon Gene / 312-07271)
[0244] Phusion high-fidelity DNA polymerase (NEB / M0530L) 0.5μL
[0245] 10 μM PCR primer R 2.5 μL
[0246] 10 μM PCR primer F 2.5 μL
[0247] For 10 μL of the solution obtained by the above [4] operation, the above mixture was added, and after incubation at 98°C for 30 seconds, PCR was performed in 18 cycles of 98°C for 10 seconds, 60°C for 20 seconds and 72°C for 20 seconds.
[0248] The base sequence of PCR primer R is represented by the following sequence number 9.
[0249] (Serial number 9)5'-CAAGCAGAAGACGGCATACGAGATCGTGATGCTCCACCGATAAATATTAGCCCGT-3'
[0250] The base sequence of PCR primer F is represented by the following sequence number 10.
[0251] (Serial No. 10)5'-AATGATACGGCGACCACCGAGATCTACACGTTCAGAGTTCTACAGTCCGA-3'
[0252] As ligases, T4 RNA ligase 2 (Truncation type, New England Biolabs Inc.) was used for ligation of the 3' adapter, and T4 RNA ligase 1 (New England Biolabs Inc.) was used for ligation of the 5' adapter.
[0253] Sequencing
[0254] The completed library was sequenced using an Illumina Nextseq 550 to determine the effective read rate.
[0255] <Results>
[0256] (Effective segment reading rate)
[0257] The effective read rates for 256 3' joints across 6 different 5' joints are shown in Tables 9 to 14. The base sequences shown in Tables 9 to 14 only display the four 5' final bases excluding the adenosine nucleotide at the 5' end; that is, only the base sequence corresponding to NNNN of sequence number 7 is displayed.
[0258] It should be noted that the base sequences numbered (numbered 1 to 78) in Tables 9 to 14 are base sequences with an effective read rate of 1.0% or higher for all six 5' connectors, and are embodiments of this disclosure. On the other hand, the base sequences without numbers are comparative examples of this disclosure. The numbers of each base sequence are the same as the numbers of the base sequences recorded in Table A of this disclosure.
[0259] [Table 17]
[0260] [Table 18]
[0261] [Table 19]
[0262] [Table 20]
[0263] [Table 21]
[0264] [Table 22]
[0265] Table 15 summarizes the 78 3' adapters that achieved an effective read rate of ≥1.0% for all six 5' adapter types. Achieving an effective read rate of ≥1.0% means that even using a small sequencer like the Illumina Miniseq, approximately 250,000 effective reads can be obtained, which is a number of effective reads sufficient for a certain level of analysis. Therefore, it can be said that 3' adapters with an effective read rate of ≥1.0% are practically excellent in terms of sequencing efficiency.
[0266] [Table 23]
[0267] Table 16 summarizes the 53 types of 3-' connectors that have an effective segment reading rate of more than 1.5% for all 6 types of 5' connectors.
[0268] [Table 24]
[0269] Table 17 summarizes the 37 3' adapters (37 types) that achieved an effective read rate of ≥2.0% for all six 5' adapter types. Achieving an effective read rate of ≥2.0% means that even using a small sequencer like the Illumina Miniseq, approximately 500,000 effective reads can be obtained; that is, a number of effective reads sufficient for high-precision analysis. Therefore, it can be said that 3' adapters with an effective read rate of ≥2.0% are particularly excellent in practical terms of sequencing efficiency.
[0270] [Table 25]
[0271] Table 18 summarizes the 17 types of 3' connectors that have an effective segment reading rate of more than 3.0% for all 6 types of 5' connectors.
[0272] [Table 26]
[0273] Table 19 summarizes the 8 types of 3-' connectors that have an effective segment reading rate of 5.0% or higher for all 6 types of 5' connectors.
[0274] [Table 27]
[0275] (Correlation between different 5' connector types and the effective segment reading rate of each 3' connector)
[0276] Plot the effective reading segments of each 3' connector using a certain 5' connector on the X-axis and the effective reading segments of each 3' connector using another certain 5' connector on the Y-axis, and calculate the correlation coefficient. The results are shown below. Figure 1 .
[0277] It should be noted that among the 256 3' linkers, only the 3' linker with a 5' end base sequence of "GGGG" (excluding the adenosine nucleotide at the 5' end) exhibits an extremely high effective read rate. Therefore, in the calculation of the correlation coefficient and... Figure 1 China excluded it.
[0278] Figure 1Among them, 5_adapter_1, 5_adapter_2, 5_adapter_3, 5_adapter_4, 5_adapter_5, and 5_adapter_6 respectively represent 5'-adapter (SEQ ID NO: 1), 5'-adapter (SEQ ID NO: 2), 5'-adapter (SEQ ID NO: 3), 5'-adapter (SEQ ID NO: 4), 5'-adapter (SEQ ID NO: 5), and 5'-adapter (SEQ ID NO: 6).
[0279] The correlation coefficients were calculated for all combinations (selecting 2 types of 5'-adapters from 6 types of 5'-adapters; 6C2 = 15 types), and all the correlation coefficients were 0.6 or more. A positive correlation was shown between the different types of 5'-adapters and the effective read segment rates of each 3'-adapter. That is, it was found that regardless of the type of 5'-adapter (i.e., the base sequence of the 5'-adapter), the more excellent 3'-adapters in the 3'-adapters of the present disclosure were equally more excellent.
[0280] Example 2 (Embodiment B): Relationship between the base sequences of the 3'-adapter and the 5'-adapter and the batch-to-batch differences of T4 RNA ligase 1
[0281] When preparing an NGS library and using T4 RNA ligase 1 for the ligation of the 3'-adapter and the 5'-adapter to the target nucleic acid, batch-to-batch differences occur. First, it was investigated which of the ligation of the 3'-adapter and the 5'-adapter was more affected by the batch-to-batch differences of T4 RNA ligase 1.
[0282] As the 3'-adapter, a 5'-terminal adenylated single-stranded DNA of 5’App / AACTGTAGGCACCATCAAT / 3’ddC (SEQ ID NO: 11) was purchased. The 3'-terminal of the 3'-adapter was capped by dideoxycytidine modification.
[0283] As the 5'-adapter, a single-stranded RNA of 5’-GUUCAGAGUUCUACAGUCCGACGAUC-3’ (SEQ ID NO: 4) was purchased.
[0284] As T4 RNA ligase 1, two types of T4 RNA ligase 1 with different batches (New England Biolabs Inc.) were prepared. Hereinafter, the two types of T4 RNA ligase 1 are respectively referred to as T4 RNA ligase 1(1) and T4 RNA ligase 1(2).
[0285] <Preparation of NGS Library>
[0286] Using the two types of T4 RNA ligase 1, NGS libraries were prepared according to the following steps under the following conditions (A) to (C).
[0287] (A) For both the ligation of the 3'-adapter and the 5'-adapter, two types of T4 RNA ligase 1 were used respectively
[0288] (B) Only the 3' adapter was ligated using two T4 RNA ligases 1, and the 5' adapter was ligated using T4 RNA ligase 1(1).
[0289] (C) Only the 5' adapter was ligated using two T4 RNA ligases 1, and the 3' adapter was ligated using T4 RNA ligase 1(1).
[0290] Using the 3' and 5' adapters described above, NGS libraries of small RNA (human serum RNA was used in this example) were prepared. The preparation method of the NGS library was performed according to Seda Eminaga et al., Quantification of microRNA Expression with Next-Generation Sequencing, Curr Protoc MolBiol. 2013 July;04: Unit 4.17. That is, the following operations were performed: [1] ligation of 3' adapters, [2] annealing of reverse transcription primers, [3] ligation of 5' adapters, [4] reverse transcription, [5] PCR, and [6] purification of PCR products.
[0291] [1] Connection of 3' connector
[0292] 1.0 μL of nuclease-free water
[0293] T4 RNA ligase reaction buffer (NEB / M0242) 1.0 μL
[0294] 100%DMSO(13445-74 / Nacalai Tesque)1.0μL
[0295] 50% PEG (NEB / B1004) 1.0μL
[0296] Total RNA (human serum-derived RNA) 4.0 μL
[0297] Add 10 μM of 3' conjugate to the above mixture and incubate at 90°C for 30 seconds.
[0298] T4 RNA ligase 1 (NEB / 0204M) 1.5 μL
[0299] SUPERase inhibitor (Thermo Fisher / AM2694) 0.5 μL
[0300] Then, the above mixture was added and incubated at 22°C for 60 minutes, thereby ligating the total RNA to the 3' adapter.
[0301] [2] Annealing of RT primers
[0302] For the solution obtained by the above [1] operation, add 1 μL of 10 μM RT primer, incubate at 90 °C for 30 seconds, and then incubate at 65 °C for 5 minutes.
[0303] The base sequence of the RT primer is represented by sequence number 8 below. The part represented by N is the UMI sequence. It should be noted that in the following sequences, N independently represents A, T, G, or C.
[0304] (Serial No. 8)5'-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNNNNATTGATGGTGCCTAC-3'
[0305] [3] Connection of 5' connector
[0306] 0.8 μL of nuclease-free water
[0307] ATP (NEB / 0204M) 1.4 μL
[0308] T4 RNA ligase 1 (NEB / 0204M) 1.5 μL
[0309] For the solution obtained by the above [2] operation, add 10 μM 5' connector to the above mixture and incubate at 20°C for 60 minutes.
[0310] [4] Reverse transcription
[0311] 3.5 μL of nuclease-free water
[0312] 0.3 μL of 25 mM dNTP mixture (Nippon Gene / 312-07271)
[0313] 5×FS buffer (Thermo Fisher / 18080093) 1.5μL
[0314] 0.1M DTT (Thermo Fisher / 18080093) 1.5μL
[0315] RNase OUT(Thermo Fisher / 10777019) 0.375μL
[0316] SuperScript III(Thermo Fisher / 18080093) 1.5μL
[0317] Then, 7.5 μL of the solution obtained from the above incubation was added to the above mixture, and incubated at 48°C for 30 seconds, and further incubated at 85°C for 5 minutes.
[0318] It should be noted that the base sequence of the RT primer is shown in Serial No. 8 above.
[0319] [5]PCR
[0320] 24.1 μL of nuclease-free water
[0321] Phusion HF buffer (NEB / M0530L) 10μL
[0322] 0.4 μL of 25 mM dNTP mixture (Nippon Gene / 312-07271)
[0323] Phusion high-fidelity DNA polymerase (NEB / M0530L) 0.5μL
[0324] 10 μM PCR primer R 2.5 μL
[0325] 10 μM PCR primer F 2.5 μL
[0326] 10 μL of the solution obtained by the above [4] operation was added to the above mixture, and after incubation at 98°C for 30 seconds, PCR was performed in 18 cycles of 98°C for 10 seconds, 60°C for 20 seconds and 72°C for 20 seconds.
[0327] The base sequence of PCR primer R is represented by the following sequence number 9.
[0328] (Serial number 9)5'-CAAGCAGAAGACGGCATACGAGATCGTGATGCTCCACCGATAAATATTAGCCCGT-3'
[0329] The base sequence of PCR primer F is represented by the following sequence number 10.
[0330] (Serial No. 10)5'-AATGATACGGCGACCACCGAGATCTACACGTTCAGAGTTCTACAGTCCGA-3'
[0331] Sequencing
[0332] The completed library was sequenced using an Illumina Nextseq 550.
[0333] The influence of batch-to-batch variability was investigated by calculating the coefficients of determination for the quantitative values of the top 100 miRNAs expressed when using two different T4 RNA ligase 1s. The coefficients of determination were calculated by plotting the quantitative values of miRNAs (mean of n=4) in NGS libraries prepared using T4 RNA ligase 1 (1) on the X-axis and the corresponding quantitative values of miRNAs (mean of n=4) in NGS libraries prepared using T4 RNA ligase 1 (2) on the Y-axis. Figures 2-4 As shown, (A) the ligation of the 3' and 5' adapters is performed using two different T4 RNA ligases. Figure 2 The coefficient of determination for (B) is 0.95, and the ligation of only the 3' adapter is performed using two T4 RNA ligases. Figure 3 The coefficient of determination for (C) was 0.96, and the ligation of only the 5' adapter was performed using two T4 RNA ligases. Figure 4 The coefficient of determination was 0.99. This indicates that batch-to-batch variation has a significant impact on 3' adapter ligation, while it is almost negligible in 5' adapter ligation. This is presumably because, when using T4 RNA ligase 1, 3' adapter ligation is less efficient than 5' adapter ligation, and a larger batch-to-batch variation was observed in 3' adapter ligation. Therefore, to reduce the impact of batch-to-batch variation, the base sequence of the 3' adapter was investigated.
[0334] Example 3 (Implementation Method B): Study on the base sequence of the 3' linker that is not easily affected by batch-to-batch variations of T4 RNA ligase 1
[0335] Next, 3' adapters less susceptible to batch-to-batch variations of T4 RNA ligase 1 were screened. Fifteen sequences were selected from those yielding an effective read rate of 1% or higher in Example 1. A total of 30 libraries were prepared using the 15 3' adapters and two batches of T4 RNA ligase 1, and the effect of batch-to-batch variations of T4 RNA ligase 1 when using each 3' adapter was investigated. It should be noted that, according to Example 2, batch-to-batch variations of T4 RNA ligase 1 were ensured to occur during 3' adapter ligation; according to Example 1, an effective read rate of 1% or higher was ensured regardless of the 5' adapter sequence among the 15 3' adapters. Therefore, 5' adapter 4 (sequence number 4) was used as the 5' adapter in this Example 3.
[0336] Similar to Example 2, the impact of batch-to-batch variability was investigated by calculating the coefficients of determination for the quantitative values of the top 100 miRNAs expressed using both T4 RNA ligases 1. The coefficients of determination are shown in the table below for each 3' adapter. (See table below.) Figure 5 and Figure 6As shown, a 3' linker (e.g., a 5' end with TCCA) was found to be strongly affected by batch-to-batch variability in T4 RNA ligase 1. Figure 5 (); the coefficient of determination is approximately 0.93). On the other hand, there are also adapter sequences that are not easily affected by batch-to-batch differences in T4 RNA ligase 1 (e.g., GACA at the 5' end). Figure 6 (The coefficient of determination is approximately 0.98).
[0337] [Table 28]
[0338] Based on the results of Examples 1 to 3, it was determined that, as a 3' adapter capable of inhibiting the formation of adapter dimers and batch-to-batch variability of T4 RNA ligase 1 (Example B), a 3' adapter with a different batch-to-batch determination coefficient of 0.960 or higher at the 5' end is suitable.
[0339] [Table 29]
[0340] (Note) The means disclosed herein include the following.
[0341] {1} A connection method, comprising: A 5'-terminal adenylated single-stranded DNA, serving as a 3' adapter, is ligated to the 3' end of each of several different types of nucleic acids. The 5'-terminal base sequence of this adenylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table A below; and After the ligation of the 5'-terminal adenylated single-stranded DNA, single-stranded nucleic acids serving as 5' adapters are ligated to the 5' ends of each of the aforementioned types of nucleic acids.
[0342] [Table 30]
[0343] {2} According to the connection method described in {1}, the base sequence at the 5' end of the 3' connector is selected from any one of the group consisting of base sequences numbered 1 to 53 shown in Table A above.
[0344] {3} The ligation method as described in {1} or {2}, wherein the ligase used in the ligation of the 3' connector and the ligase used in the ligation of the 5' connector are each independently selected from at least one of the group consisting of T4 RNA ligase 1, T4 RNA ligase 2, T4 DNA ligase and their variants.
[0345] {4} A method for inhibiting adapter dimer formation, comprising: linking a 5'-terminal adenylated single-stranded DNA, serving as a 3' adapter, to the 3' ends of multiple types of nucleic acids with different sequences, wherein the 5'-terminal base sequence of the 5'-terminal adenylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table A below; and
[0346] After the ligation of the 5'-terminal adenylated single-stranded DNA, single-stranded nucleic acids serving as 5' adapters are ligated to the 5' ends of each of the aforementioned types of nucleic acids.
[0347] [Table 31]
[0348] {5}The method for inhibiting the formation of linker dimers according to {4}, wherein the base sequence at the 5' end of the 3' linker is selected from any one of the group consisting of base sequences numbered 1 to 53 shown in Table A above.
[0349] {6} The method for inhibiting adapter dimer formation according to {4} or {5}, wherein the ligase used in the ligation of the 3' adapter and the ligase used in the ligation of the 5' adapter are each independently selected from at least one of the group consisting of T4 RNA ligase 1, T4 RNA ligase 2, T4 DNA ligase and their variants.
[0350] {7} A method for preparing an NGS library, comprising: Using any one of the ligation methods in {1} to {3}, a 5'-terminal adenylated single-stranded DNA and a single-stranded nucleic acid serving as a 5' adapter are ligated to multiple types of nucleic acids with different sequences. cDNA is synthesized from the aforementioned types of nucleic acids linked with the aforementioned 3' adapter and 5' adapter via reverse transcription; and The obtained cDNA was amplified by PCR.
[0351] {8} An NGS library preparation kit comprising 5'-terminal adenylated single-stranded DNA contained in a container, wherein the 5'-terminal base sequence of the 5'-terminal adenylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table A below.
[0352] [Table 32]
[0353] {9} According to the NGS library preparation kit described in {8}, the 5' end base sequence of the above-mentioned 5' end adenylated single-stranded DNA is selected from any one of the group consisting of base sequences numbered 1 to 53 shown in Table A above.
[0354] {10} A 3' adapter used in the preparation of a nucleic acid library, which is a single-stranded DNA with its 5' end nucleotide adenosine-modified, wherein the 5' end base sequence of the single-stranded DNA, excluding the 5' end adenosine nucleotide, is any of the base sequences shown in Table A below.
[0355] [Table 33]
[0356] {11} The 3' adapter used in the preparation of the nucleic acid library as described in {10}, wherein the base sequence at the 5' end is selected from any one of the group consisting of base sequences numbered 1 to 53 shown in Table A above.
[0357] The means of implementing embodiment B of this disclosure include the following methods.
[0358] (1) A ligation method comprising: ligating a 5'-terminal adenylated single-stranded DNA as a 3' adapter to the 3' ends of each of a plurality of nucleic acids with different sequences, wherein the 5'-terminal base sequence of the 5'-terminal adenylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table B below; and
[0359] After the ligation of the 5'-terminal adenylated single-stranded DNA, single-stranded nucleic acids serving as 5' adapters are ligated to the 5' ends of each of the aforementioned types of nucleic acids.
[0360] [Table 34]
[0361] (2) According to the ligation method described in (1), wherein the ligase used in the ligation of the above 3' connector is T4 RNA ligase 1.
[0362] (3) According to the connection method described in (1) or (2), wherein the above-mentioned multiple types of nucleic acids are small RNAs.
[0363] (4) A method for preparing an NGS library, comprising: Using any one of the ligation methods in (1) to (3), a 5'-terminal adenylated single-stranded DNA and a 5'-terminal single-stranded nucleic acid are ligated to multiple types of RNA with different sequences. cDNA is synthesized from the aforementioned types of RNA linked with the aforementioned 3' and 5' adapters via reverse transcription; and The obtained cDNA was amplified by PCR.
[0364] (5) An NGS library preparation kit comprising 5'-terminal adenylated single-stranded DNA contained in a first container, wherein the 5'-terminal base sequence of the 5'-terminal adenylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table B below.
[0365] [Table 35]
[0366] (6) The NGS library preparation kit according to (5), wherein it further comprises T4 RNA ligase 1 contained in a second container.
[0367] (7) The NGS library preparation kit according to (5) further comprises, in such a manner as to be housed in separate containers, at least one selected from the group consisting of a 5' adapter, a ligase, a reverse transcriptase, a reverse transcription primer, a PCR enzyme, and a PCR primer.
[0368] (8) A 3' adapter for use in the preparation of a nucleic acid library, which is a single-stranded DNA with its 5' end nucleotide adenosine-modified, wherein the 5' end base sequence of the single-stranded DNA, excluding the 5' end adenosine nucleotide, is any of the base sequences shown in Table B below.
[0369] [Table 36]
[0370] (9) A 3' adapter used in ligation, which is a single-stranded DNA with its 5' end nucleotide adenosine-modified, wherein the 5' end base sequence of the single-stranded DNA, excluding the 5' end adenosine nucleotide, is any of the base sequences shown in Table B below.
[0371] [Table 37]
[0372] The publications of Japanese Patent Application No. 2023-177843, filed on October 13, 2023, and Japanese Patent Application No. 2024-071785, filed on April 25, 2024, are incorporated herein by reference in their entirety. All documents, patent applications, and technical standards described in this specification are incorporated herein by reference in the same way as if each document, patent application, and technical standard were specifically and individually described and incorporated herein by reference.
Claims
1. A connection method, comprising: A 5'-terminal adenosylated single-stranded DNA, which serves as a 3' adapter, is ligated to the 3' end of each of several different types of nucleic acids. The 5'-terminal base sequence of the 5'-terminal adenosylated single-stranded DNA, excluding the adenosine nucleotide at the 5' end, is any of the base sequences shown in Table A below. as well as Following the ligation of the 5'-terminal adenylated single-stranded DNA, single-stranded nucleic acids serving as 5' adapters are ligated to the 5' ends of each of the plurality of nucleic acids. [Table 1] 。 2. The connection method according to claim 1, wherein, The ligase used in the ligation of the 3' connector and the ligase used in the ligation of the 5' connector are each independently selected from at least one of the group consisting of T4 RNA ligase 1, T4 RNA ligase 2, T4 DNA ligase and their variants.
3. The connection method according to claim 1, wherein, The base sequence at the 5' end of the 3' connector is any one of the base sequences shown in Table B below. [Table 2] 。 4. The connection method according to claim 3, wherein, The ligase used in the 3' connector ligation is T4 RNA ligase 1.
5. The connection method according to claim 3, wherein, The various types of nucleic acids are all small RNAs.
6. A method for inhibiting the formation of linker dimers, comprising: In multiple types of nucleic acids with different sequences, the 3' end of each nucleic acid is used as a 5' end adenosylated single-stranded DNA that connects to a 3' adapter. The 5' end base sequence of the 5' end adenosine nucleotide, excluding the 5' end adenosine nucleotide, is any of the base sequences shown in Table A below. as well as Following the ligation of the 5'-terminal adenylated single-stranded DNA, single-stranded nucleic acids serving as 5' adapters are ligated to the 5' ends of each of the plurality of nucleic acids. [Table 3] 。 7. The method for suppressing the formation of connector dimers according to claim 6, wherein, The ligase used in the ligation of the 3' connector and the ligase used in the ligation of the 5' connector are each independently selected from at least one of the group consisting of T4 RNA ligase 1, T4 RNA ligase 2, T4 DNA ligase and their variants.
8. A method for preparing an NGS library, comprising: The ligation method described in any one of claims 1 to 5 is used to ligate a 5'-terminal adenylated single-stranded DNA and a 5'-terminal single-stranded nucleic acid as 3' adapters to multiple types of nucleic acids with different sequences. cDNA is synthesized from the plurality of nucleic acids linked to the 3' and 5' adapters via reverse transcription; and The obtained cDNA was amplified by PCR.
9. The method for preparing an NGS library according to claim 8, wherein, The different types of nucleic acids are different types of RNA.
10. An NGS library preparation kit comprising 5'-terminal adenylated single-stranded DNA contained in a first container, wherein the 5'-terminal base sequence of the 5'-terminal adenylated single-stranded DNA, excluding the 5'-terminal adenosine nucleotide, is any of the base sequences shown in Table A below. [Table 4] 。 11. The NGS library preparation kit according to claim 10, wherein, The 5' end sequence of the single-stranded DNA, excluding the adenosine nucleotide at the 5' end, is any of the base sequences shown in Table B below. [Table 5] 。 12. The NGS library preparation kit according to claim 11, wherein, It further includes T4 RNA ligase 1 housed in a second container.
13. The NGS library preparation kit according to claim 11, wherein, It further comprises, in a manner that allows each to be housed in a separate container, at least one selected from the group consisting of a 5' adapter, a ligase, a reverse transcriptase, a reverse transcription primer, a PCR enzyme, and a PCR primer.
14. A 3' adapter for use in the preparation of a nucleic acid library, comprising a single-stranded DNA with its 5' end nucleotide adenylated, wherein the 5' end base sequence of the single-stranded DNA, excluding the 5' end adenosine nucleotide, is any of the base sequences shown in Table A below. [Table 6] 。 15. The 3' adapter used in the preparation of nucleic acid libraries according to claim 14, wherein, The 5' end sequence of the single-stranded DNA, excluding the adenosine nucleotide at the 5' end, is any of the base sequences shown in Table B below. [Table 7] 。 16. A 3' adapter for use in ligation, comprising a single-stranded DNA with its 5' end nucleotide adenylated, wherein the 5' end base sequence of the single-stranded DNA, excluding the 5' end adenosine nucleotide, is any of the base sequences shown in Table B below. [Table 8] 。