Cuttable DNA-coding library

By introducing a cleavable site in DNA strands using the USER® enzyme, the limitations of hairpin and double-strand DNA structures in DELs are overcome, enhancing PCR efficiency and adaptability, thus improving DEL synthesis and evaluation.

JP7838610B2Active Publication Date: 2026-04-01NISSAN CHEM CORP
View PDF 12 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Existing DNA-encoded libraries (DELs) face challenges in combining the advantages of hairpin-strand and double-strand DNA structures, particularly in terms of PCR efficiency and adaptability to various chemical conditions.

Method used

Introduction of a cleavable site, such as deoxyuridine, into the DNA strand using the USER® enzyme, allowing for selective cleavage, thereby integrating the benefits of both hairpin and double-strand DNA structures.

Benefits of technology

This approach enhances PCR efficiency and adaptability of DELs to a wider range of chemical conditions, improving the synthesis and evaluation processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007838610000067
    Figure 0007838610000067
  • Figure 0007838610000068
    Figure 0007838610000068
  • Figure 0007838610000069
    Figure 0007838610000069
Patent Text Reader

Abstract

To provide a technique that realizes both of the advantage of hairpin-stranded DNA and the advantage of double-stranded DNA as DNA strand structures for DNA-encoded libraries.SOLUTION: The present invention relates to a method for using a nucleic acid compound that contains a selectively cleavable site. The present invention also relates to a DNA-encoded library that contains a selectively cleavable site, a composition for synthesizing the same, and a method for using the same.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a DNA-encoded library containing a cleavable site in a DNA strand.

Background Art

[0002] A compound library is a group of compound derivatives systematically collected of compounds that may have specific activities, such as drug candidate compounds. This compound library is often synthesized based on the synthetic techniques and methodologies of combinatorial chemistry. Combinatorial chemistry is an experimental technique and a research field related thereto for efficiently synthesizing a variety of compounds in a systematic synthetic route from a series of compound libraries enumerated and designed based on combinatorial theory. As one type of compound library based on combinatorial chemistry, there is a DNA-encoded library. Hereinafter, the DNA-encoded library is appropriately abbreviated as DEL. In DEL, a DNA tag is added to each of the library-formed compounds. The DNA tag is designed with a sequence so that each structure of each compound can be identified and functions as a label of the compound (Patent Documents 1 to 3). The conventionally known DNA strand structures of DEL are typically two types: double strands and hairpin strands. Hereinafter, an overview of double-stranded DEL and hairpin-strand DEL and their advantages and disadvantages will be outlined. (1) Hairpin-strand DEL DEL using hairpin-strand DNA has a single-stranded structure in which two complementary DNA strands are connected, and is synthesized using hairpin-type DNA having a functional group for introducing various building blocks as a starting material (head piece) (Patent Document 3, Non-Patent Documents 1 and 2). (A) Advantages (a) Short DNA tags can be used. In this method, relatively short double-stranded DNA tags of about 9-13 mers with 2 mer sticky ends are often used, and these double-stranded DNA tags are introduced by a ligation reaction using DNA ligase. The use of such short DNA tags is possible because hairpin strand DNA forms a strong double helix within the molecule, and DNA regions other than the sticky ends do not interfere with the DNA tag. The use of short double-stranded DNA tags has several advantages in DEL synthesis. One advantage is that the cost of synthesizing the DNA tag is low. Another advantage is that using shorter DNA tags allows for a shorter overall DEL length when encoding the same number of reaction cycles. That is, even when encoding a larger number of cycles, the overall length of the DEL can be kept within a range that allows for efficient DNA sequence reading by next-generation sequencers. In fact, Non-Patent Literature 3 demonstrates the construction of a DEL using hairpin strand DNA that encodes as many as 6 reaction cycles. (b) High chemical stability Unlike double-stranded DNA, in hairpin strands, even if the double-stranded structure melts during a heating reaction, the double helix is ​​reformed within the original molecule under subsequent re-annealing conditions without strand exchange. Therefore, DEL using hairpin strand DNA has the advantage of being usable under a wider range of chemical conditions (Non-Patent Literature 2). In addition, generally, for the same chain length, hairpin strands form a stronger double helix than double-stranded DNA (higher Tm value). Therefore, under various chemical conditions when introducing building blocks, the chemical structures of hairpin strand DNA, especially the base region, should be more resistant to structural transformation than double-stranded DNA. (B) Disadvantages Hairpin strand DNA has a strong ability to form double helix strands, which makes it difficult to melt the double helix strands, bind primer oligonucleotides, and initiate the polymerase reaction, resulting in low PCR efficiency (Patent Document 4). (2) Double strand DEL DEL, which uses double-stranded DNA, is synthesized using single-stranded DNA (single-stranded DNA that is not a hairpin strand) or double-stranded DNA containing functional groups for introducing various building blocks as a starting material (headpiece). (A) Disadvantages In contrast to DELs that use hairpin strand DNA, relatively long single-stranded or double-stranded DNA tags of about 20-30 mers with 4-10 mers of sticky ends are often used (Patent Document 2, Non-Patent Document 4), and DELs that encode a reaction of about 3 cycles are common. (B) Strengths DEL, which uses double-stranded DNA, does not have the same problems as hairpin-stranded DNA in terms of PCR efficiency. Furthermore, unlike hairpin-stranded DNA, double-stranded DNA can be converted to single-stranded DNA through denaturation or strand exchange reactions, and has the advantage of being adaptable to a wide range of evaluation methods by being converted to DNA structures suitable for various applications. For example, highly sensitive evaluation methods that utilize the double-strand formation ability of DNA have been developed (Non-Patent Documents 5 and 6).

[0003] Thus, while hairpin strand DNA and double-stranded DNA each have advantages in DEL synthesis and evaluation, no technology is known that combines these advantages. [Prior art documents] [Patent Documents]

[0004] [Patent Document 1] International Publication No. 93 / 20243 [Patent Document 2] International Publication No. 2004 / 039825 [Patent Document 3] International Publication No. 2005 / 058479 [Patent Document 4] International Publication No. 2010 / 094036 [Non-patent literature]

[0005] [Non-Patent Document 1] Nature Chemical Biology, 2009, Vol. 5, pp. 647-654. [Non-Patent Document 2] A Handbook for DNA-Encoded Chemistry, edited by Robert A. Goodnow, Jr., John Wiley & Sons, Inc. [Non-Patent Document 3] ACS Chemical Biology, 2018, Vol. 13, pp. 53-59. [Non-Patent Document 4] Nature Chemistry, 2018, Vol. 10, pp. 441-448. [Non-Patent Document 5] Annual Review of Biochemistry, 2018, Vol. 87, pp. 479-502 [Non-Patent Document 6] ACS Combinatorial Science, 2020, Vol. 22, pp. 204-212. [Overview of the project] [Problems that the invention aims to solve]

[0006] This invention provides a DEL containing a cleavable site within a DNA strand, and a method for producing a DEL. [Means for solving the problem]

[0007] One aspect of nucleic acid chemistry, such as DNA cleavage techniques, involves the introduction of deoxyuridine into a DNA strand, which allows for selective cleavage by the USER® enzyme. As a result of diligent research, the inventors discovered that by introducing cleavable sites, such as deoxyuridine, into the DNA strand, it is possible to combine the advantages of both hairpin strand DNA and double-stranded DNA, thus completing the present invention. Therefore, the present invention is as follows.

[0008] [1] Equation (I) [ka] (wherein E and F are each independently an oligomer composed of nucleotides or nucleic acid analogs, provided that E and F contain base sequences complementary to each other to form a double-stranded oligonucleotide, LP is a loop site, L is a linker, D is a reactive functional group.) is a compound represented by a compound having at least one selectively cleavable site at at least one site of E, F, or LP. [2] A composition for use in preparing the headpiece of a compound library, comprising the compound according to [1]. [3] A composition for use in preparing the headpiece of a DNA-encoded library, comprising the compound according to [1]. [4] Formula (I) [Chemical formula] (wherein E and F are each independently an oligomer composed of nucleotides or nucleic acid analogs, provided that E and F contain base sequences complementary to each other to form a double-stranded oligonucleotide, LP is a loop site, L is a linker, D is a reactive functional group.) is a compound represented by having at least one selectively cleavable site at at least one site of E, F, or LP, and is used as the headpiece of a compound library. [5] Formula (I) [Chemical formula] <00001XX> (wherein E and F are independent of each other. It is an oligomer composed of nucleotides or nucleic acid analogs, However, E and F contain complementary base sequences and form a double-stranded oligonucleotide. LP is the loop section, L is a linker, D is a reactive functional group. It is a compound represented by the following: At least one of parts E, F, or LP has at least one selectively cleavable portion, A compound used as a headpiece in DNA coding libraries. [6] Equation (I) [ka] (In the formula, E and F are independent of each other. It is an oligomer composed of nucleotides or nucleic acid analogs, However, E and F contain complementary base sequences and form a double-stranded oligonucleotide. LP is the loop section, L is a linker, D is a reactive functional group. It is a compound represented by the following: A compound library headpiece having at least one selectively cleavable site in at least one of sites E, F, or LP. [7] Equation (I) [ka] (In the formula, E and F are independent of each other. It is an oligomer composed of nucleotides or nucleic acid analogs, However, E and F contain complementary base sequences and form a double-stranded oligonucleotide. LP is the loop section, L is a linker, D is a reactive functional group. It is a compound represented by the following: A headpiece of a DNA encoding library having at least one selectively cleavable site in at least one of the sites E, F, or LP. [8] Formula (II) [ka] (In the formula, X and Y are nucleotide chains, E and F are independent of each other. It is an oligomer composed of nucleotides or nucleic acid analogs, However, E and F contain complementary base sequences and form a double-stranded oligonucleotide. LP is the loop section, L is a linker, D is a divalent group derived from a reactive functional group, Sp is a bond or a bifunctional spacer, An is a substructure composed of at least one building block. It is a compound represented by the following: X and Y have sequences that can form a double helix in at least part of their structure. X binds to E at its 5' end. Y binds to F at its 3' end. A compound having at least one selectively cleavable site at any one of the sites E, F, or LP. [9] Formula (III) An-Sp-C-Bn (III) (In the formula, An and Sp have the same meaning as in [8], Bn represents a double-stranded oligonucleotide tag formed by oligonucleotide chain X and oligonucleotide chain Y. C is given by equation (I) [ka] (In the formula, E, LP, L, D, and F have the same meanings as in [8], except that D binds to Sp, and E and F bind to the corresponding terminals of the double-stranded oligonucleotide tag Bn.) Represented by, The compounds described in [8].

[10] An is the same as [8] and is a substructure constructed of n building blocks α1 to αn (where n is an integer from 1 to 10), Bn is a double-stranded oligonucleotide tag formed by oligonucleotide chain X and oligonucleotide chain Y, and is a substructure containing an oligonucleotide with a base sequence that can identify the structure of An. The compounds described in [8] or [9].

[11] LP, This is the loop region represented by (LP1)p-LS-(LP2)q, LS is a substructure selected from the group of compounds described in (A) to (C) below, (A) Nucleotides (B) Nucleic acid analogs (C) Trivalent C1-14 groups which may have substituents LP1 is a substructure selected individually or differently from the group of compounds described in (1) and (2) below, (1) Nucleotides (2) Nucleic acid analogs LP2 is each of q substructures selected individually or differently from the group of compounds described in (1) and (2) below. (1) Nucleotides (2) Nucleic acid analogs The total number of p and q is between 0 and 40. A compound described in any of [1], [4], [5], [8] to

[10] .

[12] The compound described in

[11] , wherein the total number of p and q is between 2 and 20.

[13] The compound described in

[11] , wherein the total number of p and q is between 2 and 10.

[14] The compound described in

[11] , wherein the total number of p and q is between 2 and 7.

[15] The compound described in

[11] , wherein the total number of p and q is 0.

[16] LP1, LP2, and LS have the following structures: (A) Nucleotides or (B) Nucleic acid analogs that meet the following requirements (B11) to (B15) (B11) Having phosphoric acid (or equivalent part) and a hydroxyl group (or equivalent part), (B12) Composed of carbon, hydrogen, oxygen, nitrogen, phosphorus, or sulfur, (B13) Molecular weight is between 142 and 1500. (B14) The number of atoms between residues is 3 to 30. (B15) The bonding pattern between atoms in residues is either all single bonds, or one or two double bonds with the remainder being single bonds. A compound described in any of

[11] to

[15] , having a structure selected individually or differently from the above.

[17] LP1, LP2, and LS have the following structures: (A) Nucleotides or (B) Nucleic acid analogs that meet the following requirements (B21) to (B25) (B21) Having phosphoric acid and hydroxyl groups, (B22) Composed of carbon, hydrogen, oxygen, nitrogen, or phosphorus, (B23) Molecular weight is between 142 and 1000. (B24) The number of atoms between residues is 3 to 15. (B25) The bonding mode between atoms in each residue is all single bonds. A compound described in any of

[11] to

[16] , having a structure selected individually or differently from the above.

[18] LP1, LP2, and LS have the following structures, respectively: (A) Nucleotides or (B) Nucleic acid analogs that meet the following requirements (B31) to (B35) (B31) Having phosphoric acid and hydroxyl groups, (B32) Composed of carbon, hydrogen, oxygen, nitrogen, or phosphorus, (B33) Molecular weight is between 142 and 700. (B34) The number of atoms between residues is 4 to 7. (B35) The bonding mode between atoms in each residue is all single bonds. A compound described in any of

[11] to

[17] , having a structure selected individually or differently from the above.

[19] LP1 and LP2 are as follows: (B41)d-Spacer, (B5) Polyalkylene glycol phosphate A compound described in any of

[11] to

[18] , which is one of the following.

[20] The compound according to any one of

[11] to

[19] , wherein LP1 and LP2 are each diethylene glycol phosphate ester or triethylene glycol phosphate ester.

[21] The compound according to any one of

[11] to

[20] , wherein LP1 and LP2 are each triethylene glycol phosphate esters.

[22] A compound according to any of

[11] to

[19] , wherein LP1 and LP2 are each d-Spacer. [twenty three] LP1 and LP2 are nucleotides, A compound described in any of

[11] to

[18] .

[24] LS is given by equations (a) to (g): [ka] (In the formula, * represents the bond position with the linker, ** represents the bond position with LP1 or LP2, and R represents a hydrogen atom or a methyl group.) A compound described in any of

[11] to

[23] , which is one of the following.

[25] LS is given by equation (h): [ka] (In the formula, * indicates the linker connection position, and ** indicates the linker connection position.) The compound described in any of

[11] to

[23] .

[26] A compound according to any of

[11] to

[23] , wherein LS is a polyalkylene glycol phosphate ester.

[27] LS is given by equations (i) to (k): [ka] (In the formula, n1, m1, p1, and q1 are each independent integers between 1 and 20, * indicates the linker connection position, and ** indicates the linker connection position.) The compound described in any of

[11] to

[23] .

[28] LS is given by equation (l): [ka] (In the formula, * indicates the linker connection position, and ** indicates the linker connection position.) The compound described in any of

[11] to

[23] .

[29] LS is (B42), (B43), or (B44): (B42) Amino C6 dT (B43)mdC(TEG-Amino) (B44) Uni-Link (Trademark Registered) Amino Modifier A compound described in any of

[11] to

[23] , which is one of the following.

[30] LS is a nucleotide, A compound described in any of

[11] to

[23] .

[31] LS is a trivalent C1-14 group which may have a (C) substituent, and (C) has the following structure: (1) C1-10 aliphatic hydrocarbons which may have substituents and which may be replaced by 1-3 heteroatoms, (2) C6-14 aromatic hydrocarbons which may have substituents, (3) A C2-9 aromatic heterocycle which may have substituents, or (4) C2-9 non-aromatic heterocycles which may have substituents A compound described in any of

[11] -

[15] and any of

[19] -

[23] , which is one of the above.

[32] LS is a trivalent C1-14 group which may have a (C) substituent, and (C) has the following structure: (1) C1-6 aliphatic hydrocarbons which may have substituents, (2) C6-10 aromatic hydrocarbons which may have substituents, (3) C2-5 aromatic heterocycles which may have substituents The compound described in any of

[11] -

[15] and

[19] -

[23] , which is one of the following:

[33] LS is a trivalent C1-14 group which may have a (C) substituent, and (C) has the following structure: (1) C1-6 aliphatic hydrocarbons, (2) Benzene, or (3)C2~5 nitrogen-containing aromatic heterocycle Here, (1) to (3) may be unsubstituted or substituted with one to three substituents selected individually or differently from substituent group ST1, wherein substituent group ST1 consists of C1-6 alkyl groups, C1-6 alkoxy groups, fluorine atoms, and chlorine atoms; however, if substituent group ST1 is substituted with an aliphatic hydrocarbon, alkyl groups are not selected from substituent group ST1. The compound described in any of

[11] -

[15] and

[19] -

[23] , which is one of the following:

[34] LS is a trivalent C1-14 group which may have a (C) substituent, and (C) has the following structure: (1) C1-6 alkyl groups, or (2) Unsubstituted or benzenes substituted with one or two C1-3 alkyl or C1-3 alkoxy groups The compound described in any of

[11] -

[15] and

[19] -

[23] , which is one of the following:

[35] LS is a trivalent C1-14 group which may have a (C) substituent, and (C) has the following structure: (1) C1-6 alkyl groups The compound described in any of

[11] -

[15] and

[19] -

[23] .

[36] E and F are oligomers composed independently of nucleotides or nucleic acid analogs, The chain lengths of E and F are 3 to 40, respectively. A compound described in any of [1], [4], [5], [8] to

[35] .

[37] E and F are oligomers composed independently of nucleotides or nucleic acid analogs, The chain lengths of E and F are 4 to 30, respectively. A compound described in any of [1], [4], [5], [8] to

[36] .

[38] E and F are oligomers composed independently of nucleotides or nucleic acid analogs, The chain lengths of E and F are 6 to 25, respectively. A compound described in any of [1], [4], [5], [8] to

[37] .

[39] E and F are oligomers composed independently of nucleotides or nucleic acid analogs, E and F contain complementary base sequences, forming a double-stranded oligonucleotide. The E and F double-stranded oligonucleotides are the protruding ends. A compound described in any of [1], [4], [5], [8] to

[38] .

[40] The compound according to

[39] , wherein the protrusion of the protruding end is two bases or longer in length.

[41] E and F are oligomers composed independently of nucleotides or nucleic acid analogs, E and F contain complementary base sequences, forming a double-stranded oligonucleotide. The double-chain oligonucleotides E and F have blunt ends. A compound described in any of [1], [4], [5], [8] to

[38] .

[42] The chain lengths of the complementary base sequences contained in E and F are each 3 bases or more. A compound described in any of [1], [4], [5], [8] to

[41] .

[43] The chain lengths of the complementary base sequences contained in E and F are each 4 bases or more. A compound described in any of [1], [4], [5], [8] to

[42] .

[44] The chain lengths of the complementary base sequences contained in E and F are each 6 bases or more. A compound described in any of [1], [4], [5], [8] to

[43] .

[45] The compound according to any of [1], [4], [5], [8] to

[44] , wherein E and F are oligomers each independently composed of nucleotides.

[46] The nucleotide is a ribonucleotide or a deoxyribonucleotide. A compound described in any of [1], [4], [5], [8] to

[45] .

[47] nucleotide is deoxyribonucleotide, A compound described in any of [1], [4], [5], [8] to

[46] .

[48] ​​The nucleotide is deoxyadenosine, deoxyguanosine, thymidine, or deoxycytidine. A compound described in any of [1], [4], [5], [8] to

[47] .

[49] The compound according to any of [1], [4], [5], [8] to

[44] , wherein E and F are oligomers each independently composed of nucleic acid analogs.

[50] L is (1) C1-20 aliphatic hydrocarbons which may have substituents and which may be replaced by 1-3 heteroatoms, or (2) C6-14 aromatic hydrocarbons which may have substituents The compound described in any of [1], [4], [5], [8] to

[49] .

[51] A compound according to any one of [1], [4], [5], [8] to

[50] , wherein L is a C1-6 aliphatic hydrocarbon which may have substituents, a C1-6 aliphatic hydrocarbon which may be replaced by one or two oxygen atoms, or a C6-10 aromatic hydrocarbon which may have substituents.

[52] A compound according to any of [1], [4], [5], [8] to

[51] , wherein L is a C1-6 aliphatic hydrocarbon substituted with substituent group ST1, or a benzene substituted with substituent group ST1, where substituent group ST1 is a group consisting of C1-6 alkyl groups, C1-6 alkoxy groups, fluorine atoms and chlorine atoms (however, when substituent group ST1 is substituted with an aliphatic hydrocarbon, alkyl groups are not selected from substituent group ST1).

[53] The compound according to any one of [1], [4], [5], [8] to

[52] , wherein L is a C1-6 alkyl group, or a benzene that is unsubstituted or substituted with one or two C1-3 alkyl groups or C1-3 alkoxy groups.

[54] A compound according to any of [1], [4], [5], [8] to

[53] , wherein L is a C1-6 alkyl group.

[55] The reactive functional group D is A compound according to any one of [1], [4], [5], [8] to

[54] , which is a C-C, amino, ether, carbonyl, amide, ester, urea, sulfide, disulfide, sulfoxide, sulfonamide, or a reactive functional group capable of forming a sulfonyl bond.

[56] The compound according to any one of [1], [4], [5], [8] to

[55] , wherein the reactive functional group of D is a C1 hydrocarbon having a leaving group, an amino group, a hydroxyl group, a precursor of a carbonyl group, a thiol group, or an aldehyde group.

[57] The compound according to any one of [1], [4], [5], [8] to

[56] , wherein the reactive functional group of D is a C1 hydrocarbon having a halogen atom, a C1 hydrocarbon having a sulfonic acid leaving group, an amino group, a hydroxyl group, a carboxyl group, a halogenated carboxyl group, a thiol group, or an aldehyde group.

[58] The compound according to any of [1], [4], [5], [8] to

[57] , wherein the reactive functional group of D is -CH2Cl, -CH2Br, -CH2OSO2CH3, -CH2OSO2CF3, an amino group, a hydroxyl group, or a carboxyl group.

[59] The compound according to any of [1], [4], [5], [8] to

[58] , wherein the reactive functional group of D is a primary amino group.

[60] The selectively cleavable sites are deoxyribonucleosides other than deoxyadenosine, deoxyguanosine, thymidine, and deoxycytidine. A compound described in any of [1], [4], [5], [8] to

[59] .

[61] Selectively cleavable sites include deoxyuridine, bromodeoxyuridine, deoxyinosine, 8-hydroxydeoxyguanosine, 3-methyl-2'-deoxyadenosine, N6-etheno-2'-deoxyadenosine, 7-methyl-2'-deoxyguanosine, 2'-deoxyxanthosine, or 5,6-dihydroxy-5,6-dihydrodeoxythymidine. A compound described in any of [1], [4], [5], [8] to

[60] .

[62] The selectively cleavable site is deoxyuridine or deoxyinosine. A compound described in any of [1], [4], [5], [8] to

[61] .

[63] The selectively cleavable site is deoxyuridine. A compound described in any of [1], [4], [5], [8] to

[62] .

[64] The selectively cleavable site is deoxyinosine. A compound described in any of [1], [4], [5], [8] to

[62] .

[65] The selectively cleavable site is the second phosphodiester bond 3' from deoxyinosine, A compound described in any of [1], [4], [5], [8] to

[59] .

[66] The selectively cleavable site is a ribonucleoside. A compound described in any of [1], [4], [5], [8] to

[59] .

[67] There is one selectively severable site. A compound described in any of [1], [4], [5], [8] to

[66] .

[68] A cross-section comprising at least one cleavable portion in E or (LP1)p, and at least one cleavable portion in F or (LP2)q, A compound described in any of [1], [4], [5], [8] to

[66] .

[69] The compound according to

[68] , wherein the cleavable site in E or (LP1)p and the cleavable site in F or (LP2)q are cleavable under different conditions.

[70] A compound according to any of [8] to

[69] , wherein An is a substructure constructed of n building blocks α1 to αn (where n is an integer from 1 to 10).

[71] A compound according to any of [8] to

[70] , wherein An is a low molecular weight organic compound.

[72] A compound according to any one of [8] to

[71] , wherein the building block of An is a compound with a molecular weight of 500 or less.

[73] A compound according to any one of [8] to

[72] , wherein the building block of An is a compound with a molecular weight of 300 or less.

[74] A compound according to any one of [8] to

[73] , wherein the building block of An is a compound with a molecular weight of 150 or less.

[75] The compound according to any one of [8] to

[74] , wherein An is an organic compound composed of one or more elements selected from the group of elements consisting of H, B, C, N, O, Si, P, S, F, Cl, Br and I.

[76] The compound according to any one of [8] to

[75] , wherein An is a low molecular weight organic compound having substituents selected individually or differently from the group of substituents consisting of aryl groups, non-aromatic cyclyl groups, heteroaryl groups and non-aromatic heterocyclyl groups.

[77] A compound according to any of [8] to

[76] , wherein An has a molecular weight of 5000 or less.

[78] A compound according to any of [8] to

[77] , wherein An has a molecular weight of 800 or less.

[79] A compound according to any of [8] to

[78] , wherein An has a molecular weight of 500 or less.

[80] A compound according to any one of [8] to

[70] , wherein An is a polypeptide.

[81] A compound described in any of [8] to

[80] , wherein Sp is a bond.

[82] Sp is a bifunctional spacer, The two-functional spacer is SpD-SpL-SpX, SpD is a divalent group derived from a reactive group that can constitute a C-C, amino, ether, carbonyl, amide, ester, urea, sulfide, disulfide, sulfoxide, sulfonamide, or sulfonyl bond. SpL may be polyalkylene glycol, polyethylene, C1-20 aliphatic hydrocarbons which may optionally be replaced by heteroatoms, peptides, oligonucleotides, or combinations thereof. SpX is a divalent group derived from a reactive group that forms an amide, amino, or sulfonamide bond. A compound described in any of [8] to

[80] .

[83] Sp is a bifunctional spacer, The two-functional spacer is SpD-SpL-SpX, SpD is a divalent group derived from a primary amino group, SpL is polyethylene glycol or polyethylene. SpX is a divalent group derived from a carboxyl group. A compound described in any of [8] to

[81] .

[84] A compound according to any one of [8] to

[83] , wherein the oligonucleotide chain X and the oligonucleotide chain Y are in a sequence capable of forming a double helix.

[85] A compound according to any one of [8] to

[84] wherein oligonucleotide chain X and oligonucleotide chain Y have complementary base sequences.

[86] A compound according to any of [8] to

[85] , wherein oligonucleotide chain X and oligonucleotide chain Y are each 1 to 200 bases in length.

[87] A compound according to any of [8] to

[86] , wherein oligonucleotide chain X and oligonucleotide chain Y are each 3 to 150 bases in length.

[88] A compound according to any of [8] to

[87] , wherein oligonucleotide chain X and oligonucleotide chain Y are each 30 to 150 bases in length.

[89] A compound according to any one of [8] to

[88] , wherein the oligonucleotide chain X and the oligonucleotide chain Y have blunt ends.

[90] A compound according to any one of [8] to

[88] , wherein the oligonucleotide chain X and the oligonucleotide chain Y have protruding ends.

[91] The compound described in

[90] , wherein the protrusions of the protruding ends are 1 to 30 bases in length.

[92] The compound according to

[90] or

[91] , wherein the protrusions of the protruding ends are 2 to 5 bases in length.

[93] A compound according to any one of

[90] to

[92] , wherein oligonucleotide chain X and oligonucleotide chain Y have protruding ends, and a specific molecular identification sequence is further bound to the protruding ends.

[94] A compound according to any one of [8] to

[93] , wherein a functional molecule is attached to either X or Y.

[95] A compound according to any one of [8] to

[93] , wherein biotin is bound to either X or Y. A compound library containing any of the compounds listed in

[96] , [1], [4], [5], [8] to

[95] . A DNA encoding library containing any of the compounds described in

[97] [1], [4], [5], [8] to

[95] .

[98] The library described in

[96] or

[97] , comprising more than 1000 different compounds.

[99] A method for producing the compound An-Sp-C-Bn, An is a substructure constructed from n building blocks α1 to αn (where n is an integer from 2 to 10). Sp is a bond or a bifunctional spacer. C is a hairpin-shaped headpiece having at least one "selectively severable portion," Bn is a substructure containing an oligonucleotide with a nucleotide sequence that can identify the structure of An. For C, the following steps: (a) binding α1-Sp, or binding Sp and α1, and (b) Attaching an oligonucleotide tag containing a base sequence that can identify the structure of α1, To obtain compound A1-Sp-C-B1, Next, for A(m-1)-Sp-CB(m-1) (where m is an integer from 2 to n), the following steps (c) and (d) are repeated in ascending order from m from 2 to n; (c) Attaching αn to part A, and (d) Attach an oligonucleotide tag containing a base sequence that can identify the structure of αn to the end of portion B. This includes obtaining the compound Am-Sp-C-Bm, A method in which steps (a) and (b), and steps (c) and (d) can be performed in any order.

[0100] A method for producing An-Sp-C-Bn which is a compound described in any of [9] to

[95] , An is a substructure constructed from n building blocks α1 to αn (where n is an integer from 2 to 10). Sp is a bond or a bifunctional spacer, C is a hairpin-shaped headpiece having at least one "selectively severable portion," Bn is a substructure containing an oligonucleotide with a nucleotide sequence that can identify the structure of An. For C, the following steps: (a) binding α1-Sp, or binding Sp and α1, and (b) Attaching an oligonucleotide tag containing a base sequence that can identify the structure of α1, To obtain compound A1-Sp-C-B1, Next, for A(m-1)-Sp-CB(m-1) (where m is an integer from 2 to n), the following steps (c) and (d) are repeated in ascending order from m from 2 to n; (c) Attaching αn to part A, and (d) Attach an oligonucleotide tag containing a base sequence that can identify the structure of αn to the end of portion B. This includes obtaining the compound Am-Sp-C-Bm, A method in which steps (a) and (b), and steps (c) and (d) can be performed in any order.

[0101] A method for producing An-Sp-C-Bn (An, Sp, C, and Bn have the same meanings as above), which is a compound described in any of [9] to

[95] , For C, the following steps: (a) binding α1-Sp, or binding Sp and α1, and (b) Attaching an oligonucleotide tag containing a base sequence that can identify the structure of α1, To obtain compound A1-Sp-C-B1, Next, for A(m-1)-Sp-CB(m-1) (where m is an integer from 2 to n), the following steps (c) and (d) are repeated in ascending order from m from 2 to n; (c) Attaching αn to part A, and (d) Attach an oligonucleotide tag containing a base sequence that can identify the structure of αn to the end of portion B. This includes obtaining the compound Am-Sp-C-Bm, A method in which steps (a) and (b), and steps (c) and (d) can be performed in any order.

[0102] At least one equation (III) An-Sp-C-Bn (III) (In the formula, An is a substructure constructed from n building blocks α1 to αn (where n is an integer from 1 to 10). Sp is a bond or a bifunctional spacer, C is a hairpin-shaped headpiece having at least one "selectively severable portion," Bn is a substructure containing an oligonucleotide with a nucleotide sequence that can identify the structure of An. A method for evaluating a compound library containing compounds represented by the following steps: (1) The compound library is brought into contact with the biological target under conditions suitable for at least one library molecule of the compound library to bind to the target. (2) Remove library molecules that do not bind to the target and select library molecules that have affinity for the biological target. (3) Cut the parts that can be selectively cut. (4) Identify the sequence of oligonucleotides that make up Bn. (5) Using the sequence determined in (4), identify the structure of one or more compounds that bind to the biological target. A method consisting of the following.

[0103] At least one equation (III) An-Sp-C-Bn (III) (In the formula, An is a substructure constructed from n building blocks α1 to αn (where n is an integer from 1 to 10). Sp is a bond or a bifunctional spacer, C is a hairpin-shaped headpiece having at least one "selectively severable portion," Bn is a substructure containing an oligonucleotide with a nucleotide sequence that can identify the structure of An. A method for evaluating a compound library containing any of the compounds described in [8] to

[92] , the following steps: (1) The compound library is brought into contact with the biological target under conditions suitable for at least one library molecule of the compound library to bind to the target. (2) Remove library molecules that do not bind to the target and select library molecules that have affinity for the biological target. (3) Cut the parts that can be selectively cut. (4) Identify the sequence of oligonucleotides that make up Bn. (5) Using the sequence determined in (4), identify the structure of one or more compounds that bind to the biological target. A method consisting of the following.

[0104] The method according to

[0102] or

[0103] , further comprising the step of amplifying the oligonucleotide constituting Bn between steps (3) and (4).

[0105] The method according to any one of

[0102] to

[0104] , wherein the step of selectively cutting a site is a step of selectively cutting a site using an enzyme.

[0106] The step of selectively cutting the parts that can be cut is The method according to any one of

[0102] to

[0104] , which is a step of selectively cutting a site that can be cut by a combination of enzymes and changes in chemical conditions.

[0107] The method according to

[0105] or

[0106] , wherein the enzyme is selected from at least one of glycosylases and nucleases.

[0108] The method according to

[0107] , wherein the enzyme is uracil DNA glycosylase.

[0109] The method according to

[0107] , wherein the enzyme is endonuclease VIII.

[0110] The method according to

[0107] , wherein the enzyme is a combination of uracil DNA glycosylase and endonuclease VIII.

[0111] The method according to

[0107] , wherein the enzyme is alkyladenine DNA glycosylase.

[0112] The method according to

[0107] , wherein the enzyme is endonuclease V.

[0113] The method according to any one of

[0106] to

[0112] , wherein the change in chemical conditions is heating to 50-100°C in a water-containing solution.

[0114] The method according to any one of

[0103] to

[0113] , wherein the change in chemical conditions is heating to 80-95°C in a water-containing solution.

[0115] The method according to any of

[0106] to

[0114] , wherein the change in chemical conditions is a basic condition with a pH of 8 to 13.

[0116] The method according to any of

[0106] to

[0115] , wherein the change in chemical conditions is a basic condition with a pH of 8 to 11.

[0117] The method according to any of

[0106] to

[0116] , wherein the change in chemical conditions is a basic condition with a pH of 9 to 10.

[0118] The method according to any one of

[0102] to

[0117] , wherein a cleavable site is provided near the end of the DNA tag, the site is cleaved as desired to generate a new protruding end, a specific molecular identification sequence is ligated to the sticky end, and the sequence of oligonucleotides constituting Bn is identified.

[0119] The method according to

[0118] , wherein a cleavable site located near the end of the DNA tag and a cleavable site contained in C are cleaved under different conditions.

[0120] A method for utilizing nucleic acids that bind to compounds having cleavable sites and hairpin structures, by cleaving the cleavable sites to obtain double-stranded nucleic acids.

[0121] The method according to

[0120] , wherein a nucleic acid that is chemically more stable than double-stranded nucleic acid and binds to a compound having cleavable sites and a hairpin structure is used, and the cleavable sites are cleaved to utilize it as a double-stranded nucleic acid.

[0122] The method according to

[0120] or

[0121] , which involves using a nucleic acid that binds to a compound having cleavable sites and a hairpin structure, subjecting the compound to a chemical structure transformation, and then using it as a double-stranded nucleic acid by cleaving the cleavable sites.

[0123] The method according to any one of

[0120] to

[0122] , wherein a nucleic acid that binds to a compound having cleavable sites and a hairpin structure is used, and after chemically transforming the nucleic acid, the cleavable sites are cleaved to utilize it as a double-stranded nucleic acid.

[0124] A method according to any one of

[0120] to

[0123] , wherein a nucleic acid that binds to a compound having cleavable sites and a hairpin structure is used, and after subjecting the nucleic acid to a nucleic acid extension reaction, the cleavable sites are cleaved to utilize it as a double-stranded nucleic acid.

[0125] A method according to any one of

[0120] to

[0124] , comprising using a nucleic acid that binds to a compound having cleavable sites and a hairpin structure, cleaving the cleavable sites to make it usable as a double-stranded nucleic acid, and then performing a PCR reaction.

[0126] A method according to any of

[0120] to

[0125] , used for evaluating the functionality of a compound.

[0127] A method according to any of

[0120] to

[0126] for use in evaluating the biological activity of a compound.

[0128] The method used for DEL, as described in any of

[0120] to

[0127] .

[0129] A method using any of

[0120] to

[0124] for the manufacture of DEL.

[0130] A method for converting a DEL compound synthesized using nucleic acids that bind to compounds having cleavable sites and hairpin structures into a DEL compound having single-stranded DNA by cleaving the cleavable sites.

[0131] A method for synthesizing a DEL compound using nucleic acids that bind to compounds having cleavable sites and hairpin structures, by cleaving the cleavable sites, converting it into a DEL having single-stranded DNA, and forming a double helix with cross-linker-modified DNA.

[0132] A method for synthesizing a crosslinker-modified double-stranded DEL compound by cleaving a DEL compound synthesized using nucleic acids that bind to compounds having cleavable sites and hairpin structures, attaching a crosslinker-modified primer, and extending the attached primer. [Effects of the Invention]

[0009] This invention provides a DEL containing cleavable regions within a DNA strand, and a composition for its synthesis, enabling the production of DEL with greater convenience than conventional methods. [Brief explanation of the drawing]

[0010] [Figure 1] An exemplary DEL manufacturing method of Form 1 is shown. Starting with a headpiece containing a first oligonucleotide strand with a cleavable region in the DNA strand, a loop region, and a second oligonucleotide strand, the manufacturing of DEL is achieved by repeatedly binding building blocks and double-strand ligation of oligonucleotide tags corresponding to the building blocks (three times in Figure 1), and optionally double-strand ligation of oligonucleotide tags including primer regions. [Figure 2] An exemplary method for using DEL in Form 1 is shown. By using a cleavable region in the first oligonucleotide chain of the headpiece, and using a cleavage method such as an enzyme to cleave the cleavable region, the DEL can be converted into a double-stranded oligonucleotide that is not bound at the loop region, allowing for highly efficient PCR. [Figure 3] An exemplary method for using DEL in Form 2 is shown. For DEL containing cleavable sites in the second oligonucleotide chain of the headpiece, PCR can be performed with high efficiency by using a cleavage method such as an enzyme to cleave the cleavable sites and convert them into double-stranded oligonucleotides that are not bound at the loop site. [Figure 4]An exemplary method for using DEL in Form 3 is shown. For DEL containing cleavable sites in both the first and second oligonucleotide strands of the headpiece, PCR can be performed with high efficiency by using a cleavage method such as an enzyme to cleave both cleavable sites and convert them into double-stranded oligonucleotides that do not have a loop attached. [Figure 5] An exemplary use of the DEL of Form 4 is shown. For a DEL containing two different cleavable sites in the first and second oligonucleotide chains of the headpiece, the first or second oligonucleotide chain can be selectively cleaved by selecting the cleavage conditions. [Figure 6] An exemplary use of the DEL in Form 5 is shown. A cleavable site is provided near the end of the DNA tag, and a new overhang can be generated by cleaving this site as desired. This overhang can be used as an adhesive end to ligate a desired nucleic acid sequence, such as UMIs (Specific Molecular Identification Sequences), thereby conferring new functionality. [Figure 7] An exemplary use of DEL in Form 6 is shown. In this invention, a cleavable site can be used in combination with a modifying group or functional molecule, and for example, it is possible to prepare DEL in which hairpin strand DNA has been converted to single-stranded DNA. For example, a double-stranded oligonucleotide chain having a functional molecule (e.g., biotin) at its 3' end is ligated to the synthesized DEL compound (A), the cleavable site is cleaved (B), and a treatment according to the function of the functional molecule is applied (C). For example, if the functional molecule is biotin, streptavidin beads with biotin affinity are used to selectively remove the biotin-bound oligonucleotide chain from the system. This makes it possible to obtain DEL having single-stranded DNA. [Figure 8]An example of the use of DEL obtained in Form 6 is shown. DELs with single-stranded DNA obtained in Form 6 can be given new functions by forming a double helix with a modified oligonucleotide (e.g., crosslinker-modified DNA such as a photoreactive crosslinker) having a desired functional site. [Figure 9] An exemplary use of DEL in Form 7 is shown. In this invention, a crosslinker can be introduced by utilizing the cleavable site. A cleavable site can be cleaved from a synthesized DEL compound (A), a modifying primer can be attached (B), and a crosslinker-modified double-stranded DEL compound can be synthesized based on the attached primer (C). The crosslinker-modified double-stranded DEL compound can significantly improve the detection sensitivity in screening DEL libraries (see Non-Patent Documents 5, 6, etc.). [Figure 10] This graph shows the conversion rate of the cleavage reaction at each incubation time when the cleavage reaction of 10 hairpin-type DEL substructures containing deoxyuridine (U-DEL1-sh, U-DEL2-sh, U-DEL3-sh, U-DEL4-sh, U-DEL5-HP, U-DEL6-HP, U-DEL7-HP, U-DEL8-HP, U-DEL9-HP, and U-DEL10-HP) was verified using USER® enzyme in Example 1. [Figure 11] Examples 2, 3, 4, 5, and 7 show schematic diagrams illustrating the synthesis procedures for various hairpin DELs (U-DEL1, U-DEL2, U-DEL4, U-DEL7, U-DEL8, U-DEL9, U-DEL10, H-DEL, U-DEL5, U-DEL11, U-DEL12, U-DEL13, I-DEL1, I-DEL2, I-DEL3, R-DEL1, and BIO-DEL). Each hairpin DEL is synthesized using a corresponding headpiece as a raw material, through a two-step double-stranded ligation with double-stranded oligonucleotides Pr_TAG and CP. [Figure 12]This graph shows the Ct values ​​measured by real-time PCR for eight types of hairpin DELs (U-DEL1, U-DEL2, U-DEL4, U-DEL7, U-DEL8, U-DEL9, U-DEL10, and H-DEL) and double-stranded DELs (DS-DEL) in Example 2, broken down by sample volume. Samples treated with USER® enzyme are indicated as “USER(+)”, and untreated samples are indicated as “USER(-)”. The deoxyuridine-containing, cleavable hairpin DELs (U-DEL1, U-DEL2, U-DEL4, U-DEL7, U-DEL8, U-DEL9, and U-DEL10) show Ct values ​​equivalent to those of double-stranded DELs (DS-DEL) after USER® enzyme treatment. [Figure 13] The image shows a gel obtained by modified polyacrylamide gel electrophoresis, illustrating the progress of the cleavage reaction of six types of deoxyuridine-containing hairpin DELs (U-DEL5, U-DEL7, U-DEL9, U-DEL11, U-DEL12, and U-DEL13) by USER® enzyme in Example 3. The numbers in the figure indicate the lane numbers. [Figure 14] The image shows a gel obtained by denatured polyacrylamide gel electrophoresis, illustrating the progress of the cleavage reaction of hairpin DELs (I-DEL1, I-DEL2, I-DEL3, and I-DEL4) containing four types of deoxyinosine by endonuclease V in Example 4. The numbers in the figure indicate the lane numbers. [Figure 15] The image shows a gel obtained by denatured polyacrylamide gel electrophoresis, illustrating the progress of the cleavage reaction of hairpin DEL (R-DEL1) containing ribonucleoside by RNaseHII in Example 5. The numbers in the figure indicate the lane numbers. [Figure 16]This is a schematic diagram showing the synthesis procedure for a model library containing 3×3×3(27) compound species using U-DEL9-HP as a starting material. In Example 6, the model library is synthesized using U-DEL9-HP as a starting material through three split-and-pool steps (cycles A, B, and C). Each cycle includes a ligation reaction of a double-stranded oligonucleotide tag and a chemical reaction for introducing building blocks. [Figure 17] These are images of gels obtained by agarose gel electrophoresis, showing the progress of the ligation reaction in each cycle during the model library synthesis of Example 6. The numbers in the figures indicate the lane numbers. [Figure 18] Figure 18A shows the chromatograph obtained from the sample after cycle C was completed in the model library synthesis of Example 6. Figure 18B shows the deconvolution results of the MS spectrum obtained from the sample after cycle C was completed in the model library synthesis of Example 6. [Figure 19] This image shows the gel obtained by modified polyacrylamide gel electrophoresis, illustrating the progress of the cleavage reaction using the USER® enzyme from the model library in Example 6. The numbers in the figure indicate the lane numbers. [Figure 20] This image shows the gel obtained by modified polyacrylamide gel electrophoresis, illustrating the progress of the cleavage reaction of the DEL compound "BIO-DEL," which has biotin at its 3' end, by USER® enzyme in Example 7. The numbers in the figure indicate the lane numbers. [Figure 21] The image shows the results of a primer extension reaction performed in Example 7 using the single-stranded DNA-containing DEL compound "SS-DEL" and the photoreactive crosslinker-modified primer "PXL-Pr," obtained by polyacrylamide gel electrophoresis. The numbers in the figure indicate the lane numbers. [Modes for carrying out the invention]

[0011] As stated above, and as a concept well known to those skilled in the art, in this invention, a compound library refers to a systematic collection of compound derivatives, such as drug candidate compounds, that may possess specific activity. This compound library is often synthesized based on combinatorial chemistry synthesis techniques and methodologies. Combinatorial chemistry is the field of experimental methods and related research for efficiently synthesizing a wide variety of compounds from a series of compound libraries enumerated and designed based on combinatorial theory through systematic synthetic routes.

[0012] As mentioned above, and as is well known to those skilled in the art, one type of compound library based on combinatorial chemistry is the DNA-coding library. The DNA-coding library is often abbreviated as DEL. Furthermore, DEL is essentially synonymous with DNA-coding compound library. In this invention, a DNA-coding library means a library in which each compound in the library is tagged with a DNA tag. The DNA tag is sequenced to identify each structure of each compound and functions as a label for the compound.

[0013] A nucleotide is generally understood as a substance in which a phosphate group is bonded to a nucleoside. While nucleotides and nucleosides are well-known terms to those skilled in the art, a nucleoside is generally understood as a substance in which a nucleic acid base, such as a purine base or pyrimidine base, is glycosidically bonded to the 1-position of a sugar, such as a pentose. Nucleosides and nucleotides are also the constituent units of nucleic acids such as DNA and RNA. Furthermore, nucleic acids are a well-known concept to those skilled in the art, and are generally understood as polymers of nucleotides. In one embodiment, the nucleic acid of the present invention is a polymer composed of nucleotides and nucleic acid analogs, as described later.

[0014] Furthermore, in this specification, nucleic acid polymers composed of nucleotides and nucleic acid analogs, as well as nucleic acid monomers such as nucleotides and nucleic acid analogs, may also be simply referred to as nucleic acids. The latter usage is also in accordance with common technical knowledge and can be understood by those skilled in the art in accordance with the appropriate context.

[0015] In a broad sense, nucleotides include not only naturally occurring nucleotides (original nucleotides) but also artificial nucleotides (various nucleic acid analogs). The broad definition of nucleotide in this invention includes the following embodiments. (A) Natural nucleoside nucleotides (Examples of such nucleosides include adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxyuridine, deoxyguanosine, deoxycytidine, inosine, or diaminopurine deoxyriboside.) (B) Nucleoside nucleotides having a nucleic acid base analog. (Examples of nucleosides having the nucleic acid base analog include 2-aminoadenosine, 2-thiothymidine, pyrrolopyrimidine deoxyriboside, 3-methyladenosine, C5-propynylcytidine, C5-propynyluridine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-methylcytidine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, 6-O-methylguanosine, or 2-thiocytidine.) (C) Nucleotides having intercalated nucleic acid bases (D) Non-natural nucleotides having ribose or 2'-deoxyribose (E) Nucleotides having a modified sugar in the sugar portion (Examples of such modified sugars include modified ribose, modified 2'-deoxyribose, 2'-O-methylribose, 2'-fluororibose, D-threoninol, arabinose, hexose, anhydrohexitol, althritol, or mannitol.) (F) Nucleic acid analogs (Examples of such nucleic acid analogs include cyclohexanyl nucleic acid, cyclohexenyl nucleic acid, morpholino nucleic acid (PMO), locked nucleic acid (LNA), glycol nucleic acid (GNA), threose nucleic acid (TNA), serinol nucleic acid (SNA), acyclic threoninol nucleic acid (aTNA), or nucleic acids in which oxygen in ribose has been replaced.) The following provides a detailed explanation of each nucleic acid analog. (F1) PMO PMO is a nucleic acid analog that has a morpholine ring in the sugar region and an uncharged phosphorodiamidate structure in the phosphate diester region. (F2)LNA LNAs are nucleic acid analogs that have a cross-linking structure in the sugar portion. The most typical example is in which the 2'-hydroxyl of ribose is cross-linked to the 4'-carbon of the same ribose sugar by a C1-6 alkylene or C1-6 heteroalkylene. Examples of cross-linking structures include methylene, propylene, ether, or amino cross-linking structures. A typical example of a nucleotide nanoparticle (LNA) is 2',4'-BNA (2'-O,4'-C-methylated nucleic acid). (F3)GNA Glycol nucleic acids are also called GNAs. Examples include R-GNA and S-GNA. In these cases, ribose is replaced by glycol units bonded to a phosphodiester bond. (F4)TNA Treose nucleic acids are also called TNAs. In this case, ribose is replaced with α-L-treophranosyl-(3'→2'). (F5)SNA Serinol nucleic acids are also called SNAs. In this case, ribose is replaced by selinol units attached to a phosphodiester bond. (F6)aTNA Acyclic threoninol nucleic acid is also called aTNA. Examples include D-aTNA and L-aTNA. In this case, ribose is replaced by threoninol units bonded to a phosphodiester bond. (F7) Sugar in which oxygen in ribose has been replaced. Specific examples include substituted oxygen compounds with S, Se, or alkylenes (for example, methylene or ethylene). (G) Modified nucleotides (Examples of nucleotides with this modified skeleton include peptide nucleic acids (also known as PNAs; in this case, the 2-aminoethyl-glycine linkage replaces the ribose and phosphodiester skeletons).) (H) Nucleotides modified with a phosphate group (Examples of nucleotides modified with the phosphate group include phosphorothioates, 5'-N-phosphoamidites, phosphoroselenates, boranophosphates, boranophosphate esters, hydrogen phosphonates, phosphoramidates, phosphorodiamidates, alkyl or aryl phosphonates, phosphotryesters, cross-linked phosphoramidates, cross-linked phosphorothioates, or cross-linked methylene phosphonates.) In the following description, the oligonucleotide, oligonucleotide chain, double-stranded oligonucleotide, double-stranded oligonucleotide chain, and double-stranded DNA of the present invention are nucleotides as defined above.

[0016] In the present invention, when the term "nucleotide" is used without particular limitation, it means a natural nucleotide. Natural nucleotide is a term well known to those skilled in the art and is not particularly limited as long as it is a nucleotide that is essentially naturally occurring. In one embodiment, the natural nucleotide in the present invention is the nucleotide described in (A) above. (Nucleic acid analog) The term "nucleic acid analog" is well known to those skilled in the art, and the structure of the nucleic acid analog in this invention is not limited as long as it has the effects of the present invention. In one embodiment, a nucleic acid analog is a compound according to the embodiments of (B) to (H) described above. In one embodiment, the nucleic acid analog in the present invention is a compound having a phosphate-equivalent moiety and a hydroxyl group-equivalent moiety in a nucleic acid monomer. More preferably, the nucleic acid analog is a compound having a phosphate moiety and a hydroxyl group. In one embodiment, the nucleic acid analog in the present invention is a compound that can be used as a monomer in a nucleic acid synthesizer. As is well known to those skilled in the art, nucleic acid oligomers can be synthesized in a nucleic acid synthesizer by using a monomer in which the phosphate group (or equivalent site) of the nucleic acid analog is phosphoramiditeted and the hydroxyl group (or equivalent site) is protected with a protecting group. Furthermore, in nucleic acid analogs, substructures other than the phosphate group (or equivalent group) and the hydroxyl group (or equivalent group) can be called nucleic acid analog residues. The structure of nucleic acid analog residues is not limited as long as they have the effects of the present invention, but as a reference, if we examine the structural characteristics of natural nucleic acids (deoxyadenosine, thymidine, deoxycytidine, deoxyguanosine), we find that their molecular weights range from approximately 322 (thymidine monophosphate) to 347 (deoxyguanosine monophosphate), and the number of atoms between the hydroxyl group oxygen atom at the 3' position and the phosphorus atom at the 5' position that constitute the nucleic acid chain (including oxygen and phosphorus atoms; hereinafter also referred to as the number of atoms between residues) is 6. In addition, the following nucleic acid analogs are known to be usable in nucleic acid synthesizers. Amino C6 dT Molecular weight: 476, Number of residue atoms: 6 mdC(TEG-Amino) Molecular weight: 526, Number of residue atoms: 6 Uni-Link (trademark registered) Amino Modifier Molecular weight: 227, Number of atoms in residue: 6 (See Nucleic Acid Research, 1992, Vol. 20, pp. 6253-6259) d-Spacer Molecular weight: 198, Number of residue atoms: 6 Triethylene glycol phosphate (Spacer9) Molecular weight: 230, Number of atoms in residue: 11

[0017] For reference, the structures of each nucleic acid analog are listed below. [ka]

[0018] Therefore, in one embodiment, the nucleic acid analog is a compound (B1) characterized by the following: (B11) It has phosphoric acid (or an equivalent part) and a hydroxyl group (or an equivalent part). (B12) Composed of carbon, hydrogen, oxygen, nitrogen, phosphorus, or sulfur. (B13) The molecular weight is between 142 and 1500. (B14) The number of atoms between residues is 5 to 30. (B15) The bonding pattern between atoms in the residues is either all single bonds, or one or two double bonds with the remainder being single bonds.

[0019] In one embodiment, the nucleic acid analog is a compound (B2) characterized by the following: (B21) Contains phosphoric acid and hydroxyl groups. (B22) Composed of carbon, hydrogen, oxygen, nitrogen, or phosphorus. (B23) Molecular weight is between 142 and 1000. (B24) The number of atoms between residues is 5 to 20. (B25) The bonding mode between atoms in each residue is all single bonds.

[0020] In one embodiment, the nucleic acid analog is a compound (B3) characterized by the following: (B31) Contains phosphoric acid and hydroxyl groups. (B32) Composed of carbon, hydrogen, oxygen, nitrogen, or phosphorus. (B33) Molecular weight is between 142 and 700. (B34) The number of atoms between residues is 5 to 12. (B35) The bonding mode between atoms in each residue is all single bonds.

[0021] In one embodiment, the nucleic acid analog is one of the following compounds: (B41), (B42), (B43), (B44), (B5), (B51), or (B52). (B41)d-Spacer (B42) Amino C6 dT (B43)mdC(TEG-Amino) (B44) Uni-Link (Trademark Registered) Amino Modifier (B5) Polyalkylene glycol phosphate (B51) Diethylene glycol phosphate or triethylene glycol phosphate (B52) Triethylene glycol phosphate

[0022] In this invention, oligonucleotides and oligonucleotide chains mean polymers of nucleotides having one or more nucleotides at the 5' end, the 3' end, and the internal position between the 5' end and the 3' end.

[0023] Complementary base sequences refer to nucleotide sequences in nucleic acids that can form complementary base pairs, which are determined by hydrogen bonds between two oligonucleotides, such as adenine and thymine (or uracil), or guanine and cytosine. The formation of complementary base pairs is also called hybridization. Complementary base pairs are generally referred to as "Watson-Crick base pairs" or "natural base pairs." However, base pairs may be Watson-Crick type, Hoogsteen type base pairs, or base pairs formed by the formation of other hydrogen bonding motifs (e.g., diaminopurine and T, 5-methyl C and G, 2-thiothymine and A, 6-hydroxypurine and C, pseudoisocytosine and G). There are no restrictions on the sequences of "mutually complementary base sequences" as long as the two oligonucleotides can form a double helix and are usable for the purposes of the present invention, and there are no restrictions on the homology of the two sequences. Preferably, the homology is 99% or more, 98% or more, 95% or more, 90% or more, 85% or more, 80% or more, 70% or more, 60% or more, or 50% or more, in order of increasing preference.

[0024] To reiterate, in this invention, hybridization refers to the act of forming a double helix with oligonucleotides or oligonucleotide chains containing complementary base sequences, and the phenomenon of oligonucleotides or oligonucleotide chains containing complementary sequences forming a double helix.

[0025] In this invention, a double helix refers to a state in which two nucleic acid strands form complementary base pairs (hybridize). The two nucleic acid strands may originate from two separate nucleic acid strands, or from two nucleic acid sequences within a single nucleic acid strand molecule.

[0026] In this invention, a double-stranded oligonucleotide and a double-stranded oligonucleotide chain refer to a secondary structure formed by the hybridization of two or more different oligonucleotide chains. The two oligonucleotides may have different chain lengths and may have regions that are not hybridized. Furthermore, the region where two strands hybridize is a double helix.

[0027] In this invention, double-stranded DNA refers to a secondary structure formed by the hybridization of two different DNA strands. The lengths of the respective DNA strands may differ, and they may have regions that are not hybridized. The DNA strands are not limited to naturally occurring deoxyribonucleotides, but refer to all oligonucleotide strands that can be amplified by DNA polymerase.

[0028] In this invention, "forming a double helix" means that the nucleic acid forms a double helix under standard conditions for handling oligonucleotides, such as a temperature of 4 to 40°C, an aqueous solvent, and a pH of 4 to 10. For example, even if a double helix does not form under certain solvents or conditions, if the nucleic acid forms a double helix under standard conditions, then the nucleic acid is considered a double-helix-forming nucleic acid.

[0029] In this invention, the Tm value refers to the temperature at which half of the DNA molecules anneal with the complementary strand.

[0030] In this invention, a blunt end means that the ends of a double-stranded oligonucleotide are paired without either end protruding.

[0031] In this invention, a protruding end means that one of the ends of a double-stranded oligonucleotide has a protruding portion. The protruding portion of the protruding end can be of any length, but is preferably 1 to 50 bases, more preferably 1 to 30 bases, even more preferably 1 to 15 bases, and most preferably 2 to 6 bases. In certain embodiments, the protruding portion can be used as a hybridized region when performing ligation of sticky ends.

[0032] PCR stands for Polymerase Chain Reaction. PCR is a method for amplifying oligonucleotide chains and is a well-known technique to those skilled in the art. In general terms, the PCR process involves (1) dissociating the double-stranded oligonucleotide chain to be amplified into two single strands by heat treatment, and (2) adjusting the temperature to a level suitable for the enzymatic reaction, then synthesizing complementary strands to each single strand using an enzyme (such as DNA polymerase) present in the reaction system. In other words, one double-stranded oligonucleotide can be amplified into two. By repeating processes (1) and (2) through temperature control, oligonucleotide chains can be amplified with high efficiency in PCR.

[0033] In this invention, "primer" refers to an oligonucleotide that can be annealed to a template oligonucleotide chain and extended by polymerase in a template-dependent manner.

[0034] In this invention, the PCR primer sequence refers to the sequence of the portion of the oligonucleotide chain that the primer anneals to, and is preferably a PCR-suitable sequence known in the art, and is preferably located at the end of the oligonucleotide chain.

[0035] In this invention, a "nick" refers to a region in a double-stranded oligonucleotide chain where internucleotide bonds are lacking, resulting in a break in the oligonucleotide chain. The 5' end of this missing region may or may not have a phosphate group.

[0036] In this invention, a gap refers to a region in a double-stranded oligonucleotide chain where one or more consecutive nucleotides are deleted, causing the oligonucleotide chain to separate. The 5' end of the deleted region may or may not have a phosphate group.

[0037] In this invention, a hairpin strand is a single-stranded structure in which two complementary nucleic acid strands are linked together, and the characteristics of hairpin strands and hairpin strand DEL are as described above. The terms "hairpin region," "hairpin structure," and "hairpin type" used in this invention are understood to be terms derived from the hairpin, which is the same concept as the aforementioned "hairpin strand."

[0038] In this invention, nucleic acid ligation and nucleic acid linking reactions refer to reactions that link the ends of nucleic acids together.

[0039] Enzymatic nucleic acid ligation refers to a reaction in which the ends of nucleic acids are linked together using enzymes.

[0040] Enzymes that can be used in nucleic acid ligation reactions include, for example, DNA ligase, RNA ligase, DNA polymerase, RNA polymerase, or topoisomerase.

[0041] In one aspect, a DNA ligase is an enzyme that connects the ends of DNA strands with phosphate diester bonds. In another aspect, a DNA ligase is understood as a ligase belonging to EC number 6.5.1.1 or 6.5.1.2. DNA ligases are also called polydeoxyribonucleotide synthases or polynucleotide ligases. Examples of DNA ligases include DNA ligases I, II, III, IV, and T4 DNA ligase.

[0042] In one aspect, RNA ligases are enzymes that connect the ends of RNA chains with phosphate diester bonds. In another aspect, RNA ligases are understood as ligases belonging to EC number 6.5.1.3. Also, in another aspect, RNA ligases belong to the poly(ribonucleotide):poly(ribonucleotide) ligase family. RNA ligases are also called polyribonucleotide synthases or polyribonucleotide ligases.

[0043] In this invention, chemical ligation refers to a reaction that joins the ends of nucleic acids without the use of enzymes.

[0044] In chemical ligation, a linkage is formed when the ends of nucleic acids, which have functional groups that are paired for the chemical reaction, react with each other. The functional groups that pair for chemical reactions include, for example, pairs of an optionally substituted alkynyl group and an optionally substituted azide group; pairs of an optionally substituted diene having a 4π electron system (e.g., an optionally substituted 1,3-unsaturated compound, e.g., optionally substituted 1,3-butadiene, 1-methoxy-3-trimethylsilyloxy-1,3-butadiene, cyclopentadiene, cyclohexadiene, or furan) and an optionally substituted dienophile or optionally substituted heterodienophile having a 2π electron system (e.g., an optionally substituted alkenyl group or an optionally substituted alkynyl group); pairs of an optionally substituted amino group and a carboxylic acid group; pairs of a phosphorothioate group and an iodo group (e.g., a 3'-terminal phosphorothioate group and a 5'-terminal iodo group); or pairs of a phosphate group and a hydroxyl group (e.g., a 5'-terminal phosphate group and a 3'-terminal hydroxyl group, or a 5'-terminal hydroxyl group and a 3'-terminal phosphate group). Chemical ligation is a concept well known to those skilled in the art, and those skilled in the art can achieve chemical ligation appropriately based on common technical knowledge. See also Artificial DNA; PNA&XNA, 2014, Vol. 5, e27896, Current Opinion in Chemical Biology, 2015, Vol. 26, pp. 80-88.

[0045] In this invention, "selectively cleavable" means that, in a given compound, only a specific site can be selectively cleaved under predetermined conditions without altering the rest of the molecular structure of the compound.

[0046] In this invention, "selectively cleavable site" means a site in a compound that can be selectively cleaved under predetermined conditions.

[0047] In one embodiment, a preferred structure for the "selectively cleavable site" in the present invention is a "selectively cleavable nucleic acid." This site may be composed of multiple nucleic acids and be cleaved by a specific sequence, or it may be composed of a single nucleic acid. When the cleavable site is a nucleic acid, it is preferable from the viewpoint that (1) established manufacturing methods such as nucleic acid synthesizers can be used, resulting in good manufacturing efficiency, and (2) since the reaction conditions for constructing the building blocks of DEL require that the nucleic acid in the DNA tag portion not be degraded, if the cleavable site is a nucleic acid, it will also not be degraded.

[0048] A more preferred structure for the aforementioned "selectively cleavable nucleic acid" is a nucleic acid containing nucleotides not included in the DNA tag sequence of DEL. If the cleavable sites are nucleotides not included in the DNA tag sequence, it can be used without limiting the DNA tag sequence in order to avoid cleaving the DNA tag portion.

[0049] Deoxyadenosine, deoxyguanosine, thymidine, and deoxycytidine are preferred nucleic acids for use in DNA tag sequences. Therefore, the preferred structure for the selectively cleavable site is a nucleic acid that is neither deoxyadenosine, deoxyguanosine, thymidine, nor deoxycytidine.

[0050] An example of a "selectively cleavable site" is a "nucleotide containing a cleavable base." For example, in DEL, the N-glycosidic bond between the base and sugar portions of a "nucleotide containing a cleavable base" is cleaved by the action of DNA glycosylase, leaving a debasic site. The phosphodiester bond adjacent to the debasic site is cleaved by changes in chemical conditions (e.g., increased temperature, basic hydrolysis, etc.) or by enzymes with depurine / depyrimidine (AP) endonuclease activity or AP lyase activity (e.g., endonuclease III, endonuclease IV, endonuclease V, endonuclease VI, endonuclease VII, endonuclease VIII, APE1 (human-derived AP endonuclease), Fpg (formamidopyridine-DNA glycosylase), etc.), forming a one-base gap or nick.

[0051] Examples of "nucleotides with cleavable bases" include deoxyuridine, bromodeoxyuridine, deoxyinosine, 8-hydroxydeoxyguanosine, 3-methyl-2'-deoxyadenosine, N6-etheno-2'-deoxyadenosine, 7-methyl-2'-deoxyguanosine, 2'-deoxyxanthosine, and 5,6-dihydroxydeoxythymidine. Other nucleotides with cleavable bases are obvious to those skilled in the art. By incorporating these "nucleotides with cleavable bases" into DEL and using a DNA glycosylase that specifically recognizes their structure, the DEL can be selectively debased.

[0052] In this invention, DNA glycosylase refers to any enzyme having glycosylase activity that recognizes any nucleic acid base portion in an oligonucleotide, cleaves the N-glycosidic bond between the base portion and the sugar portion, and creates a debase site. Examples include uracil DNA glycosylase (recognizes deoxyuridine), alkyladenine DNA glycosylase (recognizes 3-methyl-2'-deoxyadenosine, 7-methyl-2'-deoxyguanosine, and deoxyinosine), Fpg (recognizes 8-hydroxydeoxyguanosine), endonuclease VIII (recognizes degraded pyrimidine bases such as 5,6-dihydroxydeoxythymidine and uracil glycol), and SUMG1 (abbreviation for single-strand selective uracil DNA glycosylase, which recognizes deoxyuridine).

[0053] In the present invention, more preferred examples of "selectively cleavable sites" include deoxyinosine and deoxyuridine.

[0054] In the present invention, a particularly preferred example of a "selectively cleavable site" is deoxyuridine.

[0055] In one embodiment, the "selectively cleavable site" in the present invention is preferably cleaved using an enzyme. Enzymes are generally preferred because they have high substrate specificity and do not recognize the DNA tag portion of DEL or the compound portion constructed from multiple building blocks as substrates, but only recognize and act on the "selectively cleavable site". Alternatively, cleavage using the enzyme may be achieved by first structurally altering the "selectively cleavable site" with an enzyme, and then changing the chemical conditions. Examples of such enzymes include glycosylases and nucleases.

[0056] In this invention, glycosylase is an enzyme that has the function of hydrolyzing glycosidic bonds (covalent bonds formed by dehydration condensation between a sugar molecule and another organic compound). Among these, DNA glycosylase, as mentioned above, is an enzyme that recognizes the nucleic acid base portion in oligonucleotides and hydrolyzes its glycosidic bonds.

[0057] In this invention, a nuclease is an enzyme that has the function of hydrolyzing the phosphodiester bond between the sugar and phosphate of nucleic acids. Nucleases include, for example, AP endonuclease, nickel endonuclease, and ribonuclease.

[0058] As described above, AP endonuclease cleaves phosphodiester bonds adjacent to debasement sites generated by the action of any DNA glycosylase. Therefore, in the present invention, it is preferable to use DNA glycosylase and AP endonuclease in combination.

[0059] Nicking endonucleases (e.g., Nb.BbvCI, Nb.BsmI, Nb.BsrDI, etc.) recognize specific DNA sequences and produce nicks in which the phosphodiester bond is cleaved on only one strand of the double helix. Endonuclease V can also produce nicks in which the second phosphodiester bond is cleaved in the 3' direction from deoxyinosine, and is useful in carrying out the present invention.

[0060] Ribonucleases are enzymes that degrade RNA. In this invention, ribonucleosides are used as "selectively cleavable sites," and can be utilized by acting with ribonucleases. RNaseHII, a type of ribonuclease, can create nicks by cleaving the phosphodiester bond at the 5' end of ribonucleotides incorporated into the DNA sequence, and is useful in carrying out this invention.

[0061] In this invention, USER (registered trademark) means "Uracil-Specific Excision Reagent" Enzyme. USER is an endonuclease cocktail that removes uracil, containing uracil DNA glycosylase (UDG) and endonuclease VIII. USER removes uracil from double-stranded DNA, creating a single-base gap and cleaving the DNA strand. In the USER process, UDG first removes the uracil base to create a debase site. Subsequently, the endonuclease breaks down the phosphodiester bond, releasing deoxyribose without a base and creating a single-base gap. In this specification, USER® Enzyme and USER® Enzyme refer to USER® as defined above.

[0062] In this invention, a building block is a portion having a functional group that can constitute a part of a compound, and may be in the form of a compound.

[0063] In this invention, a nucleotide sequence that can identify each building block refers to a specific nucleotide sequence designed to correspond to the structure of each building block. Designing a sequence means assigning a nucleic acid nucleotide sequence to each structure, for example, nucleic acid nucleotide sequence AAA to building block structure A, nucleic acid nucleotide sequence TTT to structure B, and nucleic acid nucleotide sequence CGC to structure C. Sequences can be freely designed insofar as the objective of this invention is achieved. For example, any number of nucleotide sequences can be assigned to a single building block.

[0064] In this invention, an oligonucleotide tag is a substructure that includes an oligonucleotide containing a base sequence capable of identifying the structure of a substructure constructed by building blocks. In this invention, an oligonucleotide tag may be an oligonucleotide corresponding to each building block, or it may be a longer-chain oligonucleotide containing oligonucleotides corresponding to multiple building blocks. The nucleotides constituting the oligonucleotide tag of the present invention are not limited as long as they achieve the effects of the present invention, but it is desirable that they be nucleotides suitable for amplification by PCR and analysis by sequencer, in terms of ease of these operations. Examples of such preferred nucleotides include nucleotides having the aforementioned natural nucleic acid base as the base part and the aforementioned ribose or 2'-deoxyribose as the sugar part, and more preferred examples include deoxyadenosine, thymidine, deoxycytidine, or deoxyguanosine.

[0065] (Headpiece) In this invention, "headpiece" refers to a starting compound for the production of a compound library such as DEL. The structure of the headpiece of this invention is not limited insofar as it achieves the objectives of the invention, but in its most typical embodiment, it has at least one site to which building blocks can be attached, at least one site to which oligonucleotide tags can be attached, and further includes at least one selectively cleavable site in the structure. As described below, the DNA tag is preferably a double-stranded oligonucleotide chain, and there are preferably two sites where the oligonucleotide tag can be attached.

[0066] In one embodiment, the headpiece is a compound shown in the schematic diagram below. [ka]

[0067] In one aspect, it is desirable that the headpiece be chemically stable. Furthermore, in one embodiment, it is preferable that the headpiece has a structure that allows the DNA tag and building block to be placed in appropriate spaces. In one embodiment, it is preferable that the headpiece has a moderate degree of flexibility. Here, we will further explain appropriate spatial arrangement and flexibility (structural characteristics of the headpiece). Note that the structural characteristics of the headpiece described here may be achieved by the headpiece alone, or by combining the headpiece with a bifunctional spacer. In one embodiment, a preferred structural characteristic of the headpiece is one in which the headpiece and DNA tag do not inhibit the formation reaction of the building block, and conversely, the headpiece and building block do not inhibit the extension reaction of the DNA tag. In one embodiment, a preferred structural characteristic of the headpiece is one in which the headpiece or DNA tag portion does not affect the interaction between the building block compound (library compound) and the target (target protein, etc.). In one embodiment, a preferred structural characteristic of the headpiece is one in which the DNA tag and the building block region are oriented on opposite sides (for example, more than 90 degrees opposite). In one embodiment, a preferred structural characteristic of the headpiece is that the loop portion of the headpiece and the building block are separated by a few atoms to more than ten atoms in terms of the organic compound skeleton. In one embodiment, it is preferable that the headpiece has a moderate affinity for the DNA tag portion and the building block portion. Moderate affinity means, for example, chemical reactivity and stability that allow the bonds to be formed, maintained, and cleaved under desired conditions in order to carry out the present invention. In this invention, a bifunctional spacer means a spacer portion having at least two reactive groups that enable bonding between the building block portion and the headpiece.

[0068] In this description of the present invention, the terms "headpiece," "headpiece compound," and "compound for headpiece" refer to compounds of the same concept. In the description of the present invention, the "compound used as a headpiece" can be essentially understood in the same way as "use of the compound as a headpiece" from the perspective of use, and can be essentially understood in the same way as "method of using the compound as a headpiece" from the perspective of method. The same applies to the compound library.

[0069] Hereinafter, the structure of the preferred headpiece will be described, but the structure of the headpiece is not limited as long as the effects of the present invention are achieved.

[0070] As one aspect, the headpiece has a reactive functional group having at least one site that can be directly linked to a building block (D) or indirectly linked via a bifunctional spacer, (L) a linker extending from the reactive functional group, (E) a first oligonucleotide chain having one binding site that can be linked to one strand of an oligonucleotide tag, (F) a second oligonucleotide chain having one binding site that can be linked to the other strand of the oligonucleotide tag, and (LP) a loop site that binds to the linker and the two oligonucleotide chains, and is composed of at least one of the sites of E, F, or LP has at least one selectively cleavable site.

[0071] As one aspect, the headpiece is a compound represented by the following formula (I).

Chemical formula

[0072] In the present invention, a partial structure of a site that binds to a linker among loop sites may be referred to as a linking site or (LS). In the present invention, E-LP-F may be collectively referred to as a hairpin site.

[0073] (First and second oligonucleotide strands) Preferred embodiments of the first oligonucleotide strand (E) and the second oligonucleotide strand (F) will be described below.

[0074] The first oligonucleotide strand (E) and the second oligonucleotide strand (F) preferably form a double strand in the molecule via a loop site (LP), and the headpiece forms a hairpin structure. The preferred chain length for forming a double strand in the molecule is 3 bases or more, more preferably 4 bases or more, and even more preferably 6 bases or more. The chain lengths of E and F are each 3 to 40 in one embodiment. The chain lengths of E and F are each 4 to 40 in one embodiment. The chain lengths of E and F are each 6 to 25 in one embodiment.

[0075] The site to which the oligonucleotide tag is ligated is preferably structured to be suitable for enzymatic or chemical ligation. In one embodiment, the ligation of the headpiece and the oligonucleotide tag is carried out by double-strand ligation using an enzyme. In this case, it is preferable that the first and second oligonucleotide chains form protruding ends for ligation. The chain length of the protruding end is preferably 2 bases or more, more preferably 2 to 10 bases, and even more preferably 2 to 5 bases. Therefore, it is preferable that one of the first and second oligonucleotide chains is longer than the other by the length of the protruding end. Furthermore, for ligation by DNA ligase, it is preferable that the 5' end of the chain having the 5' end of the headpiece is phosphorylated.

[0076] Furthermore, the first and second oligonucleotide chains may contain part or all of the primer binding sequence for PCR. An appropriate chain length for the primer binding sequence is 17 to 25 bases.

[0077] (Linker) The preferred embodiment of the linker (L) is described below. As described above, the linker is a site that extends from the reactive functional group and bonds to the linking site. Typically, the linker is a divalent group (-L-) derived from the following embodiments.

[0078] In one embodiment, the linker is of the following form (L1): (L1) A C1-20 aliphatic hydrocarbon which may have substituents and may be replaced by 1-3 heteroatoms, or (2) A C6-14 aromatic hydrocarbon which may have substituents.

[0079] Other embodiments of L include the following embodiments (L2), (L3), (L4), or (L5). (L2) A C1-6 aliphatic hydrocarbon which may have substituents, a C1-6 aliphatic hydrocarbon which may be replaced by one or two oxygen atoms, or a C6-10 aromatic hydrocarbon which may have substituents. (L3) A C1-6 aliphatic hydrocarbon substituted with substituent group ST1, or a benzene substituted with substituent group ST1. Here, substituent group ST1 is composed of C1-6 alkyl groups, C1-6 alkoxy groups, fluorine atoms, and chlorine atoms. However, when substituent group ST1 is substituted with an aliphatic hydrocarbon, alkyl groups are not selected from substituent group ST1. (L4) Benzenes that are C1-6 alkyl or unsubstituted, or substituted with one or two C1-3 alkyl or C1-3 alkoxy groups. (L5) C1-6 alkyl groups.

[0080] (Reactive functional group) The following describes preferred embodiments of the reactive functional group (D). As described above, the reactive functional group has at least one site that can be directly linked to the building block or indirectly linked via a bifunctional spacer, and is a site that bonds to the linker group. Typically, the reactive functional group becomes a monovalent group (D-) in the headpiece and a "divalent group derived from the reactive functional group" (-D-) in the DEL. For example, if D is an amino group, the specific structure of (D-) is (R-HN-) (where R is a substituent as described below). For example, it reacts with an activated carboxyl group, a reactive sulfonyl group, or an isocyanate group to form an amide bond, a sulfonamide bond, or a urea bond, respectively. In this case, the specific structure of (-D-) becomes (-NR-). R is not limited insofar as the effects of the present invention are achieved, but in the following embodiments (D1) to (D5), R is preferably (1) a hydrogen atom, or (2) a C1-6 alkyl group that is unsubstituted or substituted with one to three substituents selected individually or differently from the group of substituents consisting of C1-6 alkoxy groups, fluorine atoms, and chlorine atoms. R is more preferably a hydrogen atom or a C1-3 alkyl group, and still more preferably a hydrogen atom. Also, for example, when (D-) is a methylene group having a leaving group (X-), the specific structure of (D-) is (X-CH2-), and it reacts with a nucleophile such as an amino group, a hydroxy group, or a thiol group to form a carbon-nitrogen bond, a carbon-oxygen bond, or a carbon-sulfur bond. In that case, the specific structure of (-D-) becomes (-CH2-). Also, for example, when (D-) is an aldehyde group, the specific structure of (D-) is (HOC-). The aldehyde group forms a carbon-nitrogen bond, for example, by a reductive amination reaction with an amino group, and in that case (-D-) becomes -CH2-, forms a carbon-carbon double bond, for example, by a reaction with a phosphorus ylide group, and in that case (-D-) becomes -CH=, forms a carbon-carbon triple bond, for example, by a reaction with an α-diazo phosphonate group, and in that case (-D-) becomes -C≡.

[0081] In one aspect, the site (D-) is in the following aspect (D1). (D1) A functional group capable of constituting a C-C, amino, ether, carbonyl, amide, ester, urea, sulfide, disulfide, sulfoxide, sulfonamide, or sulfonyl bond. (Literally, in this case, (-D-) becomes a C-C, amino, ether, carbonyl, amide, ester, urea, sulfide, disulfide, sulfoxide, sulfonamide, or sulfonyl bond.)

[0082] In other aspects, (D-) is in the following aspects (D2), (D3), (D4), or (D5). (D2) A C1 hydrocarbon having a leaving group, an amino group, a hydroxy group, a precursor of a carbonyl group, a thiol group, or an aldehyde group. Note that in this case, (-D-) can be, for example, -(C1 hydrocarbon)-, -NR-, -O-, -(C=O)-, -S-, -CH2-, -CH=, or -C≡, etc. (D3) C1 hydrocarbons having halogen atoms, C1 hydrocarbons having sulfonic acid leaving groups, amino groups, hydroxyl groups, carboxyl groups, halogenated carboxyl groups, thiol groups, or aldehyde groups. In this case, (-D-) could be -(C1 hydrocarbon)-, -NR-, -O-, -(C=O)-, -S-, -CH2-, -CH=, or -C≡, etc. (D4) -CH2Cl, -CH2Br, -CH2OSO2CH3, -CH2OSO2CF3, amino group, hydroxyl group, or carboxyl group. In this case, (-D-) will be -CH2-, -NR-, -O-, or -(C=O)-, respectively. (D5) Primary amino group. In this case, (-D-) becomes -NH-.

[0083] The preferred configuration of the loop portion (LP) is described below. The loop portion (LP) is preferably designed so that the first oligonucleotide chain (E) and the second oligonucleotide chain (F) form a double helix within the molecule, allowing the headpiece to form a hairpin structure. In other words, the loop portion (LP) is preferably designed to have a chain length and bond flexibility that makes the loop structure thermodynamically stable. Therefore, in one form, the loop section (LP) is as follows: LP, This is the loop region represented by (LP1)p-LS-(LP2)q, LS is a substructure selected from the group of compounds described in (A) to (C) below, (A) Nucleotides (B) Nucleic acid analogs (C) Trivalent C1-14 groups which may have substituents LP1 is a substructure selected individually or differently from the group of compounds described in (1) and (2) below, (1) Nucleotides (2) Nucleic acid analogs LP2 is each of q substructures selected individually or differently from the group of compounds described in (1) and (2) below. (1) Nucleotides (2) Nucleic acid analogs The total number of p and q ranges from 0 to 40.

[0084] A more preferred embodiment of the loop portion is as described above. The following provides further details about the structure of the loop section.

[0085] Here, nucleotides are the natural nucleotides described above, and nucleic acid analogs are as described above.

[0086] Here, LP1 is each of p substructures selected individually or differently from the group of compounds described in (1) and (2) below, and LP2 is each of q substructures selected individually or differently from the group of compounds described in (1) and (2) below. (1) Nucleotides (2) Nucleic acid analogs To select p compounds individually or distinctly means, for example, if p is 4, then LP1 can be selected individually or distinctly from the compound groups listed in (1) and (2), such as AATG, ATCG, TC(d-Spacer)G, or A(d-Spacer)(d-Spacer)C. The same applies to LP2.

[0087] Furthermore, the loop region may contain part or all of the primer binding sequence for PCR.

[0088] (Regarding LS) In one embodiment, LS is (A) a nucleotide or (B) a nucleic acid analog. When LS is (A) a nucleotide or (B) a nucleic acid analog, the loop region becomes a nucleic acid oligomer. The nucleic acid oligomer of the present invention is an oligomer formed by linking nucleotides or nucleic acid analogs as monomers. An oligomer can also be called a chain-like compound. Therefore, the nucleic acid oligomer of the present invention is any of the following: an oligonucleotide chain, a nucleic acid analog chain, or a mixed chain of nucleotides and nucleic acid analogs.

[0089] If the LS is (A) a nucleotide or (B) a nucleic acid analog, the loop region becomes a nucleic acid oligomer. In that case, the headpiece can be manufactured using a nucleic acid synthesizer, which is significantly preferable in practice.

[0090] When the LS is (A) a nucleotide or (B) a nucleic acid analog, in the manufacture of the headpiece, in one embodiment, a monomer for nucleic acid synthesis is prepared in which the linker site (L) and the reactive functional group site (D) are bound to the LS, and then the nucleic acid oligomer is synthesized. Examples of such nucleic acid synthesis monomers include the aforementioned Amino C6 dT, mdC (TEG-Amino), and Uni-Link (trademark registered) Amino Modifier. In this embodiment, for example, in the structure of the monomer mdC(TEG-Amino), the nucleotide portion corresponds to the linking site (LS), and the side chain portion extending from the base corresponds to the linker site (L) and the reactive functional group site (D). In preparation, the reactive functional group (D) may be protected with a protecting group.

[0091] In that case, one possible embodiment is the following compound (B6): (B6) A compound in which the (-LD) is bonded to the base portion of a nucleotide.

[0092] In one embodiment, the nucleic acid analog is one of the following compounds: (B61), (B62), (B63), (B64), or (B65). (B61)(-LD) is (-L1-D1) (B6) (B62)(-LD) is equal to (-L2-D2) (B6). (B63)(-LD) is (-L3-D3) (B6). (B64)(-LD) is (-L4-D4) (B6). A compound described in any of (B61) to (B64), wherein (B65)(-D) is (-D5).

[0093] When LS is (A) a nucleotide or (B) a nucleic acid analog, in the manufacture of the headpiece, one embodiment is to first synthesize the nucleic acid oligomer and then attach the linker site (L) and the reactive functional group site (D). In that case, it is preferable to include the "specific nucleic acid analog" to which the linker site binds as the linking site (LS) within the hairpin site (nucleic acid analog oligomer). Examples of such "specific nucleic acid analogs" include the aforementioned Amino C6 dT, mdC (TEG-Amino), and Uni-Link (trademark registered) Amino Modifier. In this embodiment, for example, mdC(TEG-Amino) itself corresponds to the linking site (LS), and the addition sites to which further bond from the base side chain correspond to the linker site (L) and the reactive functional group site (D).

[0094] (Regarding p and q) As described above, the chain length of the loop portion is preferably such that the first oligonucleotide chain (E) and the second oligonucleotide chain (F) form a double helix within the molecule, and the headpiece forms a hairpin structure. In one possible scenario, the total number of p and q ranges from 1 to 40. In one scenario, the total number of p and q is between 2 and 20. In one possible scenario, the total number of p and q is between 2 and 10. In one possible scenario, the total number of p and q is between 2 and 7.

[0095] In one embodiment, the loop portion of the present invention is (A) Nucleotides and consists of the following nucleic acid analogs (B41), (B42), (B43), (B44), or (B52). (B41)d-Spacer (B42) Amino C6 dT (B43)mdC(TEG-Amino) (B44) Uni-Link (Trademark Registered) Amino Modifier (B52) Triethylene glycol phosphate

[0096] In one embodiment, LS is preferably B42, B43, or B44. In another embodiment, LP1 and LP2 are preferably A, B41, or B52.

[0097] In one embodiment, the loop region is a nucleic acid oligomer with the sequences described in (X1) to (X9) below. (X1) A-B41-B42-B41-A (X2) A-B41-B43-B41-A (X3) A-B41-B44-B41-A (X4)B41-B41-B42-B41-B41 (X5)B41-B41-B43-B41-B41 (X6)B41-B41-B44-B41-B41 (X7)B52-B42-B52 (X8)B52-B43-B52 (X9)A52-A44-A52

[0098] In the aforementioned headpiece, the number of cuttable parts is preferably five or less, and more preferably one to two.

[0099] In the headpiece described above, if there are two or more cleavable sites, it is preferable that at least one cleavable site is located in the first oligonucleotide chain or between the first oligonucleotide chain and the linker binding site, and at least one cleavable site is located in the second oligonucleotide chain or between the second oligonucleotide chain and the linker binding site.

[0100] In one embodiment, in the headpiece described above, the location of the severable portion is preferably within 20 bases, more preferably within 10 bases, and even more preferably within 3 bases, starting from the binding portion between the loop portion and the first oligonucleotide chain or the second oligonucleotide chain.

[0101] Just to clarify, the preferred embodiment of the "selectively severable region" and preferred embodiments such as E, F, or LP are distinct concepts. That is, even if the location of the "selectively severable region" is included in E, the preferred embodiment of E does not necessarily apply to the "selectively severable region."

[0102] In one embodiment, the compound constituting DEL of the present invention is a compound represented by the following formula (II). [ka] (In the formula, X and Y are nucleotide chains, E and F are independent of each other. It is an oligomer composed of nucleotides or nucleic acid analogs, However, E and F contain complementary base sequences and form a double-stranded oligonucleotide. LP is the loop section, L is a linker, D is a divalent group derived from a reactive functional group, Sp is a bond or a bifunctional spacer, An is a substructure composed of at least one building block. It is a compound represented by the following: X and Y have sequences that can form a double helix in at least part of their structure. X binds to E at its 5' end. Y binds to F at its 3' end. A compound having at least one selectively cleavable site at any one of the sites E, F, or LP.

[0103] In one embodiment, preferred embodiments of E, F, LP, L, and D in the compound represented by formula (II) above are the same as preferred embodiments of E, F, LP, L, and D described with respect to formula (I) above. Preferred embodiments of X, Y, Sp, and An will be described separately.

[0104] (2 Functional Spacers) As described above, a bifunctional spacer is a spacer portion having at least two reactive groups that enable the bonding of a substructure An of the compound library to the headpiece. In one embodiment, the bifunctional spacer is SpD-SpL-SpX. SpX is a reactive group that forms a covalent bond with the reactive functional group of the headpiece. SpD is a reactive group that forms a covalent bond with the substructure An in the compound library. SpL is a chemically inert spacing portion. Furthermore, similar to the reactive functional group (D), the reactive group (SpX) is a monovalent group (-SpX) in the difunctional spacer alone (the reagent state before bonding with the headpiece), and in DEL (the state bonded with the headpiece), it becomes a "divalent group derived from the reactive group" (-SpX-) based on the aforementioned (-SpX). Similarly, the reactive group (SpD) is a monovalent group (SpD-) before bonding with An, and in DEL (the state bonded with An), it becomes a "divalent group derived from the reactive group" (-SpD-) based on the aforementioned (SpD-).

[0105] A preferred embodiment of SpX is a reactive group that forms an amino, carbonyl, amide, ester, urea, or sulfonamide bond. In one embodiment, SpX is a suitable reactive group when the reactive functional group of the headpiece is an amino group, and has the following structures: (SpX1), (SpX2), or (SpX3). (SpX1): Carboxylate group, halogenated carboxylate group, aldehyde group, or halogenated sulfonyl group (SpX2): Carboxylate group or halogenated sulfonyl group (SpX3): Carboxy group

[0106] The preferred embodiment of SpD is the same as that of D described above. In one embodiment, SpD is one of the aforementioned (D1), (D2), (D3), (D4), or (D5).

[0107] A preferred embodiment of SpL is as follows: In one embodiment, SpL is the aforementioned (L1), (L2), (L3), (L4), or (L5). In one form, SpL is (SpL1), (SpL2), or (SpL3). (SpL1) Polyalkylene glycol, polyethylene, C1-20 aliphatic hydrocarbons, peptides, oligonucleotides, or combinations thereof, which may optionally be replaced by heteroatoms. (SpL2) Polyalkylene glycol, polyethylene, C1-10 aliphatic hydrocarbons, or peptides (SpL3) Polyethylene glycol, or polyethylene

[0108] One embodiment of a bifunctional spacer is as follows: (Sp1):(D4)-(SpL1)-(SpX1) (Sp2):(D4)-(SpL2)-(SpX2) (Sp3):(D4)-(SpL3)-(SpX3) (Sp4):(D5)-(SpL1)-(SpX1) (Sp5):(D5)-(SpL2)-(SpX2) (Sp6):(D5)-(SpL3)-(SpX3)

[0109] In one embodiment, the (Sp-DL) portion of the compound constituting DEL is configured as follows: (SpDL1), (SpDL2), (SpDL3), (SpDL4), (SpDL5), (SpDL6), (SpDL7), (SpDL8), (SpDL9), or (SpDL10). (SpDL1):(D4)-(L1) (SpDL2):(D5)-(L1) (SpDL3):(D4)-(L2) (SpDL4):(D5)-(L2) (SpDL5):(Sp1)-(D5)-(L5) (SpDL6):(Sp2)-(D5)-(L5) (SpDL7):(Sp3)-(D5)-(L5) (SpDL8):(Sp4)-(D5)-(L5) (SpDL9):(Sp5)-(D5)-(L5) (SpDL10):(Sp6)-(D5)-(L5) In (SpDL1), (SpDL2), (SpDL3), and (SpDL4), Sp represents a connection.

[0110] In carrying out the present invention, it is advantageous that the headpiece can be synthesized in a nucleic acid synthesis apparatus. In such a carrying out, as described above, in one embodiment, a nucleic acid synthesis monomer in which the linker moiety (L) and the reactive functional group moiety (D) are bound to LS can be prepared, and then a nucleic acid oligomer can be synthesized. Examples of such nucleic acid synthesis monomers include the aforementioned Amino C6 dT, mdC (TEG-Amino), and Uni-Link (trademark registered) Amino Modifier. On the other hand, when using commercially available nucleic acid synthesis monomers or nucleic acid analogs usable with nucleic acid synthesizers as described above, the length of the linker region may be limited. In such cases, one embodiment is to introduce an appropriate bifunctional spacer, which makes it possible to adjust the distance between the headpiece and An, and is advantageous in carrying out the invention.

[0111] In the description of this invention, terms such as "C1-C6 alkyl group" and "C1-6 alkyl group" mean that the number of carbon atoms is 1 to 6. Similarly, when m and n are integers, the description "Cm-Cn" and "Cm-n" means that the number of carbon atoms is m to n. Therefore, "C1-C6 alkyl group" and "C1-6 alkyl group" mean an alkyl group having 1 to 6 carbon atoms, and "C1-C6 alkylene" and "C1-6 alkylene" mean an alkylene having 1 to 6 carbon atoms.

[0112] In this invention, "C1-6 alkyl" refers to a linear or branched alkyl group having 1 to 6 carbon atoms. Specific examples include methyl, ethyl, propyl, isopropyl, butyl, isobutyl, sec-butyl, tert-butyl, pentyl, and hexyl.

[0113] In this invention, "C1-3 alkyl" refers to a linear or branched alkyl group having 1 to 3 carbon atoms. Specific examples include methyl, ethyl, propyl, and isopropyl alkyl groups.

[0114] In this invention, "C1-6 alkoxy" refers to a linear or branched alkoxy having 1 to 6 carbon atoms. Specific examples include methoxy, ethoxy, propoxy, isopropoxy, butoxy, isobutoxy, sec-butoxy, tert-butoxy, pentyloxy, and hexyloxy.

[0115] In this invention, "C1-3 alkoxy" refers to a linear or branched alkoxy having one to three carbon atoms. Specific examples include methoxy, ethoxy, propoxy, and isopropoxy.

[0116] In this invention, "hydrocarbon" means a chain, branched chain, or cyclic saturated or unsaturated compound composed solely of carbon atoms and hydrogen atoms.

[0117] In this invention, "aliphatic hydrocarbon" means a hydrocarbon that is non-aromatic. "Aliphatic hydrocarbon" may be linear, branched, or cyclic, and may be saturated or unsaturated. Specific examples of structures include alkyl, alkenyl, alkynyl, cycloalkyl, or cycloalkenyl structures, or combinations thereof. In this invention, "C1-20 aliphatic hydrocarbons" means aliphatic hydrocarbons having 1 to 20 carbon atoms. In this invention, "C1-10 aliphatic hydrocarbons" means aliphatic hydrocarbons having 1 to 10 carbon atoms. In this invention, "C1-6 aliphatic hydrocarbons" means aliphatic hydrocarbons having 1 to 6 carbon atoms.

[0118] In this invention, "aromatic hydrocarbons" refers to hydrocarbons that are aromatic. In this invention, "C6-14 aromatic hydrocarbons" refers to aromatic hydrocarbons having 6 to 14 carbon atoms. Specific examples include benzene, naphthalene, and anthracene. In this invention, "C6-10 aromatic hydrocarbons" refers to aromatic hydrocarbons having 6 to 10 carbon atoms. Specific examples include benzene and naphthalene.

[0119] The aromatic heterocycle of the present invention is an aromatic heterocycle having, as heteroatoms within its ring structure, elements selected individually or differently from the group consisting of nitrogen, oxygen, and sulfur. In one embodiment, an aromatic heterocycle is a "C1-9 aromatic heterocycle" having 1 to 9 carbon atoms, and in another embodiment, a "C1-9 aromatic heterocycle" is an aromatic heterocycle with 5 to 10 members. In one aspect, an aromatic heterocycle is a "C1-5 aromatic heterocycle" having 1 to 5 carbon atoms, and in another aspect, a "C1-5 aromatic heterocycle" is an aromatic heterocycle with 5 to 6 members. In one embodiment, an aromatic heterocycle is a "C2-9 aromatic heterocycle" having 2 to 9 carbon atoms, and in another embodiment, a "C2-9 aromatic heterocycle" is an aromatic heterocycle with 5 to 10 members. In one aspect, an aromatic heterocycle is a "C2-5 aromatic heterocycle" having 2-5 carbon atoms, and in another aspect, a "C2-5 aromatic heterocycle" is an aromatic heterocycle with 5-6 members.

[0120] The nitrogen-containing aromatic heterocycle of the present invention is an aromatic heterocycle having nitrogen as a heteroatom within its ring structure. In one embodiment, a nitrogen-containing aromatic heterocycle is a "C1-5 nitrogen-containing aromatic heterocycle" having 1 to 5 carbon atoms, and in another embodiment, a "C1-5 nitrogen-containing aromatic heterocycle" is an aromatic heterocycle with 5 to 6 members. In one embodiment, a nitrogen-containing aromatic heterocycle is a "C2-5 nitrogen-containing aromatic heterocycle" having 2-5 carbon atoms, and in another embodiment, a "C2-5 nitrogen-containing aromatic heterocycle" is an aromatic heterocycle with 5-6 members.

[0121] The non-aromatic heterocycle of the present invention is a non-aromatic heterocycle having, as heteroatoms within its ring structure, elements selected individually or differently from the group consisting of nitrogen, oxygen, and sulfur. Non-aromatic heterocycles may contain partially unsaturated bonds. In one embodiment, a non-aromatic heterocycle is a "C2-9 non-aromatic heterocycle" having 2 to 9 carbon atoms, and in another embodiment, a "C2-9 non-aromatic heterocycle" is a "5-10 member non-aromatic heterocycle".

[0122] In this invention, "trivalent group of C1-14" means a trivalent group derived from a compound having 1 to 14 carbon atoms. The structure is not limited as long as the effects of the present invention are achieved.

[0123] In this invention, if it is stated that "it may be replaced by a heteroatom," a heteroatom means an atom other than carbon and hydrogen. The heteroatom is preferably an oxygen atom, a nitrogen atom, a silicon atom, a phosphorus atom, or a sulfur atom, and more preferably an oxygen atom, a nitrogen atom, or a sulfur atom. Therefore, taking propyl (-CH2-CH2-CH3) as an example of a hydrocarbon, the concept of "propyl that may be replaced by heteroatoms" refers to a structure that includes ethers ((-CH2-O-CH3) or (-O-CH2-CH3)) in which the methylene (-CH2-) in the alkyl group is replaced by oxygen, or amines ((-CH2-NH-CH3) or (-NH-CH2-CH3)) in which it is replaced by nitrogen.

[0124] If the present invention states that "substituents may be present," those substituents are not limited as long as they achieve the objectives of the present invention. The substituent is preferably a C1-6 alkyl group, a C1-6 alkoxy group, an amino group, a hydroxyl group, a nitro group, a cyano group, an oxo group, or a halogen atom. The substituent is more preferably a C1-6 alkyl group, a C1-6 alkoxy group, a fluorine atom, or a chlorine atom.

[0125] In this invention, polypeptides and peptides refer to compounds or substructures formed by the linking of amino acids. Amino acids are a general term for organic compounds that have both an amino group and a carboxyl group. The amino acids constituting the polypeptides and peptides of this invention are not particularly limited and include modified amino acids, etc. Following common usage in the field of life sciences, proline (classified as an imino acid) is also included as an amino acid in this invention. The amino acids constituting the polypeptides and peptides of this invention are preferably α-amino acids, and more preferably "amino acids that constitute proteins."

[0126] The halogen atoms of this invention include fluorine atoms, chlorine atoms, bromine atoms, and iodine atoms.

[0127] C-C, amino, ether, carbonyl, amide, ester, urea, sulfide, disulfide, sulfoxide, sulfonamide, and sulfonyl bonds are chemical bonds having chemical structures understood by their respective names. Those skilled in the art will understand, for example, that ether bonds are generally represented as "-O-", and carbonyl bonds are generally represented as "-C(=O)-". Amino, amide, and urea bonds have a hydrogen atom or other substituent on the nitrogen atom, but the structure on the nitrogen atom is not limited as long as it has the effect of the present invention. The substituent on the nitrogen atom is preferably a C1-6 alkyl group or a hydrogen atom, and more preferably a hydrogen atom. It goes without saying that CC bonds mean carbon-carbon bonds. CC bonds include single bonds, double bonds, and triple bonds. In one embodiment, in steps a and / or c of the manufacturing method of the present invention, a bond appropriately selected from the above 11 types is constructed. These 11 types of bonds are particularly basic bonding modes in organic chemistry, and the reactions for constructing them are well known to those skilled in the art. Therefore, in designing and constructing the partial structure An of the compound library of the present invention, those skilled in the art can use these 11 types of bonds in appropriate combinations.

[0128] An organic compound composed of elements selected individually or differently from the group of elements consisting of H, B, C, N, O, Si, P, S, F, Cl, Br, and I is an organic compound constructed by the bonding of the aforementioned 12 elements.

[0129] In one embodiment, the partial structure An of the compound library of the present invention is constructed from the 12 elements listed above. These 12 elements are particularly fundamental elements in organic compounds, and the reactions for constructing them are well known to those skilled in the art. Therefore, when designing and constructing the partial structure An of the compound library of the present invention, those skilled in the art can use these 12 elements in appropriate combinations.

[0130] Low molecular weight organic compounds having substituents selected individually or differently from the group consisting of aryl groups, non-aromatic cyclyl groups, heteroaryl groups, and non-aromatic heterocyclyl groups are low molecular weight organic compounds having a chemical structure understood by each name. Low molecular weight compounds are a concept well known to those skilled in the art, and examples of preferred molecular weights of low molecular weight compounds in the present invention will be mentioned separately.

[0131] The aryl group of the present invention is preferably a C6-10 aryl group, and more preferably a phenyl group.

[0132] The non-aromatic cyclyl group of the present invention is preferably a 5- to 8-membered non-aromatic cyclyl group, and more preferably a 5- or 6-membered non-aromatic cyclyl group. The non-aromatic cyclyl group may contain a partially unsaturated bond.

[0133] The heteroaryl group and non-aromatic heterocyclyl group of the present invention are groups having, as heteroatoms within the ring structure, elements selected individually or differently from the group consisting of nitrogen, oxygen, and sulfur. The heteroaryl group and non-aromatic heterocyclyl group of the present invention are preferably 5-membered to 8-membered groups, more preferably 5-membered or 6-membered groups, and the non-aromatic heterocyclyl group may contain a partially unsaturated bond.

[0134] In one embodiment, the partial structure An of the compound library of the present invention has the four groups described above. These four groups are particularly fundamental partial structures in organic compounds, and the reactions for constructing them in compounds are well known to those skilled in the art. Therefore, when designing and constructing the partial structure An of the compound library of the present invention, those skilled in the art can use these four groups in appropriate combinations.

[0135] The aforementioned preferred embodiments, namely compound libraries constructed with 11 types of bonds, 12 types of elements, and / or 4 types of groups, have particular core value. Those skilled in the art will understand that compound libraries constructed without these preferred embodiments generally have limited applications and, in many cases, limited commercial value.

[0136] The synthesis history of An refers to a record of all operations performed until An is synthesized, and in particular, the structure and sequence of building blocks used until An is synthesized. For example, when a reaction is carried out in two or more separate reaction vessels, each using different building blocks and / or different reaction conditions, the synthesis history is imprinted as sequence information of the oligonucleotide by ligating an oligonucleotide chain with a predetermined sequence to the product in each reaction vessel before or after the reaction. By repeating this operation until An is constructed, an oligonucleotide of Bn with the synthesis history of An is constructed.

[0137] Split-and-pool synthesis is a synthetic method developed by Geysen et al. during the early stages of combinatorial chemistry as a combinatorial chemical method for constructing peptide libraries using solid-phase synthesis. Split-and-pool synthesis is also known as the split-mix method.

[0138] Following the above process, let's explain using the synthesis of a peptide library using solid-phase synthesis as an example. In split-and-pool synthesis, at each peptide end-adding step, instead of cutting the sample from the solid-phase support to which the amino acids are bonded, N types of support are first mixed and homogenized, then divided equally to add the next N types of amino acids.

[0139] In other words, one type of peptide chain is generated for each carrier, and by applying all 20 natural amino acids at each stage, it becomes possible to construct a peptide library that allows for all possible combinations of peptides of a specific length.

[0140] If this peptide library is to be screened for antigen presentation or receptor binding, assays can be performed using peptides on a solid-phase support using methods such as ELISA. In other words, there is no need to cleave the sample peptide from the support; the support particles that react to the assay are picked up (for example, fluorescently labeled support particles of about 0.1 mm are picked up with an optical microscope). Then, the peptides in these particles can be analyzed using an instrumental analyzer (such as a peptide analyzer) to determine the target peptide sequence, or other combinatorial chemical identification methods (such as tagging) can be used to indirectly determine candidate peptide sequences for screening.

[0141] Furthermore, in the manufacturing method of the present invention, we will explain as an example the case in which v types of structures are synthesized when m is 2, and w types of structures are synthesized when m is 3, using split-and-pool synthesis. In this explanation, the process is repeated in the order of (c) and (d). (m=2) In the m=2 step, α2 is added to A1-Sp-C-B1 in step (c) and β2 in step (d), respectively, to produce A2-Sp-C-B2. Here, v types of α2 (α2(av)) structures and their corresponding v types of β2 (β2(av)) are prepared, and steps (c) and (d) are performed for each structure, respectively, to obtain v types of A2-Sp-C-B2 (A2(a)-Sp-C-B2(a), A2(b)-Sp-C-B2(b)...A2(v)-Sp-C-B2(v): i.e., A2(av)-Sp-C-B2(av)). In split-and-pool synthesis, the v types of A2-Sp-C-B2 are mixed and then divided into w parts. More specifically, division means dividing the mixture into w reaction vessels. (m=3) In step m=3, α3 is added to A2-Sp-C-B2 in step (c) and β3 in step (d), respectively, to produce A3-Sp-C-B3. Here, we prepare w types of α3 (α3(aw)) structures and w types of corresponding β3 (β2(aw)), and perform steps (c) and (d) on w (A2(av)-Sp-C-B2(av) mixtures). Then, through steps n=2 and n=3, (v×w) types of A3-Sp-C-B3 can be efficiently synthesized in (v+w) synthesis steps.

[0142] (Biological assessment) When the w products obtained are mixed, a mixture of (v × w) types of A3-Sp-C-B3 compound libraries is obtained. For example, by performing a drug receptor binding test on this mixture, (v × w) types of compounds can be screened in a single step. By washing away the compounds that did not bind to the drug receptor, only the bound compounds can be isolated. In DEL as in the present invention, the DNA of the isolated A3-Sp-C-B3 compound is amplified to a sequence-decipherable amount, and the structure of A3 can be determined from the sequence information.

[0143] Furthermore, terms such as compound library, building block, and split-and-pool are well-known to those skilled in the art in fields such as combinatorial chemistry, and can be used as appropriate by referring to the following literature, etc. (1) Takashi Takahashi, Takayuki Doi, "Combinatorial Chemistry," Journal of the Society of Synthetic Organic Chemistry, 2002, Vol. 60, pp. 426-433. (2) Combinatorial Chemistry Research Group (ed.), "Combinatorial Chemistry", Kagaku Dojin

[0144] A DNA-coding library (or DEL) is a compound library consisting of a group of compounds labeled with DNA or oligonucleotides that have substantially equivalent function to DNA (DNA-coding compounds). Through the split-and-pool synthesis described above, the labeled DNA is imbued with the structure or synthesis history of each compound as sequence information. Due to these characteristics, a DNA-coding library is 10 2 ~10 20By screening a mixture of various compounds and identifying the DNA sequences contained in the obtained compounds using methods known in the art (e.g., the use of next-generation sequencers and / or microarrays), it becomes possible to identify the structure of the compounds. One aspect of the screening method may involve contacting a target such as a protein with a DNA-coding library and selecting compounds that bind to the target.

[0145] While "biological target" is a term well known to those skilled in the art, in one aspect, in the present invention, "biological target" refers to a group of biological substances that can be targeted in the development of drugs such as pharmaceuticals and agrochemicals, and includes, for example, enzymes (e.g., kinases, phosphatases, methylases, demethylases, proteases, and DNA repair enzymes), proteins: proteins involved in protein-protein interactions (e.g., receptor ligands), receptor targets (e.g., GPCRs), ion channels, cells, bacteria, viruses, parasites, DNA, RNA, prions, or carbohydrates. "Biological activity evaluation" is a term well known to those skilled in the art, but in one aspect, in the present invention, "biological activity evaluation" means evaluating the presence or absence, or strength, of the biological activity (for example, the ability to bind to a biological target, the function of inhibiting enzyme activity, the function of promoting enzyme activity, etc.) of a compound. For specific examples of biological activity evaluation, please refer to the aforementioned Patent Documents 2 and 3, Non-Patent Documents 1 to 6, etc. While "functional evaluation" is a term well known to those skilled in the art, in one aspect, in the present invention, "functional evaluation" refers to the evaluation of the presence or absence, or strength, of a specific function (e.g., binding ability, biological activity, luminescence properties, etc.) of a compound.

[0146] The present invention provides several methods with several advantages regarding DEL and methods for producing DEL by using DNA strands having cleavable sites. Forms 1 to 7 are described in detail below.

[0147] Form 1 The present invention provides a DEL using the aforementioned "hairpin-shaped headpiece having a severable portion".

[0148] As illustrated in Figure 1, in Form 1, DEL is produced by starting with a headpiece containing a first oligonucleotide strand with a cleavable region in the DNA strand, a loop region, and a second oligonucleotide strand, repeatedly performing the binding of building blocks and double-strand ligation of oligonucleotide tags corresponding to the building blocks (three times in Figure 1), and optionally performing double-strand ligation of oligonucleotide tags including a primer region.

[0149] As illustrated in Figure 2, in Form 1, PCR can be performed with high efficiency by using a cleavable site in the first oligonucleotide chain of the headpiece to cleave the cleavable site using a cleavage method such as an enzyme, thereby converting it into a double-stranded oligonucleotide that is not bound at the loop site.

[0150] (Regarding Form 2) As illustrated in Figure 3, in DEL using a "hairpin-shaped headpiece with a severable portion," the severable portion may be located on the second oligonucleotide chain. The features of Form 2 are the same as those of Form 1, except for the severable portion.

[0151] (Regarding Form 3) As illustrated in Figure 4, in DEL using a "hairpin-shaped headpiece with cleavable regions," the cleavable regions may be present on both the first and second oligonucleotide chains. In this embodiment, it is expected that PCR efficiency will be further improved by cleaving the loop region from both oligonucleotide chains.

[0152] (Regarding Form 4) As illustrated in Figure 5, in the present invention, cleavable sites may be present in both the first oligonucleotide chain (E) and the second oligonucleotide chain (F), and the structures of the cleavable sites may be different. In such cases, the cleavage sites can be controlled by utilizing the differences in the characteristics of the two (or more) cleavable sites. For example, deoxyuridine may be used as the cleavable site in the first oligonucleotide chain (E), and deoxyinosine may be used as the cleavable site in the second oligonucleotide chain (F). In this case, the USER enzyme can selectively cleave the deoxyuridine in the first oligonucleotide chain (E). [ka] On the other hand, using alkyladenine DNA glycosylase and endonuclease VIII, it is possible to selectively cleave the deoxyinosine-initiated cleavage site in the second oligonucleotide chain (F). [ka] In this way, by selecting the cutting site as desired, a wider range of DEL modifications becomes possible, and a wider range of evaluation methods can be applied thereafter. This can be expected.

[0153] (Regarding Form 5) As illustrated in Figure 6, the present invention also allows for the provision of a cleavable region in the DNA tag portion (e.g., oligonucleotide chain (Y)). By providing a cleavable region near the end of the DNA tag and cleaving the region as desired, a new protruding end can be generated. [ka] The protruding end can be used as an adhesive end to ligate a desired nucleic acid sequence, such as UMIs (Urban Misidentification Sequences). [ka] After biological evaluation, the selected DEL compounds are given UMIs regions as described above, and DNA sequencing is performed, enabling analysis with reduced amplification bias from PCR. Thus, the present invention provides unprecedented performance in the manufacturing and use of DEL compounds by having a site that can selectively cleave the nucleic acid sequence.

[0154] Here, UMIs (Specific Molecular Identifiers) are molecular identifiers that, when attached to DNA contained in a sample, assign a unique DNA sequence to each individual DNA molecule (see Nature Methods, 2012, Vol. 9, pp. 72-74). By attaching such molecular identifiers before PCR amplification, it becomes possible to identify PCR duplication (sequences originating from the same molecule) when quantifying the number of DNA molecules containing a specific sequence in a sample, thereby enabling quantification with reduced PCR amplification bias.

[0155] (Regarding Form 6) As illustrated in Figure 7, the present invention allows for the use of a combination of cleavable sites and modifying groups or functional molecules, making it possible, for example, to prepare DEL in which hairpin strand DNA has been converted to single-stranded DNA. As shown in Figure 7, a DEL compound using a headpiece with a severable portion in section E is given as an example. (Step A) A double-stranded oligonucleotide chain having a solid-supported, removable modifying group (e.g., biotin) at its 3' end is ligated to the synthesized DEL compound. (Step B) Cut off the parts that can be cut. (Step C) Apply a treatment according to the function of the modifying group. For example, in the case of biotin, use streptavidin beads with biotin affinity to selectively remove the biotin-bound oligonucleotide chain from the system. This makes it possible to obtain DEL containing single-stranded DNA.

[0156] Here, a functional molecule is a molecule that possesses a specific chemical or biological function (e.g., solubility, photoreactivity, substrate-specific reactivity, target protein degradation induction properties), and by conferring this function to DEL, it becomes possible to evaluate and purify DEL according to its function.

[0157] Here, "biotin" refers to all biotin compounds that bind to avidin, including not only vitamin B7 but also, for example, desthiobiotin.

[0158] As illustrated in Figure 8, DELs having single-stranded DNA can be given new functions by forming a double helix with a modified oligonucleotide (e.g., crosslinker-modified DNA such as a photoreactive crosslinker) that has a desired functional site.

[0159] (Regarding Form 7) As illustrated in Figure 9, the present invention allows for the introduction of a crosslinker by utilizing a severable portion. As shown in Figure 9, a DEL compound using a headpiece with a severable portion in section E is given as an example. (Step A) Cut the cleavable parts of the synthesized DEL compound. (Step B) Apply a modified primer having the desired functional site (e.g., a crosslinker-modified primer such as a photoreactive crosslinker). (Step C) The applied primers are extended to synthesize a crosslinker-modified double-stranded DEL compound. In DEL evaluation, crosslinker-modified double-stranded DEL compounds can significantly improve detection sensitivity by further binding the crosslink structure to the target protein when the building block compound (library small molecule compound) binds to the target protein (see Non-Patent Documents 5, 6, etc.). In the practical application of DEL technology, which evaluates a very large number of library compounds, enhancing the affinity of library compounds and improving detection sensitivity is extremely useful. This invention provides a novel and highly efficient method for producing crosslinker-modified double-stranded DEL compounds, and is extremely useful.

[0160] The present invention will be described in more detail below with reference to examples, but the present invention is not limited to these examples. The nucleic acids of various sequences in the examples can be prepared, for example, by an automated nucleic acid synthesizer according to standard procedures. An example of an automated nucleic acid synthesizer is the nS-8II (manufactured by GeneDesign). Furthermore, contract synthesis or contract laboratories can be used for nucleic acid preparation. Examples of well-known contract laboratories to those skilled in the art include GeneDesign and LGC Biosearch Technologies. Generally... These contract laboratories prepare nucleic acids with sequences specified by the client, under confidentiality agreements, and deliver them to the client.

[0161] Example 1 [Verification of the cleavage reaction of a hairpin-shaped DEL substructure containing deoxyuridine using USER® enzyme] Compounds with the sequences shown in Table 1 were prepared using the nucleic acid automated synthesizer nS-8II (manufactured by Gene Design). As will be apparent to those skilled in the art, in the sequence notation in Table 1, each sequence unit is connected by a phosphate diester bond, where "A" means deoxyadenosine, "T" means thymidine, "G" means deoxyguanosine, "C" means deoxycytidine, "(dU)" means deoxyuridine, "(p)" means phosphate, and "(amino-C6-dT)" is represented by the following formula (1) [ka] This refers to a modified nucleic acid represented by the following formula (2): "(amino-NC6-dT)" is given by the formula (2) below. [ka] This refers to modified nucleic acids represented by the following formula (3): [ka] It refers to the group represented by the following formula (4) [ka] This refers to the group represented by . In addition, amino-NC6-dT was synthesized according to the method described in (Journal of the American Chemical Society, 1993, Vol. 115, pp. 7128-7134) and is given by the following formula (5) [ka] The nucleic acid was introduced using nucleic acid synthesis reagents.

[0162] In Table 1, "No." in the left column represents the sequence number, and "Seq." in the right column represents the sequence. The left side of the sequence represents the 5' end, and the right side represents the 3' end. The names of the compounds corresponding to each sequence number (No.) are as follows. No.1: U-DEL1-sh No.2: U-DEL2-sh No.3: U-DEL3-sh No.4: U-DEL4-sh No. 5: U-DEL5-HP No. 6: U-DEL6-HP No.7: U-DEL7-HP No.8: U-DEL8-HP No.9:U-DEL9-HP No.10:U-DEL10-HP

[0163] [Table 1]

[0164] A 0.1 mM aqueous solution of each compound with the sequence shown in Table 1 was prepared, and the cleavage reaction with USER® enzyme was investigated using the following procedure.

[0165] A 1 μL 0.1 mM aqueous solution of the compound with the sequence shown in Table 1, 10 μL of CutSmart® Buffer (New England BioLabs, catalog number B7204S), and 79 μL of deionized water were added to a PCR tube. 10 μL of USER® enzyme (New England BioLabs, catalog number M5505S) was added to the solution, and the resulting solution was incubated at 37°C.

[0166] 20 μL samples were taken of each reaction solution 1 hour and 3 hours after the start of incubation. U-DEL1-sh, U-DEL5-HP, U-DEL6-HP, U-DEL7-HP, U-DEL8-HP, U-DEL9-HP, and U-DEL10-HP were also sampled at 20 hours. U-DEL8-HP and U-DEL9-HP were further incubated at 90°C for 1 hour before being sampled at 20 μL each.

[0167] Of the sampled solutions, U-DEL1-sh, U-DEL2-sh, U-DEL3-sh, and U-DEL4-sh were analyzed under the following analytical conditions 1, while U-DEL5-HP, U-DEL6-HP, U-DEL7-HP, U-DEL8-HP, U-DEL9-HP, and U-DEL10-HP were analyzed under the following analytical conditions 2.

[0168] Analysis conditions 1: Equipment: maXis (manufactured by Bruker), UltiMate 3000 (manufactured by Dionex) Column: ACQUITY UPLC Oligonucleotide BEH C18 Column (130 Å, 1.7 μm, 2.1 × 50 mm) Column temperature: 50℃ solvent: Solution A: Water (0.75% v / v hexafluoroisopropanol; 0.038% v / v triethylamine; 5 μM ethylenediaminetetraacetic acid) Solution B: 90% v / v methanol aqueous solution (0.75% v / v hexafluoroisopropanol; 0.038% v / v triethylamine; 5 μM ethylenediaminetetraacetic acid) Gradient conditions: Measurement was started with a fixed flow rate of 0.36 mL / min and a mixing ratio of solution A and solution B of 95 / 5 (v / v). After 0.56 minutes, the mixing ratio of solution A and solution B was linearly changed to 40 / 60 (v / v) over 5.5 minutes. Detection wavelength: 260nm

[0169] Analysis conditions 2: Equipment: Waters ACQUITY UPLC / SQ Detector Column: ACQUITY UPLC Oligonucleotide BEH C18 Column (130 Å, 1.7 μm, 2.1 × 50 mm) Column temperature: 50℃ solvent: Solution A: Water (0.75% v / v hexafluoroisopropanol; 0.038% v / v triethylamine; 5 μM ethylenediaminetetraacetic acid) Solution B: 90% v / v methanol aqueous solution (0.75% v / v hexafluoroisopropanol; 0.038% v / v triethylamine; 5 μM ethylenediaminetetraacetic acid) Gradient conditions: Measurement was started with a fixed flow rate of 0.36 mL / min and a mixing ratio of solution A and solution B of 95 / 5 (v / v). After 0.56 minutes, the mixing ratio of solution A and solution B was linearly changed to 40 / 60 (v / v) over 5.5 minutes. Detection wavelength: 260nm

[0170] Tables 2 and 3 show the expected product sequences and theoretical molecular weights of the deoxyuridine moiety (debase-debased deoxyuridine moiety and cleaved fragments) in each reaction solution, as well as the molecular weights detected in each reaction solution. The notation for each column in Tables 2 and 3 is as follows:

[0171] “Entry” (far left): The experiment numbers are shown, and the substrates corresponding to each experiment number (Entry) are as follows: Entry.1:U-DEL1-sh Entry.2:U-DEL2-sh Entry.3:U-DEL3-sh Entry.4:U-DEL4-sh Entry.5:U-DEL5-HP Entry.6:U-DEL6-HP Entry.7:U-DEL7-HP Entry.8:U-DEL8-HP Entry.9:U-DEL9-HP Entry.10:U-DEL10-HP

[0172] “No.” (Second from the left): This represents the sequence number. Of the sequence numbers (No.), Nos. 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 are the substrates of the respective reaction solutions, Nos. 11, 14, 17, 20, 22, 25, 29, 31, 33, and 35 are the debase-debased forms of the deoxyuridine moiety of each substrate, and the remaining sequence numbers are fragments obtained by cleaving each substrate.

[0173] “Seq.” (Third from the left): This represents the arrangement, with the left side representing the 5' side and the right side representing the 3' side. Note that in the notation, "(B)" is the following equation (6) [ka] This represents the group (debase site) represented by , and other notations are the same as in Table 1.

[0174] “Expected MW.” (Fourth from the left): This represents the theoretical molecular weight (Da) of each sequence.

[0175] “Observed MW.” (Far right): This shows the numerical value of the molecular weight (Da) detected for each sequence. A "-" indicates that the sequence was not detected.

[0176] [Table 2]

[0177] [Table 3]

[0178] The conversion rates for the debase reaction and cleavage reaction were calculated from the area ratio of the peaks corresponding to each detected sequence. For all substrates, the debase reaction resulted in over 99% conversion at 37°C for 1 hour (the substrate peak was less than 1%, with the remaining peaks consisting only of the debased product and cleaved fragments). Figure 10 shows a graph illustrating the conversion rate of the cleavage reaction. As shown in the graph, for all substrates except U-DEL8-HP and U-DEL9-HP, more than 95% of the cleavage reaction proceeded by 20 hours at 37°C. For U-DEL8-HP and U-DEL9-HP, 100% of the cleavage reaction was completed by adding an incubation period of 1 hour at 90°C.

[0179] The results above indicate that the hairpin-type DEL substructures containing various deoxyuridines undergo a debase reaction by USER® enzyme at the deoxyuridine moiety, followed by a cleavage reaction.

[0180] Example 2 [Comparison of PCR efficiency between conventional hairpin DEL and severable hairpin DEL (hairpin-type DEL containing deoxyuridine)]

[0181] As shown in the schematic diagram in Figure 11, the compound (hairpin DEL) with the sequence shown in Table 4 was synthesized using the following procedure. Note that in the sequence notation in Table 4, "S" represents the following formula (7) [ka] This refers to the base represented by , and other notations are the same as in Table 1. The names of the compounds corresponding to each sequence number (No.) are as follows: No.37: U-DEL1 No.38: U-DEL2 No.39: U-DEL4 No.40: U-DEL7 No.41: U-DEL8 No.42: U-DEL9 No.43: U-DEL10 No.44: H-DEL [Table 4] The compound names of the raw material headpieces used to synthesize each hairpin DEL are as follows: Hairpin DEL: Raw material headpiece U-DEL1 :U-DEL1-HP U-DEL2: U-DEL2-HP U-DEL4: U-DEL4-HP U-DEL7: U-DEL7-HP U-DEL8: U-DEL8-HP U-DEL9: U-DEL9-HP U-DEL10: U-DEL10-HP H-DEL: H-DEL-HP Furthermore, the sequence numbers "No." and sequence "Seq" for U-DEL1-HP, U-DEL2-HP, U-DEL4-HP, and H-DEL-HP are as shown in Table 5 below. [Table 5]

[0182] The raw material headpieces shown in Table 5 were prepared using the nucleic acid automated synthesizer nS-8II (manufactured by Gene Design Co., Ltd.) in the same manner as in Example 1.

[0183] In a PCR tube, 2.0 μL of 1 mM aqueous solution of various raw material headpieces; 2.4 μL of 1 mM aqueous solution of Pr_TAG (prepared by annealing Pr_TAG_a and Pr_TAG_b synthesized in the same manner as in Example 1; sequences are shown in Table 6); 0.8 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 100 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) and 2.0 μL of deionized water were added. 0.8 μL of a 10-fold diluted aqueous solution of T4 DNA ligase (Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16°C for 24 hours. The sequence notation in Table 6 is the same as in Table 1. The names of the compounds corresponding to each sequence number (No.) are as follows. No.49:Pr_TAG_a No.50:Pr_TAG_b [Table 6]

[0184] The reaction solution was treated with 0.8 μL of 5 M aqueous sodium chloride solution and 17.6 μL of cooled (-20°C) ethanol, and allowed to stand at -78°C for 2 hours. After centrifugation, the supernatant was removed, and the resulting pellets were air-dried. 2.0 μL of deionized water was added to each pellet to prepare the solution.

[0185] To each of the obtained solutions, 2.4 μL of a 1 mM aqueous solution of CP (prepared by annealing CP_a and CP_b synthesized in the same manner as in Example 1; sequences are shown in Table 7); 0.8 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 100 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) and 2.0 μL of deionized water were added. 0.8 μL of a 10-fold diluted aqueous solution of T4 DNA ligase (Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16°C for 24 hours. The sequence notation in Table 7 is the same as in Table 1. The names of the compounds corresponding to each sequence number (No.) are as follows. No.51: CP_a No.52: CP_b

Table 7

[0186] The reaction solution was treated with 0.8 μL of 5 M aqueous sodium chloride solution and 17.6 μL of cooled (-20 °C) ethanol, and allowed to stand at -78 °C for 2 hours. After centrifugation, the supernatant was removed, and the obtained pellet was air-dried. 10 μL of deionized water was added to the pellet to form a solution.

[0187] 1.0 μL of the obtained solution was sampled, diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under the conditions of Analytical Condition 2 in Example 1 to identify the target substance (the theoretical molecular weights and the detected molecular weights of each sequence are shown in Table 4). After the remaining solution was lyophilized, deionized water was added to each to adjust the concentration to 20 μM.

[0188] Among the 8 types of hairpin-type DELs obtained above, H-DEL is a conventional hairpin DEL, and the remaining 7 types are cleavable hairpin DELs containing deoxyuridine. To compare the PCR efficiency before treatment with USER (registered trademark) enzyme and the PCR efficiency after treatment for various hairpin-type DELs, real-time PCR analysis was performed. In addition, DS-DEL shown in Table 7 (prepared by annealing the compounds of Sequence Nos. 47 and 48) was used as the double-stranded DEL for comparison. In the sequence listings in Table 8, “(amino-C6-L)” means the group represented by the following formula (8)

Chemical formula

Table 8

[0189] <Treatment step by USER (registered trademark) enzyme> The treatment of 8 hairpin DELs and double-stranded DEL (DS-DEL) with USER® enzyme was performed according to the following procedure.

[0190] To a PCR tube, 1 μL of various DEL 20 μM aqueous solution; 1 μL of CutSmart® Buffer (manufactured by New England BioLabs, catalog number B7204S) and 7 μL of deionized water were added. 1 μL of USER® enzyme (manufactured by New England BioLabs, catalog number M5505S) was added to the solution, and the resulting solution was incubated at 37 °C for 1 hour.

[0191] <Preparation of DEL Samples> Samples of various DELs before USER® enzyme treatment and the reaction solutions after treatment were each diluted with deionized water to prepare 0.05 pM, 0.5 pM, and 5 pM DEL samples.

[0192] <Measurement of Ct Values by Real-Time PCR> The Ct values of the various DEL samples obtained above were measured by real-time PCR, and the PCR efficiencies were compared. The conditions were as follows, and the results are shown in Figure 12. The Ct value is the number of cycles at which the fluorescence signal generated with the amplification of DNA reaches an arbitrary threshold value in real-time PCR. That is, when the initial number of DNA molecules is the same, the higher the PCR efficiency, the lower the Ct value.

[0193] Apparatus: 7500 Real-Time PCR System (manufactured by Applied Biosystems) Plate: MicroAmp 96-Well Plate (manufactured by Applied Biosystems, catalog number N8010560) PCR Reaction Solution: · TB Green Premix Ex taqII (manufactured by Takara Bio, catalog number RR820): 10 μL · Forward Primer (Table 9, SEQ ID NO: 55): 0.80 μL • Reverse primer (Table 9, SEQ ID NO: 56): 0.80 μL • ROX Reference Dye II (manufactured by Takara Bio, catalog number RR39LR): 0.40 μL • Aqueous solutions of various DEL samples (0.05 pM, 0.5 pM, 5 pM)*1: 2.0 μL Deionized water: 6.0 μL * 1: The mole amounts of the DEL sample are 0.1 amol, 1 amol, and 10 amol. Temperature conditions: After holding at 95°C for 2 minutes, the following cycle was repeated 35 times. 95℃ for 5 seconds 52℃ for 30 seconds 72℃ for 30 seconds [Table 9] Note that the arrangement notation in Table 9 is the same as in Table 1.

[0194] As shown in Figure 12, the Ct value of conventional hairpin DEL (H-DEL) did not change before and after USER® enzyme treatment, but the Ct value of cleavable hairpin DEL containing deoxyuridine (U-DEL1, U-DEL2, U-DEL4, U-DEL7, U-DEL8, U-DEL9, and U-DEL10) decreased to the same level as DS-DEL, which is a double-stranded DEL, after USER® enzyme treatment.

[0195] These results indicate that DEL cleaved with USER® enzyme exhibits improved PCR efficiency compared to before cleavage, and that cleavable hairpin DEL containing deoxyuridine is cleaved with high efficiency and selectivity by USER® enzyme.

[0196] Example 3 [Verification of the cleavage reaction of hairpin DEL containing deoxyuridine using USER(registered trademark) enzyme] <Combination of 4 types of hairpins DEL (U-DEL5, U-DEL11, U-DEL12, and U-DEL13)> The compound with the sequence shown in Table 10 (hairpin DEL) was synthesized using the following procedure. Note that in the sequence notation in Table 10, "[mdC(TEG-amino)]" is represented by the following formula (9). [ka] This refers to the base represented by , and other notations are the same as in Table 4. The names of the compounds corresponding to each sequence number (No.) are as follows: No. 57: U-DEL5 No. 58: U-DEL11 No. 59: U-DEL12 No. 60: U-DEL13 [Table 10] The compound names of the raw material headpieces used to synthesize each hairpin DEL are as follows: Hairpin DEL: Raw material headpiece U-DEL5: U-DEL5-HP U-DEL11: U-DEL11-HP U-DEL12 : U-DEL12-HP U-DEL13: U-DEL13-HP Furthermore, the sequence numbers "No." and sequence "Seq" for U-DEL11-HP, U-DEL12-HP, and U-DEL13-HP are as shown in Table 11 below. Note that the notation in Table 11 is the same as in Table 10.

[0197] [Table 11]

[0198] Of the raw material headpieces shown in Table 11, U-DEL12-HP and U-DEL13-HP were prepared using the nucleic acid automated synthesizer nS-8II (manufactured by Gene Design Co., Ltd.) in the same manner as in Example 1. U-DEL11-HP was also prepared in the same manner according to the standard method.

[0199] Similar to Example 2, two-step double-stranded ligation was performed using various raw material headpieces with double-stranded oligonucleotides Pr_TAG and CP.

[0200] A portion of the obtained solution was sampled, diluted with deionized water, and then mass spectrometry by ESI-MS was performed under the analytical conditions 3 shown below to identify the target product (the theoretical molecular weight of each sequence and the detected molecular weight are shown in Table 10). The remaining solution was freeze-dried, and then deionized water was added to each to prepare a 20 μM solution.

[0201] Analysis condition 3: Equipment: Waters ACQUITY UPLC / SQ Detector Column: ACQUITY UPLC Oligonucleotide BEH C18 Column (130 Å, 1.7 μm, 2.1 × 50 mm) Column temperature: 60℃ solvent: Solution A: Water (0.75% v / v hexafluoroisopropanol; 0.038% v / v triethylamine; 5 μM ethylenediaminetetraacetic acid) Solution B: 90% v / v methanol aqueous solution (0.75% v / v hexafluoroisopropanol; 0.038% v / v triethylamine; 5 μM ethylenediaminetetraacetic acid) Gradient conditions: Measurement was started with a fixed flow rate of 0.36 mL / min and a mixing ratio of solution A and solution B of 95 / 5 (v / v). After 0.56 minutes, the mixing ratio of solution A and solution B was linearly changed to 40 / 60 (v / v) over 5.5 minutes. Detection wavelength: 260nm Deconvolution: The ion signal was analyzed using ProMass for MassLynx Software (manufactured by Waters).

[0202] <Cleavage reaction by <USER (registered trademark) enzyme>> The cleavage reactions of hairpin DELs (U-DEL5, U-DEL7, U-DEL9, U-DEL11, U-DEL12, and U-DEL13) containing six types of deoxyuridine by <USER (registered trademark) enzyme> were examined according to the following procedure.

[0203] Into a PCR tube, 2 μL of various hairpin DEL 20 μM aqueous solutions; 2 μL of CutSmart (registered trademark) Buffer (manufactured by New England BioLabs, catalog number B7204S) and 14 μL of deionized water were added. 2 μL of <USER (registered trademark) enzyme> (manufactured by New England BioLabs, catalog number M5505S) was added to the solution, and the resulting solution was incubated at 37 °C for 16 hours and then further incubated at 90 °C for 1 hour.

[0204] <Confirmation of products after cleavage by LC-MS measurement> 5.0 μL of the obtained reaction solution was sampled, diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under analysis conditions 3. The sequences and theoretical molecular weights of the products expected after cleavage in each reaction solution, and the molecular weights detected in each reaction solution are shown in Table 12. The substrates corresponding to each experimental number (Entry) are as follows, and other notations are the same as in Table 10. Entry.1: U-DEL5 Entry.2: U-DEL7 Entry.3: U-DEL9 Entry.4: U-DEL11 Entry.5: U-DEL12 Entry.6: U-DEL13

[0205]

Table 12

[0206] In all samples, no substrate MS was detected; instead, the MS of the cleavage product was observed as the main peak.

[0207] <Confirmation of cleavage reaction by gel electrophoresis> Furthermore, a portion of the obtained reaction solution was sampled and analyzed by modified polyacrylamide gel electrophoresis under the conditions shown below. The results shown in Figure 13 confirm that the cleavage reaction proceeded with high yield for all substrates. The samples in each lane of Figure 13 are as follows. Lane 1: 20 bp DNA Ladder (Lonza, catalog number 50330) Lane 2: U-DEL5 Lane 3: Sample after cleavage reaction of U-DEL5 Lane 4: U-DEL7 Lane 5: Sample after cleavage reaction of U-DEL7 Lane 6: U-DEL9 Lane 7: Sample after cleavage reaction of U-DEL9 Lane 8: U-DEL11 Lane 9: Sample after cleavage reaction of U-DEL11 Lane 10: U-DEL12 Lane 11: Sample after cleavage reaction of U-DEL12 Lane 12: U-DEL13 Lane 13: Sample after cleavage reaction of U-DEL13 Modified polyacrylamide gel electrophoresis: Gel: Novex (trademark) 10% TBE-Urea Gel (manufactured by Invitrogen by ThermoFisher SCIENTIFIC, catalog number EC68755BOX) Loading buffer: Novex (trademark) 10% TBE-Urea Sample Buffer (2x) (manufactured by Invitrogen by ThermoFisher SCIENTIFIC, catalog number LC6876) Temperature: 60℃ Voltage: 180V Electrophoresis time: 30 minutes Staining reagent: SYBER (trademark) GreenII Nucleic Acid Gel Stain (manufactured by Takara Bio, catalog number 5770A)

[0208] These results indicate that hairpin-type DELs containing various deoxyuridines undergo cleavage reactions by USER® enzyme at the deoxyuridine moiety.

[0209] Example 4 [Verification of the cleavage reaction of hairpin DEL containing deoxyinosine by endonuclease V] <Synthesis of hairpin DELs containing four types of deoxyinosine (I-DEL1, I-DEL2, I-DEL3, and I-DEL4)> The compound with the sequence shown in Table 13 (hairpin DEL) was synthesized using the following procedure. In the sequence notation in Table 13, "I" represents deoxyinosine, and the other notations are the same as in Table 2. The names of the compounds corresponding to each sequence number (No.) are as follows: No.73: I-DEL1 No.74: I-DEL2 No.75: I-DEL3 No.76: I-DEL4 [Table 13] The compound names of the raw material headpieces used to synthesize each hairpin DEL are as follows: Hairpin DEL: Raw material headpiece I-DEL1 : I-DEL1-HP I-DEL2: I-DEL2-HP I-DEL3 : I-DEL3-HP I-DEL4: I-DEL4-HP Furthermore, the sequence numbers "No." and sequence "Seq" for I-DEL1-HP, I-DEL2-HP, I-DEL3-HP, and I-DEL4-HP are as shown in Table 14 below. Note that the notation in Table 14 is the same as in Table 13.

[0210]

Table 14

[0211] The raw material head piece shown in Table 14 was prepared according to a conventional method.

[0212] Similar to Example 2, two-step double-stranded ligation of double-stranded oligonucleotide Pr_TAG and CP was carried out using various raw material head pieces.

[0213] A part of the obtained solution was sampled, diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under analytical condition 3 to identify the target substance (the theoretical molecular weights and the detected molecular weights of each sequence are shown in Table 13). The rest of the solution was lyophilized and then deionized water was added to each to prepare a solution with a concentration of 20 μM.

[0214] <Cleavage reaction by endonuclease V> Examination of the cleavage reaction of hairpin DEL (I-DEL1, I-DEL2, I-DEL3, I-DEL4) containing four kinds of deoxyinosine by endonuclease V was carried out according to the following procedure.

[0215] To a PCR tube, 1 μL of an aqueous solution of various hairpin DEL at 20 μM, 2 μL of NEBuffer (registered trademark) 4 (manufactured by New England BioLabs, catalog number B7004), and 15 μL of deionized water were added. 2 μL of Endonuclease V (manufactured by New England BioLabs, catalog number M0305S) was added to the solution, and the resulting solution was incubated at 37 °C for 24 hours.

[0216] <Confirmation of the product after cleavage by LC-MS measurement> 8.0 μL of the obtained reaction solution was sampled, diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under analytical condition 3. Table 15 shows the expected sequence and theoretical molecular weight of the cleaved products for each reaction solution, as well as the molecular weight detected for each reaction solution. The substrates corresponding to each experimental number (Entry) are as follows, and other notations are the same as in Table 13. Entry 1: I-DEL1 Entry 2: I-DEL2 Entry 3: I-DEL3 Entry 4: I-DEL4

[0217] [Table 15]

[0218] In all samples, no substrate MS was detected; instead, the MS of the cleavage product was observed as the main peak.

[0219] <Confirmation of cleavage reaction by gel electrophoresis> Furthermore, a portion of the obtained reaction solution was sampled and analyzed by modified polyacrylamide gel electrophoresis under the same conditions as in Example 3. The results shown in Figure 14 confirm that the cleavage reaction proceeded with high yield for all substrates. The samples in each lane of Figure 14 are as follows: Lane 1: 20 bp DNA Ladder (Lonza, catalog number 50330) Lane 2: I-DEL1 Lane 3: Sample after I-DEL1 cleavage reaction Lane 4: I-DEL2 Lane 5: Sample after I-DEL2 cleavage reaction Lane 6: I-DEL3 Lane 7: Sample after I-DEL3 cleavage reaction Lane 8: I-DEL4 Lane 9: Sample after I-DEL4 cleavage reaction

[0220] These results indicate that, in hairpin-type DELs containing various deoxyinosines, the second phosphodiester bond from the deoxyinosine is cleaved in the 3' direction by endonuclease V.

[0221] Example 5 [Verification of the cleavage reaction of hairpin DEL containing ribonucleoside by RNaseHII] <Synthesis of hairpin DEL (R-DEL1) containing ribonucleoside> The compound with the sequence shown in Table 16 (hairpin DEL) was synthesized using the following procedure. In the sequence notation in Table 16, "u" represents uridine, and other notations are the same as in Table 2. The names of the compounds corresponding to the sequence number (No.) are as follows: No.87: R-DEL1 [Table 16] The compound names of the raw material headpieces used to synthesize each hairpin DEL are as follows: Hairpin DEL: Raw material headpiece R-DEL1: R-DEL1-HP Furthermore, the sequence number "No." and sequence "Seq" of R-DEL1-HP are as shown in Table 17 below. Note that the notation in Table 17 is the same as in Table 16.

[0222] [Table 17]

[0223] The raw material headpieces shown in Table 17 were prepared according to standard procedures.

[0224] Similar to Example 2, a two-step double-stranded ligation was performed using a raw material headpiece with double-stranded oligonucleotides Pr_TAG and CP.

[0225] A portion of the obtained solution was sampled, diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under Analytical Condition 3 to identify the target substance (the theoretical molecular weights and the detected molecular weights of each sequence are shown in Table 16). After the remaining solution was lyophilized, deionized water was added to each to prepare a solution at 200 μM.

[0226] <Cleavage Reaction by RNaseHII> The cleavage reaction of ribonucleoside-containing hairpin DEL (R-DEL1) by RNaseHII was examined according to the following procedure.

[0227] <! To a PCR tube were added 0.5 μL of a 200 μM aqueous solution of hairpin DEL, 4.9 μL of ThermoPol® Reaction Buffer Pack (manufactured by New England BioLabs, catalog number B9004), and 43.6 μL of deionized water. 1 μL of RNase HII (manufactured by New England BioLabs, catalog number M0288S) was added to the solution, and the resulting solution was incubated at 37 °C for 8 hours.

[0228] <Confirmation of the Product after Cleavage by LC-MS Measurement> 10 μL of the obtained reaction solution was sampled and subjected to mass spectrometry by ESI-MS under Analytical Condition 3. The sequences, theoretical molecular weights, and detected molecular weights of the expected products after cleavage are shown in Table 18. The substrates corresponding to the experiment numbers (Entry) are as follows, and other notations are the same as those in Table 16. Entry.1: R-DEL1

[0229]

Table 18

[0230] In none of the samples was the MS of the substrate detected, and the MS of the product after cleavage was observed as the main peak.

[0231] <Confirmation of the Cleavage Reaction by Gel Electrophoresis> <000188> Furthermore, a portion of the obtained reaction solution was sampled and analyzed by modified polyacrylamide gel electrophoresis under the same conditions as in Example 3. The results shown in Figure 15 confirm that the cleavage reaction proceeded with high yield for all substrates. The samples in each lane of Figure 15 are as follows: Lane 1: 20 bp DNA Ladder (Lonza, catalog number 50330) Lane 2: R-DEL1 Lane 3: Sample after R-DEL1 cleavage reaction

[0232] These results indicate that hairpin-type DELs containing ribonucleosides undergo cleavage of the phosphodiester bond at the 5' end of the ribonucleotide by RNaseHII. Example 6 [Creating a model library using U-DEL9-HP as the raw material] As shown in the schematic diagram in Figure 16, a model library containing 3×3×3(27) compound species was synthesized using U-DEL9-HP as a starting material by split-and-pool synthesis, with the following reagents. ·U-DEL9-HP • Three types of building blocks (BB1, BB2, and BB3): [ka] • Ten types of double-stranded oligonucleotide tags (tag numbers in Table 19: Pr, A1, A2, A3, B1, B2, B3, C1, C2, and C3)

[0233] In Table 19, “Tag No.” (far left) represents the tag number, “No.” (second from the left) represents the sequence number, and “Seq.” (third from the left) represents the sequence. The sequence notation is the same as in Table 1.

[0234] Each double-stranded oligonucleotide tag was prepared by annealing two oligonucleotides with sequence numbers corresponding to each tag number, as shown in Table 19.

[0235] [Table 19]

[0236] <Synthesis of compound "AOP-U-DEL9-HP"> The compound "AOP-U-DEL9-HP" with the sequence shown in Table 20 was synthesized using the following procedure. Note that in the sequence notation in Table 20, "(AOP-AminoC7)" is represented by the following formula (10). [ka] This refers to the base represented by , and other notations are the same as in Table 2.

[0237] [Table 20]

[0238] Four bioremo centrifuge tubes were filled with a 2.5 mL, 1 mM solution of U-DEL9-HP in sodium borate buffer (150 mM, pH 9.4) cooled to 10°C. To each tube, 40 equivalents of N-Fmoc-15-amino-4,7,10,13-tetraoxaoctadecanoic acid (250 μL, 0.4 M N-dimethylacetamide solution) were added, followed by 40 equivalents of 4-(4,6-dimethoxy[1,3,5]triazine-2-yl)-4-methylmorpholinium chloride hydrate (DMTMM) (200 μL, 0.5 M aqueous solution). The resulting solutions were shaken at 10°C for 5 hours.

[0239] The above solutions were treated with 295 μL of 5 M aqueous sodium chloride solution and 9.7 mL of chilled (-20°C) ethanol, respectively, and allowed to stand overnight at -78°C. After centrifugation, the supernatant was removed, and the resulting pellets were air-dried. 2.75 mL of deionized water was added to each pellet to dissolve it, 306 μL of piperidine was added at 0°C, and the mixture was shaken at 10°C for 3 hours. After centrifugation of the mixture, the precipitate was removed by filtration, and the mixture was washed twice with 1.47 mL of deionized water. The resulting filtrates were treated with 600 μL of 5 M aqueous sodium chloride solution and 19.8 mL of chilled (-20°C) ethanol, respectively, and allowed to stand overnight at -78°C. After centrifugation, the supernatant was removed, and the resulting pellets were air-dried.

[0240] The obtained pellet was mixed with 10 mL of deionized water to form a solution. A portion of the solution was sampled, diluted with deionized water, and then mass spectrometry was performed by ESI-MS under the conditions of analytical condition 2 in Example 1 to identify the target compound (the theoretical molecular weight and detected molecular weight of the compound are shown in Table 20). The remainder of the solution was freeze-dried, and then deionized water was added to prepare a 5 mM solution.

[0241] <Introduction of the double-stranded oligonucleotide tag "Pr"> The compound "AOP-U-DEL9-HP-Pr" with the sequence shown in Table 21 was synthesized by ligating the compound "AOP-U-DEL9-HP" with the double-stranded oligonucleotide tag "Pr" using the following procedure. Note that the sequence notation in Table 21 is the same as in Table 20.

[0242] [Table 21]

[0243] In a Bioremo centrifuge tube, 40 μL of a 5 mM aqueous solution of the compound "AOP-U-DEL9-HP", 160 μL of a 100 mM sodium bicarbonate aqueous solution, 240 μL of a 1 mM aqueous solution of the double-stranded oligonucleotide tag "Pr", 80 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 100 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) and 272 μL of deionized water were added. 8.0 μL of T4 DNA ligase (Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16°C for 24 hours.

[0244] The reaction solution was treated with 80 μL of 5 M aqueous sodium chloride solution and 2640 μL of cooled (-20°C) ethanol, and allowed to stand at -78°C for 2 hours. After centrifugation, the supernatant was removed, and 400 μL of deionized water was added to the resulting pellet. The resulting solution was concentrated using an Amicon® Ultra Centrifugal filter (30 kD cutoff). A portion of the resulting solution was sampled and mass spectrometry was performed by ESI-MS under the conditions of analysis condition 2 to identify the target compound (the theoretical molecular weight of the compound and the detected molecular weight are shown in Table 21). Through the above steps, 133 nmol of the compound "AOP-U-DEL9-HP-Pr" with a purity of 84.5% was obtained. The obtained compound "AOP-U-DEL9-HP-Pr" was prepared to a concentration of 1 mM by adding 100 mM aqueous sodium bicarbonate solution.

[0245] <Cycle A> In each of the three PCR tubes, 20 μL of a 1 mM solution of the compound "AOP-U-DEL9-HP-Pr" obtained above, 30 μL of a 1 mM aqueous solution of one of the double-stranded oligonucleotide tags A1-A3, 8.0 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 100 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) and 21.6 μL of deionized water were added. 0.4 μL of T4 DNA ligase (Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16°C for 18 hours.

[0246] Each reaction solution was treated with 8.0 μL of 5 M aqueous sodium chloride solution and 264 μL of chilled (-20°C) ethanol, and allowed to stand at -78°C for 30 minutes. After centrifugation, the supernatant was removed, and the resulting pellets were dissolved in 20 μL of 150 mM sodium borate buffer (pH 9.4).

[0247] To each tube, 40 equivalents of one of the building blocks BB1-BB3 (4.0 μL, 200 mM N,N-dimethylacetamide solution), followed by 40 equivalents of 4-(4,6-dimethoxy[1.3.5]triazine-2-yl)-4-methylmorpholinium chloride hydrate (DMTMM) (4.0 μL, 200 mM aqueous solution), were added, and the mixture was shaken at 10°C for 2 hours. Furthermore, to each tube, 20 equivalents of building block (2.0 μL, 200 mM N,N-dimethylacetamide solution), followed by 20 equivalents of DMTMM (2.0 μL, 200 mM aqueous solution), were added, and the mixture was shaken at 10°C for 30 minutes.

[0248] Each reaction solution was treated with 3.2 μL of 5 M aqueous sodium chloride solution and 106 μL of chilled (-20°C) ethanol, and allowed to stand at -78°C for 30 minutes. After centrifugation, the supernatant was removed, and 18 μL of deionized water was added to each of the resulting pellets. The three solutions were then mixed in a single PCR tube.

[0249] To the mixed solution, 6.0 μL of piperidine was added at 0°C and shaken at room temperature for 1 hour. The reaction solution was treated with 6.0 μL of 5 M aqueous sodium chloride solution and 198 μL of chilled (-20°C) ethanol, and allowed to stand at -78°C for 18 hours. After centrifugation, the supernatant was removed, and 400 μL of deionized water was added to the resulting pellet. The resulting solution was concentrated using an Amicon® Ultra Centrifugal filter (30 kD cutoff), and 100 mM aqueous sodium bicarbonate solution was added to adjust the concentration to 1 mM, which was then used as the starting material for the next step.

[0250] <Cycle B> Each of the three PCR tubes contained 13.7 μL of 1 mM solution of the starting material obtained in cycle A; 20.6 μL of 1 mM aqueous solution of one of the double-stranded oligonucleotide tags B1-B3; 5.5 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 100 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate); and 14.8 μL of deionized water. 0.3 μL of T4 DNA ligase (Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16°C for 16 hours.

[0251] Each reaction solution was treated with 5.5 μL of 5 M aqueous sodium chloride solution and 181 μL of chilled (-20°C) ethanol, and allowed to stand at -78°C for 30 minutes. After centrifugation, the supernatant was removed, and the resulting pellets were dissolved in 13.7 μL of 150 mM sodium borate buffer (pH 9.4).

[0252] To each tube, one of the building blocks BB1-BB3 (5.5 μL, 200 mM N,N-dimethylacetamide solution) was added (80 equivalents), followed by 80 equivalents of DMTMM (5.5 μL, 200 mM aqueous solution), and the mixture was shaken at 10°C for 1 hour. Then, to each tube, 40 equivalents of building block (2.3 μL, 200 mM N,N-dimethylacetamide solution) was added (40 equivalents of DMTMM (2.3 μL, 200 mM aqueous solution)), and the mixture was shaken at 10°C for 2 hours.

[0253] Each reaction solution was treated with 2.5 μL of 5 M aqueous sodium chloride solution and 81.4 μL of chilled (-20°C) ethanol, and allowed to stand at -78°C for 30 minutes. After centrifugation, the supernatant was removed, and 12.3 μL of deionized water was added to each of the resulting pellets. The three solutions were then mixed in a single PCR tube.

[0254] To the mixed solution, 4.1 μL of piperidine was added at 0°C and shaken at room temperature for 3 hours. The reaction solution was treated with 4.1 μL of 5 M aqueous sodium chloride solution and 136 μL of chilled (-20°C) ethanol, and allowed to stand at -78°C for 3 hours. After centrifugation, the supernatant was removed, and 400 μL of deionized water was added to the resulting pellet. The resulting solution was concentrated using an Amicon® Ultra Centrifugal filter (30 kD cutoff), and 100 mM aqueous sodium bicarbonate solution was added to adjust the concentration to 0.48 mM, which was then used as the starting material for the next step.

[0255] <Cycle C> Each of the three PCR tubes contained 14.5 μL of a 0.48 mM solution of the starting material obtained in cycle B; 10.5 μL of a 1 mM aqueous solution of one of the double-stranded oligonucleotide tags C1-C3; and 2.8 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 100 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate). 0.14 μL of T4 DNA ligase (Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16°C for 16 hours.

[0256] Each reaction solution was treated with 2.8 μL of 5 M aqueous sodium chloride solution and 92 μL of chilled (-20°C) ethanol, and allowed to stand at -78°C for 30 minutes. After centrifugation, the supernatant was removed, and the resulting pellets were dissolved in 7.0 μL of 150 mM sodium borate buffer (pH 9.4).

[0257] To each tube, one of the building blocks BB1-BB3 (2.8 μL, 200 mM N,N-dimethylacetamide solution) was added (80 equivalents), followed by 80 equivalents of DMTMM (2.8 μL, 200 mM aqueous solution), and the mixture was shaken at 10°C for 1 hour. Furthermore, to each tube, 40 equivalents of building block (1.4 μL, 200 mM N,N-dimethylacetamide solution) was added (40 equivalents of DMTMM (1.4 μL, 200 mM aqueous solution)), and the mixture was shaken at 10°C for 2 hours.

[0258] Each reaction solution was treated with 1.3 μL of 5 M aqueous sodium chloride solution and 41.4 μL of chilled (-20°C) ethanol, and allowed to stand at -78°C for 30 minutes. After centrifugation, the supernatant was removed, and 6.3 μL of deionized water was added to each of the resulting pellets. The three solutions were then mixed in a single PCR tube.

[0259] To the mixed solution, 2.1 μL of piperidine was added at 0°C and shaken at room temperature for 2 hours. The reaction solution was treated with 2.1 μL of 5 M aqueous sodium chloride solution and 69 μL of chilled (-20°C) ethanol and allowed to stand at -78°C for 3 hours. After centrifugation, the supernatant was removed and 400 μL of deionized water was added to the resulting pellet. The resulting solution was concentrated using an Amicon® Ultra Centrifugal filter (30 kD cutoff), and 100 mM aqueous sodium bicarbonate solution was added to adjust the concentration to 0.41 mM, which was then used as the starting material for the next step.

[0260] <CPのライゲーション> In a PCR tube, 12.2 μL of a 0.41 mM solution of starting material obtained in cycle C was added; 6.0 μL of a 1 mM aqueous solution of CP (the same as used in Example 2); 2.1 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 100 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) and 0.7 μL of deionized water were added. 0.1 μL of T4 DNA ligase (Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16°C for 16 hours.

[0261] The reaction solution was treated with 2.1 μL of 5 M aqueous sodium chloride solution and 69.6 μL of chilled (-20°C) ethanol, and allowed to stand at -78°C for 30 minutes. After centrifugation, the supernatant was removed, and 400 μL of deionized water was added to the resulting pellet. The resulting solution was concentrated using an Amicon® Ultra Centrifugal filter (30 kD cutoff), and deionized water was added to prepare a 20 μM solution.

[0262] <Result> Samples from each cycle after ligation of the double-stranded oligonucleotide tag were analyzed by electrophoresis using a 2.2% agarose gel (Lonza, FlashGel® cassette, catalog number 57031). The results shown in Figure 17 confirm that encoding by the double-stranded oligonucleotide tag was achieved with high efficiency in each cycle. The samples for each lane in Figure 17 are as follows: Lane 1: AOP-U-DEL9-HP-Pr Lane 2: Sample after A1 ligation of double-stranded oligonucleotides in cycle A. Lane 3: Sample after A2 ligation of double-stranded oligonucleotides in cycle A. Lane 4: Sample after A3 ligation of double-stranded oligonucleotides from cycle A. Lane 5: Sample after B1 tagging of double-stranded oligonucleotides in cycle B Lane 6: Sample after B2 ligation of double-stranded oligonucleotides in cycle B Lane 7: Sample after B3 ligation of double-stranded oligonucleotides in cycle B Lane 8: Sample after C1 ligation of double-stranded oligonucleotide tag in cycle C. Lane 9: Sample after C2 ligation of double-stranded oligonucleotide tag in cycle C. Lane 10: Sample after C3 ligation of double-stranded oligonucleotides in cycle C. Lane 11: Sample after CP ligation Lane 12: 20 bp DNA Ladder (Lonza, catalog number 50330)

[0263] The samples after the completion of cycle C were analyzed under analytical condition 3. Figure 18 shows the chromatograph and mass spectrum results. Deconvolution of the obtained mass spectrum revealed an average molecular weight of 35532.4. This result is consistent with the average molecular weight expected after the completion of cycle C (35514.2), indicating that the reactions for library synthesis (ligation of double-stranded oligonucleotide tags and introduction of building blocks) were achieved with high efficiency.

[0264] Based on the above synthesis procedure, a model library containing 3×3×3(27) compound species using U-DEL9-HP as a starting material was successfully synthesized.

[0265] <Disconnection of the obtained model library by USER(registered trademark)enzyme> The cleavage reaction using the USER(registered trademark)enzyme model library obtained above was performed using the following procedure.

[0266] 2.0 μL of a 20 μM aqueous solution of the model library, 2 μL of CutSmart® Buffer (New England BioLabs, catalog number B7204S), and 14 μL of deionized water were added to a PCR tube. 2 μL of USER® enzyme (New England BioLabs, catalog number M5505S) was added to the solution, and the resulting solution was incubated at 37°C for 16 hours, followed by incubation at 90°C for 1 hour.

[0267] A portion of the obtained reaction solution was sampled and analyzed by modified polyacrylamide gel electrophoresis under the same conditions as in Example 3. The results shown in Figure 19 confirm that the model library using U-DEL9-HP as a raw material undergoes a highly efficient cleavage reaction using USER® enzyme. The samples in each lane of Figure 19 are as follows: Lane 1: 20 bp DNA Ladder (Lonza, catalog number 50330) Lane 2: Model Library Lane 3: Sample after cleavage reaction using the USER(registered trademark)enzyme model library.

[0268] Example 7 [Conversion of DEL compounds from hairpin DNA to single-stranded DNA, and conferral of new functions] <Synthesis of the DEL compound "BIO-DEL" with biotin at its 3' end> Similar to Example 2, the DEL compound "BIO-DEL" with the sequence shown in Table 22 was synthesized using the following procedure. Note that in the sequence notation in Table 22, "(BIO)" represents the following formula (11) [ka] This refers to the base represented by , and other notations are the same as in Table 20. [Table 22]

[0269] In a PCR tube, 20 μL of a 1 mM aqueous solution of AOP-U-DEL9-HP (synthesized in Example 6) was added; 24 μL of a 1 mM aqueous solution of Pr_TAG2 (prepared by annealing Pr_TAG2_a and Pr_TAG2_b synthesized in the same manner as in Example 1; the sequences are shown in Table 23); 8 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 10 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) and 20 μL of deionized water were added. 8 μL of a 10-fold diluted aqueous solution of T4 DNA ligase (Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16°C for 22 hours. The sequence notation in Table 23 is the same as in Table 1. The names of the compounds corresponding to (No.) in each sequence number are as follows. No.114:Pr_TAG2_a No.115:Pr_TAG2_b [Table 23]

[0270] The reaction solution was treated with 8 μL of 5 M aqueous sodium chloride solution and 264 μL of chilled (-20°C) ethanol, and allowed to stand overnight at -78°C. After centrifugation, the supernatant was removed, and the resulting pellet was air-dried. The pellet was dissolved in deionized water and purified by reverse-phase HPLC using a Phenomenex Gemini C18 column. The target product was eluted using a two-way mobile phase gradient profile with 50 mM triethylammonium acetate buffer (pH 7.5) and acetonitrile / 50 mM triethylammonium acetate buffer (9:1, v / v). The fraction containing the target product was collected, mixed, and concentrated. The resulting solution was desalted using an Amicon® Ultra Centrifugal filter (3 kD cutoff), precipitated with ethanol, and then 25 μL of deionized water was added to the pellet to make a solution.

[0271] A portion of the obtained solution was sampled, diluted with deionized water, and then mass spectrometry was performed by ESI-MS under the conditions of analytical condition 2 in Example 1 to identify the target compound (the theoretical molecular weight of the compound and the detected molecular weight are shown in Table d). The remaining solution was freeze-dried, and then 100 mM sodium bicarbonate aqueous solution was added to each to prepare a 1 mM solution.

[0272] To 6.2 μL of the solution obtained above, 7.4 μL of a 1 mM aqueous solution of CP-BIO (prepared by annealing CP_a and CP-BIO_b synthesized in the same manner as in Example 1; the sequences are shown in Table 24); 2.5 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 10 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) and 6.2 μL of deionized water were added. 2.47 μL of a 10-fold diluted aqueous solution of T4 DNA ligase (Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16°C for 16 hours. The sequence notation in Table 24 is the same as in Table 23. The names of the compounds corresponding to (No.) in each sequence number are as follows. No.51: CP_a No.116: CP-BIO_b

[0273]

Table 24

[0274] The reaction solution was treated with 2.5 μL of 5 M aqueous sodium chloride solution and 81.5 μL of cooled (-20 °C) ethanol, and left standing at -78 °C for 30 minutes. After centrifugation, the supernatant was removed, the obtained pellet was air-dried, and the pellet was dissolved in deionized water. The obtained solution was desalted using an Amicon® Ultra Centrifugal Filter (3 kD cutoff).

[0275] A part of the obtained supernatant was sampled, diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under the analysis conditions 3 of Example 3 to identify the target substance (the theoretical molecular weight and the detected molecular weight are shown in Table 22). The rest of the solution was lyophilized, deionized water was added, and it was adjusted to 120 μM to obtain BIO-DEL.

[0276] <Cleavage of BIO-DEL by USER® enzyme> The cleavage reaction of the above-obtained DEL compound "BIO-DEL" by USER® enzyme was carried out according to the following procedure to synthesize a DEL compound "DS-BIO-DEL" having a double-stranded nucleic acid of the sequence shown in Table 25. The sequence listing in Table 25 is the same as that in Table 22, which means that DS-BIO-DEL is formed by a double-stranded of the oligonucleotide chains of SEQ ID NO: 118 and SEQ ID NO: 119.

[0277]

Table 25

[0278] Three PCR tubes were each filled with 10 μL of a 120 μM aqueous solution of the DEL compound "BIO-DEL," 100 μL of CutSmart® Buffer (New England BioLabs, catalog number 7240S), and 860 μL of deionized water. 30 μL of USER® enzyme (New England BioLabs, catalog number 5505S) was added to each solution, and the resulting solutions were incubated at 37°C for 24 hours.

[0279] Each of the resulting reaction solutions was desalted using an Amicon® Ultra Centrifugal filter (3kD cutoff), and deionized water was added to prepare 60 μL solutions. Each solution was then treated with 6 μL of 5 M sodium chloride aqueous solution and 198 μL of cooled (-20°C) ethanol, and allowed to stand at -78°C for 30 minutes. After centrifugation, the supernatant was removed, and the resulting pellet was dissolved in deionized water and combined into a single tube.

[0280] A portion of the obtained solution was sampled, diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under analytical conditions 3 of Example 3 to identify the target double-stranded nucleic acid-containing DEL compound "DS-BIO-DEL" (the theoretical molecular weight and detected molecular weight of the compound are shown in Table 25).

[0281] Furthermore, a portion of the obtained reaction solution was sampled and analyzed by modified polyacrylamide gel electrophoresis under the same conditions as in Example 3. The results shown in Figure 20 confirm that BIO-DEL was cleaved in high yield and converted to DS-BIO-DEL. The samples in each lane of Figure 20 are as follows: Lane 1: BIO-DEL (Concentration 1: Prepared so that BIO-DEL is approximately 40 ng) Lane 2: BIO-DEL (Concentration 2: Prepared so that BIO-DEL is approximately 80 ng) Lane 3: Sample after cleavage reaction with BIO-DEL's USER® enzyme (Concentration 1: Prepared so that the target substance is approximately 40 ng) Lane 4: Sample after cleavage reaction with BIO-DEL's USER® enzyme (concentration 2: prepared so that the target substance is approximately 80 ng) Lane 5: 20 bp DNA Ladder (Lonza, catalog number 50330)

[0282] <Preparation of single-stranded DNA-containing DEL using streptavidin beads> The double-stranded nucleic acid-containing DEL compound "DS-BIO-DEL" obtained above was treated with streptavidin beads to prepare the single-stranded DNA-containing DEL compound "SS-DEL" according to the following procedure. SS-DEL is the oligonucleotide chain of SEQ ID NO: 119 in Table 25.

[0283] 450 μL of Magnosphere (trademark) MS160 / Streptavidin (JSR Life Sciences, catalog number J-MS-S160S) was added to each of two PCR tubes. The supernatant was removed by magnetic separation, and then 900 μL of 1× binding buffer (10 mM Tris-HCl, pH 7.5; 0.5 mM ethylenediaminetetraacetic acid; 1 M sodium chloride; 0.05% v / v Tween20) was added. The supernatant was removed by magnetic separation. To the resulting particles, DS-BIO-DEL aqueous solution k (700 pmol, 450 μL) and 450 μL of 2× binding buffer (20 mM Tris-HCl, pH 7.5; 1 mM ethylenediaminetetraacetic acid; 2 M sodium chloride; 0.1% v / v Tween20) were added to each tube and mixed. The mixture was shaken at room temperature for 20 minutes.

[0284] The supernatant was removed from the mixture by magnetic separation, and the particles were washed with 900 μL of 1× binding buffer (10 mM Tris-HCl, pH 7.5; 0.5 mM ethylenediaminetetraacetic acid; 1 M sodium chloride; 0.05% v / v Tween20), and the supernatant was removed by magnetic separation, repeating this process three times. Subsequently, 900 μL of aqueous solution (0.1 M sodium hydroxide; 0.1 M sodium chloride) was added to each, and the supernatant was recovered by magnetic separation.

[0285] To each of the obtained supernatants, 900 μL of 3-(N-morpholino)propanesulfonic acid buffer (1.0 M, pH 7.0) was added, and desalting was performed using an Amicon® Ultra Centrifugal filter (3 kD cutoff). The obtained supernatants were combined into a single tube and treated with 13.6 μL of 5 M sodium chloride aqueous solution and 448 μL of cooled (-20°C) ethanol, and allowed to stand at -78°C for 60 minutes. After centrifugation, the supernatant was removed, and the resulting pellet was air-dried. 60 μL of deionized water was added to the pellet to make a solution.

[0286] A portion of the obtained solution was sampled, diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under the conditions of analysis condition 3 in Example 3. A molecular weight of 23984.8 was observed, identifying the target single-stranded DNA-containing DEL compound "SS-DEL".

[0287] <Synthesis of photoreactive crosslinker-modified primers> The photoreactive crosslinker-modified primer "PXL-Pr" with the sequence shown in Table 26 was synthesized using the following procedure. Note that in the sequence notation in Table 26, "(X)" represents the following formula (12). [ka] This refers to the base represented by , and other notations are the same as in Table 2. [Table 26]

[0288] A solution of L-Pr (synthesized in the same manner as in Example 1, sequence shown in Table 27) in sodium borate buffer (150 mM, pH 9.4) cooled to 10°C (200 μL, 1 mM) was added to a PCR tube. 40 equivalents of N-Fmoc-15-amino-4,7,10,13-tetraoxaoctadecanoic acid (20 μL, 0.4 M N-dimethylacetamide solution), followed by 40 equivalents of 4-(4,6-dimethoxy[1.3.5]triazine-2-yl)-4-methylmorpholinium chloride hydrate (DMTMM) (16 μL, 0.5 M aqueous solution), were added to the tube, and the resulting mixture was shaken at 10°C for 5 hours. The sequence notation in Table 27 is the same as in Table 8. [Table 27]

[0289] The reaction mixture was treated with 23.6 μL of 5 M aqueous sodium chloride solution and 778.8 μL of cooled (-20°C) ethanol, and allowed to stand overnight at -78°C. After centrifugation, the supernatant was removed, and the resulting pellet was air-dried. 180 μL of deionized water was added to the pellet to make a solution, then 20 μL of piperidine was added, and the mixture was shaken at 10°C for 3 hours.

[0290] The obtained solution was treated with 20 μL of 5 M aqueous sodium chloride solution and 660 μL of cooled (-20°C) ethanol, and allowed to stand at -78°C for 30 minutes. After centrifugation, the supernatant was removed, and 200 μL of deionized water was added to the resulting pellet to make a 1 mM solution.

[0291] To 100 μL of the solution obtained above, 75 μL of triethylamine hydrochloride buffer (500 mM, pH 10) was added, followed by 50 equivalents of 1-((3-(3-methyl-3H-diazilin-3-yl)propanoyl)oxy)-2,5-dioxopyrrolidine-3-sulfonate sodium (Sulfo-SDA) (25 μL, 200 mM aqueous solution), and the mixture was shaken at 37°C for 2 hours.

[0292] The obtained solution was treated with 20 μL of 5 M aqueous sodium chloride solution and 660 μL of cooled (-20°C) ethanol, and allowed to stand at -78°C for 30 minutes. After centrifugation, the supernatant was removed, and 100 μL of deionized water was added to the resulting pellet, followed by 75 μL of triethylamine hydrochloride buffer (500 mM, pH 10) and 50 equivalents of Sulfo-SDA (25 μL, 200 mM aqueous solution), and the mixture was shaken at 37°C for 1 hour and 20 minutes. Another 50 equivalents of Sulfo-SDA (25 μL, 200 mM aqueous solution) were added, and the mixture was shaken at 37°C for 40 minutes.

[0293] The obtained solution was treated with 22.5 μL of 5 M aqueous sodium chloride solution and 743 μL of chilled (-20°C) ethanol, and left to stand overnight at -78°C. After centrifugation, the supernatant was removed, and 100 μL of deionized water was added to the resulting pellet, followed by 75 μL of triethylamine hydrochloride buffer (500 mM, pH 10), and then 50 equivalents of Sulfo-SDA (25 μL, 200 mM aqueous solution), and the mixture was shaken at 37°C for 3 hours.

[0294] The obtained solution was treated with 20 μL of 5 M aqueous sodium chloride solution and 660 μL of chilled (-20°C) ethanol, and allowed to stand overnight at -78°C. After centrifugation, the supernatant was removed, and the resulting pellet was air-dried. The pellet was dissolved in 50 mM triethylammonium acetate buffer (pH 7.5) and purified by reverse-phase HPLC using a Phenomenex Gemini C18 column. The target substance was eluted using a two-way mobile phase gradient profile with 50 mM triethylammonium acetate buffer (pH 7.5) and acetonitrile / water (100:1, v / v). The fraction containing the target substance was collected, mixed, and concentrated. The resulting solution was desalted using an Amicon® Ultra Centrifugal filter (3 kD cutoff), precipitated with ethanol, and then 100 μL of deionized water was added to the pellet to make a solution.

[0295] A portion of the obtained solution was sampled, diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under the conditions of analytical condition 3 in Example 3 to identify the target product, the photoreactive crosslinker-modified primer "PXL-Pr" (the theoretical molecular weight of the compound and the detected molecular weight are shown in Table 26).

[0296] <Synthesis of photoreactive crosslinker-modified double-stranded EL> Using the SS-DEL and PXL-Pr obtained above, a primer extension reaction was carried out according to the following procedure to synthesize the photoreactive crosslinker-modified double-stranded DEL "PXL-DS-DEL" with the sequence shown in Table 28. Note that the sequence notation in Table 28 is the same as in Table 26, meaning that PXL-DS-DEL is formed by a double strand of oligonucleotide chains of SEQ ID NO: 122 and SEQ ID NO: 119. [Table 28]

[0297] 50 μL of 8 μM aqueous solution of "SS-DEL", 0.673 μL of 594 μM aqueous solution of "PXL-Pr", 80 μL of 10× NEBuffer (trademark) 2 (New England BioLabs, catalog number B7002S), and 645 μL of deionized water were added to a PCR tube. 8 μL of DNA Polymerase I, Large (Klenow) Fragment (New England BioLabs, catalog number M0210) and 16 μL of Deoxynucleotide (dNTP) Solution Mix (New England BioLabs, catalog number N0447) were added to the solution, and the resulting solution was incubated at 25°C for 90 minutes.

[0298] The obtained solution was desalted using an Amicon® Ultra Centrifugal filter (3kD cutoff). 17 μL of deionized water was added to the supernatant, followed by treatment with 6 μL of 5 M sodium chloride aqueous solution and 198 μL of cooled (-20°C) ethanol. The mixture was then allowed to stand at -78°C for 60 minutes. After centrifugation, the supernatant was removed, and the resulting pellet was air-dried. 40 μL of deionized water was added to the pellet to form a solution.

[0299] A portion of the obtained solution was sampled, diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under the conditions of analytical condition 3 of Example 3 to identify the target product, a photoreactive crosslinker-modified double-stranded DEL "PXL-DS-DEL" (the theoretical molecular weight of the compound and the detected molecular weight are shown in Table 28).

[0300] Furthermore, a portion of the obtained reaction solution was sampled and analyzed by polyacrylamide gel electrophoresis under the conditions shown below. The results shown in Figure 21 confirm that PXL-DS-DEL was produced in high yield by the primer extension reaction. The samples in each lane of Figure 21 are as follows. Lane 1: 20 bp DNA Ladder (Lonza, catalog number 50330) Lane 2: DS-BIO-DEL Lane 3: SS-DEL Lane 4: Sample after primer extension reaction of SS-DEL (PXL-DS-DEL)

[0301] Polyacrylamide gel electrophoresis: Gel: SuperSep (trademark) DNA 15% TBE gel (manufactured by Fujifilm Wako Pure Chemical Industries, catalog number 190-15481) Loading Buffer: 6 x Loading Buffer (Manufactured by Takara Bio, catalog number 9156) Temperature: room temperature Voltage: 200V Electrophoresis time: 50 minutes Staining reagent: SYBER (trademark) GreenII Nucleic Acid Gel Stain (manufactured by Takara Bio, catalog number 5770A) [Industrial applicability]

[0302] The present invention utilizes nucleic acid compounds containing selectively cleavable sites. Furthermore, the present invention provides a DNA encoding library containing selectively cleavable sites, a composition for its synthesis, and a method for using the same, enabling the production of DNA encoding libraries with greater convenience than conventional methods.

Claims

1. Equation (I) 【Chemistry 1】 (In the formula, E and F are independent of each other. It is an oligomer composed of nucleotides or nucleic acid analogs, However, E and F contain complementary base sequences and form a double-stranded oligonucleotide. LP is, 【Chemistry 2】 This is a loop section represented by, LS is a substructure selected from the group of compounds described in (A) to (C) below, (A) Nucleotides (B) Nucleic acid analogs (C) Trivalent C1-14 groups which may have substituents (LP1)p is a substructure in which LP1 is selected individually or differently from the group of compounds described in (1) and (2) below, (1) Nucleotides (2) Nucleic acid analogs (LP2)q is a substructure in which LP2 is selected individually or differently from the group of compounds described in (1) and (2) below, (1) Nucleotides (2) Nucleic acid analogs The total number of p and q is between 0 and 40. L is a linker, D is a reactive functional group, (E, F, or LP has at least one selectively cleavable portion.) The use of the compound represented by the above as a headpiece, Use as a headpiece in the preparation of libraries of compounds that are low molecular weight organic compounds or polypeptides.

2. The use according to claim 1, wherein at least one of parts E or F has at least one selectively severable portion.

3. The use according to claim 1, wherein the 5' end of E is bonded to LP and E has a cleavable portion.

4. The use according to any one of claims 1 to 3, wherein the total number of p and q is 2 to 20.

5. The use according to any one of claims 1 to 3, wherein the total number of p and q is 2 to 10.

6. The use according to any one of claims 1 to 3, wherein the total number of p and q is 2 to 7.

7. The use according to any one of claims 1 to 3, wherein the total number of p and q is 0.

8. LP1, LP2, and LS each have the following structure: (A) Nucleotides or (B) Nucleic acid analogs that meet the requirements of (B11) to (B15) below (B11) Having phosphoric acid (or equivalent part) and a hydroxyl group (or equivalent part thereof), (B12) Composed of carbon, hydrogen, oxygen, nitrogen, phosphorus or sulfur, (B13) Molecular weight is between 142 and 1500. (B14) The number of atoms between residues is 3 to 30. (B15) The bonding pattern between atoms in the residues is either all single bonds, or one or two double bonds with the remainder being single bonds. The use according to any one of claims 1 to 7, wherein the structure is selected individually or differently from those.

9. LP1, LP2, and LS each have the following structure: (A) Nucleotides or (B) Nucleic acid analogs that meet the requirements of (B21) to (B25) below (B21) Having phosphoric acid and hydroxyl groups, (B22) Composed of carbon, hydrogen, oxygen, nitrogen, or phosphorus, (B23) Molecular weight is between 142 and 1000. (B24) The number of atoms between residues is 3 to 15. (B25) The bonding mode between atoms in each residue is all single bonds. The use according to any one of claims 1 to 8, wherein the structure is selected individually or differently from those.

10. LP1, LP2, and LS each have the following structure: (A) Nucleotides or (B) Nucleic acid analogs that meet the requirements of (B31) to (B35) below (B31) Having phosphoric acid and hydroxyl groups, (B32) Composed of carbon, hydrogen, oxygen, nitrogen, or phosphorus, (B33) Molecular weight is between 142 and 700. (B34) The number of atoms between residues is 4 to 7. (B35) The bonding mode between atoms in each residue is all single bonds. The use according to any one of claims 1 to 9, wherein the structure is selected individually or differently from those.

11. LP1 and LP2 are as follows: (B41) d-Spacer, (B5) Polyalkylene glycol phosphate ester The use according to any one of claims 1 to 10, which is any one of the above.

12. The use according to any one of claims 1 to 11, wherein LP1 and LP2 are each diethylene glycol phosphate ester or triethylene glycol phosphate ester.

13. The use according to any one of claims 1 to 12, wherein LP1 and LP2 are each triethylene glycol phosphate esters.

14. The use according to any one of claims 1 to 11, wherein LP1 and LP2 are each d-Spacer.

15. LP1 and LP2 are nucleotides, The use according to any one of claims 1 to 10.

16. LS is given by equations (a) to (g): 【Chemistry 4】 (In the formula, * indicates the bond position with the linker, ** indicates the bond position with LP1 or LP2, and R is a hydrogen atom or a methyl group.) The use according to any one of claims 1 to 15, which is any one of the above.

17. LS is given by equation (h): 【Transformation 5】 (In the formula, * indicates the linker connection position, and ** indicates the linker connection position.) The use according to any one of claims 1 to 15.

18. The use according to any one of claims 1 to 15, wherein LS is a polyalkylene glycol phosphate ester.

19. LS is given by equations (i) to (k): 【Transformation 6】 (In the formula, n1, m1, p1, and q1 are each independent integers between 1 and 20, * indicates the linker connection position, and ** indicates the linker connection position with LP1 or LP2.) The use according to any one of claims 1 to 15.

20. LS is given by equation (l): 【Transformation 7】 (In the formula, * indicates the linker connection position, and ** indicates the linker connection position.) The use according to any one of claims 1 to 15.

21. LS is (B42), (B43), or (B44): (B42) Amino C6 dT (B43) mdC (TEG-Amino) (B44) Uni-Link (Registered Trademark) Amino Modifier The use according to any one of claims 1 to 15, which is any one of the above.

22. LS is a nucleotide. The use according to any one of claims 1 to 15.

23. LS is a trivalent C1-C14 group which may have a (C) substituent, and (C) has the following structure: (1) C1-10 aliphatic hydrocarbons which may have substituents and which may be replaced by 1-3 heteroatoms, (2) C6-14 aromatic hydrocarbons which may have substituents, (3) A C2-9 aromatic heterocycle which may have substituents, (4) C2-9 non-aromatic heterocycles which may have substituents The use according to any one of claims 1 to 15, which is any one of the above.

24. LS is a trivalent C1-C14 group which may have a (C) substituent, and (C) has the following structure: (1) C1-6 aliphatic hydrocarbons which may have substituents, (2) C6-10 aromatic hydrocarbons which may have substituents, (3) C2-5 aromatic heterocycles which may have substituents The use according to any one of claims 1 to 15, which is any one of the above.

25. LS is a trivalent C1-C14 group which may have a (C) substituent, and (C) has the following structure: (1) C1-6 aliphatic hydrocarbons, (2) Benzene, or (3) C2-5 nitrogen-containing aromatic heterocycle Here, (1) to (3) may be unsubstituted or substituted with one to three substituents selected individually or differently from substituent group ST1, wherein substituent group ST1 is composed of C1-6 alkyl groups, C1-6 alkoxy groups, fluorine atoms, and chlorine atoms; however, if substituent group ST1 is substituted with an aliphatic hydrocarbon, alkyl groups are not selected from substituent group ST1. The use according to any one of claims 1 to 15, which is any one of the above.

26. LS is a trivalent C1-C14 group which may have a (C) substituent, and (C) has the following structure: (1) C1-6 alkyl groups, (2) Unsubstituted or benzenes substituted with one or two C1-3 alkyl groups or C1-3 alkoxy groups The use according to any one of claims 1 to 15, which is any one of the above.

27. LS is a trivalent C1-C14 group which may have a (C) substituent, and (C) has the following structure: (1) C1-6 alkyl groups The use according to any one of claims 1 to 15.

28. E and F are oligomers composed independently of nucleotides or nucleic acid analogs. The chain lengths of E and F are 3 to 40, The use according to any one of claims 1 to 27.

29. E and F are oligomers composed independently of nucleotides or nucleic acid analogs. The chain lengths of E and F are 4 to 30, respectively. The use according to any one of claims 1 to 28.

30. E and F are oligomers composed independently of nucleotides or nucleic acid analogs. The chain lengths of E and F are 6 to 25, respectively. The use according to any one of claims 1 to 29.

31. E and F are oligomers composed independently of nucleotides or nucleic acid analogs. E and F contain complementary base sequences, forming a double-stranded oligonucleotide. The E and F double-stranded oligonucleotides are the overhanging ends. The use according to any one of claims 1 to 30.

32. The use according to claim 31, wherein the protruding portion of the protruding end has a length of two bases or more.

33. E and F are oligomers composed independently of nucleotides or nucleic acid analogs. E and F contain complementary base sequences, forming a double-stranded oligonucleotide. The E and F double-stranded oligonucleotides have blunt ends. The use according to any one of claims 1 to 30.

34. The chain lengths of the complementary base sequences contained in E and F are each three bases or longer. The use described in any one of claims 1 to 33.

35. The chain lengths of the complementary base sequences contained in E and F are each four bases or longer. The use according to any one of claims 1 to 34.

36. The chain lengths of the complementary base sequences contained in E and F are each 6 bases or longer. The use described in any one of claims 1 to 35.

37. E and F are oligomers composed of nucleotides, The use according to any one of claims 1 to 36.

38. The use according to any one of claims 1 to 37, wherein the nucleotide is a ribonucleotide or a deoxyribonucleotide.

39. The use according to any one of claims 1 to 38, wherein the nucleotide is a deoxyribonucleotide.

40. The use according to any one of claims 1 to 39, wherein the nucleotide is deoxyadenosine, deoxyguanosine, thymidine, or deoxycytidine.

41. The use according to any one of claims 1 to 36, wherein E and F are oligomers composed independently of nucleic acid analogs.

42. L, (1) C1-20 aliphatic hydrocarbons which may have substituents and which may be replaced by 1-3 heteroatoms, or (2) C6-14 aromatic hydrocarbons which may have substituents The use according to any one of claims 1 to 41.

43. The use according to any one of claims 1 to 42, wherein L is a C1-6 aliphatic hydrocarbon which may have substituents, a C1-6 aliphatic hydrocarbon which may be replaced by one or two oxygen atoms, or a C6-10 aromatic hydrocarbon which may have substituents.

44. L is a C1-6 aliphatic hydrocarbon substituted with substituent group ST1, or a benzene substituted with substituent group ST1, where substituent group ST1 is a group consisting of C1-6 alkyl groups, C1-6 alkoxy groups, fluorine atoms, and chlorine atoms (however, when substituent group ST1 is substituted with an aliphatic hydrocarbon, alkyl groups are not selected from substituent group ST1), the use according to any one of claims 1 to 43.

45. The use according to any one of claims 1 to 44, wherein L is a C1-6 alkyl group, or a benzene that is unsubstituted or substituted with one or two C1-3 alkyl groups or C1-3 alkoxy groups.

46. The use according to any one of claims 1 to 45, wherein L is a C1-6 alkyl group.

47. The reactive functional group of D is The use according to any one of claims 1 to 46, wherein the functional group is a C-C, amino, ether, carbonyl, amide, ester, urea, sulfide, disulfide, sulfoxide, sulfonamide, or a reactive functional group capable of forming a sulfonyl bond.

48. The use according to any one of claims 1 to 47, wherein the reactive functional group of D is a C1 hydrocarbon having a leaving group, an amino group, a hydroxyl group, a precursor of a carbonyl group, a thiol group, or an aldehyde group.

49. The use according to any one of claims 1 to 48, wherein the reactive functional group of D is a C1 hydrocarbon having a halogen atom, a C1 hydrocarbon having a sulfonic acid leaving group, an amino group, a hydroxyl group, a carboxyl group, a halogenated carboxyl group, a thiol group, or an aldehyde group.

50. The reactive functional group of D is -CH 2 Cl, -CH 2 Br, -CH 2 OSO 2 CH 3 ien-CH 2 OSO 2 CF 3 The use according to any one of claims 1 to 49, wherein the amino group is a hydroxyl group or a carboxyl group.

51. The use according to any one of claims 1 to 50, wherein the reactive functional group of D is a primary amino group.

52. The selectively cleavable site is a deoxyribonucleoside that is neither deoxyadenosine, deoxyguanosine, thymidine, nor deoxycytidine. The use according to any one of claims 1 to 51.

53. The selectively cleavable sites are deoxyuridine, bromodeoxyuridine, deoxyinosine, 8-hydroxydeoxyguanosine, 3-methyl-2'-deoxyadenosine, N6-etheno-2'-deoxyadenosine, 7-methyl-2'-deoxyguanosine, 2'-deoxyxanthosine, or 5,6-dihydroxy-5,6-dihydrodeoxythymidine. The use according to any one of claims 1 to 52.

54. The selectively cleavable site is deoxyuridine or deoxyinosine. The use described in any one of claims 1 to 53.

55. The site that can be selectively cleaved is deoxyuridine. The use according to any one of claims 1 to 54.

56. The site that can be selectively cleaved is deoxyinosine. The use according to any one of claims 1 to 54.

57. The selectively cleavable site is the second phosphodiester bond in the 3' direction from deoxyinosine. The use according to any one of claims 1 to 51.

58. The site that can be selectively cleaved is the ribonucleoside. The use according to any one of claims 1 to 51.

59. There is only one part that can be selectively cut. The use according to any one of claims 1 to 58.

60. At least one cleavable portion is included in E or (LP1)p, and at least one cleavable portion is included in F or (LP2)q, The use according to any one of claims 1 to 58.

61. The cleavable portion included in E or (LP1)p and the cleavable portion included in F or (LP2)q are cleavable under different conditions. The use described in claim 60.

Citation Information

Patent Citations

  • Method for synthesizing bifunctional conjugates

    JP2006515752A

  • Methods for Synthesis of Encoded Libraries

    JP2007524662A

  • Method for constructing and screening DNA-coding libraries

    JP2012517812A

  • Directed and coded oligonucleotides for combinatorial synthesis of coded probe molecules

    JP2019518031A

  • Methods and compositions for synthesis of encoded libraries

    JP2020510403A