Cleavable DNA-encoded libraries

JP2024167312A5Active Publication Date: 2025-07-31NISSAN CHEM CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024146187
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-05-25
Filing Date
2024-08-28
Publication Date
2025-07-31
Estimated Expiration
2041-05-24

AI Technical Summary

Technical Problem

Existing DNA-encoded libraries (DELs) face challenges in combining the advantages of hairpin strand DNA and double-strand DNA, such as high chemical stability and efficient PCR efficiency, respectively, due to their inherent structural limitations.

Method used

Incorporation of cleavable sites, such as deoxyuridine, into the DNA strands using nucleic acid cleavage technology, allowing for the simultaneous advantages of both hairpin and double-strand DNA structures by enabling selective cleavage and efficient PCR.

Benefits of technology

This approach enables the production of DELs with enhanced chemical stability and PCR efficiency, facilitating more convenient and efficient synthesis and evaluation of compound libraries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2024167312000001
    Figure 2024167312000001
  • Figure 2024167312000002
    Figure 2024167312000002
  • Figure 2024167312000003
    Figure 2024167312000003
Patent Text Reader

Abstract

To provide a technique that realizes both of the advantage of hairpin-stranded DNA and the advantage of double-stranded DNA as DNA strand structures for DNA-encoded libraries.SOLUTION: The present invention relates to a method for using a nucleic acid compound that contains a selectively cleavable site. The present invention also relates to a DNA-encoded library that contains a selectively cleavable site, a composition for synthesizing the same, and a method for using the same.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to DNA-encoded libraries that contain cleavable sites in the DNA strand. [Background technology]

[0002] A compound library is a systematic collection of compound derivatives that may have a specific activity, such as a drug candidate compound. These compound libraries are often synthesized based on the synthesis techniques and methodologies of combinatorial chemistry. Combinatorial chemistry is an experimental method for efficiently synthesizing a wide variety of compounds through systematic synthetic routes from a series of compound libraries that are enumerated and designed based on combinatorial theory, and a field of research related to this method. One type of compound library based on combinatorial chemistry is the DNA-encoded library. Hereinafter, the DNA-encoded library will be abbreviated as DEL. In DEL, a DNA tag is added to each compound in the library. The sequence of the DNA tag is designed so that each structure of each compound can be identified, and it functions as a label for the compound (Patent Documents 1 to 3). The two most typical DNA strand structures of DEL known so far are double stranded and hairpin stranded. Below, we will provide an overview of double-stranded DELs and hairpin-stranded DELs, as well as their advantages and disadvantages. (1) Hairpin chain DEL DELs using hairpin DNA have a single-stranded structure in which two complementary DNA strands are linked, and are synthesized using hairpin-shaped DNA having functional groups for the introduction of various building blocks as the starting material (headpiece) (Patent Document 3, Non-Patent Documents 1 and 2). (A) Advantages (a) Short DNA tags can be used. In this method, a relatively short double-stranded DNA tag of about 9 to 13 mers with a 2-mer sticky end is often used, and the double-stranded DNA tag is introduced by a ligation reaction using DNA ligase. The use of such a short DNA tag is possible because the hairpin strand DNA forms a strong double strand in the molecule, and the DNA site other than the sticky end does not interfere with the DNA tag. The use of a short double-stranded DNA tag has several advantages in DEL synthesis. One advantage is that the cost of synthesizing the DNA tag is low. Another advantage is that the use of a shorter DNA tag can reduce the total length of the DEL when encoding the same number of reaction cycles. In other words, even if a larger number of cycles is encoded, the total length of the DEL can be reduced to a range where the DNA sequence can be efficiently read by a next-generation sequencer. In fact, in Non-Patent Document 3, the construction of a DEL using a hairpin strand DNA that encodes as many as six cycles of reaction has been achieved by using a hairpin strand DNA. (b) High chemical stability Unlike double strands, in the case of hairpin strands, even if the double strand structure melts during the heating reaction, the double strand is reformed within the original molecule under the subsequent reannealing conditions without strand exchange. Therefore, DEL using hairpin strand DNA has the advantage that it can be used under a wider range of chemical conditions (Non-Patent Document 2). In general, when nucleic acid strands have the same chain length, hairpin strands form double strands more strongly (higher Tm value) than double strands. Therefore, under various chemical conditions when introducing building blocks, each chemical structure of hairpin strand DNA, especially the structure of the base portion, should resist structural conversion compared to double strands. (B) Disadvantages Due to its strong double-stranded formation ability, hairpin strand DNA has the problem that it is difficult to melt the double strand, bind a primer oligonucleotide, and initiate a polymerase reaction, resulting in low PCR efficiency (Patent Document 4). (2) Double strand DEL DELs using double-stranded DNA are synthesized using single-stranded DNA (non-hairpin single-stranded DNA) or double-stranded DNA as the starting material (headpiece) that has functional groups for the introduction of various building blocks. (A) Disadvantages In contrast to DEL, which uses hairpin DNA, in many cases, relatively long single- or double-stranded DNA tags of approximately 20 to 30 mers with 4 to 10 mer sticky ends are used (Patent Document 2, Non-Patent Document 4), and DEL that codes for approximately three cycles of reaction is common. (B) Advantages DEL using double-stranded DNA does not have the same problem as hairpin DNA in terms of PCR efficiency. Furthermore, unlike hairpin DNA, double-stranded DNA can be denatured to single-stranded DNA or subjected to strand exchange reaction, and has the advantage of being adaptable to a wide range of evaluation methods by converting it into a DNA structure suited to various applications. For example, evaluation methods with high sensitivity ratios that utilize the double-stranded ability of DNA have been developed (Non-Patent Documents 5 and 6).

[0003] Thus, hairpin DNA and double-stranded DNA each have advantages in the synthesis and evaluation of DEL, but no technology is known that combines these advantages. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] WO 93 / 20243 [Patent Document 2] International Publication No. 2004 / 039825 [Patent Document 3] International Publication No. 2005 / 058479 [Patent Document 4] International Publication No. 2010 / 094036 [Non-patent literature]

[0005] [Non-Patent Document 1] Nature Chemical Biology, 2009, Vol. 5, pp. 647-654 [Non-Patent Document 2] A Handbook for DNA-Encoded Chemistry, edited by Robert A. Goodnow, Jr., John Wiley & Sons, Inc. [Non-Patent Document 3] ACS Chemical Biology, 2018, Vol. 13, pp. 53-59 [Non-Patent Document 4] Nature Chemistry, 2018, Vol. 10, pp. 441-448 [Non-Patent Document 5] Annual Review of Biochemistry, 2018, Vol. 87, pp. 479-502 [Non-Patent Document 6] ACS Combinatorial Science, 2020, Vol. 22, pp. 204-212 Summary of the Invention [Problem to be solved by the invention]

[0006] The present invention provides DELs that contain a cleavable site in the DNA strand, and methods for producing DELs. [Means for solving the problem]

[0007] One of the techniques for nucleic acid chemistry is cleavage of nucleic acids such as DNA. For example, when deoxyuridine is introduced into a DNA chain, it can be selectively cleaved by the USER® enzyme. As a result of extensive research, the present inventors discovered that by introducing a cleavable site, such as deoxyuridine, into a DNA strand, it is possible to achieve both the advantages of hairpin strand DNA and double-stranded DNA, and thus completed the present invention. Therefore, the present invention is as follows.

[0008] [1] Formula (I) [ka] (In the formula, E and F are each independently An oligomer composed of nucleotides or nucleic acid analogs, provided that E and F contain complementary base sequences to each other and form a double-stranded oligonucleotide; LP is the loop region; L is a linker, D is a reactive functional group. A compound represented by the formula: A compound having at least one selectively cleavable site in at least one of sites E, F, or LP. [2] A composition for use in preparing a headpiece of a compound library, comprising the compound described in [1]. [3] A composition for use in preparing a headpiece of a DNA-encoded library, comprising the compound described in [1]. [4] Formula (I) [ka] (In the formula, E and F are each independently An oligomer composed of nucleotides or nucleic acid analogs, provided that E and F contain complementary base sequences to each other and form a double-stranded oligonucleotide; LP is the loop region; L is a linker, D is a reactive functional group. A compound represented by the formula: At least one selectively cleavable site in at least one of E, F, or LP; A compound used as a headpiece in a compound library. [5] Formula (I) [ka] (In the formula, E and F are each independently An oligomer composed of nucleotides or nucleic acid analogs, provided that E and F contain complementary base sequences to each other and form a double-stranded oligonucleotide; LP is the loop region; L is a linker, D is a reactive functional group. A compound represented by the formula: At least one selectively cleavable site in at least one of E, F, or LP; Compounds used as headpieces for DNA-encoded libraries. [6] Formula (I) [ka] (In the formula, E and F are each independently An oligomer composed of nucleotides or nucleic acid analogs, provided that E and F contain complementary base sequences to each other and form a double-stranded oligonucleotide; LP is the loop region; L is a linker, D is a reactive functional group. A compound represented by the formula: A headpiece of a compound library having at least one selectively cleavable site in at least one of sites E, F, or LP. [7] Formula (I) [ka] (In the formula, E and F are each independently An oligomer composed of nucleotides or nucleic acid analogs, provided that E and F contain complementary base sequences to each other and form a double-stranded oligonucleotide; LP is the loop region; L is a linker, D is a reactive functional group. A compound represented by the formula: A headpiece of a DNA-encoded library having at least one selectively cleavable site in at least one of the E, F, or LP sites. [8] Formula (II) [ka] (In the formula, X and Y are oligonucleotide strands; E and F are each independently An oligomer composed of nucleotides or nucleic acid analogs, provided that E and F contain complementary base sequences to each other and form a double-stranded oligonucleotide; LP is the loop region; L is a linker, D is a divalent group derived from a reactive functional group, Sp is a bond or a bifunctional spacer; An is a partial structure composed of at least one building block. A compound represented by the formula: X and Y have a sequence capable of forming a double strand at least in part, X is bound to E at the 5' end, Y binds to F at the 3' end, A compound having at least one selectively cleavable site in at least one of sites E, F, or LP. [9] Formula (III) An-Sp-C-Bn (III) (In the formula, An and Sp have the same meaning as in [8]. Bn represents a double-stranded oligonucleotide tag formed by oligonucleotide strand X and oligonucleotide strand Y; C is a group represented by the formula (I) [ka] (In the formula, E, LP, L, D and F have the same meaning as in [8], except that D binds to Sp, and E and F bind to the corresponding terminal sides of the double-stranded oligonucleotide tag Bn.) Represented by The compound according to [8].

[10] An is the same as [8] and is a partial structure constructed of n building blocks α1 to αn (n is an integer from 1 to 10), Bn is a double-stranded oligonucleotide tag formed of an oligonucleotide strand X and an oligonucleotide strand Y, and is a partial structure containing an oligonucleotide containing a base sequence capable of identifying the structure of An; The compound according to [8] or [9].

[11] LP, A loop region represented by (LP1)p-LS-(LP2)q, LS is a partial structure selected from the group of compounds described in (A) to (C) below, (A) Nucleotides (B) Nucleic acid analog (C) a C1-14 trivalent group which may have a substituent LP1 is a partial structure selected from the group of compounds described in the following (1) and (2) in a quantity of p alone or differently, (1) Nucleotides (2) Nucleic acid analogs LP2 is each partial structure selected from the group of compounds described in the following (1) and (2) in a quantity of q, either singly or differently, (1) Nucleotides (2) Nucleic acid analogs The sum of p and q is 0 to 40. The compound according to any one of [1], [4], [5], [8] to

[10] .

[12] The compound according to

[11] , wherein the total number of p and q is 2 to 20.

[13] The compound according to

[11] , wherein the total number of p and q is 2 to 10.

[14] The compound according to

[11] , wherein the total number of p and q is 2 to 7.

[15] The compound according to

[11] , wherein the sum of p and q is 0.

[16] LP1, LP2 and LS each have the following structure: (A) Nucleotides or (B) A nucleic acid analog satisfying the following requirements (B11) to (B15). (B11) Having a phosphoric acid group (or an equivalent portion) and a hydroxyl group (or an equivalent portion), (B12) composed of carbon, hydrogen, oxygen, nitrogen, phosphorus or sulfur; (B13) having a molecular weight of 142 to 1500; (B14) The number of atoms between residues is 3 to 30. (B15) The bonds between the atoms of the residues are all single bonds or contain one or two double bonds and the rest are single bonds. The compound according to any one of

[11] to

[15] , wherein the compound has a structure selected alone or differently from the following:

[17] LP1, LP2 and LS each have the following structure: (A) Nucleotides or (B) A nucleic acid analog satisfying the following requirements (B21) to (B25). (B21) Having phosphoric acid and hydroxyl groups, (B22) composed of carbon, hydrogen, oxygen, nitrogen or phosphorus; (B23) having a molecular weight of 142 to 1000; (B24) The number of atoms between residues is 3 to 15. (B25) The bonds between the atoms of the residues are all single bonds. The compound according to any one of

[11] to

[16] , wherein the compound has a structure selected alone or differently from the following:

[18] LP1, LP2 and LS each have the following structure: (A) Nucleotides or (B) A nucleic acid analog satisfying the following requirements (B31) to (B35). (B31) Having phosphoric acid and hydroxyl groups, (B32) Composed of carbon, hydrogen, oxygen, nitrogen or phosphorus, (B33) having a molecular weight of 142 to 700; (B34) The number of atoms between residues is 4 to 7. (B35) The bonds between the residues are all single bonds. The compound according to any one of

[11] to

[17] , wherein the compound has a structure selected alone or differently from the following:

[19] LP1 and LP2 are, respectively: (B41) d-Spacer, (B5) Polyalkylene glycol phosphate ester The compound according to any one of

[11] to

[18] ,

[20] The compound according to any one of

[11] to

[19] , wherein LP1 and LP2 are diethylene glycol phosphate or triethylene glycol phosphate, respectively.

[21] The compound according to any one of

[11] to

[20] , wherein LP1 and LP2 are each triethylene glycol phosphate ester.

[22] The compound according to any one of

[11] to

[19] , wherein LP1 and LP2 are each a d-Spacer. [twenty three] LP1 and LP2 are each a nucleotide; The compound according to any one of

[11] to

[18] .

[24] LS is expressed by the formulas (a) to (g): [ka] (In the formula, * means a bonding position with a linker, ** means a bonding position with LP1 or LP2, and R is a hydrogen atom or a methyl group.) The compound according to any one of

[11] to

[23] ,

[25] LS is expressed by the formula (h): [ka] (In the formula, * means a bonding position with a linker, and ** means a bonding position with LP1 or LP2.) The compound according to any one of

[11] to

[23] , wherein

[26] The compound according to any one of

[11] to

[23] , wherein LS is a polyalkylene glycol phosphate ester.

[27] LS is represented by the formulas (i) to (k): [ka] (In the formula, n1, m1, p1, and q1 each independently represent an integer of 1 to 20, * represents a bonding position with the linker, and ** represents a bonding position with LP1 or LP2.) The compound according to any one of

[11] to

[23] , wherein

[28] LS is represented by the formula (l): [ka] (In the formula, * means a bonding position with a linker, and ** means a bonding position with LP1 or LP2.) The compound according to any one of

[11] to

[23] , wherein

[29] LS is (B42), (B43) or (B44): (B42) Amino C6 dT (B43)mdC(TEG-Amino) (B44) Uni-Link (trademark) Amino Modifier The compound according to any one of

[11] to

[23] ,

[30] LS is a nucleotide. The compound according to any one of

[11] to

[23] .

[31] L-S is (C) a C1-14 trivalent group which may have a substituent, and (C) is the following structure: (1) C1-10 aliphatic hydrocarbons which may have a substituent and may be substituted with 1 to 3 heteroatoms; (2) C6-14 aromatic hydrocarbons which may have a substituent, (3) a C2-9 aromatic heterocycle which may have a substituent, or (4) C2-9 non-aromatic heterocycle optionally having a substituent The compound according to any one of

[11] to

[15] and

[19] to

[23] ,

[32] LS is (C) a C1-14 trivalent group which may have a substituent, and (C) is the following structure: (1) C1-6 aliphatic hydrocarbons which may have a substituent, (2) an optionally substituted C6-10 aromatic hydrocarbon, or (3) C2-5 aromatic heterocycle optionally having a substituent The compound according to any one of

[11] to

[15] and

[19] to

[23] ,

[33] LS is (C) a C1-14 trivalent group which may have a substituent, and (C) is the following structure: (1) C1-6 aliphatic hydrocarbons, (2) Benzene, or (3)C2~5 nitrogen-containing aromatic heterocycle wherein the (1) to (3) are unsubstituted or may be substituted with 1 to 3 substituents selected, either alone or differently, from the substituent group ST1, the substituent group ST1 being a group consisting of a C1 to 6 alkyl group, a C1 to 6 alkoxy group, a fluorine atom and a chlorine atom, provided that when the substituent group ST1 substitutes an aliphatic hydrocarbon, an alkyl group is not selected from the substituent group ST1; The compound according to any one of

[11] to

[15] and

[19] to

[23] ,

[34] LS is (C) a C1-14 trivalent group which may have a substituent, and (C) is the following structure: (1) a C1-6 alkyl group, or (2) Benzene that is unsubstituted or substituted with one or two C1-3 alkyl or C1-3 alkoxy groups The compound according to any one of

[11] to

[15] and

[19] to

[23] ,

[35] L-S is (C) a C1-14 trivalent group which may have a substituent, and (C) is the following structure: (1) C1-6 alkyl group The compound according to any one of

[11] to

[15] and

[19] to

[23] , wherein

[36] E and F are each independently an oligomer composed of nucleotides or nucleic acid analogs; The chain lengths of E and F are each 3 to 40; The compound according to any one of [1], [4], [5], and [8] to

[35] .

[37] E and F are each independently an oligomer composed of nucleotides or nucleic acid analogs; The chain lengths of E and F are 4 to 30, respectively. The compound according to any one of [1], [4], [5], and [8] to

[36] .

[38] E and F are each independently an oligomer composed of nucleotides or nucleic acid analogs; The chain lengths of E and F are 6 to 25, respectively. The compound according to any one of [1], [4], [5], and [8] to

[37] .

[39] E and F are each independently an oligomer composed of nucleotides or nucleic acid analogs; E and F contain complementary base sequences to each other and form a double-stranded oligonucleotide; The double-stranded oligonucleotides E and F are overhanging ends. The compound according to any one of [1], [4], [5], and [8] to

[38] .

[40] The compound according to

[39] , wherein the overhang of the protruding end is two or more bases in length.

[41] E and F are each independently an oligomer composed of nucleotides or nucleic acid analogs; E and F contain complementary base sequences to each other and form a double-stranded oligonucleotide; The double-stranded oligonucleotides E and F are blunt ended. The compound according to any one of [1], [4], [5], and [8] to

[38] .

[42] The complementary sequences in E and F are each 3 or more base pairs long. The compound according to any one of [1], [4], [5], [8] to

[41] .

[43] The complementary sequences in E and F are each 4 or more base pairs long. The compound according to any one of [1], [4], [5], and [8] to

[42] .

[44] The length of the complementary sequences in E and F is 6 or more bases. The compound according to any one of [1], [4], [5], and [8] to

[43] .

[45] The compound according to any one of [1], [4], [5], and [8] to

[44] , wherein E and F are each independently an oligomer composed of nucleotides.

[46] The nucleotide is a ribonucleotide or a deoxyribonucleotide. The compound according to any one of [1], [4], [5], and [8] to

[45] .

[47] The nucleotide is a deoxyribonucleotide. The compound according to any one of [1], [4], [5], and [8] to

[46] .

[48] ​​The method according to claim 1, wherein the nucleotide is deoxyadenosine, deoxyguanosine, thymidine, or deoxycytidine. The compound according to any one of [1], [4], [5], and [8] to

[47] .

[49] The compound according to any one of [1], [4], [5], and [8] to

[44] , wherein E and F are each independently an oligomer composed of a nucleic acid analog.

[50] L. (1) C1-20 aliphatic hydrocarbons which may have a substituent and may be substituted with 1 to 3 heteroatoms; or (2) C6-14 aromatic hydrocarbons which may have a substituent The compound according to any one of [1], [4], [5], and [8] to

[49] , wherein

[51] The compound according to any one of [1], [4], [5], and [8] to

[50] , wherein L is a C1-6 aliphatic hydrocarbon optionally having a substituent, a C1-6 aliphatic hydrocarbon optionally substituted with 1 or 2 oxygen atoms, or a C6-10 aromatic hydrocarbon optionally having a substituent.

[52] The compound according to any one of [1], [4], [5], and [8] to

[51] , wherein L is a C1-6 aliphatic hydrocarbon which can be substituted with a substituent group ST1, or a benzene which can be substituted with a substituent group ST1, where the substituent group ST1 is a group consisting of a C1-6 alkyl group, a C1-6 alkoxy group, a fluorine atom, and a chlorine atom (however, when the substituent group ST1 substitutes an aliphatic hydrocarbon, an alkyl group is not selected from the substituent group ST1).

[53] The compound according to any one of [1], [4], [5], and [8] to

[52] , wherein L is a C1-6 alkyl group, or a benzene that is unsubstituted or substituted with one or two C1-3 alkyl groups or C1-3 alkoxy groups.

[54] The compound according to any one of [1], [4], [5], and [8] to

[53] , wherein L is a C1-6 alkyl group.

[55] The reactive functional group of D is The compound according to any one of [1], [4], [5], and [8] to

[54] , which is a reactive functional group capable of forming a C-C, amino, ether, carbonyl, amide, ester, urea, sulfide, disulfide, sulfoxide, sulfonamide, or sulfonyl bond.

[56] The compound according to any one of [1], [4], [5], and [8] to

[55] , wherein the reactive functional group of D is a C1 hydrocarbon having a leaving group, an amino group, a hydroxyl group, a precursor of a carbonyl group, a thiol group, or an aldehyde group.

[57] The compound according to any one of [1], [4], [5], and [8] to

[56] , wherein the reactive functional group of D is a C1 hydrocarbon having a halogen atom, a C1 hydrocarbon having a sulfonic acid leaving group, an amino group, a hydroxyl group, a carboxy group, a halogenated carboxy group, a thiol group, or an aldehyde group.

[58] The compound according to any one of [1], [4], [5], and [8] to

[57] , wherein the reactive functional group of D is -CH2Cl, -CH2Br, -CH2OSO2CH3, -CH2OSO2CF3, an amino group, a hydroxyl group, or a carboxy group.

[59] The compound according to any one of [1], [4], [5], and [8] to

[58] , wherein the reactive functional group of D is a primary amino group.

[60] The selectively cleavable site is a deoxyribonucleoside other than deoxyadenosine, deoxyguanosine, thymidine, and deoxycytidine. The compound according to any one of [1], [4], [5], and [8] to

[59] .

[61] The selectively cleavable site is deoxyuridine, bromodeoxyuridine, deoxyinosine, 8-hydroxydeoxyguanosine, 3-methyl-2'-deoxyadenosine, N6-etheno-2'-deoxyadenosine, 7-methyl-2'-deoxyguanosine, 2'-deoxyxanthosine, or 5,6-dihydroxy-5,6-dihydrodeoxythymidine. The compound according to any one of [1], [4], [5], and [8] to

[60] .

[62] The selectively cleavable site is deoxyuridine or deoxyinosine. The compound according to any one of [1], [4], [5], and [8] to

[61] .

[63] The selectively cleavable site is deoxyuridine. The compound according to any one of [1], [4], [5], and [8] to

[62] .

[64] The selectively cleavable site is deoxyinosine. The compound according to any one of [1], [4], [5], and [8] to

[62] .

[65] The selectively cleavable site is the second phosphodiester bond 3' from the deoxyinosine, The compound according to any one of [1], [4], [5], and [8] to

[59] .

[66] The selectively cleavable site is a ribonucleoside. The compound according to any one of [1], [4], [5], and [8] to

[59] .

[67] The compound has one selectively cleavable site. The compound according to any one of [1], [4], [5], and [8] to

[66] .

[68] At least one cleavable site is included in E or (LP1)p and at least one cleavable site is included in F or (LP2)q; The compound according to any one of [1], [4], [5], and [8] to

[66] .

[69] The compound according to

[68] , wherein the cleavable site contained in E or (LP1)p and the cleavable site contained in F or (LP2)q are cleavable under different conditions.

[70] The compound according to any one of [8] to

[69] , wherein An is a partial structure constructed from n building blocks α1 to αn (n is an integer of 1 to 10).

[71] The compound according to any one of [8] to

[70] , wherein An is a low molecular weight organic compound.

[72] The compound according to any one of [8] to

[71] , wherein the building block of An is a compound having a molecular weight of 500 or less.

[73] The compound according to any one of [8] to

[72] , wherein the building block of An is a compound having a molecular weight of 300 or less.

[74] The compound according to any one of [8] to

[73] , wherein the building block of An is a compound having a molecular weight of 150 or less.

[75] The compound according to any one of [8] to

[74] , wherein An is an organic compound composed of elements selected, alone or differently, from the group consisting of H, B, C, N, O, Si, P, S, F, Cl, Br and I.

[76] The compound according to any one of [8] to

[75] , wherein An is a low molecular weight organic compound having substituents, either alone or differently, selected from the group consisting of an aryl group, a non-aromatic cyclyl group, a heteroaryl group, and a non-aromatic heterocyclyl group.

[77] The compound according to any one of [8] to

[76] , wherein An has a molecular weight of 5,000 or less.

[78] The compound according to any one of [8] to

[77] , wherein An has a molecular weight of 800 or less.

[79] The compound according to any one of [8] to

[78] , wherein An has a molecular weight of 500 or less.

[80] The compound according to any one of [8] to

[70] , wherein An is a polypeptide.

[81] The compound according to any one of [8] to

[80] , wherein Sp is a bond.

[82] Sp is a bifunctional spacer; the bifunctional spacer is SpD-SpL-SpX, SpD is a divalent group derived from a reactive group capable of forming a C-C, amino, ether, carbonyl, amide, ester, urea, sulfide, disulfide, sulfoxide, sulfonamide, or sulfonyl bond; SpL is a polyalkylene glycol, a polyethylene, a C1-20 aliphatic hydrocarbon optionally substituted with a heteroatom, a peptide, an oligonucleotide, or a combination thereof; SpX is a divalent group derived from a reactive group forming an amide, amino, or sulfonamide bond; The compound according to any one of [8] to

[80] .

[83] Sp is a bifunctional spacer; the bifunctional spacer is SpD-SpL-SpX, SpD is a divalent group derived from a primary amino group, SpL is polyethylene glycol or polyethylene; SpX is a divalent group derived from a carboxy group. The compound according to any one of [8] to

[81] .

[84] The compound according to any one of [8] to

[83] , wherein the oligonucleotide strand X and the oligonucleotide strand Y are sequences capable of forming a double strand.

[85] The compound according to any one of [8] to

[84] , wherein the oligonucleotide strand X and the oligonucleotide strand Y contain complementary base sequences.

[86] The compound according to any one of [8] to

[85] , wherein the oligonucleotide chain X and the oligonucleotide chain Y each have a length of 1 to 200 bases.

[87] The compound according to any one of [8] to

[86] , wherein the oligonucleotide chain X and the oligonucleotide chain Y each have a length of 3 to 150 bases.

[88] The compound according to any one of [8] to

[87] , wherein the oligonucleotide chain X and the oligonucleotide chain Y each have a length of 30 to 150 bases.

[89] The compound according to any one of [8] to

[88] , wherein the oligonucleotide chain X and the oligonucleotide chain Y have blunt ends.

[90] The compound according to any one of [8] to

[88] , wherein the oligonucleotide chain X and the oligonucleotide chain Y have protruding ends.

[91] The compound according to

[90] , wherein the overhang of the cohesive end has a length of 1 to 30 bases.

[92] The compound according to

[90] or

[91] , wherein the overhang of the cohesive end has a length of 2 to 5 bases.

[93] The compound according to any one of

[90] to

[92] , wherein the oligonucleotide strand X and the oligonucleotide strand Y have protruding ends, and a specific molecular recognition sequence is further bound to the protruding ends.

[94] The compound according to any one of [8] to

[93] , wherein a functional molecule is bound to either X or Y.

[95] The compound according to any one of [8] to

[93] , wherein biotin is bound to either X or Y.

[96] A compound library comprising the compound according to any one of [1], [4], [5], [8] to

[95] .

[97] A DNA-encoded library comprising any one of the compounds described in [1], [4], [5], [8]-

[95] .

[98] A library according to

[96] or

[97] , which is composed of 1,000 or more different compounds.

[99] A method for producing a compound An-Sp-C-Bn, comprising the steps of: An is a partial structure constructed by n building blocks α1 to αn (n is an integer of 2 to 10), Sp is a bond or a bifunctional spacer; C is a hairpin-shaped headpiece having at least one "selectively cleavable site"; Bn is a partial structure containing an oligonucleotide containing a base sequence capable of identifying the structure of An, For C, the following steps: (a) binding α1-Sp, or binding Sp and α1, and (b) binding an oligonucleotide tag containing a base sequence capable of identifying the structure of α1; to obtain compound A1-Sp-C-B1 by Then, repeating the following steps (c) and (d) on A(m-1)-Sp-CB(m-1) (m is an integer of 2 to n) in ascending order from 2 to n; (c) attaching αn to the A moiety; and (d) binding an oligonucleotide tag containing a base sequence capable of identifying the structure of αn to the end of the B portion; obtaining the compound Am-Sp-C-Bm by The method, wherein steps (a) and (b), and steps (c) and (d) may be performed in any order.

[0100] A method for producing the compound An-Sp-C-Bn according to any one of [9] to

[95] , comprising the steps of: An is a partial structure constructed by n building blocks α1 to αn (n is an integer of 2 to 10), Sp is a bond or a bifunctional spacer; C is a hairpin-shaped headpiece having at least one "selectively cleavable site"; Bn is a partial structure containing an oligonucleotide containing a base sequence capable of identifying the structure of An, For C, the following steps: (a) binding α1-Sp, or binding Sp and α1, and (b) binding an oligonucleotide tag containing a base sequence capable of identifying the structure of α1; to obtain compound A1-Sp-C-B1 by Then, repeating the following steps (c) and (d) on A(m-1)-Sp-CB(m-1) (m is an integer of 2 to n) in ascending order from 2 to n; (c) attaching αn to the A moiety; and (d) binding an oligonucleotide tag containing a base sequence capable of identifying the structure of αn to the end of the B portion; obtaining the compound Am-Sp-C-Bm by The method, wherein steps (a) and (b), and steps (c) and (d) may be performed in any order.

[0101] A method for producing the compound An-Sp-C-Bn (An, Sp, C and Bn have the same meanings as defined above) according to any one of [9] to

[95] , comprising the steps of: For C, the following steps: (a) binding α1-Sp, or binding Sp and α1, and (b) binding an oligonucleotide tag containing a base sequence capable of identifying the structure of α1; to obtain compound A1-Sp-C-B1 by Then, repeating the following steps (c) and (d) on A(m-1)-Sp-CB(m-1) (m is an integer of 2 to n) in ascending order from 2 to n; (c) attaching αn to the A moiety; and (d) binding an oligonucleotide tag containing a base sequence capable of identifying the structure of αn to the end of the B portion; obtaining the compound Am-Sp-C-Bm by The method, wherein steps (a) and (b), and steps (c) and (d) may be performed in any order.

[0102] At least one compound of formula (III) An-Sp-C-Bn (III) (In the formula, An is a partial structure constructed by n building blocks α1 to αn (n is an integer of 1 to 10), Sp is a bond or a bifunctional spacer; C is a hairpin-shaped headpiece having at least one "selectively cleavable site"; Bn is a partial structure containing an oligonucleotide containing a base sequence that can identify the structure of An. A method for evaluating a compound library comprising a compound represented by the formula: (1) contacting the compound library with a biological target under conditions suitable for at least one library molecule of the compound library to bind to the target; (2) removing library molecules that do not bind to the target and selecting library molecules that have affinity for the biological target; (3) selectively cleaving the cleavable site; (4) Identifying the sequences of the oligonucleotides that make up Bn; (5) using the sequence determined in (4) to identify the structure of one or more compounds that bind to the biological target; The method comprises:

[0103] At least one compound of formula (III) An-Sp-C-Bn (III) (In the formula, An is a partial structure constructed by n building blocks α1 to αn (n is an integer of 1 to 10), Sp is a bond or a bifunctional spacer; C is a hairpin-shaped headpiece having at least one "selectively cleavable site"; Bn is a partial structure containing an oligonucleotide containing a base sequence that can identify the structure of An. A method for evaluating a compound library comprising the compound according to any one of [8] to

[92] , comprising the steps of: (1) contacting the compound library with a biological target under conditions suitable for at least one library molecule of the compound library to bind to the target; (2) removing library molecules that do not bind to the target and selecting library molecules that have affinity for the biological target; (3) selectively cleaving the cleavable site; (4) Identifying the sequences of the oligonucleotides that make up Bn; (5) using the sequence determined in (4) to identify the structure of one or more compounds that bind to the biological target; The method comprises:

[0104] A method according to

[0102] or

[0103] , comprising a step of amplifying the oligonucleotides constituting Bn between steps (3) and (4).

[0105] A method according to any one of

[0102] to

[0104] , wherein the step of cleaving the selectively cleavable site is a step of cleaving the selectively cleavable site with an enzyme.

[0106] cleaving the selectively cleavable site A method according to any one of

[0102] to

[0104] , comprising a step of cleaving a selectively cleavable site by a combination of an enzyme and a change in chemical conditions.

[0107] The method according to

[0105] or

[0106] , wherein the enzyme is at least one selected from glycosylase and nuclease.

[0108] The method described in

[0107] , wherein the enzyme is uracil DNA glycosylase.

[0109] The method described in

[0107] , wherein the enzyme is endonuclease VIII.

[0110] The method described in

[0107] , wherein the enzyme is a combination of uracil DNA glycosylase and endonuclease VIII.

[0111] The method described in

[0107] , wherein the enzyme is alkyladenine DNA glycosylase.

[0112] The method described in

[0107] , wherein the enzyme is endonuclease V.

[0113] The method according to any one of

[0106] to

[0112] , wherein the change in chemical condition is heating at 50 to 100°C in a solution containing water.

[0114] The method according to any one of

[0103] to

[0113] , wherein the change in chemical condition is heating at 80 to 95°C in a solution containing water.

[0115] A method according to any one of

[0106] to

[0114] , wherein the change in chemical condition is a basic condition of pH 8 to 13.

[0116] A method according to any one of

[0106] to

[0115] , wherein the change in chemical conditions is a basic condition of pH 8 to 11.

[0117] A method according to any one of

[0106] to

[0116] , wherein the change in chemical condition is a basic condition of pH 9 to 10.

[0118] A method according to any one of

[0102] to

[0117] , comprising providing a cleavable site near the end of a DNA tag, cleaving the site as desired to generate a new sticky end, ligating a specific molecular recognition sequence to the sticky end, and identifying the sequence of an oligonucleotide constituting Bn.

[0119] The method according to

[0118] , wherein the cleavable site located near the end of the DNA tag and the cleavable site contained in C are cleaved under different conditions.

[0120] A method in which a nucleic acid that binds to a compound having a cleavable site and a hairpin structure is used, and the cleavable site is cleaved to produce a double-stranded nucleic acid.

[0121] The method described in

[0120] above, which uses a nucleic acid that is chemically more stable than double-stranded nucleic acid and binds to a compound having a cleavable site and a hairpin structure, and utilizes it as a double-stranded nucleic acid by cleaving the cleavable site.

[0122] A method according to

[0120] or

[0121] , which uses a nucleic acid that binds to a compound having a cleavable site and a hairpin structure, converts the chemical structure of the compound, and then cleaves the cleavable site to use it as a double-stranded nucleic acid.

[0123] A method according to any one of

[0120] to

[0122] , comprising using a nucleic acid that binds to a compound having a cleavable site and a hairpin structure, and then subjecting the nucleic acid to a chemical structure conversion, and cleaving the cleavable site to use the nucleic acid as a double-stranded nucleic acid.

[0124] A method according to any one of

[0120] to

[0123] , comprising using a nucleic acid that binds to a compound having a cleavable site and a hairpin structure, subjecting the nucleic acid to a nucleic acid extension reaction, and then cleaving the cleavable site to use the nucleic acid as a double-stranded nucleic acid.

[0125] A method according to any one of

[0120] to

[0124] , comprising using a nucleic acid that binds to a compound having a cleavable site and a hairpin structure, cleaving the cleavable site to make it available as a double-stranded nucleic acid, and carrying out a PCR reaction.

[0126] A method according to any one of

[0120] to

[0125] , used for evaluating the functionality of a compound.

[0127] A method according to any one of

[0120] to

[0126] , used for evaluating the biological activity of a compound.

[0128] A method according to any one of

[0120] to

[0127] used for DEL.

[0129] A method according to any one of

[0120] to

[0124] , used for producing DEL.

[0130] A method for converting a DEL compound synthesized using a nucleic acid that binds to a compound having a cleavable site and a hairpin structure into a DEL having single-stranded DNA by cleaving the cleavable site.

[0131] A method in which a DEL compound is synthesized using a nucleic acid that binds to a compound having a cleavable site and a hairpin structure, the cleavable site is cleaved, the compound is converted into a DEL having single-stranded DNA, and the compound is allowed to form a double strand with a crosslinker-modified DNA.

[0132] A method for synthesizing a DEL compound using a nucleic acid that binds to a compound having a cleavable site and a hairpin structure, cleaving the cleavable site, attaching a cross-linker-modified primer, and extending the attached primer to synthesize a cross-linker-modified double-stranded DEL compound. Effect of the Invention

[0009] The present invention provides a DEL that contains a cleavable site in the DNA strand, and a composition for synthesizing the same, making it possible to produce a DEL with greater convenience than conventional methods. [Brief description of the drawings]

[0010] [Figure 1] The exemplary DEL production method of mode 1 is shown. Starting from a headpiece including a first oligonucleotide strand containing a cleavable site in the DNA strand, a loop site and a second oligonucleotide strand, the binding of building blocks and the double-stranded ligation of the oligonucleotide tag corresponding to the building block are repeated (three times in FIG. 1), and further the double-stranded ligation of the oligonucleotide tag containing a primer region is performed as desired to produce DEL. [Diagram 2] The following shows an exemplary method of using DEL in mode 1. For DEL containing a cleavable site in the first oligonucleotide strand of the headpiece, the cleavable site can be cleaved using a cleavage means such as an enzyme to induce double-stranded oligonucleotides that are not linked at the loop site, thereby enabling PCR to be performed with high efficiency. [Diagram 3] The following shows an exemplary method of using DEL in mode 2. For DEL containing a cleavable site in the second oligonucleotide strand of the headpiece, the cleavable site can be cleaved using a cleavage means such as an enzyme to induce double-stranded oligonucleotides that are not linked at the loop site, thereby enabling PCR to be performed with high efficiency. [Figure 4]The following shows an exemplary method of using DEL in Format 3. For DEL containing cleavable sites in the first oligonucleotide strand and the second oligonucleotide strand of the headpiece, both cleavable sites are cleaved using a cleavage means such as an enzyme to induce a double-stranded oligonucleotide without a loop site, thereby enabling PCR to be performed with high efficiency. [Diagram 5] The following shows an exemplary method of using DEL in mode 4. For DEL containing two different cleavable sites in the first oligonucleotide strand and the second oligonucleotide strand of the headpiece, the cleavage conditions can be selected to selectively cleave either the first oligonucleotide strand or the second oligonucleotide strand. [Figure 6] An exemplary method of using DEL in Format 5 is shown below. A cleavable site is provided near the end of the DNA tag, and a new protruding end can be generated by cleaving the site as desired. The protruding end can be used as a sticky end to ligate a desired nucleic acid sequence, such as UMIs (unique molecular identifiers), and can be given a new function. [Figure 7] An exemplary method of using DEL in format 6 is shown. In the present invention, a cleavable site can be used in combination with a modifying group or a functional molecule, and for example, a DEL in which a hairpin strand DNA is converted into a single-stranded DNA can be prepared. For example, a double-stranded oligonucleotide chain having a functional molecule (e.g., biotin) at the 3' end is ligated to a synthesized DEL compound (A), the cleavable site is cleaved (B), and a treatment according to the function of the functional molecule is applied (C). For example, when the functional molecule is biotin, streptavidin beads having affinity for biotin are used to selectively remove the oligonucleotide chain to which biotin is bound from the system. This makes it possible to obtain a DEL having a single-stranded DNA. [Figure 8]The following shows an exemplary use of the DEL obtained in Mode 6. The DEL having the single-stranded DNA obtained in Mode 6 can be given a new function by forming a double strand with a modified oligonucleotide having a desired functional site (e.g., a crosslinker-modified DNA such as a photoreactive crosslinker). [Figure 9] An exemplary method of using DEL in format 7 is shown. In the present invention, a crosslinker can be introduced using a cleavable site. The cleavable site is cleaved from the synthesized DEL compound (A), a modified primer is added (B), and a crosslinker-modified double-stranded DEL compound can be synthesized based on the added primer (C). The crosslinker-modified double-stranded DEL compound can significantly improve the detection sensitivity in screening a DEL library (see Non-Patent Documents 5, 6, etc.). [Figure 10] FIG. 1 is a graph showing the conversion rate of the cleavage reaction at each incubation time when verifying the cleavage reaction by USER (registered trademark) enzyme of hairpin-type DEL partial structures containing deoxyuridine (10 types: U-DEL1-sh, U-DEL2-sh, U-DEL3-sh, U-DEL4-sh, U-DEL5-HP, U-DEL6-HP, U-DEL7-HP, U-DEL8-HP, U-DEL9-HP, and U-DEL10-HP) in Example 1. [Figure 11] FIG. 1 is a schematic diagram showing the synthesis procedure of various hairpin DELs (U-DEL1, U-DEL2, U-DEL4, U-DEL7, U-DEL8, U-DEL9, U-DEL10, H-DEL, U-DEL5, U-DEL11, U-DEL12, U-DEL13, I-DEL1, I-DEL2, I-DEL3, R-DEL1, and BIO-DEL) in Examples 2, 3, 4, 5, and 7. The hairpin DEL synthesis is achieved by two-step double-stranded ligation with the double-stranded oligonucleotide Pr_TAG and CP using the corresponding headpiece as a raw material. [Figure 12]FIG. 1 is a graph showing the Ct values ​​measured by real-time PCR for eight hairpin DELs (U-DEL1, U-DEL2, U-DEL4, U-DEL7, U-DEL8, U-DEL9, U-DEL10, and H-DEL) and double-stranded DEL (DS-DEL) for each sample amount in Example 2. Samples in which various DELs were treated with USER® enzyme are indicated as "USER(+)", and untreated samples are indicated as "USER(-)". Cleavable hairpin DELs containing deoxyuridine (U-DEL1, U-DEL2, U-DEL4, U-DEL7, U-DEL8, U-DEL9, and U-DEL10) show Ct values ​​equivalent to those of double-stranded DEL (DS-DEL) after treatment with USER® enzyme. [Figure 13] This is an image of a gel obtained by denaturing polyacrylamide gel electrophoresis, showing the progress of the cleavage reaction of hairpin DELs containing six kinds of deoxyuridines (U-DEL5, U-DEL7, U-DEL9, U-DEL11, U-DEL12, and U-DEL13) by USER (registered trademark) enzyme in Example 3. The numbers in the figure indicate the numbers of each lane. [Figure 14] This is an image of a gel obtained by denaturing polyacrylamide gel electrophoresis, showing the progress of the cleavage reaction of hairpin DELs (I-DEL1, I-DEL2, I-DEL3, and I-DEL4) containing four types of deoxyinosine by endonuclease V in Example 4. The numbers in the figure indicate the numbers of each lane. [Figure 15] This is an image of a gel obtained by denaturing polyacrylamide gel electrophoresis, showing the progress of the cleavage reaction of a ribonucleoside-containing hairpin DEL (R-DEL1) by RNase HII in Example 5. The numbers in the figure indicate the numbers of each lane. [Figure 16]This is a schematic diagram showing the synthesis procedure of a model library containing 3 × 3 × 3 (27) compound species using U-DEL9-HP as a starting material. In Example 6, the synthesis of the model library is achieved by three split-and-pool steps (cycles A, B, and C) using U-DEL9-HP as a starting material. Each cycle also includes a ligation reaction of a double-stranded oligonucleotide tag and a chemical reaction for introducing a building block. [Figure 17] 1 is an image of a gel obtained by agarose gel electrophoresis showing the progress of the ligation reaction for each cycle in the synthesis of the model library in Example 6. The numbers in the figure indicate the numbers of each lane. [Figure 18] 18A is a chromatograph obtained from a sample after completion of cycle C in the model library synthesis of Example 6. FIG 18B is a deconvolution result of an MS spectrum obtained from a sample after completion of cycle C in the model library synthesis of Example 6. [Figure 19] This is an image of a gel obtained by denaturing polyacrylamide gel electrophoresis, showing the progress of the cleavage reaction of the model library by the USER (registered trademark) enzyme in Example 6. The numbers in the figure indicate the numbers of each lane. [Figure 20] This is an image of a gel obtained by denaturing polyacrylamide gel electrophoresis showing the progress of the cleavage reaction of the DEL compound "BIO-DEL" having biotin at the 3' end by USER (registered trademark) enzyme in Example 7. The numbers in the figure indicate the numbers of each lane. [Figure 21] This is an image of a gel obtained by polyacrylamide gel electrophoresis, showing the results of a primer extension reaction carried out using a DEL compound "SS-DEL" having a single-stranded DNA and a photoreactive crosslinker-modified primer "PXL-Pr" in Example 7. The numbers in the figure indicate the numbers of each lane. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] As mentioned above and a concept well known to those skilled in the art, in the present invention, a compound library means a group of compound derivatives systematically collected from compounds that may have a specific activity, such as candidate pharmaceutical compounds. This compound library is often synthesized based on the synthesis techniques and methodologies of combinatorial chemistry. Combinatorial chemistry is an experimental method for efficiently synthesizing a wide variety of compounds through a systematic synthesis route from a series of compound libraries that are enumerated and designed based on combinatorial theory, and a research field related to the method.

[0012] As mentioned above and known to those skilled in the art, one type of chemical library based on combinatorial chemistry is a DNA-encoded library. The DNA-encoded library is conveniently abbreviated as DEL. DEL is also essentially synonymous with DNA-encoded chemical library. In the present invention, a DNA-encoded library refers to a library in which each compound in the library is tagged with a DNA tag whose sequence is designed to identify each structure of each compound, and which functions as a label for the compound.

[0013] A nucleotide is generally understood as a substance in which a phosphate group is bound to a nucleoside. Nucleotides and nucleosides are terms well known to those skilled in the art, and in one general embodiment, a nucleoside is understood as a nucleic acid base such as a purine base or a pyrimidine base bound to the first position of a sugar such as a pentose via glycosidic bond. Nucleosides and nucleotides are also units that constitute nucleic acids such as DNA and RNA. Nucleic acids are also a concept well known to those skilled in the art, and are generally understood as polymers of nucleotides. In one embodiment, the nucleic acid of the present invention is a polymer composed of nucleotides and nucleic acid analogs described below.

[0014] In this specification, in addition to nucleic acid polymers composed of nucleotides or nucleic acid analogs, nucleic acid monomers such as nucleotides or nucleic acid analogs may also be referred to simply as nucleic acids. The latter usage is also in accordance with common general technical knowledge, and can be understood by those skilled in the art according to the appropriate context.

[0015] In the broad sense, nucleotides include not only natural nucleotides (original nucleotides) but also artificial nucleotides (various nucleic acid analogues). The broad definition of a nucleotide in the present invention includes the following aspects. (A) Natural nucleoside nucleotides (Examples of such nucleosides include adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxyuridine, deoxyguanosine, deoxycytidine, inosine, or diaminopurine deoxyriboside.) (B) Nucleotides with nucleoside analogues (Examples of nucleosides having such nucleic acid base analogs include 2-aminoadenosine, 2-thiothymidine, pyrrolopyrimidine deoxyriboside, 3-methyladenosine, C5-propynylcytidine, C5-propynyluridine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-methylcytidine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, 6-O-methylguanosine, and 2-thiocytidine.) (C) A nucleotide having an intercalated nucleobase. (D) Unnatural nucleotides containing ribose or 2'-deoxyribose (E) A nucleotide having a modified sugar in the sugar portion. (Examples of such modified sugars include modified ribose, modified 2'-deoxyribose, 2'-O-methylribose, 2'-fluororibose, D-threoninol, arabinose, hexose, anhydrohexitol, altritol, or mannitol.) (F) Nucleic acid analogues (Examples of such nucleic acid analogs include cyclohexanyl nucleic acid, cyclohexenyl nucleic acid, morpholino nucleic acid (PMO), locked nucleic acid (LNA), glycol nucleic acid (GNA), threose nucleic acid (TNA), serinol nucleic acid (SNA), acyclic threoninol nucleic acid (aTNA), or nucleic acid in which oxygen in ribose is replaced.) Each nucleic acid analog will be described in detail below. (F1)PMO PMOs are nucleic acid analogues that have a morpholine ring in the sugar moiety and an uncharged phosphorodiamidate structure in the phosphodiester moiety. (F2)LNA LNA is a nucleic acid analogue with a bridged structure at the sugar moiety, most typically the 2'-hydroxyl of ribose is bridged to the 4'-carbon of the same ribose sugar by a C1-6 alkylene or C1-6 heteroalkylene. Examples of bridged structures include methylene, propylene, ether or amino bridged structures. Exemplary LNAs include 2',4'-BNAs (2'-O,4'-C-methano bridged nucleic acids). (F3)GNA Glycol Nucleic Acids are also called GNAs, such as R-GNAs or S-GNAs, in which the ribose is replaced by a glycol unit linked to a phosphodiester bond. (F4)TNA Threose Nucleic Acid (TNA) is a nucleic acid in which the ribose is replaced by α-L-threofuranosyl-(3'→2'). (F5)SNA Serinol Nucleic Acid is also called SNA, in which the ribose is replaced by a serinol unit linked to a phosphodiester bond. (F6)aTNA Acyclic threoninol nucleic acids are also called aTNAs, such as D-aTNA or L-aTNA, in which the ribose is replaced by a threoninol unit linked to a phosphodiester bond. (F7) A sugar in which the oxygen in ribose has been replaced Specific examples include the replacement of oxygen with S, Se, or alkylene (eg, methylene or ethylene). (G) Backbone-modified nucleotides (An example of such a backbone-modified nucleotide is a peptide nucleic acid, also called PNA, in which a 2-aminoethyl-glycine linkage replaces the ribose and phosphodiester backbone.) (H) Nucleotide with modified phosphate group (Examples of nucleotides in which the phosphate group is modified include phosphorothioates, 5'-N-phosphoramidites, phosphoroselenates, boranophosphates, boranophosphate esters, hydrogen phosphonates, phosphoramidates, phosphorodiamidates, alkyl or aryl phosphonates, phosphotriesters, bridged phosphoramidates, bridged phosphorothioates, and bridged methylene-phosphonates.) In the following description, the oligonucleotide, oligonucleotide chain, double-stranded oligonucleotide, double-stranded oligonucleotide chain, and double-stranded DNA of the present invention are nucleotides as defined above.

[0016] In the present invention, when the term "nucleotide" is used without any particular limitation, it means a natural nucleotide. The term "natural nucleotide" is a term well known to those skilled in the art, and is not particularly limited as long as it is a nucleotide that essentially exists in nature. In one embodiment, the natural nucleotide in the present invention is the nucleotide described in (A) above. (Nucleic acid analogues) The term "nucleic acid analog" is well known to those skilled in the art, and the structure of the nucleic acid analog in the present invention is not limited as long as it has the effect of the present invention. In one embodiment, the nucleic acid analog is a compound according to any one of embodiments (B) to (H) above. In one embodiment, the nucleic acid analog of the present invention is a compound having a site corresponding to a phosphate group and a site corresponding to a hydroxyl group in a nucleic acid monomer, more preferably a compound having a phosphate group and a hydroxyl group. In one embodiment, the nucleic acid analog in the present invention is a compound that can be used as a monomer in a nucleic acid synthesizer. As is well known to those skilled in the art, in a nucleic acid synthesizer, a nucleic acid oligomer can be synthesized by converting the phosphoric acid (or a corresponding site) of the nucleic acid analog into a phosphoramidite and using the nucleic acid analog as a monomer in which the hydroxyl group (or a corresponding site) is protected with a protecting group. In addition, the partial structure of a nucleic acid analog other than the phosphate site (or equivalent site) and the hydroxyl group (or equivalent site) can be called a nucleic acid analog residue. The structure of the nucleic acid analog residue is not limited as long as it has the effect of the present invention, but for reference, when checking the characteristics of the structures of natural nucleic acids (deoxyadenosine, thymidine, deoxycytidine, deoxyguanosine), they have molecular weights of about 322 (thymidine monophosphate) to 347 (deoxyguanosine monophosphate), and the number of atoms between the hydroxyl oxygen atom at the 3' position and the phosphorus atom at the 5' position constituting the nucleic acid chain (including oxygen atoms and phosphorus atoms; hereinafter also referred to as the number of atoms between residues) is 6. In addition, the following are known as nucleic acid analogs that can be used in nucleic acid synthesizers. Amino C6 dT Molecular weight: 476, Number of residue atoms: 6 mdC(TEG-Amino) Molecular weight: 526, Number of residue atoms: 6 Uni-Link (trademark) Amino Modifier Molecular weight: 227, Number of residue atoms: 6 (See Nucleic Acid Research, 1992, Vol. 20, pp. 6253-6259.) d-Spacer Molecular weight: 198, Number of residue atoms: 6 Triethylene glycol phosphate ester (Spacer 9) Molecular weight: 230, Number of residue atoms: 11

[0017] For reference, the structures of each nucleic acid analog are shown below. [ka]

[0018] Thus, in one embodiment, the nucleic acid analog is compound (B1) characterized in that: (B11) Having a phosphate group (or an equivalent portion) and a hydroxyl group (or an equivalent portion). (B12) Consisting of carbon, hydrogen, oxygen, nitrogen, phosphorus or sulfur. (B13) The molecular weight is 142 to 1500. (B14) The number of atoms between residues is 5 to 30. (B15) The bonds between the atoms of the residues are all single bonds or contain one or two double bonds and the rest are single bonds.

[0019] In one embodiment, the nucleic acid analog is compound (B2), characterized in that: (B21) Having phosphoric acid and hydroxyl groups. (B22) Composed of carbon, hydrogen, oxygen, nitrogen or phosphorus. (B23) The molecular weight is 142 to 1000. (B24) The number of atoms between residues is 5 to 20. (B25) The bonding modes of atoms between the residues are all single bonds.

[0020] In one embodiment, the nucleic acid analog is compound (B3), characterized in that: (B31) Having phosphoric acid and hydroxyl groups. (B32) Consisting of carbon, hydrogen, oxygen, nitrogen or phosphorus. (B33) The molecular weight is 142 to 700. (B34) The number of atoms between residues is 5 to 12. (B35) The bonds between the atoms of the residues are all single bonds.

[0021] In one embodiment, the nucleic acid analog is the following compound (B41), (B42), (B43), (B44), (B5), (B51), or (B52). (B41)d-Spacer (B42) Amino C6 dT (B43)mdC(TEG-Amino) (B44) Uni-Link (trademark) Amino Modifier (B5) Polyalkylene glycol phosphate ester (B51) Diethylene glycol phosphate or triethylene glycol phosphate (B52) Triethylene glycol phosphate ester

[0022] In the present invention, an oligonucleotide and an oligonucleotide chain refer to a polymer of nucleotides having one or more nucleotides at the 5'-terminus, the 3'-terminus, and at an internal position between the 5'-terminus and the 3'-terminus.

[0023] A mutually complementary base sequence refers to a sequence of nucleotides that can form a fixed pair of adenine and thymine (or uracil) or guanine and cytosine between two oligonucleotides of a nucleic acid, which is called a complementary base pair and is connected by hydrogen bonds. The formation of a complementary base pair is also called hybridization. The complementary base pair is generally a concept called "Watson-Crick base pair" or "natural base pair". However, the base pair may be a Watson-Crick type, a Hoogsteen type base pair, or a base pair formed by other hydrogen bond motifs (e.g., diaminopurine and T, 5-methyl C and G, 2-thiothymine and A, 6-hydroxypurine and C, pseudoisocytosine and G). There is no restriction on the sequence of the "mutually complementary base sequence" as long as it is a sequence in which two oligonucleotides can form a double strand and can be used for the purpose of the present invention, and there is also no restriction on the homology of the two sequences. The homology is preferably 99% or more, 98% or more, 95% or more, 90% or more, 85% or more, 80% or more, 70% or more, 60% or more, or 50% or more, in order of preference.

[0024] To reiterate, in the present invention, hybridization refers to the act of forming a double strand between oligonucleotides or oligonucleotide chains containing complementary base sequences, and the phenomenon in which oligonucleotides or oligonucleotide chains containing complementary sequences form a double strand.

[0025] In the present invention, the term "double strand" refers to a state in which two nucleic acid strands form complementary base pairs (hybridize). The two nucleic acid strands may be derived from two nucleic acid strands or may be derived from two nucleic acid sequences within a single nucleic acid strand molecule.

[0026] In the present invention, the double-stranded oligonucleotide and the double-stranded oligonucleotide chain refer to a secondary structure formed by hybridization of two or more different oligonucleotide chains. The chain lengths of the two oligonucleotides may be different, and they may have a non-hybridized region. The region where the two strands hybridize is called a double strand.

[0027] In the present invention, double-stranded DNA refers to a secondary structure formed by hybridization of two different DNA strands. The lengths of the DNA strands may be different, and they may have non-hybridized regions. The DNA strand is not limited to naturally occurring deoxyribonucleotides, but refers to any oligonucleotide strand that can be amplified by DNA polymerase.

[0028] In the present invention, "forming a double strand" refers to forming a double strand under standard conditions for handling an oligonucleotide, such as a temperature of 4 to 40° C., an aqueous solvent, and a pH of 4 to 10. For example, even if a double strand is not formed due to a specific solvent or condition, the nucleic acid is a double-stranded nucleic acid if it forms a double strand under standard conditions.

[0029] In the present invention, the Tm value refers to the temperature at which half of the DNA molecules anneal with a complementary strand.

[0030] In the present invention, blunt ends means that the ends of a double-stranded oligonucleotide are paired with neither end protruding.

[0031] In the present invention, a protruding end means that one of the ends of a double-stranded oligonucleotide has a protruding portion. The protruding portion of the protruding end can be of any length, but is preferably 1 to 50 bases, more preferably 1 to 30 bases, even more preferably 1 to 15 bases, and most preferably 2 to 6 bases in length. In a specific embodiment, the protruding portion can be used as a hybridization region when carrying out ligation of a sticky end.

[0032] PCR means polymerase chain reaction. PCR is a means of amplifying oligonucleotide chains, and is a technique well known to those skilled in the art. The PCR process is outlined as follows: (1) the double-stranded oligonucleotide chain to be amplified is dissociated into two single strands by heat treatment or the like, and (2) the temperature is adjusted to a suitable temperature for the enzyme reaction, and then a complementary strand is synthesized for each single strand by an enzyme (such as DNA polymerase) present in the reaction system. In other words, one double-stranded oligonucleotide can be amplified into two. In PCR, the processes of (1) and (2) are repeated by adjusting the temperature, thereby amplifying the oligonucleotide chain with high efficiency.

[0033] In the present invention, a primer means an oligonucleotide that can be annealed to a template oligonucleotide strand and extended by a polymerase in a template-dependent manner.

[0034] In the present invention, a primer sequence for PCR means a sequence in an oligonucleotide chain to which a primer anneals, and is preferably a sequence suitable for PCR as known in the art, and is preferably present at the end of the oligonucleotide chain.

[0035] In the present invention, a nick refers to a portion in a double-stranded oligonucleotide chain where an internucleotide bond is missing and the oligonucleotide chain is broken. The 5' side of the missing portion may or may not have a phosphate group.

[0036] In the present invention, a gap refers to a portion in a double-stranded oligonucleotide chain where one or more consecutive nucleotides are deleted, resulting in the separation of the oligonucleotide chains. The 5' side of the deleted portion may or may not have a phosphate group.

[0037] In the present invention, a hairpin strand is a single-stranded structure in which two complementary nucleic acid strands are connected, and the characteristics of the hairpin strand and the hairpin strand DEL are as described above. The terms "hairpin site," "hairpin structure," and "hairpin type" used in the present invention are understood to be terms derived from the concept of hairpin, which is the same as the above-mentioned "hairpin strand."

[0038] In the present invention, the nucleic acid linking reaction and ligation refer to a reaction for linking the ends of nucleic acids together.

[0039] Enzymatic nucleic acid joining reaction and enzymatic ligation refer to a reaction in which the ends of nucleic acids are joined together using an enzyme.

[0040] The enzyme that can be used in the nucleic acid ligation reaction is, for example, DNA ligase, RNA ligase, DNA polymerase, RNA polymerase, or topoisomerase.

[0041] In one embodiment, DNA ligase is an enzyme that connects the ends of DNA strands with a phosphodiester bond. In one embodiment, DNA ligase is understood as a ligase belonging to EC number: 6.5.1.1 or 6.5.1.2. DNA ligase is also called polydeoxyribonucleotide synthase or polynucleotide ligase. Examples of DNA ligase include DNA ligase I, II, III, IV, and T4 DNA ligase.

[0042] In one embodiment, RNA ligase is an enzyme that connects the ends of RNA chains with a phosphodiester bond. In one embodiment, RNA ligase is understood as a ligase belonging to EC number: 6.5.1.3. In one embodiment, RNA ligase belongs to the family of poly(ribonucleotide):poly(ribonucleotide) ligase. RNA ligase is also called polyribonucleotide synthase or polyribonucleotide ligase.

[0043] In the present invention, chemical ligation refers to a reaction for joining the ends of nucleic acids without the use of enzymes.

[0044] In chemical ligation, the ends of nucleic acids having chemically reactive functional groups react with each other to form a linkage. Examples of functional groups that can undergo chemical reactions include a pair of an optionally substituted alkynyl group and an optionally substituted azide group, a pair of an optionally substituted diene having a 4π-electron system (for example, an optionally substituted 1,3-unsaturated compound such as an optionally substituted 1,3-butadiene, 1-methoxy-3-trimethylsilyloxy-1,3-butadiene, cyclopentadiene, cyclohexadiene, or furan) and an optionally substituted dienophile or an optionally substituted heterodienophile having a 2π-electron system (for example, an optionally substituted alkenyl group or an optionally substituted alkynyl group), a pair of an optionally substituted amino group and a carboxylic acid group, a pair of a phosphorothioate group and an iodo group (for example, a phosphorothioate group at the 3'-terminus and an iodo group at the 5'-terminus), or a pair of a phosphate group and a hydroxyl group (for example, a pair of a phosphate group at the 5'-terminus and a hydroxyl group at the 3'-terminus, or a pair of a hydroxyl group at the 5'-terminus and a phosphate group at the 3'-terminus). Chemical ligation is a concept well known to those skilled in the art, and those skilled in the art can appropriately achieve chemical ligation based on their common technical knowledge. In addition to the above, see Artificial DNA; PNA & XNA, 2014, Vol. 5, e27896, Current Opinion in Chemical Biology, 2015, Vol. 26, pp. 80-88, etc.

[0045] In the present invention, the term "selectively cleavable" means that only a specific site in a certain compound can be selectively cleaved under predetermined conditions without causing any change in the other molecular structures of the compound.

[0046] In the present invention, the term "selectively cleavable site" refers to a site in a certain compound that can be selectively cleaved under specific conditions.

[0047] In one embodiment, a preferred structure of the "selectively cleavable site" in the present invention is a "selectively cleavable nucleic acid". The site may be a site composed of multiple nucleic acids and cleaved by a specific sequence, or a site composed of a single nucleic acid. When the cleavable site is a nucleic acid, it is preferable from the viewpoints of (1) the efficiency of production is good because an established production method such as a nucleic acid synthesizer can be used, and (2) since the reaction conditions for constructing the building block of DEL require that the nucleic acid in the DNA tag portion is not decomposed, if the cleavable site is a nucleic acid, it is not decomposed.

[0048] A more preferred structure of the "selectively cleavable nucleic acid" is a nucleic acid containing a nucleotide not included in the sequence of the DEL DNA tag. If the cleavable site is a nucleotide not included in the sequence of the DNA tag, it can be used without limiting the sequence of the DNA tag to avoid cleavage of the DNA tag portion.

[0049] The nucleic acids used in the DNA tag sequence are preferably deoxyadenosine, deoxyguanosine, thymidine, and deoxycytidine. Therefore, the preferred structure of the selectively cleavable site is a nucleic acid that is neither deoxyadenosine, deoxyguanosine, thymidine, nor deoxycytidine.

[0050] An example of a "selectively cleavable site" is a "nucleotide having a cleavable base". For example, the N-glycosidic bond between the base and sugar moiety of a "nucleotide having a cleavable base" in DEL is cleaved by the action of DNA glycosylase, leaving an abasic site. The phosphodiester bond adjacent to the abasic site is cleaved by a chemical change in conditions (e.g., temperature increase, basic hydrolysis, etc.) or by an enzyme having apurinic / apyrimidinic (AP) endonuclease activity or AP lyase activity (e.g., endonuclease III, endonuclease IV, endonuclease V, endonuclease VI, endonuclease VII, endonuclease VIII, APE1 (human AP endonuclease), Fpg (formamidopyridine-DNA glycosylase), etc.), forming a gap or nick of one base.

[0051] Examples of "nucleotides having a cleavable base" include deoxyuridine, bromodeoxyuridine, deoxyinosine, 8-hydroxydeoxyguanosine, 3-methyl-2'-deoxyadenosine, N6-etheno-2'-deoxyadenosine, 7-methyl-2'-deoxyguanosine, 2'-deoxyxanthosine, and 5,6-dihydroxydeoxythymidine. Other nucleotides having a cleavable base will be obvious to those skilled in the art. By incorporating these "nucleotides having a cleavable base" into DEL and using a DNA glycosylase that specifically recognizes the structure, DEL is selectively debased.

[0052] In the present invention, the DNA glycosylase is any enzyme having glycosylase activity, which recognizes any nucleic acid base in an oligonucleotide, cleaves the N-glycosidic bond between the base and the sugar, and creates an abasic site. Examples include uracil DNA glycosylase (recognizes deoxyuridine), alkyladenine DNA glycosylase (recognizes 3-methyl-2'-deoxyadenosine, 7-methyl-2'-deoxyguanosine, and deoxyinosine), Fpg (recognizes 8-hydroxydeoxyguanosine), endonuclease VIII (recognizes decomposed pyrimidine bases such as 5,6-dihydroxydeoxythymidine and uracil glycol), and SUMG1 (abbreviation for single-strand selective uracil DNA glycosylase, which recognizes deoxyuridine).

[0053] More preferred examples of the "selectively cleavable site" in the present invention include deoxyinosine and deoxyuridine.

[0054] A particularly preferred example of the "selectively cleavable site" in the present invention is deoxyuridine.

[0055] In one embodiment, the "selectively cleavable site" of the present invention is preferably cleaved using an enzyme. Enzymes generally have high substrate specificity and do not recognize the DNA tag portion of DEL or the compound portion constructed by multiple building blocks as substrates, but only recognize and act on the "selectively cleavable site", which is preferable. In addition, cleavage using the enzyme may be achieved by changing the chemical conditions after structurally changing the "selectively cleavable site" with the enzyme. Examples of such enzymes include glycosylase and nuclease.

[0056] In the present invention, glycosylase is an enzyme that has the function of hydrolyzing glycosidic bonds (covalent bonds formed by dehydration condensation between a sugar molecule and another organic compound). Among them, DNA glycosylase is an enzyme that recognizes the nucleic acid base portion in an oligonucleotide and hydrolyzes the glycosidic bond, as described above.

[0057] In the present invention, a nuclease is an enzyme that has the function of hydrolyzing the phosphodiester bond between the sugar and phosphate of a nucleic acid. Examples of nucleases include AP endonucleases, nicking endonucleases, and ribonucleases.

[0058] As described above, AP endonuclease cleaves the phosphodiester bond adjacent to the abasic site generated by the action of any DNA glycosylase. Therefore, in the present invention, it is preferable to use a DNA glycosylase and an AP endonuclease in combination.

[0059] Nicking endonucleases (e.g., Nb.BbvCI, Nb.BsmI, Nb.BsrDI, etc.) recognize specific DNA sequences and generate a nick by cleaving the phosphodiester bond in only one of the two strands. Endonuclease V can generate a nick by cleaving the second phosphodiester bond in the 3' direction from deoxyinosine, and is useful in carrying out the present invention.

[0060] Ribonucleases are enzymes that degrade RNA. In the present invention, ribonucleosides are used as "selectively cleavable sites" and can be used by allowing ribonucleases to act on them. RNase HII, a type of ribonuclease, can generate nicks by cleaving the phosphodiester bond at the 5' end of a ribonucleotide incorporated in a DNA sequence, and is useful in carrying out the present invention.

[0061] In the present invention, USER (registered trademark) means "Uracil-Specific Excision Reagent" Enzyme. USER is an endonuclease cocktail that contains uracil DNA glycosylase (UDG) and endonuclease VIII to remove uracil. USER removes uracil from double-stranded DNA to create a one-base gap and cleave the DNA strand. In the USER process, UDG first removes the uracil base to create an abasic site. The endonuclease then breaks the phosphodiester bond to release the base-free deoxyribose, creating a one-base gap. In the description of this specification, USER® Enzymes and USER® Enzymes are USER® as defined above.

[0062] In the present invention, a building block is a moiety that has a functional group and can constitute a part of a compound, and may be in the form of a compound.

[0063] In the present invention, the base sequence capable of identifying each building block means a specific base sequence designed to correspond to the structure of each building block. Designing a sequence means assigning a nucleic acid base sequence to each structure, for example, the nucleic acid base sequence AAA to building block structure A, the nucleic acid base sequence TTT to structure B, and the nucleic acid base sequence CGC to structure C. The sequence can be freely designed as long as the object of the present invention is achieved. For example, any number of base sequences can be assigned to one building block.

[0064] In the present invention, the oligonucleotide tag is a partial structure containing an oligonucleotide containing a base sequence capable of identifying the structure of a partial structure constructed by building blocks. In the present invention, the oligonucleotide tag may be an oligonucleotide corresponding to each building block, or may be a longer oligonucleotide containing oligonucleotides corresponding to multiple building blocks. The nucleotides constituting the oligonucleotide tag of the present invention are not limited as long as they achieve the effects of the present invention, but from the viewpoint of ease of amplification by PCR and analysis by a sequencer, it is desirable that the nucleotides are suitable for these operations. Examples of such preferred nucleotides include nucleotides having the above-mentioned natural nucleic acid bases as the base moiety and the above-mentioned ribose or 2'-deoxyribose as the sugar moiety, and more preferred examples include deoxyadenosine, thymidine, deoxycytidine, and deoxyguanosine.

[0065] (Headpiece) In the present invention, the headpiece refers to a starting compound for producing a compound library such as DEL. The structure of the headpiece of the present invention is not limited as long as it achieves the object of the present invention, but in the most typical embodiment, it has at least one site to which a building block can be linked and at least one site to which an oligonucleotide tag can be linked, and further contains at least one selectively cleavable site in the structure. As described below, the DNA tag is preferably a double-stranded oligonucleotide chain, and there are preferably two sites to which the oligonucleotide tag can be linked.

[0066] In one embodiment, the headpiece is a compound as shown in the schematic diagram below. [ka]

[0067] In one aspect, it is desirable for the headpiece to be chemically stable. In one embodiment, the headpiece preferably has a structure that allows the DNA tag and building blocks to be arranged in an appropriate space. In one embodiment, it is preferable that the headpiece has a suitable degree of flexibility. Here, we will further explain the appropriate spatial arrangement and flexibility (structural characteristics of the headpiece). The structural characteristics of the headpiece described here may be achieved by the headpiece alone, or may be achieved by binding the headpiece to a bifunctional spacer. In one embodiment, preferred structural characteristics of the headpiece are such that the headpiece or DNA tag does not inhibit the reaction of forming the building block, and conversely, the headpiece or building block does not inhibit the extension reaction of the DNA tag. In one embodiment, the preferred structural characteristics of the headpiece are such that the headpiece or DNA tag portion does not affect the interaction between the building block compound (library compound) and the target (such as a target protein). In one embodiment, the preferred structural feature of the headpiece is one that orients the DNA tag and building block sites on opposite sides (e.g., at 90 degrees or more opposite sides). In one embodiment, the preferred structural characteristic of the headpiece is one that separates the loop portion of the headpiece from the building block by several atoms to a dozen or so atoms in terms of the organic compound skeleton. In one embodiment, the headpiece preferably has a suitable affinity with the DNA tag portion and the building block portion. Suitable affinity means, for example, chemical reactivity and stability that allow the bond to be formed, maintained, and broken under desired conditions to carry out the present invention. In the present invention, a bifunctional spacer means a spacer moiety having at least two reactive groups that enable binding between a building block moiety and a headpiece.

[0068] In the description of the present invention, the terms "headpiece", "headpiece compound" and "compound for a headpiece" are terms indicating compounds of the same concept. In the description of the present invention, "a compound used as a headpiece" can be essentially understood as "use of a compound as a headpiece" from the viewpoint of use, and can be essentially understood as "a method of using a compound as a headpiece" from the viewpoint of method. The same applies to the compound library.

[0069] A preferred headpiece structure will be described below, but the headpiece structure is not limited as long as it achieves the effects of the present invention.

[0070] In one embodiment, the headpiece comprises: (D) a reactive functional group having at least one site capable of linking directly to a building block or indirectly via a bifunctional spacer; (L) a linker extending from the reactive functional group; (E) a first oligonucleotide strand having one binding site capable of being linked to one strand of an oligonucleotide tag; (F) a second oligonucleotide strand having one binding site that can be linked to the other strand of the oligonucleotide tag; and (LP) a loop region that binds the linker and the two oligonucleotide strands; It is composed of At least one of the E, F, and LP sites has at least one selectively cleavable site.

[0071] In one embodiment, the headpiece is a compound represented by formula (I): [ka] (In the formula, E and F each independently represent An oligomer composed of nucleotides or nucleic acid analogs, provided that E and F contain complementary base sequences to each other and form a double-stranded oligonucleotide; LP is the loop region; L is a linker, D is a reactive functional group. A compound represented by the formula: A compound having at least one selectively cleavable site in at least one of sites E, F, or LP.

[0072] In the present invention, the partial structure of the loop site that binds to the linker may be referred to as the linking site (LS). In the present invention, E-LP-F may be collectively referred to as a hairpin site.

[0073] (First and Second Oligonucleotide Strands) Preferred embodiments of the first oligonucleotide strand (E) and the second oligonucleotide strand (F) are described below.

[0074] The first oligonucleotide strand (E) and the second oligonucleotide strand (F) preferably form a duplex in an intramolecular manner via the loop portion (LP), and the headpiece preferably forms a hairpin structure. The chain length for intramolecular duplex formation is preferably 3 bases or more, more preferably 4 bases or more, and even more preferably 6 bases or more. In one embodiment, the chain length of E and F is 3 to 40. In one embodiment, the chain length of E and F is 4 to 40, respectively. In one embodiment, the chain length of E and F is 6 to 25.

[0075] The site to which the oligonucleotide tag is linked is preferably a structure suitable for enzymatic ligation or chemical ligation. In one embodiment, the linkage between the headpiece and the oligonucleotide tag is carried out by double-stranded ligation using an enzyme. In that case, the first and second oligonucleotide strands preferably form a protruding end for ligation. The chain length of the protruding end is preferably 2 bases or more, more preferably 2 to 10 bases, and even more preferably 2 to 5 bases. Therefore, it is preferable that one of the first and second oligonucleotide strands is longer than the other strand by the chain length of the protruding end. In addition, for ligation by DNA ligase, the 5' end of the strand having the 5' end of the headpiece among the first and second oligonucleotide strands is preferably phosphorylated.

[0076] The first and second oligonucleotide strands may contain a part or the whole of a primer binding sequence for PCR. The appropriate length of the primer binding sequence is 17 to 25 bases.

[0077] (Linker) Preferred embodiments of the linker (L) are described below. As described above, the linker is a moiety that extends from the reactive functional group and bonds to the linking moiety. Typically, the linker is a divalent group (-L-) derived from the following embodiments.

[0078] In one embodiment, the linker has the following embodiment (L1): (L1) a C1-20 aliphatic hydrocarbon which may have a substituent and which may be substituted with 1 to 3 heteroatoms, or (2) a C6-14 aromatic hydrocarbon which may have a substituent.

[0079] In other embodiments, L is the following embodiment (L2), (L3), (L4) or (L5). (L2) C1-6 aliphatic hydrocarbons which may have a substituent, C1-6 aliphatic hydrocarbons which may be substituted with 1 or 2 oxygen atoms, or C6-10 aromatic hydrocarbons which may have a substituent. (L3) A C1-6 aliphatic hydrocarbon which can be substituted with the substituent group ST1, or a benzene which can be substituted with the substituent group ST1. Here, the substituent group ST1 is a group consisting of a C1-6 alkyl group, a C1-6 alkoxy group, a fluorine atom, and a chlorine atom. However, when the substituent group ST1 substitutes an aliphatic hydrocarbon, an alkyl group is not selected from the substituent group ST1. (L4) C1-6 alkyl, or benzene which is unsubstituted or substituted with one or two C1-3 alkyl or C1-3 alkoxy groups. (L5) C1~6 alkyl.

[0080] (Reactive Functional Group) Preferred embodiments of the reactive functional group (D) will be described below. As described above, the reactive functional group has at least one site that can be directly linked to a building block or indirectly linked via a bifunctional spacer, and is a site that binds to a linker group. Typically, the reactive functional group is a monovalent group (D-) in the headpiece, and is a "divalent group derived from a reactive functional group" (-D-) based on the (D-) in the DEL. For example, when D is an amino group, the specific structure of (D-) is (R-HN-) (R is a substituent described below). For example, it reacts with an activated carboxy group, a reactive sulfonyl group, or an isocyanate group to form an amide bond, a sulfonamide bond, or a urea bond, respectively. In that case, the specific structure of (-D-) is (-NR-). R is not limited as long as the effects of the present invention are achieved, but in the following embodiments (D1) to (D5), R is preferably (1) a hydrogen atom, or (2) a C1-6 alkyl group which is unsubstituted or substituted with 1 to 3 substituents selected, either alone or differently, from the group of substituents consisting of a C1-6 alkoxy group, a fluorine atom, and a chlorine atom. R is more preferably a hydrogen atom or a C1-3 alkyl group, and even more preferably a hydrogen atom. Also, for example, when (D-) is a methylene group having a leaving group (X-), the specific structure of (D-) is (X-CH2-), and it reacts with a nucleophilic reagent such as an amino group, a hydroxyl group, or a thiol group to form a carbon-nitrogen bond, a carbon-oxygen bond, or a carbon-sulfur bond. In that case, the specific structure of (-D-) is (-CH2-). Also, for example, when (D-) is an aldehyde group, the specific structure of (D-) is (HOC-). The aldehyde group forms a carbon-nitrogen bond, for example, by a reductive amination reaction with an amino group, and in that case (-D-) becomes -CH2-, and forms a carbon-carbon double bond, for example, by a reaction with a phosphorus ylide group, and in that case (-D-) becomes -CH=, and forms a carbon-carbon triple bond, for example, by a reaction with an α-diazophosphonate group, and in that case (-D-) becomes -C≡.

[0081] In one embodiment, the site (D-) is the following embodiment (D1). (D1) Functional groups that can form C-C, amino, ether, carbonyl, amide, ester, urea, sulfide, disulfide, sulfoxide, sulfonamide, or sulfonyl bonds. (Literally, in this case, (-D-) is a C-C, amino, ether, carbonyl, amide, ester, urea, sulfide, disulfide, sulfoxide, sulfonamide, or sulfonyl bond.)

[0082] In other embodiments, (D-) is the following embodiment (D2), (D3), (D4) or (D5). (D2) A C1 hydrocarbon with a leaving group, an amino group, a hydroxyl group, a precursor of a carbonyl group, a thiol group, or an aldehyde group. In this case, (-D-) can be -(C1 hydrocarbon)-, -NR-, -O-, -(C=O)-, -S-, -CH2-, -CH=, or -C≡, etc. (D3) C1 hydrocarbons having halogen atoms, C1 hydrocarbons having sulfonic acid leaving groups, amino groups, hydroxyl groups, carboxy groups, halogenated carboxy groups, thiol groups, or aldehyde groups. In this case, (-D-) can be -(C1 hydrocarbon)-, -NR-, -O-, -(C=O)-, -S-, -CH2-, -CH=, or -C≡, etc. (D4) -CH2Cl, -CH2Br, -CH2OSO2CH3, -CH2OSO2CF3, an amino group, a hydroxyl group, or a carboxy group. In this case, (-D-) is -CH2-, -NR-, -O-, or -(C=O)-, respectively. (D5) Primary amino group. In this case, (-D-) becomes -NH-.

[0083] Preferred embodiments of the loop region (LP) are described below. The loop portion (LP) is preferably designed so that the first oligonucleotide strand (E) and the second oligonucleotide strand (F) form a double strand in the molecule and the headpiece can form a hairpin structure. That is, the loop portion (LP) preferably has a chain length and bond flexibility that make the loop structure thermodynamically stable. Thus, in one embodiment, the loop portion (LP) is: LP, A loop region represented by (LP1)p-LS-(LP2)q, LS is a partial structure selected from the group of compounds described in (A) to (C) below, (A) Nucleotides (B) Nucleic acid analog (C) a C1-14 trivalent group which may have a substituent LP1 is a partial structure selected from the group of compounds described in the following (1) and (2) in a quantity of p alone or differently, (1) Nucleotides (2) Nucleic acid analogs LP2 is each partial structure selected from the group of compounds described in the following (1) and (2) in a quantity of q, either singly or differently, (1) Nucleotides (2) Nucleic acid analogs The total number of p and q is 0 to 40.

[0084] More preferred embodiments of the loop site are as explained above. The structure of the loop region will be further explained below.

[0085] Here, the nucleotides are natural nucleotides as explained above, and the nucleic acid analogs are as explained above.

[0086] Here, LP1 is each of p partial structures selected from the group of compounds described in (1) and (2) below, either singly or differently, and LP2 is each of q partial structures selected from the group of compounds described in (1) and (2) below, either singly or differently. (1) Nucleotides (2) Nucleic acid analogs The term "p selected independently or differently" means that, for example, when p is 4, LP1 can be selected independently or differently from the group of compounds described in (1) and (2), such as AATG, ATCG, TC(d-Spacer)G, or A(d-Spacer)(d-Spacer)C. The same applies to LP2.

[0087] The loop site may also contain a part or the whole of a primer binding sequence for PCR.

[0088] (About LS) In one embodiment, the LS is (A) a nucleotide or (B) a nucleic acid analog. When LS is (A) a nucleotide or (B) a nucleic acid analog, the loop portion is a nucleic acid oligomer. The nucleic acid oligomer of the present invention is an oligomer in which nucleotides or nucleic acid analogs are linked as monomers. The oligomer can also be called a chain compound. Thus, a nucleic acid oligomer of the present invention may be either an oligonucleotide strand, a nucleic acid analog strand, or a mixed strand of nucleotides and nucleic acid analogs.

[0089] When LS is (A) a nucleotide or (B) a nucleic acid analog, the loop portion is a nucleic acid oligomer. In this case, the headpiece can be produced by a nucleic acid synthesizer, which is significantly preferable in practice.

[0090] When the LS is (A) a nucleotide or (B) a nucleic acid analog, in one embodiment of the production of a headpiece, a monomer for nucleic acid synthesis in which a linker portion (L) and a reactive functional group portion (D) are bound to the LS can be prepared, and then a nucleic acid oligomer can be synthesized. Examples of such nucleic acid synthesis monomers include the above-mentioned Amino C6 dT, mdC (TEG-Amino), and Uni-Link (registered trademark) Amino Modifier. In this embodiment, for example, in the structure of the monomer mdC(TEG-Amino), the nucleotide portion corresponds to the linking site (LS), and the side chain portion extending from the base corresponds to the linker site (L) and the reactive functional group site (D). In the preparation, the reactive functional group (D) may be protected with a protecting group.

[0091] In that case, in one embodiment, the nucleic acid analog is the following compound (B6). (B6) A compound in which the (-LD) is bound to the base portion of a nucleotide.

[0092] In one embodiment, the nucleic acid analog is the following compound (B61), (B62), (B63), (B64) or (B65). (B61) (-LD) is (-L1-D1) (B6) (B62)(-LD) is (-L2-D2) (B6). (B63)(-LD) is (-L3-D3) (B6). (B64)(-LD) is (-L4-D4) (B6). (B65) The compound according to any one of (B61) to (B64), wherein (-D) is (-D5).

[0093] When LS is (A) a nucleotide or (B) a nucleic acid analog, in one embodiment of the production of the headpiece, a nucleic acid oligomer can be synthesized first, and then the linker portion (L) and the reactive functional group portion (D) can be bound. In this case, it is preferable to incorporate the "specific nucleic acid analog" to which the linker site binds into the hairpin site (nucleic acid analog oligomer) as a linking site (LS). Examples of the "specific nucleic acid analog" include the above-mentioned Amino C6 dT, mdC (TEG-Amino), and Uni-Link (registered trademark) Amino Modifier. In this embodiment, for example, mdC(TEG-Amino) itself corresponds to the linking site (LS), and the additional site from the base side chain to which it is further bound corresponds to the linker site (L) and the reactive functional group site (D).

[0094] (About p and q) As described above, the length of the loop portion is preferably such that the first oligonucleotide strand (E) and the second oligonucleotide strand (F) form a double strand within the molecule and the headpiece forms a hairpin structure. In one embodiment, the sum of p and q is 1-40. In one embodiment, the total number of p and q is 2-20. In one embodiment, the total number of p and q is 2-10. In one embodiment, the total number of p and q is 2-7.

[0095] In one embodiment, the loop site of the present invention is (A) Nucleotides and the following nucleic acid analogues (B41), (B42), (B43), (B44) or (B52). (B41)d-Spacer (B42) Amino C6 dT (B43)mdC(TEG-Amino) (B44) Uni-Link (trademark) Amino Modifier (B52) Triethylene glycol phosphate ester

[0096] In one embodiment, LS is preferably B42, B43 or B44. In another embodiment, LP1 and LP2 are preferably A, B41 or B52.

[0097] In one embodiment, the loop site is a nucleic acid oligomer having a sequence shown in (X1) to (X9) below. (X1)A-B41-B42-B41-A (X2)A-B41-B43-B41-A (X3)A-B41-B44-B41-A (X4) B41-B41-B42-B41-B41 (X5) B41-B41-B43-B41-B41 (X6) B41-B41-B44-B41-B41 (X7) B52-B42-B52 (X8)B52-B43-B52 (X9)A52-A44-A52

[0098] In the headpiece, the number of cleavable sites is preferably 5 or less, and more preferably 1 or 2.

[0099] In the headpiece, when there are two or more cleavable sites, it is preferred that at least one cleavable site is in the first oligonucleotide strand or between the first oligonucleotide strand and the linker binding site, and at least one cleavable site is in the second oligonucleotide strand or between the second oligonucleotide strand and the linker binding site.

[0100] In one embodiment, in the headpiece, the position of the cleavable site is preferably within 20 bases, more preferably within 10 bases, and even more preferably within 3 bases from the binding site between the loop site and the first oligonucleotide strand or the second oligonucleotide strand.

[0101] Just to be clear, the preferred embodiment of the "selectively cleavable site" and the preferred embodiments of, for example, E, F, or LP are different concepts. In other words, even if the position of the "selectively cleavable site" is included in E, the preferred embodiment of E does not apply to the "selectively cleavable site".

[0102] In one embodiment, the compound constituting the DEL of the present invention is a compound represented by the following formula (II): [ka] (In the formula, X and Y are oligonucleotide strands; E and F are each independently An oligomer composed of nucleotides or nucleic acid analogs, provided that E and F contain complementary base sequences to each other and form a double-stranded oligonucleotide; LP is the loop region; L is a linker, D is a divalent group derived from a reactive functional group, Sp is a bond or a bifunctional spacer; An is a partial structure composed of at least one building block. A compound represented by the formula: X and Y have a sequence capable of forming a double strand at least in part, X is bound to E at the 5' end, Y binds to F at the 3' end, A compound having at least one selectively cleavable site in at least one of sites E, F, or LP.

[0103] In one embodiment, preferred embodiments of E, F, LP, L, and D in the compound represented by the above formula (II) are the same as the preferred embodiments of E, F, LP, L, and D described in relation to the above formula (I). Preferred embodiments of X, Y, Sp, and An are described below.

[0104] (Bifunctional spacer) As described above, the bifunctional spacer is a spacer moiety having at least two reactive groups that allow the binding of the compound library substructure An to the headpiece. In one embodiment, the bifunctional spacer is SpD-SpL-SpX. SpX is a reactive group that forms a covalent bond with a reactive functional group on the headpiece. SpD is a reactive group that forms a covalent bond with a partial structure An of the compound library. SpL is a chemically inert spacing moiety. In addition, like the reactive functional group (D), the reactive group (SpX) becomes a monovalent group (-SpX) in the bifunctional spacer alone (the state of the reagent before binding to the headpiece), and becomes a "divalent group derived from the reactive group" (-SpX-) based on the (-SpX) in DEL (the state bound to the headpiece). Similarly, the reactive group (SpD) becomes a monovalent group (SpD-) before bonding with An, and in DEL (bonded with An) becomes a “divalent group derived from the reactive group” (-SpD-) based on the (SpD-).

[0105] A preferred embodiment of SpX is a reactive group that forms an amino, carbonyl, amide, ester, urea, or sulfonamide bond. In one embodiment, SpX is a reactive group suitable for when the reactive functional group of the headpiece is an amino group, and has the following structure (SpX1), (SpX2), or (SpX3). (SpX1): Carboxy group, halogenated carboxy group, aldehyde group, or halogenated sulfonyl group (SpX2): Carboxy group or halogenated sulfonyl group (SpX3): Carboxy group

[0106] The preferred embodiment of SpD is the same as D described above. In one embodiment, SpD is (D1), (D2), (D3), (D4) or (D5) above.

[0107] Preferred embodiments of the SpL are as follows: In one embodiment, SpL is (L1), (L2), (L3), (L4) or (L5) as described above. In one embodiment, SpL is (SpL1), (SpL2) or (SpL3). (SpL1) polyalkylene glycol, polyethylene, C1-20 aliphatic hydrocarbon optionally substituted with a heteroatom, peptide, oligonucleotide, or a combination thereof. (SpL2) Polyalkylene glycol, polyethylene, C1-10 aliphatic hydrocarbon, or peptide (SpL3) Polyethylene glycol, or polyethylene

[0108] In one embodiment, the bifunctional spacer is: (Sp1):(D4)-(SpL1)-(SpX1) (Sp2):(D4)-(SpL2)-(SpX2) (Sp3):(D4)-(SpL3)-(SpX3) (Sp4):(D5)-(SpL1)-(SpX1) (Sp5):(D5)-(SpL2)-(SpX2) (Sp6):(D5)-(SpL3)-(SpX3)

[0109] In one embodiment, the (Sp-DL) portion of the compound constituting the DEL is configured as follows: (SpDL1), (SpDL2), (SpDL3) (SpDL4), (SpDL5), (SpDL6), (SpDL7), (SpDL8), (SpDL9), or (SpDL10). (SpDL1):(D4)-(L1) (SpDL2):(D5)-(L1) (SpDL3):(D4)-(L2) (SpDL4):(D5)-(L2) (SpDL5):(Sp1)-(D5)-(L5) (SpDL6):(Sp2)-(D5)-(L5) (SpDL7):(Sp3)-(D5)-(L5) (SpDL8):(Sp4)-(D5)-(L5) (SpDL9):(Sp5)-(D5)-(L5) (SpDL10):(Sp6)-(D5)-(L5) In (SpDL1), (SpDL2), (SpDL3), and (SpDL4), Sp means a bond.

[0110] In the practice of the present invention, it is advantageous if the headpiece can be synthesized by a nucleic acid synthesizer. In the practice, as described above, in one embodiment, a nucleic acid synthesis monomer in which the linker portion (L) and the reactive functional group portion (D) are bound to LS can be prepared, and then a nucleic acid oligomer can be synthesized. Examples of such a nucleic acid synthesis monomer include the above-mentioned Amino C6 dT, mdC (TEG-Amino), Uni-Link (registered trademark) Amino Modifier, etc. On the other hand, when using commercially available nucleic acid synthesis monomers or nucleic acid analogs that can be used in a nucleic acid synthesizer, the length of the linker site may be limited. In such a case, as one embodiment, by introducing an appropriate bifunctional spacer, it becomes possible to adjust the distance between the headpiece and An, which is advantageous in carrying out the invention.

[0111] In the description of the present invention, "C1-C6" and "C1-6" in terms such as "C1-C6 alkyl group" and "C1-6 alkyl group" mean that the number of carbon atoms is 1 to 6. Similarly, when m and n are integers and there is a description such as "Cm-Cn" or "Cm-n", the description means that the number of carbon atoms is m to n. Therefore, "C1-C6 alkyl group" and "C1-6 alkyl group" mean an alkyl group having 1 to 6 carbon atoms, and "C1-C6 alkylene" and "C1-6 alkylene" mean an alkylene having 1 to 6 carbon atoms.

[0112] In the present invention, "C1-6 alkyl" refers to a straight or branched alkyl group having 1 to 6 carbon atoms. Specific examples include methyl, ethyl, propyl, isopropyl, butyl, isobutyl, sec-butyl, tert-butyl, pentyl, hexyl, and the like.

[0113] In the present invention, "C1-3 alkyl" refers to a straight or branched alkyl group having 1 to 3 carbon atoms. Specific examples are methyl, ethyl, propyl, and isopropyl.

[0114] In the present invention, "C1-6 alkoxy" means a straight or branched alkoxy having 1 to 6 carbon atoms. Specific examples include methoxy, ethoxy, propoxy, isopropoxy, butoxy, isobutoxy, sec-butoxy, tert-butoxy, pentyloxy, hexyloxy, etc.

[0115] In the present invention, "C1-3 alkoxy" means a straight or branched alkoxy having 1 to 3 carbon atoms. Specific examples include methoxy, ethoxy, propoxy, and isopropoxy.

[0116] In the present invention, the term "hydrocarbon" refers to a linear, branched or cyclic, saturated or unsaturated compound composed solely of carbon and hydrogen atoms.

[0117] In the present invention, the term "aliphatic hydrocarbon" refers to a non-aromatic hydrocarbon. The "aliphatic hydrocarbon" may be linear, branched or cyclic, and may be saturated or unsaturated. Specific examples of the structure include alkyl, alkenyl, alkynyl, cycloalkyl or cycloalkenyl, or a combination thereof. In the present invention, the term "C1-20 aliphatic hydrocarbon" refers to an aliphatic hydrocarbon having 1 to 20 carbon atoms. In the present invention, the term "C1-10 aliphatic hydrocarbon" refers to an aliphatic hydrocarbon having 1 to 10 carbon atoms. In the present invention, "C1-6 aliphatic hydrocarbon" means an aliphatic hydrocarbon having 1 to 6 carbon atoms.

[0118] In the present invention, the term "aromatic hydrocarbon" refers to aromatic hydrocarbons. In the present invention, "C6-14 aromatic hydrocarbon" means an aromatic hydrocarbon having a carbon atom number of 6 to 14. Specific examples include benzene, naphthalene, and anthracene. In the present invention, "C6-10 aromatic hydrocarbon" means an aromatic hydrocarbon having 6 to 10 carbon atoms. Specific examples are benzene or naphthalene.

[0119] The aromatic heterocycle of the present invention is an aromatic heterocycle having, as a heteroatom in the ring structure, an element selected, either alone or differently, from the group consisting of nitrogen, oxygen and sulfur. In one embodiment, the aromatic heterocycle is a "C1-9 aromatic heterocycle" having 1 to 9 carbon atoms, and in one embodiment, the "C1-9 aromatic heterocycle" is a 5- to 10-membered aromatic heterocycle. In one embodiment, the aromatic heterocycle is a "C1-5 aromatic heterocycle" having 1 to 5 carbon atoms, and in one embodiment, the "C1-5 aromatic heterocycle" is a 5- or 6-membered aromatic heterocycle. In one embodiment, the aromatic heterocycle is a "C2-9 aromatic heterocycle" having 2 to 9 carbon atoms, and in one embodiment, the "C2-9 aromatic heterocycle" is a 5- to 10-membered aromatic heterocycle. In one embodiment, the aromatic heterocycle is a "C2-5 aromatic heterocycle" having 2 to 5 carbon atoms, and in one embodiment, the "C2-5 aromatic heterocycle" is a 5- or 6-membered aromatic heterocycle.

[0120] The nitrogen-containing aromatic heterocycle of the present invention is an aromatic heterocycle having nitrogen as a heteroatom in the ring structure. In one embodiment, the nitrogen-containing aromatic heterocycle is a "C1-5 nitrogen-containing aromatic heterocycle" having 1 to 5 carbon atoms, and in one embodiment, the "C1-5 nitrogen-containing aromatic heterocycle" is a 5- or 6-membered aromatic heterocycle. In one embodiment, the nitrogen-containing aromatic heterocycle is a "C2-5 nitrogen-containing aromatic heterocycle" having 2 to 5 carbon atoms, and in one embodiment, the "C2-5 nitrogen-containing aromatic heterocycle" is a 5- or 6-membered aromatic heterocycle.

[0121] The non-aromatic heterocycle of the present invention is a non-aromatic heterocycle having, as a heteroatom in the ring structure, an element selected, either alone or differently, from the group consisting of nitrogen, oxygen and sulfur. The non-aromatic heterocycle may contain partially unsaturated bonds. In one embodiment, the non-aromatic heterocycle is a "C2-9 non-aromatic heterocycle" having 2 to 9 carbon atoms, and in one embodiment, the "C2-9 non-aromatic heterocycle" is a 5- to 10-membered non-aromatic heterocycle.

[0122] In the present invention, a "C1-14 trivalent group" refers to a trivalent group derived from a compound having 1 to 14 carbon atoms. The structure is not limited as long as the effects of the present invention are achieved.

[0123] In the present invention, when it is stated that "may be replaced by a heteroatom", the heteroatom means an atom other than carbon and hydrogen. The heteroatom is preferably an oxygen atom, a nitrogen atom, a silicon atom, a phosphorus atom, or a sulfur atom, and more preferably an oxygen atom, a nitrogen atom, or a sulfur atom. Therefore, for example, taking propyl (-CH2-CH2-CH3) as an example of a hydrocarbon, "propyl which may be replaced by a heteroatom" is a concept that includes structures such as ether ((-CH2-O-CH3) or (-O-CH2-CH3)) in which the methylene (-CH2-) in the alkyl is replaced by oxygen, or amine ((-CH2-NH-CH3) or (-NH-CH2-CH3)) in which the methylene (-CH2-) is replaced by nitrogen.

[0124] In the present invention, when it is stated that "a substituent may be present", the substituent is not limited as long as the object of the present invention is achieved. The substituent is preferably a C1 to 6 alkyl group, a C1 to 6 alkoxy group, an amino group, a hydroxy group, a nitro group, a cyano group, an oxo group or a halogen atom. The substituent is more preferably a C1 to 6 alkyl group, a C1 to 6 alkoxy group, a fluorine atom or a chlorine atom.

[0125] In the present invention, the term "polypeptide" and "peptide" refers to a compound or partial structure formed by linking amino acids. "Amino acid" is a general term for organic compounds having both amino and carboxyl functional groups. The amino acids constituting the polypeptides and peptides of the present invention are not particularly limited and include modified amino acids. In accordance with common usage in the field of life science, proline (classified as an imino acid) is also included in the amino acids in the present invention. The amino acids constituting the polypeptides and peptides of the present invention are preferably alpha amino acids, and more preferably "proteinogenic amino acids".

[0126] In the present invention, the halogen atom includes a fluorine atom, a chlorine atom, a bromine atom and an iodine atom.

[0127] C-C, amino, ether, carbonyl, amide, ester, urea, sulfide, disulfide, sulfoxide, sulfonamide, and sulfonyl bonds are chemical bonds having a chemical structure that can be understood by each name. Those skilled in the art will understand that, for example, an ether bond is a bond that can be generally expressed as "-O-", and a carbonyl bond is a bond that can be generally expressed as "-C(=O)-". Although amino, amide, and urea bonds have a hydrogen atom or other substituent on the nitrogen atom, the structure on the nitrogen atom is not limited as long as the effect of the present invention is obtained. The substituent on the nitrogen atom is preferably a C1-6 alkyl group or a hydrogen atom, and more preferably a hydrogen atom. It is needless to say that a CC bond means a carbon-carbon bond. The CC bond includes a single bond, a double bond, and a triple bond. In one embodiment, in the step a and / or c of the production method of the present invention, a bond appropriately selected from the above 11 types is constructed. These 11 types of bonds are particularly basic bond patterns in organic chemistry, and the reactions for constructing them are also well known to those skilled in the art. Therefore, when designing and constructing a partial structure An of the compound library of the present invention, a person skilled in the art can use these 11 types of bonds in appropriate combination.

[0128] An organic compound composed of elements selected, either singly or differently, from the group consisting of H, B, C, N, O, Si, P, S, F, Cl, Br, and I is an organic compound constructed by bonding the above 12 elements.

[0129] In one embodiment, the partial structure An of the compound library of the present invention is constructed from the above 12 elements. These 12 elements are particularly basic elements in organic compounds, and the reactions for constructing them are also well known to those skilled in the art. Therefore, when designing and constructing the partial structure An of the compound library of the present invention, those skilled in the art can use these 12 elements in appropriate combination.

[0130] The low molecular weight organic compound having a substituent selected from the group consisting of an aryl group, a non-aromatic cyclyl group, a heteroaryl group, and a non-aromatic heterocyclyl group, is a low molecular weight organic compound having a chemical structure that can be understood from each name. The low molecular weight compound is a concept well known to those skilled in the art, and examples of preferred molecular weights of the low molecular weight compound in the present invention will be mentioned separately.

[0131] The aryl group of the present invention is preferably a C6-10 aryl group, and more preferably a phenyl group.

[0132] The non-aromatic cyclyl groups of the invention are preferably 5- to 8-membered, more preferably 5- or 6-membered non-aromatic cyclyl groups, which may contain partially unsaturated bonds.

[0133] The heteroaryl and non-aromatic heterocyclyl groups of the present invention are groups having, as heteroatoms in the ring structure, elements selected, either singly or differently, from the group consisting of nitrogen, oxygen and sulfur. The heteroaryl and non-aromatic heterocyclyl groups of the present invention are preferably 5- to 8-membered groups, more preferably 5- or 6-membered groups, and the non-aromatic heterocyclyl groups may contain partially unsaturated bonds.

[0134] In one embodiment, the partial structure An of the compound library of the present invention has the above four types of groups. These four types of groups are particularly basic partial structures in organic compounds, and the reactions for constructing them in compounds are also well known to those skilled in the art. Therefore, when designing and constructing the partial structure An of the compound library of the present invention, those skilled in the art can use these four types of groups in appropriate combination.

[0135] The compound libraries constructed in the above-mentioned preferred embodiments, i.e., 11 types of bonds, 12 types of elements, and / or 4 types of groups, have particularly core value. Therefore, those skilled in the art will understand that compound libraries constructed outside of these preferred embodiments will generally have limited uses and in many cases limited commercial value.

[0136] The synthetic history of An means a record of all the operations performed until An is synthesized, and in particular means the structure and order of the building blocks used until An is synthesized. For example, when reactions are carried out in two or more separate reaction vessels using different building blocks and / or under different reaction conditions, the synthetic history is imparted as sequence information of the oligonucleotide by linking an oligonucleotide chain of a predetermined sequence to the product in each reaction vessel before or after the reaction. By repeating such operations until An is constructed, an oligonucleotide of Bn having the synthetic history of An is constructed.

[0137] Split-and-pool synthesis is a synthesis method developed by Geysen et al. in the early days of combinatorial chemistry as a method for constructing combinatorial peptide libraries using solid-phase synthesis. Split-and-pool synthesis is also called the split-mix method.

[0138] Following the above process, let us take the synthesis of a peptide library using solid-phase synthesis as an example. In split-and-pool synthesis, at each step of peptide amplification, the N types of supports are first mixed and homogenized, and then equally divided to amplify the next N types of amino acids, without cutting out the sample from the solid-phase support on which the amino acids are peptide-bonded.

[0139] In other words, one type of peptide chain is generated for each carrier, and by applying all 20 natural amino acids at each stage, a peptide library of all possible combinations for a peptide of a particular length can be constructed.

[0140] If this peptide library is to be screened for antigen presentation or receptor binding, an assay can be performed using peptides on a solid-phase carrier by using an ELISA method or the like. In other words, there is no need to cut out the sample peptides from the carrier, and carrier particles that react to the assay can be picked up (for example, fluorescently labeled carrier particles of about 0.1 mm in size can be picked up with an optical microscope). The peptides on the particles can then be analyzed with an instrument (such as a peptide analyzer) to determine the target peptide sequence, or the peptide sequences that are candidates for screening can be indirectly determined by other combinatorial chemical identification methods (such as tagging methods).

[0141] Furthermore, in the production method of the present invention, an example will be described in which v types of structures are synthesized by split-and-pool synthesis when m is 2, and w types of structures are synthesized when m is 3. In this explanation, the steps are repeated in the order of (c) and (d). (m=2) In the step where m=2, α2 is added to A1-Sp-C-B1 in step (c) and β2 is added in step (d) to produce A2-Sp-C-B2. Here, v types of α2 (α2(av)) and v types of corresponding β2 (β2(av)) are prepared, and steps (c) and (d) are performed for each structure, respectively, to obtain v types of A2-Sp-C-B2 (A2(a)-Sp-C-B2(a), A2(b)-Sp-C-B2(b)...A2(v)-Sp-C-B2(v): i.e., A2(av)-Sp-C-B2(av)). In split-and-pool synthesis, v types of A2-Sp-C-B2 are mixed and then divided into w pieces. Division, most specifically, means dividing into w reaction vessels. (m=3) In the step m=3, α3 is added to A2-Sp-C-B2 in step (c) and β3 is added in step (d) to produce A3-Sp-C-B3. Here, w types of α3 (α3(aw)) and w types of β3 (β2(aw)) are prepared, and steps (c) and (d) are carried out for w (A2(av)-Sp-C-B2(av) mixtures). Then, through steps n=2 and 3, (v×w) types of A3-Sp-C-B3 can be efficiently synthesized in (v+w) syntheses.

[0142] (Biological evaluation) Mixing the w products thus obtained results in a mixture of (v×w) types of A3-Sp-C-B3 compound libraries. For example, if a drug receptor binding test is performed on this mixture, screening of (v×w) types of compounds can be performed in one go. Compounds that do not bind to the drug receptor can be washed away, and only the bound compounds can be isolated. In the DEL of the present invention, the DNA of the isolated A3-Sp-C-B3 compound is amplified to an amount that can be sequenced, and the structure of A3 can be understood from the sequence information.

[0143] The terms "compound library," "building block," "split-and-pool," and the like are well known to those skilled in the art in the field of combinatorial chemistry, and these can be carried out with reference to the following literature, etc., as appropriate. (1) Takashi Takahashi and Takayuki Doi, "Combinatorial Chemistry," Journal of the Society of Organic Synthesis, 2002, Vol. 60, pp. 426-433 (2) Combinatorial Chemistry Study Group (ed.), "Combinatorial Chemistry," Kagaku Dojin

[0144] A DNA-encoded library (or DEL) is a compound library consisting of a group of compounds (DNA-encoded compounds) labeled with DNA or oligonucleotides having substantially the same functions as DNA. By split-and-pool synthesis as described above, the labeled DNA is given the structure or synthetic history of each compound as sequence information. Due to these characteristics, DNA-encoded libraries can be used to generate a wide variety of compounds. 2 ~10 20The structure of the compound can be identified by screening in the form of a mixture of various compounds and identifying the DNA sequence contained in the obtained compound by a method known in the art (e.g., using a next-generation sequencer and / or using a microarray). As one embodiment of the screening method, a method can be selected in which a target such as a protein is contacted with a DNA-encoded library and a compound that binds to the target is selected.

[0145] The term "biological target" is well known to those skilled in the art. In one embodiment, in the present invention, a "biological target" refers to a group of biological substances that can be targets in the development of drugs, such as medical and agricultural chemicals, and includes, for example, enzymes (e.g., kinases, phosphatases, methylases, demethylases, proteases, and DNA repair enzymes), proteins involved in protein:protein interactions (e.g., receptor ligands), receptor targets (e.g., GPCRs), ion channels, cells, bacteria, viruses, parasites, DNA, RNA, prions, or carbohydrates. "Biological activity evaluation" is a term well known to those skilled in the art, and in one embodiment, in the present invention, "biological activity evaluation" refers to evaluating the presence or absence or strength of biological activity of a compound (e.g., ability to bind to a biological target, function to inhibit enzyme activity, function to promote enzyme activity, etc.). As specific examples of biological activity evaluation, reference can be made to the aforementioned Patent Documents 2 and 3, and Non-Patent Documents 1 to 6. "Functionality evaluation" is a term well known to those skilled in the art. In one embodiment, in the present invention, "functionality evaluation" refers to evaluating the presence or absence, or the strength or weakness, of a specific function (e.g., binding ability, biological activity, luminescence properties, etc.) of a compound.

[0146] By using a DNA strand with a cleavable site, the present invention provides several approaches for DELs and methods for producing DELs, each of which has several advantages. Modes 1 to 7 are described in detail below.

[0147] Form 1 The present invention provides a DEL that uses the above-mentioned "hairpin-shaped headpiece having a cleavable site."

[0148] As illustrated in Figure 1, in mode 1, starting from a first oligonucleotide strand containing a cleavable site in the DNA strand, and a headpiece containing a loop site and a second oligonucleotide strand, the production of DEL is achieved by repeating (three times in Figure 1) the binding of building blocks and double-stranded ligation of oligonucleotide tags corresponding to the building blocks, and further performing double-stranded ligation of an oligonucleotide tag containing a primer region, if desired.

[0149] As illustrated in Figure 2, in mode 1, for a DEL containing a cleavable site in the first oligonucleotide strand of the headpiece, a cleavable site is cleaved using a cleavage means such as an enzyme to induce a double-stranded oligonucleotide that is not linked at the loop site, thereby enabling PCR to be performed with high efficiency.

[0150] (Regarding Form 2) In DEL using a "hairpin headpiece with a cleavable site," the cleavable site may be present in the second oligonucleotide strand, as illustrated in Figure 3. The characteristics of Mode 2 are similar to Mode 1, except for the cleavable site.

[0151] (Regarding Form 3) As illustrated in Figure 4, in DEL using a "hairpin-shaped headpiece having a cleavable site," the cleavable site may be present in both the first and second oligonucleotide strands. In this embodiment, it is expected that PCR efficiency will be further improved by cleaving the loop site from both oligonucleotide strands.

[0152] (Regarding Form 4) As illustrated in Figure 5, in the present invention, the cleavable site may be present in both the first oligonucleotide strand (E) and the second oligonucleotide strand (F), and the structures of the cleavable sites may be different. In such a case, the difference in the properties of the two (or more) cleavable sites can be utilized to control the cleavage site. For example, deoxyuridine may be used as the cleavable site in the first oligonucleotide strand (E) and deoxyinosine as the cleavable site in the second oligonucleotide strand (F). In this case, the USER enzyme can be used to selectively cleave deoxyuridine in the first oligonucleotide strand (E). [ka] On the other hand, by using alkyladenine DNA glycosylase and endonuclease VIII, the cleavage site starting from deoxyinosine in the second oligonucleotide strand (F) can be selectively cleaved. [ka] In this way, by selecting the cleavage site as desired, a wider range of DEL modifications can be achieved, and a wider range of techniques can be applied for subsequent evaluation. We can expect this.

[0153] (Regarding Form 5) As illustrated in Figure 6, in the present invention, a cleavable site can also be provided in the DNA tag portion (e.g., the oligonucleotide strand (Y)). By providing a cleavable site near the end of the DNA tag and cleaving the site as desired, a new protruding end can be generated. [ka] The overhanging ends can be used as sticky ends to ligate desired nucleic acid sequences, such as UMIs (unique molecular identifiers). [ka] After biological evaluation, the selected DEL compounds are given UMIs regions as described above and subjected to DNA sequencing, enabling analysis with reduced PCR amplification bias. Thus, in the present invention, by having a selectively cleavable site in the nucleic acid sequence, it is possible to impart unprecedented performance to the DEL compound in terms of its production and use.

[0154] Here, UMIs (Unique Molecular Identification Sequences) are molecular identifiers that, when added to DNA contained in a sample, give each DNA molecule an individual DNA sequence (see Nature Methods, 2012, Vol. 9, pp. 72-74). By adding such molecular identifiers before PCR amplification, it becomes possible to distinguish PCR duplicates (sequences derived from the same molecule) when quantifying the number of DNA molecules with a specific sequence in a sample, making it possible to quantify with reduced PCR amplification bias.

[0155] (Regarding Form 6) As illustrated in FIG. 7, in the present invention, a cleavable site can be used in combination with a modifying group or a functional molecule, and it is possible to prepare, for example, a DEL in which a hairpin strand DNA is converted into a single-stranded DNA. As an example, a DEL compound using a headpiece with a cleavable site in the E portion according to FIG. (Step A) A double-stranded oligonucleotide chain having a solid-supported removable modifying group (eg, biotin) at the 3'-end is ligated to the synthesized DEL compound. (Step B) Cleavage of the cleavable site. (Step C) Add a treatment according to the function of the modifying group. For example, in the case of biotin, streptavidin beads with biotin affinity are used to selectively remove the oligonucleotide chains to which biotin is bound from the system. This makes it possible to obtain DELs with single-stranded DNA.

[0156] Here, a functional molecule is a molecule that has a specific chemical or biological function (e.g., solubility, photoreactivity, substrate-specific reactivity, target protein degradation induction property), and by attaching it to a DEL, it becomes possible to evaluate and purify the DEL according to its function.

[0157] Here, biotin refers to all biotins that bind to avidin, and includes not only vitamin B7 but also, for example, desthiobiotin.

[0158] As illustrated in Figure 8, DELs containing single-stranded DNA can be given new functions by forming a double strand with a modified oligonucleotide having a desired functional site (e.g., crosslinker-modified DNA such as a photoreactive crosslinker).

[0159] (Regarding Form 7) As illustrated in FIG. 9, in the present invention, a cleavable site can be utilized to introduce a cross-linker. As an example, a DEL compound using a headpiece with a cleavable site in the E portion according to FIG. (Step A) The synthesized DEL compound is cleaved at a cleavable site. (Step B) A modified primer having a desired functional site (e.g., a crosslinker-modified primer such as a photoreactive crosslinker) is provided. (Step C) The applied primer is extended to synthesize a crosslinker-modified double-stranded DEL compound. In the case of DEL evaluation, the crosslinker-modified double-stranded DEL compound can further bind the crosslink structure to the target protein when the building block compound (library low molecular weight compound) binds to the target protein, and can significantly improve the detection sensitivity (see Non-Patent Documents 5 and 6, etc.). In the practical application of DEL technology, which evaluates a very large number of library compounds, it is very useful to enhance the affinity of the library compounds and improve the detection sensitivity. The present invention provides a novel and highly efficient method for producing a cross-linker-modified double-stranded DEL compound, and is therefore extremely useful.

[0160] The present invention will be described in more detail below with reference to examples, but the present invention is not limited to these examples. Nucleic acids of various sequences in the examples can be prepared, for example, by an automatic nucleic acid synthesizer according to standard methods. An example of an automatic nucleic acid synthesizer is nS-8II (manufactured by Gene Design). Nucleic acids can also be prepared by contract synthesis or by using a contract laboratory. Examples of contract laboratories well known to those skilled in the art include Gene Design and LGC Biosearch Technologies. In general, These contract laboratories, under confidentiality agreements, prepare nucleic acids with sequences specified by the clients and deliver them to the clients.

[0161] Example 1 [Verification of the cleavage reaction of the hairpin-type DEL substructure containing deoxyuridine by USER® enzyme] Compounds having the sequences shown in Table 1 were prepared using an automatic nucleic acid synthesizer nS-8II (Gene Design). As will be clear to those skilled in the art, in the sequence notations in Table 1, each sequence unit is bonded with a phosphodiester bond, and "A" means deoxyadenosine, "T" means thymidine, "G" means deoxyguanosine, "C" means deoxycytidine, "(dU)" means deoxyuridine, "(p)" means phosphate, and "(amino-C6-dT)" is represented by the following formula (1): [ka] "(amino-NC6-dT)" means a modified nucleic acid represented by the following formula (2): [ka] "(dSpacer)" means a modified nucleic acid represented by the following formula (3): [ka] "(aminoC7)" means a group represented by the following formula (4) [ka] The amino-NC6-dT is a group represented by the following formula (5) synthesized according to the method described in the Journal of the American Chemical Society, 1993, Vol. 115, pp. 7128-7134: [ka] The nucleic acid was introduced using a nucleic acid synthesis reagent.

[0162] In Table 1, "No." in the left column indicates the sequence number, and "Seq." in the right column indicates the sequence. The left side of the sequence indicates the 5' side, and the right side indicates the 3' side. The names of the compounds corresponding to each sequence number (No.) are as follows: No.1: U-DEL1-sh No.2: U-DEL2-sh No.3: U-DEL3-sh No.4: U-DEL4-sh No.5: U-DEL5-HP No.6: U-DEL6-HP No.7: U-DEL7-HP No.8: U-DEL8-HP No.9:U-DEL9-HP No.10:U-DEL10-HP

[0163] [Table 1]

[0164] A 0.1 mM aqueous solution of each of the compounds having the sequences shown in Table 1 was prepared, and the cleavage reaction by USER (registered trademark) enzyme was examined according to the following procedure.

[0165] In a PCR tube, 1 μL of a 0.1 mM aqueous solution of the compound of the sequence shown in Table 1, 10 μL of CutSmart® Buffer (New England BioLabs, Catalog No. B7204S) and 79 μL of deionized water were added. 10 μL of USER® enzyme (New England BioLabs, Catalog No. M5505S) was added to the solution, and the resulting solution was incubated at 37° C.

[0166] 20 μL of each reaction solution was sampled 1 hour and 3 hours after the start of incubation. 20 μL of U-DEL1-sh, U-DEL5-HP, U-DEL6-HP, U-DEL7-HP, U-DEL8-HP, U-DEL9-HP, and U-DEL10-HP were sampled after 20 hours. 20 μL of U-DEL8-HP and U-DEL9-HP were sampled after further incubation at 90°C for 1 hour.

[0167] Of the sampled solutions, U-DEL1-sh, U-DEL2-sh, U-DEL3-sh, and U-DEL4-sh were analyzed under analytical condition 1 shown below, and U-DEL5-HP, U-DEL6-HP, U-DEL7-HP, U-DEL8-HP, U-DEL9-HP, and U-DEL10-HP were analyzed under analytical condition 2 shown below.

[0168] Analysis condition 1: Equipment: maXis (Bruker), UltiMate 3000 (Dionex) Column: ACQUITY UPLC Oligonucleotide BEH C18 Column (130Å, 1.7μm, 2.1×50mm) Column temperature: 50℃ solvent: Solution A: Water (0.75% v / v hexafluoroisopropanol; 0.038% v / v triethylamine; 5 μM ethylenediaminetetraacetic acid) Solution B: 90% v / v methanol in water (0.75% v / v hexafluoroisopropanol; 0.038% v / v triethylamine; 5 μM ethylenediaminetetraacetic acid) Gradient conditions: The measurement was started with a flow rate of 0.36 mL / min and a fixed mixture ratio of solutions A and B of 95 / 5 (v / v). After 0.56 minutes, the mixture ratio of solutions A and B was changed linearly to 40 / 60 (v / v) over a period of 5.5 minutes. Detection wavelength: 260nm

[0169] Analysis conditions 2: Instrument: Waters ACQUITY UPLC / SQ Detector Column: ACQUITY UPLC Oligonucleotide BEH C18 Column (130Å, 1.7μm, 2.1×50mm) Column temperature: 50℃ solvent: Solution A: Water (0.75% v / v hexafluoroisopropanol; 0.038% v / v triethylamine; 5 μM ethylenediaminetetraacetic acid) Solution B: 90% v / v methanol in water (0.75% v / v hexafluoroisopropanol; 0.038% v / v triethylamine; 5 μM ethylenediaminetetraacetic acid) Gradient conditions: The measurement was started with a flow rate of 0.36 mL / min and a fixed mixture ratio of solutions A and B of 95 / 5 (v / v). After 0.56 minutes, the mixture ratio of solutions A and B was changed linearly to 40 / 60 (v / v) over a period of 5.5 minutes. Detection wavelength: 260nm

[0170] The sequences and theoretical molecular weights of the products (abasic forms of the deoxyuridine moiety and cleaved fragments) assumed in each reaction solution, and the molecular weights detected in each reaction solution are shown in Tables 2 and 3. The symbols in each column in Tables 2 and 3 are as follows:

[0171] “Entry” (leftmost): The experiment number is shown, and the substrates corresponding to each experiment number (Entry) are as follows. Entry.1:U-DEL1-sh Entry.2:U-DEL2-sh Entry.3:U-DEL3-sh Entry.4:U-DEL4-sh Entry.5:U-DEL5-HP Entry.6:U-DEL6-HP Entry.7:U-DEL7-HP Entry.8:U-DEL8-HP Entry.9:U-DEL9-HP Entry.10:U-DEL10-HP

[0172] “No.”(second from the left): The sequence numbers are shown below. Of the sequence numbers (No.), No. 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 are substrates for the respective reaction solutions, No. 11, 14, 17, 20, 22, 25, 29, 31, 33, and 35 are abasic forms of the deoxyuridine moiety of each substrate, and the remaining sequence numbers are fragments obtained by cleavage of each substrate.

[0173] “Seq.” (third from the left): The sequence is shown, with the left side representing the 5' side and the right side representing the 3' side. In the sequence notation, "(B)" is the following formula (6) [ka] The other notations are the same as in Table 1.

[0174] “Expected MW.” (fourth from the left): The theoretical molecular weight (Da) of each sequence is shown.

[0175] "Observed MW." (far right): The numerical value of the detected molecular weight (Da) identified for each sequence is shown. Note that "-" indicates that it was not detected.

[0176] [Table 2]

[0177] [Table 3]

[0178] The conversion rates of the abasic and cleavage reactions were calculated from the area ratio of the peaks corresponding to each detected sequence. The abasic reactions were 99% or more complete at 37°C for all substrates after 1 hour (the substrate peak was less than 1%, and the remaining peaks were only the abasic form and cleavage fragments). A graph showing the conversion rate of the cleavage reaction is shown in Figure 10. As shown in the graph, for all substrates except U-DEL8-HP and U-DEL9-HP, the cleavage reaction had progressed to 95% or more by 20 hours at 37°C, and the cleavage reaction was also completed to 100% in U-DEL8-HP and U-DEL9-HP by adding an incubation time of 1 hour at 90°C.

[0179] The above results indicate that the partial structures of hairpin-type DEL containing various deoxyuridines undergo an abasic reaction by USER (registered trademark) enzyme at the deoxyuridine site, followed by a cleavage reaction.

[0180] Example 2 [Comparison of PCR efficiency between conventional hairpin DEL and cleavable hairpin DEL (hairpin DEL containing deoxyuridine)]

[0181] As shown in the schematic diagram of FIG. 11, the compound (hairpin DEL) having the sequence shown in Table 4 was synthesized by the following procedure. In the sequence notation in Table 4, "S" represents the following formula (7): [ka] The other symbols are the same as in Table 1. The names of the compounds corresponding to each sequence number are as follows: No.37: U-DEL1 No.38: U-DEL2 No.39: U-DEL4 No.40: U-DEL7 No.41: U-DEL8 No.42: U-DEL9 No.43: U-DEL10 No.44: H-DEL [Table 4] The compound names of the raw headpieces used to synthesize each hairpin DEL are as follows: Hairpin DEL: Raw headpiece U-DEL1 :U-DEL1-HP U-DEL2 :U-DEL2-HP U-DEL4 :U-DEL4-HP U-DEL7 :U-DEL7-HP U-DEL8 :U-DEL8-HP U-DEL9 :U-DEL9-HP U-DEL10 :U-DEL10-HP H-DEL :H-DEL-HP Furthermore, the sequence numbers "No." and "Seq" of U-DEL1-HP, U-DEL2-HP, U-DEL4-HP, and H-DEL-HP are as shown in Table 5 below. [Table 5]

[0182] The raw headpiece shown in Table 5 was prepared in the same manner as in Example 1 using an automatic nucleic acid synthesizer nS-8II (Gene Design Co., Ltd.).

[0183] In a PCR tube, 2.0 μL of 1 mM aqueous solution of various raw head pieces; 2.4 μL of 1 mM aqueous solution of Pr_TAG (prepared by annealing Pr_TAG_a and Pr_TAG_b synthesized in the same manner as in Example 1, the sequence is shown in Table 6); 0.8 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 100 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) and 2.0 μL of deionized water were added. 0.8 μL of a 10-fold diluted aqueous solution of T4 DNA ligase (manufactured by Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16 ° C. for 24 hours. The sequence notation in Table 6 is the same as in Table 1. The names of the compounds corresponding to each sequence number (No.) are as follows. No.49:Pr_TAG_a No.50:Pr_TAG_b [Table 6]

[0184] The reaction solution was treated with 0.8 μL of 5 M aqueous sodium chloride solution and 17.6 μL of chilled (−20° C.) ethanol, and then incubated at −78° C. for 2 hours. After centrifugation, the supernatant was removed and the resulting pellets were air-dried. 2.0 μL of deionized water was added to each pellet to prepare a solution.

[0185] To each of the obtained solutions, 2.4 μL of 1 mM aqueous solution of CP (prepared by annealing CP_a and CP_b synthesized in the same manner as in Example 1, the sequence is shown in Table 7); 0.8 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 100 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) and 2.0 μL of deionized water were added. 0.8 μL of a 10-fold diluted aqueous solution of T4 DNA ligase (manufactured by Thermo Fisher, catalog number EL0013) was added to the solution, and the obtained solution was incubated at 16° C. for 24 hours. The sequence notation in Table 7 is the same as in Table 1. The names of the compounds corresponding to each sequence number (No.) are as follows. No.51: CP_a No.52: CP_b

Table 7

[0186] The reaction solution was treated with 0.8 μL of 5 M aqueous sodium chloride solution and 17.6 μL of cooled (-20 °C) ethanol, and left standing at -78 °C for 2 hours. After centrifugation, the supernatant was removed, and the obtained pellet was air-dried. 10 μL of deionized water was added to the pellet to form a solution.

[0187] 1.0 μL of the obtained solution was sampled, diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under the conditions of Analytical Condition 2 in Example 1 to identify the target substance (the theoretical molecular weights and the detected molecular weights of each sequence are shown in Table 4). The rest of the solution was lyophilized, and then deionized water was added to each to adjust to 20 μM.

[0188] Among the 8 types of hairpin-type DELs obtained above, H-DEL is a conventional hairpin DEL, and the remaining 7 types are cleavable hairpin DELs containing deoxyuridine. To compare the PCR efficiency before treatment with USER (registered trademark) enzyme and the PCR efficiency after treatment of various hairpin-type DELs, real-time PCR analysis was performed. Also, DS-DEL shown in Table 7 (prepared by annealing the compounds of Sequence Nos. 47 and 48) was used as the double-stranded DEL for comparison. In the sequence listings in Table 8, “(amino-C6-L)” means the group represented by the following formula (8)

Chemical formula

Table 8

[0189] <Process of treatment with USER (registered trademark) enzyme> The treatment of 8 types of hairpin DELs and double-stranded DEL (DS-DEL) with USER® enzyme was performed according to the following procedure.

[0190] To a PCR tube, 1 μL of various DEL 20 μM aqueous solutions; 1 μL of CutSmart® Buffer (manufactured by New England BioLabs, catalog number B7204S) and 7 μL of deionized water were added. 1 μL of USER® enzyme (manufactured by New England BioLabs, catalog number M5505S) was added to the solution, and the resulting solution was incubated at 37 °C for 1 hour.

[0191] <Preparation of DEL Samples> Samples of various DELs before USER® enzyme treatment and the reaction solutions after treatment were each diluted with deionized water to prepare 0.05 pM, 0.5 pM, and 5 pM DEL samples.

[0192] <Measurement of Ct Values by Real-Time PCR> The Ct values of the various DEL samples obtained above were measured by real-time PCR to compare the PCR efficiencies. The conditions were as follows, and the results are shown in Figure 12. Note that the Ct value is the number of cycles at which the fluorescence signal generated with the amplification of DNA reaches an arbitrary threshold value in real-time PCR. That is, when the initial number of DNA molecules is the same, the higher the PCR efficiency, the lower the Ct value.

[0193] Equipment: 7500 Real-Time PCR System (manufactured by Applied Biosystems) Plate: MicroAmp 96-Well Plate (manufactured by Applied Biosystems, catalog number N8010560) PCR Reaction Solution: · TB Green Premix Ex taqII (manufactured by Takara Bio, catalog number RR820): 10 μL · Forward Primer (Table 9, SEQ ID NO: 55): 0.80 μL Reverse primer (Table 9, SEQ ID NO: 56): 0.80 μL ROX Reference Dye II (Takara Bio, catalog number RR39LR): 0.40 μL Aqueous solutions of various DEL samples (0.05pM, 0.5pM, 5pM)*1: 2.0μL Deionized water: 6.0 μL *1: The number of moles of DEL samples is 0.1 amol, 1 amol, and 10 amol. Temperature conditions: After holding at 95°C for 2 minutes, the following cycle was repeated 35 times. 95°C for 5 seconds 52°C for 30 seconds 72°C, 30 seconds [Table 9] The sequence notation in Table 9 is the same as in Table 1.

[0194] As shown in Figure 12, the Ct value of conventional hairpin DEL (H-DEL) did not change before and after USER (registered trademark) enzyme treatment, but the Ct values ​​of the cleavable hairpin DELs containing deoxyuridine (U-DEL1, U-DEL2, U-DEL4, U-DEL7, U-DEL8, U-DEL9, and U-DEL10) decreased to the same level as that of the double-stranded DEL, DS-DEL, after USER (registered trademark) enzyme treatment.

[0195] These results indicate that the PCR efficiency of DEL cleaved by USER® enzyme is improved compared to before cleavage, and that the cleavable hairpin DEL containing deoxyuridine is cleaved with high efficiency and selectivity by USER® enzyme.

[0196] Example 3 [Verification of the cleavage reaction of hairpin DEL containing deoxyuridine by USER® enzyme] <Synthesis of four hairpin DELs (U-DEL5, U-DEL11, U-DEL12, and U-DEL13)> The compound (hairpin DEL) having the sequence shown in Table 10 was synthesized by the following procedure. In the sequence notation in Table 10, "[mdC(TEG-amino)]" is represented by the following formula (9): [ka] The other symbols are the same as in Table 4. The names of the compounds corresponding to each sequence number are as follows: No.57: U-DEL5 No.58: U-DEL11 No.59: U-DEL12 No.60: U-DEL13 [Table 10] The compound names of the raw headpieces used to synthesize each hairpin DEL are as follows: Hairpin DEL: Raw headpiece U-DEL5 :U-DEL5-HP U-DEL11 :U-DEL11-HP U-DEL12 :U-DEL12-HP U-DEL13 :U-DEL13-HP Furthermore, the sequence numbers "No." and "Seq" of U-DEL11-HP, U-DEL12-HP, and U-DEL13-HP are as shown in Table 11 below. The notations in Table 11 are the same as those in Table 10.

[0197] [Table 11]

[0198] Among the raw head pieces shown in Table 11, U-DEL12-HP and U-DEL13-HP were prepared using an automatic nucleic acid synthesizer nS-8II (manufactured by Gene Design Co., Ltd.) in the same manner as in Example 1. U-DEL11-HP was also prepared according to the standard method.

[0199] As in Example 2, various starting headpieces were used to carry out two-step double-stranded ligation with the double-stranded oligonucleotide Pr_TAG and CP.

[0200] A portion of the resulting solution was sampled and diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under the following analysis condition 3 to identify the target substance (the theoretical molecular weight of each sequence and the detected molecular weight are shown in Table 10). The remaining solution was freeze-dried, and then deionized water was added to each solution to prepare a concentration of 20 μM.

[0201] Analysis condition 3: Instrument: Waters ACQUITY UPLC / SQ Detector Column: ACQUITY UPLC Oligonucleotide BEH C18 Column (130Å, 1.7μm, 2.1×50mm) Column temperature: 60℃ solvent: Solution A: Water (0.75% v / v hexafluoroisopropanol; 0.038% v / v triethylamine; 5 μM ethylenediaminetetraacetic acid) Solution B: 90% v / v methanol in water (0.75% v / v hexafluoroisopropanol; 0.038% v / v triethylamine; 5 μM ethylenediaminetetraacetic acid) Gradient conditions: The measurement was started with a flow rate of 0.36 mL / min and a fixed mixture ratio of solutions A and B of 95 / 5 (v / v). After 0.56 minutes, the mixture ratio of solutions A and B was changed linearly to 40 / 60 (v / v) over a period of 5.5 minutes. Detection wavelength: 260nm Deconvolution: The ion signal was analyzed using ProMass for MassLynx Software (manufactured by Waters).

[0202] <Cutting reaction by <USER(trademark)> enzyme The cutting reaction of hairpin DELs (U-DEL5, U-DEL7, U-DEL9, U-DEL11, U-DEL12, and U-DEL13) containing six types of deoxyuridine by <USER(trademark)> enzyme was examined according to the following procedure.

[0203] To a PCR tube, 2 μL of various hairpin DEL 20 μM aqueous solution; 2 μL of CutSmart (trademark) Buffer (manufactured by New England BioLabs, catalog number B7204S) and 14 μL of deionized water were added. 2 μL of <USER(trademark)> enzyme (manufactured by New England BioLabs, catalog number M5505S) was added to the solution, and the resulting solution was incubated at 37 °C for 16 hours and then further incubated at 90 °C for 1 hour.

[0204] <Confirmation of the product after cleavage by LC-MS measurement 5.0 μL of the obtained reaction solution was sampled, diluted with deionized water, and then mass spectrometry was performed by ESI-MS under analysis condition 3. The sequences and theoretical molecular weights of the products expected after cleavage in each reaction solution, and the molecular weights detected in each reaction solution are shown in Table 12. The substrates corresponding to each experimental number (Entry) are as follows, and other notations are the same as in Table 10. Entry.1: U-DEL5 Entry.2: U-DEL7 Entry.3: U-DEL9 Entry.4: U-DEL11 Entry.5: U-DEL12 Entry.6: U-DEL13

[0205]

Table 12

[0206] In all samples, no MS of the substrate was detected, and the MS of the cleavage product was observed as the main peak.

[0207] <Confirmation of cleavage reaction by gel electrophoresis> In addition, a portion of the reaction solution obtained was sampled and analyzed by denaturing polyacrylamide gel electrophoresis under the conditions shown below. From the results shown in Figure 13, it was confirmed that the cleavage reaction proceeded with high yield for all substrates. The samples in each lane in Figure 13 are as follows. Lane 1: 20 bp DNA ladder (Lonza 20 bp DNA Ladder, Cat. No. 50330) Lane 2: U-DEL5 Lane 3: Sample after U-DEL5 cleavage reaction Lane 4: U-DEL7 Lane 5: Sample after U-DEL7 cleavage reaction Lane 6: U-DEL9 Lane 7: Sample after U-DEL9 cleavage reaction Lane 8: U-DEL11 Lane 9: Sample after U-DEL11 cleavage reaction Lane 10: U-DEL12 Lane 11: Sample after U-DEL12 cleavage reaction Lane 12: U-DEL13 Lane 13: Sample after U-DEL13 cleavage reaction Denaturing polyacrylamide gel electrophoresis: Gel: Novex (trademark) 10% TBE-Urea gel (Invitrogen by ThermoFisher SCIENTIFIC, catalog number EC68755BOX) Loading buffer: Novex 10% TBE-Urea Sample Buffer (2x) (Invitrogen by ThermoFisher SCIENTIFIC, Catalog No. LC6876) Temperature: 60℃ Voltage: 180V Running time: 30 minutes Staining reagent: SYBER (trademark) Green II Nucleic Acid Gel Stain (Takara Bio, catalog number 5770A)

[0208] The above results indicate that hairpin-type DELs containing various deoxyuridines undergo cleavage reaction by USER (registered trademark) enzyme at the deoxyuridine site.

[0209] Example 4 [Verification of the cleavage reaction of hairpin DEL containing deoxyinosine by endonuclease V] <Synthesis of four types of deoxyinosine-containing hairpin DEL (I-DEL1, I-DEL2, I-DEL3, and I-DEL4)> The compound (hairpin DEL) having the sequence shown in Table 13 was synthesized by the following procedure. In the sequence notation in Table 13, "I" means deoxyinosine, and the other notations are the same as those in Table 2. The names of the compounds corresponding to each sequence number are as follows: No.73:I-DEL1 No.74:I-DEL2 No.75: I-DEL3 No.76: I-DEL4 [Table 13] The compound names of the raw headpieces used to synthesize each hairpin DEL are as follows: Hairpin DEL: Raw headpiece I-DEL1 : I-DEL1-HP I-DEL2 : I-DEL2-HP I-DEL3 : I-DEL3-HP I-DEL4 : I-DEL4-HP Furthermore, the sequence numbers "No." and "Seq" of I-DEL1-HP, I-DEL2-HP, I-DEL3-HP, and I-DEL4-HP are as shown in Table 14 below. The notations in Table 14 are the same as those in Table 13.

[0210]

Table 14

[0211] The raw material headpiece shown in Table 14 was prepared according to a conventional method.

[0212] Similar to Example 2, two-step double-stranded ligation with the double-stranded oligonucleotide Pr_TAG and CP was carried out using various raw material headpieces.

[0213] A part of the obtained solution was sampled, diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under analytical condition 3 to identify the target substance (the theoretical molecular weights and the detected molecular weights of each sequence are shown in Table 13). After freeze-drying the remaining solution, deionized water was added to each to prepare a concentration of 20 μM.

[0214] <Cleavage reaction by endonuclease V> The cleavage reaction of hairpin DEL (I-DEL1, I-DEL2, I-DEL3, I-DEL4) containing four kinds of deoxyinosine by endonuclease V was examined by the following procedure.

[0215] To a PCR tube, 1 μL of various hairpin DEL 20 μM aqueous solutions, 2 μL of NEBuffer (registered trademark) 4 (manufactured by New England BioLabs, catalog number B7004), and 15 μL of deionized water were added. 2 μL of Endonuclease V (manufactured by New England BioLabs, catalog number M0305S) was added to the solution, and the resulting solution was incubated at 37 °C for 24 hours.

[0216] <Confirmation of the product after cleavage by LC-MS measurement> 8.0 μL of the resulting reaction solution was sampled and diluted with deionized water, and then mass spectrometry was performed by ESI-MS under analysis condition 3. The sequence and theoretical molecular weight of the product after cleavage expected in each reaction solution, and the molecular weight detected in each reaction solution are shown in Table 15. The substrates corresponding to each experiment number (Entry) are as follows, and other notations are the same as in Table 13. Entry.1: I-DEL1 Entry.2: I-DEL2 Entry.3: I-DEL3 Entry.4: I-DEL4

[0217] [Table 15]

[0218] In all samples, no MS of the substrate was detected, and the MS of the cleavage product was observed as the main peak.

[0219] <Confirmation of cleavage reaction by gel electrophoresis> In addition, a portion of the reaction solution obtained was sampled and analyzed by denaturing polyacrylamide gel electrophoresis under the same conditions as in Example 3. From the results shown in Figure 14, it was confirmed that the cleavage reaction proceeded with high yield for all substrates. The samples in each lane in Figure 14 are as follows. Lane 1: 20 bp DNA ladder (Lonza 20 bp DNA Ladder, Cat. No. 50330) Lane 2: I-DEL1 Lane 3: Sample after I-DEL1 cleavage reaction Lane 4: I-DEL2 Lane 5: Sample after I-DEL2 cleavage reaction Lane 6: I-DEL3 Lane 7: Sample after I-DEL3 cleavage reaction Lane 8: I-DEL4 Lane 9: Sample after I-DEL4 cleavage reaction

[0220] These results indicate that in hairpin-type DELs containing various deoxyinosines, the second phosphodiester bond in the 3' direction from the deoxyinosine is cleaved by endonuclease V.

[0221] Example 5 [Verification of cleavage reaction of hairpin DEL containing ribonucleoside by RNase HII] <Synthesis of ribonucleoside-containing hairpin DEL (R-DEL1)> The compound (hairpin DEL) having the sequence shown in Table 16 was synthesized by the following procedure. In the sequence notation in Table 16, "u" means uridine, and the other notations are the same as those in Table 2. The names of the compounds corresponding to the sequence numbers are as follows: No.87:R-DEL1 [Table 16] The compound names of the raw headpieces used to synthesize each hairpin DEL are as follows: Hairpin DEL: Raw headpiece R-DEL1 : R-DEL1-HP Furthermore, the sequence number "No." and the sequence "Seq" of R-DEL1-HP are as shown in Table 17 below. Note that the notation in Table 17 is the same as that in Table 16.

[0222] [Table 17]

[0223] The raw headpieces shown in Table 17 were prepared according to a conventional method.

[0224] As in Example 2, the starting headpiece was used to carry out two-step double-stranded ligation with the double-stranded oligonucleotide Pr_TAG and CP.

[0225] A portion of the obtained solution was sampled, diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under Analytical Condition 3 to identify the target substance (the theoretical molecular weights and the detected molecular weights of each sequence are shown in Table 16). After the remaining solution was lyophilized, deionized water was added to each to prepare a concentration of 200 μM.

[0226] <RNaseHII Cleavage Reaction> The cleavage reaction of ribonucleoside-containing hairpin DEL (R-DEL1) by RNaseHII was examined according to the following procedure.

[0227] To a PCR tube, 0.5 μL of 200 μM aqueous solution of hairpin DEL; 4.9 μL of ThermoPol® Reaction Buffer Pack (manufactured by New England BioLabs, catalog number B9004) and 43.6 μL of deionized water were added. 1 μL of RNase HII (manufactured by New England BioLabs, catalog number M0288S) was added to the solution, and the resulting solution was incubated at 37 °C for 8 hours.

[0228] <Confirmation of Products after Cleavage by LC-MS Measurement> 10 μL of the obtained reaction solution was sampled and subjected to mass spectrometry by ESI-MS under Analytical Condition 3. The sequences, theoretical molecular weights, and detected molecular weights of the expected products after cleavage are shown in Table 18. The substrates corresponding to the experimental numbers (Entry) are as follows, and other notations are the same as in Table 16. Entry.1: R-DEL1

[0229]

Table 18

[0230] For all samples, the MS of the substrate was not detected, and the MS of the product after cleavage was observed as the main peak.

[0231] <Confirmation of Cleavage Reaction by Gel Electrophoresis> In addition, a portion of the obtained reaction solution was sampled and analyzed by denaturing polyacrylamide gel electrophoresis under the same conditions as in Example 3. From the results shown in Figure 15, it was confirmed that the cleavage reaction proceeded with high yield for all substrates. The samples in each lane in Figure 15 are as follows. Lane 1: 20 bp DNA ladder (Lonza 20 bp DNA Ladder, Cat. No. 50330) Lane 2: R-DEL1 Lane 3: Sample after R-DEL1 cleavage reaction

[0232] These results indicate that in hairpin-type DELs containing ribonucleosides, the phosphodiester bond at the 5' end of the ribonucleotide is cleaved by RNase HII. Example 6 [Creating a model library using U-DEL9-HP as a raw material] As shown in the schematic diagram in Figure 16, a model library containing 3 × 3 × 3 (27) compounds was synthesized by split-and-pool synthesis using U-DEL9-HP as the starting material and the following reagents. ·U-DEL9-HP Three building blocks (BB1, BB2, and BB3): [ka] 10 double-stranded oligonucleotide tags (tag numbers in Table 19: Pr, A1, A2, A3, B1, B2, B3, C1, C2, and C3)

[0233] In Table 19, "Tag No." (leftmost) indicates the tag number, "No." (second from the left) indicates the sequence number, and "Seq." (third from the left) indicates the sequence. The sequence notation is the same as in Table 1.

[0234] Each double-stranded oligonucleotide tag was prepared by annealing two oligonucleotides having the sequence numbers shown in Table 19, which correspond to each tag number.

[0235] [Table 19]

[0236] <Synthesis of compound "AOP-U-DEL9-HP"> The compound "AOP-U-DEL9-HP" having the sequence shown in Table 20 was synthesized by the following procedure. In the sequence notation in Table 20, "(AOP-AminoC7)" is represented by the following formula (10). [ka] The other symbols are the same as in Table 2.

[0237] [Table 20]

[0238] To four Violamo centrifuge tubes was added a solution of U-DEL9-HP (2.5 mL, 1 mM) in sodium borate buffer (150 mM, pH 9.4) cooled to 10° C. To each tube was added 40 equivalents of N-Fmoc-15-amino-4,7,10,13-tetraoxaoctadecanoic acid (250 μL, 0.4 M N-dimethylacetamide solution), followed by 40 equivalents of 4-(4,6-dimethoxy[1.3.5]triazin-2-yl)-4-methylmorpholinium chloride hydrate (DMTMM) (200 μL, 0.5 M aqueous solution), and the resulting solutions were shaken at 10° C. for 5 h.

[0239] The above solutions were treated with 295 μL of 5M aqueous sodium chloride solution and 9.7 mL of chilled (-20°C) ethanol, and left to stand overnight at -78°C. After centrifugation, the supernatant was removed and the resulting pellets were air-dried. The pellets were dissolved in 2.75 mL of deionized water, and 306 μL of piperidine was added at 0°C and shaken at 10°C for 3 hours. After centrifuging the mixtures, the precipitate was removed by filtration and washed twice with 1.47 mL of deionized water. The obtained filtrates were treated with 600 μL of 5M aqueous sodium chloride solution and 19.8 mL of chilled (-20°C) ethanol, and left to stand overnight at -78°C. After centrifugation, the supernatant was removed and the resulting pellets were air-dried.

[0240] The pellets thus obtained were mixed with 10 mL of deionized water to prepare a solution. A portion of the resulting solution was sampled and diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under the analytical condition 2 of Example 1 to identify the target substance (the theoretical molecular weight of the compound and the detected molecular weight are shown in Table 20). The remaining part of the solution was freeze-dried, and then deionized water was added to prepare a 5 mM solution.

[0241] <Introduction of double-stranded oligonucleotide tag "Pr"> The compound "AOP-U-DEL9-HP" and the double-stranded oligonucleotide tag "Pr" were ligated to synthesize the compound "AOP-U-DEL9-HP-Pr" having the sequence shown in Table 21 by the following procedure. The sequence notation in Table 21 is the same as in Table 20.

[0242] [Table 21]

[0243] In a Violamo centrifuge tube, 40 μL of a 5 mM aqueous solution of the compound "AOP-U-DEL9-HP", 160 μL of a 100 mM aqueous solution of sodium bicarbonate, 240 μL of a 1 mM aqueous solution of the double-stranded oligonucleotide tag "Pr", 80 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 100 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) and 272 μL of deionized water were added. 8.0 μL of T4 DNA ligase (Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16 °C for 24 hours.

[0244] The reaction solution was treated with 80 μL of 5M aqueous sodium chloride solution and 2640 μL of cooled (-20°C) ethanol, and left at -78°C for 2 hours. After centrifugation, the supernatant was removed, and 400 μL of deionized water was added to the resulting pellet. The resulting solution was concentrated using an Amicon (registered trademark) Ultra Centrifugal filter (30 kD cutoff). A portion of the resulting solution was sampled and subjected to mass spectrometry by ESI-MS under the analytical condition 2 to identify the target substance (the theoretical molecular weight of the compound and the detected molecular weight are shown in Table 21). Through the above process, 133 nmol of the compound "AOP-U-DEL9-HP-Pr" with a purity of 84.5% was obtained. A 100 mM aqueous sodium bicarbonate solution was added to the resulting compound "AOP-U-DEL9-HP-Pr" to prepare a 1 mM solution.

[0245] <Cycle A> To each of the three PCR tubes, 20 μL of a 1 mM solution of the compound "AOP-U-DEL9-HP-Pr" obtained above, 30 μL of a 1 mM aqueous solution of one of the double-stranded oligonucleotide tags A1 to A3, 8.0 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 100 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) and 21.6 μL of deionized water were added. 0.4 μL of T4 DNA ligase (Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16 °C for 18 hours.

[0246] Each reaction solution was treated with 8.0 μL of 5 M aqueous sodium chloride solution and 264 μL of chilled (−20° C.) ethanol, and allowed to stand at −78° C. for 30 minutes. After centrifugation, the supernatant was removed, and the resulting pellets were each dissolved in 20 μL of 150 mM sodium borate buffer (pH 9.4).

[0247] To each tube, 40 equivalents of one of the building blocks BB1 to BB3 (4.0 μL, 200 mM N,N-dimethylacetamide solution) was added, followed by 40 equivalents of 4-(4,6-dimethoxy[1.3.5]triazin-2-yl)-4-methylmorpholinium chloride hydrate (DMTMM) (4.0 μL, 200 mM aqueous solution), and the tubes were shaken at 10 ° C for 2 hours. Furthermore, 20 equivalents of the building block (2.0 μL, 200 mM N,N-dimethylacetamide solution) was added, followed by 20 equivalents of DMTMM (2.0 μL, 200 mM aqueous solution), and the tubes were shaken at 10 ° C for 30 minutes.

[0248] Each reaction solution was treated with 3.2 μL of 5 M aqueous sodium chloride solution and 106 μL of chilled (−20° C.) ethanol, and left to stand for 30 minutes at −78° C. After centrifugation, the supernatant was removed, and 18 μL of deionized water was added to each of the resulting pellets, and the three solutions were mixed in one PCR tube.

[0249] To the mixed solution, 6.0 μL of piperidine was added at 0° C., and the mixture was shaken at room temperature for 1 hour. The reaction solution was treated with 6.0 μL of 5 M aqueous sodium chloride solution and 198 μL of chilled (−20° C.) ethanol, and allowed to stand at −78° C. for 18 hours. After centrifugation, the supernatant was removed, and 400 μL of deionized water was added to the resulting pellet. The resulting solution was concentrated using an Amicon® Ultra Centrifugal filter (30 kD cutoff), and 100 mM aqueous sodium bicarbonate solution was added to adjust the concentration to 1 mM, and the solution was used as the starting material for the next step.

[0250] <Cycle B> To each of three PCR tubes, 13.7 μL of the 1 mM solution of the starting material obtained in cycle A, 20.6 μL of a 1 mM aqueous solution of one of the double-stranded oligonucleotide tags B1 to B3, 5.5 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 100 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) and 14.8 μL of deionized water were added. 0.3 μL of T4 DNA ligase (Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16 °C for 16 hours.

[0251] Each reaction solution was treated with 5.5 μL of 5 M aqueous sodium chloride solution and 181 μL of chilled (−20° C.) ethanol, and allowed to stand at −78° C. for 30 minutes. After centrifugation, the supernatant was removed, and the resulting pellets were each dissolved in 13.7 μL of 150 mM sodium borate buffer (pH 9.4).

[0252] To each tube, 80 equivalents of one of the building blocks BB1 to BB3 (5.5 μL, 200 mM N,N-dimethylacetamide solution) was added, followed by 80 equivalents of DMTMM (5.5 μL, 200 mM aqueous solution), and the tubes were shaken at 10 ° C for 1 hour. Furthermore, to each tube, 40 equivalents of the building block (2.3 μL, 200 mM N,N-dimethylacetamide solution) was added, followed by 40 equivalents of DMTMM (2.3 μL, 200 mM aqueous solution), and the tubes were shaken at 10 ° C for 2 hours.

[0253] Each reaction solution was treated with 2.5 μL of 5 M aqueous sodium chloride solution and 81.4 μL of chilled (−20° C.) ethanol, and allowed to stand at −78° C. for 30 minutes. After centrifugation, the supernatant was removed, and 12.3 μL of deionized water was added to each of the resulting pellets, and the three solutions were mixed in one PCR tube.

[0254] To the mixed solution, 4.1 μL of piperidine was added at 0° C., and the mixture was shaken at room temperature for 3 hours. The reaction solution was treated with 4.1 μL of 5 M aqueous sodium chloride solution and 136 μL of chilled (−20° C.) ethanol, and allowed to stand at −78° C. for 3 hours. After centrifugation, the supernatant was removed, and 400 μL of deionized water was added to the resulting pellet. The resulting solution was concentrated using an Amicon® Ultra Centrifugal filter (30 kD cutoff), and 100 mM aqueous sodium bicarbonate solution was added to adjust the concentration to 0.48 mM, which was then used as the starting material for the next step.

[0255] <Cycle C> To each of three PCR tubes, 14.5 μL of a 0.48 mM solution of the starting material obtained in cycle B, 10.5 μL of a 1 mM aqueous solution of one of the double-stranded oligonucleotide tags C1 to C3, and 2.8 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 100 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) were added. 0.14 μL of T4 DNA ligase (Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16 °C for 16 hours.

[0256] Each reaction solution was treated with 2.8 μL of 5 M aqueous sodium chloride solution and 92 μL of chilled (−20° C.) ethanol, and allowed to stand at −78° C. for 30 minutes. After centrifugation, the supernatant was removed, and the resulting pellets were each dissolved in 7.0 μL of 150 mM sodium borate buffer (pH 9.4).

[0257] To each tube, 80 equivalents of one of the building blocks BB1 to BB3 (2.8 μL, 200 mM N,N-dimethylacetamide solution) was added, followed by 80 equivalents of DMTMM (2.8 μL, 200 mM aqueous solution), and the tubes were shaken at 10°C for 1 hour. Furthermore, to each tube, 40 equivalents of the building block (1.4 μL, 200 mM N,N-dimethylacetamide solution) was added, followed by 40 equivalents of DMTMM (1.4 μL, 200 mM aqueous solution), and the tubes were shaken at 10°C for 2 hours.

[0258] Each reaction solution was treated with 1.3 μL of 5 M aqueous sodium chloride solution and 41.4 μL of chilled (−20° C.) ethanol, and left to stand for 30 minutes at −78° C. After centrifugation, the supernatant was removed, and 6.3 μL of deionized water was added to each of the resulting pellets, and the three solutions were mixed in one PCR tube.

[0259] To the mixed solution, 2.1 μL of piperidine was added at 0° C., and the mixture was shaken at room temperature for 2 hours. The reaction solution was treated with 2.1 μL of 5 M aqueous sodium chloride solution and 69 μL of chilled (−20° C.) ethanol, and allowed to stand at −78° C. for 3 hours. After centrifugation, the supernatant was removed, and 400 μL of deionized water was added to the resulting pellet. The resulting solution was concentrated using an Amicon® Ultra Centrifugal filter (30 kD cutoff), and 100 mM aqueous sodium bicarbonate solution was added to adjust the concentration to 0.41 mM, which was then used as the starting material for the next step.

[0260] <CPのライゲーション> A PCR tube was charged with 12.2 μL of 0.41 mM solution of starting material obtained in cycle C, 6.0 μL of 1 mM aqueous solution of CP (same as used in Example 2), 2.1 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 100 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) and 0.7 μL of deionized water. 0.1 μL of T4 DNA ligase (Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16° C. for 16 hours.

[0261] The reaction solution was treated with 2.1 μL of 5 M aqueous sodium chloride solution and 69.6 μL of chilled (−20° C.) ethanol, and allowed to stand at −78° C. for 30 minutes. After centrifugation, the supernatant was removed, and 400 μL of deionized water was added to the resulting pellet. The resulting solution was concentrated using an Amicon® Ultra Centrifugal filter (30 kD cutoff), and deionized water was added to adjust the concentration to 20 μM.

[0262] <Result> The samples after ligation of the double-stranded oligonucleotide tag in each cycle were analyzed by electrophoresis using a 2.2% agarose gel (Lonza, FlashGel® cassette, catalog number 57031). The results shown in FIG. 17 confirmed that encoding by the double-stranded oligonucleotide tag was achieved with high efficiency in each cycle. The samples in each lane in FIG. 17 are as follows: Lane 1: AOP-U-DEL9-HP-Pr Lane 2: Sample after ligation of double-stranded oligonucleotide tag A1 in cycle A Lane 3: Sample after ligation of double-stranded oligonucleotide tag A2 in cycle A Lane 4: Sample after ligation of double-stranded oligonucleotide tag A3 in cycle A Lane 5: Sample after ligation of double-stranded oligonucleotide tag B1 in cycle B Lane 6: Sample after ligation of double-stranded oligonucleotide tag B2 in cycle B Lane 7: Sample after ligation of double-stranded oligonucleotide tag B3 in cycle B Lane 8: Sample after ligation of double-stranded oligonucleotide tag C1 in cycle C Lane 9: Sample after ligation of double-stranded oligonucleotide tag C2 in cycle C Lane 10: Sample after ligation of double-stranded oligonucleotide tag C3 in cycle C Lane 11: Sample after CP ligation Lane 12: 20 bp DNA ladder (Lonza 20 bp DNA Ladder, Cat. No. 50330)

[0263] The sample after the completion of cycle C was analyzed under analysis condition 3. Figure 18 shows the chromatograph and mass spectrum results. By deconvolution of the obtained mass spectrum, the average molecular weight was observed to be 35532.4. This result is consistent with the average molecular weight expected after the completion of cycle C (35514.2), indicating that the reaction for library synthesis (ligation of double-stranded oligonucleotide tags and introduction of building blocks) was achieved with high efficiency.

[0264] As a result, the synthesis of a model library containing 3 × 3 × 3 (27) compound species was achieved using U-DEL9-HP as the raw material by the above synthesis procedure.

[0265] <Cleavage of the obtained model library with USER (registered trademark) enzyme> The cleavage reaction of the model library obtained above with USER (registered trademark) enzyme was carried out according to the following procedure.

[0266] 2.0 μL of 20 μM model library solution in water, 2 μL of CutSmart® Buffer (New England BioLabs, Catalog No. B7204S) and 14 μL of deionized water were added to a PCR tube. 2 μL of USER® enzyme (New England BioLabs, Catalog No. M5505S) was added to the solution, and the resulting solution was incubated at 37° C. for 16 hours, followed by an additional incubation at 90° C. for 1 hour.

[0267] A portion of the resulting reaction solution was sampled and analyzed by denaturing polyacrylamide gel electrophoresis under the same conditions as in Example 3. From the results shown in Figure 19, it was confirmed that the model library using U-DEL9-HP as the raw material underwent a highly efficient cleavage reaction with USER (registered trademark) enzyme. The samples in each lane in Figure 19 are as follows. Lane 1: 20 bp DNA ladder (Lonza 20 bp DNA Ladder, Cat. No. 50330) Lane 2: Model library Lane 3: Sample after cleavage reaction of model library with USER® enzyme

[0268] Example 7 [Conversion of hairpin DNA from DEL compounds to single-stranded DNA and the addition of new functions] <Synthesis of "BIO-DEL", a DEL compound with biotin at the 3' end> Similarly to Example 2, the DEL compound "BIO-DEL" having the sequence shown in Table 22 was synthesized by the following procedure. In the sequence notation in Table 22, "(BIO)" represents the following formula (11): [ka] and other symbols are the same as in Table 20. [Table 22]

[0269] In a PCR tube, 20 μL of 1 mM aqueous solution of AOP-U-DEL9-HP (synthesized in Example 6); 24 μL of 1 mM aqueous solution of Pr_TAG2 (prepared by annealing Pr_TAG2_a and Pr_TAG2_b synthesized in the same manner as in Example 1, the sequence is shown in Table 23); 8 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 10 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) and 20 μL of deionized water were added. 8 μL of a 10-fold diluted aqueous solution of T4 DNA ligase (manufactured by Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16 ° C. for 22 hours. The sequence notation in Table 23 is the same as in Table 1. The names of the compounds corresponding to each sequence number (No.) are as follows. No.114:Pr_TAG2_a No.115:Pr_TAG2_b [Table 23]

[0270] The reaction solution was treated with 8 μL of 5 M aqueous sodium chloride solution and 264 μL of chilled (−20° C.) ethanol, and left overnight at −78° C. After centrifugation, the supernatant was removed and the resulting pellet was air-dried. The pellet was dissolved in deionized water and purified by reversed-phase HPLC using a Phenomenex Gemini C18 column. The target substance was eluted using 50 mM triethylammonium acetate buffer (pH 7.5) and acetonitrile / 50 mM triethylammonium acetate buffer (9:1, v / v) using a binary mobile phase gradient profile. The fractions containing the target substance were collected, mixed, and concentrated. The resulting solution was desalted using an Amicon® Ultra Centrifugal filter (3 kD cutoff), and ethanol precipitation was performed, after which 25 μL of deionized water was added to the pellet to make it into a solution.

[0271] A portion of the resulting solution was sampled and diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under the analytical condition 2 of Example 1 to identify the target substance (the theoretical molecular weight of the compound and the detected molecular weight are shown in Table d). The remaining solution was freeze-dried, and then 100 mM aqueous sodium bicarbonate solution was added to each solution to adjust the concentration to 1 mM.

[0272] To 6.2 μL of the solution obtained above, 7.4 μL of 1 mM aqueous solution of CP-BIO (prepared by annealing CP_a and CP-BIO_b synthesized in the same manner as in Example 1, the sequence is shown in Table 24); 2.5 μL of 10X ligase buffer (500 mM Tris-HCl, pH 7.5; 500 mM sodium chloride; 10 mM magnesium chloride; 100 mM dithiothreitol; 20 mM adenosine triphosphate) and 6.2 μL of deionized water were added. 2.47 μL of a 10-fold diluted aqueous solution of T4 DNA ligase (manufactured by Thermo Fisher, catalog number EL0013) was added to the solution, and the resulting solution was incubated at 16° C. for 16 hours. The sequence notation in Table 24 is the same as in Table 23. The names of the compounds corresponding to each sequence number (No.) are as follows. No.51: CP_a No.116: CP-BIO_b

[0273]

Table 24

[0274] The reaction solution was treated with 2.5 μL of 5 M aqueous sodium chloride solution and 81.5 μL of cooled (-20 °C) ethanol, and allowed to stand at -78 °C for 30 minutes. After centrifugation, the supernatant was removed, the obtained pellet was air-dried, and the pellet was dissolved in deionized water. The obtained solution was desalted using an Amicon® Ultra Centrifugal Filter (3 kD cut-off).

[0275] A portion of the obtained supernatant was sampled, diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under the analysis conditions 3 of Example 3 to identify the target substance (the theoretical molecular weight and the detected molecular weight are shown in Table 22). The remainder of the solution was lyophilized, deionized water was added, and it was adjusted to 120 μM to obtain BIO-DEL.

[0276] <Cleavage of <BIO-DEL> by USER® enzyme The cleavage reaction of the above-obtained DEL compound "BIO-DEL" by USER® enzyme was carried out according to the following procedure to synthesize a DEL compound "DS-BIO-DEL" having a double-stranded nucleic acid of the sequence shown in Table 25. The sequence listing in Table 25 is the same as that in Table 22, which means that DS-BIO-DEL is formed by the double-stranded of the oligonucleotide chains of SEQ ID NO: 118 and SEQ ID NO: 119.

[0277]

Table 25

[0278] Three PCR tubes were each added with 10 μL of a 120 μM aqueous solution of the DEL compound "BIO-DEL", 100 μL of CutSmart® Buffer (New England BioLabs, Catalog No. 7240S) and 860 μL of deionized water. 30 μL of USER® enzyme (New England BioLabs, Catalog No. 5505S) was added to each solution, and the resulting solution was incubated at 37° C. for 24 hours.

[0279] The resulting reaction solutions were desalted using an Amicon (registered trademark) Ultra Centrifugal filter (3 kD cutoff), and deionized water was added to prepare a 60 μL solution. Then, each solution was treated with 6 μL of 5 M aqueous sodium chloride solution and 198 μL of chilled (−20° C.) ethanol, and allowed to stand at −78° C. for 30 minutes. After centrifugation, the supernatant was removed, and deionized water was added to the resulting pellet to prepare a solution, which was then combined into one tube.

[0280] A portion of the resulting solution was sampled and diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under analysis condition 3 of Example 3, to identify the DEL compound having the desired double-stranded nucleic acid, "DS-BIO-DEL" (the theoretical molecular weight of the compound and the detected molecular weight are shown in Table 25).

[0281] A portion of the reaction solution was sampled and analyzed by denaturing polyacrylamide gel electrophoresis under the same conditions as in Example 3. The results shown in Figure 20 confirmed that BIO-DEL was cleaved in high yield and converted to DS-BIO-DEL. The samples in each lane in Figure 20 are as follows: Lane 1: BIO-DEL (Concentration 1: Prepared to have approximately 40 ng of BIO-DEL) Lane 2: BIO-DEL (Concentration 2: Prepared so that BIO-DEL is approximately 80 ng) Lane 3: Sample after cleavage reaction using BIO-DEL's USER® enzyme (Concentration 1: Prepared so that the target product is approximately 40 ng) Lane 4: Sample after cleavage reaction using BIO-DEL's USER® enzyme (concentration 2: prepared so that the target product is approximately 80 ng) Lane 5: 20 bp DNA ladder (Lonza 20 bp DNA Ladder, Cat. No. 50330)

[0282] <Preparation of DEL containing single-stranded DNA using streptavidin beads> The DEL compound "DS-BIO-DEL" having double-stranded nucleic acid obtained above was treated with streptavidin beads to prepare a DEL compound "SS-DEL" having single-stranded DNA by the following procedure. SS-DEL is an oligonucleotide chain of SEQ ID NO: 119 in Table 25.

[0283] 450 μL of Magnosphere (trademark) MS160 / Streptavidin (JSR Life Sciences, catalog number J-MS-S160S) was added to each of two PCR tubes, and the supernatant was removed by magnetic separation. 900 μL of 1× binding buffer (10 mM Tris-HCl, pH 7.5; 0.5 mM ethylenediaminetetraacetic acid; 1 M sodium chloride; 0.05% v / v Tween20) was added and the supernatant was removed by magnetic separation. DS-BIO-DEL aqueous solution k (700 pmol, 450 μL) and 450 μL of 2× binding buffer (20 mM Tris-HCl, pH 7.5; 1 mM ethylenediaminetetraacetic acid; 2 M sodium chloride; 0.1% v / v Tween20) were added to the obtained particles, mixed, and shaken at room temperature for 20 minutes.

[0284] The supernatant was removed from the mixture by magnetic separation, and the washing of the particles with 900 μL of 1× binding buffer (10 mM Tris-HCl, pH 7.5; 0.5 mM ethylenediaminetetraacetic acid; 1 M sodium chloride; 0.05% v / v Tween 20) and the removal of the supernatant by magnetic separation were repeated three times each. Then, 900 μL of an aqueous solution (0.1 M sodium hydroxide; 0.1 M sodium chloride) was added, and the supernatant was collected by magnetic separation.

[0285] To the obtained supernatants, 900 μL of 3-(N-morpholino)propanesulfonic acid buffer (1.0 M, pH 7.0) was added, and the mixture was desalted using an Amicon® Ultra Centrifugal filter (3 kD cutoff). The obtained supernatants were combined into one tube, treated with 13.6 μL of 5 M aqueous sodium chloride solution and 448 μL of chilled (−20° C.) ethanol, and allowed to stand at −78° C. for 60 minutes. After centrifugation, the supernatant was removed, and the obtained pellet was air-dried. 60 μL of deionized water was added to the pellets to make a solution.

[0286] A portion of the resulting solution was sampled and diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under analytical condition 3 of Example 3. A molecular weight of 23,984.8 was observed, and the DEL compound "SS-DEL" having the desired single-stranded DNA was identified.

[0287] <Synthesis of photoreactive crosslinker modified primer> The photoreactive crosslinker-modified primer "PXL-Pr" having the sequence shown in Table 26 was synthesized by the following procedure. In the sequence notation in Table 26, "(X)" represents the following formula (12): [ka] The other symbols are the same as in Table 2. [Table 26]

[0288] A solution (200 μL, 1 mM) of L-Pr (synthesized as in Example 1, sequence shown in Table 27) in sodium borate buffer (150 mM, pH 9.4) cooled to 10° C. was added to a PCR tube. 40 equivalents of N-Fmoc-15-amino-4,7,10,13-tetraoxaoctadecanoic acid (20 μL, 0.4 M N-dimethylacetamide solution) was added to the tube, followed by 40 equivalents of 4-(4,6-dimethoxy[1.3.5]triazin-2-yl)-4-methylmorpholinium chloride hydrate (DMTMM) (16 μL, 0.5 M aqueous solution), and the resulting mixture was shaken at 10° C. for 5 hours. The sequence notation in Table 27 is the same as in Table 8. [Table 27]

[0289] The reaction solution was treated with 23.6 μL of 5 M aqueous sodium chloride solution and 778.8 μL of chilled (−20° C.) ethanol, and left overnight at −78° C. After centrifugation, the supernatant was removed and the resulting pellet was air-dried. 180 μL of deionized water was added to the pellet to make a solution, after which 20 μL of piperidine was added and the mixture was shaken at 10° C. for 3 hours.

[0290] The resulting solution was treated with 20 μL of 5 M aqueous sodium chloride solution and 660 μL of chilled (−20° C.) ethanol, and allowed to stand at −78° C. for 30 minutes. After centrifugation, the supernatant was removed, and 200 μL of deionized water was added to the resulting pellet to make a 1 mM solution.

[0291] To 100 μL of the solution obtained above, 75 μL of triethylamine hydrochloride buffer (500 mM, pH 10) was added, followed by 50 equivalents of 1-((3-(3-methyl-3H-diazirin-3-yl)propanoyl)oxy)-2,5-dioxopyrrolidine-3-sodium sulfonate (Sulfo-SDA) (25 μL, 200 mM aqueous solution), and the mixture was shaken at 37° C. for 2 hours.

[0292] The resulting solution was treated with 20 μL of 5M aqueous sodium chloride solution and 660 μL of chilled (−20° C.) ethanol, and allowed to stand at −78° C. for 30 minutes. After centrifugation, the supernatant was removed, and 100 μL of deionized water was added to the resulting pellet, followed by 75 μL of triethylamine hydrochloride buffer (500 mM, pH 10) and 50 equivalents of Sulfo-SDA (25 μL, 200 mM aqueous solution), and the mixture was shaken at 37° C. for 1 hour and 20 minutes. Further, 50 equivalents of Sulfo-SDA (25 μL, 200 mM aqueous solution) was added, and the mixture was shaken at 37° C. for 40 minutes.

[0293] The resulting solution was treated with 22.5 μL of 5 M aqueous sodium chloride solution and 743 μL of chilled (−20° C.) ethanol, and allowed to stand overnight at −78° C. After centrifugation, the supernatant was removed, and 100 μL of deionized water was added to the resulting pellet, followed by 75 μL of triethylamine hydrochloride buffer (500 mM, pH 10), and then 50 equivalents of Sulfo-SDA (25 μL, 200 mM aqueous solution), and the mixture was shaken at 37° C. for 3 hours.

[0294] The resulting solution was treated with 20 μL of 5 M aqueous sodium chloride solution and 660 μL of chilled (−20° C.) ethanol and allowed to stand overnight at −78° C. After centrifugation, the supernatant was removed and the resulting pellet was air-dried. The pellet was dissolved in 50 mM triethylammonium acetate buffer (pH 7.5) and purified by reversed-phase HPLC using a Phenomenex Gemini C18 column. The target substance was eluted using 50 mM triethylammonium acetate buffer (pH 7.5) and acetonitrile / water (100:1, v / v) using a binary mobile phase gradient profile. The fractions containing the target substance were collected, mixed and concentrated. The resulting solution was desalted using an Amicon® Ultra Centrifugal filter (3 kD cutoff), and ethanol precipitation was performed, after which 100 μL of deionized water was added to the pellet to make it into a solution.

[0295] A portion of the obtained solution was sampled and diluted with deionized water, and then subjected to mass spectrometry by ESI-MS under the analytical condition 3 of Example 3 to identify the target photoreactive crosslinker-modified primer "PXL-Pr" (the theoretical molecular weight of the compound and the detected molecular weight are shown in Table 26).

[0296] <Synthesis of photoreactive crosslinker-modified double-stranded DEL> Using the SS-DEL and PXL-Pr obtained above, a primer extension reaction was carried out according to the following procedure to synthesize a photoreactive crosslinker-modified double-stranded DEL "PXL-DS-DEL" having the sequence shown in Table 28. The sequence notation in Table 28 is the same as in Table 26, and this means that PXL-DS-DEL is formed by a double strand of the oligonucleotide strands of SEQ ID NO:122 and SEQ ID NO:119. [Table 28]

[0297] In a PCR tube, 50 μL of 8 μM "SS-DEL" aqueous solution, 0.673 μL of 594 μM "PXL-Pr" aqueous solution, 80 μL of 10× NEBuffer®2 (New England BioLabs, Catalog No. B7002S) and 645 μL of deionized water were added. 8 μL of DNA Polymerase I, Large (Klenow) Fragment (New England BioLabs, Catalog No. M0210) and 16 μL of Deoxynucleotide (dNTP) Solution Mix (New England BioLabs, Catalog No. N0447) were added to the solution, and the resulting solution was incubated at 25° C. for 90 minutes.

[0298] The resulting solution was desalted using an Amicon® Ultra Centrifugal filter (3 kD cutoff). 17 μL of deionized water was added to the resulting supernatant, which was then treated with 6 μL of 5 M aqueous sodium chloride solution and 198 μL of chilled (−20° C.) ethanol and allowed to stand at −78° C. for 60 minutes. After centrifugation, the supernatant was removed and the resulting pellet was air-dried. 40 μL of deionized water was added to the pellet to make it into a solution.

[0299] A portion of the resulting solution was sampled and diluted with deionized water, and then subjected to mass analysis by ESI-MS under analytical condition 3 of Example 3, identifying the target photoreactive crosslinker-modified double-stranded DEL "PXL-DS-DEL" (the theoretical molecular weight of the compound and the detected molecular weight are shown in Table 28).

[0300] A portion of the reaction solution was sampled and analyzed by polyacrylamide gel electrophoresis under the conditions shown below. The results shown in Figure 21 confirmed that PXL-DS-DEL was produced in high yield by the primer extension reaction. The samples in each lane in Figure 21 are as follows: Lane 1: 20 bp DNA ladder (Lonza 20 bp DNA Ladder, Cat. No. 50330) Lane 2: DS-BIO-DEL Lane 3: SS-DEL Lane 4: Sample after primer extension reaction of SS-DEL (PXL-DS-DEL)

[0301] Polyacrylamide gel electrophoresis: Gel: SuperSep (trademark) DNA 15% TBE gel (Fujifilm Wako Pure Chemical Industries, catalog number 190-15481) Loading buffer: 6× Loading Buffer (Takara Bio, catalog number 9156) Temperature: room temperature Voltage: 200V Running time: 50 min Staining reagent: SYBER (trademark) Green II Nucleic Acid Gel Stain (Takara Bio, catalog number 5770A) [Industrial Applicability]

[0302] The present invention can utilize a nucleic acid compound that contains a selectively cleavable site. Furthermore, the present invention provides a DNA-encoded library that contains a selectively cleavable site, a composition for synthesizing the same, and a method for using the same, which allows for the production of a more convenient DNA-encoded library than ever before.

Claims

1. Formula (I) 【Chemical 1】 (wherein E and F are each independently an oligomer composed of nucleotides or nucleic acid analogs,[ provided that E and F contain base sequences complementary to each other and form a double-stranded oligonucleotide,[ LP is 【Chemical 2】 a loop site represented by,[ LS is a partial structure selected from the group of compounds described in the following (A) to (C),[ (A) nucleotide (B) nucleic acid analog (C) a trivalent group having 1 to 14 carbon atoms which may have a substituent (LP1) p, where LP1 is each partial structure selected alone or differently in p numbers from the group of compounds described in the following (1) and (2),[ (1) nucleotide (2) nucleic acid analog (LP2) q, where LP2 is each partial structure selected alone or differently in q numbers from the group of compounds described in the following (1) and (2),[ (1) nucleotide (2) nucleic acid analog the total number of p and q is 0 to 40,[ L is a linker,[ D is a reactive functional group.) a compound represented by,[ a compound having at least one selectively cleavable site at at least one of the sites of E, F, or LP.[

2. The compound according to claim 1, having at least one selectively cleavable site at at least one of the sites of E or F.[

3. The compound according to claim 1, wherein the 5'-end of E is bound to LP and E has a cleavable site.[

4. A composition for use in the preparation of the headpiece of a compound library, comprising the compound according to any one of claims 1 to 3.[

5. A composition for use in the preparation of the headpiece of a DNA-encoded library, comprising the compound according to any one of claims 1 to 3.[

6. The compound according to claim 1, for use as the headpiece of a compound library.[

7. The compound according to claim 1, for use as the headpiece of a DNA-encoded library.[

8. A headpiece of a compound library, having the structure of the compound according to claim 1.[

9. A headpiece of a DNA-encoded library, having the structure of the compound according to claim 1.[

10. The compound according to any one of claims 1 to 3, 6, and 7, wherein the total number of p and q is 2 to 20.[

11. The compound according to any one of claims 1 to 3, 6, and 7, wherein the total number of p and q is 2 to 10.[

12. The compound according to any one of claims 1 to 3, 6, and 7, wherein the total number of p and q is 2 to 7.

13. The compound according to any one of claims 1 to 3, 6, and 7, wherein the total number of p and q is 0.

14. LP1, LP2, and LS each have the following structures: (A) nucleotide or (B) a nucleic acid analog satisfying the following (B11) to (B15) (B11) having a phosphate (or corresponding moiety) and a hydroxyl group (or corresponding moiety), (B12) composed of carbon, hydrogen, oxygen, nitrogen, phosphorus, or sulfur, (B13) having a molecular weight of 142 to 1500, (B14) having 3 to 30 atoms between residues, (B15) the bonding pattern of atoms between residues is all single bonds or includes 1 to 2 double bonds and the rest are single bonds, The compound according to any one of claims 1 to 3, 6, 7, and 10 to 13, which is a structure selected singly or differently.

15. LP1, LP2, and LS each have the following structures: (A) nucleotide or (B) a nucleic acid analog satisfying the following (B21) to (B25) (B21) having a phosphate and a hydroxyl group, (B22) composed of carbon, hydrogen, oxygen, nitrogen, or phosphorus, (B23) having a molecular weight of 142 to 1000, (B24) having 3 to 15 atoms between residues, (B25) the bonding pattern of atoms between residues is all single bonds, The compound according to any one of claims 1 to 3, 6, 7, and 10 to 14, which is a structure selected singly or differently.

16. LP1, LP2, and LS each have the following structures: (A) nucleotide or (B) a nucleic acid analog satisfying the following (B31) to (B35) (B31) having a phosphate and a hydroxyl group, (B32) composed of carbon, hydrogen, oxygen, nitrogen, or phosphorus, (B33) having a molecular weight of 142 to 700, (B34) having 4 to 7 atoms between residues, (B35) the bonding pattern of atoms between residues is all single bonds, The compound according to any one of claims 1 to 3, 6, 7, and 10 to 15, which is a structure selected singly or differently.

17. LP1 and LP2 are each one of the following: (B41) d-Spacer, (B5) polyalkylene glycol phosphate ester The compound according to any one of claims 1 to 3, 6, 7, and 10 to 16.

18. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 17, wherein LP1 and LP2 are diethylene glycol phosphate or triethylene glycol phosphate, respectively.

19. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 18, wherein LP1 and LP2 are triethylene glycol phosphate, respectively.

20. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 17, wherein LP1 and LP2 are d-Spacer, respectively.

21. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 16, wherein LP1 and LP2 are nucleotides, respectively.

22. LS is of formula (a) to formula (g): 【Chemical Formula 4】 (In the formula, * means the bonding position with the linker, ** means the bonding position with LP1 or LP2, and R is a hydrogen atom or a methyl group.) The compound according to any one of claims 1 to 3, 6, 7 and 10 to 21, which is any of the above.

23. LS is of formula (h): [Chemical Formula 5] (In the formula, * means the bonding position with the linker, ** means the bonding position with LP1 or LP2) The compound according to any one of claims 1 to 3, 6, 7 and 10 to 21, which is any of the above.

24. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 21, wherein LS is a polyalkylene glycol phosphate.

25. LS is of formula (i) to formula (k): 【Chemical Formula 6】 (In the formula, n1, m1, p1, and q1 are each independently an integer from 1 to 20, * means the bonding position with the linker, and ** means the bonding position with LP1 or LP2.) The compound according to any one of claims 1 to 3, 6, 7 and 10 to 21, which is any of the above.

26. LS is of formula (l): [Chemical Formula 7] (In the formula, * means the bonding position with the linker, ** means the bonding position with LP1 or LP2) The compound according to any one of claims 1 to 3, 6, 7 and 10 to 21, which is any of the above.

27. LS is (B42), (B43) or (B44): (B42) Amino C6 dT (B43) mdC(TEG-Amino) (B44) Uni-Link (trademark registered) Amino Modifier The compound according to any one of claims 1 to 3, 6, 7 and 10 to 21, which is any of the above.

28. LS is a nucleotide. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 21.

29. LS is a trivalent group of C1-14 which may have a (C) substituent, and (C) has the following structure: (1) a C1-10 aliphatic hydrocarbon which may have a substituent and may be replaced by 1 to 3 heteroatoms; (2) a C6-14 aromatic hydrocarbon which may have a substituent; (3) a C2-9 aromatic heterocyclic ring which may have a substituent, or (4) a C2-9 non-aromatic heterocyclic ring which may have a substituent The compound according to any one of claims 1 to 3, 6, 7, 10 to 13 and 17 to 21, which is any one of the above.

30. LS is a trivalent group of C1-14 which may have a (C) substituent, and (C) has the following structure: (1) a C1-6 aliphatic hydrocarbon which may have a substituent; (2) a C6-10 aromatic hydrocarbon which may have a substituent, or (3) a C2-5 aromatic heterocyclic ring which may have a substituent The compound according to any one of claims 1 to 3, 6, 7, 10 to 13 and 17 to 21, which is any one of the above.

31. LS is a trivalent group of C1-14 which may have a (C) substituent, and (C) has the following structure: (1) a C1-6 aliphatic hydrocarbon; (2) benzene, or (3) a C2-5 nitrogen-containing aromatic heterocyclic ring Here, the above (1) to (3) may be unsubstituted or may be substituted by 1 to 3 substituents selected singly or differently from the substituent group ST1. The substituent group ST1 is a group composed of a C1-6 alkyl group, a C1-6 alkoxy group, a fluorine atom and a chlorine atom. However, when the substituent group ST1 substitutes an aliphatic hydrocarbon, an alkyl group is not selected from the substituent group ST1. The compound according to any one of claims 1 to 3, 6, 7, 10 to 13 and 17 to 21, which is any one of the above.

32. LS is a trivalent group of C1-14 which may have a (C) substituent, and (C) has the following structure: (1) a C1-6 alkyl group, or (2) benzene which is unsubstituted or substituted by one or two C1-3 alkyl groups or C1-3 alkoxy groups The compound according to any one of claims 1 to 3, 6, 7, 10 to 13 and 17 to 21, which is any one of the above.

33. LS is a trivalent group of C1-14 which may have a (C) substituent, and (C) has the following structure: (1) a C1-6 alkyl group The compound according to any one of claims 1 to 3, 6, 7, 10 to 13 and 17 to 21, which is any one of the above.

34. E and F are each independently an oligomer composed of nucleotides or nucleic acid analogs, and the chain lengths of E and F are each 3 to 40, The compound according to any one of claims 1 to 3, 6, 7 and 10 to 33.

35. E and F are each independently an oligomer composed of nucleotides or nucleic acid analogs, and the chain lengths of E and F are each 4 to 30 The compound according to any one of claims 1 to 3, 6, 7 and 10 to 34.

36. E and F are each independently an oligomer composed of nucleotides or nucleic acid analogs, and the chain lengths of E and F are each 6 to 25 The compound according to any one of claims 1 to 3, 6, 7 and 10 to 35.

37. E and F are each independently an oligomer composed of nucleotides or nucleic acid analogs, E and F contain base sequences complementary to each other to form a double-stranded oligonucleotide, and the double-stranded oligonucleotide of E and F has a protruding end, The compound according to any one of claims 1 to 3, 6, 7 and 10 to 36.

38. The compound according to claim 37, wherein the protrusion of the protruding end has a length of 2 bases or more.

39. E and F are each independently an oligomer composed of nucleotides or nucleic acid analogs, E and F contain base sequences complementary to each other to form a double-stranded oligonucleotide, and the double-stranded oligonucleotide of E and F has a blunt end, The compound according to any one of claims 1 to 3, 6, 7 and 10 to 36.

40. The chain lengths of the base sequences complementary to each other contained in E and F are each 3 bases or more The compound according to any one of claims 1 to 3, 6, 7 and 10 to 39.

41. The chain lengths of the base sequences complementary to each other contained in E and F are each 4 bases or more The compound according to any one of claims 1 to 3, 6, 7 and 10 to 40.

42. The chain lengths of the base sequences complementary to each other contained in E and F are each 6 bases or more The compound according to any one of claims 1 to 3, 6, 7 and 10 to 41.

43. E and F are each independently an oligomer composed of nucleotides, The compound according to any one of claims 1 to 3, 6, 7 and 10 to 42.

44. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 43, wherein the nucleotide is a ribonucleotide or a deoxyribonucleotide.

45. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 44, wherein the nucleotide is a deoxyribonucleotide.

46. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 45, wherein the nucleotide is deoxyadenosine, deoxyguanosine, thymidine, or deoxycytidine.

47. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 42, wherein E and F are each independently an oligomer composed of nucleic acid analogs.

48. L is (1) a C1-C20 aliphatic hydrocarbon which may have a substituent and may be replaced by 1 to 3 heteroatoms, or (2) a C6-C14 aromatic hydrocarbon which may have a substituent The compound according to any one of claims 1 to 3, 6, 7 and 10 to 47.

49. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 48, wherein L is a C1-C6 aliphatic hydrocarbon which may have a substituent, a C1-C6 aliphatic hydrocarbon which may be replaced by one or two oxygen atoms, or a C6-C10 aromatic hydrocarbon which may have a substituent.

50. L is a C1-C6 aliphatic hydrocarbon which can be substituted by a substituent group ST1, or benzene which can be substituted by a substituent group ST1, where the substituent group ST1 is a group composed of a C1-C6 alkyl group, a C1-C6 alkoxy group, a fluorine atom and a chlorine atom (however, when the substituent group ST1 substitutes for an aliphatic hydrocarbon, the alkyl group is not selected from the substituent group ST1). The compound according to any one of claims 1 to 3, 6, 7 and 10 to 49.

51. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 50, wherein L is a C1-C6 alkyl group, or benzene which is unsubstituted or substituted by one or two C1-C3 alkyl groups or C1-C3 alkoxy groups.

52. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 51, wherein L is a C1-C6 alkyl group.

53. The reactive functional group of D is The compound according to any one of claims 1 to 3, 6, 7 and 10 to 52, which is a reactive functional group capable of forming a C—C, amino, ether, carbonyl, amide, ester, urea, sulfide, disulfide, sulfoxide, sulfonamide, or sulfonyl bond.

54. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 53, wherein the reactive functional group of D is a C1 hydrocarbon having a leaving group, an amino group, a hydroxyl group, a precursor of a carbonyl group, a thiol group, or an aldehyde group.

55. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 54, wherein the reactive functional group of D is a C1 hydrocarbon having a halogen atom, a C1 hydrocarbon having a sulfonic acid-based leaving group, an amino group, a hydroxyl group, a carboxy group, a halogenated carboxy group, a thiol group, or an aldehyde group.

56. The reactive functional group of D is -CH 2 Cl, -CH 2 Br, -CH 2 OSO 2 CH 3 , -CH 2 OSO 2 CF 3 , an amino group, a hydroxyl group, or a carboxy group, the compound according to any one of claims 1 to 3, 6, 7 and 10 to 55.

57. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 56, wherein the reactive functional group of D is a primary amino group.

58. The selectively cleavable site is a deoxyribonucleoside that is not any of deoxyadenosine, deoxyguanosine, thymidine, and deoxycytidine. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 57.

59. The selectively cleavable site is deoxyuridine, bromodeoxyuridine, deoxyinosine, 8-hydroxydeoxyguanosine, 3-methyl-2'-deoxyadenosine, N6-etheno-2'-deoxyadenosine, 7-methyl-2'-deoxyguanosine, 2'-deoxyxanthosine, or 5,6-dihydroxy-5,6-dihydrodeoxythymidine. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 58.

60. The selectively cleavable site is deoxyuridine or deoxyinosine. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 59.

61. The selectively cleavable site is deoxyuridine. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 60.

62. The selectively cleavable site is deoxyinosine. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 60.

63. The selectively cleavable site is the second phosphodiester bond in the 3' direction from deoxyinosine. The compound according to any one of claims 1 to 3, 6, 7 and 10 to 57.

64. The selectively cleavable site is a ribonucleoside, The compound according to any one of claims 1 to 3, 6, 7 and 10 to 57.

65. There is one selectively cleavable site, The compound according to any one of claims 1 to 3, 6, 7 and 10 to 64.

66. Containing at least one cleavable site in E or (LP1)p and at least one cleavable site in F or (LP2)q, The compound according to any one of claims 1 to 3, 6, 7 and 10 to 64.

67. The cleavable site contained in E or (LP1)p and the cleavable site contained in F or (LP2)q are cleavable under different conditions, The compound according to claim 66.

68. A method of using a nucleic acid that binds to a compound having a cleavable site and a hairpin structure, and cleaving the cleavable site to be used as a double-stranded nucleic acid, The method, wherein the compound having a cleavable site and a hairpin structure is the compound according to any one of claims 1 to 3, 6, 7 and 10 to 67.

69. The method according to claim 68, wherein a nucleic acid that is chemically more stable than a double-stranded nucleic acid and binds to a compound having a cleavable site and a hairpin structure is used, and the cleavable site is cleaved to be used as a double-stranded nucleic acid.

70. The method according to claim 68 or 69, wherein a nucleic acid that binds to a compound having a cleavable site and a hairpin structure is used, the compound is subjected to a chemical structure conversion, and then the cleavable site is cleaved to be used as a double-stranded nucleic acid.

71. The method according to any one of claims 68 to 70, wherein a nucleic acid that binds to a compound having a cleavable site and a hairpin structure is used, the nucleic acid is further subjected to a chemical structure conversion, and then the cleavable site is cleaved to be used as a double-stranded nucleic acid.

72. The method according to any one of claims 68 to 71, wherein a nucleic acid that binds to a compound having a cleavable site and a hairpin structure is used, the nucleic acid is further subjected to a nucleic acid extension reaction, and then the cleavable site is cleaved to be used as a double-stranded nucleic acid.

73. The method according to any one of claims 68 to 72, wherein a nucleic acid that binds to a compound having a cleavable site and a hairpin structure is used, the cleavable site is cleaved to be usable as a double-stranded nucleic acid, and a PCR reaction is performed.

74. The method according to any one of claims 68 to 73, which is used for the functional evaluation of a compound.

75. The method according to any one of claims 68 to 74, which is used for the biological activity evaluation of a compound.

76. The method according to any one of claims 68 to 75, which is used for DEL.

77. The method according to any one of claims 68 to 72, which is used for the production of DEL.

78. A method for cleaving a cleavable site of a DEL compound synthesized using a nucleic acid that binds to a compound having a cleavable site and a hairpin structure and converting it into a DEL having single-stranded DNA, wherein the compound having a cleavable site and a hairpin structure is the compound according to any one of claims 1 to 3, 6, 7, and 10 to 67.

79. A method for cleaving a cleavable site of a DEL compound synthesized using a nucleic acid that binds to a compound having a cleavable site and a hairpin structure, converting it into a DEL having single-stranded DNA, and forming a double strand with a cross-linker-modified DNA, wherein the compound having a cleavable site and a hairpin structure is the compound according to any one of claims 1 to 3, 6, 7, and 10 to 67.

80. A method for cleaving a cleavable site of a DEL compound synthesized using a nucleic acid that binds to a compound having a cleavable site and a hairpin structure, imparting a cross-linker-modified primer, extending the imparted primer, and synthesizing a cross-linker-modified double-stranded DEL compound, wherein the compound having a cleavable site and a hairpin structure is the compound according to any one of claims 1 to 3, 6, 7, and 10 to 67.