Modified nucleobases with uniform hydrogen-bonding interactions, homo- and hetero-basepair bias, and mismatch discrimination

JP2023161073A5Active Publication Date: 2025-11-25CARNEGIE MELLON UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023137843
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-06-08
Filing Date
2023-08-28
Publication Date
2025-11-25
Estimated Expiration
2039-06-07

AI Technical Summary

Technical Problem

Existing oligonucleotide molecules struggle with sequence-specific binding to DNA or RNA targets, particularly in intracellular and in vivo applications, due to issues with enzyme stability and cell permeability, and face challenges in selectively targeting RNA secondary structures like stem-loop configurations.

Method used

Development of nucleobase moieties with modified hydrogen-bonding interactions that enable uniform binding to nucleic acids, allowing for sequence-specific recognition and discrimination of mismatches, incorporated into nucleic acid backbones or analogs to form gene recognition reagents.

Benefits of technology

The modified nucleobases provide enhanced sequence discrimination and cell permeability, enabling selective targeting of RNA secondary structures and improving therapeutic and diagnostic applications by reducing nonspecific binding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2023161073000001
    Figure 2023161073000001
  • Figure 2023161073000002
    Figure 2023161073000002
  • Figure 2023161073000003
    Figure 2023161073000003
Patent Text Reader

Abstract

To provide nucleobases, polymer monomers comprising the nucleobases, nucleic acids and analogs thereof comprising the nucleobases, and methods of use thereof.SOLUTION: Described herein are divalent nucleobases that each bind two nucleic acid strands, matched or mismatched when incorporated into a nucleic acid backbone or nucleic acid analog backbone, e.g., in a γ-peptide nucleic acid (γPNA). Also provided are genetic recognition reagents comprising one or more of the divalent nucleobases and a nucleic acid backbone or nucleic acid analog backbone, such as a γPNA backbone. Uses of the divalent nucleobases and monomers and genetic recognition reagents containing the divalent nucleobases also are provided.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference of related applications

[0001] This application claims the benefits of U.S. Provisional Patent Application No. 62 / 763,299, filed on 8 June 2018, which in its entirety is incorporated herein by reference. [Background technology]

[0002] This specification describes nucleic acid bases, polymer monomers comprising said nucleic acid bases, and nucleic acids and analogs comprising said nucleic acid bases. This specification also describes methods of using nucleic acid bases, polymer monomers comprising said nucleic acid bases, and nucleic acids and analogs comprising said nucleic acid bases.

[0003] For most organisms, genetic information is encoded in double-stranded DNA in the form of Watson-Crick base pairing, where adenine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). The developmental program and physiological state are determined by which sets of this genetic information are decoded through transcription and translation. The development of molecules that can be tailor-designed to sequence-specifically bind to any portion of this genetic biomolecule (DNA or RNA), thereby enabling control of the flow of genetic information and the evaluation and manipulation of genome structure and function, is crucial for biological and biomedical research, including molecular tools for basic biological research, in efforts to elucidate the molecular basis of life. These efforts are also important for medical and therapeutic applications for the treatment and detection of genetic diseases.

[0004] Oligonucleotides are versatile molecular tools for basic biological research, as well as for use as molecular reagents for therapeutic and diagnostic applications, such as in the treatment and detection of genetic diseases. Generally, oligonucleotide molecules are short fragments (10–30 nucleotides long) of single-stranded DNA or RNA, or derivatives thereof. Oligonucleotide molecules may contain a sugar phosphodiester skeleton, or other skeletons, linked to the nucleic acid bases adenine (A), cytosine (C), guanine (G), and thymine (T) or uridine (U). Oligonucleotide molecules are designed to bind to DNA or RNA targets via Watson-Crick base pairing, where A pairs with T (or U) and C pairs with G. Oligonucleotide molecules have been used in a wide range of applications, including, for example, querying nucleic acid sequence information, manipulating RNA structure, and regulating gene expression. Many of the successes of these applications are due to the oligomers' ability to bind strictly sequence-specifically to DNA or RNA targets. Further requirements for intracellular and in vivo gene targeting include enzyme stability and cellular permeability. [Overview of the Initiative]

[0005] It comprises multiple nucleic acid base moieties attached to a nucleic acid skeleton or nucleic acid analog skeleton, wherein at least one nucleic acid base moiety: [ka] (In the formula, X1 is =O (= represents a double bond), =S, =Se, or CH3; X2 is H, CH3, CN, NC, N3, C(O)OH, or C(O)NH2; X3 is O or S; X4 is H, C(O)CH3, or C(O)OCH3; and Y is N or CH, and in (1), if X1 and X3 are O, X2 is neither H nor methyl) A gene recognition reagent is provided.

[0006] structure: [ka] (wherein X1 is =O, =S, =Se, or CH3; X2 is H, CH3, CN, NC, N3, C(O)OH, or C(O)NH2; X3 is O or S; X4 is H, C(O)CH3, or C(O)OCH3; and Y is N or CH, and in (I), if X1 and X3 are O, then X2 is neither H nor methyl) Compounds are also provided that have a nucleic acid skeleton monomer or nucleic acid analog skeleton monomer linked to the nucleic acid base portion. [Brief explanation of the drawing]

[0007] [Figure 1] Figure 1 shows the hydrogen bonding interactions between (A) modified nucleic acid bases and modified nucleic acid bases (homodouble helix), (B) natural type and natural type, and (C) modified type and natural type (heterodouble helix). [Figure 2] Figure 2 is a schematic diagram illustrating the clear advantages of providing a newly designed oligonucleotide molecule that can selectively target RNA secondary structures. (A) Despite (a' / a) sequence complementarity, PNA oligomers (chiral or non-chiral) containing modified (u, c, a, and g) nucleic acid bases cannot adopt a hairpin structure and can hybridize to their complementary stem-loop RNA target. (B) While strongly binding oligonucleotide molecules such as LNA or yPNA can penetrate stem-loop structures, their application in therapeutic and diagnostic agents poses considerable risk due to nonspecific binding. (C) Medium avidity oligonucleotides, a category to which most oligonucleotide molecules belong, cannot open stem-loop structures due to a lack of binding free energy or as a result of dynamic (hairpin) trap formation. [Figure 3] Figure 3 schematically shows examples of (A) secondary and (B) tertiary structures of RNA in which the oligonucleotide molecules described herein enable selective targeting that would otherwise be difficult to achieve with existing nucleic acid systems. [Figure 4] Figure 4 shows the structure of an exemplary nucleic acid base. [Figure 5] Figure 5 shows exemplary nucleic acid analog residues of nucleic acid analogs, including phosphorothioate DNA (PS DNA), α,β-restricted nucleic acid (α,β-CNA), 2'-methoxyl RNA, 2'-fluoroRNA, phosphorodiamidate morpholino oligomer (PMO), locked nucleic acid (LNA), 2',4'-restricted ethyl nucleic acid ((S)-cEt), 2',4'-crosslinked nucleic acid NC(NH) (BNA-NC(NH)), 2',4'-crosslinked nucleic acid NC(N-methyl) (BNA-NC(N-Me)), ((S)-5'-C-methylDNA(RNA)), and 5'-E-vinylphosphonic acid nucleic acid (E-VP) (wherein R is H, OH, F, OMe, or O(CH2)2OMe). [Figure 6] Figure 6 shows an exemplary synthesis scheme for the selected nucleic acid bases. [Figure 7] Figure 7 shows an exemplary synthesis scheme for selected cell-permeable γPNA monomers. Detailed description of the invention

[0008] The use of numerical values ​​within the various ranges specified herein is described as an approximation, as if the word “approximately” were placed before both the minimum and maximum values ​​within the described range, unless otherwise explicitly stated. In this way, even with slight variations above and below the described range, substantially the same results as values ​​within the range can be achieved. Furthermore, unless otherwise specified, the disclosure of a range is intended as a continuous range encompassing all values ​​between the minimum and maximum values. As used herein, “one (a)” and “one (an)” refer to one or more.

[0009] As used herein, the term “comprising” is open-ended and may be synonymous with “including,” “containing,” or “characterized by.” As used herein, embodiments “comprising” one or more of the described elements or processes also include, but are not limited to, embodiments “essentially consisting of” and “spanning” those described elements or processes.

[0010] The term "polymer composition" refers to a composition comprising one or more polymers. The class "polymer" includes, but is not limited to, homopolymers, heteropolymers, copolymers, block polymers, and block copolymers, and may be natural or synthetic. A homopolymer contains one type of constituent block or monomer, while a copolymer contains two or more types of monomers. An "oligomer" is a polymer comprising a small number of monomers, e.g., 3 to 100 monomer residues. Therefore, the term "polymer" includes oligomers. The terms "nucleic acid" and "nucleic acid analog" include nucleic acids, nucleic acid polymers, and oligomers.

[0011] When a polymer incorporates a monomer listed in the description, the polymer "contains" or "is derived from" that monomer. Therefore, the incorporated monomers in the polymer are not the same as the monomers before they were incorporated into the polymer, at least in that a specific crosslinking group is incorporated into the polymer backbone or a specific group is removed during the polymerization process. If a specific type of bond exists within the polymer, the polymer is said to contain that bond. The incorporated monomers are "residues." Typical monomers of nucleic acids or nucleic acid analogs, when incorporated into a polymer, are called nucleotides or nucleotide residues.

[0012] A “part” is a part of a chemical compound and includes groups such as functional groups. Therefore, the nucleic acid base part is a nucleic acid base modified by attachment to another compound part, such as a polymer monomer, for example, a nucleic acid or nucleic acid analog monomer as described herein, or a polymer such as a nucleic acid or nucleic acid analog as described herein.

[0013] "Alkyl" refers to a linear, branched, or cyclic hydrocarbon group containing 1 to approximately 20 carbon atoms, for example, but is not limited to C 1-3 , C 1-6 , C 1-10Groups, for example, but not limited to, refer to linear, methyl, ethyl, propyl, butyl, pentyl, hexyl, heptyl, octyl, nonyl, decyl, undecyl, dodecyl and other branched alkyl groups. The alkyl group is, for example, a substituted or unsubstituted C1, C2, C3, C4, C5, C6, C7, C8, C9, C 10 C 11 C 12 C 13 C 14 C 15 C 16 C 17 C 18 C 19 C 20 C 21 C 22 C 23 C 24 C 25 C 26 C 27 C 28 C 29 C 30 C 31 C 32 C 33 C<> 34 C 35 C 36 C 37 C 38 C 39 C 40 C 41 C 42 C 43 C 44 C 45 C 46 C 47 C 48 C 49 or C 50It may be a group. Not limited examples of linear alkyl groups include methyl, ethyl, propyl, butyl, pentyl, hexyl, heptyl, octyl, nonyl, and decyl. Branched alkyl groups include any linear alkyl group substituted with any number of alkyl groups. Not limited examples of branched alkyl groups include isopropyl, isobutyl, sec-butyl, and t-butyl. Not limited examples of cyclic alkyl groups include cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, and cyclooctyl groups. Cycloalkyl groups also include condensed bicyclic, cross-linked bicyclic, and spiro-bicyclic, as well as higher-order condensed, cross-linked, and spiro systems. Cycloalkyl groups can be substituted with any number of linear, branched, or cyclic alkyl groups.

[0014] "Substitutive alkyl" refers to an alkyl group substituted in one or more locations (e.g., 1, 2, 3, 4, 5, or 6 locations) by substitutions as described herein, where the substituents are attached to any available atoms to produce a stable compound. "Optionally substituted alkyl" refers to an alkyl or a substituted alkyl group. "Halogen," "halide," and "halo" refer to -F, -CI, -Br, and / or -I. "Alkylene" and "substituted alkylene" refer to divalent alkyl and divalent substituted alkyl groups, respectively, and include, but are not limited to, ethylene (-CH2-CH2-). "Optionally substituted alkylene" refers to an alkylene or a substituted alkylene.

[0015] "Alkenes or alkenyls" are linear, branched, or cyclic hydrocarbyl groups containing 2 to approximately 20 carbon atoms, for example, but not limited to, having one or more (e.g., 1, 2, 3, 4, or 5) carbon-carbon double bonds. 2-3 , C 2-6 , C 2-10Refers to a group. One or more olefins of an alkenyl group can be, for example, E, Z, cis, trans, terminal, or exomethylene. The alkenyl or alkenylene group can be, for example, substituted or unsubstituted C2, C3, C4, C5, C6, C7, C8, C9, C 10 , C 11 , C 12 , C 13 , C 14 , C 15 , C 16 , C 17 , C 18 , C 19 , C 20 , C 21 , C 22 , C 23 , C 24 , C 25 , C 26 , C 27 , C 28 , C 29 , C 30 , C 31 , C 32 , C 33 , C 34 , C 35 , C 36 , C 37 , C 38 , C 39 , C 40 , C 41 , C 42 , C 43 , C 44 , C 45 , C 46 , C 47 , C 48 , C 49 , or C 50 It can be a group. The haloalkenyl group may be any alkenyl group substituted with any number of halogen atoms.

[0016] "Substituted alkene" refers to an alkene substituted at one or more positions (e.g., 1, 2, 3, 4, or 5 positions) with substitutions as described herein, and those substituents are attached with any available atoms such that a stable compound results. "Optionally substituted alkene" refers to an alkene or a substituted alkene. Similarly, "alkenylene" refers to a divalent alkene. Examples of alkenylene include, but are not limited to, ethenylene (-CH=CH-) and all of its stereoisomeric and conformational forms. "Substituted alkenylene" refers to a divalent substituted alkene. "Optionally substituted alkenylene" refers to an alkenylene or a substituted alkenylene.

[0017] Alkyne or "alkynyl" refers to a straight-chain, branched-chain, or cyclic unsaturated hydrocarbon having the indicated number of carbon atoms and at least one triple bond. The triple bond of an alkyne or alkynyl group may be internal or terminal. Examples of (C2-C8) alkynyl groups include, but are not limited to, acetylene, propyne, 1-butyne, 2-butyne, 1-pentyne, 2-pentyne, 1-hexyne, 2-hexyne, 3-hexyne, 1-heptyne, 2-heptyne, 3-heptyne, 1-octyne, 2-octyne, 3-octyne, and 4-octyne. An alkynyl group may be unsubstituted or optionally substituted with one or more substituents as described hereinafter. An alkyne or alkynyl group may be, for example, a substituted or unsubstituted C2, C3, C4, C5, C6, C7, C8, C9, C 10 、C 11 、C 12 、C 13 、C 14 、C 15 、C 16 、C 17 、C 18 、C 19 、C 20 、C 21 、C 22 、C 23 、C 24 、C 25 、C 26 、C 27 、C 28, C 29 , C 30 , C 31 , C 32 , C 33 , C 34 , C 35 , C 36 , C 37 , C 38 , C 39 , C 40 , C 41 , C 42 , C 43 , C 44 , C 45 , C 46 , C 47 , C 48 , C 49 , or C 50 It can be a group. The haloalkynyl group may be any alkenyl group substituted with any number of halogen atoms.

[0018] The term "alkynylene" refers to a divalent alkyne. Examples of alkynylenes, though not limited to them, include ethynylene and propynylene. "Substitutive alkynylene" refers to a divalent substituted alkyne.

[0019] The term "alkoxy" refers to an -O-alkyl group having the indicated number of carbon atoms. Ethers or ether groups contain alkoxy groups. For example, (C1-C6) alkoxy groups include -O-methyl (methoxy), -O-ethyl (ethoxy), -O-propyl (propoxy), -O-isopropyl (isopropoxy), -O-butyl (butoxy), -O-sec-butyl (sec-butoxy), -O-tert-butyl (tert-butoxy), -O-pentyl (pentoxy), -O-isopentyl (isopentoxy), -O-neopentyl (neopentoxy), -O-hexyl (hexyloxy), -O-isohexyl (isohexyloxy), and -O-neohexyl (neohexyloxy). "Hydroxyalkyl" refers to an alkyl group in which one or more hydrogen atoms of the alkyl group are substituted with an -OH group (C1-C 10) refers to alkyl groups. Examples of hydroxyalkyl groups, though not limited to, include -CH2OH, -CH2CH2OH, -CH2CH2CH2OH, -CH2CH2CH2CH2OH, -CH2CH2CH2CH2CH2OH, -CH2CH2CH2CH2CH2CH2OH, and their branched forms. The term "ether" or "oxygen ether" refers to alkyl groups in which one or more carbon atoms of the alkyl group are substituted with an -O- group (C1-C 10 ) refers to alkyl groups. The term ether includes -CH2-(OCH2-CH2) q The compound contains OP1, where P1 is a protecting group, -H, or (C1-C 10 ) is alkyl. Exemplary ethers include polyethylene glycol, diethyl ether, and methylhexyl ether.

[0020] The term "thioether" refers to a group in which one or more carbon atoms of an alkyl group are replaced by an -S- group (C1-C 10 ) refers to alkyl groups. The term thioether includes -CH2-(SCH2-CH2) q -SP1 compound is included, wherein P1 is a protecting group, -H, or (C1-C 10 ) is alkyl. Exemplary thioethers include dimethyl thioether or ethyl methyl thioether.

[0021] Protecting groups (for example, to protect amines during the synthesis of the compounds herein) are known in the art and are not limited to, but include 9-fluorenylmethyloxycarbonyl (Fmoc), t-butyloxycarbonyl (Boc), benzhydryloxycarbonyl (Bhoc), benzyloxycarbonyl (Cbz), O-nitroveratryloxycarbonyl (Nvoc), benzyl (Bn), allyloxycarbonyl (alloc), trityl (Trt), dimethoxytrityl (DMT), l-(4,4-dimethyl-2,6-dioxacyclohexylidene)ethyl (Dde), diatiasuccinoyl (Dts), benzothiazole-2-sulfonyl (Bts), and monomethoxytrityl (MMT) groups.

[0022] "Aryl" refers to an aromatic monocyclic or bicyclic ring system, such as phenyl or naphthyl, either alone or in combination. "Aryl" also includes aromatic ring systems that may optionally be fused with a cycloalkyl ring. "Substitutive aryl" is an aryl independently substituted with one or more substituents attached to any available atom to produce a stable compound, the substituents as described herein. Substituents may be, for example, hydrocarbyl groups, alkyl groups, alkoxy groups, and halogen atoms. "Optionally substituted aryl" refers to an aryl or substituted aryl. An aryloxy group may be, for example, one in which the oxygen atom is substituted with any aryl group, e.g., phenoxy. An arylalkoxy group may be, for example, one in which the oxygen atom is substituted with an aralkyl group, e.g., benzyloxy.

[0023] "Arylene" refers to a divalent aryl compound, and "substituted arylene" refers to a divalent substituted aryl compound. "Arylene that may be substituted in some cases" refers to arylene or substituted arylene.

[0024] "Heteroatom" refers to N, O, P, and S. Compounds containing an N or S atom may optionally be oxidized to the corresponding N-oxide, sulfoxide, or sulfone compound. "Heterosubstituted" refers to any organic compound of any embodiment described herein in which one or more carbon atoms are substituted with N, O, P, or S.

[0025] "Cycloalkyl" refers to a monocyclic, bicyclic, tricyclic, or polycyclic 3- to 14-membered ring system that is saturated, unsaturated, or aromatic. Cycloalkyls may be attached via any of the atoms. Cycloalkyls also refer to fused rings in which a cycloalkyl is fused with an aryl or heteroaryl ring. Representative examples of cycloalkyls, but not limited to, include cyclopropyl, cyclobutyl, cyclopentyl, and cyclohexyl. Cycloalkyls may be unsubstituted or optionally substituted with one or more substituents as described below herein. "Cycloalkylene" refers to a divalent cycloalkyl. The term "optionally substituted cycloalkylene" refers to a cycloalkylene substituted with one, two, or three substituents attached via any available atoms to produce a stable compound.

[0026] "Carboxyl" or "carboxylic acid" refers to a group having the indicated number of carbon atoms and a -C(O)OH group at its terminus, thus having the structure -RC(O)OH (where R is a divalent organic group including a linear, branched, or cyclic hydrocarbon). Examples of these, not limited to C, include 1-8 Examples of carboxylic acid groups include ethane acid group, propanoic acid group, 2-methylpropanoic acid group, butanoic acid group, 2,2-dimethylpropanoic acid group, and pentanoic acid group.

[0027] "(C3-C8)aryl-(C1-C6)alkylene" refers to the C1-C6 alkylene group This refers to a divalent alkylene in which one or more hydrogen atoms are replaced by a (C3-C8) aryl group. Examples of (C3-C8)aryl-(C1-C6)alkylene groups include, but are not limited to, 1-phenylbutylene, phenyl-2-butylene, l-phenyl-2-methylpropylene, phenylmethylene, phenylpropylene, and naphthylethylene. The term "(C3-C8)cycloalkyl-(C1-C6)alkylene" refers to a divalent alkylene in which one or more hydrogen atoms of a C1-C6 alkylene group are substituted with a (C3-C8)cycloalkyl group. Examples of (C3-C8)cycloalkyl-(C1-C6)alkylene groups include, but are not limited to, 1-cyclopropylbutylene, cyclopropyl-2-butylene, cyclopentyl-1-phenyl-2-methylpropylene, cyclobutylmethylene, and cyclohexylpropylene.

[0028] In larger molecules such as nucleic acids or nucleic acid polymer chains as described herein, a portion such as a nucleic acid base moiety, functional group, guanidine-containing group, or PEG-containing group is “linked” to the rest of the molecule, meaning that it is covalently attached to the rest of the molecule either directly or via an inert crosslinking moiety. “Inert” means that the linker does not substantially affect the function of the nucleic acid or nucleic acid analog for its intended use. Examples of inert linkers, though not limited to these, include linear, branched, and / or cyclic hydrocarbyl moieties or substituted hydrocarbyl moieties having 15, 10, or fewer than 6 carbon atoms, for example, alkylene moieties, for example, methylene, ethylene, propylene, or butylene groups, or aryl groups, and heteroatoms such as S (e.g., thioether), N (tertiary amine), or O (e.g., ether) atoms, and / or crosslinking bonds, for example, but not limited to ester bonds, amide bonds, carbamic acid bonds, or carbonate bonds, which may optionally be present. A linker can act as a spacer to physically or spatially separate two components of a molecule.

[0029] This specification provides nucleic acids and their analogues, collectively known as "gene recognition reagents," which specifically bind to nucleic acid chains under physiological conditions, for example, at 37°C, in physiological saline (0.9 wt% NaCl), or in another isotonic solution such as Tris-buffered saline or PBS. Each gene recognition reagent comprises a plurality of nucleic acid base moieties, each attached to a nucleic acid backbone or nucleic acid analog backbone monomer residue, forming a nucleoside or, in the case of DNA or RNA, a nucleotide with a phosphate group, and forming part of a larger gene recognition reagent comprising at least two nucleic acid or nucleic acid monomer residues, and thus at least two nucleic acid base moieties.

[0030] In one embodiment, all nucleic acid bases of the gene recognition reagent are modified nucleic acid bases as described herein. In another embodiment, the gene recognition reagent described herein comprises at least one modified nucleic acid base as described herein, the other nucleic acid bases being natural nucleic acid bases (e.g., adenine, guanine, cytosine, thymine, or uracil) or different from those modified bases.

[0031] Therefore, in one embodiment, modified nucleic acid bases are provided. These nucleic acid bases can be incorporated into nucleic acids or nucleic acid analog monomers (e.g., nucleotides), and then incorporated into oligomers or polymers of the monomer having a desired nucleic acid base sequence. The structures of the modified nucleic acid bases are shown below: [ka] (In the formula, X1 is =O (= represents a double bond), =S, =Se, or CH3; X2 is H, CH3, CN, NC, N3, C(O)OH, or C(O)NH2; X3 is either O or S; X4 is H, C(O)CH3, or C(O)OCH3; and Y is either N or CH. (1) If X1 and X3 are O, then X2 is neither H nor methyl. This includes: Examples of such nucleic acid bases that form a complete set of nucleic acid bases capable of binding to the A, T, G, and C of natural DNA or RNA include: [ka] It includes.

[0032] Further examples of such nucleic acid bases are the following fluorescent bases: [ka] That is the case.

[0033] Any, some, or all of the above nucleic acid bases can be incorporated into nucleotide monomers and gene recognition reagents.

[0034] In one embodiment, the gene recognition reagents described herein can self-assemble on a nucleic acid template comprising the target sequence of the gene recognition reagent. For example, the first gene recognition reagent can hybridize to a first portion of the nucleic acid template. The second self-assembling gene recognition reagent can hybridize to a second portion adjacent to the first portion of the nucleic acid template. The first and second gene recognition reagents can be covalently linked or non-covalently interconnected using various functional end groups to form a continuous structure. The gene recognition reagent may comprise a first portion linked by a linker to a first end of a nucleic acid backbone or nucleic acid analog backbone, and a second portion linked by a linker to a second end of the nucleic acid backbone or nucleic acid analog backbone. This second portion may be the same as the first portion, or it may be different from the first portion. For example, International Patent Application PCT / US18 / 67096, which is incorporated herein by reference, includes the following: TTC, TTCTTC, TCT, TCTTCT, CTT, CTTCTT, CCG, CCGCCG, CGC, CGCCGC, GCC, GCCGCC, CCG, CGGCGG, GCG, GCGGCG, GGC, GGCGGC, CTG, CTGCTG, TGC, TGCTGC, GCT, GCTGCT, CAG, CAGCAG, AGC, AGCAGC, GCA, GCAGCA, CAGG, CA Self-assembling gene recognition reagents are disclosed that contain aryl groups at their ends that, when aligned adjacent to target sequences such as extended repeats, GGCAGG, AGGC, AGGCAGGC, GGCA, GGCAGGCA, GCAG, GCAGGCAG, AGAAT, GAATA, AATAG, ATAGA, TAGAA, GGCCCC, GCCCCG, CCCCGG, CCCGGC, CCGGCC, and CGGCCC, form strong non-covalent bonds between unit gene recognition reagents via a pi-stack. Another example of a self-assembling gene recognition reagent is disclosed in International Patent Application Publication WO2014 / 169216, which is incorporated herein by reference in its entirety, describing a gene recognition reagent that self-assembles and covalently bonds in a reducing environment due to terminal thioester and sulfhydryl groups.Covalent bonds are formed between the unit self-assembling gene recognition reagents when they are adjacent to a target sequence, such as a nucleic acid containing the extended repeat described above. Other terminal groups, such as affinity partners or reactive groups, can also be attached to the ends of the gene recognition reagents described herein, allowing them to self-assemble on a complementary nucleic acid template.

[0035] In this specification, in one embodiment, a compound comprising a nucleic acid base and a nucleic acid backbone or nucleic acid analog backbone monomer is provided. For the purposes of this disclosure, “nucleotide” refers to a compound or residue of a gene recognition reagent comprising at least one nucleic acid base and a backbone element (in the case of nucleic acids such as RNA or DNA, this is ribose or deoxyribose). Nucleotide monomers also include reactive groups that enable polymerization under certain conditions. In natural DNA and RNA, these reactive groups are the 5'-phosphate group and the 3'-hydroxyl group. For the chemical synthesis of nucleic acids and their analogs, as is known in the art, the base and backbone monomers may contain modifying groups such as blocked amines. “Nucleotide residue” refers to a single nucleotide incorporated into an oligonucleotide or polynucleotide. Similarly, “nucleotide base residue” refers to a nucleic acid base incorporated into a nucleotide or nucleic acid or its analog. A “gene recognition reagent” generally refers to a nucleic acid or nucleic acid analog comprising a sequence of nucleic acid bases that can hybridize to a complementary nucleic acid sequence on a nucleic acid by cooperative base pairing (e.g., Watson-Crick base pairing or Watson-Crick-like base pairing) (see Figure 1). Intramolecular base pairing does not occur between the modified bases described herein, where steric collisions occur or there are insufficient hydrogen bonds between bases that normally form base pairs in natural nucleic acids (Figure 1 Panel (A) ("Figure 1(A)")). As can be seen in Figure 1(B), normal base pairing is asymmetric, with two hydrogen bonds formed between nucleic acid bases A and T / U, and three hydrogen bonds formed between nucleic acid bases G and C. Figure 1(C) illustrates the substantial benefit of using the modified nucleic acid bases described herein in that the bonding is symmetric, and two hydrogen bonds are formed between all combinations of the modified nucleic acid bases described herein and their natural base-pairing partners. The nearly equal binding "weight" distribution of base pairs not only greatly simplifies probe design, but also provides greater sequence discrimination against base pair mismatches than is achievable with natural nucleic acid bases, as the probe's binding strength is more strongly dependent on length than on length and sequence composition.

[0036] Figure 2 illustrates the benefits of the modified nucleic acid bases described herein. Figure 2(A) shows that PNA oligomers (chiral or non-chiral) containing modified (u, c, a, and g) nucleic acid bases, despite (a' / a) sequence complementarity, cannot adopt a hairpin structure but can hybridize to their complementary stem-loop RNA targets. Figure 2(B) shows that strongly binding oligonucleotide molecules, such as locked nucleic acids (LNA) or PNAs (e.g., yPNA), can penetrate stem-loop structures, but their application in therapeutic and diagnostic agents poses considerable risks due to nonspecific binding. Figure 2(C) shows that medium avidity oligonucleotides, a category to which most oligonucleotide molecules belong, cannot open stem-loop structures due to a lack of binding free energy or as a result of dynamic (hairpin) trap formation.

[0037] Examples of secondary and tertiary structures of selectively targetable RNA are shown in Figure 3. Potential therapeutic and diagnostic targets include both coding and non-coding RNAs, as well as DNA.

[0038] Modified nucleic acid bases are provided herein in several embodiments. A nucleic acid base is a recognition moiety that specifically binds to one or more of adenine, guanine, thymine, cytosine, and uracil, for example, by Watson-Crick or Watson-Crick-like base pairing via hydrogen bonding. "Nucleic acid bases" include basic (natural) nucleic acid bases: adenine, guanine, thymine, cytosine, and uracil, as well as modified purine and pyrimidine bases, for example, but not limited to, hypoxanthine, xanthene, 7-methylguanine, 5,6,dihydrouracil, 5-methylcytosine, and 5-hydroxymethylcytosine. Figure 4 also shows an unspecified number of nucleic acid bases, including monovalent nucleic acid bases (e.g., adenine, cytosine, guanine, thymine, or uracil that bind to a single strand of nucleic acid or nucleic acid analog), and "clamp" nucleic acid bases such as "G-clamps" that bind to complementary nucleic acid bases with enhanced strength. Further purines, purine-like, pyrimidine, and pyrimidine-like nucleic acid bases are also known in the art, for example, disclosed in U.S. Patents 8,053,212, 8,389,703, and 8,653,254. Divalent nucleic acid bases are described in more detail in U.S. Patent Application Publication 2016 / 0083434A1 and International Patent Application Publication WO / 2018 / 058091, all of which are incorporated herein by reference, and can bind with two nucleic acid bases rather than one, and thus form complex trimer structures with matched or mismatched nucleic acids.

[0039] In one example, the backbone monomer is ribose monophosphate, diphosphate, or triphosphate, or deoxyribose monophosphate, diphosphate, or triphosphate, for example, 5'-monophosphate, diphosphate, or triphosphate of ribose or deoxyribose. The backbone monomer includes both structural "residue" components such as ribose in RNA and active groups modified during monomer interconnection, such as the 5'-triphosphate and 3'-hydroxyl groups of ribonucleotides, which are modified upon polymerization into RNA, leaving a phosphodiester bond. Similarly, with respect to PNA, the C-terminal carboxyl and N-terminal amine active groups of the N-(2-aminoethyl)glycine backbone monomer condense upon polymerization, leaving a peptide (amide) bond. In another embodiment, the active group is a phosphoramidite group, useful for phosphoramidite oligomer synthesis, as is widely known in the art. The nucleotide monomer may also optionally contain one or more protecting groups, such as 4,4'-dimethoxytrityl (DMT), as is described herein. Several further methods for producing synthetic gene recognition reagents are known, differing in their skeletal structure and the specific chemistry of the base addition steps. The determination of which active groups to use for linking nucleotide monomers and which groups within the bases to protect, as well as the steps required in the production of oligomers, are well within the capabilities of those skilled in the art in the chemical field and in the specific field of nucleic acid and nucleic acid analog oligomer synthesis.

[0040] As used herein, the term "nucleic acid" refers to deoxyribonucleic acid (DNA) and ribonucleic acid (RNA). Nucleic acid analogs include, but are not limited to, phosphorothioate DNA (PS DNA), α,β-restricted nucleic acids (α,β-CNA), 2'-methoxyl RNA, 2'-fluoroRNA, phosphorodiamidate morpholino oligomers (PMO), locked nucleic acids (LNA), 2',4'-restricted ethyl nucleic acids ((S)-cEt), 2',4'-crosslinked nucleic acids NC(NH) (BNA-NC(NH)), 2',4'-crosslinked nucleic acids NC(N-methyl) (BNA-NC(N-Me)), ((S)-5'-C-methylDNA(RNA)), and 5'-E-vinylphosphonic acid nucleic acids (E-VP) (wherein R is H, OH, F, OMe, or O(CH2)2OMe), as well as combinations thereof that may include ribonucleotides or deoxyribonucleotide residues. "Oligononucleotides" are short, single-stranded gene recognition reagents. Oligonucleotides can be designated by the "mer" designation, depending on the length of the chain (i.e., the number of nucleotides or nucleic acid bases). For example, an oligonucleotide with 22 nucleotides is called a 22-mer.

[0041] "Peptide nucleic acid" refers to a DNA or RNA analog, or an analogue of DNA or RNA in which the sugar phosphodiester backbone is replaced with an N-(2-aminoethyl)glycine unit. In one embodiment, the peptide nucleic acid has the following structure: [ka] (In the formula, n is 1 or greater, and R1, R2, R3, R4, R5, and R6 are independently H;CH3, CH2OH, CH(CH3)OH, CH2SH, CH(CH3)CH3, CH2CH(CH3)CH3, CH(CH3)CH2CH3, CH2CH2SCH3, CH2CH3, CH2-C6H5, CH2-C6H4OH, 1H-indole-3-ylmethyl, CH2C(O)OH, CH2CH2C(O)OH, CH2C(O)NH2, CH2CH2C(O)NH 2, 1H-imidazole-4-ylmethyl, CH2CH2CH2CH2NH2, or CH2CH2CH2NHC(NH)NH2; linear or branched (C3-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, (C3-C8) cycloalkyl(C1-C6) alkylene, guanidine-containing group, CH2-(OCH2-CH2) n -OH,CH2-(OCH2-CH2) n -NH2, CH2-(OCH2-CH2) n -SH, CH2-(OCH2-CH2) n -NHC(NH)NH2, CH2-(OCH2-CH2) n - Morpholine, CH2-(OCH2-CH2) n -Piperadin, [ka] X is a linker such as a linear linker, R1 and R2 together form a 1,3-propylene bond, R3 and R4 together form a 1,3-propylene bond, or R5 and R6 together form a 1,3-propylene bond, and in any case R7 is an independent nucleic acid base, and if n is 2 or more, it forms a nucleic acid base sequence. It has.

[0042] A linker is a part of a compound that covalently connects one part to another. The linker does not substantially adversely affect the activity of the compound as a whole, for example, in the case of the present invention, the ability of a gene recognition reagent to function in its intended use. Besides providing a covalent linkage between two parts, a linker can have beneficial effects, such as physical separation of the parts to which it is linked, for example, to optimize the spacing to avoid steric effects. Linkers can also provide several additional functions, such as providing further sites (e.g., amines protected by protecting groups) to link further parts of the compound, or modifying the hydrophobic / hydrophilicity of the entire molecule to make the molecule more rigid. A linker is attached to the rest of a compound by any preferred bonding site ("bond"), for example, by a carbon-carbon bond, ester bond, thioester bond, amine bond, ether bond, amide bond, carbonate bond, or carbamic acid bond to further parts of the compound. The linker may be a hydrocarbil containing only carbon and hydrogen (e.g., 1 to 10 methylene groups), and optionally containing one or more heteroatoms such as N, O, and / or S. With respect to the present invention, in one embodiment, one preferred linker is a PEG group-(O-CH2-CH2) n - is a divalent portion comprising n in the range of 2 to 100, for example 2 to 10 (PEG 2-10 ), or 2-5 (PEG 2-5 The range is, for example, 2(PEG2), 3, 4, 5, 6, 7, 8, 9, or 10. The PEG linker may contain one or more methylene groups at any of its ends in addition to a suitable bond for attaching the PEG group to the linkage portion.

[0043] In another embodiment, PNA is gamma PNA (γPNA), which is an oligomer or polymer of gamma-modified N-(2-aminoethyl)glycine monomer in formula (I) above, where the γ carbon is the chiral center, and generally, one of R1 or R2 attached to the gamma carbon is H and the other is not hydrogen, or R1 and R2 are different, thereby making the gamma carbon a chiral center. If both R1 and R2 are hydrogen (N-(2-aminoethyl)-glycine skeleton) or are the same, there is no such chirality with respect to the gamma carbon. Alpha PNA (αPNA) is an oligomer or polymer of alpha-modified N-(2-aminoethyl)glycine monomer in formula (I) above, where the α carbon is the chiral center. Beta PNA (βPNA) is an oligomer or polymer of beta-modified N-(2-aminoethyl)glycine monomer in formula (I) above, where the β carbon is the chiral center. The following considerations regarding γPNA also apply to αPNA and βPNA. Furthermore, any two combinations of α, β, and γ carbons (α,β, α,γ, or β,γ) or all three (α,β,γ) can form a chiral center.

[0044] In various embodiments, for example, the skeletons of one or more of R1, R2, R3, R4, R5, and R6 are pegylated with 1 to 50 oxyethylene residues, i.e., n is 1 to 50 (including) [-O-CH2-CH2-] n The PEG group can be crosslinked with any suitable crosslinking group, for example, C 1-6 The skeleton can be linked by alkyl groups, or by aryl-alkyl groups.

[0045] In other embodiments, for example, the skeletons of one or more of R1, R2, R3, R4, R5, and R6 include one or more guanidine-containing groups, for example, an alkyl or aryl-alkyl moiety ending with a guanidine moiety.

[0046] In other embodiments, to facilitate cell uptake and endosomal escape, one or more of R1, R2, R3, R4, R5, and R6 are S1A, S1B, S1C, S1D, S1E, S1F, S1G, S1H, S1I, S1J, S1K, or S1L.

[0047] "Amino acid side chain" refers to the side chain of an amino acid. Amino acids have the following structure: [ka] (In the formula, "Side" refers to the amino acid side chain.) Examples of amino acid side chains that are not limited include CH3(Ala), CH2OH(Ser), CH(CH3)OH(Thr), CH2SH(Cys), CH(CH3)CH3(Val), CH2CH(CH3)CH3(Leu), CH(CH3)CH2CH3(Ile), CH2CH2SCH3(Met), 4-CH2-C6H4OH(Tyr), CH2-C6H5(Phe), and 1H-indole-3 Examples include ylmethyl (Trp), CH2C(O)OH (Asp), CH2CH2C(O)OH (Glu), CH2C(O)NH2 (Asn), CH2CH2C(O)NH2 (Gln), 1H-imidazole-4-ylmethyl (His), CH2CH2CH2CH2NH2 (Lys), or CH2CH2CH2NHC(NH)NH2 (Arg containing a guanidine / guanidium group). In this embodiment, glycine is not represented because it lacks a side chain (Side is H).

[0048] γPNA monomers incorporated into γPNA oligomers or polymers are referred herein to as “γPNA monomer residues,” and each residue has the same or different nucleic acid bases, e.g., modified nucleic acid bases as described herein, and thus, as in the case of DNA or RNA, the order of the bases in this γPNA constitutes its “sequence.” The nucleic acid base sequence of nucleic acid or nucleic acid analog oligomers or polymers, e.g., γPNA oligomers or polymers, is linked to the complementary sequence of adenine, guanine, cytosine, thymine and / or uracil residues of the nucleic acid strand by cooperative linkage, essentially similar to the Watson-Crick linkage of complementary bases in double-stranded DNA or RNA. A “Watson-Crick-like” linkage refers to a hydrogen bond of nucleic acid bases other than G, A, T, C, or U, e.g., the linkage of divalent bases represented herein by G, A, T, C, U, or other nucleic acid bases.

[0049] Unless otherwise specified, the nucleic acids and nucleic acid analogs described herein are not described in relation to any particular nucleotide sequence. This disclosure covers modified nucleic acid bases, compositions comprising these modified nucleic acid bases, and methods of use of these modified nucleic acid bases and compounds containing them. The usefulness of any particular embodiment described herein is generally applicable, although in all cases it generally depends on a specific sequence. Nucleic acid base sequences attached to the backbone of γPNA oligomers can hybridize with complementary nucleic acid base sequences of target nucleic acids or nucleic acid analogs via Watson-Crick or Watson-Crick-like hydrogen bonds. Those skilled in the art will understand that the compositions and methods described herein are sequence-independent and describe novel, generalized compositions and related methods comprising divalent nucleic acid bases.

[0050] The gene recognition reagent can be prepared as small oligonucleotides and assembled in situ, in vivo, ex vivo, or in vitro, as described in, for example, U.S. Patent Application Publication 2016 / 0083433A1, the entire contents of which are incorporated herein by reference. This method allows small oligomers, which have higher cell or tissue permeability compared to longer sequences such as trimers, to be transferred into cells, and these oligomers, once hybridized to a template nucleic acid, can be assembled into a continuous, longer sequence. This can be achieved in vitro or ex vivo, for example, to rapidly assemble a longer sequence for use in hybridizing to a target nucleic acid.

[0051] In one embodiment, the gene recognition reagent is provided in an array. Arrays are particularly useful for performing high-throughput assays such as gene detection assays. As used herein, the term “array” means reagents arranged or attached to two or more independent, identifiable and / or addressable positions on a substrate (support), e.g., the gene recognition reagents described herein. In one embodiment, the array is a device having two or more independent, identifiable reaction chambers, such as a 96-well dish, on which a reaction comprising identified components is carried out. In one embodiment, two or more gene recognition reagents comprising one or more divalent nucleic acid bases, as described herein, are immobilized on the substrate in a spatially addressable manner so that each primer or probe is positioned at a different and (addressable) identifiable position on the substrate. One or more gene recognition reagents are covalently linked to the substrate, or otherwise bound or positioned at an addressable position on the array. Substrates include, but are not limited to, multiwell plates, silicon chips, and beads. In one embodiment, the array comprises two or more sets of beads, each bead having an identifiable marker, such as a quantum dot or fluorescent tag, so that the beads can be individually identified using a flow cytometer, for example, but not limited to this. In one embodiment, the array is a multiwell plate comprising two or more wells containing the described gene recognition reagents for binding to specific sequences. Thus, reagents such as probes and primers are bound to specific positions on the array, or otherwise attached to the surface or interior, or positioned at specific locations. The reagents may be in any preferred form, including, but not limited to, solutions, dried products, lyophilized products, or glass-solidified products. When covalently linked to a substrate such as agarose beads or silicon chips, various linking techniques are known for attaching chemical moieties such as gene recognition reagents to such substrates.

[0052] Linkers and spacers for use in linking nucleic acids, peptide nucleic acids, and other nucleic acid analogs are widely known in the fields of chemistry and arrays and are therefore not described herein. As an example, without limitation, γPNA gene recognition reagents contain reactive amines that can be reacted with carboxyl-functional, cyanogen bromide-functional, N-hydroxysuccinimide-functional, carbonyldiimidazole-functional, or aldehyde-functional agarose beads, for example, available from Thermo Fisher Scientific (Pierce Protein Biology Products), Rockford, Illinois, and various other suppliers. The gene recognition reagents described herein may be attached to the substrate with or without a linker. Apparatus for carrying out the reaction and apparatus for reading the array are widely known and available, and informatics and / or statistical software or other computer implementation processes for analyzing array data and / or identifying genetic risk factors from data obtained from patient samples are known in the art.

[0053] Certain modified nucleic acid bases described herein fluoresce due to their ring structure. These compositions can be used as fluorescent dyes, or their internal fluorescence can be used as probes by binding to a target sequence in an in-situ assay or gel or blot, for example, to visualize the target sequence.

[0054] According to one aspect of the present invention, a method for detecting a target sequence in a nucleic acid is provided, comprising contacting a sample comprising a gene recognition reagent composition, such as those described herein, with a nucleic acid, and detecting the binding of the gene recognition reagent to the nucleic acid. In one embodiment, the gene recognition reagent is used such that a substrate, e.g., a nucleic acid sample immobilized on an array and labeled (e.g., fluorescently or radioactively labeled), is contacted with the immobilized gene recognition reagent, and the amount of labeled nucleic acid specifically bound to the gene recognition reagent is measured. In one variant, a nucleic acid comprising the gene recognition reagent or the target sequence of the gene recognition reagent is bound to the substrate, and the labeled nucleic acid or labeled gene recognition reagent comprising the target sequence of the gene recognition reagent binds to the immobilized gene recognition reagent or nucleic acid, respectively, to form a complex. In one embodiment, the nucleic acid of the complex comprises a partial target sequence so that the nucleic acid comprising the complete target sequence is advantageous in competition for the gene recognition reagent with the nucleic acid that forms the complex. The complex is then exposed to a nucleic acid sample, and the loss of the bound label from the complex can be detected and quantified according to standard methods for assisting the quantification of nucleic acid markers in a nucleic acid sample. These are just two of the many analytical assays that may be available to detect or quantify the presence of specific nucleic acids in nucleic acid samples.

[0055] With respect to compositions such as nucleic acids or gene recognition reagents as described herein, “immobilized” means attached to a substrate of any physical structure or chemical composition. Immobilized compositions are immobilized by any method useful for their end use. Compositions are immobilized by covalent or non-covalent methods, such as by covalent bonds between amine groups and linkers or spacers, or by non-covalent bonds including van der Waals forces and / or hydrogen bonds. “Label” is a chemical portion useful for the detection or purification of a molecule or composition containing the label. Labels are, for example, but are not limited to: 14 C, 32 P, 35The label may be a radioactive portion such as S, a fluorescent dye such as fluorescein isothiocyanate or cyanine dye, an enzyme, or a ligand for binding to other compounds, such as biotin for binding to streptavidin, or an epitope for binding to an antibody. Numerous such labels and methods of use are known to those skilled in the art of immunology and molecular biology. However, since certain nucleic acid bases described herein are fluorescent, incorporating such bases into nucleotide residues of nucleic acids or nucleic acid analogs, or covalently linking divalent nucleic acid bases to nucleic acids, nucleic acid analogs, binding reagents, ligands, or other detection reagents, allows for the detection and / or quantification of reagents in samples, reaction mixtures, arrays, etc.

[0056] In yet another aspect of the present invention, a method for isolating and purifying nucleic acids containing a target sequence is provided. In one non-limited embodiment, a gene recognition reagent as described herein is immobilized on a substrate such as beads (e.g., but not limited to agarose beads, beads containing a fluorescent marker for sorting, or magnetic beads), a porous matrix, a surface, or a tube. A nucleic acid sample is brought into contact with this immobilized gene recognition reagent, and the nucleic acid containing the target sequence binds to the gene recognition reagent. The bound nucleic acid is then washed to remove unbound nucleic acid, and the bound nucleic acid is then eluted and precipitated by any useful method widely known in the field of molecular biology, or otherwise concentrated.

[0057] In a further embodiment, a kit is provided. The kit comprises at least one container in any form containing a cartridge for automated nucleic acid, nucleic acid analogue, or PNA synthesis, and may comprise one or more containers in the form of individual, independent, and optionally independently addressable compartments for use in an automated sequencing apparatus for producing nucleic acids and / or nucleic acid analogues. The containers may be disposable or may contain contents sufficient for multiple uses. The kit may also comprise an array. The kit may optionally comprise one or more additional reagents for use in the production or use of gene recognition reagents in any embodiment described herein. The kit comprises a container containing any divalent nucleic acid base in any form described herein, or a monomer or gene recognition reagent in any manner described herein. Different nucleic acid bases, monomers or gene recognition reagents are generally sealed in separate containers, which may be separate compartments within a cartridge.

[0058] In several embodiments, the compounds and gene recognition reagents are used for therapeutic purposes, and therefore these compounds and gene recognition reagents are formulated as pharmaceuticals, pharmaceutical compositions, or dosage forms for human and veterinary use, comprising a composition for human and veterinary use, comprising a therapeutically effective amount of the compound or gene recognition reagent and excipients for therapeutic delivery, e.g., but not limited to oral, topical, intravenous, intramuscular, or subcutaneous administration, such as a vehicle or diluent. The composition can be formulated conventionally using a solid or liquid vehicle, diluent, and additive suitable for the desired mode of administration. Orally, these compounds can be administered in the form of tablets, capsules, granules, powders, etc. These compositions may optionally comprise one or more additional active agents, as is widely known in the fields of pharmacy, medicine, veterinary medicine, or biology. The compounds described herein may be administered in any effective mode. Further examples of delivery routes include, but are not limited to, local delivery, e.g., transdermal delivery, inhalation delivery, enema, intraocular delivery, intraocular delivery, and nasal delivery; enteral delivery, e.g., oral delivery, delivery by gastric feeding tube or swallowing, and rectal delivery; and parenteral delivery, e.g., intravenous delivery, intraarterial delivery, intramuscular delivery, intracardiac delivery, subcutaneous delivery, intraosseous delivery, intradermal delivery, subarachnoid delivery, intraperitoneal delivery, transdermal delivery, iontophoresis delivery, transmucosal delivery, epidural delivery, and intravitreous delivery. Therapeutic / pharmaceutical compositions are prepared according to accepted pharmaceutical procedures, as is widely known.

[0059] Any of the compounds described herein may be formulated or otherwise manufactured as a composition suitable for use, such as a pharmaceutical dosage form or medicinal product, in which the compound or gene recognition reagent is the active ingredient. For example, medicinal products described herein may be oral tablets, capsules, caplets, liquid-filled or gel-filled capsules, etc. The composition may comprise a pharmaceutically acceptable carrier or excipient. An “excipient” is an inert substance used as a carrier for the active ingredient of a drug. Although “inert,” excipients can promote and assist in increasing the delivery, stability, or bioavailability of the active ingredient in a medicinal product. Not limited examples of useful excipients, as available in the pharmaceutical / formulation field, include antifouling agents, binders, rheological modifiers, coatings, disintegrants, emulsifiers, oils, buffers, salts, acids, bases, fillers, diluents, solvents, flavoring agents, colorants, flow enhancers, lubricants, preservatives, antioxidants, adsorbents, vitamins, and sweeteners. [Examples]

[0060] Examples - Synthesis of modified nucleic acid bases and cell-permeable γPNA The compounds and gene recognition reagents described herein are synthesized according to methods known in the fields of chemical and organic synthesis. Exemplary synthesis schemes are shown in Figures 6 and 7.

[0061] The following numbered items describe non-limiting embodiments and aspects of the present invention.

[0062] Article 1: comprising multiple nucleic acid base portions attached to a nucleic acid skeleton or nucleic acid analog skeleton, wherein at least one nucleic acid base portion: [ka] (In the formula, X1 is =O, =S, =Se, or CH3; X2 is H, CH3, CN, NC, N3, C(O)OH, or C(O)NH2; X3 is either O or S; X4 is H, C(O)CH3, or C(O)OCH3; and Y is either N or CH. (1) If X1 and X3 are O, then X2 is neither H nor methyl. It is a gene recognition reagent.

[0063] Paragraph 2: At least one nucleic acid base portion: [ka] The gene recognition reagent described in item 1.

[0064] Paragraph 3: At least one nucleic acid base portion: [ka] The gene recognition reagent described in item 1.

[0065] Section 4: At least one nucleic acid base portion: [ka] The gene recognition reagent described in item 1.

[0066] Section 5: At least one nucleic acid base portion: [ka] The gene recognition reagent described in item 1.

[0067] Section 6: At least one nucleic acid base portion: [ka] The gene recognition reagent described in item 1.

[0068] Paragraph 7: At least one nucleic acid base portion: [ka] The gene recognition reagent described in item 1.

[0069] Paragraph 8: At least one nucleic acid base portion: [ka] The gene recognition reagent described in item 1.

[0070] Item 9: A gene recognition reagent according to any one of items 1 to 8, wherein the backbone is selected from one of the following: DNA, RNA, peptide nucleic acid (PNA), phosphorothioate DNA (PS DNA), α,β-restricted nucleic acid (α,β-CNA), 2'-methoxyl RNA, 2'-fluoroRNA, locked nucleic acid (LNA), 2',4'-restricted ethyl nucleic acid ((S)-cEt), 2',4'-crosslinked nucleic acid NC(NH) (BNA-NC(NH)), 2',4'-crosslinked nucleic acid NC(N-methyl) (BNA-NC(N-Me)), 2'-(R)-(S)-5'-C-methyl DNA, or 2'-R-5'-E-vinylphosphonic acid nucleic acid (E-VP) (wherein R is H, OH, F, OMe, or O(CH2)2OMe backbone).

[0071] Item 10: A gene recognition reagent as described in Item 1, wherein the backbone is a peptide nucleic acid (PNA) backbone.

[0072] Item 11: The gene recognition reagent described in Item 10, wherein the skeleton is pegylated by one or more PEG moieties of 2 to 50 (-O-CH2-CH2-) residues linked to the skeleton.

[0073] Item 12: The gene recognition reagent according to item 10, wherein the skeleton comprises one or more guanidine moieties linked to the skeleton.

[0074] Item 13: A gene recognition reagent as described in Item 1, wherein the skeleton is a gamma peptide nucleic acid (γPNA) skeleton.

[0075] Section 14: The skeleton consists of residues [ka] (In the formula, n is 1 or greater, and R1, R2, R3, R4, R5, and R6 are independently H; CH3, CH2OH, CH(CH3)OH, CH2SH, CH(CH3)CH3, CH2CH(CH3)CH3, CH(CH3)CH2CH3, CH2CH2SCH3, CH2CH3, CH2-C6H5, CH2-C6H4OH, 1H-indole-3-ylmethyl, CH2C(O)OH, CH2CH2C(O)OH, CH2C(O)NH2, CH2CH2C(O)NH 2, 1H-imidazole-4-ylmethyl, CH2CH2CH2CH2NH2, or CH2CH2CH2NHC(NH)NH2; linear or branched (C3-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, (C3-C8) cycloalkyl(C1-C6) alkylene, guanidine-containing group, CH2-(OCH2-CH2) n -OH,CH2-(OCH2-CH2) n -NH2, CH2-(OCH2-CH2) n -SH, CH2-(OCH2-CH2) n -NHC(NH)NH2, CH2-(OCH2-CH2) n - Morpholine, CH2-(OCH2-CH2) n -Piperadin, [ka] (wherein X is a linker, R1 and R2 together form a 1,3-propylene bond, R3 and R4 together form a 1,3-propylene bond, or R5 and R6 together form a 1,3-propylene bond, and in any case R7 is independently a nucleic acid base.) A gene recognition reagent as described in paragraph 1, comprising a PNA skeleton.

[0076] Section 15: The gene recognition reagent described in Section 14, wherein at least one of R1, R2, R3, R4, R5, and R6 is S1A, S1B, S1C, S1D, S1E, S1F, S1G, S1H, S1I, S1J, S1K, or S1L.

[0077] Item 16: A gene recognition reagent as described in Item 14 or 15, wherein the α-carbon, β-carbon, or γ-carbon is a chiral center.

[0078] Item 17: A gene recognition reagent as described in Item 16, wherein the γ-carbon is a chiral center.

[0079] Paragraph 18: R2 is [ka] The gene recognition reagent described in item 17.

[0080] Item 19: The gene recognition reagent described in Item 16, wherein R1, R3, R4, R5, and R6 are H.

[0081] Item 20: A gene recognition reagent according to any one of items 1 to 19, comprising terminal groups linked to a nucleic acid skeleton or nucleic acid analog skeleton for the self-assembly of two or more adjacent gene recognition reagents on a nucleic acid template.

[0082] Item 21: A gene recognition reagent as described in Item 20, wherein the terminal group is a fused polycyclic aromatic moiety of 2 to 5 rings.

[0083] Item 22: The gene recognition reagent described in Item 20, wherein the terminal group is a sulfhydryl group or a thioester group.

[0084] Item 23: A gene recognition reagent according to any one of items 1 to 22, wherein multiple nucleic acid base portions form a sequence comprising TTC, TTCTTC, TCT, TCTTCT, CTT, CTTCTT, CCG, CCGCCG, CGC, CGCCGC, GCC, GCCGCC, CCG, CGGCGG, GCG, GCGGCG, GGC, GGCGGC, CTG, CTGCTG, TGC, TGCTGC, GCT, GCTGCT, CAG, CAGCAG, AGC, AGCAGC, GCA, GCAGCA, CAGG, CAGGCAGG, AGGC, AGGCAGGC, GGCA, GGCAGGCA, GCAGGCAG, AGAAT, GAATA, AATAG, ATAGA, TAGAA, GGCCCC, GCCCCG, CCCCGG, CCCGGC, CCGGCC, and CGGCCC, or any of the above consecutive repeats.

[0085] Item 24: A gene recognition reagent according to any one of items 1 to 23, wherein multiple nucleic acid base portions are composed of sequences complementary to the target sequence of the nucleic acid.

[0086] Item 25: A gene recognition reagent according to any one of items 1 to 24, having a nucleic acid base portion of 3 to 25.

[0087] Section 26: Structure: [ka] (In the formula, X1 is =O, =S, =Se, or CH3; X2 is H, CH3, CN, NC, N3, C(O)OH, or C(O)NH2; X3 is either O or S; X4 is H, C(O)CH3, or C(O)OCH3; and Y is either N or CH. (I) If X1 and X3 are O, then X2 is neither H nor methyl. A compound comprising a nucleic acid skeleton monomer or nucleic acid analog skeleton monomer linked to a nucleic acid base portion having a specific characteristic.

[0088] Paragraph 27: At least one nucleic acid base portion: [ka] The compound described in item 26.

[0089] Article 28: At least one nucleic acid base portion: [ka] The compound described in item 26.

[0090] Paragraph 29: At least one nucleic acid base portion: [ka] The compound described in item 26.

[0091] Paragraph 30: At least one nucleic acid base portion: [ka] The compound described in item 26.

[0092] Item 31: A compound according to any one of items 26 to 30, wherein the backbone monomer is DNA, RNA, peptide nucleic acid (PNA), phosphorothioate DNA (PS DNA), α,β-restricted nucleic acid (α,β-CNA), 2'-methoxyl RNA, 2'-fluoroRNA, locked nucleic acid (LNA), 2',4'-restricted ethyl nucleic acid ((S)-cEt), 2',4'-crosslinked nucleic acid NC(NH) (BNA-NC(NH)), 2',4'-crosslinked nucleic acid NC(N-methyl) (BNA-NC(N-Me)), 2'-(R)-(S)-5'-C-methylDNA, or 2'-R-5'-E-vinylphosphonic acid nucleic acid (E-VP) (wherein R is H, OH, F, OMe, or O(CH2)2OMe backbone monomer).

[0093] Item 32: A compound according to any one of items 26 to 31, wherein the backbone monomer is a peptide nucleic acid (PNA) backbone monomer.

[0094] Section 33: The compound described in Section 32, wherein the backbone monomer is pegylated by one or more PEG moieties of 2 to 50 (-O-CH2-CH2-) residues linked to the backbone monomer.

[0095] Item 34: The compound according to item 32, wherein the skeleton comprises one or more guanidine moieties linked to the skeleton.

[0096] Item 35: The compound described in item 26, wherein the backbone monomer is a gamma peptide nucleic acid (γPNA) backbone monomer.

[0097] Section 36: Skeletal monomers, structure: [ka] (In the formula, R1, R2, R3, R4, R5, and R6 are independently H; CH3, CH2OH, CH(CH3)OH, CH2SH, CH(CH3)CH3, CH2CH(CH3)CH3, CH(CH3)CH2CH3, CH2CH2SCH3, CH2CH3, CH2-C6H5, 1H-indole-3-ylmethyl, CH2-C6H4OH, CH2C(O)OH, CH2CH2C(O)OH, CH2C(O)NH2, CH2CH2C(O)NH2, 1H- Imidazole-4-ylmethyl, CH2CH2CH2CH2NH2, or CH2CH2CH2NHC(NH)NH2; linear or branched (C3-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, (C3-C8) cycloalkyl(C1-C6) alkylene, guanidine-containing group, CH2-(OCH2-CH2) n -OH,CH2-(OCH2-CH2) n -NH2, CH2-(OCH2-CH2) n -SH, CH2-(OCH2-CH2) n -NHC(NH)NH2, CH2-(OCH2-CH2) n - Morpholine, CH2-(OCH2-CH2) n -Piperadin, [ka] (wherein X is a linker, R1 and R2 together form a 1,3-propylene bond, R3 and R4 together form a 1,3-propylene bond, or R5 and R6 together form a 1,3-propylene bond, and in any case R7 is independently a nucleic acid base.) The compound described in item 26, which has a PNA skeleton having the following characteristics.

[0098] Item 37: The compound described in item 36, wherein at least one of R1, R2, R3, R4, R5, and R6 is S1A, S1B, S1C, S1D, S1E, S1F, S1G, S1H, S1I, S1J, S1K, or S1L.

[0099] Item 38: The compounds described in item 36 or 37, wherein the α-carbon, β-carbon, or γ-carbon is a chiral center.

[0100] Item 39: A compound described in Item 38, wherein the γ-carbon is a chiral center.

[0101] Paragraph 40: R2 [ka] The compound described in paragraph 39.

[0102] Item 41: The compound described in item 37, wherein R1, R3, R4, R5, and R6 are H.

[0103] Item 42: A kit comprising a container containing one of the compounds described in any one of items 36 to 41.

[0104] Section 43: The kit according to Section 42, comprising monomer-bound adenine, guanine, cytosine, and one or both of uracil and thymine in separate containers.

[0105] Item 44: A kit comprising a container containing a gene recognition reagent as described in any one of items 1 through 25.

[0106] Item 45: A kit as described in any one of items 42 to 44, wherein the container is one or more compartments in a cartridge for use in an automated device.

[0107] Item 46: An array comprising a gene recognition reagent as described in any one of items 1 to 19.

[0108] Paragraph 47: A method for detecting a target sequence in nucleic acids, comprising contacting a sample containing nucleic acids with a gene recognition reagent described in any one of paragraphs 1 to 25, and detecting the binding of the gene recognition reagent to the nucleic acids.

[0109] Paragraph 48: A method for isolating and purifying a nucleic acid sample containing a target sequence, comprising contacting the nucleic acid sample with a gene recognition reagent described in any one of paragraphs 1 to 29, separating the nucleic acid sample from the gene recognition reagent, leaving the nucleic acid bound to the gene recognition reagent, and separating the gene recognition reagent from the nucleic acid bound to the gene recognition reagent.

[0110] Paragraph 49: The method according to Paragraph 48, comprising: immobilizing a gene recognition reagent on a substrate; contacting a nucleic acid with the substrate; washing the substrate to remove unbound nucleic acids from the substrate, but leaving the nucleic acids bound to the substrate bound; and eluting the bound nucleic acids from the substrate.

[0111] Item 50: A composition comprising a gene recognition reagent or compound as described in any one of items 1 to 49 and a pharmaceutically acceptable excipient.

[0112] The present invention has been described in relation to specific exemplary embodiments, dispersible compositions, and their uses. However, those skilled in the art will recognize that various substitutions, modifications, or combinations of any of the exemplary embodiments can be made without departing from the spirit and scope of the invention. Accordingly, the present invention is not limited by the description of the exemplary embodiments, but rather by the claims appended to the original application.

Claims

1. comprising a plurality of nucleobase moieties attached to a nucleic acid backbone or nucleic acid analog backbone, wherein at least one nucleobase moiety is: 【Chemistry 1】 (In the formula, X 3 is O or S; and Y is N) and a gene recognition reagent, wherein the nucleic acid backbone or nucleic acid analog backbone is selected from one of DNA, RNA, peptide nucleic acid (PNA), phosphorothioate DNA (PS DNA), α,β-constrained nucleic acid (α,β-CNA), 2'-methoxyl RNA, 2'-fluoro RNA, locked nucleic acid (LNA), 2',4'-constrained ethyl nucleic acid ((S)-cEt), 2',4'-bridged nucleic acid NC(N-H) (BNA-NC(N-H)), 2',4'-bridged nucleic acid NC(N-methyl) (BNA-NC(N-Me)), 2'-(R)-(S)-5'-C-methyl DNA, or 2'-R-5'-E-vinylphosphonate nucleic acid (E-VP), where R is H, OH, F, OMe, or O(CH 2 ) 2 OMe.

2. A gene recognition reagent as described in claim 1, wherein the backbone is a nucleic acid analog backbone, and the nucleic acid analog backbone is a peptide nucleic acid (PNA) backbone.

3. A gene recognition reagent as described in claim 2, wherein the PNA backbone comprises one or more guanidine moieties linked to the PNA backbone.

4. A gene recognition reagent as described in claim 1, wherein the backbone is a nucleic acid analog backbone, and the nucleic acid analog backbone is a gamma-peptide nucleic acid (gamma-PNA) backbone.

5. the backbone is a nucleic acid analog backbone, the nucleic acid analog backbone is a PNA backbone; The PNA backbone and the plurality of nucleobase moieties together comprise: 【Chemistry 2】 (wherein n′ is 2 or more, and R 1 , R 2 , R 3 , R 4 , R 5 , and R 6 are independently H, CH 3 , C.H. 2 OH, CH(CH 3 ) OH, CH 2 SH, CH (CH 3 ) CH 3 , C.H. 2 CH (CH 3 ) CH 3 , CH(CH 3 ) CH 2 CH 3 , C.H. 2 CH 2 SCH 3 , C.H. 2 CH 3 , C.H. 2 -C 6 H 5 , C.H. 2 -C 6 H 4 OH, 1H-indol-3-ylmethyl, CH 2 C(O)OH, CH 2 CH 2 C(O)OH, CH 2 C(O)NH 2 , C.H. 2 CH 2 C(O)NH 2 , 1H-imidazol-4-ylmethyl, CH 2 CH 2 CH 2 CH 2 NH 2 , C.H. 2 CH 2 CH 2 NHC (NH) NH 2 , linear or branched chain (C 3 -C 8 ) alkyl, (C 2 -C 8 ) alkenyl, (C 2 -C 8 ) alkynyl, (C 3 -C 8 ) aryl, (C 3 -C 8 ) cycloalkyl, (C 3 -C 8 ) aryl (C 1 -C 6 ) alkylene, (C 3 -C 8 ) cycloalkyl (C 1 -C 6 ) alkylene, guanidine-containing group, CH 2 -(OCH 2 -CH 2 ) n -OH, CH 2 -(OCH 2 -CH 2 ) n -NH 2 , C.H. 2 -(OCH 2 -CH 2 ) n -SH, CH 2 -(OCH 2 -CH 2 ) n -NHC(NH)NH 2 , C.H. 2 -(OCH 2 -CH 2 ) n -morpholine, CH 2 -(OCH 2 -CH 2 ) n -piperazine, where n is 1 to 50; 【Transformation 3】 wherein X is a linker: 【Chemistry 4】 ) and R 1 and R 2 together form a 1,3-propylene bond, and R 3 and R 4 together form a 1,3-propylene bond, or R 5 and R 6 together form a 1,3-propylene bond, and R 7 are independently nucleobases of said plurality of nucleobase moieties) The gene recognition reagent according to claim 1 , wherein

6. R 1 , R 2 , R 3 , R 4 , R 5 , and R 6 6. The gene recognition reagent according to claim 5, wherein at least one of the following is S1A, S1B, S1C, S1D, S1E, S1F, S1G, S1H, S1I, S1J, S1K, or S1L.

7. 7. The gene recognition reagent according to claim 5, wherein the α-carbon, β-carbon, or γ-carbon is the chiral center.

8. 8. The gene recognition reagent according to claim 7, wherein the γ-carbon is a chiral center.

9. R 2 but: 【Transformation 5】 The gene recognition reagent according to claim 8,

10. The plurality of nucleobase moieties are TTC, TTCTTC, TCT, TCTTCT, CTT, CTTCTT, CCG, CCGCCG, CGC, CGCCGC, GCC, GCCGCC, CGG, CGGCGG, GCG , GCGGCG, GGC, GGCGGC, CTG, CTGCTG, TGC, TGCTGC, GCT, GCTGCT, CAG, CAGCAG, AGC, AGCAGC, GCA, GCAGCA, CAGG, C The gene recognition reagent according to any one of claims 1 to 8, which forms a sequence comprising AGGCAGG, AGGC, AGGCAGGC, GGCA, GGCAGGCA, GCAG, GCAGGCAG, AGAAT, GAATA, AATAG, ATAGA, TAGAA, GGCCCC, GCCCCG, CCCCGG, CCCGGC, CCGGCC, or CGGCCC, or consecutive repeats of any of the foregoing.

11. A gene recognition reagent described in any one of claims 1 to 10, wherein the multiple nucleic acid base portions are arranged in a sequence complementary to a target sequence of a nucleic acid.

12. The gene recognition reagent according to any one of claims 1 to 11, which has 3 to 25 nucleic acid base moieties.

13. comprising nucleic acid backbone monomers or nucleic acid analog backbone monomers, The backbone monomer has the structure: 【Transformation 6】 (In the formula, X 3 is O or S; and Y is N) and linked to a nucleobase moiety having the formula: The compound wherein the nucleic acid backbone monomer or nucleic acid analog backbone monomer is DNA, RNA, peptide nucleic acid (PNA), phosphorothioate DNA (PS DNA), α,β-constrained nucleic acid (α,β-CNA), 2'-methoxyl RNA, 2'-fluoro RNA, locked nucleic acid (LNA), 2',4'-constrained ethyl nucleic acid ((S)-cEt), 2',4'-bridged nucleic acid NC(N-H) (BNA-NC(N-H)), 2',4'-bridged nucleic acid NC(N-methyl) (BNA-NC(N-Me)), 2'-(R)-(S)-5'-C-methyl DNA, or 2'-R-5'-E-vinylphosphonate nucleic acid (E-VP), where R is H, OH, F, OMe, or O(CH 2 ) 2 OMe.

14. The compound of claim 13, wherein the backbone monomer is a nucleic acid analog backbone monomer, and the nucleic acid analog backbone monomer is a peptide nucleic acid (PNA) backbone monomer.

15. The compound of claim 14, wherein the PNA backbone monomer comprises one or more guanidine moieties linked to the PNA backbone monomer.

16. The compound of claim 13, wherein the backbone monomer is a nucleic acid analog backbone monomer, and the nucleic acid analog backbone monomer is a gamma-peptide nucleic acid (gamma-PNA) backbone monomer.

17. The backbone monomer of claim 17, wherein the backbone monomer is a nucleic acid analog backbone monomer, the nucleic acid analog backbone monomer having the structure: 【Transformation 7】 (In the formula, R 1 , R 2 , R 3 , R 4 , R 5 , and R 6 are independently H, CH 3 , C.H. 2 OH, CH(CH 3 ) OH, CH 2 SH, CH (CH 3 ) CH 3 , C.H. 2 CH (CH 3 ) CH 3 , CH(CH 3 ) CH 2 CH 3 , C.H. 2 CH 2 SCH 3 , C.H. 2 CH 3 , C.H. 2 -C 6 H 5 , 1H-indol-3-ylmethyl, CH 2 -C 6 H 4 OH, CH 2 C(O)OH, CH 2 CH 2 C(O)OH, CH 2 C(O)NH 2 , C.H. 2 CH 2 C(O)NH 2 , 1H-imidazol-4-ylmethyl, CH 2 CH 2 CH 2 CH 2 NH 2 , C.H. 2 CH 2 CH 2 NHC (NH) NH 2 , linear or branched chain (C 3 -C 8 ) alkyl, (C 2 -C 8 ) alkenyl, (C 2 -C 8 ) alkynyl, (C 3 -C 8 ) aryl, (C 3 -C 8 ) cycloalkyl, (C 3 -C 8 ) aryl (C 1 -C 6 ) alkylene, (C 3 -C 8 ) cycloalkyl (C 1 -C 6 ) alkylene, guanidine-containing group, CH 2 -(OCH 2 -CH 2 ) n -OH, CH 2 -(OCH 2 -CH 2 ) n -NH 2 , C.H. 2 -(OCH 2 -CH 2 ) n -SH, CH 2 -(OCH 2 -CH 2 ) n -NHC(NH)NH 2 , C.H. 2 -(OCH 2 -CH 2 ) n -morpholine, CH 2 -(OCH 2 -CH 2 ) n -piperazine, where n is 1 to 50; 【Transformation 8】 wherein X is a linker: 【Chemistry 9】 ) and R 1 and R 2 together form a 1,3-propylene bond, and R 3 and R 4 together form a 1,3-propylene bond, or R 5 and R 6 together form a 1,3-propylene bond) 14. The compound of claim 13, which is a PNA backbone monomer having the formula:

18. R 1 , R 2 , R 3 , R 4 , R 5 , and R 6 18. The compound of claim 17, wherein at least one of: S1A, S1B, S1C, S1D, S1E, S1F, S1G, S1H, S1I, S1J, S1K, or S1L.

19. 19. The compound of claim 17 or 18, wherein the α-carbon, β-carbon, or γ-carbon is a chiral center.

20. 20. The compound of claim 19, wherein the γ-carbon is a chiral center.

21. R 2 but 【Chemistry 10】 21. The compound of claim 20, wherein:

22. A method for in vitro detection of a target sequence in a nucleic acid, comprising contacting a gene recognition reagent according to any one of claims 1 to 12 with a sample comprising a nucleic acid, and detecting binding between the gene recognition reagent and the nucleic acid.

23. A method for in vitro isolation and purification of nucleic acids containing a target sequence, comprising contacting a nucleic acid sample with the gene recognition reagent of any one of claims 1 to 12, separating the nucleic acid sample from the gene recognition reagent, leaving the nucleic acid bound to the gene recognition reagent bound to the gene recognition reagent, and separating the gene recognition reagent from the nucleic acid bound to the gene recognition reagent.

24. The gene recognition reagent is immobilized on a substrate, 24. The method of claim 23, comprising contacting nucleic acid with the substrate, washing the substrate to remove unbound nucleic acid from the substrate, but leaving bound nucleic acid on the substrate bound, and eluting the bound nucleic acid from the substrate.

25. A composition comprising the gene recognition reagent according to any one of claims 1 to 12 or the compound according to any one of claims 13 to 21, and a pharmaceutically acceptable excipient.