Bidirectional cyclic peptide type DNA (Deoxyribose Nucleic Acid) coding compound and library construction method and application thereof
By constructing a DNA-encoded cyclic peptide library using a two-way synthesis method, the problem of insufficient macrocyclic diversity in traditional methods is solved, and the efficient synthesis of large-size cyclic peptide structures and enhanced target affinity are achieved, thus expanding the chemical space of the compound library.
Patent Information
- Application Number
- CN202410722743.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-05
- Publication Date
- 2025-12-05
AI Technical Summary
Traditional methods for synthesizing cyclic peptide DNA-encoded molecular libraries can only synthesize small-sized cyclic structures, resulting in insufficient overall diversity of macrocycles and inadequate target affinity, making it difficult to meet the needs of drug development.
A DNA-encoded cyclic peptide library was constructed using a two-way synthesis method. Large-sized cyclic structures were synthesized through nucleic acid compatibility reactions to enhance the interaction with the target and improve the diversity of macrocyclic structures.
This technology enables the synthesis of large-size cyclic peptide structures in a limited number of chemical steps, enhancing affinity for targets, expanding the chemical space of the compound library, and increasing the likelihood of discovering lead compounds.
Smart Images

Figure BDA0004877685340000021 
Figure BDA0004877685340000023 
Figure BDA0004877685340000025
Abstract
Description
Technical Field
[0001] This invention relates to the fields of chemistry and biotechnology, and more specifically, to bidirectional cyclic peptide DNA-encoded compounds, their library construction methods, and applications. Background Technology
[0002] Cyclic peptides are a special class of compounds, occupying a chemical space between small molecules and biomacromolecules, possessing unique physicochemical properties and drug-making advantages. They combine the advantages of traditional small molecule synthesis—convenient operation, low cost, low toxicity, and potential for oral drug development and transmembrane targeting of intracellular targets—with the advantages of biomacromolecules—high affinity and selectivity for their targets. In particular, cyclic peptides exhibit unique advantages in targeting disease-related protein-protein interaction (PPI) targets. With the continuous development of chemical-biological technologies and researchers' deeper understanding of the physicochemical properties, metabolic stability, and permeability of cyclic peptides, they have attracted a new wave of attention.
[0003] High-throughput screening based on physical compound libraries is one of the most important methods for discovering cyclic peptide lead compounds. DNA-encoded compound library technology is a new generation of high-throughput screening technology platform. It utilizes combinatorial techniques to rapidly synthesize large numbers of compounds, while combining chemical biology techniques to read out the structures of compounds with affinity from a mixed library of molecules. This technology platform has advantages such as low cost, short time, and low labor costs, and has been widely adopted by pharmaceutical companies and universities to find lead compound structures targeting specific targets. Therefore, we designed and developed a cyclic peptide DNA-encoded compound library with drug-like properties.
[0004] Traditional synthesis methods for DNA-encoded cyclic peptide libraries employ a one-way synthetic approach, limiting the number of chemical synthesis steps to synthesize only small-sized cyclic structures. This results in insufficient diversity of macrocyclic structures and inadequate affinity for targets. Therefore, we designed and developed a two-way synthetic method for constructing DNA-encoded cyclic peptide libraries. This method offers the advantage of synthesizing large-sized cyclic structures using a limited number of chemical steps. It enhances both macrocyclic structural diversity and target affinity, thereby increasing the likelihood of drug development. Summary of the Invention
[0005] One object of the present invention is to provide a novel bidirectional cyclic peptide DNA-encoded compound and a method for constructing its molecular library.
[0006] In a first aspect of the present invention, a method for preparing a bidirectional cyclic peptide-type DNA-encoded compound is provided, comprising the following steps:
[0007] (a) Provides a compound of formula M0, wherein the compound of formula M0 contains a Linker (L), A2 and A3, wherein the Linker is connected to an -A1-N segment;
[0008]
[0009] in,
[0010] N (or ) is a single-stranded / double-stranded deoxyribonucleotide sequence, a single-stranded / double-stranded ribonucleotide sequence, or a combination thereof;
[0011] A1 is a chemical bond or chemical structure connecting the Linker and N;
[0012] Linker (or L) is a multifunctional linker framework;
[0013] A2 and A3 are reactive groups;
[0014] (b) Provide the composite block shown in M1;
[0015]
[0016] In the formula,
[0017] For the core structure of the composite building blocks;
[0018] G1 and G2 are reactive groups;
[0019] (c) Reacting the compound represented by formula M0 with the synthetic building block represented by formula M1 to form a compound represented by formula Q containing a nucleic acid fragment;
[0020]
[0021] In the formula,
[0022] m and n are each independent integers from 1 to 100;
[0023] Each O1 and O2 is an independent chemical bond or chemical structure;
[0024] N (or A1, Linker (or L), G2 is defined as above;
[0025] (d) The compound represented by formula Q is subjected to a cyclization reaction with the synthetic building block represented by formula M1 to form a compound represented by formula R containing a nucleic acid fragment:
[0026]
[0027] In the formula,
[0028] P1 and P2 are chemical structures or chemical bonds formed by the reaction of two reactive groups G2 on the central structure of the synthetic building block in the compound represented by formula Q with two reactive groups (G1 and G2) on the synthetic building block represented by formula M1.
[0029] N (or A1, Linker (or L), m, n, O1 and O2 are defined as above.
[0030] In another preferred embodiment, m and n are each independently an integer from 1 to 50. More preferably, m and n are each independently 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.
[0031] In another preferred example, m and n are the same.
[0032] In another preferred example, m and n are not the same.
[0033] In another preferred embodiment, O1 is formed by a nucleic acid-compatible chemical reaction between A2 in the compound represented by formula M0 and G1 in the synthetic building block represented by formula M1;
[0034] O2 is formed by a nucleic acid-compatible chemical reaction between A3 in the compound represented by formula M0 and G2 in another synthetic building block represented by formula M1.
[0035] In another preferred embodiment, when m and n are 1, O1 is the chemical bond formed between A2 in the compound of formula M0 and G1 in the compound of formula M1, and O2 is the chemical bond formed between A3 in the compound of formula M0 and another G2 in the compound of formula M1.
[0036] When both m and n are greater than 1, O1 is the chemical bond formed between G2 on the central structure of the synthetic building block in compound Q and G1 on compound M1, and O2 is the chemical bond formed between another G2 on the central structure of the synthetic building block in compound Q and another G1 on compound M1.
[0037] In another preferred embodiment, P1 is formed by G2 in the compound represented by formula Q and G1 in the synthetic building block represented by formula M1 through a nucleic acid-compatible chemical reaction;
[0038] P2 is formed by a nucleic acid-compatible chemical reaction between another G2 in the compound represented by formula Q and G2 in the same synthetic building block represented by formula M1.
[0039] In another preferred embodiment, step (c) further includes: reacting two or one G2 of the compound Q obtained from the reaction with the synthetic building block shown in the formula M1, repeating the reaction multiple times to obtain a compound Q having multiple synthetic building blocks.
[0040] In another preferred embodiment, A1, O1, O2, P1, and P2 are each independently a chemical structure or chemical bond selected from the group consisting of:
[0041]
[0042] In the formula, It is an aromatic ring or an aromatic heterocyclic ring.
[0043] In another preferred embodiment, A2, A3, G1, and G2 are each independently selected from the group consisting of: H, -N3, aldehyde, hydroxyl, carboxyl, terminal alkene, terminal alkyne, chlorine, bromine, iodine, substituted or unsubstituted mercapto, substituted or unsubstituted -S-SH, substituted or unsubstituted phenyl, substituted or unsubstituted 5-7 heteroaryl, substituted or unsubstituted amino, substituted or unsubstituted -NH-C(O)-OH, substituted or unsubstituted -NH-C(O)H, substituted or unsubstituted C 1-4 Alkyl, substituted or unsubstituted C 2-6 alkenyl, substituted or unsubstituted C 2-6 Alkyne, substituted or unsubstituted C 3-7 Cycloalkyl, 4-7 membered cycloalkenyl, substituted or unsubstituted C 5-9 Cycloalkynyl, substituted or unsubstituted -C(O)O-C1-C6 alkyl, substituted or unsubstituted sulfonamide;
[0044] The substitution refers to the substitution of one or more (1, 2, 3, 4, 5, 6) hydrogen atoms on a group by a substituent selected from the group consisting of: oxo (=O), cyano, halogen, -SO3Na, substituted or unsubstituted C atoms. 1-6 Alkyl, C 1-4 Alkoxy, C 2-6 Alkenyl, substituted or unsubstituted 4-7 membered heterocyclic alkenyl, substituted or unsubstituted C 2-6 Alkynyl, substituted or unsubstituted phenyl, substituted or unsubstituted benzyl, substituted or unsubstituted 4-7 membered heterocyclic, substituted or unsubstituted 6-8 membered heterocyclic alkynyl, substituted or unsubstituted 5-7 membered heteroaryl, halogenated C 1-4 Alkyl, halophenyl; wherein the substitution refers to substitution by one or more (1, 2, 3, 4, 5, 6) substituents selected from the group consisting of: halogen, cyano, oxo (=O), -SO3Na, -SO2(C 1-6 Alkyl), C 1-4 Alkoxy, C 1-4 Alkoxy, nitro, phenyl, C 1-4Alkoxy-substituted phenyl, C 1-6 Amide group.
[0045] In another preferred example, A2, A3, G1, and G2 are each independently selected from the following group:
[0046]
[0047]
[0048] Where X is F, Cl, Br or I;
[0049] It is an aromatic ring or an aromatic heterocyclic ring.
[0050] In another preferred embodiment, It is a C6-20 aromatic ring or a 5-20 membered aromatic heterocycle.
[0051] In another preferred embodiment, It is a C6-10 aromatic ring or a 5-9 membered aromatic heterocycle.
[0052] In another preferred embodiment, The substitution sites of the substituents are ortho, meta, or para.
[0053] In another preferred embodiment, the Linker (or L) is a trivalent group formed by a combination of chemical elements and / or chemical bonds selected from the group consisting of:
[0054] (1) Chemical elements: C, H, O, N, P, S, Si; and
[0055] (2) Chemical bonds: C C, C = C, CY C = Y, YY, Y = Y; where Y is independently selected from H, O, N, P, S, and Si.
[0056] In another preferred embodiment, the Linker (or L) is selected from the group consisting of:
[0057]
[0058] Where q is 1, 2, 3, or 4.
[0059] In another preferred embodiment, the central structure of the synthetic building block is a structure composed of chemical elements and / or chemical bonds selected from the group consisting of:
[0060] (1) Chemical elements: C, H, O, N, P, S, Si, halogens; and
[0061] (2) Chemical bonds: C C, C = C, CY C = Y, YY, Y = Y; where Y is independently selected from H, O, N, P, S, and Si.
[0062] In another preferred embodiment, the synthetic building block is a natural amino acid or a non-natural amino acid.
[0063] In another preferred embodiment, in steps (c) and (d), the synthetic building blocks are linked to each other or to the linking groups (i.e., A2 and A3) in formula M0 via a reaction selected from the group consisting of: amidation, reductive amination, nucleophilic substitution, Suzuki coupling reaction, Heck coupling reaction, Sonogashira coupling reaction, S-aromatization reaction, CuAAC reaction, photoinduced reaction, preferably, via amidation reaction.
[0064] In a second aspect of the invention, a bidirectional cyclic peptide DNA-encoded compound of formula R is provided.
[0065]
[0066] In the formula,
[0067] N (or ) is a single-stranded / double-stranded deoxyribonucleotide sequence, a single-stranded / double-stranded ribonucleotide sequence, or a combination thereof;
[0068] A1 is a chemical bond or chemical structure connecting the Linker and N;
[0069] Linker (or L) is a multifunctional linker framework;
[0070] m and n are each independent integers from 1 to 100;
[0071] Each O1 and O2 is an independent chemical bond or chemical structure;
[0072] For the central structure of the composite building blocks; and
[0073] P1 and P2 are chemical bonds or chemical structures, respectively.
[0074] In a third aspect, a method for constructing a bidirectional cyclic peptide DNA-encoded compound library is provided.
[0075] The bidirectional cyclic peptide DNA-encoded compound library contains t bidirectional cyclic peptide DNA-encoded compounds having the structure shown in Formula R below, where t is a positive integer ≥1000;
[0076]
[0077] in,
[0078] N (or A1, Linker (or L), m, n, O1, O2, P1, and P2 are as defined in the second aspect of this invention;
[0079] The method includes:
[0080] (a) Using the method of claim 1, synthesizing the bidirectional cyclic peptide DNA-encoded compound of formula R; and
[0081] (b) Combining t compounds of formula R to construct the nucleic acid-encoded compound library.
[0082] In another preferred embodiment, t is ≥10,000, more preferably ≥100,000, even more preferably ≥1,000,000, even more preferably ≥5,000,000, and most preferably ≥10,000,000.
[0083] In a fourth aspect of the present invention, a bidirectional cyclic peptide DNA-encoded molecular library is provided, the bidirectional cyclic peptide DNA-encoded molecular library comprising t bidirectional cyclic peptide DNA-encoded compounds having the structure shown in Formula R below, where t is a positive integer ≥1000;
[0084]
[0085] in,
[0086] N (or A1, Linker (or L), m, n, O1, O2, P1, and P2 are defined as in the second aspect of this invention.
[0087] It should be understood that, within the scope of this invention, the above-described technical features of this invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be described in detail here. Attached Figure Description
[0088] Figure 1 The mass spectra of the product in step (1) of Example 1 are shown.
[0089] Figure 2 The mass spectra of the product in step (2) of Example 1 are shown.
[0090] Figure 3 The mass spectra of the product in step (3) of Example 1 are shown.
[0091] Figure 4 The mass spectra of the product in step (4) of Example 1 are shown.
[0092] Figure 5 The mass spectra of the product in step (5) of Example 1 are shown. Detailed Implementation
[0093] Through extensive and in-depth research, the inventors have developed a novel bidirectional synthetic method for the first time, enabling the efficient construction of large-size cyclic peptide DNA-coding molecular libraries. The method provided by this invention is simple to operate, operates under mild conditions, has high substrate universality, and has a wide range of applications, thus promoting the development of the medical field. Based on these findings, the inventors completed this invention.
[0094] the term
[0095] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0096] As used herein, the terms “containing” or “including (comprise)” can be open-ended, semi-closed, or closed. In other words, the terms also include “consistently made of” or “composed of”.
[0097] Group definition
[0098] Definitions of standard chemical terms can be found in the references (including Carey and Sundberg, "Advanced Organic Chemistry 4th Edition," Vols. A (2000) and B (2001), Plenum Press, New York). Unless otherwise stated, conventional methods within the scope of the art, such as mass spectrometry, NMR, IR, UV / VIS spectroscopy, and pharmacological methods, are used. Unless specifically defined, the terminology used herein in the relevant descriptions of analytical chemistry, organic synthetic chemistry, and pharmaceutical and medicinal chemistry is known in the art. Standard techniques can be used in chemical synthesis, chemical analysis, drug preparation, formulation and delivery, and in the treatment of patients. For example, reactions and purifications can be carried out using the manufacturer's instructions for use of kits, or in accordance with methods known in the art or the description of this invention. The techniques and methods described above can generally be carried out according to conventional methods well known in the art, based on the descriptions in the various summary and more specific references cited and discussed in this specification. In this specification, groups and their substituents can be selected by those skilled in the art to provide stable structural moieties and compounds.
[0099] In this document, unless otherwise specified, in all compounds of the present invention, each chiral carbon atom may optionally be in the R configuration or S configuration, or a mixture of R and S configurations. Unless otherwise specified, the structural formulas described in this invention are intended to include all isomers (e.g., enantiomers, diastereomers, and geometric isomers (or conformational isomers): for example, R and S configurations containing an asymmetric center, (Z) and (E) isomers of double bonds, and (Z) and (E) conformational isomers. Therefore, any single stereochemical isomer of the compounds of the present invention, or a mixture of its enantiomers, diastereomers, or geometric isomers (or conformational isomers), is within the scope of this invention.
[0100] "Substituted amino" refers to an amino group that is substituted by one or two alkyl, alkylcarbonyl, aromatic cycloalkyl, or heteroaromatic cycloalkyl groups as defined below, such as monoalkylamino, dialkylamino, alkylamide, aromatic cycloalkylamino, or heteroaromatic cycloalkylamino.
[0101] In this application, as a group or part of other groups (e.g., in halogen-substituted alkyl groups), the term "alkyl" refers to a fully saturated straight-chain or branched hydrocarbon chain group composed only of carbon and hydrogen atoms, having, for example, 1 to 12 (preferably 1 to 8, more preferably 1 to 6) carbon atoms, and connected to the rest of the molecule by single bonds, such as including but not limited to methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, sec-butyl, tert-butyl, n-pentyl, 2-methylbutyl, 2,2-dimethylpropyl, n-hexyl, heptyl, 2-methylhexyl, 3-methylhexyl, octyl, nonyl, and decyl. For the purposes of this invention, the term "alkyl" preferably refers to an alkyl group containing 1 to 6 carbon atoms.
[0102] In this application, as part of a group or other group, the term "alkenyl" refers to a straight or branched hydrocarbon chain group consisting only of carbon and hydrogen atoms, containing at least one double bond, having, for example, 2 to 14 (preferably 2 to 10, more preferably 2 to 6) carbon atoms connected to the rest of the molecule by single bonds, such as, but not limited to, vinyl, propenyl, allyl, but-1-enyl, but-2-enyl, pent-1-enyl, pent-1,4-dienyl, etc.
[0103] As used herein, or as part of other groups, the term "alkynyl" refers to a straight or branched hydrocarbon chain group consisting only of carbon and hydrogen atoms, containing at least one carbon-carbon triple bond, having, for example, 2 to 14 (preferably 2 to 10, more preferably 2 to 6) carbon atoms connected to the rest of the molecule by single bonds, such as, but not limited to, ethynyl, 1-propynyl, 1-butynyl, heptynyl, octyynyl, etc.
[0104] In this application, as part of a group or other group, the term "carbocyclic" refers to a stable non-aromatic monocyclic or polycyclic hydrocarbon group consisting only of carbon and hydrogen atoms. It may include fused ring systems, bridged ring systems, or spirocyclic systems, having 3 to 15 carbon atoms, preferably 3 to 10 carbon atoms, more preferably 3 to 8 carbon atoms, and may be saturated or unsaturated and can be linked to the rest of the molecule via a single bond through any suitable carbon atom. Preferably, the carbocyclic group is a cycloalkyl group. The cycloalkyl group is preferably a C3-10 cycloalkyl group. Unless otherwise specifically indicated in this specification, the carbon atoms in the carbocyclic group may optionally be oxidized. Examples of carbocyclic groups include, but are not limited to, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, cyclooctyl, adamantyl, 2,3-dihydroindenyl, octahydro-4,7-methylene-1H-indenyl, 1,2,3,4-tetrahydro-naphthyl, 5,6,7,8-tetrahydro-naphthyl, cyclopentenyl, cyclohexenyl, cyclohexadienyl, 1H-indenyl, 8,9-dihydro-7H-benzocyclohepten-6-yl, and 6,7,8,9-tetrahydro-5H-benzocyclo Heptenyl, 5,6,7,8,9,10-hexahydro-benzocyclooctenyl, fluorenyl, bicyclo[2.2.1]heptyl, 7,7-dimethyl-bicyclo[2.2.1]heptyl, bicyclo[2.2.1]heptyl, bicyclo[2.2.2]octyl, bicyclo[3.1.1]heptyl, bicyclo[3.2.1]octyl, bicyclo[2.2.2]octenyl, bicyclo[3.2.1]octenyl and octahydro-2,5-methylene-benzocyclopentadienyl, etc.
[0105] In this application, as part of a group or other group, the term "heterocyclic group" means a stable 3- to 20-membered non-aromatic cyclic group consisting of 2 to 14 carbon atoms and 1 to 6 heteroatoms selected from nitrogen, phosphorus, oxygen, and sulfur. Unless otherwise specifically indicated in this specification, a heterocyclic group can be a monocyclic, bicyclic, tricyclic, or more ring system, which may include fused ring systems, bridged ring systems, or spirocyclic systems; the nitrogen, carbon, or sulfur atoms in the heterocyclic group may optionally be oxidized; the nitrogen atom may optionally be quaternized; and the heterocyclic group may be partially or fully saturated. The heterocyclic group may be connected to the remainder of the molecule via a carbon atom or a heteroatom and by a single bond. In heterocyclic groups containing fused rings, one or more rings may be aromatic or heteroaromatic groups as defined below, provided that the connection point with the remainder of the molecule is a non-aromatic ring atom. For the purposes of this invention, the heterocyclic group is preferably a stable 4- to 11-membered non-aromatic monocyclic, bicyclic, bridged, or spirocyclic group containing 1 to 3 heteroatoms selected from nitrogen, oxygen, and sulfur, and more preferably a stable 4- to 8-membered non-aromatic monocyclic, bicyclic, bridged, or spirocyclic group containing 1 to 3 heteroatoms selected from nitrogen, oxygen, and sulfur. Examples of heterocyclic groups include, but are not limited to: pyrrolidinyl, morpholinyl, piperazinyl, homopiperazinyl, piperidinyl, thiomorpholinyl, 2,7-diaza-spiro[3.5]nonane-7-yl, 2-oxa-6-aza-spiro[3.3]heptane-6-yl, 2,5-diaza-bicyclo[2.2.1]heptane-2-yl, aziridine, pyranyl, tetrahydropyranyl, thiaranyl, tetrahydrofuranyl, oxazinyl, dioxocyclopentyl, tetrahydroisoquinolinyl, decahydroisoquinolinyl, imidazolinyl, imidazoalkyl, quinazinyl, thiazoalkyl, isothiazyl, isoxazylalkyl, dihydroindolyl, octahydroindolyl, octahydroisoindolyl, pyrrolidinyl, pyrazolyl, phthalimide, etc.
[0106] As used herein, the term "aromatic ring (group)" as a group or part of another group refers to a conjugated hydrocarbon ring system group having 6 to 18 carbon atoms (preferably 6 to 10 carbon atoms). For the purposes of this invention, the aromatic ring (group) can be a monocyclic, bicyclic, tricyclic, or more cyclic system, and can be fused with carbocyclic or heterocyclic groups as defined above, provided that the aromatic ring (group) is connected to the rest of the molecule via single bonds through atoms on the aromatic ring. Examples of aromatic rings (groups) include, but are not limited to, phenyl, naphthyl, anthraceneyl, phenanthryl, fluorenyl, 2,3-dihydro-1H-isoindolyl, 2-benzoxazolinone, 2H-1,4-benzoxazine-3(4H)-one-7-yl, etc.
[0107] As used herein, as part of a group or other group, the term "aromatic heterocyclic (group)" means a 5- to 16-membered conjugated cyclic group having 1 to 15 carbon atoms (preferably 1 to 10 carbon atoms) and 1 to 6 heteroatoms selected from nitrogen, oxygen, and sulfur. Unless otherwise specifically indicated in this specification, aromatic heterocyclic (groups) can be monocyclic, bicyclic, tricyclic, or more ring systems, and can be fused with carbocyclic or heterocyclic groups as defined above, provided that the aromatic heterocyclic (group) is connected to the remainder of the molecule via single bonds through atoms on the heteroaromatic ring. The nitrogen, carbon, or sulfur atoms in the aromatic heterocyclic (group) may optionally be oxidized; the nitrogen atom may optionally be quaternized. For the purposes of this invention, the heteroaromatic ring (group) is preferably a stable 5- to 12-membered aromatic group containing 1 to 5 heteroatoms selected from nitrogen, oxygen and sulfur, more preferably a stable 5- to 10-membered aromatic group containing 1 to 4 heteroatoms selected from nitrogen, oxygen and sulfur, or a 5- to 6-membered aromatic group containing 1 to 3 heteroatoms selected from nitrogen, oxygen and sulfur. Examples of aromatic heterocyclic groups include, but are not limited to, thiophene, imidazolyl, pyrazolyl, thiazolyl, oxazolyl, oxadiazolyl, isoxazolyl, pyridinyl, pyrazinyl, pyrazinyl, benzimidazolyl, benzopyrazolyl, indolyl, furanyl, pyrrolithyl, triazolyl, tetrazolyl, triazinyl, inazinyl, isoindolyl, indazole, isoindazole, purine, quinolinyl, isoquinolinyl, diazonyl, naphthidyl, quinoxolinyl, pteridyl, carbazolyl, carbazolyl, phenanthridine, phenanthroxolinyl, acridineyl, phenazinyl, isothiazolyl, benzothiazolyl, benzothiophene. Oxatriazolyl, cyclolinyl, quinazolinyl, phenylthioyl, indene, o-diazaphenyl, isoxazolyl, phenoxazinyl, phenthiazinyl, 4,5,6,7-tetrahydrobenzo[b]thiophenyl, naphthopyridyl, [1,2,4]triazolo[4,3-b]pyrazine, [1,2,4]triazolo[4,3-a]pyrazine, [1,2,4]triazolo[4,3-c]pyrimidine, [1,2,4]triazolo[4,3-a]pyridine, imidazo[1,2-a]pyridine, imidazo[1,2-b]pyrazine, imidazo[1,2-a]pyrazine, etc.
[0108] Bidirectional cyclic peptide DNA-encoded compounds and their preparation methods
[0109] This invention provides a bidirectional cyclic peptide DNA-encoding compound with the structure shown in Formula R and a method for preparing the same. Typically, the method of this invention includes the following steps: under a nucleic acid-compatible reaction environment, reacting the two reactive groups A2 and A3 in the compound of Formula M0 with a synthetic building block of Formula M1 containing reactive groups (G1 and G2), respectively, to obtain the compound shown in Formula Q; reacting the two reactive groups in the synthetic building block of Formula M1 with one reactive group present on a different synthetic building block of the compound shown in Formula Q, respectively, to obtain the bidirectional cyclic peptide DNA-encoding compound shown in Formula R.
[0110]
[0111] Specifically, a representative construction method includes the following steps:
[0112] (a) Provides a compound of formula M0, wherein the compound of formula M0 contains a Linker (L) having -A1-N segments, A2 and A3 attached thereto;
[0113]
[0114] (b) Provide the composite block shown in M1;
[0115]
[0116] (c) Reacting the compound represented by formula M0 with the synthetic building block represented by formula M1 to form a compound represented by formula Q1 having a nucleic acid fragment;
[0117]
[0118] (d) Reacting the compound shown in Formula Q1 with the synthetic building block shown in Formula M1 to form the compound shown in Formula Q2 containing a nucleic acid fragment:
[0119]
[0120] (e) Repeat reaction (d) 1 to r times to produce the nucleic acid-linked compounds shown in formulas Q3, Q4, and Q5:
[0121]
[0122] Among them, O1, O2, O3, O4, O5, O6, O7, O8, O9 and O 10 Each has an independent chemical structure that is the same or different (with at least two connection sites);
[0123] r is an integer from 2 to 100; preferably an integer from 3 to 50; even more preferably an integer from 3 to 5.
[0124] (f) The compounds represented by formulas Q1, Q2, Q3, Q4, and Q5 can be reacted with the synthetic building block represented by formula M1 to generate bidirectional cyclic peptide DNA-encoded compounds represented by formulas R1, R2, R3, R4, and R5, respectively.
[0125]
[0126]
[0127] In the formula, N (or A1, A2, A3, Linker (or L), m, n O1 and O2 are defined as above; O3, O4, O5, O6, O7, O8, O9, O 10 O 11 O 12 The definitions are similar to those of O1 and O2;
[0128] For example, O 11 The chemical bond formed between G2 and G1 through a nucleic acid-compatible chemical reaction; O 12 The chemical bond formed between G2 and G2 through a nucleic acid-compatible chemical reaction.
[0129] An exemplary reaction flow diagram provided by the present invention is shown below:
[0130]
[0131] R1 and R2 are not particularly limited and are determined by the specific compound selected. They are all groups or combinations thereof that are common in the art, such as alkyl, alkenyl, alkynyl, amino, cyano, thio, mercapto, hydroxyl, aryl, heteroaryl, cycloalkyl, heterocyclic, etc.
[0132] application
[0133] The bidirectional cyclic peptide DNA-encoding compound represented by Formula R of the present invention can also be used to prepare bidirectional cyclic peptide DNA-encoding molecular libraries for large-scale and high-throughput screening.
[0134] Typically, the high-throughput screening process for bidirectional cyclic peptide DNA-coding molecular libraries of the present invention includes:
[0135] (S1) A method for constructing a library of nucleic acid-compatible bidirectional cyclic peptide DNA-encoded compounds was developed.
[0136] (S2) The method developed in (S1) is used to prepare a bidirectional cyclic peptide DNA-encoded molecule library.
[0137] (S3) The molecular library prepared in (S2) is co-incubated with the target protein to collect bidirectional cyclic peptide DNA-encoded compounds that have a certain affinity for the target. The nucleotide sequences of these compounds are amplified by PCR and sequenced in high throughput to determine the specific structure of the collected compounds.
[0138] (S4) The coded compound collected in the chemical synthesis step (S3) is used to verify the activity of the compound through pharmacological experiments.
[0139] Compared with the prior art, the main advantages of the present invention include:
[0140] (1) A method for constructing a bidirectional cyclic peptide DNA-encoding molecular library under a nucleic acid-compatible environment was developed for the first time. The synthesis method is mild, easy to operate, has strong substrate universality, is inexpensive and readily available, and has a wide range of applications.
[0141] (2) This method was first applied to the construction of DNA-encoded molecular libraries, which expanded the size of macrocycles, increased the diversity of macrocycle structures, broadened the range of chemical spaces covered, and improved the probability of discovering lead compounds.
[0142] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Experimental methods in the following embodiments, unless otherwise specified, are generally performed under conventional conditions as described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or as recommended by the manufacturer. Unless otherwise stated, percentages and parts are by weight.
[0143] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as are familiar to those skilled in the art. Furthermore, any methods and materials similar to or equivalent to those described herein may be applied to the methods of this invention. The preferred embodiments and materials described herein are for illustrative purposes only.
[0144] Example 1: Preparation of a 96×96 DNA-encoded compound library
[0145] In this exemplary DNA-encoded compound library, a fixed-encoding nucleic acid sequence (referred to as the starting primer, customized by Suzhou Genewiz Biotechnology Co., Ltd., and purified by HPLC) is ligated into the starting fragment HP of the compound library.
[0146]
[0147] HP can be simplified to:
[0148] Starting primer on the upper strand: 5'-PO4 2- -AAATCGATGTG-3' (Serial Number: 01)
[0149] Starting primer lower strand: 5'-PO4 2- -CATCGATTTGG-3' (Serial Number: 02)
[0150] Step (1): Connect HP to Linker under nucleic acid compatibility conditions
[0151] HP was dissolved in 8 μL of sodium borate buffer (pH = 9.4, 250 mM). Then, 3 μL of 200 mM N-Fmoc-N'-Boc-L-2,3-diaminopropionic acid (dissolved in DMA, freshly prepared), 3 μL of 200 mM 2-(7-azabenzotriazole)-N,N,N',N'-tetramethylurea hexafluorophosphate (HATU) (dissolved in DMA, freshly prepared), and 3 μL of 200 mM N,N-diisopropylethylamine (DIPEA) (dissolved in DMA, freshly prepared) were added. The reaction was carried out at room temperature. After the reaction, 5 M sodium chloride aqueous solution and frozen ethanol were added. The mixture was incubated at -78°C for 0.5 hours, then centrifuged at 4°C. The supernatant was removed, and the resulting DNA precipitate was dissolved in 8 μL of phosphate buffer (pH = 5.5, 250 mM) and heated at 60°C for 12 hours. After the reaction, 5M sodium chloride aqueous solution and frozen ethanol were added, and the mixture was incubated at -78°C for 0.5 hours, followed by centrifugation at 4°C. The supernatant was removed, and the resulting DNA ligation product was dissolved in distilled water. The identification results are as follows: Figure 1 As shown.
[0152]
[0153] Step (2): The product from step (1) is ligated to the starting primer via an enzymatic reaction.
[0154] 600 μL of the 1 mM product from step (1) was mixed thoroughly with 368 μL of annealed starting primer aqueous solution (660 nmol). T4 buffer and T4 ligase were added at 0 °C, and the mixture was reacted at 16 °C for 16 hours. After the reaction was complete, 300 μL of 5 M sodium chloride aqueous solution and 8500 μL of frozen ethanol were added. The mixture was incubated at -78 °C for 0.5 hours, then centrifuged at 4 °C. The supernatant was removed, and the resulting DNA precipitate was dissolved in 585 μL of distilled water. The identification results are as follows: Figure 2 As shown.
[0155]
[0156] Step (3): Synthesis of the first cycle of the DNA-encoded compound library
[0157] Ninety-six ligation reactions were set up. The solution of the product of step (2) with a concentration of 1 mM was dispensed into 96 consecutive wells of a 96-well plate, 5.5 μL per well. Then, 3.1 μL each of the upper and lower strands of the first cycle labeled nucleotide duplex (hereinafter referred to as the first cycle labeled nucleotide duplex, customized by Suzhou Genewiz Biotechnology Co., Ltd., purified by HPLC) with a concentration of 1.8 mM was added to each well.
[0158]
[0159] X is one of the four deoxyribonucleotides: A, T, C, and G.
[0160] Following step (2), add T4 buffer solution and T4 ligase. After the reaction is complete, precipitate with ethanol as described above to obtain DNA precipitate. Dissolve the DNA precipitate in 8 μL of sodium borate buffer solution (pH = 9.4, 250 mM), add 3 μL of 200 mM Fmoc-amino acid (dissolved in DMA, freshly prepared), 3 μL of 200 mM 2-(7-azabenzotriazole)-N,N,N',N'-tetramethylurea hexafluorophosphate (HATU) (dissolved in DMA, freshly prepared), and 3 μL of 200 mM N,N-diisopropylethylamine (DIPEA) (dissolved in DMA, freshly prepared). React at room temperature to ligate the corresponding small molecule compound from the first cycle.
[0161]
[0162] X is one of the four deoxyribonucleotides: A, T, C, and G.
[0163] After the reaction was complete, 5M sodium chloride solution and cold ethanol were added to the reaction solution. The mixture was incubated at -78°C for 0.5 hours, then centrifuged at 4°C to remove the supernatant. The resulting DNA precipitate was further dissolved in a 5% piperidine aqueous solution. After the reaction was complete, ethanol was used for precipitation as described above to obtain the precipitate. After the reaction was complete, all reaction solutions were mixed together, and the product was desalted and purified using a 500 μL 10K ultrafiltration tube (Amicon Ultra Centrifugal). The identification results are as follows: Figure 3 As shown.
[0164]
[0165] X is one of the four deoxyribonucleotides: A, T, C, and G.
[0166] The information on the small molecules and their corresponding nucleotide double strands in the first cycle is shown in Table A below:
[0167] Table A
[0168]
[0169]
[0170]
[0171]
[0172]
[0173] Step (4): Synthesis of the second cycle of the DNA-encoded compound library
[0174] The ligation of nucleic acid sequences and amino acid monomers in the second cycle is similar to step (3). After the reaction is complete, all reaction solutions are mixed together and ethanol precipitation is performed as described above. The resulting DNA precipitate is further dissolved in 200 μL of distilled water, and the product is desalted and purified using a 500 μL 10K ultrafiltration tube (Amicon Ultra Centrifugal). The identification results are as follows. Figure 4 As shown.
[0175]
[0176] X is one of the four deoxyribonucleotides: A, T, C, and G; among them, R1 and R2 are not particularly limited to common groups in the art and are determined by the specific compound being linked.
[0177] The specific information on the small molecules and their corresponding nucleotide double strands in the second cycle is shown in Table B below:
[0178] Table B
[0179]
[0180]
[0181]
[0182]
[0183]
[0184] Step (5): Nucleic acid-compatible cyclization reaction
[0185] The product of step (4) with a concentration of 1 mM was dissolved in 8 μL of sodium borate buffer solution (pH = 9.4, 250 mM), and 3 μL of 200 mM N-acetyl-S-(tert-butylthio)cysteine (dissolved in DMA, freshly prepared), 3 μL of 200 mM 2-(7-azabenzotriazole)-N,N,N',N'-tetramethylurea hexafluorophosphate (HATU) (dissolved in DMA, freshly prepared), and 3 μL of 200 mM N,N-diisopropylethylamine (DIPEA) (dissolved in DMA, freshly prepared) were added. The mixture was reacted at room temperature.
[0186]
[0187] X is one of the four deoxyribonucleotides: A, T, C, and G.
[0188] After the reaction was complete, 5M sodium chloride aqueous solution and frozen ethanol were added. The mixture was incubated at -78°C for 0.5 hours, then centrifuged at 4°C. The supernatant was removed, and the resulting DNA precipitate was dissolved in sodium borate buffer (pH 9.4, 250mM). 100mM tris(2-carboxyethyl)phosphonic acid hydrochloride (dissolved in distilled water, freshly prepared) and 100mM 1,4-di(bromomethyl)benzene (dissolved in DMA, freshly prepared) were added, and the reaction was carried out at room temperature. After the reaction was complete, 5M sodium chloride aqueous solution and frozen ethanol were added. The mixture was incubated at -78°C for 0.5 hours, then centrifuged at 4°C. The supernatant was removed, and the resulting DNA precipitate was dissolved in distilled water. The product was desalted and purified using a 500μL 10K ultrafiltration tube (Amicon Ultra Centrifugal). The resulting DNA-encoded compound library (identification results are shown in the figure) was obtained. Figure 5 (As shown) is used in subsequent protein affinity screening.
[0189]
[0190] X is one of the four deoxyribonucleotides: A, T, C, and G.
[0191] All documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference. Furthermore, it should be understood that after reading the foregoing teachings of this invention, those skilled in the art can make various alterations or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims.
Claims
1. A method for preparing a bi-directional cyclic peptide-type DNA encoded compound, characterized in that, comprising the steps of: (a) providing a compound of Formula M0, wherein the compound of Formula M0 comprises a Linker (L), A2, and A3, wherein the Linker is attached to -A1-N; wherein, N (or ) is a single-stranded / double-stranded deoxyribonucleotide sequence, a single-stranded / double-stranded ribonucleotide sequence, or a combination thereof; A1 is a chemical bond or a chemical structure connecting the Linker and N; Linker (or L) is a multifunctional linker backbone; A2 and A3 are reactive groups; (b) providing a synthetic building block of Formula M1; wherein, To synthesize the center structure of the block; G1, G2 are reactive groups; (c) reacting the compound of Formula M0 with the synthetic building block of Formula M1 to form a compound of Formula Q comprising a nucleic acid fragment; wherein, m, n are each independently an integer from 1 to 100; each O1 and O2 is independently a chemical bond or a chemical structure; N (or ), A1, Linker (or L), G2 is as defined above; (d) ring-closing reaction of the compound of Formula Q with the synthetic building block of Formula M1 to form a compound of Formula R comprising a nucleic acid fragment: wherein, P1 and P2 are each a chemical structure or a chemical bond formed by the reaction of two reactive groups G2 on the central structure of the compound of Formula Q with two reactive groups (G1 and G2) on the synthetic building block of Formula M1; N (or ), A1, Linker (or L), m, n, O1, O2are as defined above.
2. The method of claim 1, wherein, A1, O1, O2, P1, P2 are each independently a chemical structure or a chemical bond selected from the group consisting of: in which is an aromatic or heteroaromatic ring.
3. The method of claim 1, wherein, A2, A3, G1, and G2 are each independently selected from the following groups: H, -N3, aldehyde, hydroxyl, carboxyl, terminal alkene, terminal alkyne, chlorine, bromine, iodine, substituted or unsubstituted mercapto, substituted or unsubstituted -S-SH, substituted or unsubstituted phenyl, substituted or unsubstituted 5-7 heteroaryl, substituted or unsubstituted amino, substituted or unsubstituted -NH-C(O)-OH, substituted or unsubstituted -NH-C(O)H, substituted or unsubstituted C 1-4 Alkyl, substituted or unsubstituted C 2-6 alkenyl, substituted or unsubstituted C 2-6 Alkyne, substituted or unsubstituted C 3-7 Cycloalkyl, 4-7 membered cycloalkenyl, substituted or unsubstituted C 5-9 Cycloalkynyl, substituted or unsubstituted -C(O)O-C1-C6 alkyl, substituted or unsubstituted sulfonamide; The substituent is one or more (1, 2, 3, 4, 5, 6) hydrogens on the group being replaced by a substituent selected from the group consisting of oxo (=0), cyano, halogen, -SO3Na, substituted or unsubstituted C 1-6 alkyl, C 1-4 alkoxy, C 2-6 alkenyl, substituted or unsubstituted 4-7 membered heterocycloalkenyl, substituted or unsubstituted C 2-6 alkynyl, substituted or unsubstituted phenyl, substituted or unsubstituted benzyl, substituted or unsubstituted 4-7 membered heterocyclyl, substituted or unsubstituted 6-8 membered heterocycloalkynyl, substituted or unsubstituted 5-7 membered heteroaryl, halogenated C 1-4 alkyl, halogenated phenyl; wherein the substituent is one or more (1, 2, 3, 4, 5, 6) hydrogens on the group being replaced by a substituent selected from the group consisting of halogen, cyano, oxo (=0), -SO3Na, -SO2(C 1-6 alkyl), C 1-4 alkoxy, C 1-4 alkoxy, nitro, phenyl, C 1-4 alkoxy-substituted phenyl, C 1-6 amido.
4. The method of claim 1, wherein, A2, A3, G1, and G2 are each independently selected from the group consisting of: wherein, X is F, Cl, Br, or I; is an aromatic or heteroaromatic ring.
5. The method of claim 1, wherein, the Linker (or L) is selected from the group consisting of: wherein, q is 1, 2, 3, or 4.
6. The method of claim 1, wherein, In steps (c) and (d), the synthetic building blocks or the synthetic building block and the linking groups (i.e., A2 and A3) in Formula M0 are connected by a reaction selected from amidation, reductive amination, nucleophilic substitution, Suzuki coupling, Heck coupling, Sonogashira coupling, S- aromatization, CuAAC reaction, photoinduced reaction, preferably by amidation.
7. A bidirectional cyclic peptide type DNA encoded compound of Formula R, wherein, N (or ) is a single-stranded / double-stranded deoxyribonucleotide sequence, a single-stranded / double-stranded ribonucleotide sequence, or a combination thereof; A1 is a chemical bond or a chemical structure connecting the Linker and N; Linker (or L) is a multifunctional linker backbone; m, n are each independently an integer from 1 to 100; each O1 and O2 is independently a chemical bond or a chemical structure; To synthesize the center structure of the block; and P1 and P2 are each a chemical bond or a chemical structure.
8. A method for constructing a library of bidirectional cyclic peptide type DNA encoded compounds, characterized in that, the library of bidirectional cyclic peptide type DNA encoded compounds comprises t bidirectional cyclic peptide type DNA encoded compounds of Formula R, t is a positive integer ≥ 1000; wherein, N (or ), A1, Linker (or L), m, n, O1, O2, P1, P2, G2 are as defined in claim 7; the method comprises: (a) synthesizing a bidirectional cyclic peptide type DNA encoded compound of Formula R using the method of claim 1; and (b) combining t compounds of Formula R to construct the library of nucleic acid encoded compounds.
9. The method of claim 8, wherein, t is > 10,000, preferably > 100,000, more preferably > 1000,000, more preferably > 5000,000, most preferably > 10,000,000.
10. A library of bi-directional cyclic peptide-like DNA encoded molecules, characterized in that, The bidirectional cyclic peptide type DNA coded molecular library comprises t bidirectional cyclic peptide type DNA coded compounds having the following formula R structure, t is a positive integer > 1000; wherein, N (or ), A1, Linker (or L), m, n, O1, O2, P1, P2 are as defined in claim 7.