Combination synthesis of oligonucleotide-directed and recorded coding probe molecules

CN115820806BActive Publication Date: 2026-08-11INSITRO INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2017-06-08
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

现有方法仍然受到合成探针分子的连续反应步骤的低效率的影响

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115820806B_ABST
    Figure CN115820806B_ABST
Patent Text Reader

Abstract

This disclosure relates to the combinatorial synthesis of oligonucleotide-guided and recorded coding probe molecules, and to multifunctional molecules comprising molecules of formula (I): (I)([(B1)M-D-L1]Y-H1)O-G-(H2-[L2-E-(B2)K]W)P, wherein G, H1, H2, D, E, B1, B2, M, K, L1, L2, O, P, Y, and W are defined herein. This disclosure also relates to methods for preparing these multifunctional molecules and methods for using these multifunctional molecules to identify coding molecules capable of binding to target molecules.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the international application filed on June 8, 2017, with international application number PCT / US2017 / 036582, which entered the Chinese national phase on December 7, 2018, application number 201780035706.3, and invention title "Combinatorial Synthesis of Oligonucleotide-Directed and Recorded Encoding Probe Molecules".

[0002] Cross-referencing

[0003] This application claims priority to U.S. Provisional Application Serial No. 62 / 351,046, filed June 16, 2017, which is incorporated herein by reference in its entirety. Technical Field

[0004] This invention relates to multifunctional molecules, and methods for preparing and using such multifunctional molecules. The invention also provides methods for using said multifunctional molecules to identify coding molecules capable of binding to target molecules or possessing other desired properties such as target molecule selectivity or cell permeability. Background Technology

[0005] There are essentially three methods for discovering molecules with desired functions, such as drugs. These are found in nature, are rationally designed, and are discovered through trial and error. In many cases, trial and error is arguably the most promising, but it can be extremely inefficient. The key to making trial and error more effective is creating combinatorial libraries of molecules that can be synthesized in large quantities and tested for the desired properties. The need to efficiently discover new molecules through trial and error has led to the rise of combinatorial chemistry.

[0006] There are three main problems with synthesizing and testing combinatorial libraries. First, many methods for preparing probe molecules from combinatorial libraries are limited by the type and number of sequential chemical subunits or structural units that can be assembled. Second, many methods for assembling sequential structural units are limited by the reaction efficiency of each step. Third, it should be understood that, to maintain efficiency, a large number of probe molecules should be tested simultaneously to determine if they possess the desired properties. It should also be understood that a library with sufficient diversity of molecular shapes may only contain a few copies of any given molecule. Low copy numbers hinder the identification of probe molecules with the desired properties. Therefore, each probe molecule should be labeled with a unique identifier so that researchers can identify the desired probe molecule.

[0007] Researchers have developed DNA-encoded probe molecules to address some of these problems. Some researchers use DNA oligonucleotides as templates to guide one or more steps in combinatorial synthesis. Others use DNA oligonucleotides to record combinatorial synthesis and uniquely label probe molecules, enabling the identification of molecular probes that remain bound to target molecules using PCR (polymerase chain reaction) amplification. Furthermore, other researchers use DNA oligonucleotides to guide one or more steps in combinatorial synthesis and label probe molecules with unique markers.

[0008] Despite the success of many of these methods, several problems remain. Existing methods are still hampered by the inefficiency of the sequential reaction steps in synthesizing probe molecules. They also struggle to detect probe molecules that are tightly bound to the target molecule but present in low quantities. This difficulty can lead to false negatives. A method for producing oligonucleotide probe molecules is needed to improve the efficiency of the sequential reaction steps. Furthermore, the ability to identify probe molecules that bind to the target molecule, even if these probe molecules may be present in low quantities, is also required. Summary of the Invention

[0009] This disclosure relates to encoded molecules. In some embodiments, the encoded molecule is a molecule of formula (I).

[0010] (I)([(B1) M —D—L1] Y —H1) O —G—(H2—[L2—E—(B2) K ] W ) P

[0011] in

[0012] G is an oligonucleotide comprising at least two coding regions and at least one terminal coding region, wherein the at least two coding regions are single-stranded and the at least one terminal coding region is single-stranded or double-stranded;

[0013] H1 is a hairpin structure containing an oligonucleotide, wherein H1 terminates at the 5' end and is attached to one end of the oligonucleotide G;

[0014] H2 is a hairpin structure containing an oligonucleotide, wherein H2 terminates at the 3' end and is attached to one end of the oligonucleotide G;

[0015] D is the first structural unit;

[0016] E is the second structural unit, where D and E may be the same or different;

[0017] B1 is a positional structure unit and M represents an integer from 1 to 20;

[0018] B2 is a positional structural unit and K represents an integer from 1 to 20, where B1 and B2 are the same or different, and M and K are the same or different;

[0019] L1 is a connector that operatively connects H1 to D;

[0020] L2 is the connector that operatively connects H2 to E;

[0021] O is an integer between 0 and 1;

[0022] P is an integer between 0 and 1;

[0023] The condition is that at least one of O and P is 1;

[0024] Y is an integer from 1 to 5;

[0025] W is an integer from 1 to 5; and

[0026] At least one of the positional structural unit B1 at position M and the positional structural unit B2 at position K is identified by one of the coding regions, and at least one of the first structural unit D and the second structural unit E is identified by the at least one end coding region.

[0027] In some embodiments of the molecule of formula (I), G comprises the components of formula (C). N —(Z N —C N+1 ) A The sequence is represented by (), where C is a coding region, Z is a non-coding region, N is an integer from 1 to 20, and A is an integer from 1 to 20; wherein each non-coding region contains 4 to 50 nucleotides and is optionally double-stranded. In some embodiments of the molecule of formula (I), one of O or P is 0. In some embodiments of the molecule of formula (I), at least one of Y and W is an integer from 1 to 2. In some embodiments of the molecule of formula (I), each coding region contains 6 to 50 nucleotides. In some embodiments of the molecule of formula (I), at least one of H1 and H2 contains 20 to 90 nucleotides. In some embodiments of the molecule of formula (I), each coding region contains 12 to 40 nucleotides. In some embodiments of the molecule of formula (I), P is 0, Y is an integer from 1 to 2, and each coding region contains 12 to 40 nucleotides. In some embodiments of the molecule of formula (I), O is 0, W is an integer from 1 to 2, and each coding region contains 12 to 40 nucleotides.

[0028] This disclosure relates to a method for identifying probe molecules capable of binding to or selecting target molecules, comprising exposing the target molecule to a pool of probe molecules, wherein the probe molecules are molecules of formula (I) above, removing at least one probe molecule that does not bind to the target molecule, amplifying at least one oligonucleotide G from at least one probe molecule that was never removed from the target molecule to form a copy sequence, sequencing at least one oligonucleotide G of the copy sequence to identify at least two coding regions of the probe molecule to further identify at least one of each positional structural unit B1 at position M and positional structural unit B2 at position K, and identifying at least one terminal coding region of the copy molecule to further identify at least one of a first structural unit D and a second structural unit E of the probe molecule.

[0029] This disclosure relates to a method for forming a formula (I) molecule. In some embodiments, the method for forming a formula (I) molecule includes:

[0030] At least one hybridization array is provided, the at least one hybridization array comprising at least one single-stranded anticodon oligomer immobilized on the at least one hybridization array, wherein the at least one single-stranded anticodon oligomer immobilized on the at least one hybridization array is capable of hybridizing with the coding region of the molecule of formula (II):

[0031] (II)([(B1) (M-1) —D—L1] Y —H1) O —G—(H2—[L2—E—(B2) (K-1) ] W ) P

[0032] in

[0033] G is an oligonucleotide comprising at least two coding regions and at least one terminal coding region, wherein the at least two coding regions are single-stranded and the at least one terminal coding region is single-stranded or double-stranded;

[0034] H1 is a hairpin structure containing an oligonucleotide, wherein H1 terminates at the 5' end and is attached to one end of the oligonucleotide G;

[0035] H2 is a hairpin structure containing an oligonucleotide, wherein H2 terminates at the 3' end and is attached to one end of the oligonucleotide G;

[0036] D is the first structural unit;

[0037] E is the second structural unit, where D and E may be the same or different;

[0038] B1 is a positional structure unit and M represents an integer from 1 to 20;

[0039] B2 is a positional structural unit and K represents an integer from 1 to 20, where B1 and B2 are the same or different, and M and K are the same or different;

[0040] L1 is a connector that operatively connects H1 to D;

[0041] L2 is the connector that operatively connects H2 to E;

[0042] O is an integer between 0 and 1;

[0043] P is an integer between 0 and 1;

[0044] The condition is that at least one of O and P is 1;

[0045] Y is an integer from 1 to 5;

[0046] W is an integer from 1 to 5; and

[0047] At least one of the positional structural unit B1 at position M and the positional structural unit B2 at position K is identified by one of the coding regions, and at least one of the first structural unit D and the second structural unit E is identified by the at least one end coding region.

[0048] The molecular pool of formula (II) is sorted into sub-pools by hybridizing the coding region of the sub-pool of the molecule with the at least one single-stranded anticodon oligomer fixed on the at least one hybridization array;

[0049] Optionally, the step of releasing the subpools of the molecule of formula (II) from the at least one hybridization array into a separate container;

[0050] Provide at least one of structural units B1 and B2; and

[0051] At least one of the structural units B1 and B2 is reacted with the molecule of formula (II) to form a sub-pool of the molecule of formula (I):

[0052] (I)([(B1) M —D—L1] Y —H1) O —G—(H2—[L2—E—(B2) K ] W ) P ,

[0053] in

[0054] G is an oligonucleotide comprising at least two coding regions and at least one terminal coding region, wherein each coding region is single-stranded and the at least one terminal coding region is single-stranded or double-stranded;

[0055] H1 is a hairpin structure containing an oligonucleotide, wherein H1 terminates at the 5' end and is attached to one end of the oligonucleotide G;

[0056] H2 is a hairpin structure containing an oligonucleotide, wherein H2 terminates at the 3' end and is attached to one end of the oligonucleotide G;

[0057] D is the first structural unit;

[0058] E is the second structural unit, where D and E may be the same or different;

[0059] B1 is a positional structure unit and M represents an integer from 1 to 20;

[0060] B2 is a positional structural unit and K represents an integer from 1 to 20, where B1 and B2 are the same or different, and M and K are the same or different;

[0061] L1 is a connector that operatively connects H1 to D;

[0062] L2 is the connector that operatively connects H2 to E;

[0063] O is an integer between 0 and 1;

[0064] P is an integer between 0 and 1;

[0065] The condition is that at least one of O and P is 1;

[0066] Y is an integer from 1 to 5;

[0067] W is an integer from 1 to 5; and

[0068] At least one of the positional structural unit B1 at position M and the positional structural unit B2 at position K is identified by one of the coding regions, and at least one of the first structural unit D and the second structural unit E is identified by the at least one end coding region.

[0069] In some embodiments of the method for forming a molecule of formula (I), the molecule of formula (II) is prepared by:

[0070] An oligonucleotide pool is provided, wherein the oligonucleotide G' comprises at least two coding regions and at least one terminal coding region, wherein the at least two coding regions are single-stranded, the at least one terminal coding region is single-stranded, and the at least one terminal coding region at the 5' and / or 3' ends of the oligonucleotide G' is different;

[0071] Provide at least one charged carrier anticodon, said at least one charged carrier anticodon having the formula ([(B1)]).(M-1) —D—L1] Y —H1) and / or (H2—[L2—E—(B2)) (K-1) ] W );

[0072] Combine the oligonucleotide G' pool with the at least one carrier anticodon;

[0073] At least one oligonucleotide G' is bonded to the 5' end of H1, and / or at least one oligonucleotide G' is bonded to the 5' end of H2, to form a molecular pool of formula (II):

[0074] (II)([(B1) (M-1) —D—L1] Y —H1) O —G—(H2—[L2—E—(B2) (K-1) ] W ) P , where G, H1, H2, D, E, B1, B2, L1, L2, O, P, Y and W are as defined in equation (I) above, and M and K are 1.

[0075] In some embodiments of the method for forming a molecule of formula (I), the method further includes removing a portion of an oligonucleotide from at least one terminal coding region of the molecule of formula (I) or (II). In some embodiments of the method for forming a molecule of formula (I), the method further includes linking H1 to G and linking H2 to at least one of G. In some embodiments of the method for forming a molecule of formula (I), G comprises oligonucleotides derived from formula (C). N —(Z N —C N+1 ) A The sequence is represented by (), where C is the coding region, Z is the non-coding region, N is an integer from 1 to 20, and A is an integer from 1 to 20; wherein each non-coding region contains 4 to 50 nucleotides and is optionally double-stranded. In some embodiments of the method of forming a molecule of formula (I), one of O or P is 0. In some embodiments of the method of forming a molecule of formula (I), at least one of Y and W is an integer from 1 to 2. In some embodiments of the method of forming a molecule of formula (I), at least one of H1 and H2 contains 20 to 90 nucleotides. In some embodiments of the method of forming a molecule of formula (I), P is 0, Y is an integer from 1 to 2, and each coding region contains 12 to 40 nucleotides; or O is 0, W is an integer from 1 to 2, and each coding region contains 12 to 40 nucleotides. Attached Figure Description

[0076] The foregoing summary of the invention and the following detailed description of its embodiments will be better understood when read in conjunction with the accompanying drawings. For illustrative purposes, some embodiments, which may be preferred, are shown in the drawings. It should be understood that the depicted embodiments are not limited to the precise details shown.

[0077] Figure 1 This is a diagram illustrating one implementation scheme of a method for preparing multifunctional molecules.

[0078] Figure 2 This is a diagram illustrating one implementation of a method for preparing multiple multifunctional molecules.

[0079] Figure 3 This is a flowchart illustrating one implementation of a method for forming a molecule of formula (I).

[0080] Figure 4 This is a flowchart illustrating one implementation of a method for forming a molecule of formula (I). Detailed Implementation

[0081] Unless otherwise stated, all measurements are in standard metric units.

[0082] Unless otherwise stated, all instances of the words “a” or “the” can refer to one or more words that are modified by them.

[0083] Unless otherwise stated, the phrase “at least one” refers to one or more objects. For example, “at least one of H1 and H2” means H1, H2, or both.

[0084] Unless otherwise stated, the term "about" refers to ±10% of the described, non-percentage number, rounded to the nearest integer. For example, about 100mm would include 90 to 110mm. Unless otherwise stated, the term "about" refers to ±5% of a percentage. For example, about 20% would include 15% to 25%. When the term "about" is discussed in relation to a range, it refers to an appropriate amount less than the lower limit and greater than the upper limit. For example, about 100 to about 200mm would include 90 to 220mm.

[0085] Unless otherwise stated, the terms "hybrid" and "hybridized" include Watson-Crick base pairings, which for DNA include guanine-cytosine and adenine-thymine (GC and AT) pairings, and for RNA include guanine-cytosine and adenine-uracil (GC and AU) pairings. For nucleotide complementary strands referred to as anticodons or anticoding regions, these terms are used in cases of selective nucleotide chain recognition.

[0086] The phrases “selective hybridization” and “selective sorting” refer to the selectivity of complementary strands relative to non-complementary strands being 5:1 to 100:1 or higher.

[0087] The term "multifunctional molecule" refers to a molecule of this disclosure containing an oligonucleotide and at least one coding motif.

[0088] The term "coding part" refers to one or more parts of a multifunctional molecule that contain only structural units, such as a first structural unit, a second structural unit, and positional structural units B1 and B2. The term "coding part" does not include hairpin structures or connectors, even though these structures may be added as part of the synthesis process of the coding part.

[0089] The term "coded molecule" refers to a molecule that is formed or created when the coding portion of a multifunctional molecule is removed or separated from the rest of the multifunctional molecule.

[0090] The term "probe molecule" refers to a molecule used to determine which coding portion of a multifunctional molecule or which coding molecule can bind to a target molecule or select desired properties such as target molecule selectivity or cell permeability.

[0091] The term "target molecule" refers to a molecule or structure. For example, structures include macromolecular complexes such as ribosomes and liposomes.

[0092] The term "probe molecule" can include multifunctional molecules.

[0093] The term "encoded probe molecule" is used interchangeably with the term "multifunctional molecule".

[0094] The term "polydisplay" refers to a multifunctional molecule that has at least two coded parts.

[0095] In this disclosure, hyphens or dashes in a molecular formula indicate that the parts of the formula are directly connected to each other by covalent bonds or hybridization.

[0096] Unless otherwise stated, all ranges of nucleotides and integer values ​​include all intermediate integers as well as endpoints. For example, a range of 5 to 10 oligonucleotides should be understood to include 5, 6, 7, 8, 9, and 10 nucleotides.

[0097] In some embodiments, this disclosure relates to a multifunctional molecule comprising at least one oligonucleotide moiety and at least one coding moiety, wherein the oligonucleotide moiety is synthesized using combinatorial chemistry to guide or encode the at least one coding moiety. In some embodiments, the oligonucleotide moiety of the multifunctional molecule can identify at least one coding moiety of the multifunctional molecule. In some embodiments, the multifunctional molecule comprises a hairpin structure attached to at least one end of the multifunctional molecule, wherein the hairpin structure allows multiple displays of multiple coding moieties of the multifunctional molecule. Not wishing to be bound by theory, it is believed that multiple displays of multiple coding moieties allow the multifunctional molecule of this disclosure to be more effectively selected to have desired properties, even if the multifunctional molecule may be present in low quantities, low concentrations, low relative concentrations compared to other probe molecules, or with a lesser degree of desired properties. In some embodiments, the multifunctional molecule of this disclosure comprises at least one oligonucleotide or oligonucleotide moiety comprising at least two coding regions and at least one terminal coding region, wherein the at least two coding regions and at least one terminal coding region correspond to and can be used to identify the sequence of structural units in the coding moiety. In some embodiments, at least one oligonucleotide or oligonucleotide moiety can be amplified by PCR to produce a copy of at least one oligonucleotide or oligonucleotide moiety, and the original or copy can be sequenced to determine the identity of at least two coding regions and at least one terminal coding region of the multifunctional molecule. In some embodiments, the identity of the at least two coding regions and at least one terminal coding region can be associated with a series of combinatorial chemical steps used to synthesize the coding region of the multifunctional molecule corresponding to the PCR copy.

[0098] In some embodiments, this disclosure also relates to methods for forming multifunctional molecules, and methods for exposing target molecules to multifunctional molecules to identify which coding portion and therefore which coding molecule exhibits desired properties, including but not limited to the ability to bind to one or more target molecules, the ability not to bind to other anti-target molecules, the ability to resist chemical changes caused by enzymes, the ability to be easily chemically altered by enzymes, the ability to have a degree of water solubility, and the ability to be permeable to cells.

[0099] In some embodiments, the molecule of formula (I) is a multifunctional molecule. In some embodiments of the molecule of formula (I), G is an oligonucleotide that directs or selects the encoding portion of the synthesis. In some embodiments of the molecule of formula (I), (B1) M —D and E—(B2) KEach represents a coding portion. In some embodiments of the molecule of formula (I), the molecule contains an oligonucleotide portion and at least one coding portion. It should be understood that many structural features of the oligonucleotide G in the present document are discussed in relation to guiding or encoding the synthesis of at least one coding portion of the molecule of formula (I). It should be understood that many structural features of the oligonucleotide G of formula (I) are discussed in relation to the ability of the oligonucleotide G or its PCR copy to identify the synthetic steps for preparing the molecule of formula (I), and thus the sequence and / or identity of the structural units and the chemical reactions for forming the coding portion of the molecule of formula (I).

[0100] In some embodiments of the molecule of formula (I), G comprises an oligonucleotide or an oligonucleotide. In some embodiments, the oligonucleotide contains at least two coding regions, wherein about 1% to about 100%, including about 50% to about 100%, including about 90% to about 100%, of the coding regions are single-stranded. In some embodiments, the oligonucleotide G contains at least one terminal coding region, one or both of which are single-stranded. In some embodiments, the oligonucleotide G contains at least one terminal coding region, one or both of which are double-stranded.

[0101] In some embodiments of the molecule of formula (I), the oligonucleotide G contains at least two coding regions, comprising 2 to about 21 coding regions, comprising 3 to 10 coding regions, or comprising 3 to 5 coding regions. In some embodiments, if the number of coding regions is less than 2, the number of possible coding portions that can be synthesized becomes too small to be practical. In some embodiments, if the number of coding regions exceeds 20, the low efficiency of synthesis interferes with accurate synthesis.

[0102] In some embodiments of the molecule of formula (I), at least two coding regions contain about 6 to about 50 nucleotides, including about 12 to about 40 nucleotides, including about 8 to about 30 nucleotides. In some embodiments, if the coding region contains less than about 6 nucleotides, the coding region cannot accurately guide the synthesis of the coding moiety. In some embodiments, if the coding region contains more than about 50 nucleotides, the coding region may become cross-reactive. This cross-reactivity can interfere with the ability of the coding region to accurately guide and identify the synthetic steps used to synthesize the coding moiety of the molecule of formula (I).

[0103] In some embodiments of the formula (I) molecule, the purpose of the oligonucleotide G is to guide the synthesis of at least one coding portion of the formula (I) molecule by selective hybridization with the complementary anticoding strand. In some embodiments, the coding region is single-stranded to facilitate hybridization with the complementary strand. In some embodiments, 70% to 100%, including 80% to 99%, including 80% to 95%, of the coding region is single-stranded. It should be understood that the complementary strand of the coding region (if present) may be added after the step encoding the coding portion of the formula (I) molecule during synthesis.

[0104] In some embodiments, the oligonucleotide may contain both natural and non-natural nucleotides. Suitable nucleotides include natural nucleotides of DNA (deoxyribonucleic acid), including adenine (A), guanine (G), cytosine (C), and thymine (T), and natural nucleotides of RNA (ribonucleic acid), including adenine (A), uracil (U), guanine (G), and cytosine (C). Other suitable bases include natural bases such as deoxyadenosine, deoxythymidine, deoxyguanosine, deoxycytidine, inosine, and diaminopurine; base analogs such as 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, C5-propynylcytidine, C5-propynyluridine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-methylcytidine, 7-deazoadenosine, 7-deazoguanosine, and 8-oxoadenosine. 8-O-guanosine, O(6)-methylguanine, 4-((3-(2-(2-(3-aminopropoxy)ethoxy)ethoxy)propyl)amino)pyrimidin-2(1H)-one, 4-amino-5-(hept-1,5-diyn-1-yl)pyrimidin-2(1H)-one, 6-methyl-3,7-dihydro-2H-pyrrolo[2,3-d]pyrimidin-2-one, 3H-benzo[b]pyrimido[4,5-e][1,4] Oligonucleotides can be azido-2(10H)-ones and 2-thiocytidines; modified nucleotides, such as 2'-substituted nucleotides, including 2'-O-methylated bases and 2'-fluorobases; and modified sugars, such as 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose; and / or modified phosphate ester groups, such as thiophosphates and 5'-N-phosphoramide bonds. It should be understood that oligonucleotides are polymers of nucleotides. The terms "polymer" and "oligopolymer" are used interchangeably herein. In some embodiments, oligonucleotides do not necessarily contain continuous bases. In some embodiments, oligonucleotides may be dispersed with linker portions or non-nucleotide molecules.

[0105] In some embodiments of the molecule of formula (I), the oligonucleotide G contains about 60% to 100%, including about 80% to 99%, of which about 80% to 95% is DNA nucleotide. In some embodiments, the oligonucleotide contains about 60% to 100%, including about 80% to 99%, of which about 80% to 95% is RNA nucleotide.

[0106] In some embodiments of the molecule of formula (I), the oligonucleotide G contains at least two coding regions, wherein the at least two coding regions overlap to extend together, provided that the overlapping coding regions share only about 30% to 1%, including about 20% to 1%, and include about 10% to 2% of the same nucleotides. In some embodiments of the molecule of formula (I), except for the terminal coding region, about 40% to 100%, including about 60% to 100%, and including about 80% to 100%, of the oligonucleotide G is single-stranded. In some embodiments of the molecule of formula (I), the oligonucleotide G contains at least two coding regions, wherein the at least two coding regions are adjacent. In some embodiments of the molecule of formula (I), the oligonucleotide G contains at least two coding regions, wherein the at least two coding regions are separated by nucleotide regions that do not direct or record the synthesis of the coding portion of the molecule of formula (I).

[0107] The term "non-coding region," when present, refers to an oligonucleotide region that cannot hybridize with the complementary strand of a nucleotide to direct the synthesis of the coding portion of a molecule of formula (I) or does not correspond to any anticoding oligonucleotide used to sort the molecule of formula (I) during synthesis. In some embodiments, the non-coding region is optional. In some embodiments, the oligonucleotide contains 1 to about 20 non-coding regions, including 2 to about 9 non-coding regions, and including 2 to about 4 non-coding regions. In some embodiments, the non-coding region contains about 4 to about 50 nucleotides, including about 12 to about 40 nucleotides, and including about 8 to about 30 nucleotides.

[0108] In some embodiments of the molecule of formula (I), one purpose of the non-coding region is to separate the coding region to avoid or reduce cross-hybridization, as cross-hybridization can interfere with the accurate coding of the coding portion of the molecule of formula (I). In some embodiments, one purpose of the non-coding region is to add functionality to the molecule of formula (I) beyond mere hybridization or coding. In some embodiments, one or more non-coding regions may be oligonucleotide regions modified with markers such as fluorescent or radiolabeled markers. These markers can facilitate the visualization or quantification of the molecule of formula (I). In some embodiments, one or more non-coding regions are modified with processing-enhancing functional groups or ligands. In some embodiments, one or more non-coding regions are double-stranded, which reduces cross-hybridization. In some embodiments, it should be understood that the non-coding region is optional. In some embodiments, suitable non-coding regions do not interfere with the PCR amplification of oligonucleotides.

[0109] In some embodiments, one or more of the coding region or terminal coding region may be a region of oligonucleotide G modified with a marker such as a fluorescent marker or a radiolabel. These markers can facilitate the visualization or quantification of the molecule of formula (I). In some embodiments, one or more of the coding region or terminal coding region is modified with processing-enhancing functional groups or chain ties.

[0110] In some embodiments of the molecule of formula (I), G includes the components of formula (C). N —(Z N —C N+1 ) A The sequence is represented by ), where C is the coding region, Z is the non-coding region, N is an integer from 1 to 20, and A is an integer from 1 to 20. In some embodiments, about 70% to 100%, including about 80% to 99%, of the non-coding region contains 4 to 50 nucleotides. In some embodiments, G includes about 70% to 100%, including about 80% to 99%, of the non-coding region being double-stranded.

[0111] In some embodiments of the molecule of formula (I), the oligonucleotide contains at least one, including one or two terminal coding regions. In some embodiments, the terminal coding regions are nucleotide sequences that do not directly bind to hairpin structures and terminate at the 5' or 3' end. In some embodiments, the terminal coding regions are nucleotide sequences that directly bind to hairpin structures. It should be understood that, based on the potential orientation of the nucleotides, the oligonucleotide will have 5' and 3' orientations, even if both ends of the oligonucleotide are bound by hairpin structures.

[0112] In some embodiments, one purpose of the terminal coding region is to facilitate selective hybridization of hairpin structures containing complementary sequences with the ends of oligonucleotides during the synthesis of the molecule of formula (I). In some embodiments, the terminal coding region contains about 6 to about 50 nucleotides, including about 12 to about 40 nucleotides, and about 8 to about 30 nucleotides. In some embodiments, if the terminal coding region contains less than about 6 nucleotides, the number of available non-cross-reactive sequences is too small, which may interfere with the accurate encoding of the coding portion of the molecule of formula (I). In some embodiments, if the terminal coding region contains more than about 50 nucleotides, the terminal coding region may become cross-reactive and lose too much specificity to selectively hybridize with only one hairpin structure. This cross-reactivity may interfere with the ability of the coding region to accurately encode the addition of the first structural unit D and / or the second structural unit E. In some embodiments of the molecule of formula (I), the terminal coding region is single-stranded or double-stranded.

[0113] In some embodiments of the molecule of formula (I), H1 and H2 are each independently a hairpin structure. As used in this disclosure, the term "hairpin structure" refers to a molecular structure containing 60% to 100% nucleotides by mass percentage and capable of hybridizing with the terminal coding region of oligonucleotide G. In some embodiments of the hairpin structure, the hairpin structure forms a single continuous polymer chain and contains at least one overlapping portion (commonly referred to as a "stem") containing a nucleotide sequence that hybridizes with a complementary sequence of the same hairpin structure. In some embodiments of the hairpin structure, a bridge structure connects two independent oligonucleotide chains; said bridge structure may comprise a polyethylene glycol (PEG) polymer of 2 to 20 PEG units, comprising 3 to 15 PEG units, comprising 6 to 12 PEG units. In some embodiments of the hairpin structure, the bridge structure may comprise an alkane chain of up to 30 carbons or a polyglycine chain of up to 20 units, or some other chain with reactive functional groups. In some embodiments of the molecule of formula (I), the overlapping portion of H1 and / or H2 is bound to or linked to the terminal coding region of oligonucleotide G. In some implementations, H1 and H2 each independently contain one, two, three, or four rings.

[0114] In some embodiments of the molecule of formula (I), H1 and H2 each independently comprise about 20 to about 90 nucleotides, about 32 to about 80 nucleotides, and about 45 to about 80 nucleotides. In some embodiments, H1 and H2 each independently contain 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, including 1 to 5, including 2 to 4, including 2 to 3 nucleotides modified with suitable functional groups to facilitate reaction with linker molecules or optionally with structural units, including cases where H1 and H2 are synthesized independently using bases such as, but not limited to, 5'-dimethoxytriphenylmethyl-5-ethynyl-2'-deoxyuridine, 3'-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramide (also known as 5-ethynyl-dU-CE phosphoramide, available from Glen Research, Sterling VA). In some embodiments, H1 and H2 each independently comprise nonnucleotides having suitable functional groups to facilitate reaction with linker molecules or optionally with structural units, including but not limited to 3-dimethoxytriphenylmethyloxy-2-(3-(5-hexylenamido)propamido)propyl-1-O-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramide (also known as alkyne modifier serinel phosphoramide, from Glen Research, Sterling VA), and baseless alkyne CEP (from IBA GmbH, Goettingen, Germany).In some embodiments, H1 and H2 each independently comprise nucleotides having modified bases with linkers, such as H1 and H2 independently synthesized using bases such as, but not limited to, 5'-dimethoxytriphenylmethyl-N6-benzoyl-N8-[6-(trifluoroacetylamino)-hex-1-yl]-8-amino-2'-deoxyadenosine-3'-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramide (also known as amino modifier C6 dA, available from Glen Research, Sterling VA), 5'-dimethoxytriphenylmethyl-N2-[6-(trifluoroacetylamino)-hex-1-yl]-2'-deoxyguanosine-3'-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramide (also known as amino modifier C6 dG, available from Glen Research, Sterling VA). Research, Sterling, VA), 5'-dimethoxytriphenylmethyl-5-[3-methacrylate]-2'-deoxyuridine, 3'-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramide (also known as carboxyl dT, purchased from Glen Research, Sterling, VA), 5'-dimethoxytriphenylmethyl-5-N-((9-fluorenylmethoxycarbonyl)-aminohexyl)-3-acryloimide]-2'-deoxyuridine, 3'-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramide (also known as Fmoc-amino modifier C6 dT, ...[3-meth-acrylate]-2'-deoxyuridine, 3'-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramide (also known as Fmoc-amino modifier C6 dT, Glen Research, Sterling, VA), 5'-dimethoxytriphenylmethyl-5-[3-meth-acrylate]-2'-deoxy Glen Research, Sterling, VA), 5'-dimethoxytriphenylmethyl-5-(oct-1,7-diynyl)-2'-deoxyuridine, 3'-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphamide (also known as C8 alkyne dT, Glen Research, Sterling VA), 5'-(4,4'-dimethoxytriphenylmethyl)-5-[N-(6-(3-benzoylthiopropionyl)-aminohexyl)-3-acrylamido]-2'-deoxyuridine, 3'-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphamide (also known as S-Bz-thiol modifier C6-dT, Glen Research, Sterling VA) and 5-carboxylic acid dC CEP (from IBA GmbH, Goettingen, Germany), N4-TriGl-amino2'-deoxycytidine (from IBA) GmbH, Goettingen, Germany). Functional groups applicable to modified nucleotides and non-nucleotides in H1 and H2 include, but are not limited to, primary amines, secondary amines, carboxylic acids, primary alcohols, esters, thiols, isocyanates, chloroformates, sulfonyl chlorides, thiocarbonates, heteroaryl halides, aldehydes, chloroacetic acid esters, aryl halides, halides, boric acids, alkynes, azides, and alkenes.

[0115] In some embodiments, one or more of the hairpin structures H1 and H2 may be modified with markers such as fluorescent or radioactive markers. These markers may facilitate the visualization or quantification of the molecule of formula (I). In some embodiments, one or more of the hairpin structures H1 and H2 may be modified with processing-enhancing functional groups or chain ties.

[0116] In some embodiments of the molecule of formula (I), the advantage of the hairpin structures of H1 and H2 is that one or both can allow multiple displays of multiple coding portions at one or both ends of the molecule of formula (I). Without wishing to be bound by theory, it is believed that multiple displays of multiple coding portions at one or both ends of the multifunctional molecules of this disclosure provide improved selectivity under certain conditions.

[0117] In some embodiments of the molecule of formula (I), Y and W are each independently 1, 2, 3, 4, or 5. In some embodiments, if O is 0, then W is an integer from 2 to 5, including 2, 3, 4, or 5. In some embodiments, if P is 0, then Y is an integer from 2 to 5, including 2, 3, 4, or 5. In some embodiments, if O and P are each 1, then W and Y are each independently an integer from 1 to 5, including 1, 2, 3, 4, or 5. In some embodiments, W and Y are each independently a measure or indication of multiple displays of the coded portion of the molecule of formula (I), wherein the coded portion should be understood as unit (B1). M —D and E—(B2) K Generally, the higher the aggregation values ​​of Y and W, the higher the multiple display of the molecule in formula (I).

[0118] In some embodiments of the molecule of formula (I), D is the first structural unit. In some embodiments, when D is present, D is encoded by or selected by the terminal coding region of G directly connected to H1. In some embodiments, the terminal coding region of G closest to the location of D corresponds to and can be used to identify the first structural unit D.

[0119] In some embodiments of the molecule of formula (I), E is the second structural unit. In some embodiments, when E is present, E is encoded by or selected by the terminal coding region of G directly connected to H2. In some embodiments, the terminal coding region of G closest to E corresponds to and can be used to identify the first structural unit E. In some embodiments, the first structural unit D and the second structural unit E may be the same or different. It should be understood that both the first and second structural units are structural units.

[0120] In some embodiments of the molecule of formula (I), B represents a positional structural unit. As used herein, the phrase "positional structural unit" refers to one of a series of individual structural units that are linked together as subunits forming a larger molecule. In some embodiments, (B1) M and (B2) K Each of these represents an independent structural unit that combines with other units to form a polymer chain having numbers M and K, respectively. For example, where M is 10, then (B) 10 This refers to the following structural unit chain: B 10 —B9—B8—B7—B6—B5—B4—B3—B2—B1. For example, when M is 3 and K is 2, equation (I) can be precisely expressed by the following equation:

[0121] ([((B1)3—(B1)2—(B1)1—D—L1] Y —H1) O —G—(H2—[L2—E—(B2)1—(B2)2] W ) P .

[0122] It should be understood that M and K each independently serve as location markers for each individual unit of B.

[0123] The precise definition of the term "structural unit" in this disclosure depends on its context. A "structural unit" is a chemical structural unit capable of chemically linking with other chemical structural units. In some embodiments, a structural unit has one, two, or more reactive chemical groups that allow the structural unit to undergo a chemical reaction that links it to other chemical structural units. It should be understood that some or all of the reactive chemical groups of a structural unit may be lost when it undergoes a reaction to form chemical bonds. For example, a structural unit in solution may have two reactive chemical groups. In this example, a structural unit in solution may react with the reactive chemical groups of a structural unit that is part of a chain of structural units to increase the length of the chain or to branch from the chain. When a structural unit is mentioned in the context of solution or as a reactant, it is understood to contain at least one reactive chemical group, but may contain two or more reactive chemical groups. When a structural unit is mentioned in the context of a polymer, oligomer, or a molecule larger than the structural unit itself, it is understood to have the structure of a structural unit that is a (monomer) unit of a larger molecule, even if one or more reactive chemical groups have reacted.

[0124] The types of molecules or compounds that can be used as structural units are generally unrestricted, as long as one structural unit can react with another to form a covalent bond. In some embodiments, the structural unit has a chemically reactive group to act as a terminal unit. In some embodiments, the structural unit has 1, 2, 3, 4, 5, or 6 suitable reactive chemical groups. In some embodiments, the first structural unit D, the second structural unit E, and the positional structural unit B each independently have 1, 2, 3, 4, 5, or 6 suitable reactive chemical groups. Suitable reactive chemical groups for structural units include primary amines, secondary amines, carboxylic acids, primary alcohols, esters, thiols, isocyanates, chloroformates, sulfonyl chlorides, thiocarbonates, heteroaryl halides, aldehydes, haloacetic esters, aryl halides, azides, halides, trifluoromethanesulfonates, dienes, dienophiles, boric acids, alkynes, and alkenes.

[0125] Any coupling chemistry can be used to link structural units, provided that the coupling chemistry is compatible with the presence of oligonucleotides. Exemplary coupling chemistry includes the formation of amides by reacting an amine (e.g., a DNA-linked amine) with an Fmoc-protected amino acid or other various substituted carboxylic acids; the formation of ureas by reacting an amine (including a DNA-linked amine) with an isocyanate and another amine (ureation); the formation of carbamates by reacting an amine (including a DNA-linked amine) with a chloroformate (carbamylation) and an alcohol; the formation of sulfonamides by reacting an amine (including a DNA-linked amine) with a sulfonyl chloride; the formation of thioureas by reacting an amine (including a DNA-linked amine) with a thiocarbonate and another amine (thioureation); the formation of aniline by reacting an amine (including a DNA-linked amine) with a heteroaryl halide (SNAr); the formation of secondary amines by reacting an amine (including a DNA-linked amine) with an aldehyde followed by reduction (reductive amination); the formation of peptides by acylation of an amine (including a DNA-linked amine) with a chloroacetate followed by replacement of the chloride with another amine (SN2 reaction); and the formation of peptides by reacting an amine (including a DNA-linked amine) with an aldehyde followed by reduction (reductive amination); and the formation of peptides by reacting an amine (including a DNA-linked amine) with a chloroacetate followed by replacement of the chloride with another amine (SN2 reaction). Acylation of amines (including DNA-linked amines) with carboxylic acids substituted with aryl halides, followed by substitution of the halide with a substituted alkyne (Sonogashira reaction), forms alkyne-containing compounds; acylation of amines (including DNA-linked amines) with carboxylic acids substituted with aryl halides, followed by substitution of the halide with a substituted boric acid (Suzuki reaction), forms biaryl compounds; reaction of amines (including DNA-linked amines) with cyanuric chloride, followed by reaction with another amine, phenol, or thiol (cyanuric acid acylation, aromatic substitution), forms substituted triazines; acylation of amines (including DNA-linked amines) with carboxylic acids substituted with suitable leaving groups such as halogens or trifluoromethanesulfonates, followed by substitution of the leaving group with another amine (SN2 / SN1 reaction), forms secondary amines; and cyclic compounds are formed by substituting amines with compounds containing alkenes or alkynes and reacting the product with azides or alkenes (Diehls-Alder and Huisgen reactions). In some embodiments of the reaction, the molecules reacting with the amine group include primary amines, secondary amines, carboxylic acids, primary alcohols, esters, thiols, isocyanates, chloroformates, sulfonyl chlorides, thiocarbonates, heteroaryl halides, aldehydes, chloroacetic esters, aryl halides, alkenes, halides, boric acids, alkynes, and alkenes, with a molecular weight of about 30 to about 330 Daltons.

[0126] In some embodiments of the coupling reaction, the first structural unit can be added by using any of the above-described chemicals with a secondary reactive group, such as amines, thiols, halides, boric acids, alkynes, or alkenes, to substituted amines (including DNA-linked amines). The secondary reactive group can then react with the structural unit having a suitable reactive group. Exemplary secondary reactive group coupling chemistry includes acylation of amines (including DNA-linked amines) with Fmoc-amino acids, followed by removal of protecting groups and reductive amination of the newly deprotected amine with aldehydes and borohydrides; reductive amination of amines (including DNA-linked amines) with aldehydes and borohydrides, followed by reaction of the substituted amine with cyanuric chloride, and then substitution of another chloride from the triazine with a thiol, phenol, or another amine; acylation of amines (including DNA-linked amines) with carboxylic acids substituted with heteroaryl halides, followed by an SNAr reaction with another amine or thiol to replace the halide and form aniline or thioether; acylation of amines (including DNA-linked amines) with carboxylic acids substituted with haloaryl groups, followed by substitution of the halide with an alkyne in a Sonogashira reaction; or substitution of the halide with an aryl group in a borate ester-mediated Suzuki reaction.

[0127] In some implementations, coupling chemistry is based on suitable bonding reactions known in the art. See, for example, March, Advanced Organic Chemistry, 4th Edition, New York: John Wiley and Sons (1992), Chapters 10–16; Carey and Sundberg, Advanced Organic Chemistry, Part B, Plenum (1990), Chapters 1–11; and Coltman et al., Principles and Applications of Organotransition Metal Chemistry, University Science Books, Mill Valley, CA (1987), Chapters 13–20; each of which is incorporated herein by reference in its entirety.

[0128] In some embodiments, in addition to one or more reactive groups for connecting the structural unit, the structural unit may also include one or more functional groups. One or more of these additional functional groups may be protected against undesirable reactions of these functional groups. Protecting groups suitable for a wide variety of functional groups are known in the art (Greene and Wuts, Protective Groups in Organic Synthesis, 2nd ed., New York: John Wiley and Sons (1991), which is incorporated herein by reference in its entirety). Particularly useful protecting groups include tert-butyl esters and ethers, acetals, triphenylmethyl ethers and amines, acetyl esters, trimethylsilyl ethers, trichloroethyl ethers and esters and carbamates.

[0129] The type of structural unit is generally unrestricted, as long as it is compatible with one or more reactive groups capable of forming covalent bonds with other structural units. Suitable structural units include, but are not limited to, peptides, carbohydrates, glycolipids, lipids, proteoglycans, glycopeptides, sulfonamides, nucleoproteins, ureas, carbamates, vinylogous polypeptides, amides, vinylogous sulfonamide peptides, esters, carbohydrates, carbonates, peptidyl phosphonates, azatides, peptide-like oligopeptides (oligomeric N-substituted glycine), ethers, ethoxymethyl acetal oligomers, thioethers, ethylene, ethylene glycol, disulfides, arylsulfides, nucleotides, morpholine, imines, pyrrolidones, ethyleneimine, acetates, styrene, acetylene, vinyl groups, phospholipids, siloxanes, isocyanates, isocyanates, and methacrylates. In some embodiments, formula (I) (B1) M Or (B2) K Each of these structural units, independently representing M or K units, includes polypeptides, polysaccharides, polyglycolipids, polylipids, polyproteoglycans, polyglycopeptides, polysulfonamides, polynuclear proteins, polyureas, polyurethanes, polyintercalated polypeptides, polyamides, polyintercalated sulfonamide peptides, polyesters, polysaccharides, polycarbonates, polypeptide-based phosphonates, polypolyacylhydrazides, polypeptides (oligomeric N-substituted glycine), polyethers, polyethoxyformaldehyde oligomers, polysulfides, polyethylene, polyethylene glycol, polydisulfide, polyaryl sulfides, polynucleotides, polymorpholine, polyimide, polypyrrolidone, polyethyleneimine, polyacetylene, polystyrene, polyacetylene, polyethylene, polyphospholipids, polysiloxanes, polyisocyanates, polyisocyanates, and polymethacrylates. In some embodiments of the molecule of formula (I), about 50% to about 100%, including about 60% to about 95%, and including about 70% to about 90% of the structural units having a molecular weight of about 30 to about 500 Daltons, including about 40 to about 350 Daltons, including about 50 to about 200 Daltons.

[0130] It should be understood that structural units with two reactive groups will form linear oligomeric or polymeric structures, or linear nonpolymeric molecules, containing each structural unit as a unit. It should also be understood that structural units with three or more reactive groups can form molecules with branches at each structural unit having three or more reactive groups.

[0131] In some embodiments of the molecule of formula (I), L1 and L2 each independently represent a linker. The term "linker molecule" refers to a molecule having two or more reactive groups capable of reacting to form a linker. The term "linker" refers to a molecular portion that operatively connects or covalently bonds a hairpin structure to a structural unit. The term "operatively connected" refers to two or more chemical structures connected or covalently bonded together in such a way that the connection is maintained during various operations (including PCR amplification) that the multifunctional molecule is expected to undergo.

[0132] In some embodiments of the molecule of formula (I), L1 is a connector that operatively links H1 to D. In some embodiments of the molecule of formula (I), L2 is a connector that operatively links H2 to E. In some embodiments, L1 and L2 are each independently bifunctional molecules that link H1 to D by reacting one reactive functional group of L1 with a reactive group of H1 and another reactive functional group of L1 with a reactive functional group of D, and that link H2 to E by reacting one reactive functional group of L2 with a reactive group of H2 and another reactive functional group of L2 with a reactive functional group of E. In some embodiments of the molecule of formula (I), L1 and L2 are each independently formed by reacting chemically reactive groups of H1 and D or H2 and E with a commercially available connector molecule, said commercially available connector molecule including PEG (e.g., azido-PEG-NHS, or azido-PEG-amine, or diazido-PEG), or alkane chain moiety (e.g., 5-azidopentanoic acid, (S)-2-(azidomethyl)-1-Boc-pyrrolidine, 4-azidoaniline, or 4-azido-but-1-acid N-hydroxysuccinimide); thiol reactive connectors, such as those that are PEG (e.g., SM(PEG)n). NHS-PEG-maleimide), alkane chains (e.g., 3-(pyridin-2-yldithio)-propionic acid-Osu or 6-(3'-[2-pyridinyldithio]-propamido)hexanoic acid sulfosuccinimide); and imides for oligonucleotide synthesis, such as amino modifiers (e.g., 6-(trifluoroacetylamino)-hexyl-(2-cyanoethyl)-(N,N-diisopropyl)-phosphoramide), and thiol modifiers (e.g., 5-triphenylmethyl-6-mercaptocyanide). Hexyl-1-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramide, or chemi-co-reactive modifiers (e.g., 6-hexyn-1-yl-(2-cyanoethyl)-(N,N-diisopropyl)-phosphoramide, 3-dimethoxytriphenylmethyloxy-2-(3-(3-propynyloxypropamido)propamido)propyl-1-O-succinoyl, long-chain alkylamino CPG, or 4-azido-but-1-acid N-hydroxysuccinimide); and compatible combinations thereof.

[0133] In some embodiments, the multifunctional molecule is a molecule of formula (IA), which is a subspecies of the molecule of formula (I):

[0134] (IA) [(B1) M —D—L1] Y —H1—G,

[0135] Wherein G, H1, D, B1, M, L1, and Y are as defined above with respect to formula (I). In some embodiments, the multifunctional molecule is a molecule of formula (IB), which is a subspecies of the molecule of formula (I):

[0136] (IB) [(B1) M —D—L1] Y —H1—G—H2,

[0137] Wherein G, H1, H2, D, B1, M, L1, and Y are as defined above with respect to formula (I). In some embodiments, the multifunctional molecule is a molecule of formula (IC), which is a subspecies of the molecule of formula (I):

[0138] (IC) [(B1) M —D—L1] Y —H1—G—H2—[L2] W ,

[0139] Wherein G, H1, H2, D, B1, M, L1, L2, W, and Y are as defined above with respect to formula (I). In some embodiments, the multifunctional molecule is a molecule of formula (ID), which is a subspecies of the molecule of formula (I):

[0140] (ID) [(B1) M —D—L1] Y —H1—G—H2—[L2—E] W ,

[0141] Wherein G, H1, H2, D, B1, E, M, L1, L2, W, and Y are as defined above with respect to formula (I). In some embodiments, the multifunctional molecule is a molecule of formula (IE), which is a subspecies of the molecule of formula (I):

[0142] (IE) G—H2—[L2—E—(B2) K ] W ,

[0143] Wherein G, H2, E, B2, K, L2, and W are as defined above with respect to formula (I). In some embodiments, the multifunctional molecule is a molecule of formula (IF), which is a subspecies of the molecule of formula (I):

[0144] (IF) H1—G—H2—[L2—E—(B2) K ] W

[0145] G, H1, H2, E, B2, K, L2, P, and W are as defined above with respect to formula (I). In some embodiments, the multifunctional molecule is a molecule of formula (IG), which is a subspecies of the molecule of formula (I):

[0146] (IG) [L1] Y —H1—G—H2—[L2—E—(B2) K ] W ,

[0147] Wherein G, H1, H2, E, B2, K, L1, L2, Y, and W are as defined above with respect to formula (I). In some embodiments, the multifunctional molecule is a molecule of formula (IH), which is a subspecies of the molecule of formula (I):

[0148] (IH) [D—L1] Y —H1—G—H2—[L2—E—(B2) K ] W ,

[0149] G, H1, H2, D, E, B, M, L1, L2, P, Y, and W are as defined above with respect to equation (I).

[0150] This disclosure relates to a method for synthesizing multifunctional molecules, including molecules of formula (I). In some embodiments of the method, at least one terminal coding region on an oligonucleotide, such as G', is capable of hybridizing with at least one carrier anticodon, said carrier anticodon comprising:

[0151] [(B1) (M-1) —D—L1] Y —H1 and / or H2—[L2—E—(B2)] (K-1) ] W ,

[0152] To form a molecule of formula (II):

[0153] (II)([(B1) (M-1) —D—L1] Y —H1) O —G—(H2—[L2—E—(B2) (K-1) ] W ) P ,

[0154] Wherein B1, M, D, L1, Y, H1, O, H2, L2, E, B2, K, W, and P are as defined above with respect to formula (I), and G' is an oligonucleotide containing at least two coding regions and at least one terminal coding region or containing an oligonucleotide containing at least two coding regions and at least one terminal coding region, wherein the at least two coding regions are single-stranded, the at least one terminal coding region is single-stranded, and at least one terminal coding region at the 5' and / or 3' ends of the oligonucleotide G' is different.

[0155] It should be understood that, [(B1)] (M-1) —D—L1] Y —H1 is (D—L1) Y —H1, where M is 1. It should be understood that H2—[L2—E—(B2)] (K-1) ] W It is H2—(L2—E) W , where K is 1.

[0156] like Figure 1 and Figure 3 As shown, in some embodiments, the advantage of this method of forming the molecule of formula (II) is that the terminal coding region of G' can encode or guide the addition of a first portion (including the first structural unit D and / or the second structural unit E) of the coding portion of the molecule. For example, each terminal coding region in the molecule of formula (I) will uniquely identify the first structural unit D and / or the second structural unit E, because the identities of the first structural unit D and / or the second structural unit E of the carrier anticodon capable of selectively hybridizing with the terminal coding region are known.

[0157] like Figure 2As shown, in some embodiments of the method for synthesizing the (I) molecule, the method uses a series of “sorting and reaction” steps, wherein a mixture of multifunctional molecules containing different combinations of coding regions is sorted into subpools by selective hybridization of one or more coding regions of the multifunctional molecule with an anti-coding oligomer immobilized on a hybridization array. In some embodiments of the method, the advantage of sorting the multifunctional molecules into subpools is that this separation allows each subpool to react with positional structural units B (including B1 and / or B2) under separate reaction conditions, and then the subpools of multifunctional molecules are combined or mixed for further chemical processing. In some embodiments of the method, the sorting and reaction process can be repeated to add a series of positional structural units. In some embodiments of the method, the advantage of using a sorting and reaction method to add structural units is that the identity of each positional structural unit of the coding portion of the molecule can be associated with the coding regions used to selectively separate or sort the multifunctional molecules prior to adding the structural units. In some embodiments, each coding region uniquely identifies the structural unit based on its position, because the identity of the coding region can be associated with the identity of the reaction process used to add each structural unit (which will include the identity of the positional structural unit being added). In some embodiments, the method can synthesize multifunctional molecules, including molecules of formula (I), wherein at least one of the first structural unit D and the second structural unit E is identified by or corresponds to at least one terminal coding region, and at least one of each positional structural unit B1 at position M and positional structural unit B2 at position K is identified by or corresponds to a coding region. It should be understood that molecules of formulas (I) and (II) may include one or more coding regions and terminal coding regions that are identical among different molecules in the pool; however, it should also be understood that the vast majority (if not all) of the molecules in the pool have different combinations of coding regions and terminal coding regions. In some embodiments of the method, the advantage of a molecular pool with different combinations of coding regions and terminal coding regions is that different combinations can encode multifunctional molecules with multiple different coding portions.

[0158] In some embodiments of the method for synthesizing the (I) molecule, the method includes the step of providing at least one carrier anticodon, wherein the at least one carrier anticodon has the formula:

[0159] ([(B1) (M-1) —D—L1] Y —H1) and / or (H2—[L2—E—(B2)) (K-1) ] W ),

[0160] Wherein B1, M, D, L1, Y, H1, H2, L2, E, B2, K, and W are as defined with respect to the molecule of formula (I). The term “provided” is generally unrestricted and may include synthetic or commercially available molecules. The term “carrier anticodon” refers to a hairpin structure operatively linked to a first or second structural unit via a linker and having an oligonucleotide anticodon region capable of selectively binding to at least one terminal coding region of oligonucleotide G’ or G. In some embodiments of the method, the purpose of the carrier anticodon is to allow the terminal coding region to encode, direct, or select the addition of the first and / or second structural units. In some embodiments of the method, the purpose of the carrier anticodon is to bind or link a structural unit comprising the first or second structural unit to at least one end of oligonucleotide G’ to form a molecule of formula (II), which will allow oligonucleotide G to encode or direct the synthesis of positional structural units B1 and / or B2.

[0161] In some embodiments of the method for synthesizing formula (I) molecules, the combination step is generally unrestricted, as long as the oligonucleotide G' pool and at least one carrier anticodon are allowed to interact or mix under conditions that allow selective hybridization.

[0162] In some embodiments of the method for synthesizing formula (I) molecule, the step of attaching the 5' end of at least one oligonucleotide G' to the 3' end of H1 includes selectively hybridizing the terminal coding region of the 5' end of G' with the anticodon of H1. In some embodiments of the method for synthesizing formula (I) molecule, the step of attaching the 3' end of at least one oligonucleotide G' to the 5' end of H2 includes selectively hybridizing the terminal coding region of the 3' end of G' with the anticodon of H2.

[0163] In some embodiments of the method for synthesizing molecule (I), the method includes linking the 5' end of at least one oligonucleotide G to the 3' end of H1 to form a covalent bond, and / or linking the 3' end of at least one oligonucleotide G to the 5' end of H2. This linking step can be performed during or after the formation of molecule (II). In some embodiments of the method, the benefit of forming the covalent bond during or after the linking step includes improved treatment during other chemical processing steps.

[0164] In some embodiments of the method for synthesizing the molecule of formula (I), the method may further include the step of removing all or part of the oligonucleotide from at least one terminal coding region of the molecule of formula (I) or (II) during or after the formation of the molecule of formula (II). In some embodiments of the method for synthesizing the molecule of formula (I), the method may further include the step of removing all or part of the oligonucleotide from at least one terminal coding region of the molecule of formula (I) or (II), wherein at least one terminal coding region is double-stranded and the oligonucleotide may be removed from the hairpin structure of H1 and / or H2, including removal from the anticodon of the hairpin structure of H1 and / or H2, or removal from the terminal coding region of G or G'. In a particular embodiment, the benefit of removing all or part of the oligonucleotide from at least one terminal coding region of the molecule of formula (I) or (II) may include improved chemical processing during subsequent steps.

[0165] like Figure 3 and Figure 4 As schematically depicted, in certain embodiments of the method for forming a molecule of formula (I), the molecule of formula (I) can be synthesized via multiple synthetic routes by changing the order of processing steps. For example, in Figure 3 In this process, two carrier anticodons can be added in the same step. Subsequent sorting and reaction steps can then be performed to add structural units B1 and B2 under the same reaction conditions. In this embodiment, (B1) M and (B2) K The encoding regions may be the same, where B1 = B2 and M = K.

[0166] Or, in Figure 4 In the middle, a carrier anticodon ([(B1)) (M-1) —D—L1] Y —H1) can combine and bond with G' to form a molecule of formula (I), where p is 0. Subsequent sorting and reaction steps can then be performed to add the structural unit B1. Once (B1) is complete... M The formation of the encoding part allows the second carrier anticodon (H2—[L2—E—(B2)) to be used. (K-1) ] W (B1) is combined and bonded with G to form different molecules of formula (I), where P is 1. Subsequent sorting and reaction steps can then be performed to add structural unit B2, using the same or different positional structural units and reaction conditions. In some embodiments, (B1) M and (B2) K The differences lie in the type and sequence of reactions used to select and sort molecular pools (II) and (I), respectively. It should be understood that in some embodiments, the carrier anticodon (H2—[L2—E—(B2)]) may be added first.(K-1) ] W (B2) K Then add ([(B1)) (M-1) —D—L1] Y —H1) to form (B1) M .

[0167] It should also be understood that it is not necessary to include (B1). M Or (B2) K The carrier anticodon is added after the complete formation or encoding of the carrier. Alternatively, the carrier anticodon can be added subsequently after at least one of B1 or B2 has been added to the previously added carrier anticodon. After the addition of the second carrier anticodon, subsequent sorting and reaction steps can be performed to add structural units B1 and B2 to form (B1). M Or (B2) K In these implementations, at least a portion of the first coded portion will match a subsequently added coded portion, but both coded portions (B1) M Or (B2) K The difference will at least lie in the positional structural units added before the carrier anticodon added after bonding.

[0168] In some embodiments, the method of forming a molecule of formula (I) includes making a molecule of formula (II):

[0169] (II)([(B1) (M-1) —D—L1] Y —H1) O —G—(H2—[L2—E—(B2) (K-1) ] W ) P ,

[0170] Reacts with one or more of the positional structural units B1 and / or B2 to form a molecule of formula (I):

[0171] (I)([(B1) M —D—L1] Y —H1) O —G—(H2—[L2—E—(B2) K ] W ) P ,

[0172] H1, H2, D, E, B1, M, B2, K, L1, L2, O, P, Y, and W are as defined with respect to equation (I).

[0173] In some embodiments, the method includes providing at least one hybridization array. The step of providing the hybridization array is generally unrestricted and includes manufacturing or commercially available hybridization arrays using techniques known in the art. In some embodiments of the method, the hybridization array comprises a substrate having at least two separate regions having immobilized anticodon oligomers on its surface. In some embodiments, each region of the hybridization array contains a different immobilized anticodon oligomer, wherein the anticodon oligomer is an oligonucleotide sequence capable of hybridizing with one or more coding regions of a molecule of formula (I) or (II). In some embodiments of the method, the hybridization array uses two or more chambers. In some embodiments of the method, the chambers of the hybridization array contain particles, such as beads, having immobilized anticodon oligomers on their surfaces. In some embodiments of the method, the advantage of immobilizing the molecule of formula (I) or (II) on the array is that this step allows for sorting or selective separation of molecules into molecular pools based on the specific oligonucleotide sequence of each coding region. In some embodiments, the separated molecular pools can then be individually released from the array or removed into a reaction chamber for further chemical processing. In some embodiments, the release step is optional and generally unrestricted, and may include dehybridizing the molecules by heating, using a denaturing agent, or exposing the molecules to a buffer solution with pH ≥ 12. In some embodiments, chambers or regions containing arrays of different immobilized oligonucleotides may be positioned to allow the contents of each chamber or region to flow into a series of pores for further chemical processing.

[0174] In some embodiments, the method includes reacting at least one structural unit B comprising B1 and / or B2 with a molecule of formula (II) to form a subpool of a molecule of formula (I) or (II), wherein B1 and / or B2 are as defined above with respect to formula (I). In some embodiments, structural units B1 and / or B2 may be added to the container before, during, or after the molecule of formula (I) or (II). It should be understood that the container may contain solvents and co-reactants under acidic, basic, or neutral conditions, depending on the coupling chemistry used to react structural units B1 and / or B2 with the molecule of formula (II) or (I).

[0175] A method for identifying probe molecules capable of binding to or selecting target molecules is disclosed. In some embodiments, the method includes exposing the target molecule to a pool of multifunctional molecules (e.g., molecules of formula (I)) to determine whether one of the multifunctional molecules is capable of binding the target molecule. In some embodiments, the term "exposure" includes any manner of contacting the target molecule with the probe molecules (including molecules of formula (I)). In some embodiments, probe molecules that do not bind to the target molecule are removed by a removal method comprising washing the unbound probe molecules off the target molecule with an excess solvent. In some embodiments, the target molecule is immobilized on a surface. In some embodiments, the target molecule includes proteins, enzymes, lipids, oligosaccharides, and nucleic acids having a tertiary structure.

[0176] In some embodiments of the method, the amplification step includes generating a copy sequence of oligonucleotide G of formula (I) using PCR techniques known in the art. In some embodiments of the method, the copy sequence contains copies of at least two coding regions of formula (I) and at least one terminal coding region of formula (I). In some embodiments, one benefit of amplifying oligonucleotide G from at least one probe molecule includes the ability to detect which coding portion of a multifunctional molecule is capable of binding to a target molecule, even if the multifunctional molecule cannot be easily removed from the target molecule. In some embodiments, the benefit of amplification is that it allows for the generation of libraries of molecules with great diversity. This great diversity comes at the cost of a low number of molecules of any given formula (I). PCR amplification allows the identification of oligonucleotide sequences present in very small quantities by increasing these numbers until easily detectable numbers are reached. DNA sequencing and analysis of the copy sequence can then identify coding portions of multifunctional molecules of formula (I) capable of binding to targets or associated with coding portions of multifunctional molecules of formula (I) capable of binding to targets.

[0177] More details about the challenge

[0178] The high cost of drug discovery and the growing need to discover molecules with unique and desirable properties for use in medicine, research, biotechnology, agriculture, food production, and industry have spurred the rise of the field of combinatorial chemistry.

[0179] For specific applications, discovering molecules with highly desired properties may not always be straightforward. For example, molecules that bind to target proteins or biomacromolecules or multimolecular structures can be very difficult to rationally design. When faced with the challenge of discovering molecules (for which first-principles structural design is impossible or inefficient), combinatorial chemistry itself has become a viable tool. Combinatorial chemistry achieves discovery through the following general process: (a) researchers make the best available hypotheses about the more general properties and structures that molecules may possess, in accordance with the criteria of the desired application; (b) researchers design and synthesize a very large number of molecules with the hypothesized general properties or structures, called a library; and (c) the library is tested to determine whether any library member possesses the properties required for the desired application.

[0180] In situations with limited information, the hypotheses that can be made about the structure of the desired molecule will be more loosely defined and less well-defined compared to situations where there is ample knowledge to inform these hypotheses. When more data informs structural hypotheses, libraries with less complexity or diversity, such as 1e... 4 -1e 7 A single member, closely focused on the chemical shape space region that the hypothesis is rich in the desired structure, may be more successful. In cases where little or no data is available, a library with far greater complexity, capable of sampling and deeply sampling a much larger shape space region (e.g., 1e) may be required. 5 -1e 14 (A unique member) to achieve success.

[0181] Combinatorial chemistry allows for the synthesis of compound libraries on this scale via separation-collection or sorting and reaction chemosynthesis methods. A typical separation-collection library begins with functional groups incorporated into the chains of a polymer solid support such as polystyrene beads. Thousands to millions of beads are grouped into a series of containers, and the beads in each container react with different chemical subunits or structural units. When the reaction is complete, all beads are collected, thoroughly mixed, and regrouped into a new series of reaction containers for a second step of chemosynthesis using the same or different groups of structural units. The separation-collection reaction process is repeated until synthesis is complete. The number of compounds prepared by this method is limited only by the number of beads that can be handled in the process and the number of structural units used in each step. These two parameters will limit the complexity of such a library. For example, if there are 5 chemical subunits in each of the 4 steps, then 5 4 =625 members will constitute the library. Similarly, if there are 52 structural units in the first step, 3 in the second step, 384 in the third step, and 96 in the fourth step, then 52 × 5 × 384 × 96 = 3,833,856 library members will be generated.

[0182] The molecular libraries can then be tested to determine which of them possesses the desired properties for the selected application. Identification of these molecules can be challenging because the molecular weights produced on a single bead can be very small, making identification difficult. It is generally understood in combinatorial chemistry that, for libraries appropriately tailored to the amount of structural data available to guide library design, larger libraries are expected to be more likely to have highly desired members. However, for any given amount of library produced, the greater the complexity, or the greater the number of unique molecules in the library, the lower the copy number, or copy number per member. Therefore, as the complexity of the library and the probability of having successful members increase, the total number of successful members and the ability of combinatorial chemists to correctly identify it decrease.

[0183] Combinatorial chemists then face the constraint of optimizing the library to have sufficient complexity to contain the desired members, while also preparing enough copies of each library member to ensure accurate identification of the desired members. Typically, as the complexity of the library increases, the size of the solid support must also decrease; as the size of the support decreases, the amount of sample available for analysis and identification also decreases.

[0184] Typically, given sufficient resources, 10 [units of something] can be synthesized on polystyrene beads from a single-bead-single-compound library. 10 A very large combined library of unique members. But if each polystyrene bead is a sphere with a volume of 0.1 microliters, then 10 10 A single member library would have a volume greater than 1 cubic meter, enough to fill a typical hot tub or spa, and perhaps even overflow. While industrial chemical processes are typically carried out on this scale, processes of this complexity are rarely performed at this scale. Libraries of this size also present challenges in testing them and generating molecular targets for those tests. Such testing could easily require one kilogram of purified protein, and the cost of producing so many drug targets would be astronomical for many drug targets.

[0185] DNA-encoded combinatorial chemistry libraries attempt to improve this situation. The fact that PCR can amplify single DNA template strands very accurately and in large quantities, and that the amplified strands can be easily sequenced, makes it possible to reduce the size of solid supports to a single DNA molecule. Therefore, by tethering combinatorial chemistry library members to DNA strands in a manner that establishes a correspondence between DNA sequences and library member identities, it is possible to prepare extremely large libraries (e.g., 10^6). 6 -10 14The library population is then divided into individual members (each possessing a unique trait) and the ability to identify successful molecules from the population. A selection experiment is then performed. "Selection" is an experiment that physically isolates members of the library population that possess the desired trait from those that do not. The DNA encoding the trait-positive library members is then amplified by PCR, and sequencing of the DNA identifies the trait-positive library members. In this way, libraries of immense complexity can be synthesized, and trait-positive individuals can be identified from very small sample sizes.

[0186] Able to return 10 6 -10 8 Novel DNA sequencing technologies, yielding unique sequences, significantly improve the analysis of DNA-encoded libraries. "Deep sequencing" data enables robust statistical analysis of highly complex chemical libraries. These types of analyses not only identify specific individual members of the library suitable for a chosen application but also reveal previously unknown general traits that assign "fitness" to library members for the application. Typically, deep sequencing of the DNA library precedes a selection experiment designed to physically separate nearby individuals more suited to the application from those less suited. The resulting population is then deep sequenced, and a comparison of the two datasets reveals which individuals are more suited due to their increased relative frequency within the population. Less suited individuals are identified because their relative frequency decreases within the population. However, DNA-encoded combinatorial chemistry approaches can yield libraries with a complexity far exceeding that of even the most powerful deep sequencing technologies currently available. While deep sequencing significantly improves the practicality and success rate of DNA-encoded combinatorial libraries, it still provides only statistically undersampling of theoretically usable data.

[0187] The problem of undersampling data is further complicated by the fact that not every step in a combinatorial chemistry process proceeds with perfect efficiency. Loss of fidelity is observed because some reactions are not completed and some produce byproducts. Therefore, the DNA sequences returned by deep sequencing do not always fully represent the actual molecules they encode, but may sometimes represent truncated products or products altered by side reactions.

[0188] Complicating the undersampling problem is the issue of synthetic fidelity. Not every reaction used to prepare combinatorial libraries is entirely efficient. This means that some DNA in a DNA-coding library is not tethered to the molecules it encodes, but rather to truncated products resulting from the incomplete incorporation of one or more structural units, or to similar compounds produced by incorporation byproducts or side reactions. Consequently, data analysis is affected because some genotypes surviving the selection are observed to represent molecules that are not encoded by them.

[0189] Directed Evolution

[0190] The purpose of this disclosure is to identify molecules suitable for desired applications by generating DNA-coding libraries with high complexity and fidelity, thereby maximizing current sequencing technologies and overcoming undersampling problems through multi-generational selection. Another purpose of this disclosure is to achieve more accurate synthesis by allowing for a more precise first step in the synthesis of purified libraries.

[0191] Typically, the method disclosed herein operates as follows. A molecular population is encoded in a pool of DNA genes (oligonucleotides G or G'). The DNA sequence (oligonucleotide G or a copy thereof) is then used as a template for the synthesis or translation of small molecule library members (coding regions), establishing a covalent linker between the DNA genotype and its corresponding small molecule phenotype. The genotype-phenotype fusion population (the molecular library of formula (I)) can then be subjected to selection pressure, and the DNA (at least two coding regions and at least two terminal coding regions) of these individuals that survive the selection is amplified by PCR. This second-generation survivor population (copy sequences) can be (a) deeply sequenced and analyzed to identify suitable individuals (coding regions with the desired characteristics), and / or (b) the population is retranslated for a second round of more rigorous selection for the same or different characteristics. The second-generation selected survivors are then (a) deeply sequenced and analyzed to identify suitable individuals and / or (b) retranslated and subjected to further rounds of selection, sequencing, and analysis until molecules with suitable fitness are obtained.

[0192] The ability of the method disclosed herein to generate libraries of formula (I) molecules that can be retranslated and reselected over multiple generations is a significant advantage because it allows for directed multigenerational “evolution.” While the initial population with sufficient complexity and the surviving population after the first round of selection may be undersampled by current best-in-class, affordable deep sequencing methods, and while statistical analysis of sequencing data for identifying fit individuals may also be undersampled, multigenerational selection enables a complete analysis of the population. Since many unique, less fit individuals will be eliminated from the population through successive rounds of selection, the actual complexity of the population will decrease as it becomes enriched with more fit individuals. Therefore, each round of sequencing will require increasingly important sampling, enabling very robust computational analysis. The ability to perform multigenerational selection greatly improves this analysis.

[0193] The DNA 'genes' or oligonucleotides G constituting the library have a somewhat familiar structure. Like genes in cells, the information is arranged in a linear sequence. However, unlike typical genes, in this synthetic biological system, the "codons," or coding regions, are typically about 20 bases long. In natural systems, codons are read in a linear order from one end of the gene to the other. In one embodiment of this disclosure, the coding regions containing the gene can be read in any order, provided that (a) the sequence begins with a terminal codon, (b) it is predetermined, and (c) the predetermined order is maintained in all successive progeny of the selection activity. Furthermore, the gene may optionally incorporate non-coding regions between each coding region. If the non-coding regions have unique restriction sites, these non-coding regions facilitate accurate translation of the library, as well as mutagenesis or gene shuffling of the library. Oligonucleotides G in a given library will generally all have the same number of codons arranged in a similar manner.

[0194] In some implementations, for a given oligonucleotide G library, there are typically predetermined sets and numbers of sequences to be used in each coding region. In some implementations, the coding sequence used in one coding region is not used in any other coding region. In some implementations, the coding sequences are assembled into genes in a combinatorial manner, thereby representing all possible combinations of the coding sequences. In other implementations, the number of combinations of coding sequences is significantly reduced in the initial gene library, and after a selection event, a subset of the population will undergo a gene shuffling or crossover procedure to achieve the evolution and selection of new phenotypes.

[0195] In some implementations, these “crossover” events are facilitated by placing unique restriction sites within non-coding regions. Partial digestion of the population in one or more non-coding regions, followed by reconnection, allows for reclassification of combinations of coding sequences between the two coding regions that underwent digestion. This shuffling or crossover event can occur between a pair of coding regions or between multiple pairs of coding regions, or the library can be divided into pools, with each pool experiencing crossover events between different combinations of coding regions. In some implementations, this capability effectively allows for genetic recombination between two individuals who have survived selection and proven fit to encode a genotype different from either parent. Typically, recombination of fit genotypes produces new, more fit genotypes for selection.

[0196] Identification of the coding part

[0197] In some embodiments, this disclosure provides a multifunctional molecule, which is a molecular probe that has a correspondence between a DNA gene sequence in oligonucleotide G and the identity of the coding portion or molecule encoded by the gene.

[0198] In some implementations, the correspondence is established as follows to prepare gene libraries in a manner where coding regions are single-stranded and any non-coding regions are double-stranded.

[0199] Independently, reaction site adaptors are prepared. The reaction site adaptors (hairpin structures) will be described in more detail below. In short, in some embodiments, the reaction site adaptors are typically DNA hairpins functionalized with reactive functional groups and include a stem region and an anticoding region. In some embodiments, because terminal coding regions are present in the library, many reaction site adaptors with many different anticoding sequences will be provided and prepared. In some embodiments, reaction site adaptors with anticoding sequences complementary to the sequence of each terminal coding region will be prepared. In some embodiments, the reactive functional group reaction site on each reaction site adaptor having its own sequence will react with a first structural unit to produce a loaded reaction site adaptor (vector anticodon), and will optionally be purified to remove unreacted reaction site adaptors, providing significant advantages in translation fidelity. Thus, the anticoding sequence of the reaction site adaptor corresponds to a specific structural unit. This chemistry will be described in more detail below. The loaded reaction site adaptor pool can be incubated with the gene library to specifically anneal the loaded reaction site adaptors (vector anticodons) to their complementary terminal coding sequences. The annealed adaptor / library complexes can then be ligated together, for example, using T4 DNA ligase. In this way, the terminal coding regions of the gene will correspond to specific structural units.

[0200] In some implementations, the correspondence between the next selected coding region sequence and the next structural unit is established by sorting the library into subpools based on the coding sequence of the selected coding region. In some implementations, this sorting is achieved by sequence-specific hybridization of single-stranded coding sequences with complementary oligonucleotides immobilized on an array of solid supports (called a hybridization array).

[0201] The construction of the hybridization array is described below. In short, in some embodiments, the hybridization array is an array containing spatially separated features of solid supports. In some embodiments, ssDNA oligonucleotides are covalently tethered to these supports, wherein sequences complementary to the coding region sequence are sorted. In some embodiments, library members with complementary coding sequences can be specifically immobilized by passing a molecular library of formula (I) or formula (II) carrying multiple coding sequences through or through a solid support carrying a given inverse coding sequence. In some embodiments, passing the library through or through an array of solid supports (each solid support carrying a different immobilized inverse coding sequence) sorts the library into subpools based on the coding sequence. In some embodiments, each sequence-specific subpool can then be independently reacted with a specific structural unit (positional structural unit) to establish a sequence-to-structural unit correspondence. This synthesis will be described in more detail below and can be performed on the hybridization array or after the subpools have been eluted from the array into a suitable environment (e.g., a separate container).

[0202] The correspondence between coding sequences and structural units for all other internal non-terminal coding regions can be established in the same way, the only difference being that different hybridization arrays with different sets of inverse coding sequences are used when appropriate.

[0203] The coding region in oligonucleotide G can also encode other information. In some implementations, after library translation is complete, library sorting may be required based on the index coding region sequence. In some implementations, the index coding region sequence can encode the intended purpose or the selection history of its corresponding sub-pool. For example, libraries for multiple targets can be translated simultaneously and then selected into sub-pools by index coding. Thus, sub-pools intended for different targets and / or for selection under different conditions can be separated from each other and ready for their respective applications. Therefore, the selection history of library members that have undergone multiple rounds of selection for various characteristics can be recorded in the index region.

[0204] One intended purpose of a response site adaptor is to establish a correspondence between sequences in a gene library of oligonucleotide G and specific structural units. Broadly speaking, this is achieved as described above. Typically, the key elements of a response site adaptor are an inverse coding sequence that allows the adaptor to hybridize with a specific coding region of oligonucleotide G, a reactive functional group that can be covalently linked to the first or second structural unit D or E, a linker that covalently tethers these structural units to the response site adaptor, and a mechanism that directly or indirectly covalently links the structural unit-loaded response site adaptor or the respond site adaptor to the gene (oligonucleotide G). In some embodiments, the response site adaptor is a DNA hairpin comprising a stem, a loop, and a 3' or 5' overhang containing an inverse coding sequence. In some embodiments, the stem may contain one or more unique restriction sites that, upon cleavage with a restriction enzyme, facilitate the release of a very tight binding agent from the target, or contain a favorable priming sequence to enable more pure gene amplification by PCR.

[0205] In some embodiments, the loop region performs the following task: generating a directional change in the reactive site adaptor strand to match the orientation of the oligonucleotide G strand, allowing it to attach to the end of the oligonucleotide G. Due to the nature of DNA, the coding sequence in the oligonucleotide G will have the opposite orientation to the anticoding sequence in the reactive site adaptor. DNA ligation only occurs between the two ends of nucleic acid strands with the same orientation. That is, a strand with a 5' to 3' orientation can attach to another strand with a 5' to 3' orientation. The loop region performs the following task: generating a directional change in the reactive site adaptor strand to match the orientation of the oligonucleotide G strand, allowing it to attach to the end of the oligonucleotide G. In some embodiments, the loop of the hairpin structure H1 and / or H2 comprises 3 to 12 DNA bases. In some embodiments, it is a polyethylene glycol adapter containing 6-12 PEG units. In some embodiments, the reactive site adaptor will be enzymatically ligated to the oligonucleotide G. In some embodiments, the terminus of oligonucleotide G will be functionalized with an azide or alkyne, and the terminus of the reactive site adaptor will be functionalized with an alkyne or azide, and the linkage will be achieved via copper-mediated "click" chemistry. Those skilled in the art will understand that numerous equivalent chemistry methods are applicable to covalently tethering the reactive site adaptor to oligonucleotide G.

[0206] In some embodiments, the reactive site comprises a free amine tethered to the modified nucleobase via a PEG connector or alkyl linker or via a homopolymer or heteropolymer linker. The reactive site can function when attached at any point on the hairpin of the reactive site adaptor. In some embodiments, the reactive site is attached at the 5' position of the loop on the stem. In some embodiments, the reactive site is attached at the 3' position of the loop. In some embodiments, the reactive site is attached to one side of the loop identical to the inverse coding sequence and close to the initial base of the inverse coding sequence. In some embodiments, the reactive site is attached to the opposite side of the loop opposite the inverse coding sequence and away from it along the stem. In some embodiments, more than one reactive site is attached to the adaptor. One advantage of having more than one reactive site is an increased probability of correctly synthesizing the coding moiety. Another advantage of having more than one reactive site is that the diversity of the coding moiety during selection can generate affinity, allowing weaker binders to be found during selection. In some embodiments with two or more reactive sites, the reactive sites will be attached to the same side of the loop and close to opposite ends of the stem. In some embodiments, the reactive sites will be located on opposite sides of the loop and near opposite ends of the stem. In some embodiments with more than one reactive site, both reactive sites will be located within the loop region. In some embodiments, the reactive site connector will comprise more than one stem and more than one loop. In such embodiments, multiple reactive sites may be connected along the stem, on the loop, or in a combination of both. In some embodiments, the arrangement and positioning of the reactive sites are designed to facilitate better selection outcomes, which is achieved by adjusting the distance between the reactive sites according to the size of the target in question to promote better affinity.

[0207] In some implementations, the reactive site adaptor will have a mechanism that allows it to easily break into smaller fragments. This fragmentation can facilitate better downstream processing of the library; for example, the presence of hairpins connecting one or both ends of the template library strand can complicate PCR amplification by interfering with the ability of primers to anneal properly with the molecule of formula (I). In another embodiment, fragmentation that loads the reactive site adaptor after ligation can reduce the space volume near the reactive site and improve chemical properties or increase yield during hybridization. Fragmentation methods include, but are not limited to, restriction digestion at a unique restriction site in the stem region, or incorporation of dU bases during the synthesis of the reactive site adaptor followed by treatment with uracil DNA glycosylase to produce a pyrimidine-free site, followed by alkaline hydrolysis of the strand. Fragmentation methods can tether the reactive site directly or indirectly to the template strand, or can remove it from the template strand.

[0208] In some embodiments, the loading reactive site adaptor is specifically hybridized and ligated to the terminal coding regions at both ends of the template strand. When the terminal coding regions encode the same first and second structural units, these embodiments offer the advantage of providing multiple possibilities for the correct coding moiety to be synthesized. Furthermore, these embodiments provide opportunities for multiple coding moieties to generate affinity and improve selection efficiency, particularly for weaker binders. When the first and second structural units differ, these embodiments allow the same number of gene template strands to synthesize twice the number of unique coding moies. When the first and second structural units are the same for some formula (I) molecules and different for others, it allows for analysis of the relative contribution of the two different structural units to the overall fitness conferred throughout the coding moieties.

[0209] Many chemistry methods are applicable to this invention. Theoretically, any chemical reaction that does not chemically alter DNA can be used. Known DNA-compatible reactions include, but are not limited to: Wittig reaction, Heck reaction, Homer-Wads-worth-Emmons reaction, Henry reaction, Suzuki coupling, Sonogashira coupling, Huisgen reaction, reductive amination, reductive alkylation, peptide bond reaction, peptide-like bond formation reaction, acylation, SN2 reaction, SNAr reaction, sulfonation, ureation, thioureation, carbamylation, formation of benzimidazole, imidazolidinone, quinazolinone, isoindolinone, thiazole, imidazopyridine, diol cleavage to form glyoxal, Diels-Alder reaction, indole-styrene coupling, Michael addition, olefin-alkyne oxidative coupling, aldol reactions, Fmoc deprotection, trifluoroacetamide deprotection, Alloc deprotection, Nvoc deprotection, and Boc deprotection. (See Handbook for DNA-Encoded Chemistry (edited by Goodnow RA, Jr.), pp. 319–347, Wiley, NY, 2014; March, Advanced Organic Chemistry, 4th ed., New York: John Wiley and Sons (1992), Chapters 10–16; Carey and Sundberg, Advanced Organic Chemistry, Part B, Plenum (1990), Chapters 1–11; and Coltman et al., Principles and Applications of Organotransition Metal Chemistry, University Science Books, Mill Valley, Calif. (1987), Chapters 13–20; each of which is incorporated herein by reference in its entirety.)

[0210] Those skilled in the art will understand that a wide variety of different combination scaffolds can be incorporated into the multifunctional molecules of this disclosure. Examples of general types of scaffolds include, but are not limited to, the following: (a) chains of end-to-end connected bifunctional structural units, peptides and peptide-like structures being two examples of such scaffolds; it should be understood that not every bifunctional structural unit in the chain has the same pair of functional groups, and some structural units may have only one functional group, such as terminal structural units; (b) branches of bifunctional structural units that include some trifunctional structural units and may or may not include monofunctional structural units; (c) molecules comprising a single multifunctional structural unit and a set of monofunctional structural units; in one embodiment, such a molecule may have a multifunctional structural unit as a central core, wherein other monofunctional structural units are added as diversification elements; (d) molecules comprising two or more multifunctional structural units to which a set of monofunctional or bifunctional structural units are attached as diversification elements; and (e) any of the above-described scaffolds that includes the formation of a ring by reacting a portion of a connector or structural unit installed in an earlier step with a portion of a structural unit or connector installed in a later step. Other scaffolds or chemical structures can also be incorporated, and these universal structural scaffolds are limited only by the ingenuity of practitioners in designing the chemical pathways through which they are synthesized.

[0211] In some implementations, ion exchange chromatography facilitates the chemical reaction of the DNA-tethered matrix in two ways. For reactions carried out in aqueous solvents, the reactants can be poured onto an ion exchange resin such as... or Purification can be easily achieved using the SuperQ 650M. In some implementations, DNA is bound to the resin via ion exchange, and unused reactants, byproducts, and other reaction components can be washed away with an aqueous buffer, an organic solvent, or a mixture of both. A real problem exists for reactions that work best in organic solvents: DNA is very poorly soluble in organic solvents, and these reactions have low yields. In these cases, the library DNA can be immobilized on an ion exchange resin, residual water can be washed away with an aqueous-miscible organic solvent, and the reaction can be carried out in an organic solvent that may or may not be miscible with water. See, for example, R.M. Franzini et al., Bioconjugate Chemistry 2014 25(8), 1453-1461, and references therein. Many types and kinds of ion exchange media exist, each with different properties that may be more or less suitable for different chemistry or applications, and they are available from many companies such as... SIGMA and Wait for commercial purchase. It should be understood that there are many possible means and media by which library DNA can be fixed or dissolved for chemical reactions to install structural units, remove protective groups, or activate portions for further modification, which are not listed here.

[0212] In some embodiments, the hybridization array includes means for sorting heterogeneous mixtures of ssDNA sequences by sequence-specific hybridization of the ssDNA sequence with complementary oligonucleotides immobilized in a position-addressable form. See, for example, U.S. Patent No. 5,759,779. It should be understood that hybridization arrays can take many physical forms. In some embodiments, the hybridization array has the ability to contact a heterologous sample or ssDNA (i.e., a library of compounds of formula (I)) with complementary oligonucleotides immobilized on the array surface. The complementary oligonucleotides are immobilized on the array surface in a manner that enables, allows, or facilitates sequence-specific hybridization of the ssDNA with the immobilized oligonucleotides, thereby also immobilizing the ssDNA. In some embodiments, ssDNA immobilized by a common sequence can be independently removed from the array to form subpools.

[0213] In some embodiments, the hybridization array comprises a chassis consisting of a rectangular plastic sheet 0.1 to 100 mm thick, with a series of holes cut into it, referred to as a "feature". In some embodiments, filter membranes are adhered to the bottom and top of the sheet. In some embodiments, in the feature, what is trapped between the filter membranes is a solid surface or an assembly of solid surfaces, referred to as a "solid support". In some embodiments, in any given feature, a single sequence of oligonucleotide is immobilized on the solid support.

[0214] In some embodiments, type I or II molecular libraries can be sorted on an array by allowing an aqueous solution of the library to flow through and through these features. In some embodiments, library members become immobilized within a feature when they come into contact with an oligonucleotide in a feature carrying a complementary sequence. In some embodiments, after hybridization, the features of the array can be positioned in a receiving container such as a 96-well or 384-well plate. In some embodiments, an alkaline solution that causes DNA dehybridization can be added to each feature, and this solution will carry the thus mobile library into the receiving container. Other dehybridization methods are also possible, such as using heat buffers or denaturing agents. Therefore, in some embodiments, molecular libraries can be sorted into subpools in a sequence-specific manner.

[0215] It should be understood that the aforementioned chassis may comprise plastic, ceramic, glass, polymer, or metal. It should be understood that the solid support may comprise resin, glass, metal, plastic, polymer, or ceramic, and the support may be porous or non-porous. It should be understood that a higher surface area on the solid support allows for the immobilization of a larger amount of complementary oligonucleotides and can capture a larger number of library pools within the features. It should be understood that the solid supports may be held within their respective features by means other than a filter membrane, such as glue, adhesive, or covalent bonding of the support to the chassis and / or other supports. It should be understood that these features may or may not be holes in the chassis, but rather independent structures that can be removed from or placed within the chassis. It should be understood that the shape of the chassis does not need to be a rectangle with features arranged in two dimensions, but may be a cylinder or rectangular prism with features arranged in one or three dimensions. See, for example, U.S. Patent No. 5,759,779.

[0216] The library of molecules of formula (I) can be considered a phenotypic population tethered to their respective genotypes. Such a population can withstand selection pressures, which remove less fit individuals and allow more fit members to survive. The oligonucleotide G genotypes of the second-generation population, i.e., those that survived selection, can be amplified by PCR, retranslated, and subjected to another, more stringent selection for the same trait, or for some orthogonal traits. Deep sequencing or next-generation sequencing technologies can also typically be used to sequence the surviving subpopulations, and the sequencing data can be analyzed to identify the most suitable coding portions (phenotypes).

[0217] Many options are possible. The most typical option is to identify individuals in the population that can bind to the target protein. In some implementations, this selection is performed by fixing the target protein to a solid support, such as NUNC. The surface of the wells in the plate, or by biotinylating the target and immobilizing it on streptavidin-coated magnetic beads. In some embodiments, after target immobilization, the population of molecules of formula (I) is incubated together with the target on the support. All those individuals capable of binding the target are thus immobilized. The solid support is washed with a suitable buffer to remove the non-binding agent. In some embodiments, the DNA encoding the binding agent can be amplified by PCR and sent for sequencing for retranslation and another round of selection.

[0218] In some embodiments, selection can be made by choosing individuals that bind to a target protein to exclude different anti-target proteins or a group of anti-target proteins. In this case, one selection method requires fixing the target and anti-target to a solid support in separate containers. In some embodiments, the library is first incubated with the anti-target, and so are individuals capable of binding to the anti-target. In some embodiments, the non-binding agent is carefully removed from the container and transferred to a container containing the target. In this way, the population selected for the ability to bind to the target is first depleted among individuals capable of binding to the anti-target, and individuals whose fitness is characterized by the ability to bind to the target or exclude the anti-target are selected.

[0219] In some implementations, a second approach to identifying the coding portion of a target that selectively binds to another target is to select the two targets in parallel and then eliminate the coding portions that exhibit affinity for both targets during sequencing data analysis.

[0220] In some embodiments, selection of binders with low dissociation rates can also be achieved by using a mixture of immobilized and free targets. In some embodiments, the library is incubated with the immobilized target, thereby allowing binder binding. An excess of free target is then added and incubated for a predetermined amount of time. During this period, any binders released from the immobilized target and subsequently re-bound have a high probability of re-binding to the free target. After washing away the non-binding agents, the free target and any substances bound to it are also washed away. The only binders remaining after the free target are those with dissociation rates longer than the predetermined incubation time for the free target.

[0221] The selection methods described in the preceding paragraphs can be found in the literature on phage display, ribosome display, and mRNA display. See, for example, Amstutz, Patrick, et al., Cell biology: a laboratory handbook, 3rd edition. ELSEVIER, Amsterdam (2006): 497-509, and its references.

[0222] In principle, any property can be selected, provided that a mechanism can be constructed to selectively amplify individuals in a population that possess the property compared to those that do not. Pharmacologically relevant properties other than target binding can be selected in principle, and examples include, but are not limited to, selections for water solubility, cell membrane penetrance, and non-toxicity.

[0223] It should also be understood that synthesizing libraries in sufficient quantities can allow for more than one selection in a given round. In some embodiments, the survivor subpopulation after selection based on target affinity can be isolated, optionally purified, and subjected to a second selection based on affinity for the same or different targets, or on orthogonal properties. In some embodiments, the survivor pool is then amplified by PCR and sequenced, or it is amplified and retranslated for further selection.

[0224] In some implementations, sequencing data are analyzed by comparing the performance of library members in the population before and after selection. In some implementations, members that perform less after selection are generally considered less fit, and members that perform more after selection are considered more fit. Additionally, data may be analyzed to determine which individual structural units confer fit, which structural units confer fit when bound within the same coding moiety, and which triples of structural units confer fit. In some implementations, data may be analyzed to determine which structural elements within different structural units and different coding moies confer fit to the selected library members. In some implementations, these analyses inform which members should be synthesized for independent testing and suggest similar molecules that should be prepared and tested, which may not be native members of the library. In some implementations, a three-dimensional docking algorithm may also inform these processes.

[0225] In some embodiments, library members identified in data analysis can be synthesized with or without an oligonucleotide moiety, typically using the same or similar synthetic conditions as those used in library preparation. In some embodiments, these independently synthesized samples can then be subjected to various tests characterizing their physical and chemical properties and demonstrating their general suitability for the desired task. In some embodiments, these properties include, but are not limited to, the dissociation constant or KD that measures the tightness of binding of the library member to its target, such as water solubility measured by water:octanol partitioning, and extracellular penetrance measured in CaCo cells.

[0226] In some embodiments, the identified library members that bind to a biomolecule can be used to determine the biological function of that biomolecule. In some embodiments, the functions of many proteins are not yet known, and the methods of this disclosure provide a readily available pathway for discovering molecular probes to help elucidate these functions. In some embodiments, library members identified by the methods of this disclosure can be used to help determine whether a biomolecule is particularly suitable for small molecule discovery and targeting for drug intervention.

[0227] In some implementations, the effect on the function of the biomolecules binding to the library member can be determined in in vitro or in vivo assays, in cell-based or cell-free assays. For biomolecules with known functions, the effect of the identified library member on that function can be assessed. If the biomolecule is an enzyme, the effect on its activity rate can be assessed. If it is a signaling protein, the effect on cellular function, including cell viability, cellular gene expression, or cellular phenotypic expression, can be assessed. If the target is a viral protein, the effect of the library member on viral proliferation and viability can be assessed.

[0228] In some implementations, the effects of selected library members on animal and human and plant health in in vivo experiments can also be assessed.

[0229] In some embodiments, identified library members can also be used as affinity reagents for purifying biomolecular targets. In some embodiments, the identified coding moiety can be immobilized on a solid support, and a heterogeneous solution containing the target can flow through the solid support. In some embodiments, the target binds to the coding moiety and is immobilized. In some embodiments, all other components of the mixture can be washed away, leaving a purified target sample.

[0230] The invention is illustrated by the following examples, but is not limited thereto. Those skilled in the art will recognize many equivalent techniques for implementing the steps or steps listed herein.

[0231] Example

[0232] An implementation scheme for the molecule of construction formula (I) is as follows.

[0233] Example 1: Constructing an 8×10 9 Gene library of each member

[0234] Codons for the gene library are provided. Ninety-six double-stranded DNA ("dsDNA") sequences are provided or purchased from gene synthesis companies such as Genscript (Piscataway NJ), Synbio Technologies (Monmouth Junction NJ), Biomatik (Wilmington DE), and Epoch Life Sciences (Sugarland TX). These sequences contain five coding regions, each with 20 bases. Each coding region is flanked by 20-base non-coding regions (creating a total of six non-coding regions). All coding region sequences are unique and selected to not cross-react with other coding and non-coding regions. The five non-coding regions in the DNA molecule have different sequences, but the sequence at each position is conserved across all DNA. All coding and non-coding regions are designed to have similar melting temperatures (between 58°C and 62°C). The computer design of the coding and non-coding regions is as follows. The DNA sequences are randomly generated in a computer.

[0235] Once generated, the sequence melting temperature and thermodynamic properties (ΔH, ΔS, and ΔG at melting) are calculated using the nearest neighbor method. If the calculated Tm and other thermodynamic properties are not within the predetermined range required for the library, the sequence is excluded. Acceptable sequences are analyzed using a sequence similarity algorithm. Sequences predicted by the algorithm to be sufficiently non-homologous are considered non-cross-reactive and are retained. Others are excluded. Coding and non-coding regions are sometimes selected from an empirical list of oligonucleotides shown as non-cross-hybridizing. See Giaever G, Chu A, Ni L, Connelly C, Riles L, et al., (2002) Functional profiling of the Saccharomyces cerevisiae genome. Nature 418:387-391. This reference lists 10,000 non-cross-reactive oligonucleotides. Their respective Tm is calculated, and those falling within the predetermined range are analyzed using a sequence homology algorithm. Those sufficiently non-homologous are retained.

[0236] Each non-coding region contains a unique restriction site. The non-coding region at the 5' end of the template strand contains a SacI recognition site at bases 13-18 from the 5' end. The non-coding region at the 3' end of the coding strand contains an EcoRI restriction site at bases 14-19 from the 3' end of the template strand. The second, third, fourth, and fifth non-coding regions from the 5' end of the template strand contain HindIII, NcoI, NsiI, and SphI recognition sites at bases 8-13, respectively.

[0237] The DNA was subjected to restriction digestion to uncouple all codons from each other. The DNA sequence was then assembled and dissolved in a solution from New England Biolabs (NEB, Massachusetts). In the buffer solution, the concentration was approximately 20 μg / ml. An internal restriction enzyme from NEB was added. and The enzyme was digested at 37°C for 1 hour according to the manufacturer's instructions. The enzyme was then heat-inactivated at 80°C for 20 minutes. After inactivation, the reaction was maintained at 60°C for 30 minutes, then cooled to 45°C and maintained for 30 minutes, and then cooled to 16°C.

[0238] Codon recombination was performed to generate a gene library. To reassemble the individual codons generated during the digestion reaction into full-length genes, T4 DNA ligase from NEB was added to the reaction at 50 U / ml, dithiothreitol (DTT, Thermo Fisher Scientific, Massachusetts) at 10 mM, and 5'-adenosine triphosphate (ATP, from NEB) at 1 mM, according to the manufacturer's protocol. The ligation reaction was carried out for 2 hours, and the products were purified by agarose gel electrophoresis. Because the sticky ends generated at one site of the provided gene through digestion will anneal with the sticky ends of all other digestion products at the same site, complete recombination occurs. Therefore, the 96 genes, each containing 5 codons, provided will yield 965 genes. Since there are 96 coding sequences at each of the 5 coding positions, there are 96... 5 =8×10 9 A group or library member.

[0239] Preparation of gene libraries for translation .

[0240] Gene libraries were amplified by PCR. The T7 promoter was attached to the 5' end of the non-template strand by extension PCR, using these reactants for 50 μL reactions: High-fidelity DNA polymerase (“ Polymerase (NEB), 10 μL; deoxynucleotide (dNTP) solution mixture, 200 μM final concentration; forward primer, final concentration 750 nM; reverse primer, final concentration 750 nM; template (sufficient template should be used to adequately oversample the library); dimethyl sulfoxide (DMSO), 2.5 μL; Polymerase, 2 μL. PCR was performed using an annealing temperature of 57 °C and an extension temperature of 72 °C. Annealing was performed for 5 seconds per cycle; extension was performed for 5 seconds per cycle. Products were analyzed by agarose gel electrophoresis.

[0241] DNA was transcribed into RNA. The transcription reaction was performed in 250 μL of unpurified PCR product using the following reactants: PCR product, 25 μL; RNase-free water, 90 μL; nucleoside triphosphates (NTPs), each at a final concentration of 6 mM; 5xT7 buffer, 50 μL; NEB T7 RNA polymerase, 250 units; optionally, additional... Ribonuclease inhibitor (Promega Corporation, WI) to 200 U / ml; optionally, pyrophosphatase to 10 μg / ml. 5xT7 buffer contains: 1 M HEPES-KOH (4-(2-hydroxyethyl)-1-piperazine ethanesulfonic acid) pH 7.5; 150 mM magnesium acetate; 10 mM spermidine; 200 mM DTT. The reaction is carried out at 37°C for 4 hours. RNA is purified by lithium chloride precipitation. The transcription reaction is diluted with 1 volume of water. LiCl is added to 3 M. The mixture is vortexed at 4°C at maximum g for at least 1 hour. The supernatant is discarded and retained. Clean globules will be a clear, glassy gel that is difficult to dissolve. Alternating gentle heating (at 70°C for one minute) and gentle vortexing will cause the globules to resuspend. Analyze by agarose gel electrophoresis, quantifying as quickly as possible and freezing to avoid degradation. See, for example, Analytical Biochemistry 195, pp. 207-213 (1991); and Analytical Biochemistry 220, pp. 420-423 (1994).

[0242] RNA was reverse transcribed into DNA. This was done using technology from Thermo Fisher Scientific. III. Reverse transcriptase and the provided first-strand buffer were used to reverse transcribe single-stranded RNA ("ssRNA") in a two-step procedure. Step 1 was performed using the following components at these final concentrations: dNTPs, 660 μM each; RNA template, ~5 μM; primers, 5.25 μM. The Step 1 components were heated to 65°C for 5 minutes and then frozen for at least 2 minutes. The final concentrations of the Step 2 components were: first-strand buffer, 1x; DTT, 5 mM; RNase inhibitor (NEB), 0.01 U / μL. III reverse transcriptase, 0.2 U / μl. Combine the components from step 2, warm to 37°C, and add the mixture from step 2 to the mixture from step 1 after freezing the components from step 1 for 2 minutes. Incubate the combined portion at 37°C for 12 hours. Perform agarose gel electrophoresis after the reaction. Sample the reaction for known starting RNA and known products or known product analogs such as PCR product libraries. Add ethylenediaminetetraacetic acid ("EDTA") to all samples, heat to 65°C for 2 minutes, rapidly cool, and then perform electrophoresis on an agarose gel. ssRNA should be resolved from the complementary DNA ("cDNA") product. Purify the cDNA product by adding 1.5 volumes of isopropanol and ammonium acetate to 2.5 M, then centrifuging at 48,000 g for 1 hour. Resuspend the cDNA globules in distilled water ("dH2O") and hydrolyze the RNA strands by adding LiOH to pH 13. Heat the solution to 95°C for 10 minutes. Add 1.05 equivalents of non-coding region-specific primers, neutralize the pH with tris(hydroxymethyl)aminomethane ("Tris") and acetic acid, and slowly cool the reaction to room temperature. Then concentrate and buffer exchange to NEB. In buffer solution.

[0243] Remove the terminal non-coding region. The reverse transcribed ssDNA product containing complementary oligonucleotides that make the non-coding region double-stranded is suspended in NEB at a concentration of 100 μg / ml. In buffer solution. The restriction enzyme from NEB... and Add 1 μg of DNA to the digest. Incubate the digest at 37°C for 1 hour, then heat-inactivate the enzyme at 65°C for 20 minutes.

[0244] Prepare reactive site adaptors for translation.

[0245] Reactive site adaptors are provided. Two sets of 96 reactive site adaptors are provided, each adaptor comprising a hairpin loop, a stem, and a slack end containing an inverse coding sequence. One set has a 3' slack end inverse coding sequence that specifically hybridizes to the 3' end coding region of the template strand after removal of the 3' end uncoding region; the other set has a 5' slack end inverse coding sequence that specifically hybridizes to the 5' end coding region of the template strand after removal of the 5' end uncoding region. The set with the 3' slack end has a 5' phosphoryl group. In this embodiment, the stem region of each set has the same sequence as the corresponding end uncoding region previously removed by restriction digestion. The loop region of each set carries a base modified with the reactive site N4-TriGl-amino2'-deoxycytidine (from IBA, Goettingen, Germany). The adaptors described herein can be purchased from DNA oligonucleotide synthesis companies such as Sigma Aldrich, Integrated DNA Technologies of Coralville, IA, or Eurofins MWG of Louisville, KY.

[0246] Loading of reaction site initiators. Two sets of 96 reaction site initiators were provided in separate wells and dissolved in TE buffer (Promega, MA). 15 μl of... SuperQ-650M (Sigma-Aldrich, St. Louis, MO) ion exchange resin was placed in each well of a filter plate and washed with 100 μl of 10 mM HOAc. Aliquots of the linkers at each reaction site, proportional to the amount of the template chain, were transferred to individual wells of the filter plate, where they were immobilized on the resin. The immobilized linkers were washed with dH₂O, then with piperidine, and finally with dimethylformamide ("DMF"). Ninety-six reaction solutions were prepared, each containing: 50 μl of DMF, 75 mM of Fmoc-protected amino acids, 75 mM of 4-(4,6-dimethoxy-1,3,5-triazin-2-yl)-4-methylmorpholinon tetrafluoroborate, and 90 mM of N-methylmorpholine. These mixtures were allowed to activate the acid at room temperature for ten minutes, then added to the resin and reacted for 30 minutes. The resin was then washed with 4 × 100 μl of DMF, and the coupling step was repeated with freshly prepared reaction mixture. The resin was washed again with DMF, and the Fmoc protecting groups were removed by adding 50 μl of a 20% piperidine DMF solution to each well and incubating at room temperature for 2 hours. The resin was washed again with 4 × 100 μl of DMF, followed by washing with 3 × 100 μl of dH₂O. The solution was then treated with 1.5 M NaCl, 50 mM KOH, and 0.01% TRITON. TMX-100 elutes the loaded reaction site initiators from the resin. The solution is neutralized by adding Tris to 15 mM and HOAc to pH 7.4. The loaded reaction site initiators are then collected and passed through ZEBA. TM 7K MWCO (ThermoFisher Scientific, MA) desalination cartridges are used for desalination.

[0247] Translation of the document library

[0248] The loading reaction site linker was connected to the library. ZEBA was used at 25°C. TM The restriction-digested template library was buffer-exchanged in a 30K MWCO (ThermoFisher Scientific, MA) centrifuge to 50 mM Tris-HCl, 10 mM MgCl2, and 25 mM NaCl (pH 7.5). 1.1 equivalents of a 3' end-specific loading site adaptor and 1.1 equivalents of a 5' end-specific loading site adaptor were added, and the mixture was diluted to a template concentration of 1 μM with the same buffer. The reaction was incubated at 65°C for 10 minutes, cooled to 45°C over 1 hour, and incubated at 45°C for 4 hours. After cooling to room temperature, DTT was added to 10 mM, ATP to 1 mM, and T4 DNA ligase to 50 U / mL. The ligation reaction was carried out at room temperature for 12 hours, then the enzyme was heat-inactivated at 65°C for 10 minutes, and the reaction was slowly cooled to room temperature. The reactants were buffer exchanged and concentrated using a 30K molecular weight cutoff (MWCO) centrifuge to 150 mM NaCl, 20 mM citrate, 15 mM Tris, 0.02% sodium dodecyl sulfate ("SDS"), 0.05% Tween 20 (from Sigma-Aldrich), pH 7.5.

[0249] Preparation of hybridization arrays. The hybridization arrays consist of ~2mm thick TECAFORM. TM The chassis is constructed of acetal copolymer with holes cut by a computer numerical control (CNC) machine. A 40-micron nylon mesh from ELKOFILTERING is adhered to the bottom of the chassis using NP200 double-sided tape from Nitto Denko. Then, CM (azide-functionalized nylon mesh) is used. Solid supports of resin (Sigma Aldrich) were used to fill the pores. The resin was functionalized using an 8-PEG-azide PEG-amine purchased from Broadpharm (San Diego, CA). 45ml CM packages were used. The resin was placed in a sintering funnel and washed with DMF. It was then suspended in 90 ml of DMF and reacted with 4.5 mM azido-PEG-amine, 75 mM EDC, and 7.5 mM HOAt at room temperature for 12 hours. The resin was washed with DMF, water, and isopropanol and stored in 20% ethanol at 4°C. A 40-micron nylon mesh was then adhered to the top of the substrate. The azido groups allow for the use of click chemistry to tether alkyne-linked oligonucleotides to a solid support. The array was placed in an array-well plate adapter and the adapter was fixed to the well plate so that the captured oligonucleotides could be aligned and "clicked" onto the azido groups. Above. A 30 μl solution (100 mM) containing 1 nmol alkyne oligonucleotide, copper sulfate, 625 μM tris(3-hydroxy-propyl-triazolyl-methyl)amine ("THPTA") (ligand), 3.1 mM aminoguanidine, 12.5 mM ascorbate, and 12.5 mM phosphate buffer (pH 7) was added to each well of the array-plate connector, allowing it to adsorb onto the plate. On the support. After 10 minutes, the solution was centrifuged from the array and transferred into the plate, then repositioned back onto the array, one well aligned with the other, for a second round of reaction. After the second 10-minute reaction, the reaction solution was centrifuged into the well plate and the plate was set aside. The array was thoroughly washed with 1 mM EDTA and stored in phosphate buffered solution ("PBS") containing 0.05% sodium azide. The reaction solutions were each diluted to 100 μl with dH2O, loaded onto diethylaminoethyl (DEAE) ion exchange resin, and washed with dH2O to remove all reagents and reaction byproducts except for any unincorporated oligonucleotides. These solutions were analyzed by high-performance liquid chromatography (HPLC) to determine the extent of incorporation due to the disappearance of the starting material. One array contains an oligonucleotide complementary to a coding position in the template library. A separate array was fabricated for each coding position.

[0250] Library sorting was performed using sequence-specific hybridization. The prepared hybridization libraries were placed in 1x hybridization buffer (2x saline sodium citrate (SSC), +15mM Tris, pH 7.4, +0.005%). Dilute to 13 ml in X100 (0.02% SDS, 0.05% sodium azide). Add 10 μg of transfer RNA ("tRNA") to block nonspecific nucleic acid binding sites. Select an array corresponding to the desired coding position in the template library. Place the array in a chamber with a 1-2 mm gap on both sides and pour in 13 ml of library solution. Seal the chamber and gently shake at 37°C for 48 hours. Optionally, place the array in a device that allows the solution containing the library to be directionally pumped through various features in a pre-patterned path as a means of faster sorting of the library on the array.

[0251] The sorted library was eluted from the hybridization array. The array was washed by opening the chamber and replacing the hybridization solution with fresh 1x hybridization buffer, followed by shaking at 37°C for 30 minutes. Washing was repeated 3 times with hybridization buffer, and then 2 times with 1 / 4x hybridization buffer. The library was then eluted from the array. The array was placed in an array-plate connector, and 30 μl of 10 mM NaOH and 0.005% HCl were added to each well. X-100 and incubate for 2 minutes. Spin the solution through the array into a plate in a centrifuge. Perform the elution program three times. Neutralize the sorted library solution by sequentially adding 9 μl of 1M Tris pH 7.4 and 9 μl of 1M HOAc to each well.

[0252] A peptide-like coupling chemistry step was performed on the sorted library. 15 μl of SuperQ 650M resin was added to each well of a filter plate, and the sample was washed with 100 μl of 10 mM HOAc. The sorted library was transferred from the well plate, aligned with each other, after being eluted from the hybridization array and rotated into the well plate containing ion-exchange resin. The resin and library were washed with 1 x 90 μl of 10 mM HOAc, 2 x 90 μl of dH2O, 2 x 90 μl of DMF, and 1 x 90 μl of piperidine. Separately, a methanol solution containing 100 mM sodium chloroacetate and 150 mM 4-(4,6-dimethoxy-1,3,5-triazin-2-yl)-4-methylmorpholinium chloride was prepared. 40 μl of this solution was added to each resin well, and the reaction was carried out at room temperature for 30 min. The resin was washed with 3 x 90 μl of methanol, and then the coupling was repeated with 3 x 90 μl of methanol and 3 x 90 μl of DMSO. Additionally, a 2M (or saturated if necessary) DMSO solution of a primary amine was prepared. 40 μl of a primary amine solution was added to each resin well, and the reaction was carried out at 37 °C for 12 h. The resin was washed with 3 x 90 μl of DMSO, 3 x 90 μl of 10 mM acetic acid (HOAc), and 3 x 90 μl of dH2O. The solution was then washed with 1.5 M NaCl, 50 mM NaOH, and 0.005% DMSO. X-100 eluted the DNA library from the ion exchange resin at 3 × 30 μl aliquots. All reactants were collected and neutralized by adding Tris to 15 mM and HOAc to pH 7.4. The solution was then concentrated and buffered into 1X hybridization buffer.

[0253] Complete the synthesis of the library. Perform three additional sorting and synthesis steps using the library sorting scheme described above for the hybridization array, and the peptide or peptide-like chemistry scheme described above, or the schemes described below for other chemical steps in Examples 10-32, and fully translate the library.

[0254] Libraries were prepared for selection. Once the translation of the libraries was complete, 1x DREAMT AQ was used to extract the libraries using a template of 1.0 μM or less. TM Buffer, 1000x dNTPs [template], 0.2 U / μl DREAMTAQ TM Polymerase and an equimolar amount of MgCl2 supplement for each dNTP are combined in dH2O, optionally converting the single-stranded region into a double-stranded one. Note that the 3' end reaction site adaptor will serve as the primer for this reaction. The mixture is heated to 95°C for 2 minutes, then annealed at 57°C for 10 seconds and extended at 72°C for 10 minutes. The reaction is purified by ethanol precipitation.

[0255] Select a ligand that binds to the target protein. Immobilize 5 μg of streptavidin in 100 μl PBS onto MAXISORP. TM Incubate the plate in four wells overnight at 4°C with shaking. Wash the wells with 4 x 340 μl PBST. Block two wells with 200 μl casein and the other two wells with 5 mg / ml BSA at room temperature for 2 hours. Wash the wells with 4 x 340 μl PBST. Add 5 μg of biotinylated target protein in 100 μl PBS to the wells blocked with casein and to the wells blocked with BSA, and incubate at room temperature with shaking for 1 hour (for the protocol on protein biotinylation, see Elia, G. 2010. Protein Biotinylation. Current Protocols in Protein Science. 60:3.6:3.6.1-3.6.21). Add 100 μl of translational library aliquots in PBS (PBST) containing Tween 20 to each well that did not receive the target protein, and add 100 μl PBST to the two wells that received the target protein. Incubate the samples at room temperature with shaking for 1 hour. Carefully aspirate the buffer from the wells containing only the immobilized target protein and PBST. Carefully transfer the buffer containing the library from the target-free wells to the target-free wells. Add 100 μl of PBST to the target-free wells. Incubate all wells at room temperature with shaking for 4 hours. Carefully remove the library using a pipette and store it. Wash the wells with 4 x 340 μl PBST. To elute library members tightly bound to the target protein, add excess biotin from 100 μl of PBST to the wells and incubate at 37°C for 1 hour. Carefully aspirate the buffer and use it as a template for the PCR reaction.

[0256] Analysis of the selection results. PCR products from libraries before and after selection were deep sequenced using primers and protocols requested by the DNA sequencing service providers. Providers included Seqmatic in Fremont, CA, and Elim BioPharm in Hay Ward, CA. The coding sequences of the terminal and internal coding regions of each sequencing strand were analyzed to deduce the structural units used for synthesizing the coding portions. The relative frequencies of library members identified before and after selection indicate the degree to which selection enriched library members within the population. Analysis of various chemical subgroups containing library members that survived selection shows the degree to which these subgroups conferred fitness upon the library members and were used to evolve more suitable molecules or predict similar molecules for independent synthesis and analysis.

[0257] Example 2: Preparation and translation of a library with a single reactive site adaptor.

[0258] The library with a single reactive site linker was prepared exactly as in Example 1, except that the following steps were omitted: (a) removing a terminal non-coding region and (b) linking the corresponding reactive site linker.

[0259] Example 2a: Preparation and translation of a library with a single reactive site adaptor at the 5' end of G. To prepare a library with a single reactive site adaptor at the 5' end of the coding strand of G, the library can be prepared exactly as in Example 1, except that in the step "Removal of Terminal Uncoding Region," the only restriction endonuclease added should be SacI. This removes the 5' terminal uncoding region, allowing the 5' loading reactive site adaptor to hybridize and ligate properly to the template strand. Using only a restriction endonuclease specific to the recognition site in the 5' terminal uncoding region, and omitting the restriction endonuclease specific to the recognition site in the 3' terminal uncoding region, leaves the 3' terminal uncoding region in situ, thus preventing the 3' reactive site adaptor from ligating to that end. The 5' loading reactive site adaptor is added in the step "Ligating the Loaded Reactive Site Adaptor to the Library" as in Example 1. The addition of the 3' loading reactive site adaptor in this step is omitted. As described in Examples 3, 4a, and 4b, the responder in these examples may contain more than one stem, more than one loop, and more than one adapter. Those skilled in the art will understand that the 5' uncoding region described in Examples 6a and 6b can be used to remove the 5' uncoding region as described above. Those skilled in the art will understand that other restriction sites can be designed into the 5' coding region, and different restriction enzymes can be used for this purpose.

[0260] Example 2b: Preparation and translation of a library with a single reactive site adaptor at the 3' end of G. To prepare a library with a single reactive site adaptor at the 3' end of the coding strand of G, the library can be prepared exactly as in Example 1, except that the only restriction endonuclease added in the step "Removal of the Terminal Uncoding Region" should be EcoRI. This removes the 3' terminal uncoding region, allowing the 3' loading reactive site adaptor to hybridize and ligate properly to the template strand. Using only a restriction endonuclease specific to the recognition site in the 3' terminal uncoding region, and omitting the restriction endonuclease specific to the recognition site in the 5' terminal uncoding region, leaves the 5' terminal uncoding region in situ, thus preventing the 5' reactive site adaptor from ligating to that end. As in Example 1, the 3' loading reactive site adaptor is added in the step "Ligating the Loaded Reactive Site Adaptor to the Library". The addition of the 5' loading reactive site adaptor in this step is omitted. As described in Examples 3, 4a, and 4b, the responder in these examples may contain more than one stem, more than one loop, and more than one adapter. Those skilled in the art will understand that the method described in Example 6b can be used to remove the 3' uncoding region as described above. Those skilled in the art will understand that other restriction sites can be designed into the 3' coding region, and different restriction enzymes can be used for this purpose.

[0261] Example 2c. Preparation and translation of a library with reactive site adaptors linked at different points during synthesis. A library with two reactive site adaptors can be prepared, wherein one reactive site adaptor is linked to a template oligonucleotide G, on which some positional structural units are mounted, and then a second reactive site adaptor is linked to G. First, the method of Example 2a or Example 2b is used. Second, one or more positional structural units are mounted using the step of Example 1, “sorting the library by sequence-specific hybridization and eluting the sorted library from the hybridization array,” and any chemical steps for mounting positional structural units as described in Example 1, such as peptide coupling or peptide-like coupling, or any chemical steps described in Examples 10-32. Third, the second reactive site adaptor is linked to G. Those skilled in the art will understand that the second terminal non-coding region can be removed immediately before linking the second reactive site adaptor, or the chemical steps for mounting the positional structural units can be intermediate between the removal of the second terminal non-coding region and the linking of the second reactive site adaptor.

[0262] Example 3: Preparation and translation of libraries with two or more response sites per response site in the adaptor.

[0263] Libraries with multiple reactive sites on a single adaptor can be prepared exactly as described in Example 1 or Example 4a, the difference being the provision of adaptor hairpins with two (or more) reactive site-modified bases as described in Example 1 or 4a. Several arrangements of the reactive site-modified bases are possible, including placing a base with a reactive site near either end of the stem, or placing two reactive sites in the loop region. Multiple reactive sites can be placed on the adaptor when using only one adaptor, or when using two adaptors. Such reactive site adaptors are synthetic or purchased from DNA oligonucleotide synthesis companies, such as IDT of Coralville, IA, or EurofinsMWG of Louisville, KY.

[0264] Example 4a: Preparation and translation of a library with alternative hairpins in the reactive site adaptor.

[0265] Numerous forms of hairpins can be manufactured and used in various situations using the same scheme described in Example 1. Where smaller hairpins are advantageous, the stem can contain as few as 5 base pairs. Furthermore, hairpins containing 6-PEG linkers between complementary stem sequences can replace larger DNA loops. See Durand, M. et al., "Circular dichroism studies of an oligodeoxyribonucleotide containing a hairpin loop made of a hexaethyleneglycol chain: conformation and stability." Nucleic acids research 18.21(1990):6353-6359. For situations where multiple displays are advantageous, the distance between the multiple coding portions on a given hairpin and the placement of those coding portions may be important.

[0266] The distance between coding portions is variable and tailored to the needs of each project. The distance between coding portions can be increased or decreased by increasing or decreasing the number of bases constituting the stem by placing a linker in or near a loop region of a hairpin and a second linker in or near an anti-coding sequence or in either strand of a double-stranded stem region. Similarly, the distance between coding portions can be increased or decreased by placing a linker in or near a loop region, keeping the number of nucleotides in the stem constant, but changing the position of the second linker along the length of the stem. Optionally, if both linkers are placed in the stem region, the number of nucleotides separating them can be changed to suit the project's needs. The two linkers can optionally be placed in loop regions, and the number of bases between them can be varied to suit the project.

[0267] If the placement of the coding portion on the hairpin is important—for example, if the coding portion placed along the stem has different proximity to the target molecule compared to the coding portion placed in the loop—a hairpin with multiple loops and stems is used. In one embodiment, the hairpin may have 2 or 3 loops and 2 stems. The hairpin may include an inverse coding region connected to a first chain of a first stem region, which is connected to a first loop region, which is connected to a first chain of a second stem region, which is connected to a second loop region, which is connected to a second chain of the second stem region, which is optionally connected to a third loop region and then to the second chain of the first stem region, or directly to the second chain of the first stem region. Depending on the specific project requirements, one or more connectors are placed in one or more loops and in one or more stem regions.

[0268] Those skilled in the art will understand that numerous hairpin tertiary structures are possible, incorporating many secondary structures, including but not limited to inner rings, protrusions, and cross-shaped structures, as described below: Svoboda, P. et al., Cellular and Molecular Life Sciences CMLS, April 2006, Vol. 63, No. 7, pp. 901-908; Bikard et al., Microbiology And Molecular Biology Reviews, December 2010, pp. 570-588; Kari et al., DNA Computing Volume 3892 of the series Lecture Notes in Computer Science, pp. 158-170; Domaratzki, Theory Comput Syst (2009) 44:432-454; Brazda et al., BMC Molecular Biology 2011 12:33. Those skilled in the art will understand that hairpin oligonucleotide sequences incorporating such secondary and tertiary structures are synthesized by many DNA synthesis companies, such as Sigma-Aldrich, Integrated DNA Technologies (Coralville, Iowa), and Eurofins MWG (Louisville, KY). It should be understood that reactive sites for attaching the adapter, or modified bases with both adapter and reactive sites, can be placed at any desired location within the hairpin during the synthesis process. Those skilled in the art will understand that hairpins with more secondary structures and / or more information will tend to contain longer nucleotide sequences.

[0269] Example 4b: Preparation and translation of a library with alternative hairpins in the reactive site adaptor.

[0270] Numerous forms of hairpins can be manufactured and used in various situations using the same protocol described in Example 1. The stem region sequence of the reaction site adaptor may contain one or more restriction sites to allow cleavage in or near the stem region. Restriction digestion at these sites can release a binding agent that is very tight with the immobilized target and promote PCR amplification by removing the loop region, which will allow the primers to anneal properly. Other information may also be encoded in the reaction site adaptor hairpin DNA. One example is a series of distinct bases incorporated into the loop region. These distinct bases will help identify library members enriched in the selection due to amplification bias or as artificial products when amplified after selection. Another example is a specific sequence indicating information about the selection or synthetic history of the molecule, similar to the index sequence described in Example 7. The hairpin may also contain fluorescently labeled bases or base analogues, radiolabeled bases or base analogues, for quantifying and analyzing various aspects of the library and its synthesis or properties. The hairpin may also contain bases or modified bases with processing-promoting functional groups such as biotin. Such hair clips can be purchased from reputable suppliers of custom DNA oligonucleotides, such as IDT, Sigma Aldrich, Coralville, IA, or Eurofins MWG, Louisville, KY.

[0271] Example 4c. Using other chemistry to ligate the reaction site adaptor to the template strand. The reaction site adaptor was annealed to the terminal coding region of the template gene and ligated with T4 DNA ligase according to Example 1. Other methods of covalently tethering the reaction site adaptor can be used, including chemical or enzymatic methods. The reaction site adaptor was loaded chemically using reagents such as water-soluble carbodiimide and cyanogen bromide, as done in: Shabarova et al., (1991) Nucleic Acids Research, 19, 4247-4251; Federova et al., (1996) Nucleosides and Nucleotides, 15, 1137-1147; Grya Znov, Sergei M. et al., J. Am. Chem. Soc, Vol. 115: 3808-3809 (1993); and Carriero and Damlia (2003) Journal of Organic Chemistry, 68, 8328-8338. Optionally, chemical ligation can be performed using a 5M cyanogen bromide acetonitrile solution at a 1:10 v / v ratio with 5' phosphorylated DNA in a buffer (pH 7.6) containing 1M MES and 20mM MgCl2, with the reaction carried out at 0°C for 5 minutes. Ligation can also be performed using the manufacturer's protocol via topoisomerase, polymerase, and ligase.

[0272] Example 5: Preparation and translation of a library with a single-chain terminal coding region.

[0273] A library with a small spatial volume is prepared by removing oligonucleotides from the terminal coding region of the reaction site adaptor to make the terminal coding region and optionally all or part of the stem region single-stranded. This is performed exactly as in Example 1, with the following exceptions. Deoxyuridine is incorporated into the provided reaction site adaptor at a location in the reverse coding sequence and at a location in the stem between the end of the reverse coding region and the nearest adapter. After sequence-specific hybridization and ligation of the loaded reaction site adaptor to the template strand, the library is buffer-exchanged to 1x UDG reaction buffer from NEB, uracil-DNA glycosylase ("UDG") is added at a concentration of 20 U / ml, and incubated at 37°C for 30 min according to the manufacturer's protocol. Subsequently, it is heated to 95°C at pH 12 for 20 min to hydrolyze the pyrimidine-free sites in the hairpin. Small ssDNA fragments resulting from this are removed by size exclusion using a buffer maintained at 65°C.

[0274] Example 5a. Oligonucleotides are removed from the terminal coding region of the reactive site adaptor at several optional points during the execution of Example 1. Oligonucleotides may optionally be removed entirely according to Example 1, except that the procedure in Example 5 is performed after loading the reactive site adaptor but before adding the first positional structural unit. Oligonucleotides may optionally be removed entirely according to Example 1, except that the procedure in Example 5 is performed after adding the first positional structural unit but before adding any subsequent positional structural units. Oligonucleotides may optionally be removed entirely according to Example 1, except that the procedure in Example 5 is performed after adding all positional structural units. Those skilled in the art will understand that the task of cleaving the DNA strand at the desired location is accomplished in a variety of ways, and numerous commercially available enzymes and disclosed methods exist to facilitate this task; for example, New England Biolabs sells at least 10 nicking endonucleases and discloses their usage methods. The specific examples given herein are exemplary and do not preclude other methods for achieving the task of making the terminal coding region and optionally part of the hairpin single-stranded.

[0275] Example 6a: Using UDG to remove the 5' end non-coding region.

[0276] The restriction digestion used in Example 1 to remove the 5' uncoding region was eliminated and replaced with UDG treatment and subsequent alkaline hydrolysis of the pyrimidine-free site. The library was prepared exactly as described in Example 1, except that the oligonucleotide initiating reverse transcription incorporated a dU base at or near the 3' end of the primer. Following reverse transcription and base hydrolysis of the RNA strand, UDG removes uracil, producing a pyrimidine-free site, which is then cleaved by heat and alkali (see Example 5 for the use of UDG and reaction conditions), yielding a terminal coding region ready for ligation into an adaptor loading the reaction site. Those skilled in the art will understand that various methods exist for cleaving single-stranded or double-stranded DNA at the desired location, and numerous commercially available enzymes and disclosed protocols exist to facilitate this task. The specific examples given herein are exemplary and do not preclude other methods for achieving the task of removing the 5' uncoding region.

[0277] Example 6b: Removal of the 5' untranslated region or the 3' untranslated region using the restriction enzyme NdeI.

[0278] Restriction digestion for removing the 5' or 3' coding region, or both, is achieved by including an NdeI recognition site in the terminal non-coding region and performing restriction digestion after the reverse transcription step. NdeI has the ability to cleave RNA / DNA hybrids as well as single-stranded DNA. Therefore, NdeI is used for cleavage before or after alkaline hydrolysis of the RNA strand, or both before and after alkaline hydrolysis of the RNA strand.

[0279] Example 6c: Removal of the 5' uncoding region using RNA bases in reverse transcription primers. The exact procedure of Example 1 is used to remove the 5' uncoding region, except that the primers used in the step "Reverse Transcription of RNA into DNA" contain RNA bases. According to Example 1, after the RNA strand of the reverse transcription product is hydrolyzed, the RNA bases in the DNA primers are also hydrolyzed, thereby removing the portion of the DNA primer that serves as the RNA base 5'.

[0280] Example 7: Index molecule of formula (I).

[0281] Coding regions are reserved or added for use as indexing regions. After the library is prepared and translated according to Example 1, it is sorted on a hybridization array by reserving coding regions for indexing. Sub-pools generated by this sorting are used for different purposes, selected for different characteristics, different targets, or the same target under different conditions. Optionally, the products of different selections are independently amplified by PCR, re-pooled with other sub-pools, and re-translated as in Example 1.

[0282] Example 8a: Preparation of gene libraries by alternative methods.

[0283] Example 1 describes restriction digestion of all internal non-coding regions of the provided library gene sequence, followed by ligation, while simultaneously reclassifying all codons. This process is optionally performed stepwise rather than simultaneously. The same reaction conditions are used in step “Restriction Digestion of DNA to Uncouple All Codons from Each Other” in Example 1, except that a single restriction endonuclease is added instead of all endonucleases. The restriction digestion products are then re-ligated using the same reaction conditions as in step “Reclassifying Codons to Gene Library” in Example 1. The ligation products are purified by agarose gel electrophoresis, amplified by PCR, and then digested with the next restriction enzyme. This process is repeated until the gene library is complete.

[0284] Example 8b: Preparation of gene libraries by alternative methods.

[0285] Examples 1 and 8a describe the reclassification of all codon combinations by restriction digestion at all internal non-coding regions followed by ligation. In some embodiments, it is advantageous to reclassify incomplete codon combinations to produce a population with significantly lower complexity. Such a gene library is produced by dividing a mixture of the 96 gene sequences described in Example 1 into several aliquots. Each aliquot is then restriction-digested using different combinations of 1-3 restriction enzymes, under the reaction conditions present in step “Restriction Digestion of DNA to Uncouple All Codons from Each Other” in Example 1 or in Example 8a. After heat inactivation of the restriction enzymes, the individual digestion products are re-ligated according to the protocol in step “Reclassification of Codon Combinations to Produce a Gene Library” in Example 1. The pooled products are purified by agarose gel electrophoresis, amplified by PCR, and the remaining library preparation, translation, and selection are performed according to Example 1.

[0286] Example 8c: Gene shuffling or cross-reaction of the library. After translation and selection of the library as described in Example 1, Example 8a, or Example 8b, gene shuffling will produce new posterior phenotypes not previously present in the library, or posterior phenotypes resampled from surviving phenotypes in the selection. The selected library is amplified by PCR. The PCR product is aliquoted into a plurality of aliquots, and each aliquot is processed according to the protocol described in Example 8, or optionally according to the protocol described in step “Restriction digestion of DNA to uncouple all codons from each other and recombination and reclassification of codons to produce a gene library” in Example 1. The digested / religated products are collected, purified, and amplified as described in Example 1 or 8b, and subsequent rounds of library preparation, translation, and selection are performed according to Example 1.

[0287] Example 9: Preparation and translation of libraries with alternative reactive site functional groups and linkers.

[0288] Libraries can be prepared in several ways using different initial reactive sites derived from free amines. One method involves blocking existing initial reactive site functional groups with bifunctional molecules containing the desired initial reactive site functional groups. Libraries are prepared exactly as described in Example 1, except that in the step “Loading Reactive Site Linkers,” the peptide coupling reaction conditions listed in that step are used, each linker requiring a different initial reactive site. A bifunctional compound containing a carboxylic acid and the desired initial reactive site functional group is used to form a peptide bond with the amine containing the initial reactive site functional group. For example, 5-hydroxyvalerate can react with free amines to form a peptide bond, establishing a hydroxyl functional group as the initial reactive site for library synthesis.

[0289] The second approach involves incorporating different bases modified with different reactive sites, which can or facilitate the installation of other desired initial reactive site functional groups. One such base is 5-ethynyl-dU-CE phosphoramidamide (“ethynyl-dU”), sold by Glen Research in Virginia. Alternatively, a bifunctional linker compound with an azide and the desired initial reactive site functional group can be used for modification. For example, 5-azidopentanoic acid can be reacted with the alkynyl moiety in a “click” reaction (Huisgen reaction) under the conditions present in Example 25 to establish a carboxylic acid as the initial reactive site functional group. As another representative but not inclusive example, 5-azido-1-pentanal can be reacted with the alkynyl moiety in a “click” reaction (Huisgen reaction) to establish an aldehyde as the initial reactive site functional group. As another representative example, 4-azido,1-bromomethylbenzene can be reacted with the alkynyl moiety in a “click” reaction (Huisgen reaction) to establish a benzyl halide as the initial reactive site functional group. The base is optionally used as the initial reactive site for the alkyne in a library synthesis using a chemistry suitable for alkynes selected from Examples 10-33. Ideal initial reactive sites include, but are not limited to, amines, azides, carboxylic acids, aldehydes, alkenes, acryloyl groups, benzyl halides, α-carbonyl halides, and 1,3-dienes.

[0290] The third method involves incorporating bases modified with both the linker and the initial reactive site functional groups during the synthesis of the reactive site hairpin. For example, incorporating 5'-dimethoxytriphenylmethyl-N6-benzoyl-N8-[6-(trifluoroacetylamino)-hex-1-yl]-8-amino-2'-deoxyadenosine-3'-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramide (also known as amino-modifying agent C6 dA, purchased from Glen Research, Sterling, VA) at a critical position during hairpin synthesis establishes a free amine as the initial reactive site functional group and a 6-carbon alkyl chain as a linker, as does incorporating 5'-dimethoxytriphenylmethyl-N2-[6-(trifluoroacetylamino)-hex-1-yl]-2'-deoxyguanosine-3'-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramide (also known as amino-modifying agent C6 dG, purchased from Glen Research, Sterling, VA). During hairpin synthesis, the incorporation of 5'-dimethoxytriphenylmethyl-5-[3-methacrylate]-2'-deoxyuridine, 3'-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramide (also known as carboxyl dT, purchased from Glen Research, Sterling, VA) at a critical position establishes a carboxylic acid as the initial reactive site functional group and a 2-carbon chain as a linker. During hairpin synthesis, the incorporation of 5'-dimethoxytriphenylmethyl-5-N-((9-fluorenylmethoxycarbonyl)-aminohexyl)-3-acryloimide]-2'-deoxyuridine, 3'-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramide (also known as Fmoc-amino modifier C6 dT, Glen Research, Sterling, VA) at a critical position establishes an Fmoc-protected amine as the initial reactive site functional group and a 6-carbon alkyl chain as a linker. During hairpin synthesis, the incorporation of 5'-dimethoxytriphenylmethyl-5-(oct-1,7-diynyl)-2'-deoxyuridine, 3'-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphite (also known as C8 alkyne dT, Glen Research, Sterling VA) at a critical position establishes the alkyne as the initial reactive site functional group and the 8-carbon chain as the linker. During hairpin synthesis, the incorporation of 5'-(4,4'-dimethoxytriphenylmethyl)-5-[N-(6-(3-benzoylthiopropionyl)-aminohexyl)-3-acrylamido]-2'-deoxyuridine, 3'-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramide (also known as S-Bz thiol modifier C6-dT, Glen Research, Sterling VA) at key positions establishes the thiol as the initial reactive site functional group and the 14-atom chain as a linker.During hairpin synthesis, the incorporation of N4-TriGl-amino2'-deoxycytidine (from IBA GmbH, Goettingen, Germany) at a critical position establishes an amine as the initial reactive site functional group and a 3-ethylene glycol unit chain as a linker.

[0291] Suitable adapters perform two key functions: (i) they covalently tether hairpins to structural units, and (ii) they do not interfere with other key functions in the synthesis or use of the formula (I) molecule. Therefore, in some embodiments, adapters are alkyl or PEG chains because (a) they are highly flexible, allowing the coding moiety to be appropriately and freely presented to the target molecule during selection, and (b) they are relatively chemically inert and generally do not undergo side reactions during the synthesis of the formula (I) molecule. To adequately perform most, but not all, tasks, adapters need not contain a total length greater than approximately eight PEG units. Those skilled in the art will understand that much longer adapters and / or much stiffer adapters, such as peptide α-helices, will be useful and attractive when the library DNA must be selected as far away as possible from the target molecule, target structure, or target surface. Other desired adapters may include polyglycine, polyalanine, or peptides. Adapters incorporating fluorophores, radiolabels, or functional portions of the formula (I) molecule in a manner orthogonal to or complementary to the binding of the coding moiety are also used. For example, in some cases, it may be necessary to incorporate biotin into the adapter to immobilize the library. Incorporating a known ligand into one binding pocket of a target molecule may also be useful as a means of selecting the encoding portion of a second binding pocket that can bind to the same target molecule.

[0292] The library can be prepared by optionally using different connectors and different chemistry on different reactive site linkers. One or more connectors on the 5' reactive site linker may have one type of connector and one type of reactive site functional group, while the 3' reactive site linker may have different connectors and the same reactive site functional group, or different connectors and different reactive site functional groups. Any connectors and functional groups described herein are applicable to this embodiment, provided that the chemistry required for the subsequent mounting position structural unit is compatible with the functional groups on the first structural unit D and the second structural unit E, which react with the reactive site functional groups on their respective hairpins.

[0293] This compatibility has two modes. In the first mode, different chemistry is used to load the reaction site linker, but both the first structural unit D and the second structural unit E are capable of undergoing the same chemical transformation in the next or subsequent downstream steps. In the second mode, different chemistry is used to load the reaction site linker, and subsequent downstream steps require different chemistry. This second mode requires that the functional groups on the newly formed 5' coding portion, the functional groups on the 5' entry position structural unit, and the chemistry used for the coupling do not react with the functional groups present on the newly formed 3' coding portion. Similarly, this second mode requires that the functional groups on the newly formed 3' coding portion, the functional groups on the 3' entry position structural unit, and the chemistry used for the coupling do not react with the functional groups present on the newly formed 5' coding portion. The steps of installing structural units using orthogonal chemistry on the 3' and 5' reaction site linkers can be performed in any order. Furthermore, those skilled in the art will understand that, among the diverse structural units installed in a given synthetic step, not performing the installation of any structural unit is an important element of diversity. Suitable chemistry for these steps includes, but is not limited to, the chemistry described in Examples 10-32 and Example 1.

[0294] Example 10: Synthesis of the coding portion using Suzuki coupling chemistry.

[0295] The reactive site on the adaptor with aryl iodine as the reactive site, the DNA library which is used as the structural unit loaded on the adaptor or as a partial translation molecule, is dissolved in water at 1 mM. Add 50 equivalents of boric acid in the form of a 200 mM dimethylacetamide stock solution, 300 equivalents of sodium carbonate in the form of a 200 mM aqueous solution, 0.8 equivalents of palladium acetate in the form of a 10 mM dimethylacetamide stock solution, and a premixed 20 equivalents of 3,3',3' phosphonanetriyltris(benzenesulfonic acid) trisodium salt in the form of a 100 mM aqueous solution. React the mixture at 65°C for 1 hour, then purify by ethanol precipitation. Dissolve the DNA library in buffer to 1 mM and add 120 equivalents of sodium sulfide in the form of a 400 mM aqueous solution, then react at 65°C for 1 hour. Dilute the product to 200 μl with dH₂O and purify by ion-exchange chromatography (see Gouliaev, AH, Franch, TPO, Godskesen, MA, and Jensen, KB (2012) Bi-functional Complexes and methods for making and using such complexes. Patent Application WO 2011 / 127933 A1).

[0296] Example 11: Synthesis of the coding portion using Sonogashira coupling chemistry.

[0297] A reactive site on an adaptor containing an aryl iodine group, serving as a structural unit for loading the adaptor or as a partial translation molecule, was dissolved in water at 1 mM. 100 equivalents of an alkyne in 200 mM dimethylacetamide stock solution, 300 equivalents of pyrrolidine in 200 mM dimethylacetamide stock solution, 0.4 equivalents of palladium acetate in 10 mM dimethylacetamide stock solution, and 2 equivalents of 3,3',3'-phosphonanetriyltris(benzenesulfonic acid) trisodium salt in 100 mM aqueous solution were added. The reaction was carried out at 65 °C for 2 hours, followed by purification by ethanol precipitation or ion-exchange chromatography.(See (1) Liang, B., Dai, M., Chen, J. and Yang, Z. (2005) Cooper-free sonogashira coupling reaction with PdCl2 in water under aerobic conditions. J. Org. Chem. 70, 391-393; (2) Li, N., Lim, RKV, Edwardraja, S. and Lin, Q. (2011) Copper-free Sonogashira cross-coupling for functionalization of alkyne encoded proteins in aqueous medium and in bacterial (2) Cells (for the functionalization of alkyne-encoded proteins in aqueous media and bacterial cells, copper-free Sonogashira crosscoupling). J. Am. Chem. Soc. 133, 15316-15319; (3) Marziale, AN, Schlüter, J. and Eppinger, J. (2011) An efficient protocol for copper-free palladium-catalyzed Sonogashira crosscoupling inaqueous media at low temperatures. Tetrahedron Lett. 52, 6355-6358; (4) Kanan, MW, Rozenman, MM, Sakurai, K., Snyder, TM and Liu, DR (2004) Reaction discovery enabled by DNA-templated synthesis and in vitro selection. Nature 431, 545-549.

[0298] Example 12: Using carbamylation to synthesize the coding portion.

[0299] A DNA library containing an amine as a reactive site on an adaptor, serving as a structural unit for loading the adaptor or as a partial translation molecule, was dissolved in water at 1 mM. Triethylamine (1:4 v / v) and 50 equivalents of di-2-pyridyl carbonate in 200 mM dimethylacetamide stock solution were added. The reaction was carried out at room temperature for 2 hours, followed by the addition of 40 equivalents of amine in 200 mM dimethylacetamide stock solution at room temperature for another 2 hours. The product was purified by ethanol precipitation or ion-exchange chromatography. (See (1) Artuso, E., Degani, I. and Fochi, R. (2007) Preparation of mono-, di-, and trisubstituted ureas by carbonylation of aliphatic amines with S,S-dimethyl dithiocarbonate. Synthesis 22, 3497-3506; (2) Franch, T., Lundorf, MD, Jacobsen, SN, Olsen, EK, Andersen, AL, Holtmann, A., Hansen, AH, Sorensen, AM, Goldbech, A., De Leon, D. et al., Enzymatic encoding methods for efficient synthesis of large libraries. WIPO WO 2007 / 062664 A2, 2007.)

[0300] Example 13: Using thiourea to synthesize the coding portion.

[0301] A DNA library containing an amine as a reactive site on an adaptor, serving as a structural unit for loading the adaptor or as a partial translation molecule, was dissolved in water at 1 mM. 20 equivalents of 2-pyridylthiocarbonate in the form of a 200 mM dimethylacetamide stock solution were added at room temperature and reacted for 30 minutes. Then, 40 equivalents of amine in the form of a 200 mM dimethylacetamide stock solution were added at room temperature, and the mixture was slowly heated to 60 °C and reacted for 18 hours. The product was purified by ethanol precipitation or ion-exchange chromatography. (See Deprez-Poulain, RF, Charton, J., Leroux, V. and Deprez, BP (2007). Convenient synthesis of 4H-1,2,4-triazole-3-thiols using di-2-pyridylthionocarbonate. Tetrahedron Lett. 48, 8157-8162.)

[0302] Example 14: Synthesis of the coding moiety using the reducing monoalkylation of amines.

[0303] A DNA library containing an amine as a reactive site on an adaptor, serving as a structural unit on the adaptor or as a partial translation molecule, was dissolved in water at 1 mM. 40 equivalents of an aldehyde in the form of a 200 mM dimethylacetamide stock solution were added, and the mixture was reacted at room temperature for 1 hour. Then, 40 equivalents of sodium borohydride in the form of a 200 mM acetonitrile stock solution were added, and the mixture was reacted at room temperature for 1 hour. The product was purified by ethanol precipitation or ion-exchange chromatography. (See Abdel-Magid, AF, Carson, KG, Harris, BD, Maryanoff, CA, and Shah, RD (1996). Reductive amination of aldehydes and ketones with sodium triacetoxyborohydride. J. Org. Chem. 61, 3849-3862.)

[0304] Example 15: Synthesis of the coding part using SNAr of heteroaryl compounds.

[0305] A DNA library containing an amine as a reactive site on an adaptor, serving as a structural unit for loading the adaptor or as a partial translation molecule, was dissolved in water at 1 mM. 60 equivalents of a heteroaryl halide in a 200 mM dimethylacetamide stock solution were added, and the mixture was reacted at 60 °C for 12 h. The product was purified by ethanol precipitation or ion-exchange chromatography. (See Franch, T., Lundorf, MD, Jacobsen, SN, Olsen, EK, Andersen, AL, Holtmann, A., Hansen, AH, Sorensen, AM, Goldbech, A., De Leon, D. et al., Enzymatic encoding methods for efficient synthesis of large libraries. WIPO WO 2007 / 062664 A2, 2007.)

[0306] Example 16: Synthesis of the coding portion using Horner-Wadsworth-Emmons chemistry.

[0307] A DNA library containing an aldehyde as a reactive site on an adaptor, serving as a structural unit on the adaptor or as a partial translation molecule, was dissolved at 1 mM in borate buffer at pH 9.4. 50 equivalents of ethyl 2-(diethoxyphosphoryl)acetate in 200 mM dimethylacetamide stock solution and 50 equivalents of cesium carbonate in 200 mM aqueous solution were added, and the mixture was reacted at room temperature for 16 hours. The product was purified by ethanol precipitation or ion-exchange chromatography. (See Manocci, L., Leimbacher, M., Wichert, M., Scheuermann, J., and Neri, D. (2011) 20 years of DNA-encoded chemical libraries. Chem. Commun. 47, 12747-12753.)

[0308] Example 17: Using sulfonation to synthesize the coding portion.

[0309] A DNA library containing an amine-containing reactive site on an adaptor, serving as a structural unit on the adaptor or as a partial translation molecule, was dissolved in 1 mM borate buffer at pH 9.4. 40 equivalents of sulfonyl chloride in the form of a 200 mM dimethylacetamide stock solution were added, and the mixture was reacted at room temperature for 16 hours. The product was purified by ethanol precipitation or ion-exchange chromatography. (See Franch, T., Lundorf, MD, Jacobsen, SN, Olsen, EK, Andersen, AL, Holtmann, A., Hansen, AH, Sorensen, AM, Goldbech, A., De Leon, D., et al., Enzymatic encoding methods for efficient synthesis of large libraries. WIPO WO 2007 / 062664 A2, 2007.)

[0310] Example 18: Synthesis of the coding part using trichloro-nitro-pyrimidine.

[0311] DNA libraries containing amines as reactive sites on adaptors, serving as structural units on adaptors or as partial translation molecules, are dissolved at 1 mM in borate buffer at pH 9.4. 20 equivalents of trichloronitropyrimidine (TCNP) in 200 mM dimethylacetamide stock solution are added at 5°C. The reaction mixture is heated to room temperature over one hour and purified by ethanol precipitation. Alternatively, DNA libraries can be dissolved at 1 mM in borate buffer at pH 9.4, and 40 equivalents of amine in 200 mM dimethylacetamide stock solution and 100 equivalents of pure triethylamine are added, and the mixture is reacted at room temperature for 2 hours. The library is purified by ethanol precipitation. Alternatively, DNA libraries can be immediately dissolved in borate buffer for immediate reaction, or aggregated, re-sorted on an array, dissolved in borate buffer, and then reacted with 50 equivalents of amine in 200 mM dimethylacetamide stock solution and 100 equivalents of triethylamine, and reacted at room temperature for 24 hours. The product was purified by ethanol precipitation or ion-exchange chromatography. (See Roughley, SD and Jordan, AM (2011). The medicinal chemist's toolbox: an analysis of reactions used in the pursuit of drug candidates. J. Med. Chem. 54, 3451-3479.)

[0312] Example 19: Synthesis of the coding portion using trichloropyrimidine.

[0313] A DNA library containing an amine as a reactive site on an adaptor, serving as a structural unit for loading the adaptor or as a partial translation molecule, is dissolved at 1 mM in borate buffer at pH 9.4. 50 equivalents of 2,4,6-trichloropyrimidine in 200 mM DMA stock solution are added, and the reaction is carried out at room temperature for 3.5 hours. The DNA is precipitated in ethanol and then redissolved at 1 mM in borate buffer at pH 9.4. 40 equivalents of amine in 200 mM acetonitrile stock solution are added, and the reaction is carried out at 60–80 °C for 16 hours. The product was purified by ethanol precipitation, and the DNA library was then immediately dissolved in borate buffer for immediate reaction, or it was pooled, re-sorted on an array, and then dissolved in borate buffer, followed by reaction with 60 equivalents of boric acid in 200 mM dimethylacetamide (DMA) stock solution, 200 equivalents of sodium hydroxide in 500 mM aqueous solution, 2 equivalents of palladium acetate in 10 mM DMA stock solution, and 20 equivalents of trisodium tris(3-sulfophenyl)phosphine (TPPTS) in 100 mM aqueous solution, and reacted at 75 °C for 3 h. The DNA was precipitated in ethanol, then dissolved in water at 1 mM, and reacted with 120 equivalents of sodium sulfide in 400 mM aqueous stock solution at 65 °C for 1 h. The product was purified by ethanol precipitation or ion-exchange chromatography.

[0314] Example 20: Using Boc deprotection to synthesize the coding portion.

[0315] A Boc-protected amine was used as the reactive site on the adaptor. DNA libraries, either as structural units loading the adaptor or as partial translation molecules, were dissolved in 0.5 mM borate buffer at pH 9.4 and heated to 90°C for 16 hours. The products were purified by ethanol precipitation, size exclusion chromatography, or ion exchange chromatography. (See Franch, T., Lundorf, MD, Jacobsen, SN, Olsen, EK, Andersen, AL, Holtmann, A., Hansen, AH, Sorensen, AM, Goldbech, A., De Leon, D. et al., Enzymatic encoding methods for efficient synthesis of large libraries. WIPO WO 2007 / 062664A2, 2007.)

[0316] Example 21: Synthesis of the coding portion using the hydrolysis of tert-butyl ester.

[0317] A DNA library containing a reactive site on an adaptor with tert-butyl ester as the reactive site, or serving as a structural unit on the adaptor or as a partial translation molecule, was dissolved in borate buffer at 1 mM and reacted at 80 °C for 2 hours. The product was purified by ethanol precipitation, size exclusion chromatography, or ion exchange chromatography. (See Franch, T., Lundorf, MD, Jacobsen, SN, Olsen, EK, Andersen, AL, Holtmann, A., Hansen, AH, Sorensen, AM, Goldbech, A., De Leon, D. et al., Enzymatic encoding methods for efficient synthesis of large libraries. WIPO WO 2007 / 062664 A2, 2007.)

[0318] Example 22: Using Alloc deprotection to synthesize the coding portion.

[0319] Alloc-protected amines were used as reactive sites on the adaptor, and DNA libraries, either as structural units loaded on the adaptor or as partial translation molecules, were dissolved in 1 mM borate buffer at pH 9.4. Ten equivalents of tetra(triphenylphosphine)palladium in 10 mM DMA stock solution and 10 equivalents of sodium borohydride in 200 mM acetonitrile stock solution were added, and the mixture was reacted at room temperature for 2 hours. The product was purified by ethanol precipitation or ion-exchange chromatography. (See Beugelmans, R., Neuville, MB-C, Chastanet, J., and Zhu, J. (1995) Palladium-catalyzed reductive deprotection of Alloc: Transprotection and peptide bondformation. Tetrahedron Lett. 36, 3129.)

[0320] Example 23: Synthesis of the coding portion using the hydrolysis of methyl ester / ethyl ester.

[0321] A DNA library containing a reactive site on an adaptor with methyl or ethyl ester as the reactive site, serving as a structural unit on the adaptor or as a partial translation molecule, was dissolved in 1 mM borate buffer and reacted with 100 equivalents of NaOH at 60 °C for 2 hours. The product was purified by ethanol precipitation, size exclusion chromatography, or ion exchange chromatography. (See Franch, T., Lundorf, MD, Jacobsen, SN, Olsen, EK, Andersen, AL, Holtmann, A., Hansen, AH, Sorensen, AM, Goldbech, A., De Leon, D. et al., Enzymatic encoding methods for efficient synthesis of large libraries. WIPO WO 2007 / 062664A2, 2007.)

[0322] Example 24: Synthesis of the coding portion using the reduction of nitro groups.

[0323] A DNA library containing a nitro group as a reactive site on an adaptor, serving as a structural unit on the adaptor or as a partial translation molecule, is dissolved in water at 1 mM. 10% (v / v) of Raney nickel slurry and 10% (v / v) of hydrazine in a 400 mM aqueous solution are added, and the mixture is reacted at room temperature with shaking for 2–24 hours. The product is purified by ethanol precipitation or ion-exchange chromatography. (See Balcom, D. and Furst, A. (1953) Reductions with hydrazine hydrate catalyzed by Raney nickel. J. Am. Chem. Soc. 76, 4334-4334.)

[0324] Example 25: Using "click" chemistry to synthesize the coding portion.

[0325] The reactive site on the adaptor containing an alkyne or azide group, the structural unit loading the adaptor, or the DNA library as a partial translation molecule, is dissolved at 1 mM in 100 mM phosphate buffer. Copper sulfate is added to 625 μM, THPTA (ligand) to 3.1 mM, aminoguanidine to 12.5 mM, ascorbate to 12.5 mM, and azide to 1 mM (if the DNA contains an alkyne) or alkyne to 1 mM (if the DNA contains an azide). The reaction is carried out at room temperature for 4 hours. The product is purified by ethanol precipitation, size exclusion chromatography, or ion exchange chromatography. (See Hong, V., Presolski, Stanislav I., Ma, C. and Finn, MG (2009), Analysis and Optimization of Copper-Catalyzed Azide-Alkyne Cycloaddition for Bioconjugation. Angewandte Chemie International Edition, 48: 9879-9883.)

[0326] Example 26: Synthesis of a combination containing the coding portion of benzimidazole.

[0327] A DNA library containing an aryl o-diamine as a reactive site on an adaptor, serving as a structural unit for loading the adaptor or as a partial translation molecule, was dissolved at 1 mM in borate buffer at pH 9.4. 60 equivalents of an aldehyde in the form of a 200 mM DMA stock solution were added, and the mixture was reacted at 60°C for 18 hours. The product was purified by ethanol precipitation or ion-exchange chromatography. (See (1) Mandal, P., Berger, SB, Pillay, S., Moriwaki, K., Huang, C, Guo, H., Lich, JD, Finger, J, Kasparcova, V., Votta, B. et al., (2014) RIP3 induces apoptosis independent of pronecrotic kinase activity. Mol. Cell 56, 481-495; (2) Gouliaev, AH, Franch, TP-O., Godskesen, MA and Jensen, KB (2012) Bi-functional Complexes and methods for making and using such complexes. Patent application WO 2011 / 127933 A1; (3) Mukhopadhyay, C and Tapaswi, PK (2008) Dowex 50W: A highly (Dowex 50W: An efficient and recyclable green catalyst for the construction of the 2-substituted benzimidazole moiety in aqueous medium. Catal. Commun. 9, 2392-2394.)

[0328] Example 27: Synthesis of a coding portion containing an imidazolidine ketone.

[0329] A DNA library containing a reactive site on an adaptor with an α-amino-amide as the reactive site, serving as a structural unit on the adaptor or as a partial translation molecule, was dissolved at 1 mM in a 1:3 methanol:borate buffer at pH 9.4. 60 equivalents of an aldehyde in the form of a 200 mM DMA stock solution were added, and the mixture was reacted at 60°C for 18 hours. The product was purified by ethanol precipitation or ion-exchange chromatography. (See (1) Barrow, JC, Rittle, KE, Ngo, PL, Selnick, HG, Graham, SL, Pitzenberger, SM, McGaughey, GB, Colussi, D., Lai, M.-T., Huang, Q. et al., (2007) Design and synthesis of 2,3,5-substituted imidazolidin-4-one inhibitors of BACE-1. Chem. Med. Chem. 2, 995-999; (2) Wang, X.-J, Frutos, RP, Zhang, L., Sun, X., Xu, Y., Wirth, T., Nicola, T., Nummy, LJ, Krishnamurthy, D., Busacca, CA, Yee, N. and Senanayake, CH (2011) Asymmetric synthesis of LFA-1 inhibitor BIRT2584 on metric ton scale. Org. Process Res. Dev. 15, 1185-1191; (3) Blass, BE, Janusz, JM, Wu, S., Ridgeway, JMII, Coburn, K., Lee, W., Fluxe, AJ, White, RE, Jackson, CM and Fairweather, N. 4-Imidazolidinones as KV 1.5 Potassium channel inhibitors. WIPO WO2009 / 079624 A1, 2009.

[0330] Example 28: Synthesis of a quinazolinone-containing coding portion.

[0331] A DNA library containing a reactive site on an adaptor with 2-aniline-1-benzamide as the reaction site, or as a structural unit loading the adaptor or as a partial translation molecule, was dissolved in 1 mM borate buffer at pH 9.4. 200 equivalents of NaOH in 1 M aqueous solution and an aldehyde in 200 mM DMA stock solution were added, and the reaction was carried out at 90 °C for 14 h. The product was purified by ethanol precipitation or ion-exchange chromatography. (See Witt, A. and Bergmann, J. (2000) Synthesis and reactions of some 2-vinyl-3H-quinazolin-4-ones. Tetrahedron 56, 7245-7253.)

[0332] Example 29: Synthesis of a coding portion containing isoindolinetone.

[0333] A DNA library containing a reactive site on an adaptor with an amine as the reactive site, or as a structural unit loading the adaptor or as a partial translation molecule, was dissolved at 1 mM in borate buffer at pH 9.4. 4-Bromo,2-enyl methyl ester in the form of a 200 mM DMA stock solution was added, and the mixture was reacted at 60 °C for 2 hours. The product was purified by ethanol precipitation or ion-exchange chromatography. (See Chauleta, C, Croixa, C, Alagillea, D., Normand, S., Delwailb, A., Favotb, L., Lecronb, J.-C and Viaud-Massuarda, MC (2011). Design, synthesis and biological evaluation of new thalidomide analogues as TNF-α and IL-6 production inhibitors. Bioorg. Med. Chem. Lett. 21, 1019-1022.)

[0334] Example 30: Synthesis of a thiazole coding portion.

[0335] A reactive site on an adaptor containing thiourea as the reaction site, a DNA library serving as a structural unit on the adaptor or as a partial translation molecule, was dissolved in 1 mM borate buffer at pH 9.4. 50 equivalents of bromoketone in 200 mM DMA stock solution were added, and the mixture was reacted at room temperature for 24 hours. The product was purified by ethanol precipitation or ion-exchange chromatography. (See Potewar, TM, Ingale, SA, and Srinivasan, KV (2008). Catalyst-free efficient synthesis of 2-aminothiazoles in water at ambient temperature. Tetrahedron 64, 5019-5022.)

[0336] Example 31: Synthesis of a coding portion containing imidazopyridine.

[0337] A DNA library containing an aryl aldehyde as a reactive site on an adaptor, serving as a structural unit for loading the adaptor or as a partial translation molecule, was dissolved at 1 mM in borate buffer at pH 9.4. 50 equivalents of 2-aminopyridine in 200 mM DMA stock solution and 2500 equivalents of NaCN in 1 M aqueous solution were added, and the mixture was reacted at 90 °C for 10 hours. The product was purified by ethanol precipitation or ion-exchange chromatography. (See (1) Alexander Lee Satz, Jianping Cai, Yi Chen, Robert Goodnow, Felix Gruber, Agnieszka Kowalczyk, Ann Petersen, Goli Naderi-Oboodi, Lucja Orzechowski and Quentin Strebel. DNA Compatible Multistep Synthesis and Applications to DNA Encodedlibraries. Bioconjugate Chemistry 2015 26(8), 1623-1632; (2) Beat, GN, Liu, Y. and Plouvier, BMCPCT International Application 2001096335, December 20, 2001; (3) Inglis, SR, Jones, RK, Booker, GW and Pyke, SM (2006) Synthesis of N-benzylated-2-aminoquinolines as ligands for the Tec SH3 domain. Synthesis of N-benzylated 2-aminoquinoline ligands with SH3 domain. (Bioorg. Med. Chem. Lett. 16, 387-390.)

[0338] Example 32: Using various other chemicals to synthesize the coding portion.

[0339] The Handbook for DNA-Encoded Chemistry (edited by Goodnow RA, Jr.), 2014 Wiley, New York, lists 31 types of compatible chemical reactions. These include the SNAr reaction of trichlorotriazine, oxidation of diols to glyoxal compounds, Msec deprotection, Ns deprotection, Nvoc deprotection, pentenoyl deprotection, indole-styrene coupling, Diels-Alder reaction, Wittig reaction, Michael addition, Heck reaction, Henry reaction, and 1,3-dipolar cycloaddition of nitroketones with activated alkenes. The formation of azoles, deprotection of trifluoroacetamide, oxidative coupling of olefins and alkynes, ring-closing metathesis, and aldol reactions are disclosed in this reference. Other reactions with the potential to function in the presence of DNA and suitable for use are also disclosed.

[0340] Example 33: Using different restriction enzymes in library preparation.

[0341] It should be understood that the restriction enzymes mentioned in other embodiments are representative, and other restriction enzymes may serve the same purpose under equal or advantageous circumstances.

[0342] Example 34. Another method for preparing a gene library. The library is prepared using a coding region of about 4 to about 40 nucleotides. The library is prepared and translated according to Example 1, with the following exceptions. The library is constructed by purchasing two sets of oligonucleotides: a coding strand set and an anticoding strand set. Each set contains as many subsets as the present coding region, and each subset contains as many different sequences as the different coding sequences of the coding region. Each oligonucleotide in each subset of the coding strand oligonucleotides contains a coding sequence and optionally a 5' noncoding region. Each oligonucleotide in each subset of the anticoding strand oligonucleotides contains an anticoding sequence and optionally a complementary sequence to the 5' noncoding region. To facilitate downstream ligation in this process, all oligonucleotides except for the 5' ends of the coding and anticoding strands are purchased with 5' phosphorylated, or phosphorylated with T4 PNK from NEB according to the manufacturer's protocol. The subset of oligonucleotides with the coding sequence at the 5' end of the coding strand is combined with the subset with the anticoding sequence at the 3' end in T4 DNA ligase buffer from NEB, allowing hybridization between the two sets. The resulting products include a single-stranded 5' overhang non-coding region on the coding strand, a double-stranded coding region, and an optional single-stranded 5' overhang non-coding region on the anticoding strand. This hybridization procedure is performed separately for each coding / anticoding pair of the oligonucleotide subset. For example, a subset encoding the second coding region from the 5' end is hybridized with its complementary anticoding subset, a subset encoding the third coding region from the 5' end is hybridized with its complementary subset, and so on. The hybridized subset pairs are pooled and optionally purified by agarose gel electrophoresis. If the genes in the library have non-coding regions longer than 1 base, and if the non-coding regions between coding regions are unique, equimolar amounts of each hybridized subset pair are added to a single container. The single-stranded non-coding regions are hybridized and ligated to each other using the manufacturer's protocol via T4 DNA ligase from NEB. If the non-coding regions are longer than 1 base but not unique, two adjacent hybridized subsets are added to a container, the single-stranded non-coding regions are annealed, and ligated using T4 DNA ligase. After the reaction is complete, the product is optionally purified by agarose gel electrophoresis, and a third hybridization subset adjacent to one end of the ligated product is added, annealed, and ligated. This process is repeated until the library construction is complete. It should be understood that this method constructs libraries containing any number of coding regions. For the present purpose, libraries with more than 20 coding regions may be impractical for reasons unrelated to library construction. It should be understood that those skilled in the art typically perform blunt-end ligation, and coding regions are ligated without inserting non-coding regions; however, for hybridization subsets that do not have non-coding regions at either end, ligation provides both sense and antisense products. By preparing the library and sequentially sorting it on all hybridization arrays, the product with the correct sense is purified from the product with the antisense. The portion of the library captured on the array at each hybridization step has the correct sense.It should be understood that non-coding regions containing only unique restriction site sequences are an attractive option for this method.

[0343] Example 34. Constructing a gene library with corresponding terminal coding regions. Construct a library in which the 5' terminal coding region and the 3' terminal coding region encode the same structural unit or the same pair of different structural units. This can be achieved if each member of a gene library with a given 5' terminal coding sequence has only one 3' terminal coding sequence. Construct such a library using the method of Example 33, except that hybridization subset pairs for the 5' terminal coding region are not pooled, and the 3' terminal coding regions are not pooled. Pool and ligate all internal coding regions as in Example 33. Divide the product of ligating all internal coding regions into aliquots, and add one aliquot to each subset sequence of the 5' terminal hybridization and ligate. The ligation product in each well has a single 5' terminal coding sequence, but is a combined mixture of all sequences at all internal coding regions. These ligation products with a single 5' terminal coding sequence are independently transferred to wells containing a single 3' terminal hybridization subset sequence and ligated. The product in each well is a gene containing a combined mixture of a single 5' terminal coding sequence, a single 3' terminal coding sequence, and all sequences at all internal coding regions. It should be understood that there are other ways to generate the same resulting library.

[0344] Example 35. Using an alternative method to perform the selection of binding target molecules. Selection of library members capable of binding target molecules was performed according to Example 1, the difference being that the target molecules were immobilized on a plastic plate, such as... plate, Plates or other plates commonly used to immobilize biomolecules for ELISA, or where target molecules are biotinylated and immobilized on streptavidin-coated or neutral avidin-coated surfaces, or avidin-coated surfaces, including magnetic beads, beads made of synthetic polymers, beads made of polysaccharides or modified polysaccharides, plate wells, tubes, and resins. It should be understood that the selection of library members with the desired traits will be performed in buffers compatible with DNA, compatible with maintaining any target molecule in its native conformation, compatible with any enzymes used in the selection or amplification process, and compatible with the identification of trait-positive library members. These buffers include, but are not limited to, buffers made with phosphates, citrates, and TRIS. Such buffers may also include, but are not limited to, salts of potassium, sodium, ammonium, calcium, magnesium, and other cations, as well as salts of chloride, iodide, acetate, phosphate, citrate, and other anions. Such buffers may include, but are not limited to, surfactants, such as… TRITON TM And Chaps (3-[(3-cholamidopropyl)dimethylammonium]-1-propanesulfonate).

[0345] Example 36. Selection of binders with low dissociation rates. As described in Example 1, selection was performed to identify individuals in a library population capable of binding target molecules. Individuals binding target molecules with low dissociation rates were selected as follows. Target molecules were immobilized by biotinylation and incubated together with a streptavidin-coated surface, or optionally immobilized on a plastic surface without biotinylation. The library population is incubated with the immobilized target in an appropriate buffer for 0.1 to 8 hours. The incubation duration depends on the estimated copy number of each individual library member in the sample and the number of immobilized target molecules. Higher individual copy numbers and target molecule loads may result in shorter durations. Lower copy numbers and / or target molecule loads may result in longer durations. The goal is to ensure that each individual in the population has the opportunity to fully interact with the target. After incubating the library with the immobilized target, it is assumed that the binder in the library is bound to the target. At this point, excess unimmobilized target is added to the system, and incubation continues for approximately 1 to approximately 24 hours. Any individual bound to the immobilized target with a high dissociation rate may be released from the immobilized target and, upon rebinding, separate into those bound to the free target and those bound to the immobilized target. Individuals bound with a low dissociation rate will remain bound to the immobilized target. The immobilized surface is washed to preferentially remove non-binding agents and binding agents with fast dissociation rates, thereby selecting individuals with low dissociation rates. DNA encoding a low-dissociation-rate binding agent is amplified according to Example 1.

[0346] Example 37. Selection using a mobile target. Selection is performed where the target molecule is biotinylated and then incubated with the library for an appropriate duration. The mixture is then immobilized on a surface such as streptavidin, thus immobilizing the target and any library members that bind to the target. The surface is washed to remove the non-binding agent. Amplification of the DNA encoding the binding agent is performed according to Example 1.

[0347] Example 38. Target-specific selection. Selection is performed to identify individuals in the library population that bind to the desired target molecule, excluding other anti-target molecules. The anti-target molecule (or multiple anti-target molecules, if more than one) is biotinylated and immobilized on a streptavidin-coated surface, or optionally immobilized on a plastic surface such as... Plates or other plates suitable for binding proteins are used for ELISA-like assays. In separate containers, target molecules are immobilized by biotinylation and incubated with a streptavidin-coated surface, or optionally immobilized on a plastic surface such as… The library is first incubated with the anti-target. This depletes the population of individuals binding to the anti-target molecules. After incubation with the anti-target, the library is transferred to a container containing the desired target and incubated for an appropriate duration. Washing removes the non-binding agent. Amplification of the DNA encoding the low dissociation rate binding agent is performed according to Example 1. The identified target binding agent has a higher probability of selectively binding to the target compared to the anti-target. Optionally, target affinity selection is performed by immobilizing the target, over-adding free mobile anti-target, then adding the library and incubating for an appropriate duration. Under this protocol, individuals with affinity for the anti-target are preferentially bound by the anti-target because it is present in excess and can therefore be removed during surface washing. Amplification of the DNA encoding the binding agent is performed according to Example 1.

[0348] Example 39. Selection Based on Differential Mobility. Selection is based on the ability of individual library members to interact with the target molecule or macromolecular structure in the complex formed when library members interact with the target molecule or macromolecular structure. Interaction between the target molecule or structure and the library members is allowed, and then the mixture is passed through a size exclusion medium, resulting in the physical separation of library members that do not interact with the target molecule or structure from those that do interact, because the complex of the interacting library member and the target molecule or structure will be larger than that of the non-interacting library member, and therefore moves through the medium with a different mobility. It should be understood that in the absence of a size exclusion medium, the difference in mobility can be a function of diffusion, and mobility can be induced by various means, including but not limited to gravitational flow, electrophoresis, and diffusion.

[0349] Example 40. General Strategies for Other Options. Those skilled in the art will understand that selection can be made for virtually any trait, provided that the designed assay (a) physically separates individuals in a library population possessing the desired trait from those not possessing it, or (b) allows DNA-coding individuals in a library population possessing the desired trait to be preferentially amplified compared to DNA-coding library members not possessing said trait. Many target molecule immobilization methods are known in the art, including labeling target molecules with His tags and immobilizing them on a nickel surface, labeling target molecules with flag tags and immobilizing them with anti-flag antibodies, or labeling target molecules with adapters and covalently immobilizing them on a surface. It should be understood that the order in which library members are allowed to bind to the target and the order in which the target is immobilized is as indicated or achievable by the immobilization method used. It should be understood that selection is made where immobilization or physical separation of trait-positive and trait-negative individuals is not required. For example, trait-positive individuals recruit factors capable of amplifying their DNA, while trait-negative members do not. Trait-positive individuals are labeled with PCR primers, while trait-negative individuals are not labeled. Any procedure is suitable for differentially amplified individuals with positive traits.

[0350] Example 41a. Chemistry for loading reaction site integrators. It should be understood that any chemistry described in Examples 10-32 is suitable for loading reaction site integrators. The structural units are loaded onto the reaction site in aqueous solution, in an aqueous / organic mixture, or when fixed to a solid support. The chemistry for loading the structural units onto the reaction site integrators is not limited to reactions carried out when the reaction site integrators are fixed to a solid support such as DEAE or Super Q650M; nor is it limited to reactions carried out in the solution phase.

[0351] Example 41b. Deletion of structural units is a coded element of diversity. During library synthesis, diversity arises when multiple structural units are independently mounted on various library sub-pools with different sequences. Deletion of structural units is an optional element of diversity. Deletion of structural units is encoded exactly as in Example 1, except that in the required chemical steps, one or more sequence-specific sub-pools of the library are not chemically treated to mount the structural units. In this case, the sequences of those sub-pools thus encode the deletion of structural units.

[0352] Example 42. Hybridization arrays containing other materials. Hybridization arrays can perform two key tasks: (a) they can sort heterogeneous mixtures of at least a portion of single-stranded DNA through sequence-specific hybridization, and (b) the arrays can enable or allow the independent removal of sorted subpools from the array. Arrays in which anti-coding oligonucleotides are immobilized can be characterized by any three-dimensional orientation arrangement that meets the above criteria, but two-dimensional rectangular grid arrays are currently the most attractive because a large number of commercially available laboratory instruments are already mass-produced in this format (e.g., 96-well plates, 384-well plates).

[0353] A solid support in an array of immobilized anti-coding oligonucleotides can achieve four tasks: (a) it can permanently immobilize the anti-coding oligonucleotides, (b) it can enable or allow capture of library DNA by sequence-specific hybridization with the immobilized oligonucleotides, (c) it can have low background or non-specific binding to the library DNA, and (d) it can be chemically stable to processing conditions, including steps performed at high pH. CM It has been proven that the amine of azido-PEG-amine and CM Peptide bonds are formed between carboxyl groups on the resin surface, which is then functionalized with azide-PEG-amine (with 9 PEG units). The reverse-encoded oligonucleotide with alkynyl modifier is "clicked" onto the azide in a copper-mediated 1,3-dipolar cycloaddition (Huisgen).

[0354] Other suitable solid supports include hydrophilic beads, or polystyrene beads with a hydrophilic surface coating, polymethyl methacrylate beads with a hydrophilic surface coating, and other beads with hydrophilic surfaces that also have reactive functional groups such as carboxylic esters, amines, or epoxides, to which appropriately functionalized reverse-coding oligonucleotides are immobilized. Other suitable supports include monolithic materials and hydrogels. See, for example, J Chromatogr A. June 14, 2002; 959(1-2):121-9; J Chromatogr A. April 29, 2011; 1218(17):2362-7; J Chromatogr A. December 9, 2011; 1218(49):8897-902; Trends in Microbiology, Vol. 16, No. 11, 543-551; J. Polym. Sci. A Polym. Chem, 35:1013-1021; J. Mol. Recognit. 2006; 19:305-312; J. Sep. Sci. 2004, 27, 828-836. Generally, solid supports with larger surface areas capture more library DNA, and beads with smaller diameters produce much higher back pressure and resistance to flow. These limitations have been partially mitigated by using porous supports or hydrogels with very high surface area but low back pressure. Typically, positively charged beads produce a greater degree of non-specific DNA binding.

[0355] The chassis of a hybridization array performs three tasks: (a) it must maintain physical separation between features, (b) it must enable or allow the library to flow through or through the features, and (c) it must enable or allow the independent removal of sorted library DNA from different features. The chassis can be constructed from any material that is sufficiently rigid, chemically stable under processing conditions, and compatible with any means required to fix supports within the features. Typical materials for the chassis include plastics such as... Or polyetheretherketone (PEEK), ceramics, and metals such as aluminum or stainless steel.

[0356] Example 43. Gene Library Parameters. A gene library can contain 2 to 20 coding regions. The number of available coding sequences in each internal coding region is limited only by the number of features with available immobilized inverse coding sequences. Given the abundance of industry-standard 24-well, 96-well, and 384-well laboratory instruments, using these numbers of coding sequences in coding regions is convenient, but coding regions with, for example, 768 or 1536 coding sequences are also practical. Terminal coding sequences are not sorted on the array, therefore the number of sequences used in terminal coding regions needs to conform to the well counts in industry-standard plates and laboratory instruments. In principle, using 96 or 960 or more different coding sequences in terminal coding regions would be practical.

Claims

1. A probe molecule, wherein the probe molecule is according to formula (I), (I)([(B1) M —D—L1] Y —H1) O —G—(H2—[L2—E—(B2) K ] W ) P in G is an oligonucleotide comprising at least two coding regions for encoding a positional structural unit and at least one terminal coding region for encoding a first structural unit or a second structural unit, wherein the at least two coding regions are single-stranded and the at least one terminal coding region is single-stranded or double-stranded; H1 is a hairpin structure containing an oligonucleotide, comprising a loop portion, a stem portion, and a 5' single-stranded portion, wherein H1 contains a 5' end complementary to one end of the oligonucleotide G; H2 is a hairpin structure containing an oligonucleotide, comprising a loop portion, a stem portion, and a 3' single-stranded portion, wherein H2 contains a 3' end complementary to one end of the oligonucleotide G; D is the first structural unit; E is the second structural unit, where D and E may be the same or different; B1 is a positional structure unit and M represents an integer from 1 to 20; B2 is a positional structural unit and K represents an integer from 1 to 20, where B1 and B2 are the same or different, and M and K are the same or different; L1 is a connector that covalently binds D to the loop or stem portion of H1; L2 is a connector that covalently binds E to the loop or stem portion of H2; O is an integer between 0 and 1; P is an integer between 0 and 1; The condition is that at least one of O and P is 1; When the value of O is 1, Y is an integer from 1 to 5; When the value of P is 1, W is an integer from 1 to 5; and At least one of the positional structural unit B1 at position M and the positional structural unit B2 at position K is identified by one of the at least two coding regions, and at least one of the first structural unit D and the second structural unit E is identified by the at least one end coding region. At least one of H1 or H2 contains a unique restriction site or at least one deoxyuridine (dU) base.

2. The probe molecule according to claim 1, wherein the multiple positional structural units at B1 are different.

3. The probe molecule according to claim 1 or 2, wherein the multiple positional structural units at B2 are different.

4. The probe molecule according to any one of claims 1 to 3, wherein L1 covalently binds D to the loop portion or stem portion of H1, and wherein L2 covalently binds E to the loop portion or stem portion of H2.

5. The probe molecule according to any one of claims 1 to 4, wherein at least one of Y or W is at least 2 and / or both O and P are 1.

6. The probe molecule according to any one of claims 1 to 5, wherein both O and P are 1.

7. The probe molecule according to any one of claims 1 to 6, wherein G comprises the components of formula (C N —(Z N —C N+1 ) A The sequence is represented by ), where C is the coding region, Z is the non-coding region, N is an integer from 1 to 20, and A is an integer from 1 to 20; Each non-coding region contains 4 to 50 nucleotides and is optionally double-stranded.

8. The probe molecule according to any one of claims 1 to 7, wherein one of O or P is 0.

9. The probe molecule according to any one of claims 1 to 8, wherein each coding region contains 6 to 50 nucleotides.

10. The probe molecule according to any one of claims 1 to 9, wherein at least one of H1 and H2 comprises 20 to 90 nucleotides.

11. The probe molecule according to any one of claims 1 to 10, wherein each coding region contains 12 to 40 nucleotides.

12. The probe molecule according to any one of claims 1 to 11, wherein P is 0, Y is 2, and each coding region contains 12 to 40 nucleotides.

13. The probe molecule according to any one of claims 1 to 12, wherein O is 0, W is 2, and each coding region contains 12 to 40 nucleotides.

14. The probe molecule according to any one of claims 1 to 13, wherein H1 or H2 of the probe molecule contains a unique restriction site.

15. The probe molecule according to any one of claims 1 to 14, wherein at least one of H1 and H2 of the probe molecule comprises at least one dU base.

16. A method for analyzing a probe molecule according to claim 14, the method comprising: The probe molecule is treated with a restriction enzyme that recognizes the unique restriction site; and Polymerase chain reaction (PCR) is performed on the treated probe molecules. The method described herein is for purposes other than diagnosing or treating a disease.

17. A method for analyzing a probe molecule according to claim 15, the method comprising: The probe molecule was treated with uracil DNA glycosylation enzyme; and Polymerase chain reaction (PCR) is performed on the treated probe molecules. The method described herein is for purposes other than diagnosing or treating a disease.

Citation Information

Patent Citations

  • Polynucleotide-array assay and methods

    US5759779A

  • BI-functional complexes and methods for making and using such complexes

    WO2011127933A1

  • Method for the synthesis of a bifunctional complex

    CN102838654A

  • Nucleic acid encoding reaction

    CN103890245A