Techniques for discovering and selecting aptamers

JP2026517323APending Publication Date: 2026-05-29ILLUMINA INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ILLUMINA INC
Filing Date
2024-04-15
Publication Date
2026-05-29

Smart Images

  • Figure 2026517323000001_ABST
    Figure 2026517323000001_ABST
Patent Text Reader

Abstract

A library of coded aptamer candidate regions having nucleotide modifications is described. Each aptamer candidate region comprises a first conserved primer region and a second conserved primer region. The aptamer candidate also comprises a variable region located between the first and second conserved primer regions and containing at least one modified nucleotide. The coding region contains a nucleotide sequence specific to the modification type of at least one modified nucleotide, so that the sequence of each aptamer candidate can be used to identify the associated modification type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Reference to Electronic Sequence Listing) This application includes a sequence listing electronically submitted in XML format, which is hereby incorporated by reference in its entirety. The XML copy created on March 30, 2023, is designated as "IP-2332-PRV.xml" and has a size of 11,047 bytes. The sequence listing contained in this XML file is part of this specification and is hereby incorporated by reference in its entirety.

Background Art

[0002] The technology disclosed herein generally relates to techniques for the discovery, generation, and / or selection of aptamers. In particular, the disclosed technology relates to techniques for improved and more diverse aptamer candidates for use in an aptamer library for selecting aptamers for aptamer-based assays or therapeutic applications.

[0003] The subject matter discussed in this section should not be assumed to be prior art merely as a result of mention in this section. Similarly, the problems mentioned in this section, or problems associated with the subject matter provided as background, should not be assumed to have been previously recognized in the prior art. The subject matter of this section merely represents different approaches and, in itself, may also correspond to embodiments of the claimed technology.

[0004] Protein expression patterns help define the identity and state of cells. RNA transcripts are often used as a surrogate for protein expression, but the relationship between the abundance of a protein and its mRNA is not one-to-one. There are differences caused by regulation of post-transcriptional degradation, translational degradation, and proteolysis. Therefore, direct nucleic acid sequencing of RNA transcripts may not provide an accurate estimate of protein expression.

[0005] Aptamers are nucleic acids that bind to molecular targets, such as proteins, with high affinity and specificity. Advances in aptamer selection and design include the systematic evolution of ligands by exponential enrichment (SELEX). In SELEX, high-affinity nucleic acids for different target analytes can be isolated from combinatorial libraries, enabling high-throughput characterization of aptamer-target binding and multiplexed assays for analytes in complex biological samples. When an aptamer binds to an analyte target, the binding event can be detected to characterize the presence and concentration of various analytes in the biological sample. However, the available panel of aptamers for use in protein detection is a function of the aptamer discovery process, and not all desired protein targets can produce aptamers with appropriate binding characteristics. Therefore, the generation and discovery of improved aptamers would be beneficial in expanding the range of available protein targets. [Overview of the project]

[0006] In one embodiment, the disclosure provides an aptamer candidate library comprising a plurality of at least partially single-stranded nucleic acids. Each of the plurality of at least partially single-stranded nucleic acids comprises: a first conserved primer region conserved among the plurality of at least partially single-stranded nucleic acids; a second conserved primer region conserved among the plurality of at least partially single-stranded nucleic acids; a variable region located between the first and second conserved primer regions, wherein the variable region is variable among the plurality of at least partially single-stranded nucleic acids, and each variable region comprises at least one modified nucleotide; and a coding region comprising a nucleotide sequence specific to the modification type of at least one modified nucleotide.

[0007] In one embodiment, the Disclosure provides an aptamer selection method comprising providing a library of aptamer candidates of different aptamer subgroups, wherein each aptamer candidate of an individual subgroup of the different aptamer subgroups comprises: a first conserved primer region conserved among the aptamer candidates; a second conserved primer region conserved among the aptamer candidates; a variable region located between the first and second conserved primer regions, wherein the variable region is variable among the aptamer candidates and contains at least one modified nucleotide; and a coding region comprising a nucleotide sequence that is specific to the modification type of at least one modified nucleotide and uniquely identifies the modification type from different modification types of other subgroups of the different subgroup. The method also comprises selecting aptamer candidates based on binding to a target molecule and amplifying the selected aptamer candidates using primers based on the first and second conserved primer regions. This includes sequencing the amplified aptamer candidates, determining the sequence of the variable region, and identifying the modification type.

[0008] In one embodiment, the disclosure provides an aptamer candidate library comprising a plurality of partially double-stranded aptamer candidates including a first strand and a second strand. The first strand comprises a double-stranded portion including a first complementary region that hybridizes to a second complementary region of the second strand, and a single-stranded portion comprising a first conserved primer region conserved between the first and second strands of a plurality of partial double-stranded aptamer candidates, a second conserved primer region conserved between the first and second strands of a plurality of partial double-stranded aptamer candidates, a variable region located between the first and second conserved primer regions, the variable region being variable between the first and second strands of a plurality of partial double-stranded aptamer candidates and comprising at least one modified nucleotide, and a coding region comprising a nucleotide sequence specific to the modification type of at least one modified nucleotide.

[0009] The foregoing description is provided to enable the fabrication and use of the disclosed technology. Various modifications to the disclosed embodiments are apparent, and the general principles defined herein may be applied to other embodiments and uses without departing from the spirit and scope of the disclosed technology. Accordingly, the disclosed technology is not intended to be limited to the embodiments shown, but rather to be given the broadest scope consistent with the principles and features disclosed herein. The scope of the disclosed technology is defined by the appended claims. [Brief explanation of the drawing]

[0010] These and other features, aspects, and advantages of the disclosed embodiments will be better understood by reading the following detailed description with reference to the accompanying drawings, and similar features are represented in similar parts across the drawings. [Figure 1] This is a schematic diagram of the prior art SELEX method. [Figure 2] This is a schematic diagram of a SELEX method for selecting aptamers from a candidate library having multiple modifications, according to one embodiment of the present disclosure. [Figure 3] This is a schematic diagram of amplification and characterization of a selected aptamer candidate according to one embodiment of the present disclosure. [Figure 4] This is a schematic diagram of a SELEX method for selecting aptamers that facilitate double-strand formation, according to one embodiment of the present disclosure. [Figure 5] This is a schematic diagram of amplification and characterization of selected aptamer double-stranded candidates according to one embodiment of the present disclosure. [Figure 6] This disclosure shows exemplary uridine modifications that may be incorporated into an aptamer according to certain embodiments of this disclosure. [Figure 7] This is a schematic diagram of a conserved nucleotide region and a variable nucleotide region in an aptamer nucleotide sequence according to one embodiment of the present disclosure. [Figure 8] This is a schematic diagram of a SELEX selective cycle having aptamer amplification according to one embodiment of the present disclosure. [Figure 9] This is a schematic diagram of a SELEX selective cycle having aptamer amplification according to one embodiment of the present disclosure. [Figure 10] This is a block diagram of a sequencing device configured to acquire sequencing data according to this technique. [Modes for carrying out the invention]

[0011] The following considerations are presented to enable those skilled in the art to fabricate and use the disclosed technology and are provided in relation to specific uses and their requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and uses without departing from the spirit and scope of the disclosed technology. Accordingly, the disclosed technology is not intended to be limited to the embodiments shown, but is given the broadest scope consistent with the principles and features disclosed herein.

[0012] Aptamers are short, single-stranded nucleic acid molecules (ssDNA or ssRNA) or modified nucleic acids that can bind with high affinity to their specific target molecules. In addition to their use in high-throughput multi-omics applications, the identification of aptamers with high specific binding to the target molecule can be used in therapeutic applications. For example, pegaptanib, a pegylated anti-vascular endothelial growth factor aptamer, is approved for clinical use. Aptamers are selected using a method known as phylogenetic evolution of ligands by exponential enrichment (SELEX). SELEX involves up to 10 times the amount of nucleic acid molecules. 15 A library of individual candidate sequences (e.g., candidate aptamers) is synthesized. However, conventional SELEX does not always produce aptamers with the desired specificity and affinity for the target molecule.

[0013] Generating high-quality aptamers for relevant targets remains a major challenge. One of the biggest limitations is the use of structures with limited chemical diversity, which makes the process of finding high-affinity specific aptamers more difficult. Unlike peptides, single-stranded DNA molecules do not have a wide range of chemical groups. The limited chemical diversity of nucleotide-based libraries limits the finding of successful aptamers. Aptamers generally have four nucleotide building block components. Furthermore, oligonucleotides are negatively charged polymers, which also limits the chemical characteristics of candidate aptamers. In contrast, antibodies have the advantage of achieving high selectivity and sensitivity by combining 20 amino acids in their sequence.

[0014] To overcome this limited range of starting materials, several strategies have been reported in recent years, including LOOPER-SELEX, Click-SELEX, X-SELEX, SELEX post-modification, and the use of modified nuclear bases (i.e., SOMAmers) in the SELEX process. These aptamers, with their expanded chemical diversity, exhibit enhanced binding properties compared to their unmodified versions. Figure 1 shows an exemplary prior art method for aptamer selection. The first step is to generate a library of single-stranded DNA (ssDNA). This may include ssDNA with one modified nucleotide, typically dUTP, modified with an aromatic molecule. Most aptamers identified in this method are not sufficiently selective for target proteins and require further modification with the help of computer and mutation studies. Therefore, after selecting the best aptamer sequences, a second step is performed, involving computer modeling and the creation of a new library with new modifications on the selected sequences. Screening new libraries, or successive new libraries, is time-consuming and inefficient, often involving 5 to 15 rounds of selection. Therefore, existing methods for developing aptamers with high selectivity and affinity are insufficient.

[0015] This specification provides techniques for aptamer generation, which involve incorporating encoded small organic molecules or other modifications into single-stranded or partially single-stranded SELEX libraries. The combination of two molecular evolution-based techniques—aptamer discovery and DNA-encoding libraries—facilitates the generation of aptamers with great chemical diversity, such as small molecule-DNA hybrid aptamers, overcoming the lack of chemical diversity in existing methods and one of the greatest challenges in aptamer discovery. Generating aptamer candidates with increased chemical diversity has the advantage of producing aptamers that are highly selective and specific to a given target molecule (i.e., protein). The use of DNA encoding allows each nucleotide modification type to be identified via its associated encoded DNA sequence, and the identified candidate aptamers can be characterized by both their binding sequence and their modification type and / or modification site. The disclosed aptamer discovery technique, which utilizes a DNA-coding library in combination with the SELEX process, can introduce up to millions of modified fragments in a single selection round, overcoming the limitations of conventional aptamers and offering the following advantages: high-throughput screening, expanded chemical space, easy synthesis, stability, fast generation time, low production cost, and high specificity and selectivity.

[0016] Figure 2 shows an exemplary workflow for the selection of encoded aptamers by the present technique. In the illustrated embodiment, an aptamer candidate library 12 is formed by combining different aptamer candidate subgroups 14 (illustrated as different subgroups 14a, 14b, 14c, 14d, 14e, 14f, 14g and collectively referred to as the aptamer candidate library 12), and the different subgroups 14 have different types of nucleotide modifications relative to each other. It should be understood that the illustration is by way of example and any number of different subgroups 14 may be used to form the aptamer candidate library 12. Each subgroup 14 can be formed separately, and then the separate subgroups 14 are pooled with other subgroups 14 to generate the aptamer library 12. Thus, the specific modification or group of modifications used to generate one subgroup 14 does not modify the other subgroups 14.

[0017] In one example, the different types of modifications of each subgroup 14 can include different types of base modifications and / or modified bases of different identities. For example, an aptamer candidate 20a in the first subgroup 14a can be generated using a specific type of modified uridine (dUTP) 22a. An aptamer candidate 20b in the second subgroup 14b can be generated using a different type of modified uridine (dUTP) 22b such that the modified uridine 22a is chemically distinguishable from the modified uridine 22b. An aptamer candidate 20c in the third subgroup 14c can be generated using a modified adenosine 22c, and so on. The identity of the modified nucleotides 22 (A, C, T, U, G) can be varied or combined to increase the diversity of the library. The individual subgroups 14 of the library 12 can be formed using a single type of modified nucleotide 22 or two or more types of modified nucleotides 22. The types of chemical modifications can include those discussed herein (see Figure 6) and can include modifications introduced via click chemistry.

[0018] Within each subgroup 14, the aptamer candidates 20 can also differ from one another with respect to variable nucleotide regions (see Figure 3), which may, in some cases, be randomly generated. Thus, aptamer candidates 20a may contain mixtures of different sequences, so that the aptamer candidates 20a of subgroup 14a differ from one another at least partially, and so on. Since each aptamer candidate 20 within a particular subgroup 14 has at least different nucleotide sequences within its variable region, the incorporation of modified nucleotides 22 (and any related molecules) is also variable. That is, if the modified nucleotide 22 is uridine, one variable sequence of aptamer candidate 20 may contain 10 or more U (e.g., thymidine) sites, while another variable sequence may contain 1, 2, 3, or 4 U sites. In addition, these sites may be located in different positions for different aptamer candidates 20. Furthermore, in some cases, the incorporation may not be complete. Therefore, if the variable region has 10 available sites, additional diversity can be created by the incomplete incorporation of modified nucleotides 22 via mixing with unmodified nucleotides during synthesis. However, complete incorporation of modified nucleotides 22 at all available sites may result in a more direct characterization of candidates in subsequent steps.

[0019] Therefore, a particular subgroup 14a may be generated using one type of modified nucleotide 22a, or a limited subset of modified nucleotide 22a, but the aptamer candidates 20a generated in subgroup 14a will differ from one another, at least based on the diversity of the variable nucleotide sequences of the aptamer candidates 20a. For example, the first aptamer candidate 20a of subgroup 14a is sequence AAAU as part of a randomly generated variable region. * It may contain GC, and modified nucleotide 22a, dUTP is incorporated at the fourth available position. The second aptamer candidate 20a of subgroup 14a has sequence U as part of a randomly generated variable region. * GCGCU *It may include, and is incorporated at the first and sixth positions where modified nucleotide 22a and dUTP are available. It should be understood that these sequences are for illustrative purposes only. The structures of these different aptamer candidates 20a can be determined based on sequencing, and in certain embodiments, can be determined based on the assumption of incorporation of modified nucleotide 22a at all available sites.

[0020] The disclosed embodiments facilitate an increase in the efficiency of candidate selection in a SELEX-type workflow by enabling screening of different types of modifications together. Aptamer candidates 20 having desired binding characteristics are resolved through sequencing and characterized, and if present, determine which related nucleotide modifications may have contributed to the binding activity. Thus, instead of screening only one subgroup 14 at a time, the combined library 12 can be screened together. As shown in FIG. 2, the aptamer candidates 20 of the library are contacted with a target molecule 30, such as a protein. Unbound aptamer candidates 20 can be separated from any bound aptamer candidates 20, and the bound aptamer candidates 20 shown in FIG. 2 can be retained as aptamer candidates 20f having modified nucleotide 22f. For example, the target molecule 30 can be immobilized on a substrate and the unbound aptamer candidates 20 can be washed away. Thereafter, the bound aptamer candidates 20 can be separated from the target molecule 30 for amplification and subsequent sequencing, and can be eluted, for example, based on a change in buffer conditions.

[0021] None of the nucleotide modifications are retained after the amplification step. However, the nucleotide modification information is preserved via the DNA sequence by including a unique code for each subgroup 14 that is uniquely associated with the modification type for the entire subgroup 14 (e.g., the identity and type of the modified nucleotide 22). Thus, different subgroups 14 (e.g., different subgroups 14a, 14b, 14c, 14d, 14e, 14f, 14g) can be distinguished from one another using their unique codes. In the illustrated embodiment, sequencing of the amplification product of any retained candidate aptamer, shown as example aptamer candidate 20f, as well as identification of both the binding sequence and the unique codes for the associated modification type, can be used as input for one or more additional selection cycles. Additional cycles may include negative selection for binding activity to other target molecules. However, as discussed herein, the ability to distinguish different subgroups 14 and their associated modifications within a single library 12 via encoded nucleic acid sequences allows for the joint screening of a wider variety of aptamer candidates 20 in library 12, thus increasing library selection efficiency and potentially reducing the number of selection cycles.

[0022] Figure 3 shows an exemplary arrangement of retained or selected ssDNA aptamer candidates 20f during the selection cycle of Figure 2. As described in Figure 2, selected aptamer candidates 20 can be isolated from library 12 based on their binding to the target molecule 30. Any binding aptamer candidate 20 at the isolation stage may not be characterized. To identify the features contributing to binding, retained aptamer candidates 20f are amplified and sequenced.

[0023] Aptamer candidate 20f includes a variable region 100 and a coding region 102. The variable region 100 and coding region 102 are adjacent to primer regions 110 and 112. Since aptamer candidate 20 is at least partially single-stranded, the first primer region 110 can represent a primer-binding site which is the reverse complement of the first primer 120, while the second primer region 112 can correspond to the sequence of the second primer 122 which binds to the amplified chain generated from the first primer 120. In some embodiments, aptamer candidate 20 is 50 to 200 nucleotides long. The variable region 100 involved in target binding may be 20 to 120 nucleotides long in some embodiments. The coding region 102 may be 5 to 30 nucleotides long in some embodiments. The primer regions 110 and 112 may be 10 to 30 nucleotides long in some embodiments. To prevent overlap between the randomly generated variable region 100 and the coding region 102, the coding region 102 may be formed from only available nucleotides, for example, T, C, or a subset of A, T, C.

[0024] The amplification product is sequenced to generate a variable region nucleic acid sequence 130 and a unique coding sequence 132. The unique coding sequence 132 can be used to identify subgroups 14f of aptamer candidate 20f and their corresponding modified nucleotide types. That is, the coding region 102 can be conserved for all members of subgroup 14 in one embodiment. In the illustrated embodiment, the modified nucleotide 22f may include a mixture of different modifications, indicated as modified nucleotides 22f', 22f''. The candidate aptamer 22f containing these modifications can produce a sufficient amplification product, and the presence of the modifications does not hinder amplification. As illustrated, the primer regions 110, 112 are the 5' and 3' of the variable region 100 and the coding region 102. The coding region 102 may be the 5' or 3' of the variable region 100. Furthermore, as discussed herein, the coding region 102 may include two or more different coding regions 102, each encoding a different modification aspect.

[0025] As shown in Figure 4, the aptamer candidate library 12 can be expanded exponentially by enabling DNA double-strand formation between the aptamer candidates 20. In one embodiment, individual aptamer candidates 20 partially hybridize with each other to form a partially double-stranded aptamer candidate 150. Each individual aptamer candidate 20 can be combined with a second aptamer 20 having different, similar, or identical modifications, or having no modifications at all, thus expanding the library. 2 To increase diversity, the advantage of this method is that it can be extended to more structures exhibiting two or more variable regions. Furthermore, although the illustrated embodiment shows a partially double-stranded aptamer candidate 150 formed from two different aptamer candidates 20, additional and / or more complex structures may be formed. For example, three, four, five, or more strands of aptamer candidate 20 may be joined to a single structure via complementary region hybridization.

[0026] In one embodiment, a partially double-stranded candidate aptamer 150 is formed by combining different library subgroups (e.g., subgroups 14a, 14b, and 14c) with each other, allowing complementary regions on each aptamer candidate 20 to hybridize to form a double-stranded region 154. Thus, in embodiments, aptamer candidate 20a may include complementary regions, each designed to be complementary to the corresponding regions on aptamer candidate 20b and / or aptamer candidate 20c. The complementary regions are designed to avoid self-complementarity and may facilitate the formation of partially double-stranded regions between subgroups 14 rather than combinations within subgroups. However, it should be understood that combinations within subgroups are also included within this disclosure.

[0027] In one embodiment, library 12 may include all or some potential combinations of two or more subgroups 14. In another embodiment, library 12 may include a combination of one or more subgroups 14 having incorporated modified nucleotides 22 and an unmodified subgroup 14. Furthermore, the described arrangement includes a partially double-stranded candidate aptamer 150 from which each single-stranded branch 156 has a terminal double-stranded region 154 extending therefrom, but other arrangements are also contemplated. In one embodiment, the double-stranded region 154 is an internal region having a single-stranded branch 156 having a divisible variable region extending in both the 5' and 3' directions. Furthermore, one or more of each single-stranded branch 156 may be adjacent to different double-stranded regions 154 that ligate to an additional strand.

[0028] Once formed, the library 12 can be used to select partially double-stranded candidate aptamers 150 that have binding activity or specific affinity to a particular target molecule 30 of interest, as generally considered with respect to Figure 2. For example, individual partially double-stranded candidate aptamers 150ab, formed from aptamer candidates 20a and 20b that partially hybridize with each other, bind to the target molecule 30 and are retained during selection. As illustrated, binding activity may also be mediated by single-stranded branches 156, and the double-stranded region 154ab may not be directly involved in binding. Thus, region 154ab may be conserved among multiple partially double-stranded candidate aptamers 150 formed from different aptamer candidates 20a and 20b. After selection, the partially double-stranded candidate aptamers 150ab are unwound and sequenced.

[0029] To maintain the identity of the double-stranded candidate aptamer 150 formed from two different aptamer candidates 20, chemical ligation can be performed to ligate the component oligonucleotides together, and this can then be identified by sequencing the combined construct after PCR amplification. Chemical ligation may be click chemistry ligation or enzyme-mediated ligation.

[0030] Figure 5 shows an exemplary arrangement of the partially double-stranded candidate aptamer 150ab that was retained or selected during the selection cycle in Figure 4. As described in Figure 4, the selected partially double-stranded candidate aptamer 150ab can be isolated from Library 12 based on its binding to the target molecule 30. To identify the features contributing to binding, the retained partially double-stranded candidate aptamer 150ab is unwound or denatured to isolate the strands of component aptamer candidates 20a and 20b, which are then amplified and sequenced.

[0031] The component aptamer candidates 20a and 20b include variable regions 100 and coding regions 102 adjacent to primer regions 110 and 112, as generally considered with respect to Figure 3. The primer regions 110 and 112 may be conserved between the component aptamer candidates 20a and 20b, so that a single set of primers can be used to amplify the chains of both strands 20a and 20b, or any selected partially double-stranded candidate aptamer 150. In addition, the component aptamer candidates 20 used for partial double-stranding also include complementary regions 160 that hybridize with each other, for example, being complementary to each other. In one embodiment, the complementary regions are located outside the amplification regions so that only the active single-stranded portion of the aptamer candidate 20 is amplified for efficiency. However, other configurations are also conceivable, and one or both of the primer regions 110 and 112 may be part of the complementary region 160. In the partially double-stranded candidate aptamer 150ab, the variable region 100a of the first-strand aptamer 20a may differ from the variable region 100b of the second-strand aptamer 20b. Furthermore, the respective coding regions 102a and 102b differ, reflecting different modification types. Finally, the complementary regions 160a and 160b are reverse complements to facilitate partial double-strand formation.

[0032] The amplification products of both component aptamer candidates 20a and 20b are both sequenced to generate the first strand variable region nucleic acid sequence 162, the first strand intrinsic coding sequence 164, the second strand variable region nucleic acid sequence 162, and the second strand intrinsic coding sequence 168. Using the intrinsic coding sequences 164 and 168, the corresponding modified nucleotide types, shown as subgroups 14a and 14b of aptamer candidates 20a and 20b, and 22a'a'' and 22b, can be identified. That is, the coding region 102 can be conserved for all members of subgroup 14 in one embodiment.

[0033] Figure 6 shows functional groups that can be added to modified nucleotide 22. These are illustrative examples, and it should be understood that other modifications are also within the scope of the disclosed embodiments. For example, modified nucleotide 22 may include phosphate modifications: methylphosphonate (neutral), phosphorothioate (anionic), guanidinopropylphosphorumidate (cationic). Additionally or alternatively, modified nucleotide 22 may include sugar modifications: 2'-F; 2'-amino; 2'-OMe; 2'-azide; conformationally fixed sugars (LNA) X=O, LNA; X=NR, amino-LNA; X=S, thio-LNA. Additionally or alternatively, modified nucleotide 22 may include phosphate substitutions: triazole (neutral); guanidinium (cationic). Additionally or alternatively, modified nucleotide 22 may include purine modifications: 2,6-diaminopurine; 3-deaza-adenine; 7-deaza-guanine; 8-azido-adenine. Additionally or alternatively, modified nucleotide 22 may include pyrimidine modifications: 2-thio-thymidine; 5-carboxamide-uracil 5-methyl-cytosine; 5-ethinyl-uracil (click chemistry site).

[0034] As provided herein, modified nucleotides may contain non-standard functional groups incorporated by click chemistry by introducing azide modifications. These non-standard functional groups have the following characteristics: based on fragmentary small molecules, e.g., fragments known to bind weakly to target proteins, their position is encoded in the DNA sequence with a binary code based on T(0),C(1) to avoid G-quadrichains, as is generally considered with respect to Figure 7.

[0035] In one embodiment, the modified nucleotide may include modification with almost any azide-containing functional group that can be conjugated to an alkyne modification (C5 site) dU via a readily copper-catalyzed azide-alkyne cycloaddition (CuAAC) or "click chemistry". In one embodiment, the modification may be generated based on a starting 5-ethynyluridine or 5-ethynyluracil (Jena Bioscience) containing an alkyne that can subsequently be ligated to an azide-containing molecule via click ligation.

[0036] Primer extension has been shown to act on the modified template to yield good to excellent yields of full-length DNA. Therefore, it is possible to use the selection round as in conventional SELEX (see Figure 8). In addition, an adapted SELEX method involving the combination of ssDNA to double strands can also be performed (see Figure 4), further increasing the chemical space of the SELEX library.

[0037] As provided herein, the aptamer candidate technique can combine two molecular evolution techniques used to find binding molecules with high specificity and selectivity for a target molecule of interest by merging a DNA-coding small organic ligand with ssDNA SELEX. Step 1 involves the synthesis and design of a hybrid small molecule-DNA library and the selection of the target molecule. A standard method using phosphoramidite chemistry is used for the synthesis of the hybrid small molecule-DNA library. Binding of small organic molecules or fragments is investigated using two strategies. The first method utilizes post-modification clickSELEX. In this method, the chemical group is introduced into the DNA library via click chemistry before the selection step and then removed during the amplification step. By doing so, the problem of enzymatic incompatibility associated with larger nucleic acid base modifications is avoided. Thus, the disclosed embodiments allow for easy embodiment of a number of different chemical functional groups adapted to the imposed requirements. Mayer et al. have reported a rapid and cost-effective protocol for the large-scale, high-fidelity generation of nucleic acid base-modified nucleic acids. Following solid-phase synthesis of DNA containing EdU, a click reaction is performed with the DNA still bound to the solid phase, and then deprotection and purification are carried out according to standard procedures. 5-ethynyl 2'-deoxyuridine (EdU) is used as a linker for ssDNA and small molecules.

[0038] Figure 7 shows an exemplary aptamer candidate 20. The location of the "clickable" nucleotide is encoded in the DNA sequence using binary coding bases on the TC base to avoid the G-quadricardium. The library member design is as follows: Regions a and a' correspond to conserved primer sequences 110, 112 for PCR amplification. Region b corresponds to the first portion 170 of coding region 102 and has a unique short DNA sequence of 6 nucleotides (TC coding) to identify the location in region c (variable region 100) which has a modified nucleotide 22 that binds to a clickable nucleotide, e.g., molecular fragment - "small molecule". Region c may be a random sequence of 40-50 nucleotides containing non-clickable and clickable functionalized nucleotides. Region d corresponds to the second portion 172 of coding region 102 which contains the code for the identity of aromatic non-clickable and clickable nucleotides (encoded as ATCG bases).

[0039] Figure 8 shows a template DNA strand library (approximately 10) composed of random sequences adjacent to the fixed sequence regions required for PCR. 14 An exemplary selection cycle is shown, beginning with the chemical synthesis of the sequence. PCR amplification of this template library in the presence of a 3' primer containing a biotin capture molecule at its 5' end yields biotinylated dsDNA that can be captured using a reaction with streptavidin. Note that this step helps ensure that the sequence to be selected is amplified. The dsDNA is then treated with streptavidin immobilized on beads, and the complementary strand is removed from the beads by heating or increasing the pH to produce an ssDNA library in solution. CE-SELEX works in a different manner, without requiring a biotin-functionalized primer.

[0040] Figure 9 shows a selection cycle starting from a DNA template library prepared by automated DNA synthesis. The amide linkage of the modified dUTP derivative is stable under basic conditions for extended periods at ambient temperature. This modified ssDNA library can then undergo a selection step for binding to the protein of interest or any other target. The selected sequences then enter PCR with TTP, dATP, dCTP, and dGTP, avoiding PCR amplification using modified dUTP. Using the same 3'-primer containing 5'-biotin, an enriched library is created, ready to initiate the next selection cycle. An additional next-generation sequencing (NGS) step may be performed at the end of the selection when adding any extra small molecule fragments.

[0041] Small fragments can be selected to contribute to the binding affinity between the aptamer and its target through hydrogen bonding, structural compatibility, aromatic ring stacking, electrostatic and hydrophobic interactions, and van der Waals forces. For this purpose, the first choice is to add ultrasmall molecules as fragments. Small molecules are selected considering their commercial availability and hydrophobicity. The library consists of 1000 different fragments combined with random sequences of nucleotides containing modified bases dUTP, resulting in a hybrid library.

[0042] In step 2, the optimal conditions for PCR amplification of the library are identified. Positive and negative selection rounds are used for the identification of hybrid aptamers. High-throughput sequencing SELEX (HTS-SELEX) is used to enable sequencing of the library across all selection rounds. Identification of DNA-coding molecular fragments is performed in each round, and therefore the enriched sequences are visible at a much earlier stage, making the process more time-efficient. The synthetic alkyne-modified DNA library is functionalized with azide fragments by CuAAC click chemistry. After selection and removal of unbound library members, the binding sequences are isolated and amplified by PCR using alkyne-modified triphosphate instead of thymidine (azide modification can be used alternatively). Thus, the modification in the elongation strand is removed and the alkyne moiety is reintroduced. After PCR, single-stranded DNA is prepared by λ-exonuclease digestion of the 5'-phosphorylated antisense strand. Next, the modifications are reintroduced using CuAAC-click chemistry, and the resulting library is used in the next SELEX cycle.

[0043] Step 3 involves identifying the best-binding molecule. Standard NGS sequencing techniques are used to determine the DNA sequence of the aptamer. The identity of the small molecule fragment is revealed by its unique DNA coding. Since HTS-SELEX is used in each selection round, early recognition of the best-binding aptamer is expected.

[0044] Step 4 involves characterizing the binding properties of the hybrid aptamers. The binding affinity of the selected aptamers is measured using ITC (Isothermal Calorimetry), fluorescence polarization spectroscopy, and SPR (Surface Plasmon Resonance). ITC measurements provide information related to the thermodynamics of the binding event between the ligand and the target molecule. SPR is used to determine the dynamics (Kon / Koff) of the binding event. Finally, fluorescence polarization spectroscopy may be performed by introducing fluorophores into the aptamer structure.

[0045] In some embodiments, the disclosed techniques are used to generate sequence data from amplified aptamer candidates 20. Figure 10 is a schematic diagram of a sequencing device 200 that may be used in conjunction with the disclosed embodiments for obtaining sequencing data from candidate aptamers as provided herein. The sequencing device 200 may be implemented according to any sequencing technique, including those incorporating the synthetic sequencing described in U.S. Patent Applications Nos. 2007 / 0166705, 2006 / 0188901, 2006 / 0240439, 2006 / 0281109, 2005 / 0100900, U.S. Patent No. 7,057,026, International Publication Nos. 05 / 065814, 06 / 064199, and 07 / 010251, the entire disclosure of which is incorporated herein by reference. Alternatively, sequencing by ligation techniques may be used in the sequencing device 200. Such techniques involve incorporating oligonucleotides using DNA ligase and identifying such oligonucleotide incorporation, and are described in U.S. Patents 6,969,488, 6,172,218, and 6,306,597, the entirety of which disclosures are incorporated herein by reference.

[0046] In the described embodiment, the sequencing device 200 includes a separate sample substrate 202, e.g., a flow cell or sequencing cartridge, and an associated computer 204. However, as described above, these may be implemented as a single device. In the described embodiment, a biological sample is loaded onto the substrate 210 and imaged to generate sequence data. For example, a reagent interacting with the biological sample fluoresces at a specific wavelength in response to an excitation beam generated by the imaging module 212, thereby returning radiation for imaging. For example, the fluorescent component may be generated by a fluorescently tagged nucleic acid that hybridizes to a complementary molecule of the component or to a fluorescently tagged nucleotide incorporated into an oligonucleotide using polymerase. As will be understood by those skilled in the art, the wavelength at which the sample dyes are excited and the wavelength at which they fluoresce will depend on the absorption and emission spectra of the particular dye. Such returned radiation may propagate through a directional optical system. This retrobeam may be directed to a detection optical system of the imaging module 212, which may generally be a camera or other optical detector.

[0047] The imaging module detection optics can be obtained based on any preferred technique and may be, for example, a charged coupled device (CCD) sensor that generates pixelated image data based on photons that affect location within the device. However, it will be understood that any of a variety of other detectors may also be used, including but not limited to detector arrays configured for time delay integration (TDI) operation, complementary metal oxide semiconductor (CMOS) detectors, avalanche photodiode (APD) detectors, Geiger-mode photon counters, or any other preferred detectors. TDI mode detection can be coupled with line scanning, as described in U.S. Patent No. 7,329,860, incorporated herein by reference. Other useful detectors are described, for example, in the references previously provided herein in the context of various nucleic acid sequencing methodologies.

[0048] The imaging module 212 may be under processor control, for example, via the processor 214, and may include I / O control 216, an internal bus 218, non-volatile memory 220, RAM 222, and any other memory structure in which the memory can store executable instructions, as well as other suitable hardware components that may be similar to those described with respect to Figure 10. Furthermore, the associated computer 204 may also include a processor 224, I / O control 226, a communication module 234, and a memory architecture including RAM 228 and non-volatile memory 230 in which executable instructions 232 can be stored. The hardware components may be connected by an internal bus 194 which may also be connected to the display 236. In embodiments in which the array determination device 200 is implemented as an all-in-one device, certain redundant hardware elements may be eliminated.

[0049] The processor 214 may be programmed to assign individual sequencing reads to subgroup 14 based on relevant unique coding sequences or sequences, in accordance with the techniques provided herein. Each sequence read may include both a variable region nucleic acid sequence and a unique coding sequence. The sequencing data includes base calls for each base of the sequencing read.

[0050] As used herein, an aptamer may refer to a non-naturally occurring nucleic acid having a specific binding affinity to a target molecule. In certain embodiments, an aptamer candidate 20 is provided, which is a nucleic acid of unknown binding ability that can be determined to have sufficient binding affinity to be selected as an aptamer for a particular application during screening and selection. In certain cases, aptamer candidate 20 and the aptamer may be used interchangeably in this disclosure. Binding of an aptamer to a target molecule may result in catalytically altering the target molecule, reacting with the target molecule in a manner that modifies or alters the target molecule or its functional activity, covalently binding to the target molecule (as in suicide inhibitors), and facilitating a reaction between the target molecule and another molecule. In one embodiment, the target molecule is a non-polynucleotide three-dimensional chemical structure that binds to the aptamer via a mechanism primarily independent of Watson / Crick base pairing or triple helix bonding. In some embodiments, the aptamer is not a nucleic acid with a known physiological function to which it is bound by the target molecule.

[0051] An aptamer comprises nucleic acids identified from a candidate mixture of nucleic acids (e.g., a pool of candidate aptamers). Furthermore, the pool of candidate aptamers may contain one or more aptamers of interest for a particular target molecule. An aptamer can be identified from the candidate aptamers as a ligand for a given target molecule by contacting the candidate mixture with the target, where nucleic acids having increased affinity for the target compared to other nucleic acids in the candidate mixture may be allocated from the remainder of the candidate mixture. The specific binding affinity of an aptamer to its target generally refers to aptamer binding to that target with a much higher degree of affinity than binding to other non-target components in the mixture or sample. Different aptamers may have either the same or different numbers of nucleotides. An aptamer may be DNA or RNA, and may be single-stranded, double-stranded, or contain double-stranded regions.

[0052] The disclosed aptamers and / or aptamer candidates can be used to generate aptamers that can modify the bioactivity of a target through binding to and / or crosslinking. In one embodiment, an aptamer for a specific target associated with or related to a particular disease process is identified. This aptamer can be used as a diagnostic reagent either in vitro or in vivo. In another embodiment, an aptamer for a target associated with a disease state may be administered to an individual and used to treat the disease in vivo. The aptamers identified herein can be used in any diagnostic, imaging, high-throughput screening, or target validation technique or procedure, or in assays that can use aptamers, oligonucleotides, antibodies, and ligands, but are not limited to these.

[0053] In certain embodiments of this disclosure, an aptamer candidate 20 may include one or more conserved primer regions, such as a first conserved primer region and a second conserved primer region. The conserved region is conserved among at least several other aptamer candidates 20 such that the conserved region has the same or similar nucleotide sequences when compared among the aptamer candidates 20. In some embodiments, the conserved region is conserved in both nucleotide sequence and position among different aptamer candidates 20. In some embodiments, the conserved region has the same sequence such that two or more different aptamer candidates 20 have the same sequence in the conserved region. In some embodiments, the conserved region has fewer than two different nucleotides among different aptamer candidates 20 having the conserved region. In some embodiments, the conserved region may be a universal region. For example, for a given library 12, all aptamer candidates in library 12 may have the same first conserved primer region and second conserved primer region. Thus, a primer based on the first and second preserved primer regions can be used to amplify any selected aptamer candidate 20 of the library 12. In one embodiment, the preserved regions may be preserved only within a particular subgroup 14, and all aptamer candidates 20 in that particular subgroup 14 have a preserved code region or a preserved complementary region, while other aptamer candidates 20 in different subgroups 14 have different code regions and / or complementary regions.

[0054] The preserved primer region may include a region having the sequence of the Universal Illumina® Capture Primer or a region that specifically hybridizes with the Universal Illumina® Capture Primer. The Universal Illumina® Capture Primer includes, for example, P5 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 1) or P7 (5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 2)), or a fragment thereof. The region that specifically hybridizes with the Universal Illumina® Capture Primer may include, for example, the reverse complement sequence of the Illumina® Capture Primer P5 ("Anti-P5": 5'-TCGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 3) or P7 ("Anti-P7": 5'-TCGTATGCCGTCTTCTGCTTG-3' (SEQ ID NO: 4)), or a fragment thereof.

[0055] The preserved primer region may additionally or alternatively include a region having the sequence of an Illumina® sequencing primer or a fragment thereof, or a region that specifically hybridizes with an Illumina® sequencing primer or a fragment thereof. An Illumina® sequencing primer includes, for example, SBS3 (5'-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3' (SEQ ID NO: 5)) or SBS8 (5'-CGGTCTCGGCATTCCTGCTGAACCGCTCTTCCGATCT-3' (SEQ ID NO: 6)). A region that specifically hybridizes with an Illumina® sequencing primer or a fragment thereof includes, for example, an Illumina® sequencing primer SBS3 ("Anti-SBS3": 5'-AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT-3' (SEQ ID NO: 7)) or SBS8 ("Anti-SBS8": The sequence may contain the reverse complement sequence of 5'-AGATCGGAAGAGCGGTTCAGCAGGAATGCCGAGACCG-3 (SEQ ID NO: 8) or a fragment thereof. Incorporation of the sequencing primer sequence into the candidate aptamer may be direct or via subsequent amplification, ligation, or other sequencing library preparation steps.

[0056] This document provides for the use of the embodiments, enabling a person skilled in the art to practice the disclosed embodiments, including by using the examples, fabricating and using any device or system, and performing any incorporated methods. The patentable scope is defined by the claims and may include other embodiments conceivable by a person skilled in the art. Such other embodiments are intended to be within the claims if they include structural elements that are no different from the literal wording of the claims, or if they include equivalent structural elements that differ only slightly from the literal wording of the claims.

Claims

1. A aptamer candidate library, A plurality of at least partially single-stranded nucleic acids, wherein each of the plurality of at least partially single-stranded nucleic acids is A first conserved primer region conserved among the plurality of at least partially single-stranded nucleic acids, A second conserved primer region conserved among the plurality of at least partially single-stranded nucleic acids, A variable region located between the first conserved primer region and the second conserved primer region, wherein the variable region is variable among the plurality of at least partially single-stranded nucleic acids, and each variable region contains at least one modified nucleotide. A aptamer candidate library comprising a plurality of at least partially single-stranded nucleic acids, each comprising a coding region, wherein the coding region comprises a nucleotide sequence in which the coding region is specific to the modification type of the at least one modified nucleotide.

2. The library according to claim 1, wherein the plurality of at least partially single-stranded nucleic acids comprise a plurality of modification types, each corresponding to a different coding region having a different nucleotide sequence.

3. The library according to claim 2, wherein the variable regions of the plurality of individual at least partially single-stranded nucleic acids comprise only a single modification type of at least one modified nucleotide.

4. The library according to claim 3, wherein the variable region comprises two or more of the at least one modified nucleotide having the single modification type of the at least one modified nucleotide.

5. The library according to claim 1, wherein the modification type includes modification by molecular fragments.

6. The library according to claim 1, wherein the modification type includes chemically modified uridine.

7. The library according to claim 1, wherein the modification type includes modification via a click chemistry reaction.

8. The library according to claim 7, wherein at least one of the modified nucleotides is modified with an alkyne.

9. The library according to claim 1, wherein the plurality of individual at least partially single-stranded nucleic acids further comprises complementary regions that form double-stranded regions with the plurality of other at least partially single-stranded nucleic acids.

10. The library according to claim 9, wherein the complementary region is not located between the first storage primer region and the second storage primer region.

11. The library according to claim 1, wherein the code area is located between the first storage primer area and the second storage primer area.

12. The library according to claim 1, wherein the coding region includes a first portion for identifying the modification of the modified nucleotide and a second portion for identifying the position of the modified nucleotide within the variable region.

13. The library according to claim 12, wherein the second portion includes a coding sequence encoded by only two different nucleotides.

14. The library according to claim 12, wherein the first portion and the second portion are adjacent to the variable region and are located between the first storage primer region and the second storage primer region.

15. A method for selecting aptamers, Aptamer candidates of different aptamer subgroups, wherein each aptamer candidate of each subgroup of the said different aptamer subgroups is A first preserved primer region preserved among the aptamer candidates, A second preserved primer region preserved among the aforementioned aptamer candidates, A variable region located between the first conserved primer region and the second conserved primer region, wherein the variable region is variable among the aptamer candidates, and the variable region contains at least one modified nucleotide. To provide a library of aptamer candidates comprising: a coding region comprising a nucleotide sequence that is specific to the modification type of at least one modified nucleotide and uniquely identifies the modification type from different modification types of other subgroups of the different subgroups; Selecting aptamer candidates based on their binding to the target molecule, The selected aptamer candidate is amplified using a primer based on the first and second preserved primer regions. A method for selecting an aptamer, comprising: sequencing the amplified aptamer candidate to determine the sequence of the variable region and identifying the modification type.

16. The method according to claim 15, wherein the code regions are stored within each subgroup such that all aptamer candidates in each subgroup have the same code region.

17. The method according to claim 15, comprising separating the amplified aptamer candidate using a capture molecule.

18. A aptamer candidate library, A plurality of partial double-stranded aptamer candidates comprising a first chain and a second chain, wherein the first chain is A double-stranded portion including a first complementary region that hybridizes to a second complementary region of the second strand, A single-strand portion, and the single-strand portion is A first conserved primer region conserved between the first and second strands of the plurality of partial double-stranded aptamer candidates, A second conserved primer region is conserved between the first and second strands of the plurality of partial double-stranded aptamer candidates, A variable region located between the first conserved primer region and the second conserved primer region, wherein the variable region is variable between the first and second strands of the plurality of partial double-stranded aptamer candidates and includes at least one modified nucleotide. A aptamer candidate library comprising a plurality of partially double-stranded aptamer candidates, each comprising a coding region, the coding region comprising a nucleotide sequence in which the coding region comprises

19. The second chain, A second single-strand portion, wherein the second single-strand portion includes the first preservation primer region and the second preservation primer region, A second chain variable region, which is different from the variable region of the first chain, The library according to claim 18, comprising: a second chain code region, which is different from the code region of the second chain.

20. The library according to claim 19, wherein the second chain includes a second modification type different from the modification type of the first chain.