Barcode free library screen
The method uses MS/MS and computational analysis to decode barcode-free libraries, addressing barcode interference and throughput limitations, enabling efficient identification and structure elucidation of binders in large libraries, particularly for nucleic acid-binding targets.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- LEIDEN UNIVERSITY
- Filing Date
- 2025-10-30
- Publication Date
- 2026-05-15
AI Technical Summary
Current affinity selection technologies face limitations in decoding large libraries of small molecules due to the instability and interference of DNA barcodes, particularly when targeting nucleic acid-binding proteins, and lack efficient automated methods for identifying isobaric compounds, limiting throughput and accuracy.
A method utilizing tandem mass spectrometry (MS/MS) to analyze barcode-free libraries, separating binders from non-binders, and employing computational analysis to elucidate the molecular structure of binders directly from their MS/MS spectra, enabling high-throughput decoding of large libraries without the need for DNA or peptide tags.
Enables the identification and structural elucidation of binders in large libraries exceeding 100,000 members, overcoming barcode-related interference and achieving higher throughput and accuracy, especially for DNA-binding targets.
Smart Images

Figure NL2025050557_15052026_PF_FP_ABST
Abstract
Description
Barcode Free Library Screen
[0001] This invention relates to an affinity selection-based hit discovery.BACKGROUND
[0002] The discovery of high affinity ligands is crucial for virtually any drug discovery campaign. The pharmaceutical industry relies on vast collections of individual compounds (typically 0.5 - 4 million) and high throughput screening (HTS) facilities in order to identify novel ligands for drug targets (7). HTS have delivered several starting points for therapeutic compounds, but these libraries cost up to a few billion US dollars and the required infrastructure is large and complex, limiting the availability of such massive platforms mainly to big pharma and few academic settings.
[0003] Affinity selection technology enables the screening of large libraries in a single experiment (2-6). These libraries are typically panned against immobilized target proteins, allowing for the separation of binders from non-binders. A crucial step in this process is the decoding of each hit, which is most commonly achieved using DNA or RNA barcodes attached to each ligand. Display technologies, such as phage display or mRNA display, utilize the natural translation machinery to convert genetic code into peptidic compounds, linked to their encoding oligonucleotides (7-9). DNA-encoded libraries (DELs) feature small molecules linked to unique DNA sequences, enabling the exploration of a drug-like chemical space in the affinity selection setting (70-77).
[0004] While the DNA barcode is an essential component for hit decoding, it also represents the primary limitation in DEL technology in terms of information stability, information density and synthesis complexity. In the library preparation, chemical reactions must be alternated with enzymatic ligation steps, and all transformations need to be both water- and DNA-compatible (78). Many standard chemical reactions involve conditions that degrade DNA and could compromise the chemical barcode (79). Suitable compatible reaction conditions have to be optimized and validated (20). Furthermore, the DNA tag, which is typically > 50-times larger than the small molecule, can significantly affect the selection process, by restricting the binding pose diversity of each library member or by potentially interacting with the target and leading to false positives. This limitation becomes particularly problematic when the target protein has nucleic acid binding sites, making the screening for ligands of crucial drug targets like transcription factors or RNA-binding proteins very challenging (27).
[0005] Consequently, barcode-free selections appear as highly desirable technology. However, current approaches that rely on mass spectrometry (MS) to identify selected compounds from tag-free libraries can only process a few hundred to a few thousand compounds per sample. This limitation arises from the absence of automated de novo decoding software capable of distinguishing and annotating isobaric compounds (5, 22-25). Larger librarysizes have been achieved only for peptidic compounds, which, compared to small molecules have unfavourable drug-like parameters (2, 3, 26-30).
[0006] It is an aim of certain embodiments of the invention to provide a method for elucidating the molecular structure of members of a library of small molecules that bind to a target in an affinity selection. The method may be particularly suited to elucidating the molecular structure of members of a library comprising small molecules that comprises over 104members.BRIEF SUMMARY OF THE DISCLOSURE
[0007] The first aspect of the invention provides a method for identifying which members of a barcode-free library of small molecules bind to a target and elucidating their molecular structure, the method comprising: a) mixing the members of at least one subset of the library with the target; b) separating the members that do not bind to the target (nonbinders) from the members that bind to the target (binders); c) analysing the binders using tandem mass spectrometry (MS / MS) to generate MS / MS spectra of the binders; and d) analysing the MS / MS spectra by computational analysis to elucidate the molecular structure of the binders.
[0008] The method of the present invention eliminates the need to tag library members with barcodes such as DNA or peptide tags. Instead, compounds can be directly decoded based on their MS / MS spectra, enabling higher throughput and reducing both time and cost compared to traditional barcode-dependent approaches, which require barcode synthesis and conjugation. Furthermore, the method avoids common issues associated with barcodes, such as interference with target binding. For example, it enables the identification of DNA-binding molecules, which is typically not feasible with DNA barcode-based methods due to competitive or obstructive interactions.The library
[0009] It may be that step a) comprises mixing a subset of the library with the target. This would be more typical where the library comprises a very large number of members (e.g. greater than 1 ,000,000 or greater than 5,000,000). It maybe that steps a) to d) are carried out on a first subset of the library and then steps a) to d) are carried out on a second subset of the library. This may be continued on further subsets of the library until the entire library has been subjected to steps a) to d). The subset of the library will always comprise a plurality of members, e.g. greater than 100 members or greater than 1000 members.
[0010] It may be that step a) comprises mixing the entire library with the target, i.e. the entire library is subjected to steps a) to d) in a single sitting. The library may comprise greater than 30,000 members, e.g. greater than 50,000 members.
[0011] The library may be a combinatorial library. Thus, the method may further comprise synthesising the library via combinatorial synthesis.
[0012] The library may be enumerated.
[0013] The library is a barcode free library.
[0014] It may be that the at least one subset of the library comprises at least 3000 members. It may be that the at least one subset of the library comprises at least 10,000 members. It may be that the at least one subset of the library comprises at least 50,000 members. It may be that the at least one subset of the library contains at least 100,000 members. It may be that the at least one subset of the library contains at least 250,000 members. It may be that each subset of the library comprises no more than 5,000,000 members. It may be that each subset of the library comprises no more than 1 ,000,000 members. It may be that each subset of the library comprises no more than 750,000 members.
[0015] Typically, each member of the library will be fragmentable by collision-induced dissociation. It may be that at least some members of the library comprise at least one amide bond, C-N bond or C-0 bond. It may be that all members of the library comprise at least one amide bond, C-N bond or C-0 bond. It may be that at least some members of the library comprise at least one amide bond. It may be that all members of the library comprise at least one amide bond. It may be that at least some members of the library comprise an amino acid residue. It may be that all members of the library comprise an amino acid residue.
[0016] It may be that at least one subset of the library comprises members having a moiety selected from:AA is independently at each occurrence selected from an amino acid residue; and Ar1and Ar2are independently selected from an optionally substituted aryl.
[0017] At least some members of the library are small molecules. It may be that all members of the library are small molecules. It may be that the members of the library include isobaric compounds.
[0018] It may be that at least some members of the library are peptides.
[0019] It may be that at least some members of the library are oligonucleotides.
[0020] It may be that at least some members of the library obey Lipinski’s rule of 5. It may be that all members of the library obey Lipinski’s rule of 5. It may be that at least some members of the library do not obey Lipinski’s rule of 5. It may be that all members of the library do not obey Lipinski’s rule of 5.
[0021] The library may comprise a range of structural motifs. This might be the case, for example, if the method is being used for target identification. Alternatively, the library may comprise a range of compounds that share a structure motif, e.g. a particular chemical core. It may be that the library comprises a series of analogues of a compound that was found to bind with the target in a previous iteration of the method of the first aspect. Thus, the method may be used in hit-to-lead optimisation.The target
[0022] Typically, the target will any biological target that has the potential to bind a small molecule. The target may be a protein. The target may be an enzyme. The target may however be a non-biological target, e.g. a polymer, or other material.
[0023] It may be that the target is immobilised. It may be that the target is immobilised on a solid support. The solid support may be magnetic beads. It may be that the target is biotinylated and immobilised on a streptavidin coated solid support.
[0024] The target may be a carbonic anhydrase, e.g. carbonic anhydrase IX (CAIX). The target may be an exonuclease, e.g. Exonuclease 1 (EXO1).Step a)
[0025] Typically, step a) comprises mixing all members of the at least one subset (e.g. all members of the library) with the target in a single vessel.
[0026] Typically, step a) comprises mixing the members of the at least one subset of the library with the target in solution. It may be that the members of the at least one subset of the library are in a first solution, the target is in the vessel and the solution of members of the at least one subset of the library is added to the vessel. The vessel may also contain a solvent, in which the target is dissolved or suspended. The solvent in which the target is dissolved or suspended and the solvent in which the at least one subset of the library is dissolved may be the same solvent. The solvent in which the target is dissolved or suspended and the solvent in which the at least one subset of the library is dissolved may be a different solvent.
[0027] Step a) may comprise mixing the members of the at least one subset of the library with the target in solution, wherein the concentration of the members in the solution is greater than 10 fmol / member. The concentration of the members in solution may be greater than 100 fmol / member. The concentration of the members in solution may be from 100 fmol / member to 1 pmol / member.Step b)
[0028] Typically, the target is immobilised. Where this is the case, step b) may comprise washing the immobilised target to remove any nonbinders.
[0029] Where the target is not immobilised, step b) may comprise separating the non-binders from the target-binder pairs using size exclusion chromatography.
[0030] Step b) may further comprise unbinding the binders from the target after the nonbinders have been separated from the binders, e.g. after the nonbinders have been washed away from the immobilised target. The binders may be unbound through elution.Step c)
[0031] Step c) may comprise analysing the binders using chromatographic tandem mass spectrometry to generate MS / MS spectra. Step c) may comprise analysing the binders using liquid chromatography-tandem mass spectrometry (LC-MS / MS) to generate MS / MS spectra. Step d)
[0032] Typically, step d) is automated.
[0033] The computational analysis may elucidate the structures de novo. More typically, however, the structures (e.g. SMILES) of the at least one subset of the (enumerated) library will be accessible and the computational analysis elucidates the structures by fitting the generated spectra to the best matches in the at least one subset of the enumerated library. The computational analysis may require that the structure elucidated in step d) matches the structure of a member in the enumerated library.
[0034] It may be that the computational analysis determines how closely the members of the at least one subset of the library fit the generated MS / MS spectra. In this embodiment, the library will typically be an enumerated library. The computational analysis may score and rank the members of the at least one subset of the library based on how closely they fit the generated MS / MS spectra. It may be that the computational analysis filters the members of the at least one subset of the library and disregards members that do not fit the MS / MS spectra generated in step c). By “disregard” it is meant that the computational analysis will not determine the molecular structure of the binders to be that of a member that does not fit the MS / MS spectra generated in step c).
[0035] It may be that the computational analysis requires that the binder must have an MS1 m / z that matches that of a library member.
[0036] It may be that computational analysis uses statistics, combinatorics, rule-based predictions, or Quantum Chemistry computations. The computational analysis may be based on a machine learning model. The machine learning model may predict fragment peaks based on MS / MS spectra of at least a proportion of members of the at least one subset of the library. The computational analysis may compute fragmentation trees to interpret the MS / MS spectra.
[0037] The method may further comprise predicting MS / MS spectra of each member of the at least one subset of the library and the computational analysis compares MS / MS spectra generated in step c) to the predicted spectra. The MS / MS spectra may be predicted via fragmentation rules, machine learning or Quantum Chemistry simulations. The MS / MS spectra of each member of the at least one subset of the library may be predicted based on MS / MS spectra of at least a proportion (e.g. a proportion) of members of the at least one subset of the library. The method may therefore further comprise obtaining MS / MS spectra of at least a proportion of at least a proportion (e.g. a proportion) of members of the at least one subset of the library. The computational analysis may require that the binder must show at least one predicted fragment peak in their MS2.
[0038] It may be that the computational analysis predicts the presence or absence of substructures of the binders via rule-based considerations, a machine learning model, statistics or combinatorial optimization based on the MS / MS spectra generated in step c). When the library is a combinatorial library, the substructures may be particular building blocks from the combinatorial library. It may be that the substructures are molecular fingerprints.The method
[0039] The method may comprise steps a1) and b1): a1) mixing the same members of the at least one subset of the library with the same components of the reaction of step a) but without the target present; b1) separating the members that do not bind non-specifically to the components (nonbinders) from the members that bind to the components (non-specific binders); wherein step c) comprises subtracting the MS / MS spectra of the non-specific binders from the data set of MS / MS spectra of the binders.
[0040] Where the target is immobilised, e.g. on magnetic beads, the components used in step a1) will typically comprise the immobilisation means, e.g. the magnetic beads.
[0041] The method may comprise steps c2) and d2): c2) analysing the nonbinders obtained in step b) using tandem mass spectrometry (MS / MS) to generate MS / MS spectra; and d2) analysing the MS / MS spectra using computational analysis to elucidate the molecular structure of the nonbinders, wherein step d2) is automated.BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Embodiments of the invention are further described hereinafter with reference to the accompanying drawings, in which:Figure 1 outlines the hit discovery with SELs. A) SEL workflow: large libraries with diverse scaffolds are prepared via solid phase split and pool synthesis with a wide compatibility range of chemical transformation. The libraries are then released from beads and screened insolution against immobilized targets. Binders are enriched from the rest of the library and analyzed via nano LC-MS / MS. Hit compounds are identified and decoded by SIRIUS-COMET (Combinatorial Mass Encoding decoding Tool) software. B) Summary of SEL features making it an excellent platform for rapid and straightforward early drug discovery.Figure 2 shows the amino acid building blocks used to synthesize SELs 1 , 2 and 3. Fmoc protected variants of the amino acids were used in the library synthesis. Building blocks marked with a * are not compatible with SEL 3 and were excluded for that design. Protecting groups indicated in red are removed during cleavage.Figure 3 shows carboxylic acid building blocks used to synthesize SEL 1. Protecting groups indicated in red are removed during cleavage.Figure 4 shows the primary amine building blocks used to synthesize SEL 2.Figure 5 shows the aldehyde building blocks used to synthesize SEL 2.Figure 6 shows the aryl bromide building blocks used to synthesize SEL 3. Figure 7 shows the boronic acid building blocks used to synthesize SEL 3. Figure 8 shows the two consecutive filtering steps of COMET for removing MS / MS spectra that do not match to a library compound.Figure 9 outlines the library design and synthesis. A) SEL 1 is prepared from two amino acid building blocks decorated with a carboxylic acid decorator. Standard solid phase amide bond formation protocols result in products with excellent crude purity. B) SEL 2, resulting in a benzimidazole scaffold with three decorators is synthesized in a 7-step procedure. Optimization of all steps and reaction scope for the nucleophilic aromatic substitution and the heterocyclization are reported in detail in Fig. 10-12 and summarized in the histograms. C) SEL 3 is prepared from an amino acid, and a Suzuki-Miyaura reaction between an aryl bromide and a boronic acid. Cross-coupling optimization and reaction scope for the aryl bromides and boronic acids are reported in detail in Table 3 and Fig. 13 and 14 and summarized in the histograms. D) Based on the building block scope for all scaffolds (crude purity > 55%) and drug-like properties of the building blocks, three high diversity libraries were prepared. Selected examples from each library emphasize the chemical diversity of the generated structures. The violin plots show the compliance of our libraries with several parameters important for drug development (the violin plots on the left are for SEL 1 , the middle violin plots are for SEL 2, and the violin plots and the right are for SEL 3).Figure 10 shows the optimization of trisubstituted benzimidazole synthesis on solid support. A) Optimization of the nucleophilic aromatic substitution reaction between 4-fluoro-3- nitrobenzoic acid and the primary amines aniline and pentan-3-amine B) Optimization of the reduction of nitro group. C) Optimization of the cyclization between compound 5 and isovaleraldehyde.Figure 11 shows the reaction scope for nucleophilic aromatic substitutions in the benzimidazole scaffold (primary amines), a) Reaction conditions: Compound 8 (5 pmol, 1.0 eq.), amine (50 pmol, 10 eq.), DIPEA (50 pmol, 10 eq.) in DMF (100 pL), 80°C, overnight, b) Conversion was determined by LC-MS analysis.Figure 12 shows the reaction scope for heterocyclization to the benzimidazole scaffold (aldehydes)Figure 13 shows the reaction scope of different aryl bromides in the Suzuki-Miyaura reaction, a) Reaction conditions: Compound 1 (5 pmol, 1.0 eq.), phenylboronic acid (10 pmol, 2 eq.), PdCI2(10 mol%), XPhos (20 mol%), K2CO3(10 pmol, 2 eq.), 9:1 DMF / H2O (100 pL), 80 °C o.n. b) Conversion was determined by LC-MS analysis.Figure 14 shows the reaction scope of different boronic acids / esters in the Suzuki- Miyaura reaction, a) Reaction conditions: Compound 1 (5 pmol, 1.0 eq.), phenylboronic acid (10 pmol, 2 eq.), PdCI2(10 mol%), XPhos (20 mol%), K2CO3(10 pmol, 2 eq.), 9:1 DMF / H2O (100 pL), 80 °C o.n. b) Conversion was determined by LC-MS analysisctBu removed with TFA.Figure 15 compares libraries with selected building blocks having optimized druglike properties to libraries with randomly chosen building blocks. Left: before selection (1000*1000*1000=1 billion membered library). Right = After selection (54*54*44 =128.304 membered library).Figure 16: Decoding with SIRIUS:COMET. A) The scatter plot shows the mass degeneracy of each library design. For specific masses there can be up to 500 possible different compounds. Tandem mass spectrometry is essential to enable the unequivocal annotation of unique structures. B) For each library scaffold a small sublibrary with < 500 unique compounds was synthesized and measured by LC-MS / MS to mimic the outcome of an affinity selection outcome. C) Given that in these libraries all compounds are known we performed semimanual analysis to obtain the actual number of detectable compounds. Given that each nanoLC-MS / MS run produced ~ 80,000 spectra, such manual analysis would be unfeasible for samples with unknown content. Automated analysis with off the shelf SIRIUS 6 resulted in good recalls of library compounds, but with some false positives (annotated compounds that are not actually present in the library). D) Based on the manually annotated test libraries, fragmentation frequencies of each library scaffold were calculated and represented as histograms. E,F) The SIRIUS-COMET workflow consist of generating a fragments list based on the calculated fragmentation frequencies, filtering MS / MS scans based on the minimum number of matching peaks and structure annotation by SIRIUS 6. The SIRIUS- COMET workflow decreased falsely annotated structures while retaining high recall rates of correct compounds.Figure 17 depicts the chemical structures of hit 20 and 21 , with the extracted ion chromatogram (EIC) and LC-MS / MS chromatogram from the affinity selection of SEL 1 against CAIX.Figure 18 depicts the chemical structures of hit 24 and 25, with the extracted ion chromatogram (EIC) and LC-MS / MS chromatogram from the affinity selection of SEL 2 against CAIX.Figure 19 depicts the chemical structures of hit 35, with the extracted ion chromatogram (EIC) and LC-MS / MS chromatogram from the affinity selection of SEL 3 against CAIX.Figure 20 outlines the Identification and validation of ligands against CAIX and EXO1. A) SEL 1 ,2 and 3 were panned against immobilized CAIX. After analysis of the selected samples via nanoLCMS / MS and SIRIUS-COMET, we plotted building block frequencies for each library scaffold and position. Aromatic sulfonamides were strongly enriched from all libraries. B) Examples of hit compounds: the chromatograms show extracted ion counts of their corresponding masses in CAIX sample and bead only control, demonstrating their selective enrichment. Biotinylated version of these examples were resynthesized and BLI association / dissociation curves confirm there binding to CAIX. C) In the affinity selection against EXO1 only one clear hit was identified. Based on an isobaric building block pair, two possible structures could match the annotation. BLI showed that only compound 6 binds to EXO1.Figure 21 depicts the chemical structures of isomers 43 and 44, with the extracted ion chromatogram (EIC) and LC-MS / MS chromatogram from the affinity selection of SEL 1 against EXO1.Figure 22 shows the building block analysis after affinity selection with SEL 1 against EXO1.DETAILED DESCRIPTIONDefinitions
[0043] The term “member”, as in a member of a library of small molecules, refers to molecules with a defined chemical structure. A library comprising n members therefore comprises molecules having n different structures. The members may all be small molecules. It may be that a portion of the library of small molecules are not small molecules, providing that at least some of the members of the library are small molecules.
[0044] A “small molecule” may be considered to be a molecule (or salt thereof) having a molecular mass below 1500 gmol-1. It may be that a ‘small molecule’ has a molecular mass below 1250 gmol-1. It may be that a ‘small molecule’ has a molecular mass below 1000 gmol-1. A ‘small molecule’ may be a chemical (e.g. organic) molecule (or salt thereof) comprising less than 100 atoms. A ‘small molecule’ may be a chemical (e.g. organic) molecule (or salt thereof) comprising less than 50 heavy atoms (heavy atoms being atoms other than hydrogen). Smallmolecules may be organic, or they may be organometallic. It may be that at least some of the small molecules are organic small molecules. It may be that all small molecules are organic small molecules.
[0045] A “target” is any molecule or material that has the potential to bind a small molecule. Typically, the target will be a biological target, i.e. any part of the living organisms to which a small molecule can bind in order to bring about a physiological change. Common biological targets are nuclear receptors, ion channels, G-protein coupled receptors and enzymes.
[0046] An “enumerated” library is one of which the chemical structure of every member is known and catalogued.
[0047] A “molecular fingerprint” is a unique pattern or representation of a molecule's chemical structure and properties. A molecular fingerprint may provide information on the entire chemical structure of a molecule. A molecular fingerprint may also provide information on part of the chemical structure of a molecule, e.g. a moiety an / or fragment of a chemical molecule.
[0048] An “immobilised” target is a target that is physically confined or localized in a certain defined region of space with retention of its biological activity, and which can be used repeatedly and continuously. For example, a target may be immobilised on a solid support.
[0049] The term “solid support” or “solid phase” means, unless otherwise stated, a material that is substantially insoluble in a selected solvent system, or which can be readily separated (e.g., by e.g., by filtration, centrifugation, or magnetic capture) from a selected solvent system in which it is soluble. In the present disclosure, solid supports that are substantially insoluble in a selected solvent system may be preferred. Such solid supports are not limited to a specific type of support, and a large number of such solid supports are available and are known to one of ordinary skill in the art. Exemplary solid supports include, but are not limited to, solid and semisolid matrixes, such as aerogels and hydrogels, resins, particles, beads (including magnetic beads, such as coated magnetic beads), biochips (including thin film coated biochips), microfluidic chip, a silicon chip, multi-well plates (also referred to as microtiter plates or microplates), membranes, conducting and nonconducting metals, glass (including microscope slides) and magnetic supports. Where the solid support comprises beads, monodisperse (magnetic or non-magnetic) beads may be preferred, as monodisperse beads may provide more consistent performance in assays. A solid support may comprise magnetic beads.
[0050] A solid support may be used to immobilise a biological target. Solid supports useful for this purpose include streptavidin coated magnetic beads, which can be used to immobilise biotinylated targets.
[0051] A solid support may be used to synthesise molecules on in combinatorial synthesis. Solid supports useful for this purpose include resin beads functionalized with an -NH2 group, e.g. TentaGel®.
[0052] X® denotes a solid support.
[0053] An amino acid residue may be a residue of a proteinogenic or a non-proteinogenic amino acid. Proteinogenic amino acids include alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, selenocysteine or pyrrolysine. The amino acid may be the a-amino acid, e.g. a-leucine, or the p-amino acid, e.g. p-alanine. Typically, the amino acid will be the L-amino acid. The amino acid may be a non-proteinogenic amino acid. An amino acid residue may be the residue of an amino acid selected from Figure 2.
[0054] For the absence of doubt, in the following moieties:the amino acid residue will be connected to the adjacent carbonyl group via the nitrogen atom of its amine group, or a nitrogen atom of one of its amine groups (thereby forming a peptide bond) and will be connected to the adjacent amine group via its carbonyl carbon (thereby forming a peptide bond).
[0055] Aryl groups may be any aromatic carbocyclic ring system (i.e. a ring system containing 2(2n + 1 )TT electrons). Aryl groups may have from 6 to 12 carbon atoms in the ring system. Aryl groups may be phenyl groups. Aryl groups may be naphthyl groups or biphenyl groups. The aryl group may be unsubstituted or substituted.
[0056] In tandem mass spectrometry, also known as MS / MS or MS2, the molecules of a given sample are ionized and the first spectrometer (designated MS1) separates these ions by their mass-to-charge ratio (often given as m / z or m / Q). Ions of a particular m / z-ratio coming from MS1 are selected and then fragmented into smaller fragment ions, e.g. by collision-induced dissociation, ion-molecule reaction, or photodissociation. These fragments are then introduced into the second mass spectrometer (MS2), which in turn separates the fragments by their m / z- ratio and detects them. The fragmentation step makes it possible to identify isobaric compounds (i.e., compounds having the same mass but different chemical structures), which is not possible in regular mass spectrometry. Additional fragmentation steps (multiple-stage MS or MSn) can obtain further molecular information. Tandem mass spectrometry can be coupled with liquid chromatography, i.e. liquid chromatography-tandem mass spectrometry (LC-MS / MS).
[0057] For the absence of doubt, step c), “analysing the binders using tandem mass spectrometry (MS / MS) to generate MS / MS spectra of the binders” comprises subjecting the binders to tandem mass spectrometry (MS / MS) to generate MS / MS spectra of the binders themselves. In order to generate MS / MS spectra, the binder should therefore comprise at leasttwo bonds that are susceptible to fragmentation during tandem mass spectrometry (e.g. by collision-induced dissociation).
[0058] Typically, therefore, each member of the at least one subset comprises at least two bonds that are susceptible to fragmentation during tandem mass spectrometry. It may be that each member of the library comprises at least two bonds that are susceptible to fragmentation during tandem mass spectrometry.
[0059] Computational analysis is the use of computers to analyse data. The computational analysis required in step d) interprets the MS / MS spectra generated in step c) in order to elucidate the molecular structure of the binders. It may be that the computational analysis interprets the MS / MS spectra by converting this into molecular structure information and elucidates the molecular structure the binders based on this information.
[0060] In the method of the present invention, step d is typically automated, meaning that the MS / MS spectra generated in step c) are analysed automatically without human input in order to elucidate the molecular structure of the binders.
[0061] The computational analysis may be implemented by software. Typically, the software is metabolomics software. It may be that the software is not proteomics software. The software may include a SIRIUS software, e.g. SIRIUS 4, which is a Java software used to analyse molecules from tandem mass spectrometry data.
[0062] Fragmentation trees annotate the MS / MS spectra of a molecule via an assumed / predicted fragmentation process.
[0063] A “combinatorial library” is a library comprising members that have been synthesised by combinatorial chemistry. Typically, members of a combinatorial library will share a common molecular scaffold(s), or core structure(s). It may be that the library is divided into subsets. The members in each subset will typically share a common moiety and / or substituent (in addition to their common molecular scaffold).
[0064] Combinatorial chemistry involves the synthetic assembly of structural building blocks in various possible combinations to produce large libraries of small molecules.
[0065] A barcode library, or encoded chemical library, is a library comprising members that are conjugated to a molecular tag that allows the molecular structure of the member to be determined. One such library is a DNA-encoded chemical library (DECL, or DEL) in which members are conjugated to a DNA fragment that serve as identification barcodes. Another such library is a peptide encoded library, in which members are tagged with a unique peptide sequence that serves as an identification barcode. The barcode (in a barcode free library) may therefore be a nucleic acid (e.g. DNA) barcode or a peptide barcode. In other words, it may be that the none of the members of the library comprise a barcode, e.g. a nucleic acid (e.g. DNA) barcode, or a peptide barcode.
[0066] SMILES is the “Simplified Molecular Input Line Entry System,” which is used to translate a molecule’s three-dimensional structure into a string of symbols that is more easily understood by computer software.
[0067] The abbreviations used herein have the following meaning:BB Building blockBLI Biolayer interferometryCAIX Carbonic anhydrase IXDCM DichloromethaneDEL DNA encoded libraryDIPEA Di-isopropylethylamineDMF DimethylformamideDMSO DimethylsulfoxideEq EquivalentEXO1 Exonuclease 1FA Formic acidFBS Fetal bovine serumFmoc FluorenylmethoxycarbonylHexafluorophosphate Azabenzotriazole TetramethylHATU UraniumHBA Hydrogen bond acceptorHBD Hydrogen bond donorHEPES (4-(2-hydroxyethyl)-1 -piperazineethanesulfonic acid)Me MethylMeCN AcetonitrileMeOH MethanolMW Molecular weightMQ MilliQNMR Nuclear magnetic resonancePBS Phosphate-buffered salinePEL Peptide encoded libraryPpm Parts per million r.t. Room temperatureSA StreptavidinSEL Self-encoded library t-Bu Tert-butylTFA Trifluoroacetic acidTIPS TriisopropylsilaneTLC Thin layer chromatographyTPSA Topological polar surface areaTween-20 Polysorbate 20
[0068] Throughout the description and claims of this specification, the words “comprise” and “contain” and variations of them mean “including but not limited to”, and they are not intended to (and do not) exclude other moieties, additives, components, integers or steps. Throughout the description and claims of this specification, the singular encompasses the plural unless the context otherwise requires. In particular, where the indefinite article is used, the specification is to be understood as contemplating plurality as well as singularity, unless the context requires otherwise.
[0069] Features, integers, characteristics, compounds, chemical moieties or groups described in conjunction with a particular aspect, embodiment or example of the invention are to be understood to be applicable to any other aspect, embodiment or example described herein unless incompatible therewith. All of the features disclosed in this specification (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. The invention is not restricted to the details of any foregoing embodiments. The invention extends to any novel one, or any novel combination, of the features disclosed in this specification (including any accompanying claims, abstract and drawings), or to any novel one, or any novel combination, of the steps of any method or process so disclosed.
[0070] The reader's attention is directed to all papers and documents which are filed concurrently with or previous to this specification in connection with this application and which are open to public inspection with this specification, and the contents of all such papers and documents are incorporated herein by reference.
[0071] The invention may be as described in one of the following numbered clauses:1 . A method for identifying which members of a library of small molecules bind to a target and elucidating their molecular structure, the method comprising: a) mixing the members of at least one subset of the library with the target; b) separating the members that do not bind to the target (nonbinders) from the members that bind to the target (binders); c) analysing the binders using tandem mass spectrometry (MS / MS) to generate MS / MS spectra; and d) analysing the MS / MS spectra by computational analysis to elucidate the molecular structure of the binders.2. The method of clause 1 , wherein step c) comprises analysing the binders using liquid chromatography-tandem mass spectrometry (LC-MS / MS) to generate MS / MS spectra.3. The method of clause 1 , wherein step b) further comprises unbinding the binders from the target after the nonbinders have been separated from the binders.4. The method of any preceding clause, wherein the library is a combinatorial library.5. The method of clause 4, wherein the method further comprises synthesising the library via combinatorial synthesis.6. The method of any preceding clause, wherein the library is enumerated.7. The method of any preceding clause, wherein the library is a barcode free library.8. The method of any preceding clause, wherein the at least one subset comprises at least10,000 members.9. The method of any preceding clause, wherein the at least one subset contains at least 100,000 members.10. The method of any preceding clause, wherein the at least one subset comprises no more than 5,000,000 members.11. The method of any preceding clause, wherein step a) comprises mixing the members of the entire library with the target.12. The method of any preceding clause, wherein the members are fragmentable by collision-induced dissociation.13. The method of any preceding clause, wherein the members comprise an amino acid residue.14. The method of any preceding clause, wherein the computational analysis determines how closely the members of the at least one subset of the library fit the generated MS / MS spectra.15. The method of any preceding clause, wherein the computational analysis scores and ranks the members of the at least one subset of the library based on how closely they fit the generated MS / MS spectra.16. The method of any preceding clause, wherein the computational analysis filters the members of the at least one subset of the library and disregards members that do not fit the MS / MS spectra generated in step c).17. The method of any preceding clause, wherein the computational analysis requires that the binder must have an MS1 m / z that matches that of a library member.18. The method of any preceding clause, wherein the computational analysis uses statistics, combinatorics, rule-based predictions, or Quantum Chemistry computations to elucidate the molecular structure of the binders, or wherein the computational analysis is based on a machine learning model.19. The method of any preceding clause, wherein the method further comprises predicting MS / MS spectra of each member of the at least one subset of the library and the computational analysis compares the MS / MS spectra generated in step c) to the predicted spectra, optionally wherein the MS / MS spectra is predicted via fragmentation rules or Quantum Chemistry simulations or is predicted based on a machine learning model.20. The method of clause 19, wherein the MS / MS spectra of each member of the at least one subset of the library is predicted based on MS / MS spectra of at least a proportion of members of the at least one subset of the library.21. The method of clause 19 or clause 20, wherein the computational analysis requires that the binder must show at least one predicted fragment peak in its MS2.22. The method of any preceding clause, wherein computational analysis predicts the presence or absence of substructures of the binders using rule-based considerations, a machine learning model, statistics or combinatorial optimization based on the MS / MS spectra generated in step c).23. The method of any clause 22, wherein the library is a combinatorial library, and the substructures are particular building blocks from the combinatorial library.24. The method of clause 22 or clause 23, wherein substructures are molecular fingerprints.25. The method of any preceding clause wherein the target is an enzyme.26. The method of any preceding clause, wherein the target is immobilised.27. The method of clause 26, wherein the target is immobilised on a solid support.28. The method of clause 27, wherein the solid support is magnetic beads.29. The method of clause 27 or clause 28, wherein the target is biotinylated and immobilised on a streptavidin coated solid support.30. The method of any one of clauses 26 to 29, wherein step b) comprises washing the immobilised target to remove any nonbinders.31. The method of any preceding clause, wherein step a) comprises mixing the members of the at least one subset of the library with the target in solution, wherein the concentration of the members in the solution is greater than 10 fmol / member.32. The method of clause 31 , wherein the concentration of the members in solution is from 100 fmol / member to 1 pmol / member.33. The method of any preceding clause, wherein the computational analysis is implemented by software.EXAMPLES
[0072] Here, we report an affinity selection platform that can screen barcode free small molecule libraries with 104-106members, all at once (Fig. 1). The approach features the combinatorial synthesis of drug-like compounds on solid phase beads, allowing for a wide rangeof chemical transformations and circumventing the complexity and limitation of DEL preparation. Compounds are annotated using their tandem MS fragmentation spectra, obviating the need for barcoding tags and at the same time allowing for precise identification of many isobaric, but structurally diverse compounds. We show the feasibility of our decoding strategy on a diverse set of chemical scaffolds, prepared by a variety of chemical transformations (including crosscouplings, heterocyclizations, amide formation, nucleophilic aromatic substitution, and more). We then performed de novo discovery selections, panning libraries with 30-500k members against the two oncology targets CAIX and Exonuclease-1 (EXO1), resulting in identification of nanomolar binders for both targets. Notably, EXO1 is a DNA-processing enzyme, and therefore very challenging for DEL selections. Overall, our SEL platform presents ideal features in terms of information density, information stability and unbiased screening capabilities.
[0073] In the following sections we describe in detail the individual parts of the platform, including library design and synthesis, decoding software and methodology and the de novo ligand selection campaigns.Material and methodsReagents and Supplies
[0074] Chemicals: Reagents and solvents were purchased from Sigma-Aldrich (Merck), Fisher Scientific, BLDpharm, Fluorochem or VWR and were used without further purification unless stated otherwise. Carboxylic acids, aldehydes and Fmoc-protected amino acids building blocks were partially purchased from Chemspace. Tentagel® S NH2 90 pm (S30902) and TentaGel® M NH230 pm (M30352) were purchased from Rapp-Polymere.
[0075] The amino acid building blocks used to synthesize SEL 1 , 2 and 3 are shown in Fig. 2. The carboxylic acid building blocks used to synthesize SEL 1 are shown in Fig. 3. The primary amine building blocks used to synthesize SEL 2 are shown in Fig. 4. The aldehyde building blocks used to synthesize SEL 2 are shown in Fig. 5. The aryl bromide building blocks used to synthesize SEL 3 are shown in Fig. 6. The boronic acid building blocks used to synthesize SEL 3 are shown in Fig. 7.
[0076] Materials for affinity selection: Dynabeads MyOne Streptavidin T 1 was purchased from ThermoFisher Scientific. A KingFisher™ Duo Prime Purification System was used to perform our affinity selection experiments. The protocols were developed with Bindlt 4.0 Software.
[0077] Proteins: Biotinylated Human Carbonic Anhydrase IX (38-414), His.Avitag (CA9- H82E3) and Human Carbonic Anhydrase IX (38-414), His.Avitag (CA9-H5226) were purchased from ACROBiosystems. EXO1 was obtained from Sylvie Noordermeer (Department of Human Genetics, Leiden University Medical Center.Instrumentation
[0078] LC-MS analysis: Compound purity was determined by LC-MS, using the LCMS-2020 system (Shimadzu) with a Gemini 3 pm C18 110 A column (50 x 3 mm) using the following parameters: flow rate = 0.55 mL / min, scan range = 160- 800 m / z, column temperature (°C) = 40. The following gradient of 10-90% MeCN / H2O (0.1% formic acid) over 15 min and measuring UV absorbance at 254 nm was used, unless stated otherwise. Compounds were dissolved in H2O:MeCN:f-BuOH (1 :1 :1) before injection.
[0079] Purification: The automatic flash chromatography was performed on a Biotage Selekt System with pre-packed flash cartridges (Biotage® Star Bio C18 - Duo 300 A 20 pm). MQ Millipore water 0.1% TFA (buffer A) and Acetonitrile 0.1% TFA (buffer B) were used as mobile phase.
[0080] LC-MS / MS analysis: Analysis was performed on an Vanquish™ Neo UHPLC system (Thermo Scientific) connected to an Orbitrap Exploris 240 mass spectrometer (Thermo Fisher Scientific). Samples were run on a Double nanoViper PepMap Neo column (2 pm particle size, 15 cm x 75 pm, Thermo Fisher Scientific, DNV75150PN) following a PepMap Neo Trap Cartridge (5 pm C18 300 pm X 5 mm). The standard nano-LC method was run without temperature control and a flow rate of 300 nL / min with the following gradient: 0% solvent B ramping linearly to 60% B in A over 78 min, with solvent A = water (0.1% FA), and solvent B = 80% acetonitrile, 20% water (0.1% FA). Positive spray ionization was set at 1900 V. The following parameters were used for MS1 collection: resolution = 120.000, scan range = 275-650 m / z for SEL 1 , 300-800 m / z for SEL 2 and 275-700 for SEL 3, maximum injection times = 300 ms, RF Lens = 80%, microscans = 1 , AGC target = standard, Ion Transfer Tube Temp (°C) = 280, Charge State = +1 . For the data dependent MS / MS event the following settings were used: resolution = 15000, number of Dependent Scans= 20, isolation window (m / z) = 1.5, intensity threshold = 1 E5, MIPS mode = small molecule, absolute collision energy = 15, 25 eV, AGC Target (%) = 50, microscans = 1 , RF Lens(%) = 70, dynamic exclusion = on (exclude after n times = 3, Exclusion duration (s) = 15, excluding isotopes, 10 ppm mass tolerance).General Procedures
[0081] Manual solid-phase synthesis (SPS):
[0082] Rink linker: TentaGel S NH2resin (90 pm, 0.26 mmol / g, 1.0 eq.) was functionalized with a Rink linker using the following protocol before attaching other building blocks.
[0083] TentaGel S NH2resin (90 pm, 0.26 mmol / g, 1.0 eq.) was loaded onto a fritted syringe and swelled for 5 min in DMF. A solution of Fmoc-protected amino-acid (3.0 eq), HATU (0.4 M, 2.98 eq) and DIPEA (9.0 eq) in DMF was added to the resin. After 20 min the resin was washed with DMF (5x) and Fmoc deprotection was performed by washing with piperidine (20% in DMF, 1x) before incubating with piperidine (20% in DMF) for 10 min. The resin was washed with DMF (5x) before adding the next Fmoc-protected amino-acid. The protocol was repeated until completion of the synthesis.
[0084] Cleavage: The compounds were cleaved of the resin by incubating at room temperature for 1 hour with a solution of TFA:H2O:TIS (95:2.5:2.5) and washed with TFA:H2O:TIS (95:2.5:2.5). The volume was reduced by evaporating the TFA solution with a N2stream. The reaction mixture was dissolved in H2O:MeCN:f-BuOH (1 :1 :1) before injection into a LC-MS.
[0085] Protein biotinylation: EZ-Link™ Sulfo-NHS-LC-Biotin was used in accordance with the user guide from themofisher.
[0086] Affinity selection: A KingFisher™ Duo Prime Purification System was used to perform our affinity selection experiments. The protocols were developed with Bindlt 4.0 Software. Each experiment was performed in duplicates with protein (150 pmol) immobilized on DynabeadsMyOne Streptavidin T1 (1 mg) and library (100 fmol / member) with the King Fisher protocol as described in Table 1 (below unless stated otherwise).Table 1 : The KingFisher program used for affinity selection.Table 2: Overview of the affinity selection buffers.
[0087] Biolayer interferometry (BLI): Purified biotinylated compounds were dissolved to 1 pM in 1x PBS, 0.02% Tween-20, 1 mg / ml BSA (0.1% (w / v)) (kinetic buffer) used for immobilization onto streptavidin Octet SA Biosensors (SATORIUS). Biolayer interferometry (BLI) assays were performed in 96 well plates (GreinerBio-One, polypropylene, flat-bottom, chimney well) using an Octet R4 system (SATORIUS). Wells were filled with 200 pL with kinetic buffer, compound solution or CAIX solution.
[0088] Biotinylated compound was immobilized onto the streptavidin biosensor for 60 s. Sensors were then dipped into kinetic buffer for 60 s, CAIX solution (500 nM, 250 nM, 125 nM, 62.5 nM) for 600 s and into kinetic buffer for 600s. Measurements were carried out at 30 °C.
[0089] Sample preparation after AS: Samples from the affinity selection procedure were lyophilized and resuspended in 50 pL MQ 0.1% FA. The StageTips were prepared as described by Rappsilber et al. using C18 material from Empore SPE 47 mm discs (66883-U, Merck).43The StageTips were pre-conditioned with 200 pL MeOH, 200 pL of 0.1% (v / v) FA in MeCN and 200 pL of 0.1% (v / v) FA in MQ, respectively by centrifuging for 3 min at 300 rpm. The samples were then loaded on the StageTips and washed with 200 pL of 0.1% (v / v) FA in MQ.Compounds were eluted by adding 200 pL of 0.1% (v / v) FA in MeCN:MQ (7:3) and subsequently lyophilized before resuspending in 10 pL 0.1% (v / v) FA in UPLC-MS grade water. The samples were centrifuged for 5 min at 15.000 rpm. Afterwards, 9 pL was transferred to a LC-MS vial and 8 pL was injected into the LC-MS / MS system.Software
[0090] Proteowizard (MSconvert)44: Raw files (.raw) from the Orbitrap Exploris 240 mass spectrometer were converted to .mzML files using MSConvert selecting Peak Picking>MS levels 1-2. Resulting .mzML files were analyzed by using SIRIUS 6.0.4-SNAPSHOT software.
[0091] SIRIUS - COMET (6.0.4-SNAPSHOT): After importing the .mzML files, the COMET filter was applied using the following settings: scaffold formula: C8H3N2O for the SEL 2 and left blank for libraries 1 and 3, MS1 mass accuracy (ppm) = 5 ppm, considered fragment types = SEL 1 : S[0;1],S[1 ;2],0,2, SEL 2: S[0;2],S[1 ;2],0,1 , SEL 3: S[1 :2],0, minimum number of matching peaks = 1 , number of considered peaks = 5, number of allowed hydrogen shifts = 1 , MS2 mass accuracy (ppm) = 5, unless stated otherwise. The generated fragment list served as input for the building blocks field and was created using python scripts. Scans matching COMET were subsequently annotated using CSLFingerlD and ranked according to Epimetheus.
[0092] combinatorial Mass Encoding decoding Tool (COMET): After analyzing the final sample from the affinity selection experiment via LC-MS / MS, the obtained MS / MS spectra are annotated with potential hit compounds. To improve this annotation step, we developed COMET (Combinatorial Mass Encoding decoding Tool) for excluding those MS / MS spectra which are unrelated library molecules, either due to a non-matching precursor mass or a non-matching fragmentation pattern. Therefore, COMET aims to decrease the proportion of false positives in the final set of proposed compounds.
[0093] COMET consists of two consecutive filtering steps for removing MS / MS spectra that do not match to a library compound (Fig. 8). The first filter considers only the precursor mass of each MS / MS spectrum and excludes all those spectra whose precursor mass does not match to the mass of any library compound (i.e. all spectra for which no library candidate exists). In contrast, the second filtering step is based on our observation that some of the covalent bonds connecting the building blocks are cleaved more frequently during the collision-induced dissociation leading to corresponding fragment peaks in the spectra (Fig. 16D). Therefore, we derived certain fragmentation rules for each library to generate the most prominent fragments by cleaving only these specific bonds. For newly synthesized libraries, COMET allows the specification of new fragmentation rules. If such fragmentation information isn't known beforehand, it will generate all possible building block fragments by combinatorially cleaving each of those bonds. These fragmentation rules can now be used to check for each remaining MS / MS spectrum whether at least one corresponding candidate structure exists that can explain at least k fragment peaks in the spectrum by applying these fragmentation rules. However, if no candidate structure fulfils this condition, the spectrum will be removed.
[0094] KNIME 5.1.2 software was used for library enumeration and molecular property calculations. The following extensions were installed RDkit Nodes version 4.7., CDK version 1.5.6 and ChemAxon version 4.7.0Building Block Scoring & Selection
[0095] Fmoc-amino acid and carboxylic acid BBs were selected based on their druglike properties (molecular weight (MW), logP, hydrogen bond donors (HBD), hydrogen bond acceptors (HBA), and total polar surface area (TPSA)), using certain python scripts.
[0096] Chemspace provided lists of their in-stock Fmoc-amino acids and carboxylic acids, containing 1.7K and 6.3K building blocks (BBs), respectively. These lists were then refined to 1 ,000 BBs each by applying filters for price, availability, and compatibility with our library design. The refined lists were then used as input for the main_final_selection.py script, where the virtual library was enumerated based on the provided reaction SMARTS.
[0097] For each library member the molecular properties MW, HBD, HBA, logP and TPSA were calculated. The scoring of each library member was determined by abiding the following thresholds: MW<500, HBA<10, HBD<5, logP<5 and TPSA<140. Per matching properties it would score one point, resulting in a maximum score of 5 per library member. The scores were grouped per BB and the ranked BBs were used to select the most drug-like BBs for our library. Selected BBs can be seen in Section 10.1. List of building blocks.Optimization Studies for SEL 2
[0098] Fig. 10 outlines the optimization of trisubstituted benzimidazole synthesis on a solid support.
[0100] Starting material 9 was synthesized according to general procedure: Manual solidphase synthesis (SPS). TentaGel S NH2resin (90 pm, 0.24 mmol / g loading, 19 mg, 5 pmol, 1.0 eq.) was transferred to an eppendorf tube. Amine 9a-9cm (Fig. 11) (50 pmol, 10 eq.) and DIPEA (8.71 pL, 50 pmol, 10 eq.) in DMF (100 pL, 0.5 M) were added to the resin. The mixture was shaken at 600 rpm overnight at 80 °C. The resin was washed with DMF (5 x 2 mL) and DCM (5 x 2 mL) before cleaving with a solution of TFA:H2O:TIS (95:2.5:2.5) for 1 h and washed once with a solution of TFA:H2O:TIS (95:2.5:2.5). LC-MS samples were prepared by adding 10 pL of the TFA mixture with 90 pL of MeCN:MQ:tBuOH (1 :1 :1) The reaction mixture was characterized by LC-MS.
[0101] General procedure heterocyclization
[0102] Starting material 11 was synthesized as described in “General procedure nucleophilic aromatic substitution”. A solution of 1 .0 M SnCI2(62.5 eq.) in DMF was added to the resin (1 .0 eq.). The mixture was incubated at r.t. overnight, whereafter the resin was washed 50% MilliQ (in DMF, 5x), with DMF (5x,) and DCM (5x). TentaGel S NH2resin (90 pm, 0.24 mmol / g loading, 19 mg, 5 pmol, 1.0 eq.) was transferred to an eppendorf tube. A solution the appropriate aldehyde 13a-13co (Fig. 12) ((0.25 M, 25 pmol, 5 eq.) and p-TsOH*H2O (4.76 mg, 25 pmol, 5 eq.) in DMF (103 pL) was added to the resin. The mixture was incubated at r.t., overnight on 600 rpm. After incubation, the resin was washed DMF (5x) and DCM (5x). The resin was incubated for 1 hour with a solution of TFA:H2O:TIS (95:2.5:2.5) and washed once with a solution of TFA:H2O:TIS (95:2.5:2.5). LC-MS samples were prepared by adding 10 pL of the TFA mixture with 90 pL of MeCN:MQ:tBuOH (1 :1 :1) The reaction mixture was characterized by LC-MS.Optimization Studies for SEL 3
[0103] Table 3 shows the optimization of the Suzuki-Miyaura reaction on a solid support.Table 3a) Reaction conditions: compound 14 (5 pmol, 1.0 eq), phenylboronic acid (10 pmol, 2.0 eq),PdCI2(0.5 pmol, 10 mol%), XPhos (0.1 pmol, 20 mol%), K2CO3(10 pmol, 2 eq), solvent 9:1 DMF / H2O(100 pL), 80 °C b) Conversion was determined by LC-MS analysis.
[0104] General procedure Suzuki-Miyaura coupling16 19
[0105] Starting material 16 was synthesized according to general procedure: Manual solidphase synthesis (SPS). TentaGel S NH2resin (90 pm, 0.24 mmol / g loading, 4.8 mg, 5 pmol, 1.0 eq.) was transferred to an eppendorf tube. Boronic acid (10 pmol, 2 eq.), PdCI2(0.09 mg, 0.5 pmol, 0.1 eq.), XPhos (0.48 mg, 1 pmol, 0.2 eq.) and K2CO3(1.4 mg, 10 pmol, 2.0 eq.) were dissolved in DMF:H2O (9:1 , 100 pL) and added to the resin. The reaction was stirred at 700 rpm on 80°C for 2 hours. The resin was washed DMF (5x) and DCM (5x). The resin was incubated for 1 hour with a solution of TFA:H2O:TIS (95:2.5:2.5) and washed once with a solution of TFA:H2O:TIS (95:2.5:2.5). LC-MS samples were prepared by adding 10 pL of the TFA mixture with 90 pL of MeCN:MQ:tBuOH (1 :1 :1). The reaction mixture was characterized by LC-MS.Library Synthesis and Characterization
[0106] TentaGel S NH2resin (90 pm, 0.24 mmol / g loading, 2.625 g, 630 pmol, 1.0 eq.) was transferred to a 20mL fritted syringe and was subsequently swelled with DMF for 10 min. A solution of Fmoc-Rink-OH (1.019 g, 1.89 mmol, 3.0 eq), 0.4M HATU (1.877 mmol, 4.693 mL, 2.98 eq) and DIPEA (988 pL, 5670 pmol, 9.0 eq) in DMF was added to the resin and reacted for 2h. The resin was washed with DMF (5 x 2 mL) and 20% piperidine in DMF (1x, 2 mL) before incubating with 20% piperidine in DMF for 10 min.
[0107] The resin was divided over 62 x 2 mL fritted syringes. Stock solutions of amino acids were made with Fmoc-protected amino acids (100 pmol) in 0.4M HATU (250 pL). A solution of Fmoc-protected amino-acid (10 pmol, 3.0 eq, (75 pL from premade stock solution), DIPEA (15.68 pL, 90.00 pmol, 9.0 eq) and DMF (75 pL) were added to the resin (10 pmol). After 2h theresin was combined in a 20 mL fritted syringe and was washed with DMF (5 x 2 mL) and 20% piperidine in DMF (1x, 2 mL) before incubating with 20% piperidine in DMF for 10 min. The resin was washed with DMF (5 x 2 mL) before adding the next Fmoc-protected amino-acid.
[0108] The resin was divided over 62 x 2 mL fritted syringes. Stock solutions of amino acids were made with Fmoc-protected amino acids (100 pmol) in 0.4M HATU (250 pL). A solution of Fmoc-protected amino-acid (10 pmol, 3.0 eq , (75 pL from premade stock solution), DIPEA (15.68 pL, 90.00 pmol, 9.0 eq) and DMF (75 pL) were added to the resin (10 pmol). After 2h the resin was combined in a 20 mL fritted syringe and was washed with DMF (5 x 2 mL) and 20% piperidine in DMF (1x, 2 mL) before incubating with 20% piperidine in DMF for 10 min. The resin was washed with DMF (5 x 2 mL) before adding the next Fmoc-protected amino-acid.
[0109] The resin (350 umol, 1.0 eq) was divided over 130 eppendorf tubes. A solution of carboxylic acid (8.08 pmol, 3.0 eq), HATU (80.2 pL, 0.1 M, 8.02 pmol, 2.98 eq) and DIPEA (4.22 pL, 24.23 pmol, 9.0 eq) in DMF was added to the resin (2.69 pmol). The reactions were stirred overnight at r,t. The resin was pooled and washed with DMF (5x) and DCM (5x).
[0110] The resin was incubated for 1.5 hour with a solution of TFA:H2O:TIS (95:2.5:2.5) and washed once with a solution of TFA:H2O:TIS (95:2.5:2.5). TFA was evaporated and the library was purified with using reverse phase column chromatography with a step-wise gradient 00-70- 100% MeCN.SEL 2 Benzimidazole
[0111] TentaGel S NH2resin (90 pm, 0.24 mmol / g loading, 1.3150 g, 315 pmol, 1.0 eq.) was transferred to a 20mL fritted syringe and was subsequently swelled with DMF for 10 min. A solution of Fmoc-Rink-OH (509.9 mg, 0.945 mmol, 3.0 eq), 0.4M HATU (0.939 mmol, 2.346 mL, 2.98 eq) and DIPEA (494 pL, 2.835 mmol, 9.0 eq) in DMF was added to the resin and reacted for 2h.The resin was washed with DMF (5 x 2 mL) and 20% piperidine in DMF (1x, 2 mL) before incubating with 20% piperidine in DMF for 10 min.
[0112] The resin was divided over 62 x 2 mL fritted syringes. Stock solutions of amino acids were made with Fmoc-protected amino acids (100 pmol) in 0.4M HATU (250 pL). A solution of Fmoc-protected amino-acid (3.0 pmol, 3.0 eq , (37.5 pL from premade stock solution), DIPEA (7.84 pL, 45.00 pmol, 9.0 eq) and DMF (37.5 pL) were added to the resin (5.00 pmol). After 2hthe resin was combined in a 20 mL fritted syringe and was washed with DMF (5 x 2 mL) and 20% piperidine in DMF (1x, 2 mL) before incubating with 20% piperidine in DMF for 10 min. The resin was washed with DMF (5 x 2 mL) before adding the next Fmoc-protected amino-acid.
[0113] The resin was pooled and a solution of 4-fluoro-3-nitrobenzoic acid (175 mg, 945 pmol, 3.0 eq), HATU (2.346 mL, 939 pmol, 2.98 eq) and DIPEA (494 pL, 2.835 mmol, 9.0 eq) in DMF was added to the resin. After 1 h the reaction as washed with DMF (5 x 2 mL).
[0114] The resin was divided over 52, 2 mL fritted syringes. TentaGel S NH2resin (30 pm, 0.24 mmol / g loading, 5.96 pmol, 1 .0 eq.) was transferred to an eppendorf. Amines (59.62 pmol, 10 eq.) and DIPEA (10.55 pL, 59.62pmol, 10 eq.) in DMF (150pL, 0.4 M). The mixture was shaken at 600 rpm overnight at 80 °C. The resin was washed with DMF (5 x 2 mL) and DCM (5 x 2 mL).
[0115] The resin was incubated for 1.5 hour with a solution of TFA:H2O:TIS (95:2.5:2.5) and washed once with a solution of TFA:H2O:TIS (95:2.5:2.5). TFA was evaporated and the library was purified with using reverse phase column chromatography with a step-wise gradient 00-70- 100% MeCN.SEL 3: Suzuki-Miyaura library
[0116] TentaGel S NH2resin (30 pm, 0.24 mmol / g loading, 1.104 g, 265 pmol, 1.0 eq.) was transferred to a 20mL fritted syringe and was subsequently swelled with DMF for 10 min. A solution of Fmoc-Rink-OH (429 mg, 0.795 mmol, 3.0 eq), 0.4M HATU (0.790 mmol, 1.974 mL, 2.98 eq) and DIPEA (415 pL, 2.385 mmol, 9.0 eq) in DMF was added to the resin and reacted for xh. The resin was washed with DMF (5 x 2 mL) and 20% piperidine in DMF (1x, 2 mL) before incubating with 20% piperidine in DMF for 10 min.
[0117] The resin was divided over 60 eppendorf tubes. A solution of Fmoc-protected aminoacid (12.82 pmol, 3.0 eq), 0.4M HATU (12.74 pmol, 31.8 pL, 2.98 eq), DIPEA (6.70 pL, 38.47 pmol, 9.0 eq) and DMF (90 pL) were added to the resin (4.27 pmol). After 4h the resin was combined in a 20 mL fritted syringe and was washed with DMF (5 x 2 mL) and 20% piperidine in DMF (1x, 2 mL) before incubating with 20% piperidine in DMF for 10 min. The resin was combined in a 20 mL fritted syringe and was washed with DMF (5x) .
[0118] The resin was divided over 10 eppendorf tubes. A solution of aryl bromide (26.50 pmol, 3.0 eq), 0.4M HATU (78.97 pmol, 197 pL, 2.98 eq), DIPEA (41.50 pL, 238.5 pmol, 9.0 eq) and DMF (400 pL) were added to the resin (26.50 pmol). After 4h was, the resin was washed with DMF (5 x 2 mL). Aryl bromide AB10 was deprotected by washing with 20% piperidine in DMF (1x, 2 mL) before incubating with 20% piperidine in DMF for 10 min. A solution of benzoic acid(9.7 mg, 79.50 pmol, 3 eq.), DIPEA (41.50 pL, 238 pmol, 9.0 eq) and DMF (400 pL) were added to the resin containing AB10 (26.50 pmol). The resin was combined in a 20 mL fritted syringe and was washed with DMF (5x).
[0119] The resin was divided over 53 eppendorf tubes. Boronic acid (10 pmol, 2.0 eq.), K2CO3 (1.4 mg, 10 pmol, 2.0 eq.), PdCI2(89 pg, 0.5 pmol, 10 mol%), Xphos (0.48 mg, 1 pmol, 20 mol%) and DMF (100 pL) were added to the resin (5.0 pmol). The reaction was stirred at 800 rpm at 80 °C for 21 h. The resin was washed with DMF (5 x 2 mL) and DCM (5x) and incubated for 1 :45 h with a solution of TFA:H2O:TIS (95:2.5:2.5). The resin was washed once with a solution of TFA:H2O:TIS (95:2.5:2.5) whereafter the TFA was evaporated. The library was purified with using reverse phase column chromatography with a stepwise gradient 00-70-100% MeCN.Discussion
[0120] Aiming at affinity selection with large and diverse self-encoded libraries, we established solid phase synthesis protocols for the preparation of combinatorial libraries with different scaffold designs. By exploring different scaffolds, we aimed both at increasing diversity and at investigating the amenability of different molecular architectures for MS / MS based decoding. In order to obtain high quality combinatorial libraries, each reaction step needs to be efficient and high yielding. Self-encoded library 1 (SEL 1) is formed by the sequential attachment of two amino acid building blocks, followed by the addition of a carboxylic acid decorator using reaction conditions optimized for Fmoc-based solid phase peptide synthesis (Fig. 9A). Self-encoded library 2 (SEL 2) is based on a benzimidazole core decorated on three different positions. Based on previously described methodologies (31-33) and following systematic optimization, we established an efficient route towards trifunctional benzimidazoles (Fig. 9B and Fig. 10). The benzimidazole decorators which confer diversity to the library are based on an amino acid building block, a primary amine and an aldehyde. We tested the scope of the nucleophilic aromatic substitution with a set of 92 primary amines, of which a large fraction resulted in reasonable conversions for combinatorial synthesis (65%). (Fig. 9B and Fig. 11). We then investigated the heterocyclization efficiency using 95 aldehydes, with 65 resulting in a > 55% conversion to the final trifunctional compounds (Fig. 9B and Fig. 12). Self-encoded library 3 (SEL3) results from an amino acid building block linked to an aryl bromide and subsequently cross-coupled to a boronic acid, utilizing the palladium catalyzed Suzuki-Miyaura reaction (34). We optimized reaction conditions on a selected scaffold and then tested the scope of 19 bifunctional aryl bromides of which 9 resulted in a > 65% conversion (Fig. 9C and Fig. 13). Out of 85 boronic acids, 50 resulted in a > 65% conversion (Fig. 9C and Fig. 14). Representative crude LC-MS traces show the quality of the synthesis for each library scaffold (Fig. 9A-C).
[0121] Using a virtual library scoring script we selected building blocks to generate libraries with optimized drug-like properties. For SEL 1 , we anticipated no synthetic limitations, so wedecided to select our initial set of building blocks based on their drug-like properties, while limiting isobaric fragments. We filtered a comprehensive building block (BB) catalog (Fmoc- amino acids and carboxylic acids from Chemspace) by availability and price, narrowing it down to 1000 BBs per position. With these we enumerated a virtual library with a billion members. Each member was scored based on five Lipinski parameters: molecular weight (mw), logP, hydrogen bond donors (HBD), hydrogen bond acceptors (HBA), and topological polar surface area (TPSA). Each library members received a point for each satisfied parameter, which was then translated to a combined score per building block. This ranking allowed us to select and purchase top-scoring building blocks (62 amino acids and 130 carboxylic acids). Through solidphase split and pool synthesis, we generated a half million membered SEL 1 . Compared to the original enumerated library, all Lipinski parameters of our SEL 1 were significantly improved (Fig. 15). For SEL 2 and 3, we used one of the optimized amino acids for BB1 and selected BBs with yields greater than 55% in scope analysis for the other positions (resulting in 216,008 members for the benzimidazole scaffold and 31 ,800 members for the Suzuki based library). Overall, the majority of all compounds in the three libraries satisfy the requirements for drug-like properties (Fig. 9D).Library Decoding Software
[0122] A crucial step in the affinity selection workflow is the accurate identification of hit compounds. The final sample from an affinity selection process is always of unknown complexity and may contain up to a few hundred compounds. With the high degree of mass degeneracy in or libraries (i.e., the presence of multiple isobaric compounds with different molecular structures) structure annotation based on MS / MS fragmentation spectra is essential for unequivocal compound identification (Fig. 16A). To investigate the decodability of our libraries and simulate the final sample from an affinity selection, we prepared defined subsets of 245-500 compounds for each scaffold and analyzed each subset via nanoLC-MS / MS (Fig. 16B). Based on the known list of library compounds we analyzed the data and counted the detectable compounds - the number of detectable compounds might be lower than the theoretical number because of loss during sample preparation of highly polar / unipolar structures). Overall, each nanoLC-MS / MS run produced approximately 80,000 MS1 and MS2 scans, including mainly spectra resulting from background noise: for a real affinity selection sample of unknown content and complexity the manual analysis of such a dataset would be totally impractical.
[0123] We first attempted the automated structure annotation for our test samples with off the shelf metabolomics software. While typical metabolomics workflows use spectral databases as an input to increase recall rates and confidence and correct structure annotation, our libraries representing novel chemical matter do not have such spectral databases. To this end we utilized SIRIUS 6 with CSLFingerlD, considered best-in-class for reference spectra freestructure annotation of small molecules (35-37). SIRIUS annotates compounds by scoring predicted molecular fingerprints against fingerprints of database structures (e.g. PubChem). In an affinity selection experiment, contrary to a regular metabolomics analysis, the complete space of potential structures is already known and the fully enumerated library can be used as a structure database to score compounds against. To this end, we created custom structure databases in SIRIUS, consisting of the fully enumerated library SMILES for each library (500k, 200k and 30k respectively). We then imported the measured nanoLC-MS / MS runs into SIRIUS and performed a standard SIRIUS structure annotation workflow. While a good fraction of library compounds were correctly detected and annotated (82%, 77% and 90%, respectively for SEL 1 , 2 and 3), the total number of annotated scans (up to 2800) largely exceeded the number of molecules actually present in the library (Fig. 16C). Manual inspection in principle can help picking out correct annotations but with thousands of ‘false-positives’ this off the shelf automated annotation workflow would still be impractical for real affinity selections samples with unknown content.
[0124] To increase the proportion of genuine library molecules in the final set of proposed compounds and to improve correct compound annotation, we set off to establish fragmentation rules to predict likely MS / MS patterns for our library compounds. We thus analyzed the fragmentation spectra of our test library compounds and calculated fragmentation frequencies of bonds connecting the various building blocks (Fig. 16D). For each of the three library scaffolds we defined the most prominent recurring fragmentation modes. Based on these patterns, we generated a combinatorial fragmentor to create a list of predicted fragments for each library member (Fig. 16E). We then implemented a filter in SIRIUS, stipulating that only scans with an MS1 precursor mass matching a library compound’s mass and containing at least one predicted fragment peak in their MS2 would be selected for full annotation (Fig. 16F). This filter drastically reduced the number of total scans (compare Figure 16C and 16F), while maintaining a high correct recall and annotation rate of 66-74%. Overall, this SIRIUS- COmbinatorial Mass Encoding decoding Tool (SIRIUS-COMET) enables the high fidelity annotation and decoding of satisfying numbers of compounds from all three library scaffolds, with significantly reduced numbers hallucinated annotations.Affinity Selection Against CAIXProcedure
[0125] MyOne Streptavidin T1 Dynabeads (100 pL of 10 mg / mL stock per well, measured in duplicates) were washed with 3x 1 mL 10% FBS, 1x PBS, 0.02% Tween-20. The beads were incubated with biotinylated CAIX (100 pL, 1.5 pM per well) in 10% FBS, 1x PBS, 0.02% Tween- 20 for 1 h at 4°C. The beads were washed 2x 1 mL 10% FBS, 1x PBS, 0.02% Tween-20, 400 pM d-biotin and 1x 1 mL 10% FBS, 1x PBS, 0.02% Tween-20 before incubating with the library (100 pL, 100 fmol / member per well) in 10% FBS, 1x PBS for 1 h at 4°C. The beads werewashed with 5x 1 mL of 1x PBS and subsequently eluted with 2x MeCN:MQ (1 :1) 0.1%FA (100 pL per well).
[0126] Sample preparation after AS: Samples from the affinity selection procedure were lyophilized and resuspended in 50 pL MQ 0.1 %FA. The StageTips were prepared as described by Rappsilber et al. using C18 material from Empore SPE 47 mm discs (66883-U, Merck).1The StageTips were pre-conditioned with 200 pL MeOH, 200 pL of 0.1% (v / v) FA in MeCN and 200 pL of 0.1% (v / v) FA in MQ, respectively by centrifuging for 3 min at 300 rpm. The samples were then loaded on the StageTips and washed with 200 pL of 0.1% (v / v) FA in MQ. Compounds were eluted by adding 200 pL of 0.1% (v / v) FA in MeCN:MQ (7:3). The samples were lyophilized before resuspending in 10 pL 0.1% (v / v) FA in UPLC-MS grade water. The samples were centrifuged for 5 min at 15.000 rpm. Afterwards, 9 pL was transferred to a LC-MS vial and 8 pL was injected into the LC-MS / MS system.Screening of SEL 1 against CAIX
[0127] Using SIRIUS:COMET filter with a least matching 1 out of the 5 biggest peaks and 5 ppm accuracy in combination with CSI:FingerlD structure annotation identified 74 unique structures containing building block CA77. The structures of the 74 hit compounds found after AS-MS against CAIX with SEL 1 are provided below:
[0128] The top ranked structures by Epimetheus were taken and deconstructed into the respective building blocks.
[0129] The chemical structures of hit 20 and 21 identified using SIRIUS:COMET, with the extracted ion chromatogram (EIC) and LC-MS / MS chromatogram from the affinity selection of SEL 1 against CAIX are shown in Fig. 17.Screening of SEL 2 against CAIX
[0130] Using SIRIUS:COMET filter with a least matching 1 out of the 5 biggest peaks and 5 ppm accuracy in combination with CSI:FingerlD structure annotation identified 30 unique structures containing building block AM42. The structures of the 30 hit compounds found after AS-MS against CAIX with SEL 2 are provided below:
[0131] The top ranked structures by Epimetheus were taken and deconstructed into the respective building blocks.
[0132] Chemical structures of hit 24 and 25 identified using SIRIUS:COMET, with the extracted ion chromatogram (EIC) and LC-MS / MS chromatogram from the affinity selection of SEL 2 against CAIX is shown in Fig. 18.Screening of SEL 3 against CAIX
[0133] Using SIRIUS:COMET filter with a least matching 1 out of the 5 biggest peaks and 5 ppm accuracy in combination with CSI:FingerlD structure annotation identified 47 unique structures containing building block BA53. The structures of the 47 hit compounds found after AS-MS against CAIX with SEL 3 are provided below:
[0134] The top ranked structures by Epimetheus were taken and deconstructed into the respective building blocks.Hit-Identification (LC-MS / MS)
[0135] The chemical structures of hit 38 identified using SIRIUS:COMET with the extracted ion chromatogram (EIC) and LC-MS / MS chromatogram from the affinity selection of SEL 3 against CAIX is shown in Fig. 19.Discussion
[0136] With the three SELs in hand and the automated SIRIUS-COMET software established we initiated de novo ligand discovery experiments with the oncology drug target carbonic anhydrase IX (CAIX). CAIX has been used previously to benchmark novel DNA-encoded libraries (77, 16, 17) and is a therapeutically relevant target for cancer treatment, particularly in hypoxic tumors, due to its role in tumor cell survival and pH regulation in the tumor microenvironment (38). We immobilized biotinylated CAIX on streptavidin coated magnetic beads and incubated it with the half million membered SEL 1. After washing away non-binders we eluted potential hits with H2O / MeCN / FA (50:50:0.01) and analyzed the resulting sample with our nanoLC-MS / MS SIRIUS-COMET workflow. To exclude unspecific binders, we also ran the library against unfunctionalized magnetic beads and implemented background subtraction into SIRIUS-COMET that removes spectra from CAIX runs if they result from features present also in the control runs. From the resulting annotated spectra, we plotted the building block frequencies for each position and found a significant enrichment of 4-sulfamoylbenzoic acid in the carboxylic acid position (Fig. 20A). This finding aligns well with known carbonic anhydrase binders, where the aromatic sulfonamide forms a strong interaction with a zinc ion in the binding pocket (39). Among the annotated structures, 74 different hits contained this substructure and these compounds showed clear extracted ion peaks in the CAIX sample, while they were not detectable in the control runs (see Fig. 20B for selected examples). We also performed the selection with SEL 2 and 3 and identified aromatic sulfonamides as the most enriched BBs (30 and 47 individual compounds, for SEL 2 and 3 respectively). Adjacent positions also showed preferences for specific BBs (Fig. 20A), indicating the combination of the aryl sulfonamide with these BBs might result in preferred scaffolds for CAIX binding. We selected hits from the three different libraries for resynthesis and binding validation. All compounds demonstrated single digit nanomolar association with CAIX, as shown by biolayer interferometry (Fig. 20B). Taken together, our platform enables selection of high affinity binders from libraries with diverse scaffold architecture.
[0137] When screening diverse libraries by affinity selections, achieving high enrichments, i.e., removing non-binders from the final pool of selected compounds, and therefore increasing the binder / non-binder ratio compared to the full library, is essential. The enrichment can be calculated as follows: [(# found binders) / (# compounds after selection)] :[(# total binders) / (# totallibrary)]. Given that actual total number of binders presents in the library is unknown, usually the maximum enrichment is calculated, postulating the number of found binders equal to the total number of binders (29, 30). The enrichments obtained in our three library selections are 2.2 x 103(74 / 228:74 / 499,720), 2.9 x 103(30 / 75:30 / 216,008) and 6 x 102(47 / 51 :47 / 31 ,800) respectively for SEL 1 , 2 and 3. It is noteworthy that SEL 3 achieved an almost perfect enrichment as 47 of the 51 identified scans contain binders. These enrichment scores are in good agreement with typical phage display and peptide AS-MS selections which also achieve enrichments in the order of magnitude of 103(2, 40).
[0138] We also investigated the effect of library concentration on ligand identification. Considering solubility limits of any library, a lower initial concentration of each member for the affinity selection would enable the screening of larger libraries. We performed the selection CAIX ligands with SEL1 at 1 pmol / member, 100 fmol / member and 10 fmol / member. For 1 pmol / member to 100 fmol / member there is only a 1.3-fold decrease in the absolute number of hits. Lowering the concentration to 10 fmol / member did result in only 1 hit recovered. See Table 4.
[0139] Table 4. Selection against CAIX with a decreasing library concentration.
[0140] Given that the selection results in similar outcomes at 1 pmol / member and 100 fmol / member, it indicates that we might be able to run affinity selections with 10-times larger libraries, i.e., 5-million members, using 100 fmol / member.
[0141] To explore whether the hit discovery process could be further accelerated and streamlined, we conducted affinity selection using a pooled combination of all three libraries. This combined library encompasses approximately 750,000 members, representing a higher degree of molecular diversity. Encouragingly, sulfonamide-based CAIX binders were identified across all three library scaffolds. Although the pooled approach yielded fewer hits compared to individual library selections (90 vs. 150 total hits), it demonstrates the potential for applying this workflow to high-diversity libraries with varied scaffold architectures.Affinity Selection Against EXO1Procedure
[0142] MyOne Streptavidin T1 Dynabeads (100 pL of 10 mg / mL stock per well, measured in duplicates) were washed with 3x 1 mL 50 mM HEPES, 100 mM KCI, 5 mM MgCI2, 1 mM DTT, 10% FBS, 0.02% Tween-20, pH = 7.5. The beads were incubated with biotinylated CAIX (100pL, 1.5 pM per well) in 50 mM HEPES, 100 mM KCI, 5 mM MgCI2, 1 mM DTT, 10% FBS, 0.02% Tween-20, pH = 7.5 for 1 h at 4°C. The beads were washed 2x 1 mL 50 mM HEPES, 100 mM KCI, 5 mM MgCI2, 1 mM DTT, 10% FBS, 0.02% Tween-20, 400 pM biotin pH = 7.5 and 1x 1 mL 50 mM HEPES, 100 mM KCI, 5 mM MgCI2, 1 mM DTT, 10% FBS, 0.02% Tween-20, pH = 7.5 before incubating with the library (100 pL, 100 fmol / member per well) in 50 mM HEPES, 100 mM KCI, 5 mM MgCI21 mM DTT, 10% FBS, pH = 7.5 for 1 h at 4°C. The beads were washed with 5x 1 mL of 50 mM HEPES, 100 mM KCI, 5 mM MgCI2, 1 mM DTT, pH = 7.5 and subsequently eluted with 2x MeCN:MQ (1 :1) 0.1%FA (100 pL per well).
[0143] Sample preparation after AS: Samples from the affinity selection procedure were lyophilized and resuspended in 50 pL MQ 0.1 %FA. The StageTips were prepared as described by Rappsilber et al. using C18 material from Empore SPE 47 mm discs (66883-U, Merck).1The StageTips were pre-conditioned with 200 pL MeOH, 200 pL of 0.1% (v / v) FA in MeCN and 200 pL of 0.1% (v / v) FA in MQ, respectively by centrifuging for 3 min at 300 rpm. The samples were then loaded on the StageTips and washed with 200 pL of 0.1% (v / v) FA in MQ. Compounds were eluted by adding 200 pL of 0.1% (v / v) FA in MeCN:MQ (7:3). The samples were lyophilized before resuspending in 10 pL 0.1% (v / v) FA in UPLC-MS grade water. The samples were centrifuged for 5 min at 15.000 rpm. Afterwards, 9 pL was transferred to a LC-MS vial and 8 pL was injected into the LC-MS / MS system.Screening of SEL 1 against EX01
[0144] Chemical structures of isomers 40 and 41 identified using SIRIUS:COMET, with the extracted ion chromatogram (EIC) and LC-MS / MS chromatogram from the affinity selection of SEL 1 against EXO1 are provided in Fig. 21.Discussion
[0145] As a second target to test the affinity selection hit decoding workflow, we used Exonuclease 1 (EXO1). EXO1 plays a crucial role in DNA mismatch repair and maintaining genomic stability. BRCA1 -deficient tumours are dependent on EXO1 for survival and makes it a promising target for drug discovery (47). Fig. 20C shows the results of SEL selection against biotinylated EXO1. While no single building block was clearly enriched (Fig. 22), the SIRIUS- COMET workflow identified MS spectra matching the library design and not being present in the control sample. The fragmentation spectra pointed towards either compound 6 or compound 7, with no possibility of differentiation because of the presence of two isobaric BBs in our library design. We resynthesized both compounds and tested their association to EXO1 by BLI. Only compound 6 showed association to EXO1 (130 nM), while compound 7 did not show any binding. In a DNA cleavage assay we did not detect activity of compound 6. Given that affinity selections may lead to identification of non-competitive binders unless hits are specifically biased, we reason that 6 binds to a different site from the DNA on the EXO1 protein.Nevertheless, our experiments support the potential of our platform to rapidly identify novel ligands for biomolecular targets.Discussion
[0146] Herein we have described the use of barcode-free self-encoded libraries (SELs) for affinity selection-based hit discovery. Our approach leverages tandem MS fragmentation to accurately reconstruct the molecular structure of compounds selected from vast libraries, eliminating the need for large barcodes as required by previous platforms like DELs. The SEL strategy achieves the maximal possible information density for library selections (defining information density as [molecule mass] / [molecule + decoding tag mass]), significantly larger than commonly used DNA encoded molecules. To support the broad validity of our strategy, we have tested the decoding on three different library architectures: after developing a new software platform, SIRIUS-COMET, we could efficiently decode compounds from all three scaffolds via tandem MS in an automated fashion. The diversity in chemical connectivities and the drug-likeness of our libraries is of great promise for the potential expansion of our technology to many more interesting library architectures. We also performed affinity selection with libraries containing up to 750 thousand members. In these selections we decoded and validated several nanomolar binders for CAIX and a binder for EXO1 , demonstrating the applicability of our methodology to drug discovery campaigns.
[0147] In large libraries, mass degeneracy, i.e., the presence of many isobaric compounds is inevitable, necessitating the development of new software for automated high-fidelity decoding of library members and affinity selection hits. Our SIRIUS-COMET workflow, inspired by advances in metabolomics software, was developed to address this challenge. Unlike classical proteomics and metabolomics, which rely on spectral databases for high-confidence compound identification, synthetic SELs lack such databases, requiring a more challenging de novo compound annotation. The SIRIUS-COMET workflow with an integrated combinatorial fragmentor simplifies the combinatorial complexity of the library compounds into predictable fragments and enables the correct annotation of library compounds with good recall rates and low numbers of hallucinated annotations. This new software tool aids the analysis of complex affinity selection samples resulting from the screening of large libraries.
[0148] Beyond superior information density, SELs offer considerable advantages in synthesis complexity and scope. DELs require an alternation of enzymatic ligation steps and chemical reactions performed on the oligonucleotide linked scaffold. While an impressive number of compatible reactions have been developed (18), there are always concerns of potential degradation of the DNA barcode (19). Rbssler et al. recently reported the DNA lability under a number of conditions, and shown, on the other side, how peptide synthesized on a solid support can withstand many harsh conditions, required for small molecule synthesis (19). In that study the authors point out advantages of peptides as encoding tags of small molecules. Thesynthesis of a such a peptide encoded compounds requires > 40 reaction steps in addition to the utilization of orthogonal cleavable linkers. SELs can be synthesized in as few as five straightforward reaction steps on solid supports. Without an encoding tag, SELs can undergo any reaction condition compatible with the small molecule itself. We successfully demonstrated several transformations, including amide bond formations, nucleophilic aromatic substitutions, heterocyclizations, cross couplings and acid and base mediated protecting group removals. We predict that many more transformations will be applied for the preparation of SELs in the future. Notably, high-diversity SELs derived from our library scaffolds can be synthesized in under a week using standard organic synthesis techniques, making this approach accessible to almost any laboratory. Given that most institutions already have mass spectrometry equipment and the necessary facilities, our streamlined synthesis, selection, and analysis workflow positions SELs to democratize rapid early-stage drug discovery.
[0149] The absence of a barcoding tag in SELs can offer significant advantages during affinity selection. Barcoding tags, often larger than the molecules themselves, can interact undesirably with the target, leading to false positives, especially in DEL selections involving nucleic acidbinding proteins. Moreover, the bulky tags in DELs and PELs limit the conformational flexibility of library molecules, restricting how they can bind to their targets. SEL compounds, free from these constraints, avoid these pitfalls and potentially sample a much broader conformational diversity during affinity selection. In addition, mis-encoding can be circumvented: with the synthesis of barcoded compounds proceeding with distinct steps for barcode and molecule construction, any deletion or truncation on either part of the construct will not correspond to the corresponding entity on the other part of the molecule, leading to a mismatch between actual molecule and barcode (17, 42). Self-encoded molecules, self-containing all the information about their structure, obviate such problematic.
[0150] With SEL library sizes in the range of one million compounds, we match the scale of several recent successful DELs (11, 12). While DEL libraries can theoretically reach billions of compounds, recent evidence suggests that DEL selections work best with inputs exceeding 106individual copies per compounds, effectively capping optimal library sizes at a few million members (16, 43). Unlike DELs, SEL hits cannot be amplified, so the input quantity for a selection must account for the sensitivity limits of the mass spectrometer used for detection and decoding. However, modern instruments like Orbitrap spectrometers can typically detect as little as ~1 fmol per molecule, or even less. In this study, we determined that 100 fmol per member Peris an ideal input for affinity selection, in principle allowing for screenings with at least 5 million compounds simultaneously, remaining within the library solubility limits.
[0151] In summary, our findings demonstrate a barcode-free technology with large selfencoded libraries for early drug discovery. We anticipate that this approach will see widespread adoption in both academic and industrial research settings.References and Notes1. R. MacArron, M. N. Banks, D. Bojanic, D. J. Burns, D. A. Cirovic, T. Garyantes, D. V. S. Green, R. P. Hertzberg, W. P. Janzen, J. W. Paslay, U. Schopfer, G. S. Sittampalam, Impact of high-throughput screening in biomedical research. Nature Reviews Drug Discovery 2011 10:3 10, 188-195 (2011).2. L. Q. Koh, Y. W. Lim, Z. P. Gates, Affinity Selection from Synthetic Peptide Libraries Enabled by De Novo MS / MS Sequencing. Int J Pept Res Ther 28, 1-14 (2022).3. J. M. Mata, E. van der Nol, S. J. Pomplun, Advances in Ultrahigh Throughput Hit Discovery with Tandem Mass Spectrometry Encoded Libraries. J Am Chem Soc 145, 19129— 19139 (2023).4. A. Gironda-Mart nez, E. J. Donckele, F. Samain, D. Neri, DNA-Encoded Chemical Libraries: A Comprehensive Review with Succesful Stories and Future Challenges. ACS Pharmacol Transl Sci 4, 1265-1279 (2021).5. R. Prudent, D. A. Annis, P. J. Dandliker, J. Y. Ortholand, D. Roche, Exploring new targets and chemical space with affinity selection-mass spectrometry. Nat Rev Chem 5, 62-71 (2021).6. R. Barderas, E. Benito-Pena, The 2018 Nobel Prize in Chemistry: phage display of peptides and antibodies. Anal Bioanal Chem 411 , 2475-2479 (2019).7. C. Heinis, T. Rutherford, S. Freund, G. Winter, Phage-encoded combinatorial chemical libraries based on bicyclic peptides. Nat Chem Biol 5, 502-507 (2009).8. Y. Huang, M. M. Wiedmann, H. Suga, RNA Display Methods for the Discovery of Bioactive Macrocycles. Chem Rev 119, 10360-10391 (2019).9. R. W. Roberts, J. W. Szostak, RNA-peptide fusions for the in vitro selection of peptides and proteins. Proc Natl Acad Sci U S A 94, 12297-12302 (1997).10. S. Brenner, R. A. Lerner, Encoded combinatorial chemistry. Proc Natl Acad Sci U S A 89, 5381-5383 (1992).11. N. Favalli, G. Bassi, C. Pellegrino, J. Millul, R. De Luca, S. Cazzamalli, S. Yang, A. Trenner, N. L. Mozaffari, R. Myburgh, M. Moroglu, S. J. Conway, A. A. Sartori, M. G. Manz, R. A. Lerner, P. K. Vogt, J. Scheuermann, D. Neri, Stereo- and regiodefined DNA-encoded chemical libraries enable efficient tumour-targeting applications. Nat Chem 13, 540-548 (2021).12. D. L. Usanov, A. I. Chan, J. P. Maianti, D. R. Liu, Second-generation DNA-templated macrocycle libraries for the discovery of bioactive small molecules. Nat Chem 10, 704-714 (2018).13. M. A. Clark, R. A. Acharya, C. C. Arico-Muendel, S. L. Belyanskaya, D. R. Benjamin, N. R. Carlson, P. A. Centrella, C. H. Chiu, S. P. Creaser, J. W. Cuozzo, C. P. Davie, Y. Ding, G. J. Franklin, K. D. Franzen, M. L. Gefter, S. P. Hale, N. J. V. Hansen, D. I. Israel, J. Jiang, M. J. Kavarana, M. S. Kelley, C. S. Kollmann, F. Li, K. Lind, S. Mataruse, P. F. Medeiros, J. A.Messer, P. Myers, H. O’Keefe, M. C. Oliff, C. E. Rise, A. L. Satz, S. R. Skinner, J. L. Svendsen, L. Tang, K. Van Vloten, R. W. Wagner, G. Yao, B. Zhao, B. A. Morgan, Design, synthesis and selection of DNA-encoded small-molecule libraries. Nat Chem Biol 5, 647-654 (2009).14. J. W. Mason, Y. T. Chow, L. Hudson, A. Tutter, G. Michaud, M. V. Westphal, W. Shu, X. Ma, Z. Y. Tan, C. W. Coley, P. A. Clemons, S. Bonazzi, F. Berst, K. Briner, S. Liu, F. J. Zech, S. L. Schreiber, DNA-encoded library-enabled discovery of proximity-inducing small molecules. Nat Chem Biol 20, 170-179 (2024).15. Y. Huang, L. Meng, Q. Nie, Y. Zhou, L. Chen, S. Yang, Y. M. E. Fung, X. Li, C. Huang, Y. Cao, Y. Li, X. Li, Selection of DNA-encoded chemical libraries against endogenous membrane proteins on live cells. Nat Chem 13, 77-88 (2021).16. S. Oehler, L. Lucaroni, F. Migliorini, A. Elsayed, L. Prati, S. Puglioli, M. Matasci, K. Schira, J. Scheuermann, D. Yudin, M. Jia, N. Ban, D. Bushnell, R. Kornberg, S. Cazzamalli, D. Neri, N. Favalli, G. Bassi, A DNA-encoded chemical library based on chiral 4-amino-proline enables stereospecific isozyme-selective protein recognition. Nat Chem 15, 1431-1443 (2023).17. M. Keller, D. Petrov, A. Gloger, B. Dietschi, K. Jobin, T. Gradinger, A. Martinelli, L. Plais, Y. Onda, D. Neri, J. Scheuermann, Highly pure DNA-encoded chemical libraries by dual-linker solid-phase synthesis. Science 384, 1259-1265 (2024).18. S. Chines, C. Ehrt, M. Potowski, F. Biesenkamp, L. Griitzbach, S. Brunner, F. van den Broek, S. Bali, K. Ickstadt, A. Brunschweiger, Navigating chemical reaction space - application to DNA-encoded chemistry. Chem Sci 13, 11221-11231 (2022).19. S. L. Rbssler, N. M. Grob, S. L. Buchwald, B. L. Pentelute, Abiotic peptides as carriers of information for the encoding of small-molecule library synthesis. Science (1979) 379, 939-945 (2023).20. K. Gbtte, S. Chines, A. Brunschweiger, Reaction development for DNA-encoded library technology: From evolution to revolution? Tetrahedron Lett 61 , 151889 (2020).21 . M. J. Henley, A. N. Koehler, Advances in targeting ‘undruggable’ transcription factors with small molecules. Nat Rev Drug Discov 20, 669-688 (2021).22. I. Muckenschnabel, R. Falchetto, L. M. Mayr, I. Filipuzzi, SpeedScreen: label-free liquid chromatography-mass spectrometry-based high-throughput screening for the discovery of orphan protein ligands. Anal Biochem 324, 241-249 (2004).23. A. Annis, C. C. Chuang, N. Nazef, “ALIS: An affinity selection-mass spectrometry system for the discovery and characterization of protein-ligand interactions” in Mass Spectrometry in Medicinal Chemistry (John Wiley & Sons, Ltd, 2007)vol. 36, pp. 121-156.24. P. M. Sabale, M. Imiolek, P. Raia, S. Barluenga, N. Winssinger, Suprastapled Peptides: Hybridization-Enhanced Peptide Ligation and Enforced a-Helical Conformation for Affinity Selection of Combinatorial Libraries. J Am Chem Soc 143, 18932-18940 (2021).25. R. N. Muchiri, R. B. Breemen, Affinity selection-mass spectrometry for the discovery of pharmacologically active compounds from combinatorial libraries and natural products. Journal of Mass Spectrometry 56 (2021).26. S. Pomplun, Z. P. Gates, G. Zhang, A. J. Quartararo, B. L. Pentelute, Discovery of Nucleic Acid Binding Molecules from Combinatorial Biohybrid Nucleobase Peptide Libraries. J Am Chem Soc 142, 19642-19651 (2020).27. A. Vinogradov, Z. P. Gates, C. Zhang, A. J. Quartararo, K. H. Halloran, B. L. Pentelute, Library Design-Facilitated High-Throughput Sequencing of Synthetic Peptide Libraries. ACS Comb Sci 19, 694-701 (2017).28. S. Pomplun, M. Jbara, A. J. Quartararo, G. Zhang, J. S. Brown, Y. C. Lee, X. Ye, S. Hanna, B. L. Pentelute, De Novo Discovery of High-Affinity Peptide Binders for the SARS-CoV- 2 Spike Protein. ACS Cent Sci 7, 156-163 (2021).29. A. J. Quartararo, Z. P. Gates, B. A. Somsen, N. Hartrampf, X. Ye, A. Shimada, Y. Kajihara, C. Ottmann, B. L. Pentelute, Ultra-large chemical libraries for the discovery of high- affinity peptide binders. Nat Commun 11 , 3183 (2020).30. Z. P. Gates, A. A. Vinogradov, A. J. Quartararo, A. Bandyopadhyay, Z.-N. Choo, E. D. Evans, K. H. Halloran, A. J. Mijalis, S. K. Mong, M. D. Simon, E. A. Standley, E. D. Styduhar, S. Z. Tasker, F. Touti, J. M. Weber, J. L. Wilson, T. F. Jamison, B. L. Pentelute, Xenoprotein engineering via synthetic libraries. Proceedings of the National Academy of Sciences 115, E5298-E5306 (2018).31. D. Vourloumis, M. Takahashi, K. B. Simonsen, B. K. Ayida, S. Barluenga, G. C. Winters, T. Hermann, Solid-phase synthesis of benzimidazole libraries biased for RNA targets. Tetrahedron Lett 44, 2807-2811 (2003).32. J. P. Mayer, G. S. Lewis, C. McGee, D. Bankaitis-Davis, Solid-phase synthesis of benzimidazoles. Tetrahedron Lett 39, 6655-6658 (1998).33. D. Tumelty, M. K. Schwarz, K. Cao, M. C. Needels, Solid-phase synthesis of substituted benzimidazoles. Tetrahedron Lett 40, 6185-6188 (1999).34. J. W. Guiles, S. G. Johnson, W. V Murray, Solid-phase Suzuki coupling for C-C bond formation. Journal of Organic Chemistry 61 , 5169-5171 (1996).35. K. Diihrkop, M. Fleischauer, M. Ludwig, A. A. Aksenov, A. V. Melnik, M. Meusel, P. C. Dorrestein, J. Rousu, S. Bbcker, SIRIUS 4: a rapid tool for turning tandem mass spectra into metabolite structure information. Nat Methods 16, 299-302 (2019).36. K. Diihrkop, H. Shen, M. Meusel, J. Rousu, S. Bbcker, Searching molecular structure databases with tandem mass spectra using CSLFingerlD. Proc Natl Acad Sci U S A 112, 12580-12585 (2015).37. E. L. Schymanski, C. Ruttkies, M. Krauss, C. Brouard, T. Kind, K. Diihrkop, F. Allen, A. Vaniya, D. Verdegem, S. Bbcker, J. Rousu, H. Shen, H. Tsugawa, T. Sajed, O. Fiehn, B.Ghesquiere, S. Neumann, Critical Assessment of Small Molecule Identification 2016: automated methods. J Cheminform 9, 1-21 (2017).38. H. M. Becker, Carbonic anhydrase IX and acid transport in cancer. British Journal of Cancer 2019 122:2 122, 157-167 (2019).39. C. T. Supuran, J.-Y. Winum, Carbonic anhydrase IX inhibitors in cancer therapy: an update. Future Med Chem 7, 1407-1414 (2015).40. T. Clackson, J. A. Wells, In vitro selection from protein and peptide libraries. Trends Biotechnol 12, 173-184 (1994).41. B. van de Kooij, A. Schreuder, R. Pavani, V. Garzero, S. Uruci, T. J. Wendel, A. van Hoeck, M. San Martin Alonso, M. Everts, D. Koerse, E. Callen, J. Boom, H. Mei, E. Cuppen, M.S. Luijsterburg, M. A. T. M. van Vugt, A. Nussenzweig, H. van Attikum, S. M. Noordermeer, EXO1 protects BRCA1 -deficient cells against toxic DNA lesions. Mol Cell 84, 659-674. e7 (2024).42. A. L. Satz, What Do You Get from DNA-Encoded Libraries? ACS Med Chem Lett 9, 408-410 (2018).43. A. L. Satz, R. Hochstrasser, A. C. Petersen, Analysis of Current DNA Encoded Library Screening Data Indicates Higher False Negative Rates for Numerically Larger Libraries. ACS Comb Sci 19, 234-238 (2017).
Claims
CLAIMS1 . A method for identifying which members of a barcode-free library of small molecules bind to a target and elucidating their molecular structure, the method comprising: a) mixing the members of at least one subset of the library with the target; b) separating the members that do not bind to the target (nonbinders) from the members that bind to the target (binders); c) analysing the binders using tandem mass spectrometry (MS / MS) to generate MS / MS spectra of the binders; and d) analysing the MS / MS spectra by computational analysis to elucidate the molecular structure of the binders.
2. The method of claim 1 , wherein the at least one subset comprises at least 10,000 members.
3. The method of claim 2, wherein the at least one subset contains at least 50,000 members.
4. The method of any preceding claim, wherein the at least one subset comprises no more than 5,000,000 members.
5. The method of any preceding claim, wherein step c) comprises analysing the binders using liquid chromatography-tandem mass spectrometry (LC-MS / MS) to generate MS / MS spectra.
6. The method of any preceding claim, wherein step b) further comprises unbinding the binders from the target after the nonbinders have been separated from the binders.
7. The method of any preceding claim, wherein the library is a combinatorial library.
8. The method of claim 7, wherein the method further comprises synthesising the library via combinatorial synthesis.
9. The method of any preceding claim, wherein the library is enumerated.
10. The method of any preceding claim, wherein the small molecules have a molecular mass below 1000 gmol-1.11 . The method of any preceding claim, wherein step a) comprises mixing the members of the entire library with the target.
12. The method of any preceding claim, wherein the members are fragmentable by collision- induced dissociation.
13. The method of any preceding claim, wherein the members comprise an amino acid residue.
14. The method of any preceding claim, wherein the computational analysis determines how closely the members of the at least one subset of the library fit the generated MS / MS spectra.
15. The method of any preceding claim, wherein the computational analysis scores and ranks the members of the at least one subset of the library based on how closely they fit the generated MS / MS spectra.
16. The method of any preceding claim, wherein the computational analysis filters the members of the at least one subset of the library and disregards members that do not fit the MS / MS spectra generated in step c).
17. The method of any preceding claim, wherein the computational analysis requires that the binder must have an MS1 m / z that matches that of a library member.
18. The method of any preceding claim, wherein the computational analysis uses statistics, combinatorics, rule-based predictions, or Quantum Chemistry computations to elucidate the molecular structure of the binders, or wherein the computational analysis is based on a machine learning model.
19. The method of any preceding claim, wherein the method further comprises predicting MS / MS spectra of each member of the at least one subset of the library and the computational analysis compares the MS / MS spectra generated in step c) to the predicted spectra, optionally wherein the MS / MS spectra is predicted via fragmentation rules or Quantum Chemistry simulations or is predicted based on a machine learning model.
20. The method of claim 19, wherein the MS / MS spectra of each member of the at least one subset of the library is predicted based on MS / MS spectra of at least a proportion of members of the at least one subset of the library.21 . The method of claim 19 or claim 20, wherein the computational analysis requires that the binder must show at least one predicted fragment peak in its MS2.
22. The method of any preceding claim, wherein computational analysis predicts the presence or absence of substructures of the binders using rule-based considerations, a machine learning model, statistics or combinatorial optimization based on the MS / MS spectra generated in step c).
23. The method of any claim 22, wherein the library is a combinatorial library, and the substructures are particular building blocks from the combinatorial library.
24. The method of claim 22 or claim 23, wherein substructures are molecular fingerprints.
25. The method of any preceding claim wherein the target is an enzyme.
26. The method of any preceding claim, wherein the target is immobilised.
27. The method of claim 26, wherein the target is immobilised on a solid support.
28. The method of claim 27, wherein the solid support is magnetic beads.
29. The method of claim 27 or claim 28, wherein the target is biotinylated and immobilised on a streptavidin coated solid support.
30. The method of any one of claims 26 to 29, wherein step b) comprises washing the immobilised target to remove any nonbinders.31 . The method of any preceding claim, wherein step a) comprises mixing the members of the at least one subset of the library with the target in solution, wherein the concentration of the members in the solution is greater than 10 fmol / member.
32. The method of claim 31 , wherein the concentration of the members in solution is from 100 fmol / member to 1 pmol / member.
33. The method of any preceding claim, wherein the computational analysis is implemented by software.