Barcode Free Library Screen
Patent Information
- Application Number
- NL2039004
- Authority / Receiving Office
- NL · NL
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-11-05
- Publication Date
- 2026-06-08
- Estimated Expiration
- 2044-11-04
Smart Images

Figure 00000046_0000 
Figure 00000052_0000 
Figure 00000053_0000
Abstract
Description
[0001] This invention relates to an affinity selectionbased hit discovery. BACKGROUND
[0002] The discovery of high affinity ligands is crucial for virtuallyany drug discoverycampaign. The pharmaceutical industry relies on vast collections of individualcompounds (typically 0.5 4 million) and high throughput screening (HTS) facilities in order to identify novel ligands for drug targets (1). HTS have delivered several starting points fortherapeutic compounds, but these libraries cost up to a few billion US dollars and the required infrastructure is large and complex, limiting the availability of such massive platforms mainly to big pharma and fewacademic settings.
[0003] Affinity selection technology enables the screening of large libraries in a single experiment (26).These libraries are typicallypanned against immobilized target proteins, allowing for the separation of binders from nonbinders. A crucial step in this process is the decoding of each hit, which is mostcommonly achieved usingDNA orRNA barcodes attached to each ligand. Display technologies, such as phage display ormRNA display, utilize the natural translation machinery to convert genetic code into peptidic compounds, linked to theirencoding oligonucleotides (79). DNAencoded libraries (DELs) feature small molecules linked to unique DNAsequences, enabling the exploration of a druglike chemical space in the affinity selection setting (1017).
[0004] While the DNA barcode is an essentialcomponent for hit decoding, it also represents the primary limitation in DEL technology in terms of information stability, information densityand synthesis complexity. In the library preparation, chemical reactions must be alternated with enzymatic ligation steps, and all transformations need to be both water and DNAcompatible (18). Many standard chemical reactions involve conditions that degrade DNA and could compromise the chemical barcode (19). Suitable compatible reaction conditions have to be optimized and validated (20). Furthermore, the DNA tag, which is typically > 50times largerthan the small molecule, can significantly affect the selection process, by restricting the binding pose diversity of each librarymember or by potentially interactingwith the target and leading to false positives. This limitation becomes particularly problematicwhen the target protein has nucleic acid binding sites, making the screening for ligands of crucial drug targets like transcription factors or RNAbinding proteins very challenging (21).
[0005] Consequently, barcodefree selections appear as highly desirable technology. However, current approaches that relyon mass spectrometry (MS) to identify selected compounds from tag free libraries can only process a few hundred to a fewthousand compounds per sample. This limitation arises from the absence of automated de novo decoding software capable of distinguishing and annotating isobaric compounds (5, 2225). Larger library sizes have been achieved only for peptidic compounds, which, compared to small molecules have unfavourable druglike parameters (2, 3, 2630).
[0006] It is an aim of certain embodiments of the invention to provide a method for elucidating the molecular structure ofmembers of a library of small molecules that bind to a target in an affinity selection. The method may be particularly suited to elucidating the molecular structure of members of a library comprising small molecules that comprises over l04 members. BRIEFSUMMARYOFTHE DISCLOSURE
[0007] The first aspect of the invention provides a method for identifyingwhich members of a library comprising small molecules bind to a target and elucidating their molecular structure, the method comprising: a) mixing the members of at least one subset of the librarywith the target; b) separating the members that do not bind to the target (nonbinders) from the members that bind to the target (binders); c) analysing the binders usingtandem mass spectrometry (MS / MS) to generate MS / MS spectra; and d) analysing the MS / MS spectra bycomputational analysis to elucidate the molecular structure of the binders. The library
[0008] Itmay be that step a) comprises mixing a subset of the librarywith the target. Thiswould be more typicalwhere the librarycomprises a very large number ofmembers (e.g. greaterthan 1,000,000 or greater than 5,000,000). lt maybe that steps a) to d) are carried out on a first subset of the libraryand then steps a) to d) are carried out on a second subset of the library. This may be continued on further subsets of the library until the entire library has been subjected to steps a) to d). The subset of the library will always comprise a plurality of members, e.g. greater than 100 members or greaterthan 1000 members.
[0009] Itmay be that step a) comprises mixing the entire librarywith the target, i.e. the entire library is subjected to steps a) to d) in a single sitting. The librarymaycomprise greater than 30,000 members, e.g. greaterthan 50,000 members.
[0010] The librarymay be a combinatorial library. Thus, the method may further comprise synthesisingthe library via combinatorial synthesis.
[0011] The librarymay be enumerated.
[0012] Typically, the library is a barcode free library.
[0013] Itmay be that the at least one subset of the librarycomprises at least 10,000 members. It may be that the at least one subset of the library contains at least 100,000members. Itmay be that the at least one subset of the library contains at least 250,000members. Itmay be that each subset of the librarycomprises no more than 5,000,000 members. Itmay be that each subset of the librarycomprises no more than 1,000,000 members. Itmay be that each subset of the library comprises no more than 750,000 members.
[0014] Typically, each member of the library will be fragmentable by collisioninduced dissociation. Itmay be that at leastsome members of the librarycomprise at least one amide bond, CN bond orCO bond. Itmay be that allmembers of the librarycomprise at least one amide bond, CN bond orCO bond. Itmay be that at leastsome members of the library comprise at least one amide bond. Itmay be that allmembers of the librarycomprise at least one amide bond. It may be that at leastsome members of the librarycomprise an amino acid residue. Itmay be that allmembers of the librarycomprise an amino acid residue.
[0015] Itmay be that at least one subset of the library comprises members having a moiety selected from: O A $A / A\NH2 A ; il H A \N A\NH2 0 ; and / or Ar1 â Ar2 / \H / \NHZ 0 , wherein: AA is independently at each occurrence selected from an amino acid residue; and Ar1 and Ar2 are independently selected from an optionally substituted aryl.
[0016] At leastsome members of the library are small molecules. Itmay be that allmembers of the library are small molecules.
[0017] Itmay be that at leastsome members of the library are peptides.
[0018] Itmay be that at leastsome members of the library are oligonucleotides.
[0019] Itmay be that at leastsome members of the libraryobey Lipinskis rule of 5. Itmay be that allmembers of the libraryobey Lipinskis rule of 5. Itmay be that at leastsome members of the librarydo not obey Lipinskis rule of 5. Itmay be that allmembers of the librarydo not obey Lipinskis rule of 5.
[0020] The librarymaycomprise a range of structural motifs. This might be the case, for example, if the method is being used for target identification. Alternatively, the librarymay comprise a range ofcompounds that share a structure motif, e.g. a particular chemical core. It may be that the librarycomprises a series of analogues of a compound thatwas found to bind with the target in a previous iteration of the method of the first aspect. Thus, the method may be used in hittolead optimisation. The target
[0021] Typically, the target will any biological target that has the potential to bind a small molecule. The targetmay be a protein. The targetmaybe an enzyme. The targetmay however be a nonbiological target, e.g. a polymer, or other material.
[0022] Itmay be that the target is immobilised. Itmay be that the target is immobilised on a solid support. The solid supportmay be magnetic beads. Itmay be that the target is biotinylated and immobilised on a streptavidin coated solid support.
[0023] The targetmay be a carbonic anhydrase, e.g. carbonic anhydrase IX (CAIX). The target may be an exonuclease, e.g. Exonuclease1 (EXO1). Step a)
[0024] Typically, step a) comprises mixing allmembers of the at least one subset (e.g. all members of the library) with the target in a single vessel.
[0025] Typically, step a) comprises mixing the members of the at least one subset of the library with the target in solution. Itmay be that the members of the at least one subset of the library are in a first solution, the target is in the vessel and the solution ofmembers of the at least one subset of the library is added to the vessel. The vesselmay also contain a solvent, in which the target is dissolved or suspended. The solvent in which the target is dissolved or suspended and the solvent in which the at least one subset of the library is dissolved may be thesame solvent. The solvent in which the target is dissolved or suspended and the solvent in which the at least one subset of the library is dissolved may be a different solvent.
[0026] Step a) maycomprise mixing the members of the at least one subset of the librarywith the target in solution, wherein the concentration of the members in the solution is greater than 10 fmol / member. The concentration of the members in solution may be greaterthan 100 fmol / member. The concentration of the members in solution may be from 100 fmol / member to 1 pmol / member. Step b)
[0027] Typically, the target is immobilised. Where this is the case, step b) maycomprisewashing the immobilised target to remove any nonbinders.
[0028] Where the target is not immobilised, step b) maycomprise separating the nonbinders from the targetbinder pairs using size exclusion chromatography.
[0029] Step b) may further comprise unbinding the binders from the target after the nonbinders have been separated from the binders, e.g. after the nonbinders have beenwashed awayfrom the immobilised target. The binders may be unbound through elution. Step c)
[0030] Step c) maycomprise analysingthe binders using chromatographictandem mass spectrometry to generate MS / MS spectra. Step c) may comprise analysing the binders using liquid chromatographytandem mass spectrometry (LCMS / MS) to generate MS / MS spectra. Step cl)
[0031] Typically, step d) is automated.
[0032] The computational analysis may elucidate the structures de novo. More typically, however, the structures (e.g. SMILES) of the at least one subset of the (enumerated) library will be accessible and the computational analysis elucidates the structures by fitting the generated spectra to the best matches in the at least one subset of the enumerated library. The computational analysis may require that the structure elucidated in step d) matches the structure of a member in the enumerated library.
[0033] Itmay be that the computational analysis determines how closelythe members of the at least one subset of the library fit the generated MS / MS spectra. In this embodiment, the librarywill typically be an enumerated library. The computational analysis may score and rankthe members of the at least one subset of the library based on how closely they fit the generated MS / MS spectra. Itmay be that the computational analysis filters the members of the at least one subset of the libraryand disregards members that do not fit the MS / MS spectra generated in step c). By disregard it is meant that the computational analysis will not determine the molecular structure of the binders to be that of a member that does not fit the MS / MS spectra generated in step c).
[0034] Itmay be that the computational analysis requires that the binder must have an MS1 m / z that matches that of a librarymember.
[0035] Itmay be that computational analysis uses statistics, combinatorics, rulebased predictions, orQuantum Chemistry computations. The computational analysis may be based on a machine learning model. The machine learningmodelmay predictfragment peaks based on MS / MS spectra of at least a proportion ofmembers of the at least one subset of the library. The computational analysis maycompute fragmentation trees to interpret the MS / MS spectra.
[0036] The method may further comprise predicting MS / MS spectra of each member of the at least one subset of the libraryand the computational analysis compares MS / MS spectra generated in step c) to the predicted spectra. The MS / MS spectra may be predicted via fragmentation rules, machine learning orQuantum Chemistry simulations. The MS / MS spectra of each member of the at least one subset of the librarymay be predicted based on MS / MS spectra of at least a proportion (e.g. a proportion) ofmembers of the at least one subset of the library. The method may therefore furthercomprise obtainingMS / MS spectra of at least a proportion of at least a proportion (e.g. a proportion) ofmembers of the at least one subset of the library. The computational analysis may require that the bindermust show at least one predicted fragment peak in their MS2.
[0037] |tmay be that the computational analysis predicts the presence or absence of substructures of the binders via rulebased considerations, a machine learningmodel, statistics or combinatorialoptimization based on the MS / MS spectra generated in step c). When the library is a combinatorial library, the substructures may be particular building blocks from the combinatorial library. |tmay be that the substructures are molecular fingerprints. Themethod
[0038] The method maycomprise steps a1) and b1): a1) mixing thesame members of the at least one subset of the library with thesame components of the reaction of step a) butwithout the target present; b1 ) separating the members that do not bind nonspecifically to the components (nonbinders) from the members that bind to the components (nonspecific binders); wherein step c) comprises subtractingthe MS / MS spectra of the nonspecific binders from the data set of MS / MS spectra of the binders.
[0039] Where the target is immobilised, e.g. on magnetic beads, the components used in step a1) will typicallycomprise the immobilisation means, e.g. the magnetic beads.
[0040] The method maycomprise steps c2) and d2): c2) analysing the nonbinders obtained in step b) usingtandem mass spectrometry (MS / MS) to generate MS / MS spectra; and d2) analysing the MS / MS spectra usingcomputational analysis to elucidate the molecular structure of the nonbinders, wherein step d2) is automated. BRIEF DESCRIPTION OFTHEDRAWINGS
[0041] Embodiments of the invention are further described hereinafterwith reference to the accompanying drawings, in which: Figure 1 outlines the hit discovery with SELs. A) SELworkflow: large libraries with diverse scaffolds are prepared via solid phase split and pool synthesis with a wide compatibility range of chemical transformation. The libraries are then released from beads and screened in solution against immobilized targets. Binders are enriched from the rest of the libraryand analyzed via nano LCMS / MS. Hit compounds are identified and decoded bySIRIUSCOM ET (COmbinatorial Mass Encoding decoding Tool) software. B) Summary of SEL features making it an excellent platform for rapid and straightfonNard early drug discovery. Figure 2shows the amino acid building blocks used to synthesize SELs 1, 2 and 3. Fmoc protected variants of the amino acids were used in the library synthesis. Building blocks marked with a * are not compatible with SEL 3 and were excluded for that design. Protectinggroups indicated in red are removed during cleavage. Figure 3shows carboxylic acid building blocks used to synthesize SEL 1. Protecting groups indicated in red are removed during cleavage. Figure 4shows the primaryamine building blocks used to synthesize SEL 2. Figure 5shows the aldehyde building blocks used to synthesize SEL 2. Figure 6shows the aryl bromide building blocks used to synthesize SEL 3. Figure 7shows the boronic acid building blocks used to synthesize SEL 3. Figure 8 shows the two consecutive filtering steps ofCOMET for removingMS / MS spectra that do not match to a librarycompound. Figure 9 outlines the library design and synthesis. A) SEL1 is prepared from two amino acid building blocks decorated with a carboxylic acid decorator. Standard solid phase amide bond formation protocols result in products with excellent crude purity. B) SEL 2, resulting in a benzimidazole scaffold with three decorators is synthesized in a 7step procedure. Optimization of all steps and reaction scope forthe nucleophilic aromatic substitution and the heterocyclization are reported in detail in Fig. 1012 and summarized in the histograms. C) SEL 3 is prepared from an amino acid, and a SuzukiMiyaura reaction between an aryl bromide and a boronic acid. Cross coupling optimization and reaction scope for the aryl bromides and boronic acids are reported in detail in Table 3 and Fig. 13 and 14 and summarized in the histograms. D) Based on the building blockscope for all scaffolds (crude purity > 55%) and druglike properties of the building blocks, three high diversity libraries were prepared. Selected examples from each libraryemphasize the chemical diversity of the generated structures. The violin plots showthe compliance of our libraries with several parameters important for drug development. Figure 10 shows the optimization of trisubstituted benzimidazole synthesis on solid support. A) Optimization of the nucleophilic aromatic substitution reaction between 4fluoro3 nitrobenzoic acid and the primaryamines aniline and pentan3amine B) Optimization of the reduction of nitro group. C) Optimization of the cyclization between compound 5 and isovaleraldehyde. Figure 11 shows the reaction scope for nucleophilic aromatic substitutions in the benzimidazole scaffold (primary amines). a) Reaction conditions: Compound 8 (5 pmol, 1.0 eq.), amine (50 pmol, 10 eq.), DIPEA (50 pmol, 10 eq.) in DMF (100 pL), 80°C, overnight. b) Conversion was determined by LCMS analysis. Figure 12 shows the reaction scope for heterocyclization to the benzimidazole scaffold (aldehydes) Figure 13 shows the reaction scope of different aryl bromides in the SuzukiMiyaura reaction. a) Reaction conditions: Compound 1 (5 pmol,1.0 eq.), phenylboronic acid (10 pmol, 2 eq.), PdCl2 (10 mol%), XPhos (20 mol%), K2C03 (10 pmol, 2 eq.), 9:1 DMF / H20 (100 pL), 80 °C on. b) Conversion was determined by LCMS analysis. Figure 14 shows the reaction scope of different boronic acids / esters in the SuzukiMiyaura reaction. a) Reaction conditions: Compound 1 (5 pmol,1.0 eq.), phenylboronic acid (10 pmol, 2 eq.), PdCl2 (10 mol%), XPhos (20 mol%), K2C03 (10 pmol, 2 eq.), 9:1 DMF / H20 (100 pL), 80 °C on. b) Conversion was determined by LCMS analysis CtBu removed with TFA. Figure 15 compares libraries with selected building blocks having optimized druglike properties to libraries with randomlychosen building blocks. Left: before selection (1000*1000*1000=1 billion membered library). Right = After selection (54*54*44 =128.304 membered library). Figure 16: Decodingwith S|R|US:COMET. A) The scatter plotshows the mass degeneracy of each library design. For specific masses there can be up to 500 possible differentcompounds. Tandem mass spectrometry is essential to enable the unequivocal annotation of unique structures. B) Foreach library scaffold a small sublibrarywith < 500 uniquecompoundswas synthesized and measured by LCMS / MS to mimic the outcome of an affinity selection outcome. C) Given that in these libraries allcompounds are knownwe performed semimanual analysis to obtain the actual number of detectable compounds. Given that each nanoLCMS / MS run produced ~ 80,000 spectra, such manual analysiswould be unfeasible for samples with unknown content. Automated analysis with off the shelf SIRIUS 6 resulted in good recalls of librarycompounds, but with some false positives (annotated compounds that are not actually present in the library). D) Based on the manually annotated test libraries, fragmentation frequencies of each library scaffold were calculated and represented as histograms. E,F) The SIRIUSCOMET workflow consist of generating a fragments list based on the calculated fragmentation frequencies, filtering MS / MS scans based on the minimum number of matching peaks and structure annotation by SIRIUS 6. The SIRIUSCOM ETworkflowdecreased falselyannotated structures while retaining high recall rates of correctcompounds. Figure 17 depicts the chemical structures of hit 20 and 21, with the extracted ion chromatogram (EIC) and LCMS / MS chromatogram from the affinity selection ofSEL1 against CAIX. Figure 18 depicts the chemical structures of hit 24 and 25, with the extracted ion chromatogram (EIC) and LCMS / MS chromatogram from the affinity selection of SEL 2 against CAIX. Figure 19 depicts the chemical structures of hit 35, with the extracted ion chromatogram (EIC) and LCMS / MS chromatogram from the affinity selection of SEL 3 against CAIX. Figure 20 outlines the Identification and validation of ligands against CAIXand EXO1. A) SEL 1,2 and 3were panned against immobilized CAIX. After analysis of the selected samples via nanoLCMS / MS and SIRIUSCOMET, we plotted building block frequencies for each library scaffold and position. Aromatic sulfonamideswere strongly enriched from all libraries. B) Examples of hit compounds: the chromatograms show extracted ion counts of their corresponding masses in CAIX sample and bead only control, demonstrating their selective enrichment. Biotinylated version of these examples were resynthesized and BLI association / dissociation curves confirm there binding to CAIX. C) In the affinity selection against EXO1 onlyone clear hitwas identified. Based on an isobaric building block pair, two possible structures could match the annotation. BLI showed that onlycompound 6 binds to EXO1. Figure 21 depicts the chemical structures of isomers 43 and 44, with the extracted ion chromatogram (EIC) and LCMS / MS chromatogram from the affinity selection ofSEL1 against EXO1. Figure 22 shows the building block analysis after affinity selection with SEL1 against EXO1. DETAILED DESCRIPTION Definitions
[0042] The term member, as in a member of a library of small molecules, refers to molecules with a defined chemical structure. A library comprising n members therefore comprises molecules having n different structures. The members may all be small molecules. |tmay be that a portion of the library of small molecules are not small molecules, providing that at leastsome of the members of the library are small molecules.
[0043] A small molecule may be considered to be a molecule (or salt thereof) having a molecularmass below 1500 gmol'1. |tmay be that a small molecule has a molecularmass below 1250 gmol'1. |tmay be that a small molecule has a molecular mass below 1000 gmol'1. A small molecule may be a chemical (e.g. organic) molecule (or salt thereof) comprising less than 100 atoms.A small molecule may be a chemical (e.g. organic) molecule (or salt thereof) comprising less than 50 heavyatoms (heavyatoms beingatoms otherthan hydrogen). Small molecules may be organic, or theymay be organometallic. |tmay be that at leastsome of the small molecules are organic small molecules. |tmay be that all small molecules are organic small molecules.
[0044] A target is any molecule or material that has the potential to bind a small molecule. Typically, the target will be a biological target, i.e. any part of the living organisms to which a small molecule can bind in order to bring about a physiological change. Common biological targets are nuclear receptors, ion channels, Gprotein coupled receptors and enzymes.
[0045] An enumerated library is one ofwhich the chemical structure of everymember is known and catalogued.
[0046] A molecularfingerprint is a unique pattern or representation of a molecule's chemical structure and properties. A molecular fingerprint may provide information on the entire chemical structure of a molecule. A molecular fingerprintmay also provide information on part of the chemical structure of a molecule, e.g. a moiety an / orfragment of a chemical molecule.
[0047] An immobilised target is a target that is physically confined or localized in a certain defined region of space with retention of its biological activity, and which can be used repeatedly and continuously. For example, a targetmay be immobilised on a solid support.
[0048] The term solid support or solid phase means, unless otherwise stated, a material that is substantially insoluble in a selected solvent system, orwhich can be readily separated (e.g., by e.g., by filtration, centrifugation, or magnetic capture) from a selected solvent system in which it is soluble. In the present disclosure, solid supports that are substantially insoluble in a selected solvent system may be preferred. Such solid supports are not limited to a specific type of support, and a large number of such solid supports are available and are known to one of ordinary skill in the art. Exemplary solid supports include, but are not limited to, solid and semisolid matrixes, such as aerogels and hydrogels, resins, particles, beads (including magnetic beads, such as coated magnetic beads), biochips (including thin film coated biochips), microfluidic chip, a silicon chip, multiwell plates (also referred to as microtiter plates or microplates), membranes, conducting and nonconducting metals, glass (including microscope slides) and magnetic supports. Where the solid support comprises beads, monodisperse (magnetic or nonmagnetic) beads may be preferred, as monodisperse beads may provide more consistent performance in assays. A solid supportmaycomprise magnetic beads.
[0049] A solid supportmay be used to immobilise a biological target. Solid supports useful for this purpose include streptavidin coated magnetic beads, which can be used to immobilise biotinylated targets.
[0050] A solid supportmay be used to synthesise molecules on in combinatorial synthesis. Solid supports useful forthis purpose include resin beads functionalized with an NH2group, e.g. TentaGel®.
[0051] YO denotes a solid support.
[0052] An amino acid residue may be a residue of a proteinogenic or a nonproteinogenic amino acid. Proteinogenic amino acids include alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, selenocysteine or pyrrolysine. The amino acid may be the damino acid, e.g. dleucine, or the Bamino acid, eg. Balanine. Typically, the amino acid will be the Lamino acid. The amino acid may be a nonproteinogenic amino acid. An amino acid residue may be the residue of an amino acid selected from Figure 2.
[0053] For the absence of doubt, in the following moieties: i; o M A , A ArZ / ArTAWHz Ä N:©\)(A\NH2 A A NH2 A O , and O , the amino acid residue will be connected to the adjacent carbonylgroup via the nitrogen atom of its amine group, or a nitrogen atom of one of its amine groups (therebyforming a peptide bond) and will be connected to the adjacent amine group via its carbonyl carbon (therebyforming a peptide bond).
[0054] Arylgroupsmay be anyaromatic carbocyclic ringsystem (i.e. a ringsystem containing 2(2n + 1)Tt electrons). Aryl groups may have from 6 to 12 carbon atoms in the ring system. Aryl groups may be phenyl groups. Arylgroups may be naphthyl groups or biphenyl groups. The arylgroup may be unsubstituted or substituted.
[0055] In tandem mass spectrometry, also known as MS / MS or MS2, the molecules of a given sample are ionized and the first spectrometer (designated MS1 ) separates these ions bytheir mass tocharge ratio (often given as m / z or m / Q). Ions of a particular m / zratio coming from MS1 are selected and thenfragmented into smallerfragment ions, e.g. bycollisioninduced dissociation, ion molecule reaction, or photodissociation. These fragments are then introduced into the second mass spectrometer (MS2), which in turn separates the fragments by their m / zratio and detects them. The fragmentation step makes it possible to identify isobaric compounds (i.e., compounds having the same mass but different chemical structures), which is not possible in regular mass spectrometry. Additional fragmentation steps (multiplestage MS or MS) can obtain further molecular information. Tandem mass spectrometrycan be coupled with liquid chromatography, i.e. liquid chromatographytandem mass spectrometry (LCMS / MS).
[0056] Computational analysis is the use of computers to analyse data. The computational analysis required in step d) interprets the MS / MS spectra generated in step c) in order to elucidate the molecular structure of the binders. Itmay be that the computational analysis interprets the MS / MS spectra by converting this into molecular structure information and elucidates the molecular structure the binders based on this information.
[0057] In the method of the present invention, step d is typicallyautomated, meaning that the MS / MS spectra generated in step c) are analysed automaticallywithouthuman input in order to elucidate the molecular structure of the binders.
[0058] The computational analysis may be implemented by software. The software may include a SIRIUS software, e.g. SIRIUS 4, which is a Java software used to analyse molecules from tandem mass spectrometry data.
[0059] Fragmentation trees annotate the MS / MS spectra of a molecule via an assumed / predicted fragmentation process.
[0060] A combinatorial library is a library comprisingmembers that have been synthesised by combinatorial chemistry. Typically, members of a combinatorial library will share a common molecular scaffold(s), or core structure(s). Itmaybe that the library is divided into subsets. The members in each subset will typically share a common moiety and / or substituent (in addition to theircommon molecular scaffold).
[0061] Combinatorial chemistry involves the synthetic assembly of structural building blocks in various possible combinations to produce large libraries of small molecules.
[0062] A barcode library, or encoded chemical library, is a library comprisingmembers that are conjugated to a molecular tag that allows the molecular structure of the member to be determined. One such library is a DNAencoded chemical library (DECL, or DEL) in which members are conjugated to a DNAfragment that serve as identification barcodes.
[0063] SMILES is the Simplified Molecular Input Line EntrySystem, which is used to translate a molecules threedimensional structure into a string of symbols that is more easily understood by computer software.
[0064] The abbreviations used herein have the following meaning: BB Building block BLI Biolayer interferometry CAIX Carbonic anhydrase IX DCM Dichloromethane DEL DNA encoded library DIPEA Diisopropylethylamine DMF Dimethylformamide DMSO Dimethylsulfoxide Eq Equivalent EXO1 Exonuclease1 FA Formic acid FBS Fetal bovine serum Fmoc Fluorenylmethoxycarbonyl Hexafluorophosphate Azabenzotriazole Tetramethyl HATU Uronium HBA Hydrogen bond acceptor HBD Hydrogen bond donor HEPES (4(2hydroxyethyl)1piperazineethanesulfonic acid) Me Methyl MeCN Acetonitrile MeOH Methanol MW Molecularweight MQ MilliQ NMR Nuclear magnetic resonance PBS Phosphatebuffered saline PEL Peptide encoded library Ppm Parts per million r.t. Room temperature SA Streptavidin SEL Selfencoded library tBu Tertbutyl TFA Trifluoroacetic acid TIPS Triisopropylsilane TLC Thin layer chromatography TPSA Topological polar surface area Tween20 Polysorbate 20
[0065] Throughout the description and claims of this specification, the words comprise and contain and variations ofthem mean including but not limited to, and they are not intended to (and do not) exclude other moieties, additives, components, integers or steps. Throughout the description and claims of this specification, the singularencompasses the plural unless the context otherwise requires. In particular,where the indefinite article is used, the specification is to be understood as contemplating plurality as well as singularity, unless the context requires othenNise.
[0066] Features, integers, characteristics, compounds, chemical moieties orgroups described in conjunction with a particular aspect, embodiment orexample of the invention are to be understood to be applicable to any other aspect, embodiment or example described herein unless incompatible therewith. All of the features disclosed in this specification (including any accompanying claims, abstract and drawings), and / or all of the steps of anymethod or process so disclosed, may be combined in any combination, except combinations where at leastsome of such features and / or steps are mutually exclusive. The invention is not restricted to the details of anyforegoing embodiments. The invention extends to any novel one, orany novel combination, of the features disclosed in this specification (including anyaccompanying claims, abstract and drawings), or to any novel one, or any novel combination, of the steps of anymethod or process so disclosed.
[0067] The reader's attention is directed to all papers and documents which are filed concurrentlywith or previous to this specification in connection with this application and which are open to public inspection with this specification, and the contents of all such papers and documents are incorporated herein by reference. EXAMPLES
[0068] Here,we report an affinity selection platform that can screen barcode free small molecule libraries with 104106 members, all at once (Fig. 1). The approach features the combinatorial synthesis of druglike compounds on solid phase beads, allowing for a wide range of chemical transformations and circumventing the complexityand limitation of DEL preparation. Compounds are annotated using theirtandem MS fragmentation spectra, obviatingthe need for barcodingtags and at thesame time allowing for precise identification ofmany isobaric, but structurally diverse compounds. We showthe feasibility of ourdecoding strategyon a diverse set of chemical scaffolds, prepared by a variety of chemical transformations (including crosscouplings, heterocyclizations, amide formation, nucleophilic aromatic substitution, and more).We then performed de novo discovery selections, panning libraries with 30500k members against the two oncology targets CAIX and Exonuclease1 (EXO1), resulting in identification of nanomolar binders for both targets. Notably, EXO1 is a DNAprocessing enzyme, and therefore very challenging for DEL selections. Overall, our SEL platform presents ideal features in terms of information density, information stabilityand unbiased screening capabilities.
[0069] In the following sectionswe describe in detail the individual parts of the platform, including library design and synthesis, decoding software and methodologyand the de novo ligand selection campaigns. Materialand methods Reagents andSupplies
[0070] Chemicals: Reagents and solvents were purchased from SigmaAldrich (Merck), Fisher Scientific, BLDpharm, Fluorochem orVWR and were used without further purification unless stated otherwise. Carboxylic acids, aldehydes and Fmocprotected amino acids building blocks were partially purchased from Chemspace. Tentagel® S NH2 90pm (S30902) and TentaGel® M NH2 30 pm (M30352) were purchased from RappPolymere.
[0071] The amino acid building blocks used to synthesize SEL 1, 2 and 3 are shown in Fig. 2. The carboxylic acid building blocks used to synthesize SEL1 are shown in Fig. 3. The primaryamine building blocks used to synthesize SEL 2 are shown in Fig. 4. The aldehyde building blocks used to synthesize SEL 2 are shown in Fig. 5. The aryl bromide building blocks used to synthesize SEL 3 are shown in Fig. 6. The boronic acid building blocks used to synthesize SEL 3 are shown in Fig. 7.
[0072] Materials for affinity selection: Dynabeads MyOne Streptavidin T1 was purchased from ThermoFisher Scientific. A KingFisherT" Duo Prime Purification System was used to perform our affinity selection experiments. The protocols were developed with Bindlt 4.0 Software.
[0073] Proteins: Biotinylated Human CarbonicAnhydrase IX (38414), His,Avitag (CA9 H82E3) and Human Carbonic Anhydrase IX (38414), His,Avitag (CA9H5226) were purchased from ACROBiosystems. EXO1 was obtained from Sylvie Noordermeer (Department ofHuman Genetics, Leiden University Medical Center. Instrumentation
[0074] LC-MS analysis: Compound puritywas determined by LCMS, using the LCMS2020 system (Shimadzu) with a Gemini 3pm C18110Âcolumn (50 >< 3mm) using the following parameters: flow rate = 0.55 mL / min, scan range = 160 800 m / z, column temperature (°C) = 40. The followinggradient of 1090% MeCN / H20 (0.1% formic acid) over 15 min and measuringUV absorbance at 254nmwas used, unless stated otherwise. Compounds were dissolved in H20:MeCN:tBuOH (1 :1 :1) before injection.
[0075] Purification: The automatic flash chromatographywas performed on a Biotage Selekt System with prepacked flash cartridges (Biotage® Sfar Bio C18 Duo 300Â20 pm). MQ Millipore water 0.1%TFA (bufferA) and Acetonitrile 0.1%TFA (buffer B) were used as mobile phase.
[0076] LC-MS / MS analysis: Analysiswas performed on an VanquishT" Neo UHPLC system (Thermo Scientific) connected to an Orbitrap Exploris 240 mass spectrometer (Thermo Fisher Scientific). Samples were run on a Double nanoViper PepMap Neo column (2pm particle size, 15 cm x 75 pm,Thermo Fisher Scientific, DNV75150PN) following a PepMap Neo Trap Cartridge (5pm C18 300pm X 5 mm). The standard nanoLC method was run withouttemperature control and a flow rate of 300 nL / min with the following gradient:0% solvent B ramping linearly to 60% B inA over 78 min, with solventA = water (0.1% FA), and solvent B = 80% acetonitrile, 20% water (0.1% FA). Positive spray ionizationwas set at 1900 V. The following parameters were used for MS1 collection: resolution = 120.000, scan range = 275650 m / z forSEL 1, 300800 m / z forSEL 2 and 275700 for SEL 3, maximum injection times = 300 ms, RF Lens = 80%, microscans = 1,AGC target = standard, Ion TransferTubeTemp (°C) = 280, Charge State = +1. For the data dependent MS / MS event the following settings were used: resolution = 15000, number of Dependent Scans= 20, isolation window (m / z) = 1.5, intensitythreshold = 1E5, MIPS mode = small molecule, absolute collision energy = 15, 25 eV,AGC Target (%) = 50, microscans = 1, RF Lens(%) = 70, dynamic exclusion = on (exclude after n times = 3, Exclusion duration (s)=15, excluding isotopes, 10ppm mass tolerance). General Procedures
[0077] Manual solid-phase synthesis (SPS):
[0078] Rink linker: TentaGel S NH2 resin (90 pm, 0.26 mmol / g, 1.0 eq.)was functionalized with a Rink linker using the following protocol before attaching other building blocks.
[0079] TentaGel S NH2 resin (90 pm, 0.26 mmol / g,1.0 eq.)was loaded onto a fritted syringe and swelled for 5 min in DM F.A solution of Fmocprotected aminoacid (3.0 eq), HATU (0.4 M, 2.98 eq) and DIPEA (9.0 eq) in DMFwas added to the resin. After 20 min the resinwaswashed with DMF (5x) and Fmoc deprotection was performed bywashingwith piperidine (20% in DMF, 1x) before incubatingwith piperidine (20% in DMF) for 10 min. The resin waswashed with DMF (5x) before adding the next Fmocprotected aminoacid. The protocolwas repeated until completion of the synthesis.
[0080] Cleavage: The compounds were cleaved of the resin by incubating at room temperature for 1 hourwith a solution ofTFA:H20:T|S (95:2.5:2.5) and washed with TFA:H20:T|S (95:25:25). The volumewas reduced by evaporating theTFA solution with a N2 stream. The reaction mixturewas dissolved in H20:MeCN:tBuOH (1:1 :1) before injection into a LCMS.
[0081] Protein biotinylation: EZLinkT" SulfoNHSLCBiotin was used in accordance with the userguide from themofisher.
[0082] Affinity selection:A KingFisherT" Duo Prime Purification System was used to perform our affinity selection experiments. The protocols were developed with Bindlt 4.0 Software. Each experimentwas performed in duplicates with protein (150 pmol) immobilized on Dynabeads MyOne Streptavidin T1 (1 mg) and library (100 fmol / member) with the King Fisher protocol as described in Table 1 (below unless stated otherwise). Table 1:The KingFisher program used for affinity selection. ___ Buffer Volume Release Release Mixing Temp Mixing Collec Collec (pL) time (s) speed time speed t ttime (min) count (s) Bead 100 5 r.t. Bottom 5 10 .Il-III..- Bead A 1000 30 Mediu 3 r.t. Medium 5 10 washing m (3X) Protein 100 30 Mediu 10°C Medium 5 10 .IIIIIIIII Biotin C 1000 30 Mediu 5 r.t. Medium 5 10 blocking m (2X) Bead A 1000 30 Mediu 3 r.t. Medium 5 10 .IIIIIIIII Library E 1000 30 Mediu 10°C Medium 5 10 IIIIIIIIII Bead F 1000 30 Mediu 0.5 r.t. Medium 5 10 washing m (5X) Elution G 100 30 Mediu 3 r.t. Medium 5 10 IIIIIIIIII Table 2: Overview of the affinity selection buffers. A 10% FBS, 1x PBS, 0.02% 50mM HEPES, 100mM KCl, 5mM MgCl2,1mM DTF, III |__ C 10% FBS,1x PBS, 0.02% 50mM HEPES, 100mM KCl, 5mM MgCl2,1mM DTF, III 10% FBS, 1x PBS 50mM HEPES, 100mM KCl, 5mM MgCl2,1mM DTF, III E 100 fmol / member library in 100 fmol / member library in bufferD III F 1x PBS 50mM HEPES, 100mM KCl, 5mM MgCl2,1mM DTF, III
[0083] Biolayer interferometry (BLI): Purified biotinylated compoundswere dissolved to 1 pM in 1x PBS, 0.02%Tween20, 1 mg / ml BSA (0.1% (w / v)) (kinetic buffer) used for immobilization onto streptavidin OctetSA Biosensors (SATORIUS). Biolayer interferometry (BLI) assayswere performed in 96 well plates (GreinerBioOne, polypropylene, flat-bottom, chimneywell) using an Octet R4 system (SATORIUS). Wells were filled with 200 pLwith kinetic buffer, compound solution orCAIX solution.
[0084] Biotinylated compound was immobilized onto the streptavidin biosensor for60 s. Sensors were then dipped into kinetic buffer for60 s, CAIX solution (500 nM, 250 nM, 125 nM, 62.5 nM) for 600 s and into kinetic buffer for 600s. Measurements were carried out at 30 °C.
[0085] Sample preparation afterAS: Samples from the affinity selection procedurewere lyophilized and resuspended in 50 pLMQ 0.1% FA. The StageTips were prepared as described by Rappsilber et al. usingC18 materialfrom Empore SPE 47mm discs (66883U, Merck).43 The StageTipswere preconditioned with 200 pL MeOH, 200 pL of 0.1% (v / v) FA in MeCN and 200 pL of 0.1% (v / v) FA in MQ, respectively by centrifuging for 3 min at 300 rpm. The samples were then loaded on the StageTips and washed with 200 pL of0.1% (v / v) FA in MQ. Compoundswere eluted byadding 200 pL of0.1% (v / v) FA in MeCN:MQ (7:3) and subsequently lyophilized before resuspending in 10 pL 0.1% (v / v) FA in UPLCMS grade water. The samples were centrifuged for 5 min at 15.000 rpm. AftenNards, 9 pLwas transferred to a LCMS vial and 8 pLwas injected into the LCMS / MS system. Software
[0086] Proteowizard (MSconvert)44: Raw files (.raw) from the Orbitrap Exploris 240 mass spectrometerwere converted to.szL files using MSConvert selecting Peak Picking>MS levels 1 2. Resulting.szL files were analyzed by using SIRIUS 6.0.4SNAPSHOT software.
[0087] SIRIUS COMET (6.0.4SNAPSHOT): After importing the .szL files, theCOMET filterwas applied using the following settings: scaffold formula:C8H3N2O for the SEL 2 and left blank for libraries 1 and 3, MS1 mass accuracy (ppm) = 5 ppm, considered fragment types = SEL 1: S[0;1],S[1;2],0,2, SEL 2: S[0;2],S[1;2],0,1, SEL 3: S[1:2],0, minimum number of matching peaks = 1, number of considered peaks = 5, number of allowed hydrogen shifts = 1, MS2 mass accuracy (ppm) = 5, unless stated othenNise. The generated fragment list served as input for the building blocks field and was created using python scripts. Scans matchingCOMETwere subsequently annotated using CS|:Finger|D and ranked according to Epimetheus.
[0088] COmbinatorial Mass Encoding decodingTool (COMET): After analyzingthe finalsample from the affinity selection experiment via LCMS / MS, the obtained MS / MS spectra are annotated with potential hitcompounds. To improve this annotation step,we developed COMET (COmbinatorial Mass Encoding decoding Tool) for excluding those MS / MS spectra which are unrelated library molecules, either due to a nonmatching precursormass or a nonmatching fragmentation pattern. Therefore,COMETaims to decrease the proportion of false positives in the final set of proposed compounds.
[0089] COMET consists oftwo consecutive filtering steps for removing MS / MS spectra that do not match to a librarycompound (Fig. 8). The first filter considers onlythe precursor mass of each MS / MS spectrum and excludes all those spectra whose precursormass does not match to the mass of any librarycompound (i.e. all spectra forwhich no library candidate exists). In contrast, the second filtering step is based on our observation thatsome of the covalent bonds connecting the building blocks are cleaved more frequently during the collisioninduced dissociation leading to correspondingfragment peaks in the spectra (Fig. 16D). Therefore,we derived certain fragmentation rules for each library to generate the most prominent fragments by cleaving only these specific bonds. For newly synthesized libraries,COMET allows the specification ofnew fragmentation rules. If such fragmentation information isn't known beforehand, it will generate all possible building blockfragments by combinatorially cleaving each of those bonds. These fragmentation rules can now be used to check for each remaining MS / MS spectrum whether at least one corresponding candidate structure exists that can explain at least kfragment peaks in the spectrum by applying these fragmentation rules. However, if no candidate structure fulfils this condition, the spectrum will be removed.
[0090] KNIME 5.1.2 softwarewas used for libraryenumeration and molecular property calculations. The following extensions were installed RDkit Nodes version 4.7., CDKversion 1.5.6 and ChemAxon version 4.7.0 Building Block Scoring& Selection
[0091] Fmocamino acid and carboxylic acid BBs were selected based on their druglike properties (molecularweight (MW), logP, hydrogen bond donors (HBD), hydrogen bond acceptors (HBA), and total polar surface area (TPSA)), using certain python scripts.
[0092] Chemspace provided lists of their instock Fmocamino acids and carboxylic acids, containing 1 .7K and 6.3K building blocks (BBs), respectively. These lists were then refined to 1,000 BBs each by applying filters for price, availability, and compatibilitywith our library design. The refined lists were then used as input for the main_final_selection.py script, where the virtual library was enumerated based on the provided reaction SMARTS.
[0093] For each librarymember the molecular properties MW, HBD, HBA, logP and TPSAwere calculated. The scoring of each librarymemberwas determined by abidingthe following thresholds: MW<500, HBA<10, HBD<5, logP<5 and TPSA<140. Per matching properties itwould score one point, resulting in a maximum score of 5 per librarymember. The scores were grouped per BB and the ranked BBs were used to select the most druglike BBs for our library. Selected BBs can be seen in Section 10.1. List of building blocks. Optimization Studies forSEL2
[0094] Fig. 10 outlines the optimization of trisubstituted benzimidazole synthesis on a solid support.
[0095] General procedure nucleophilic aromatic substitution 1_H2N- FZC H o Elïîj. Û H o 02 NÄJLNHZ 2. TFA:MQ:T|PS NËJLNO OZN o = H o = 8 9
[0096] Starting material 9was synthesized according to general procedure: Manual solidphase synthesis (SPS). TentaGel S NH2 resin (90 pm, 0.24 mmol / g loading, 19 mg, 5 pmol, 1.0 eq.)was transferred to an eppendorf tube.Amine 9a-9cm (Fig. 11) (50 pmol, 10 eq.) and DIPEA (8.71 pL, 50 pmol, 10 eq.) in DMF (100 pL, 0.5 M) were added to the resin. The mixturewas shaken at 600 rpm overnight at 80 °C. The resin waswashed with DMF (5 x 2 mL) and DCM (5 x 2 mL) before cleaving with a solution ofTFA:H20:T|S (95:2.5:2.5) for 1h and washed once with a solution ofTFA:H20:T|S (95:2.5:2.5). LCMS sampleswere prepared byadding 10 pL of theTFA mixture with 90 pL of MeCN:MQ:tBuOH (1 :1 :1 ) The reaction mixturewas characterized by LCMS.
[0097] General procedure heterocyclization HN H o 1. p-TsOH Q H o HzNÛïl / NËLËO 2. TFA:MQ:TIPS ngNgLNHz 11 13
[0098] Starting material 11 was synthesized as described in General procedure nucleophilic aromatic substitution. A solution of 1.0 M SnCl2 (62.5 eq.) in DMFwas added to the resin (1.0 eq.). The mixturewas incubated at r.t. overnight, whereafter the resin waswashed 50% MilliQ (in DM F, 5x), with DMF (5x,) and DCM (5x). TentaGel S NH2 resin (90 pm, 0.24 mmol / g loading, 19 mg, 5 pmol, 1.0 eq.)was transferred to an eppendorf tube.A solution the appropriate aldehyde 13a-13co (Fig. 12) ((0.25 M, 25 pmol, 5 eq.) and pTsOH*H20 (4.76 mg, 25 pmol, 5 eq.) in DMF (103 pL)was added to the resin. The mixturewas incubated at r.t., overnight on 600 rpm. After incubation, the resin waswashed DMF (5x) and DCM (5x). The resin was incubated for 1 hourwith a solution of TFA:H20:TIS (95:2.5:2.5) and washed once with a solution ofTFA:H20:TIS (95:2.5:2.5). LCMS samples were prepared byadding 10 pL of theTFA mixture with 90 pL of MeCN:MQ:tBuOH (1:1 :1) The reaction mixturewas characterized by LCMS. Optimization Studies forSEL3
[0099] Table 3shows the optimization of the SuzukiMiyaura reaction on a solid support. Table 3 H O 1. Ph(BOH)2 H O û / / \VFN\Ë)I\NH2 2_ TFA:MQ:TIPS NÉJLN'O Br 0 - H o - 14 15 ___ ____ __ __ __- __ a) Reaction conditions: compound 14 (5 pmol,1.0 eq), phenylboronic acid (10 pmol, 2.0 eq), PdCl2 (0.5 pmol, 10 mol%), XPhos (0.1 pmol, 20 mol%), K2C03 (10 pmol, 2 eq), solvent 9:1 DMF / H20 (100 pL), 80 °C b) Conversion was determined by LCMS analysis.
[00100] General procedure Suzuki-Mivaura coupling H o 1. PhB(OH)2, PdCIgV o BFÛWVLNO_» .ÛYLNQ o s H 2. TFA:MQ:TIPS o 5 H 16 19
[00101] Starting material 16was synthesized according to general procedure: Manual solidphase synthesis (SPS). TentaGel S NH2 resin (90 pm, 0.24 mmol / g loading, 4.8 mg, 5 pmol,1.0 eq.)was transferred to an eppendorf tube. Boronic acid (10 pmol, 2 eq.), PdCl2 (0.09 mg, 0.5 pmol, 0.1 eq.), XPhos (0.48 mg, 1 pmol, 0.2 eq.) and K2C03(1.4 mg, 10 pmol, 2.0 eq.)were dissolved in DMF:H20 (9:1, 100 pL) and added to the resin. The reaction was stirred at 700 rpm on 80°C for 2 hours. The resin waswashed DMF (5x) and DCM (5x). The resinwas incubated for 1 hourwith a solution of TFA:H20:T|S (95:2.5:2.5) and washed once with a solution ofTFA:H20:T|S (95:2.5:2.5). LCMS samples were prepared byadding 10 pL of theTFA mixture with 90 pL of MeCN:MQ:tBuOH (1 :1 :1 ). The reaction mixturewas characterized by LCMS. LibrarySynthesis and Characterization SEL 1: COOH-AA-AA 1' FmocHlËlÿg / OH 1 .JîOH o EQLULE'ZEIH o Eûïuf'ïâîm ! o u ULNO>_ _ _ _ , \'er ULNO+» o Ir HHL L, H 2. DMF.P|per|d|ne :\J) o H 2. DMFzPIperIdIne \,) o
[00102] TentaGel S NH2 resin (90 pm, 0.24 mmol / g loading, 2.625 g, 630 pmol,1.0 eq.)was transferred to a 20mL fritted syringe and was subsequently swelled with DMF for 10 min.A solution of FmocRinkOH (1.019 g, 1.89 mmol, 3.0 eq), 0.4M HATU (1.877 mmol, 4.693 mL, 2.98 eq) and DIPEA (988 pL, 5670 pmol, 9.0 eq) in DMFwas added to the resin and reacted for 2h. The resin was washed with DMF (5x 2 mL) and 20% piperidine in DMF (1x, 2 mL) before incubatingwith 20% piperidine in DMF for 10 min.
[00103] The resin was divided over 62 x 2mL fritted syringes. Stock solutions ofamino acids were made with Fmocprotected amino acids (100 pmol) in 0.4M HATU (250 pL). A solution of Fmoc protected aminoacid (10 pmol, 3.0 eq, (75 pLfrom premade stock solution), DIPEA (15.68 pL, 90.00 pmol, 9.0 eq) and DMF (75 pL)were added to the resin (10 pmol). After2h the resin was combined in a 20mL fritted syringe and waswashed with DMF (5x 2 mL) and 20% piperidine in DMF (1x, 2 mL) before incubatingwith 20% piperidine in DMF for 10 min. The resin waswashed with DMF (5 x 2 mL) before addingthe next Fmocprotected aminoacid.
[00104] The resin was divided over 62 x 2mL fritted syringes. Stock solutions ofamino acids were made with Fmocprotected amino acids (100 pmol) in 0.4M HATU (250 pL). A solution of Fmoc protected aminoacid (10 pmol, 3.0 eq , (75 pLfrom premade stock solution), DIPEA (15.68 pL, 90.00 pmol, 9.0 eq) and DMF (75 pL) were added to the resin (10 pmol). After2h the resin was combined in a 20mL fritted syringe and waswashed with DMF (5x 2 mL) and 20% piperidine in DMF (1x, 2 mL) before incubatingwith 20% piperidine in DMF for 10 min. The resin waswashed with DMF (5 x 2 mL) before addingthe next Fmocprotected aminoacid.
[00105] The resin (350 umol, 1.0 eq)was divided over 130 eppendorf tubes.A solution of carboxylic acid (8.08 pmol, 3.0 eq), HATU (80.2 pL, 0.1 M, 8.02 pmol, 2.98 eq) and DIPEA (4.22 pL, 24.23 pmol, 9.0 eq) in DMFwas added to the resin (2.69 pmol). The reactionswere stirred overnight at r,t. The resin was pooled and washed with DMF (5x) and DCM (5x).
[00106] The resin was incubated for 1.5 hourwith a solution ofTFA:H20:T|S (95:2.5:2.5) and washed once with a solution ofTFA:H20:T|S (95:2.5:2.5). TFAwas evaporated and the librarywas purified with using reverse phase column chromatographywith a stepwise gradient 00701 00% MeCN. SEL 2: Benzimidazole 1 Hai- F 315.3 __ __ _ 1 mon-up H :) o mam on \: :; ' HN H H ?thle lLNO 5 s-c lIENÛìN 'J'LNO .) n :".MQ IIIS HJ:jjr° NH: O H Di.-1! ii 0 '1 O H 0
[00107] TentaGel S NH2 resin (90 pm, 0.24 mmol / g loading, 1.3150 g, 315 pmol,1.0 eq.)was transferred to a 20mL fritted syringe and was subsequently swelled with DMF for 10 min.A solution of FmocRinkOH (509.9 mg, 0.945 mmol, 3.0 eq), 0.4M HATU (0.939 mmol, 2.346 mL, 2.98 eq) and DIPEA (494 pL, 2.835 mmol, 9.0 eq) in DMFwas added to the resin and reacted for 2h. The resin waswashed with DMF (5x 2 mL) and 20% piperidine in DMF (1x, 2 mL) before incubating with 20% piperidine in DMF for 10 min.
[00108] The resin was divided over 62 x 2mL fritted syringes. Stock solutions ofamino acids were made with Fmocprotected amino acids (100 pmol) in 0.4M HATU (250 pL). A solution of Fmoc protected aminoacid (3.0 pmol, 3.0 eq , (37.5 pL from premade stock solution), DIPEA (7.84 pL, 45.00 pmol, 9.0 eq) and DMF (37.5 pL) were added to the resin (5.00 pmol). After 2h the resin was combined in a 20mL fritted syringe and waswashed with DMF (5x 2 mL) and 20% piperidine in DMF (1x, 2 mL) before incubatingwith 20% piperidine in DMF for 10 min. The resin waswashed with DMF (5 x 2 mL) before addingthe next Fmocprotected aminoacid.
[00109] The resin was pooled and a solution of 4fluoro3nitrobenzoic acid (175 mg, 945 pmol, 3.0 eq), HATU (2.346 mL, 939 pmol, 2.98 eq) and DIPEA (494 pL, 2.835 mmol, 9.0 eq) in DMFwas added to the resin. After 1 h the reaction aswashed with DMF (5 x 2 mL).
[00110] The resin was divided over 52, 2mL fritted syringes. TentaGel S NH2 resin (30 pm, 0.24 mmol / g loading, 5.96 pmol,1.0 eq.)was transferred to an eppendorf. Amines (59.62 pmol, 10 eq.) and DIPEA (10.55 pL, 59.62pmol, 10 eq.) in DMF (150pL, 0.4 M). The mixturewas shaken at 600 rpm overnight at 80 °C. The resin waswashed with DMF (5 x 2 mL) and DCM (5 x 2 mL).
[00111] The resin was incubated for 1.5 hourwith a solution ofTFA:H20:T|S (95:2.5:2.5) and washed once with a solution ofTFA:H20:T|S (95:2.5:2.5). TFAwas evaporated and the librarywas purified with using reverse phase column chromatographywith a stepwise gradient 00701 00% MeCN. SEL 3: SuzukiMivaura library BIG / <0 1. .B(OH)2 o HATU, DIPEOAH o O ÊÏACFIÊÉXËhÊÎÂrKËÊÎÎ o o H2N\ JLNO _* Br_°SN- JLNii2;> .N- JLNH2 H r.t., 20 min 2. TFA:MQ:TIPS
[00112] TentaGel S NH2 resin (30 pm, 0.24 mmol / g loading, 1.104 g, 265 pmol,1.0 eq.)was transferred to a 20mL fritted syringe and was subsequently swelled with DMF for 10 min.A solution of FmocRinkOH (429 mg, 0.795 mmol, 3.0 eq), 0.4M HATU (0.790 mmol, 1.974 mL, 2.98 eq) and DIPEA (415 pL, 2.385 mmol, 9.0 eq) in DMFwas added to the resin and reacted for xh. The resin waswashed with DMF (5x 2 mL) and 20% piperidine in DMF (1x, 2 mL) before incubatingwith 20% piperidine in DMF for 10 min.
[00113] The resin was divided over 60 eppendorf tubes.A solution of Fmocprotected aminoacid (12.82 pmol, 3.0 eq), 0.4M HATU (12.74 pmol, 31.8 pL, 2.98 eq), DIPEA (6.70 pL, 38.47 pmol, 9.0 eq) and DMF (90 pL) were added to the resin (4.27 pmol). After4h the resin was combined in a 20 mL fritted syringe and waswashed with DMF (5x 2 mL) and 20% piperidine in DMF (1x, 2 mL) before incubatingwith 20% piperidine in DMF for 1 0 min. The resin was combined in a 20 mL fritted syringe and waswashed with DMF (5x) .
[00114] The resin was divided over 10 eppendorf tubes.A solution of aryl bromide (26.50 pmol, 3.0 eq), 0.4M HATU (78.97 pmol, 197 pL, 2.98 eq), DIPEA (41.50 pL, 238.5 pmol, 9.0 eq) and DMF (400 pL)were added to the resin (26.50 pmol). After4h was, the resin waswashed with DMF (5 x 2 mL). Aryl bromideAB10was deprotected bywashingwith 20% piperidine in DMF (1x, 2 mL) before incubatingwith 20% piperidine in DMF for 1 0 min.A solution of benzoic acid (9.7 mg, 79.50 pmol, 3 eq.), DIPEA(41.50 pL, 238 pmol, 9.0 eq) and DMF (400 pL)were added to the resin containingAB10 (26.50 pmol). The resinwas combined in a 20 mL fritted syringe and waswashed with DMF (5x).
[00115] The resin was divided over 53 eppendorf tubes. Boronic acid (10 pmol, 2.0 eq.), K2C03 (1.4 mg, 10 pmol, 2.0 eq.), PdCl2 (89 pg, 0.5 pmol, 10 mol%), Xphos (0.48 mg, 1 pmol, 20 mol%) and DMF (100 pL)were added to the resin (5.0 pmol). The reaction was stirred at 800 rpm at 80 °C for 21 h. The resinwaswashed with DMF (5 x 2 mL) and DCM (5x) and incubated for 1:45 h with a solution ofTFA:H20:T|S (95:25:25). The resin waswashed once with a solution ofTFA:H20:T|S (95:2.5:2.5) whereafter the TFAwas evaporated. The librarywas purified with using reverse phase column chromatographywith a stepwise gradient 00701 00% MeCN. Discussion
[00116] Aiming at affinity selection with large and diverse selfencoded libraries,we established solid phase synthesis protocols forthe preparation of combinatorial libraries with different scaffold designs. By exploring different scaffolds,we aimed both at increasing diversityand at investigating the amenability of different molecular architectures for MS / MS based decoding. In order to obtain high quality combinatorial libraries, each reaction step needs to be efficient and high yielding. Self encoded library 1 (SEL 1) is formed bythe sequential attachment oftwo amino acid building blocks, followed bythe addition of a carboxylic acid decorator using reaction conditions optimized for Fmocbased solid phase peptide synthesis (Fig. 9A). Selfencoded library 2 (SEL 2) is based on a benzimidazole core decorated on three different positions. Based on previously described methodologies (3133) and following systematic optimization,we established an efficient route towards trifunctional benzimidazoles (Fig. 9B and Fig. 10). The benzimidazole decorators which confer diversity to the library are based on an amino acid building block, a primaryamine and an aldehyde.We tested the scope of the nucleophilic aromatic substitution with a set of 92 primary amines, ofwhich a large fraction resulted in reasonable conversions for combinatorial synthesis (65%). (Fig. 9B and Fig. 11).We then investigated the heterocyclization efficiency using 95 aldehydes, with 65 resulting in a > 55% conversion to the final trifunctionalcompounds (Fig. 9B and Fig. 12). Selfencoded library 3 (SEL3) results from an amino acid building block linked to an aryl bromide and subsequently crosscoupled to a boronic acid, utilizing the palladium catalyzed SuzukiMiyaura reaction (34).We optimized reaction conditions on a selected scaffold and then tested the scope of 19 bifunctional aryl bromides ofwhich 9 resulted in a > 65% conversion (Fig. 9C and Fig. 13). Out of 85 boronic acids, 50 resulted in a > 65% conversion (Fig. 9C and Fig. 14). Representative crude LCMS traces showthe quality of the synthesis for each library scaffold (Fig. 9A-C).
[00117] Using a virtual library scoring scriptwe selected building blocks to generate libraries with optimized druglike properties. For SEL 1,we anticipated no synthetic limitations, sowe decided to select our initial set of building blocks based on their druglike properties, while limiting isobaric fragments.We filtered a comprehensive building block (BB) catalog (Fmocamino acids and carboxylic acids from Chemspace) by availabilityand price, narrowing itdown to 1000 BBs per position. With thesewe enumerated a virtual librarywith a billion members. Each memberwas scored based on five Lipinski parameters: molecularweight (mw), logP, hydrogen bond donors (HBD), hydrogen bond acceptors (HBA), and topological polar surface area (TPSA). Each library members received a point for each satisfied parameter, whichwas then translated to a combined score per building block. This ranking allowed us to select and purchase topscoring building blocks (62amino acids and 130 carboxylic acids).Through solidphase split and pool synthesis,we generated a half million membered SEL 1 . Compared to the original enumerated library, all Lipinski parameters of our SEL 1 were significantly improved (Fig. 15). For SEL 2 and 3,we used one of the optimized amino acids for BB1 and selected BBs with yields greater than 55% in scope analysis for the other positions (resulting in 216,008 members for the benzimidazole scaffold and 31,800 members forthe Suzuki based library). Overall, the majority of allcompounds in the three libraries satisfythe requirements for druglike properties (Fig. 9D). LibragzDecodingSoftware
[00118] A crucial step in the affinity selection workflow is the accurate identification of hit compounds. The final sample from an affinity selection process is always of unknown complexity and may contain up to a few hundred compounds. With the high degree of mass degeneracy in or libraries (i.e., the presence of multiple isobaric compounds with different molecular structures) structure annotation based on MS / MS fragmentation spectra is essential for unequivocal compound identification (Fig. 16A). To investigate the decodability of our libraries and simulate the final sample from an affinity selection,we prepared defined subsets of 245500 compounds for each scaffold and analyzed each subset via nanoLCMS / MS (Fig. 16B). Based on the known list of librarycompoundswe analyzed the data and counted the detectable compounds the number of detectable compounds might be lowerthan the theoretical number because of loss duringsample preparation of highly polar / unipolar structures). Overall, each nanoLCMS / MS run produced approximately 80,000 MS1 and MS2 scans, including mainly spectra resultingfrom background noise: for a real affinity selection sample of unknown content and complexitythe manual analysis of such a datasetwould be totally impractical.
[00119] We first attempted the automated structure annotation for our test samples with off the shelf metabolomics software. While typical metabolomics workflows use spectral databases as an input to increase recall rates and confidence and correct structure annotation, our libraries representing novel chemical matter do not have such spectral databases. To this endwe utilized SIRIUS 6 with CSI:Finger|D, considered bestinclassfor reference spectra free structure annotation of small molecules (3537). SIRIUS annotates compounds by scoring predicted molecular fingerprints against fingerprints of database structures (e.g. PubChem). In an affinity selection experiment, contrary to a regular metabolomics analysis, the complete space of potential structures is alreadyknown and the fullyenumerated librarycan be used as a structure database to score compounds against. To this end,we created custom structure databases in SIRIUS, consisting of the fullyenumerated librarySMILES foreach library (500k, 200k and 30k respectively).We then imported the measured nanoLCMS / MS runs into SIRIUS and performed a standard SIRIUS structure annotation workflow. While a good fraction of librarycompounds were correctly detected and annotated (82%, 77% and 90%, respectively for SEL 1, 2 and 3), the total number of annotated scans (up to 2800) largelyexceeded the number of molecules actually present in the library (Fig. 16C). Manual inspection in principle can help picking out correct annotations butwith thousands of falsepositives this off the shelf automated annotation workflowwould still be impractical for real affinity selections samples with unknown content.
[00120] To increase the proportion ofgenuine library molecules in the final set of proposed compounds and to improve correctcompound annotation,we set off to establish fragmentation rules to predict likelyMS / MS patterns for our librarycompounds.We thus analyzed the fragmentation spectra of our test librarycompounds and calculated fragmentation frequencies of bonds connecting the various building blocks (Fig. 16D). For each of the three library scaffoldswe defined the most prominent recurringfragmentation modes. Based on these patterns,we generated a combinatorial fragmentor to create a list of predicted fragments for each library member (Fig. 16E).We then implemented a filter in SIRIUS, stipulating that onlyscans with an MS1 precursor mass matching a librarycompounds mass and containing at least one predicted fragment peak in their MS2would be selected for full annotation (Fig. 16F). This filter drastically reduced the number of total scans (compare Figure 16C and 16F), while maintaining a high correct recall and annotation rate of 6674%. Overall, this SIRIUSCOmbinatorial Mass Encoding decoding Tool (SIRIUSCOM ET) enables the high fidelity annotation and decoding of satisfyingnumbers of compounds from all three library scaffolds, with significantly reduced numbers hallucinated annotations. Affinity Selection Against CAIX Procedure
[00121] MyOne Streptavidin T1 Dynabeads (100 pL of 10 mg / mL stock per well, measured in duplicates) werewashed with 3x1 mL 10% FBS,1x PBS, 0.02%Tween20. The beads were incubated with biotinylated CAIX (100 pL, 1.5 pM per well) in 10% FBS,1x PBS, 0.02%Tween20 for 1h at 4°C. The beads werewashed 2x1 mL 10% FBS,1x PBS, 0.02%Tween20, 400 pM dbiotin and 1x1 mL 10% FBS,1x PBS, 0.02%Tween20 before incubating with the library (100 pL, 100 fmol / member per well) in 10% FBS,1x PBS for 1h at 4°C. The beads werewashed with 5x1 mL of 1x PBS and subsequently eluted with 2xMeCN:MQ (1 :1) 0.1%FA (100 pL per well).
[00122] Sample preparation afterAS: Samples from the affinity selection procedurewere lyophilized and resuspended in 50 pLMQ 0.1%FA. The StageTips were prepared as described by Rappsilber et al. usingC18 materialfrom Empore SPE 47mm discs (66883U, Merck).1 The StageTipswere preconditioned with 200 pL MeOH, 200 pL of 0.1% (v / v) FA in MeCN and 200 pL of 0.1% (v / v) FA in MQ, respectively by centrifuging for 3 min at 300 rpm. The samples were then loaded on the StageTips and washed with 200 pL of0.1% (v / v) FA in MQ. Compoundswere eluted by adding 200 pL of0.1% (v / v) FA in MeCN:MQ (7:3). The sampleswere lyophilized before resuspending in 10 pL 0.1% (v / v) FA in UPLCMS grade water. The samples were centrifuged for 5 min at 15.000 rpm. AftenNards, 9 pLwas transferred to a LCMS vial and 8 pLwas injected into the LCMS / MS system. ScreeningofSEL 1 againstCAIX
[00123] UsingS|R|US:COMET filterwith a least matching 1 out of the 5 biggest peaks and 5ppm accuracy in combination with CSI:FingerID structure annotation identified 74 unique structures containing building blockCA77. The structures of the 74 hitcompounds found afterASMS against CAIXwith SEL1 are provided below: