Method for determining interaction geometries

EP4744049A1Pending Publication Date: 2026-05-20IMPERIAL COLLEGE INNVOATIONS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
IMPERIAL COLLEGE INNVOATIONS LTD
Filing Date
2024-07-12
Publication Date
2026-05-20

AI Technical Summary

Technical Problem

Current methods for determining the geometry of molecular interactions, such as protein-ligand interactions, are limited by the need for crystallization, high costs, and slow processes, particularly for intrinsically disordered targets, and lack sufficient sensitivity and throughput for drug discovery applications.

Method used

The use of Electron-Vibration-Vibration (EVV) two-dimensional infrared (2DIR) spectroscopy in combination with quantum chemistry calculations based on density functional theory to rank structural hypotheses by comparing experimental and predicted spectra, allowing for the determination of molecular complex geometries.

Benefits of technology

This approach significantly reduces the time and cost of drug discovery by enabling rapid determination of multiple structures with small quantities of target molecules, providing precise geometric information on drug-target interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000022_0001
    Figure IMGF000022_0001
  • Figure IMGF000036_0001
    Figure IMGF000036_0001
  • Figure IMGF000036_0002
    Figure IMGF000036_0002
Patent Text Reader

Abstract

A method of determining the geometry of interaction between a protein or nucleic acid and one or more chemical entities, the method comprising the steps of: (a) providing an experimentally obtained multi-dimensional spectrum for a sample comprising the protein or nucleic acid and the one or more chemical entities; (b) providing two or more structural hypotheses; (c) obtaining a computationally generated multi-dimensional spectrum of the same type as the experimentally obtained multi-dimensional spectrum for each of the two or more structural hypotheses; (d) determining which of the computationally generated multi-dimensional spectra best fits the experimentally obtained multi-dimensional spectrum; and (e) based on step (d), identifying the geometry of interaction between the protein or nucleic acid and one or more chemical entities.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR DETERMINING INTERACTION GEOMETRIES

[0002] Field

[0003] The present invention relates to a method of determining the geometry of interaction between a macromolecule of interest such as protein or nucleic acid and one or more chemical entities.

[0004] Background

[0005] There are a wide range of techniques that can be used to gain structural insight into molecular structure and the geometry of molecular interactions. These include X-ray crystallography, cryo-electron microscopy and two-dimensional NMR spectroscopies.. Molecular recognition of one molecule for another is a key component of biological mechanism. Molecular recognition enables selected and specific interactions and binding to be enhanced whilst reducing interactions with other molecular species. A good example is the binding of an enzyme substrate or co-factor to the protein active site. Other examples include the binding of one protein to another for purposes of regulation, and the binding between proteins and nucleic acids for regulation of gene expression.. It is the specific spatial arrangement of atoms and functional groups which determine the stability and specificity of the interactions which underpin molecular recognition. Sometimes dynamics also play a role, causing potential binding motifs and structures to appear transiently until stabilised by an interaction.

[0006] One specific application is in the small-molecule drug discovery and development process. Small molecule drug discovery has a significant iterative element in its early stages, where multiple potential drug molecules are synthesised in an attempt to improve the characteristics of the putative drug molecule. This can be in hit to lead development, in lead optimisation or at any point in the process where improvements in the performance of the drug are required. As part of these iterative processes, it is essential to be able detect and characterise the interaction between a drug and a target (for example, protein-ligand interactions). Beyond this, it is also important to verify the location of the drug, detect the contacts made between the drug and target, and to determine the geometry of the overall complex formed by the drug and target. These data allow chemists to compare drug candidates against each other and to identify which elements of the molecules should be altered to improve its binding characteristics. Thus obtaining structural information on the drug-protein interaction is an essential component of most drug discovery and development processes.

[0007] Despite the power and utility of the suite of methods currently used for structural analyses, there are many cases where practical limitations mean that the desired structural information cannot be retrieved in sufficient quantity, with sufficient precision or even at all using existing techniques. This includes, for example, cases where the protein cannot be crystallized, where the crystal structure has regions of disorder which cannot be resolved, where the precision of the structure at the ligand contact is insufficient for driving the drug discovery process, or where the region of interest in the protein is relatively unstructured in its native condition.

[0008] At present, a drug development campaign can involve synthesising several thousands of molecular variants and performing subsequent synthesis steps based on occasional crystal structures of lead series candidate molecules. Although electron microscopy is starting to replace x-ray crystallography for structure determination de-novo, it remains too slow and expensive to be useable for drug discovery campaigns. Further, is unclear whether electron microscopy can be used for intrinsically disordered targets, which are currently thought to comprise -30% of all proteins. Overall, due to the cost, time, and other practical limitations of current methods, there exists a need for improved methods for gaining structural insight into molecular interactions.

[0009] Spectroscopic methods may be used to develop general assay methods which are sufficiently competent in binding detection, quantification of binding, site verification, and structural analysis. In particular, variants of two-dimensional nuclear magnetic resonance (2DNMR) spectroscopy methods are used for structural analysis and a range of optical spectroscopies are used for binding detection, often in a fluorescence assay format. However 2DNMR methods generally do not have the sensitivity or throughput to be used as a screening method for analysing drug binding, and optical spectroscopies do not, in general, have sufficient structural competence to be used for structural analysis.

[0010] Summary

[0011] Despite the relatively limited use of optical methods in structural analysis, the present inventor has employed a variant of two-dimensional infrared (2DIR) spectroscopy called Electron-Vibration-Vibration (EVV) to determine the geometries of molecular interactions. The inventor has discovered that when EEV 2DIR spectroscopy is used in combination with quantum chemistry calculations based on density functional theory, it is possible to rank a set of structural hypotheses for the geometry of molecular complexes by comparing experimental 2DIR spectra with predicted 2DIR spectra calculated from the structural hypotheses. Ranking a set of structural hypotheses can include determining the geometry of a molecular complex where the set of structural hypotheses includes the actual geometry. The method developed by the inventors is particularly advantageous in that it enables the cost and time of the drug discovery process to be reduced significantly. This is because it is possible to determine many structures a day using the invention and that this can be done using relatively small quantities of target molecule, typically but not always a protein.

[0012] The geometry in this context is synonymous with the pose of the molecule in a drug binding site as it is generally known. In both drug-target interactions and in other forms of biomolecular recognition or non- biological interactions, the geometry comprises the set of contacts between the chemical entities. The set of contacts includes hydrogen bonds, salt bridges, dipole-dipole interactions and dispersion forces including pi-pi interactions. There is provided a method of determining the geometry of interaction between a protein or nucleic acid, and one or more chemical entities, the method comprising the steps of:

[0013] (a) providing an experimentally obtained multi-dimensional spectrum for a sample comprising the protein or nucleic acid and the one or more chemical entities;

[0014] (b) providing two or more structural hypotheses;

[0015] (c) obtaining a computationally generated multi-dimensional spectrum of the same type as the experimentally obtained multi-dimensional spectrum for each of the two or more structural hypotheses;

[0016] (d) determining which of the computationally generated two-dimensional spectra best fits the experimentally obtained multi-dimensional spectrum; and

[0017] (e) based on step (d), identifying the geometry of interaction between the protein or nucleic acid, and one or more chemical entities.

[0018] There is also provided a system for determining the geometry of interaction between a protein or nucleic acid and one or more chemical entities, the system comprising:

[0019] (a) a spectrometer arranged to experimentally obtain a multi-dimensional spectrum for a sample comprising the protein or nucleic acid and the one or more chemical entities; and

[0020] (b) one or more processors arranged to: generate or receive two or more structural hypotheses each comprising a computationally generated multi-dimensional spectrum; determine which of the computationally generated multi-dimensional spectra best fits the experimentally obtained multi-dimensional spectrum; and identify the geometry of interaction between the protein or nucleic acid and one or more chemical entities based on the determination of which of the computationally generated multi-dimensional spectra best fits the experimentally obtained multidimensional spectrum.

[0021] There is also provided a computer-readable medium comprising computer-executable instructions which, when executed by a processor, cause the processor to perform any of the methods described herein.

[0022] The inventors have advantageously found it is possible to determine the geometry of interaction by comparing two or more structural hypotheses to determine which of these is the best fit to the experimental data. This approach allows for investigation of a geometry of interaction of interest which would not be possible from the experimental data alone.

[0023] Detailed Description of Embodiments In a first aspect, there is provided a method of determining the geometry of interaction between a protein or nucleic acid and one or more chemical entities, the method comprising the steps of:

[0024] (a) providing an experimentally obtained multi-dimensional spectrum for a sample comprising the protein or nucleic acid and the one or more chemical entities;

[0025] (b) providing two or more structural hypotheses;

[0026] (c) determining a computationally generated multi-dimensional spectrum of the same type as the experimentally obtained multi-dimensional spectrum for each of the two or more structural hypotheses;

[0027] (d) determining which of the two or more structural hypotheses best fits the experimentally obtained multi-dimensional spectrum; and

[0028] (e) based on step (d), identifying the geometry of interaction between the protein or nucleic acid and one or more chemical entities.

[0029] The method may be carried out in any suitable order.

[0030] The method may be carried out in the order (a), followed by (b), followed by (c), followed by (d). The method may be carried out in the order (b), followed by (a), followed by (c), followed by (d).

[0031] Step (b) or steps (b) and (c) can be carried out before step (a).

[0032] The term “geometry of interaction” as used herein refers to the arrangement or spatial configuration of two molecules, or part of molecules, or within a molecule. The specific configuration of two molecules (or parts thereof) may arise due to interactions such as hydrogen bonds, electrostatic interactions and / or dispersion forces. The specific configuration may arise due to the cumulative sum of ferees from these interactions and other effects such as thermal fluctuations. The configuration is usually considered to be a geometry which minimises the free energy of the interacting chemical entities.

[0033] The geometry of interaction is suitably the relative orientation of and / or the distance between two molecules, or parts of molecules, e.g. which are not connected by a covalent bond. The geometry of interaction can comprise or consist of the distances between interacting chemical groups. As used herein, interacting chemical groups are moieties that give rise to a feature on the spectrum (e.g. as a result of their proximity).

[0034] A “chemical entity” as used herein refers to any molecule that is capable of forming an interaction (as set out above) with a protein or molecule of interest. This includes interactions between chemical groups, such as amino acid side chains or the nucleotide bases of nucleic acids, which interact with each other without being directly covalently bonded. The chemical entity may be another part of the protein or molecule of interest, e.g. DNA molecule or a part thereof, such as a nucleotide base. Thus the chemical groups and / or chemical entities that interact can be separate molecular species or the interactions can be within one chemical entity. This includes but is not limited to interactions within a protein, within a nucleic acid, or between a protein and another chemical entity or between a nucleic acid and another chemical entity.

[0035] It will therefore be appreciated that alternatively or additionally, the present disclosure describes a method of determining the geometry of interaction between a molecule of interest or part thereof and one or more chemical entities, the method comprising the steps of:

[0036] (a) providing an experimentally obtained multi-dimensional spectrum for a sample comprising the molecule of interest or part thereof and the one or more chemical entities;

[0037] (b) providing two or more structural hypotheses;

[0038] (c) obtaining a computationally generated multi-dimensional spectrum of the same type as the experimentally obtained multi-dimensional spectrum for each of the two or more structural hypotheses;

[0039] (d) determining which of the structural hypotheses best fits the experimentally obtained multi-dimensional spectrum; and

[0040] (e) based on step (d), identifying the geometry of interaction between the molecule of interest or part thereof and one or more chemical entities.

[0041] The molecule of interest may be any suitable molecule. The molecule of interest is typically a macromolecule. It may be a polymer or biopolymer. It may be part of a molecule. The molecule of interest may be a protein. It may be part of a protein. It may be a nucleic acid, for example RNA, or DNA.

[0042] Thus, the methods described herein may be used to determine the geometry of interaction between a nucleic acid, such as a pre-microRNA, and one or more chemical entities.

[0043] The methods described herein may be used to determine the geometry of interaction between a ligand in a cryptic or allosteric site and a protein.

[0044] The methods described herein may be used to determine the geometry of interaction between two or more polymer chains, for example to determine the relative orientation of said chains.

[0045] The one or more chemical entities may be one or more drug molecules (or part thereof). As used herein, a drug molecule may be a drug candidate in a drug development or a chemical entity that is suitable for use a drug or for use in drug development

[0046] Suitably there is only a single drug molecule and the method is a method of determining the geometry of interaction between a protein and a drug. The geometry of interaction between a protein and a drug (ligand) can also be referred to as the ligand pose.

[0047] Therefore, the method of the first aspect may be a method of determining the geometry of interaction between a protein (or part thereof) and a drug molecule (or part thereof). As known in the art, drugs are considered to bind to receptors and may therefore alternatively or additionally be referred to as ligands.

[0048] The geometry of interaction may be located at a target binding site. The present disclosure may relate to a method of determining the geometry of interaction of ligand-receptor complex structure.

[0049] The method of the first aspect includes (a) providing an experimentally obtained multi-dimensional spectrum for a sample.

[0050] As used herein “multi-dimensional spectrum” refers to a spectrum that is obtained by spectroscopy which is not one-dimensional, i.e. it refers to a spectrum obtained by spectroscopy in two or more dimensions.

[0051] Traditional spectroscopy involves the measurement of a sample's response as a function of a single variable, such as time or frequency, while exposing it to a range of frequencies or wavelengths. This provides insights into the sample's absorption, emission, or scattering properties at a specific energy or frequency.

[0052] In contrast, multi-dimensional spectroscopy introduces additional dimensions, often incorporating multiple variables such as time delays, frequencies, or pulse sequence parameters. By manipulating these parameters, multi-dimensional spectra dependent on multiple variables are acquired, and correlations between different variables are explored, enabling a more thorough investigation of the sample's behaviour. Correlations can be generated through a variety of mechanisms, including, for example, coupling between molecular vibrations and or between electronic states and between vibrations and electronic states.

[0053] Two-dimensional (2D) spectroscopy obtains a two-dimensional spectrum by exciting the sample with a series of pulses and measuring the response as a function of at least two of the excitation frequencies and detection frequencies. When short optical pulses are used, multiple correlations can be excited simultaneously due to the Fourier relationship between the time and frequency domains. This Fourier relationship can be exploited to enable measurement of the response in the time domain. Thus sample responses and the resulting correlations can be measured in the time-domain, the frequency domain or both. Cross-peaks emerge in the resulting spectrum, providing information about the couplings and interactions between different electronic states and / or vibrational modes within the sample.

[0054] The method described herein may also involve measuring a multi-dimensional spectrum. Alternatively or additionally, the method may involve input of an experimental multi-dimensional spectrum that has already been measured.

[0055] Methods of measuring multi-dimensional spectra are known in the art. The multi-dimensional spectrum may be obtained using two-dimensional infrared ‘2DIR’ spectroscopy.

[0056] The multi-dimensional spectrum may be obtained using electron-vibration-vibration (EVV) two- dimensional infrared ‘EVV 2DIR’ spectroscopy alternatively known as DOubly Vibrationally Enhanced ‘DOVE’ spectroscopy .

[0057] EVV 2DIR spectroscopy measures the coupling of a 1 quantum (fundamental vibration) transition to a 2 quantum (combination band or overtone) transition. If the combination band which is excited contains the 1 quantum transition which is also excited, then Raman scattering between the two levels is an allowed transition. The visible beam, therefore, effectively probes the coupling between the two excited vibrations, one of which is a vibrational fundamental frequency and one of is a combination band or overtone. The resulting measured correlations can be plotted in different ways. If the frequencies of the two excitation beams are used as the axes, then the spectra produced have columns of features corresponding to fundamental transitions which couple to many other vibrations. They also have diagonal features which are in effect lines, where the second 1 quantum transition, which is contained within the combination band, is constant. The frequency of the second vibration within the combination band is easily estimated by subtracting the frequency of the first excitation beam, wa, from that of the second excitation beam, wp. This assumes that the frequency shift in the combination band due to coupling itself is small or negligible, which is not always the case. If the spectra are plotted by plotting the estimated excitation frequency of the second one quantum transition against the first one quantum transition, then the spectra produced have columns and rows of features, each corresponding to one quantum transitions that couple to multiple other one quantum transitions. However the spectra are plotted, the rows, columns and diagonal features in the two-dimensional spectra are sometimes approximate due to factors such as interference between spectral features and frequency shifts due to additional couplings or processes. These rows and columns can be helpful in confirming and assigning cross-peaks to specific interactions by providing additional confidence in multiple couplings and thereby creating enhanced confidence in the comparison of the calculated and experimental spectra. Sometimes only single peaks are present in a row or column due to a variety of factors such as some couplings having either short lifetimes or being too weak to detect.

[0058] Suitable spectrometers for measuring EVV 2DIR spectra are known in the art.

[0059] Suitable spectrometers comprise an ultrafast laser system, which is used to generate one or more intense laser beams. The generated laser beam(s) may be split into three arms, two of which drive optical parametric amplifiers to produce tuneable mid-IR outputs. The third beam is used to generate Raman scattering in the sample and so should not be resonant with vibrational absorptions of the sample. This third beam is usually at the frequency of the initially generated intense laser beam. All three beams may be linearly polarized parallel to the plane of propagation or in different polarisations, spatially overlapped in the sample at diameters chosen to maximise the signal intensity without damaging the sample. Different pulse sequences and different combinations of polarisation may be used, as known in the art.

[0060] Advantageously, the inventor has found that measuring a EW 2DIR spectrum means that the sequence can be tuned over a wide spectral range, making a greater part of the coupling spectrum accessible than is the case with two dimensional infrared spectrometers designed to measure vibrational couplings in other ways, which do not employ a Raman scattering process as part of the mechanism for generating the two dimensional spectrum. Moreover, a wider piece of the spectrum can be measured in a given experiment, typically measuring about one hundred spectral features for large biomolecules such as proteins and nucleic acids. These multiple features are highly advantageous in assigning experimental features by means of calculating spectra or other means. It also provides a multitude of spectral features, often aligned in rows or columns, which are used for comparing structural hypotheses, thus increasing the ability of the technique to distinguish structural hypotheses.

[0061] The time delays between the three optical pulses have to be chosen to minimize the non-resonant background generated by nonlinear frequency mixing within the sample, whilst maximizing the signal arising from vibrational couplings. The timings chosen are generally a compromise between these two competing factors and these compromises are generally determined empirically and is a function of pulse durations, pulse frequencies and sometimes the lifetimes of the couplings of interest. Typical timing delays range from 0.5ps to 2ps, but could be as from as short as 0.1 ps up to 10ps in particular circumstances.

[0062] In some designs of an EVV 2DIR spectrometer, shorter infrared pulses can be used to excite multiple vibrational states simultaneously, as is known in the art.

[0063] According to a second aspect, there is provided a system for determining the geometry of interaction between a protein or nucleic acid and one or more chemical entities, the system comprising

[0064] (a) a spectrometer arranged to experimentally obtain a multi-dimensional spectrum for a sample comprising the protein or nucleic acid and one or more chemical entities; and

[0065] (b) a processor arranged to: generate or receive two or more structural hypotheses determine a computationally generated multi-dimensional spectrum for each of the two or more structural hypotheses; determine which of the two or more structural hypotheses best fits the experimentally obtained multi-dimensional spectrum; and identify the geometry of interaction between the protein or nucleic acid and one or more chemical entities based on the determination of which of structural hypotheses best fits the experimentally obtained multi-dimensional spectrum. The following features according to the first and second aspects will be described below.

[0066] For the avoidance of doubt, the sample will contain the molecule of interest, e.g. protein or nucleic acid, and the chemical entity for which the geometry of interaction is being determined. The interaction may be between a first part of the molecule of interest and a second part of the molecule of interest.

[0067] The interactions can be one or more selected from: nucleic acid-drug, protein-nucleic acid, proteinprotein, nucleic acid-nucleic acid, intra-protein, intra-nucleic acid and non-biological chemical interactions within the same molecule or between molecules such as the interaction between chains in an artificial polymer.

[0068] When the method of the first aspect is a method of determining the geometry of interaction between a protein and one or more drug molecules, the sample comprises the protein and the one or more drug molecule.

[0069] When the method of the first aspect is a method of determining the geometry of interaction between a protein and a drug molecule, the sample comprises the protein and the drug molecule. The sample may therefore be termed a “protein-drug” sample.

[0070] The sample may comprise some water. The sample may comprise a buffer suitable for the sample e.g. phosphate buffer. The sample may comprise salt (e.g. NaCI). The sample may also contain solvent, such as dimethyl sulfoxide (DMSO) or other species needed to put the sample into a particular state or condition.

[0071] The sample may be deposited on a substrate in the form of a microarray comprising at least one spot of deposited material. The microarray may be maintained at controlled humidity (e.g. within a humidity range) before and during measurement of the sample.

[0072] As known in the art, traditional microarrays are orderly microscopic spots, otherwise known as features. These are deposited such that they attach to a solid surface, i.e. a substrate. The spots can be 1 mm to 1 pm. Typically the spots are 200 to 500 pm.

[0073] The substrate may be transparent to visible light or reflective of visible light. The substrate may be glass. The substrate may be plastic. The substrate may additionally be transparent to IR light if measurement of IR transmission is also helpful in which case the substrate may be calcium fluoride or another IR transparent material.

[0074] The sample may scatter a substantial amount of incident visible light.. The source of the incident visible light is the third laser beam. It is this beam that is measured after it has passed through the sample. In the absence of mie scattering or Rayleigh scattering all of the light emerges, but some of it undergoes inelastic (Raman) scattering. Because the Raman scattering is generated by a coherent process and emerges at a slightly different wavelength to the incident visible beam, the Raman scattered signal (beam) occurs at a slightly different angle from the input beam, but because this is coherent Raman scattering, it is still a beam, not a 360 degree scatter.

[0075] The Raman scattered signal can be transmitted through the substrate prior to detection, or it can be reflected by the substrate prior to detection. If the sample is highly elastically scattering too (i.e. a powder which produces a lot of mie scattering), then a lens or mirror needs to be used in addition to collect as much of the scattered Raman signal as possible.

[0076] Because in EVV 2DIR spectroscopy the Raman scattered wavelength is at a different wavelength from any of the input beams, the Raman signal is easily separated from any scattered light by use of a filter or spectrograph or both. This makes the technique more robust in the presence of mie or Rayleigh scattering than other forms of 2DIR spectroscopy, where the output wavelength overlaps with one or more input wavelengths.

[0077] The scattered light may be collected by an optical component or components forming a lens or mirror before being detected.

[0078] The method of the first aspect comprises (b) providing two or more structural hypotheses and for each structural hypothesis obtaining a computationally generated multi-dimensional spectrum of the same type as the experimentally obtained multi-dimensional spectrum.

[0079] For example, in the context of drug discovery for which the method described herein may find particular utility, a structural hypothesis is the pose of the ligand (of a drug molecule) with respect to the protein, more specifically within the known or putative binding site. As known in the art, pose is used to denote the relative orientation of the ligand with respect to the protein structure.

[0080] The structural hypotheses may be created using a docking model.

[0081] In such models, multiple structural permutations are tested and according a set criteria. The ranking process of a typical docking model usually has two main elements, a search algorithm to explore conformational space and a scoring algorithm to rank their likelihood of being a possible pose.

[0082] Alternatively or additionally, the structural hypotheses may be created through use of an ab initio protein structure prediction algorithm, such as alphafold code or the like, followed by use of a docking model.

[0083] Alternatively or additionally, the structural hypotheses may be based on spectra obtained by X-ray crystallography, NMR or electron microscopy or any other method able to provide some indication of the pose of the ligand. The structural hypotheses therefore provide possible relative orientations of the molecule of interest or part thereof and one or more chemical entities.

[0084] A computational spectrum of the same type of the experimentally obtained spectrum can therefore be generated or calculated for each structural hypothesis. Therefore, reference herein to “computationally generated spectra” refers to spectra generated on the basis of the structural hypotheses.

[0085] Thus step (c) may additionally or alternatively include:

[0086] (c) calculating a computationally generated multi-dimensional spectrum of the same type as the experimentally obtained multi-dimensional spectrum for each structural hypothesis.

[0087] Thus when reference is made herein to “spectra” in general, we mean to refer to the experimentally obtained multi-dimensional spectrum and each of the computationally generated multi-dimensional spectra.

[0088] The computational spectrum may be calculated using density functional theory (DFT). However it will be appreciated that other methods of calculating computational spectrum are known in the art. Suitable electronic structure calculations should include calculation of vibrational normal modes and the derivatives of the potential energy surface (leading to mechanical anharmonicity) and polarizability (leading to electrical anharmonicity). The computational cost of the calculation is a key factor in determining what kind of QM calculation is used. It is possible to reduce the calculational cost of the calculations by making further approximations. For example in many cases, the electronic coupling is much more significant than the mechanical coupling, so that only that need be calculated. It is also possible to calculate the coupling using a dipole approximation.

[0089] Step (d) may comprise ranking the two or more structural hypotheses.

[0090] The method can alternatively be worded as a method of ranking two or more structural hypotheses to determine the geometry of interaction between a molecule of interest, such as a protein, and one or more chemical entities.

[0091] Step (d) may comprise comparing the fit (correlation) of the experimentally obtained multi-dimensional spectrum to a computationally generated multi-dimensional spectrum determined from a first structural hypothesis , and so on, until a comparison of each of the fits of the experimentally obtained spectrum to each of the structural hypotheses has been performed. The experimentally obtained multi-dimensional spectrum may contain artefacts. The computationally generated multi-dimensional spectrum may contain artefacts. Such artefacts may alter and / or distort the spectrum.

[0092] The experimentally obtained multi-dimensional spectrum may be cleaned to remove artefacts.

[0093] Thus the method of the first aspect may comprise cleaning the experimentally obtained multidimensional spectrum to remove artefacts.

[0094] For example, the experimentally obtained spectrum may include artefacts in regions of low signal to noise. The experimentally obtained spectrum may also include non-resonant background, which may be removed. The non-resonant background is identified in regions between peaks of the spectra, where there is no overlap between them. At these frequencies, the non-zero signal is caused by non-resonant background. These regions can be used to establish the non-resonant background across the sample.

[0095] As the non-resonant background is identical for identical pulse sequences and sample composition, it need only be determined once for each case. The experimentally obtained spectrum may be normalised at a particular frequency. The intensities of the cross peaks measured in the experimental spectrum is dependent on the intensity of the infrared laser beams at each combination of frequencies. As the method of generating the infrared laser beams generates beams of different intensities at different frequencies, then these variations need to be normalised before comparing with a calculated spectrum. In order to do this the intensities of the beams must be measured across the spectral region to be used so that normalisation can be achieved. One way to measure this and to also take into account the potential for varying overlap of the beams at different infrared frequencies, is to first measure the 2DIR spectrum for a non-resonant substrate. The intensity of this non-resonant spectrum then provides the relative intensities of the product of the laser beams for each pair of frequencies of the 2D spectrum and in addition accounts for any change in beam overlap due to variation in beam properties as the wavelength output of that optical parametric amplifiers is varied.

[0096] The experimentally determined spectra may also be pre-processed to enhance the fidelity of comparison with calculated spectra. One form of pre-processing is to measure the experimental spectrum of a ligand-binding target, such as a protein or nucleic acid, both with and without the drug or ligand present. These two spectra are then combined to produce a difference or ratio spectrum. It is this difference or ratio spectrum which is then compared with the calculated difference or ratio spectrum in order to rank the structural hypotheses.

[0097] The computationally generated multi-dimensional spectra may each be cleaned to remove artefacts.

[0098] Thus the method of the first aspect may comprise cleaning the computationally generated multidimensional spectrum to remove artefacts. For example, the computationally generated spectra may include fermi resonances of artefactually high intensity which can create anomalously strong coupling features in the spectrum. These resonances may be ignored or the intensities may be limited using standard methods known in the art. One way to remove artefactually high intensity fermi resonances is to decompose the results of the spectrum calculation into its numerator and denominator. Artefactually high fermi resonances are caused by a low value denominator approaching zero and so a denominator threshold can be set to remove these resonances from the calculation. Alternatively or additionally, calculated features above a certain threshold can be removed. More sophisticated methods of cleaning calculated data include setting a maximum value for fermi resonances are known in the art.

[0099] Other forms of pre-processing of the calculated spectra may also be used to aid effective comparison of the calculated spectra with the experimental spectra. This includes taking account of the a scaling factor between calculated frequencies and measured frequencies, which is a known limitation of vibrational frequency calculations and is often used in comparing calculated linear one dimensional vibrational spectra with their experimental counterparts.

[0100] The computationally generated spectra may also be pre-processed to enhance the fidelity of comparison with experimental spectra. One form of pre-processing is to calculate the spectrum of a ligand-binding target, or a sub element such as the ligand binding site, both with and without the ligand present. These two spectra are then combined to produce a difference or ratio spectrum. It is this difference or ratio spectrum which is then compared with the experimentally determined difference or ratio spectrum in order to rank the structural hypotheses.

[0101] The experimentally obtained spectrum may be decomposed into individual features by a fitting process. In this process the spectra are represented by two dimensional gaussian spectral features each of which has a width and an amplitude. A Levenberg-Marquardt, Simplex or similar non-linear fitting algorithm is used to obtain a fit. Trial and error and iteration determines the minimum number of gaussian features required to fit the entire experimental spectrum. Fitting can be performed for both target spectra with ligand and without ligand present to allow a fitted version of the difference or ratio spectrum to be used for purposes of comparison.

[0102] The computationally generated spectra may each be decomposed into the same individual features by the same fitting process as for the experimentally obtained spectrum or by simply assigning each calculated cross peak to a gaussian feature of known intensity and fixed spectral width. The calculated cross peaks are often an amalgam of multiple couplings in the same region of the spectrum, added and smoothed by the finite spectral resolution of the spectroscopic method. Decomposition of the calculated spectra into a set of Gaussian features can be performed for both target spectra with ligand and without ligand present to allow a decomposed version of the difference or ratio spectrum to be used for purposes of comparison. The experimentally obtained spectrum may be fitted with two-dimensional Gaussian functions. The number of Gaussian functions needed will match the number of distinct couplings present. The computationally generated spectra may each be fitted with two-dimensional Gaussian functions.

[0103] A comparison of the fit may be performed to ascertain the closeness of the computationally generated spectra to the experimentally obtained spectrum.

[0104] Thus the method of the first aspect may comprise

[0105] (d1 a) decomposing the experimentally obtained spectra and the two or more computationally generated spectra into individual features.

[0106] In the method of the first aspect, (d) may include a reference or control. This may, for example, be the molecule of interest itself i.e., without the chemical entity present. This reference or control may be the protein itself, i.e. without the chemical entity present, for example without the one or more drug molecules present. This reference or control may be the nucleic acid itself, i.e. without the chemical entity present. The reference or control may be the chemical entity, such as the drug molecule, in the presence of a macromolecule such as a protein which it does not form a specific interaction with. For example, for a protein - drug complex the reference may be drug in the presence of the protein Bovine Serum albumin (BSA) . The reference or control may be the molecule of interest (such as DNA) in a different chemical state, such as at a different pH or in the presence or absence of ions (e.g. in the absence of potassium ions).

[0107] Therefore, in the method of the first aspect, (d) may comprise

[0108] (d2a) comparing an experimentally obtained spectrum of a sample comprising a protein or nucleic acid without the one or more chemical entities to an experimentally generated spectrum of a sample comprising the protein or nucleic acid and the one or more chemical entities and / or vice versa; and / or

[0109] (d2b) comparing a computationally generated spectrum of the protein or nucleic acid without the one or more chemical entities to a computationally generated spectrum of the protein or nucleic acid and the one or more chemical entities and / or vice versa.

[0110] The method of the first aspect may comprise (d2a) and (d2b).

[0111] The method of the first aspect may comprise:

[0112] (d2a’) analysing the difference (or ratio) of an experimentally obtained spectrum of a sample comprising a protein without the one or more chemical entities to an experimentally generated spectrum of a sample comprising the protein or nucleic and the one or more chemical entities or vice versa;

[0113] (d2b’) analysing the difference (or ratio) of a computationally generated spectrum of the protein without the one or more chemical entities to a computationally generated spectrum of the protein or nucleic acid and the one or more chemical entities or vice versa; and

[0114] (d2c’) determining the correlation between the difference of (d2a’) and (d2b’).

[0115] The higher the correlation found in (d2c’), the higher the ranking of the structural hypothesis associated with a specific computationally generated spectrum.

[0116] The method of the first aspect may comprise:

[0117] (1) calculating: a multi-dimensional spectrum of the protein or nucleic acid and the one or more chemical entities, and either: a multi-dimensional spectrum of the protein or nucleic acid without the one or more chemical entities, or a multi-dimensional spectrum of the one or more chemical entities without the protein or nucleic acid;

[0118] (2) identifying the changes in the multi-dimensional spectra calculated in (1) that are due to interactions between modes of the one or more chemical entities and modes of the protein or nucleic acid;

[0119] (3) repeating (1) and (2) for each structural hypothesis;

[0120] (4) measuring: a multi-dimensional spectrum of the protein or nucleic acid and the one or more chemical entities, and either: a multi-dimensional spectrum of a protein or nucleic acid without the one or more chemical entities, or a multi-dimensional spectrum of the one or more chemical entities without the protein or nucleic acid.

[0121] (5) comparing the experimentally measured changes provided in (4) with the calculated changes in (2) for each structural hypothesis. (6) ranking each structural hypothesis.

[0122] The method may be performed in any suitable order.

[0123] The method may be performed in the order of (1), (2), (3), (4), (5) followed by (6), or in the order (4), (1), (2), (3), (5) followed by (6).

[0124] The method of the first aspect comprises

[0125] (e) based on step (d), identifying the geometry of interaction between the protein or nucleic acid and one or more chemical entities.

[0126] The various methods described above may be implemented by a computer program.

[0127] Therefore, according to a third aspect, there is provided a computer-readable medium comprising computer-executable instructions which, when executed by a processor, cause the processor to perform the method according to the first aspect.

[0128] The computer program may include computer code arranged to instruct a computer to perform the functions of one or more of the various methods described above. The computer program and / or the code for performing such methods may be provided to an apparatus, such as a computer, on one or more computer readable media or, more generally, a computer program product. The computer readable media may be transitory or non-transitory. The one or more computer readable media could be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, or a propagation medium for data transmission, for example for downloading the code over the Internet. Alternatively, the one or more computer readable media could take the form of one or more physical computer readable media such as semiconductor or solid state memory, magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a rigid magnetic disc, and an optical disk, such as a CD-ROM, CD-R / W or DVD.

[0129] In an implementation, the modules, components and other features described herein can be implemented as discrete components or integrated in the functionality of hardware components such as ASICS, FPGAs, DSPs or similar devices.

[0130] A “hardware component” is a tangible (e.g., non-transitory) physical component (e.g., a set of one or more processors) capable of performing certain operations and may be configured or arranged in a certain physical manner. A hardware component may include dedicated circuitry or logic that is permanently configured to perform certain operations. A hardware component may be or include a special-purpose processor, such as a field programmable gate array (FPGA) or an ASIC. A hardware component may also include programmable logic or circuitry that is temporarily configured by software to perform certain operations.

[0131] Accordingly, the phrase “hardware component” should be understood to encompass a tangible entity that may be physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain manner or to perform certain operations described herein.

[0132] In addition, the modules and components can be implemented as firmware or functional circuitry within hardware devices. Further, the modules and components can be implemented in any combination of hardware devices and software components, or only in software (e.g., code stored or otherwise embodied in a machine-readable medium or in a transmission medium).

[0133] Unless specifically stated otherwise, as apparent from the following discussion, it is appreciated that throughout the description, discussions utilizing terms such as " receiving”, “determining”, “comparing ”, “enabling”, “maintaining,” “identifying,” or the like, refer to the actions and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system memories or registers or other such information storage, transmission or display devices.

[0134] The term “at least two” is synonymous with “two or more”, i.e. two, three, four, five, six, or more.

[0135] As used herein the term “about” generally encompasses or refers to a range of values that one skilled in the art would consider equivalent to the recited values (i.e. having the same function or result). Where the term “about” is used in relation to a numerical value, it can represent (in increasing order of preference) a 10%, 5%, 2%, 1 % or 0% deviation from that value.

[0136] It will be appreciated that all embodiments described herein are considered to be broadly applicable and combinable with any and all other consistent embodiments, as appropriate. Such combinations are considered to fall within the scope of the appended claims.

[0137] Embodiments will now be further described by way of reference to the following Examples which are present for the purposes of illustration only. In the Examples, reference is made to a number of Figures in which:

[0138] Figure 1 shows experimental EVV 2DIR spectra of (A) FGFR1 , (B) FGFR1 + SU5402, (C) difference spectrum of FGFR1 + SU5402 and FGFR1 , (D) BSA, (E) BSA + SU5402 and (F) difference spectrum of BSA + SU5402 and BSA. The difference spectrum from BSA contains spectral features of SU5402 in a non-specifically bound state, whilst the difference spectrum from FGFR1 contains features arising from SU5402 in the protein-bound state. The difference spectrum from FGFR1 contains all drug peaks present in the difference spectrum from BSA as well as an additional 6. These can be ascribed to be binding dependent. The frequency locations of these 6 binding-dependent peaks have been marked on both difference spectra.

[0139] Figure 2 shows frequencies of the binding-dependent peaks seen in the FGFR1 / SU5402 difference spectrum, rounded to the nearest 5 cm1.

[0140] Figure 3 shows average peak responses as a function of the SU5402 / FGFR1 molar ratio for the bindingdependent and nonbinding difference spectrum peaks. The levelling off of the 6 binding-dependent signals as a function of increasing drug concentration confirms them as being binding-dependent because of the saturation of available protein binding sites. As the signal depends on the square of the number of molecules, the square root of the fitted peak intensity is plotted against the SU5402 / FGFR1 molar ratio.

[0141] Figure 4 shows the geometry of the SU5402 molecule along with the five partial amino acid residues and one water molecule used for calculating the EVV 2DIR spectrum of SU5402 bound to FGFR1 . The carbon atoms of the SU5402 molecule are shown in orange, whilst the included protein carbon atoms are shown in pink.

[0142] Figure 5 shows (a) calculated spectrum of SU5402 and the FGFR1 -binding site, (b) experimental bound FGFR1 / SU5402 difference spectrum, (c) calculated SU5402 in the free-space spectrum, (d) experimental non-specifically bound BSA / SU5402 difference spectrum. From comparisons, the bindingdependent peak at 1660 / 3240 can be assigned to the coupling of mode E, an amide mode of the protein, and mode C of SU5402. Mode B is present in the calculated spectra, with and without the inclusion of the binding site, meaning experimental binding dependence of its peaks must arise from lifetime effects. The black dotted line is the overtone diagonal and all features lying on this line are overtones, which is why it passes through the intersection of lines of the same colour, representing the same mode coupled to itself as an overtone.

[0143] Figure 6 shows (a) the calculated EW spectrum of SU-5402 and its proximal amino acid residues and water molecule, (b) the experimental 1 :1 SU-5402 / FGFR1 difference spectrum for comparison. The letter labels show how the experimental peaks have been assigned onto the calculated peaks. Pink circles have been placed around the binding sensitive peaks.

[0144] Figure 7 shows the modes assigned to the a and p modes of the six binding dependent cross peaks.

[0145] Figure 8 shows the structure of SU-5402 and FGFR1 binding site atoms used for calculating the EVV spectrum. The atoms showing motion in Mode 1 are highlighted in pink. As can be seen in the Figure all of the vibrational motion of Mode 1 lies on the atoms of the SU-5402 molecule. Motion is seen on all SU-5402 atoms, except for those of the propionic acid group. The contribution of largest amplitude to this mode is an asymmetric ring deformation on the pyrrole group comprised of atoms labelled 38 — 42.

[0146] This mode forms the a mode of peaks D, E, F, G and H.

[0147] Figure 9 shows the structure of SU-5402 and FGFR1 binding site atoms used for calculating the EVV spectrum. The atoms showing motion in Mode 2 are highlighted in pink. As can be seen in the Figure, the vibrational motion of Mode 2 is split between the SU-5402 molecule and the protein binding site. The SU-5402 molecule shows motion on the pyrrole group, the amide portion of the oxindole group, and the alkene carbon linking the two. The protein shows motion on the amide back bone groups formed of atoms labelled 14, 13, 16, and 66 and 9, 8, 11 and 61 . The vibrational mode is likely conjugated between the SU-5402 molecule and the protein residues via hydrogen bonding interactions between the oxindole amide and the protein amide backbone. The contribution of largest amplitude to this mode is oxindole amide motion of the SU-5402 molecule. This mode forms the a mode of peak R.

[0148] Figure 10 shows the structure of SU-5402 and FGFR1 binding site atoms used for calculating the EW spectrum. The atoms showing motion in Mode 3 are highlighted in pink. As can be seen in the Figure, the vibrational motion of Mode 3 is split between the SU-5402 molecule and the protein binding site. The SU-5402 molecule shows motion on the pyrrole group, the amide portion of the oxindole group, and the alkene carbon linking the two. The protein shows motion on the amide back bone groups formed of atoms labelled 14, 13, 16, and 66 and 9, 8,11 and 61 . The vibrational mode is likely conjugated between the SU-5402 molecule and the protein residues via hydrogen bonding interactions between the oxindole amide and the protein amide backbone. The contribution of largest amplitude to this mode is oxindole amide motion of the SU-5402 molecule. This mode forms the p mode of peak H.

[0149] Figure 11 shows the structure of SU-5402 and FGFR1 binding site atoms used for calculating the EW spectrum. The atoms showing motion in Mode 4 are highlighted in pink. As can be seen in the Figure , the vibrational motion of Mode 4 is split between the SU-5402 molecule and the protein binding site. The SU-5402 molecule shows motion on the pyrrole group and its methyl, the amide containing five membered ring of the oxindole group, and the alkene carbon linking the two. The protein shows motion on the amide back bone groups formed of atoms labelled 14, 13, 16, and 66 and 18, 19, 21 and 68 and the methyl group of carbon 20. The vibrational mode is likely conjugated between the SU-5402 molecule and the protein residues via hydrogen bonding interactions between the oxindole amide and the protein amide backbone. The contribution of largest amplitude to this mode is oxindole amide motion of the SU- 5402 molecule. This mode forms the p mode of peaks G and R.

[0150] Figure 12 shows the structure of SU-5402 and FGFR1 binding site atoms used for calculating the EW spectrum. The atoms showing motion in Mode S are highlighted in pink. As can be seen in the Figure, the vibrational motion of Mode 5 is split between the SU-5402 molecule and the protein binding site. The SU-5402 molecule shows motion across all its atoms, except for those of the propionic acid group. The protein shows motion on parts of the amide back bone groups formed of atoms labelled 14, 13, 16, and 66 and 18, 19, 21 and 68 and the three methyl groups of carbons 15, 20 and 22. The vibrational mode is likely conjugated between the SU-5402 molecule and the protein residues via hydrogen bonding interactions between the oxindole amide and the protein amide backbone. The contribution of largest amplitude to this mode is oxindole amide motion of the SU-5402 molecule. This mode forms the p mode of peak F.

[0151] Figure 13 shows the structure of SU-5402 and FGFR1 binding site atoms used for calculating the EW spectrum. The atoms showing motion in Mode 6 are highlighted in pink. As can be seen in the Figure, the vibrational motion of Mode 6 is split between the SU-5402 molecule and the protein binding site. The SU-5402 molecule shows motion across all its atoms, except for those of the propionic acid group. The protein shows motion on the amide back bone groups formed of atoms labelled 5, 6, 7, 8 and 9 and the methyl group of carbon 2. Motion is also seen on the hydrogen atoms 66 and 60 The vibrational mode is likely conjugated between the SU-5402 molecule and the protein residues via hydrogen bonding interactions between the oxindole amide and the protein amide backbone. The contribution of largest amplitude to this mode is oxindole amide motion of the SU-5402 molecule. This mode forms the p mode of peak E.

[0152] Figure 14 shows the structure of SU-5402 and FGFR1 binding site atoms used for calculating the EW spectrum. The atoms showing motion in Mode 7 are highlighted in pink. As can be seen in the Figure, the vibrational motion of Mode 7 is predominantly on the SU-5402 molecule with small motions on some protein residues. The SU-5402 molecule shows motion across all of its atoms, including the propionic acid group. The protein shows motion only on the amide back bone hydrogens labelled 60 and 61 . The vibrational mode is likely conjugated between the SU-5402 molecule and the protein residues via hydrogen bonding interactions between the oxindole amide and the protein amide backbone. The contribution of largest amplitude to this mode is a ring deformation of the pyrrole group of the SU-5402 molecule. This mode forms the p mode of peak D.

[0153] Figure 15A shows the chemical structure of a G-quartet with metal ion coordination to Oe oxygen. The structure shows Hoogsteen hydrogen bonding motif between four guanine bases of a planar G-quartet.

[0154] Figure 15B shows the structure of the intramolecular, parallel, propeller-type G-quadruplex formed by Myc2345 in the presence of K+. The K+lie between two G-quartets (which has not been shown here for better visualization of nucleotide folded structure). Yellow rectangles represent guanine bases in the C2-endo / anti conformation.

[0155] Figure 15C is a schematic showing laser pulses used in EW 2DIR experiments. The third probe beam (eoy), non-resonant in the electronic domain, is used to look at the polarisation produced by the two resonant IR beams, while the two infrared laser beams, u>aand up cause resonant vibrations of the system. The signal is anti-Stokes to probe beam photons, a>s= a)y+ a>p - a>a. Figure 15D depicts a pulse sequence showing time delays between different pulses. The time delays used for all experimental measurements described in this work are Tap ~ Tpy ~ °'4 ps.

[0156] Figure 16 shows experimental EW 2DIR spectra of G4-forming DNA sequence, Myc2345, with varying [DNA]:[K+]. a-d, Processed EW 2DIR spectra of 1 mM Myc2345 in the presence of 0 mM (a), 5 mM (b), 10 mM (c) and 100 mM (d) K+. All spectra have been normalised with respect to the highest intensity peak at coa= 1525 cm1, cop - coa= 1450 cm1. In EW 2DIR spectroscopy, when considering the situation where combination bands and overtones are excited by one IR pulse at frequency cop while a one- quantum transition is excited by a second IR pulse at frequency coa, the coupling of one quantum transitions to their overtones is expected to fall on the dotted “overtone diagonal” line, although due to other factors such as anharmonicity, they do sometimes fall slightly below it.28Off-diagonal features arise from the coupling of vibrational modes at frequencies coaand cop - coa, with cop exciting combination bands with frequencies that are approximately the sum of the two coupled modes. Horizontal and vertical dashed lines in (a) highlight the row and column of cross-peaks with frequencies which are dependent on potassium ion concentration, which have been assigned to the coupling of a guanine carbonyl stretch mode with various other modes. The guanine carbonyl stretch mode giving rise to the row of cross-peaks (GCO I) is higher in frequency than the guanine carbonyl stretch mode giving rise to the column of cross-peaks (GCO II).

[0157] Figure 17 shows fitted EW 2DIR spectra of G4-forming DNA sequence, Myc2345, with varying [DNA]:[K+]. a-d, Processed EW 2DIR spectra of 1 mM Myc2345 in the presence of 0 mM (a), 5 mM (b), 10 mM (c) and 100 mM (d) K+fitted with a sum of 102 Gaussian functions. Markers denote the central frequencies of the 102 Gaussian functions used in the fitting of the respective spectra. The diagonal dotted line shows the overtone diagonal. Features along the diagonal arise from the coupling of a fundamental mode at u>aand its overtone at Off-diagonal features arise from the coupling of vibrational modes at frequencies u>aand a>p - a>a. Spectral changes observed as a result of K+-induced structural changes in the DNA include a decrease in cross-peak intensity (blue circles); an increase in cross-peak intensity (red circles); the red-shifting of guanine carbonyl stretch frequencies at -1700 cm-1(green diamonds); and the appearance new vibrational couplings (purple circles).

[0158] Figure 18 shows a line cut comparison of raw and fitted EW 2DIR spectra of Myc2345. Raw (dotted line) and fitted (solid line) EW 2DIR spectra of 1 mM Myc2345 in the presence of 100 mM K+at o)a= 1376 cnr1(purple) and a>a= 1482 cm1. A comparison of raw and fitted data at other a>aand a>p - a>afrequencies can be found in Figure 26.

[0159] Figure 19 shows the frequency shifts and intensity changes observed in cross-peaks as a result of K+- induced formation of full G4 structure, a-b, Frequency-potassium ion concentration dependence of the guanine carbonyl stretch mode, GCO I (a) and GCO II (b). c-i, Relative intensity of peaks exhibiting an increase (c) and a decrease (d-i) in cross-peak intensity as a function of potassium ion concentration. Intensity is presented as a fraction of the intensity in the absence of K+ions. Peaks presented here show an overall intensity difference of >20%. Intensity-ion concentration dependence of the remaining peaks is presented in Figure 26.

[0160] Figure 20 shows the new coupling peaks 92 (a>a= 1691 cm-1, a>p - a>a= 1132 cm-1) and 101 (a>a= 1705 cm1, a>p - a)a= 1675 cm1) appear due to K+-induced G4 formation, a, Experimental spectrum of 1 mM Myc2345 in the presence of 100 mM K+as a fraction of the spectrum of Myc2345 in the absence of ions. Dashed, black boxes showing regions where new coupling peaks 92 (bottom) and 101 (top) appear, b, Relative intensity of new peaks 92 and 101 as a function of [K+] concentration. Intensity is presented as a fraction of the intensity in the absence of K+ions, c, Raw (top) and fitted (bottom) spectra of 1 mM Myc2345 with 5 (left), 10 (centre) and 100 mM K+(right), as a fraction of the spectrum of 1 mM Myc2345 in the absence of ions, showing the gradual appearance of peak 101 as [K+] is increased, d, Raw (top) and fitted (bottom) spectra of 1 mM Myc2345 with 5 (left), 10 (centre) and 100 mM K+(right), as a fraction of the spectrum of 1 mM Myc2345 in the absence of ions, showing the gradual appearance of peak 92 as [K+] is increased. Vertical and horizontal dashed lines show the peak positions according to the Gaussian fit.

[0161] Figure 21 shows the computational EVV 2DIR spectra of a G-quartet and a G-quadruplex. a, Computational EVV 2DIR spectrum of a non-planar G-quartet. Four combinations of guanine carbonyl stretch modes at -1700 cm1exist in a G-quartet. The guanine carbonyl stretch mode giving rise to the highest intensity couplings along the row of cross-peaks is the in-phase guanine carbonyl stretch mode at 1694 cnr1as depicted in (f). Two degenerate out-of-phase guanine carbonyl stretch modes at 1704 cm'1, as depicted in (g) and (h), give rise to the highest intensity couplings along the column of crosspeaks. b, Computational EW 2DIR spectrum of a G-quadruplex. There are 12 combinations of guanine carbonyl stretches in a G4. The guanine carbonyl stretch mode giving rise to the highest intensity couplings along the row of cross-peaks at 1692 cm1involves the same motions as the 1694 cm1G- quartet mode, as seen in (f), across all three quartets of the G4. Similarly, two degenerate guanine carbonyl stretch modes give rise to the highest intensity couplings along the column of cross-peaks at 1679 cm1. They involve the same motions as the 1704 cm1G-quartet modes, as shown in (g) and (h), but across all three quartets in the G4. Animations of the G4 vibrational modes at 1692 cm-1and 1679 cm'1can be found in the Supplementary Information, c-d, Top-down view (c) and side view (d) of the G- quartet structure used to calculate the EW 2DIR spectrum in (a), e, Structure of G4 used to calculate the EVV 2DIR spectrum in (b). f, In-phase guanine carbonyl stretch mode of a G-quartet at 1694 cm1, g-h, Two degenerate out-of-phase guanine carbonyl stretch modes of a G-quartet at 1704 cm1.

[0162] Figure 22 shows the computational and experimental EVV 2DIR ratio spectra (G-quadruplex / G- quartet). a, Experimental EVV 2DIR ratio spectra of 1 mM Myc2345 in the presence of 100 mM K+divided by 1 mM Myc2345 in the absence of ions, b, Computational EVV 2DIR ratio spectrum of G- quadruplex / G-quartet. The calculated spectra include only electrical coupling, but a similar overall change is observed when mechanical coupling is included. Alternating increase and decrease in coupling observed around -1700 cm'1(dashed lines) are due to the red-shifts of guanine carbonyl stretch frequencies. Positive features in the experimental difference spectrum at u>a= 1691 cm1, - a)a= 1132 crrr1and a>u= 1705 cm1, - a>u= 1675 cm1(black crosses) are due to the appearance of new peaks 92 and 101 , respectively. The remaining difference features are due to intensity changes in cross-peaks as a result of structural changes in the DNA as potassium ion is introduced. The experimental ratio spectrum is consistent with structural differences between a G-quartet and a G4. Alternating columns of red and blue features on the right-hand side of both spectra is a qualitative indication that the essential features associated with those columns are reflected in both calculation and experiment as is the absence of that pattern in the corresponding rows.

[0163] Figure 23 shows a schematic of the proposed equilibrium between an unstacked G-quartet formed in Myc2345 in the absence of potassium ions and a G4 formed in the presence of potassium ions. Left: A depiction of a possible way in which Myc2345 in the absence of potassium ions might form unstacked G-quartets. Right: Schematic of G4 structure of Myc2345 in the presence of potassium ions. For clarity, only guanine bases involved in the quartets and bases in the end caps are included. New coupling peaks 92 and 101 suggest G / T in the end caps couple to the core G4 structure.

[0164] Figure 24 shows a comparison of raw and fitted experimental EW 2DIR spectra of 1 mM Myc2345 in the presence of 0 and 100 mM K+. Line cuts comparing raw (dotted line) and fitted (solid line) EVV 2DIR spectra of 1 mM Myc2345 with 0 K+(orange) and 100 mM K+(purple) at frequencies wa=1385 cm-1 (top left), Wa =1525 cm-1 (bottom left), wp - wa=1560 cm-1 (top right) and wp - wa=1560 cm-1 (bottom right).

[0165] Figure 25 shows the chi-squared values of the 2D Gaussian fitting of the experimental EVV 2DIR spectra of 1 mM Myc2345 in the presence of varying concentrations of potassium ions.

[0166] Figure 26 shows the relative intensity of cross-peaks as a function of potassium ion concentration. Intensity is presented as a fraction of the intensity in the absence of K+ions. The experimental EVV 2DIR spectra of Myc2345-K+were fitted with a total of 102 2D Gaussian functions. The intensity traces of cross-peaks observed to be increasing or decreasing with [K+], and the intensity traces of new vibrational couplings can be found in Figures 19 and 20b, respectively.

[0167] Figure 27 shows the computational EVV 2DIR spectrum of isolated guanine base, a, Computational EVV 2DIR spectrum of a guanine base. An isolated guanine base has one carbonyl stretch mode at 1758 cm1, which gives rise to a row and column of cross-peaks, b, Depiction of guanine carbonyl stretch mode in an isolated guanine base.

[0168] Figure 28 shows a sequence comparison of the G4-forming sequence in the human c-myc promoter region from different publications, a, Pu27, the 27-nt purine-rich wild type sequence found in the NHE of the c-myc promoter (SEQ ID NO: 1). b, Myc2345 sequence used for the EW 2DIR measurements in this work (SEQ ID NO: 2). c, Wild type Myc2345 sequence as used in NMR structure PDB:7KBV (SEQ ID NO: 3).10d, MycL1 - altered Myc2345 sequence used in CD measurements in Ref.11(SEQ ID NO: 4). Bases in red signify those that differ from the Myc2345 sequence used in the EVV 2DIR measurements.

[0169] Figure 29 shows the G-quadruplex and G-quartet structure used for the calculation of the computational EVV 2DIR spectrum, a-e, Structure of the G-quadruplex. Figure (b) shows a top-down view of the G- quadruplex. The terminal quartets (purple) are aligned with each other, and the central quartet (green) is -45° rotated with respect to the terminal quartets. Figures (d) and (e) present the terminal and central quartets of the G-quadruplex, respectively, f-g, Structure of the G-quartet from (f) top-down view and (g) side view.

[0170] Figure 30 shows a comparison of structures of G-quartet and terminal quartets in a G4. Structure of (a- b) G-quartet and (c) a terminal quartet of G4, showing the intersection of molecular planes of adjacent guanine bases. The guanine molecular planes intersect at an angle of 21 ° in the G-quartet and 19° in the terminal quartets of the G4.

[0171] Figure 31 shows processing of EVV 2DIR spectrum of glass coverslip used to normalise EVV 2DIR spectra of Myc2345. EW 2DIR spectrum of glass coverslip (a) raw, (b) after smoothing by averaging over nearest neighbours, (c) after interpolation along the vertical axis, (d) after applying a low-pass filter to the Fourier transform.

[0172] Example 1 : Use of Electron-Vibration-Vibration Two-Dimensional Infrared (EVV 2DIR) to detect Drug Binding to a Target Protein

[0173] In this example, it is shown that even in the most congested part of EW 2DIR spectral space, in the vicinity of the overtone coupling region, it is possible to identify spectral features that are only present when SU5402 is specifically bound to its target protein. Two independent methods are also demonstrated for verifying the specificity of binding and show that the stoichiometry of binding is broadly as expected and use a counter example of bovine serum albumin (BSA), which does not have a binding site for SU5402. Finally, quantum mechanical calculations of the vibrational couplings in the binding site, in the presence of the bound drug, are used to identify the binding-dependent features. From this, it was possible to assign at least one contact, demonstrating the potential for calculation to be used as an assignment tool even in a chemical system as complex as a drug bound to a protein.

[0174] Methods

[0175] Protein Sample Preparation

[0176] FGFR1 (100 pM) kinase domain construct was buffer-exchanged and concentrated using centrifugal concentrators (Corning Spin-X OF 30k MWCO) from a storage buffer of 40 mM N-(2-hydroxyethyl)- piperazine-N'-ethanesulfonic acid, pH 7.5, 200 mM NaCI, and 1 mM tris(2-carboxyethyl)phosphine to a final solution of 1 , 5 mM phosphate buffer, pH 8.0, and 5 mM NaCI. Lyophilized powder of BSA (heat shock fraction, Sigma-Aldrich) was dissolved in 5 mM phosphate buffer, pH 8.0, 5 mM NaCI, to achieve a 1 mM concentration. 2% / vol additions of SU-5402 solutions in dimethyl sulfoxide (DMSO) (or pure DMSO) were made to samples of the 1 mM protein solutions to yield final FGFR1 solutions containing 0, 0.5, 0.75, 1 , 1.5 and 2 mM SU-5402 and BSA solutions containing 0 and 1 mM SU-5402. For EW 2DIR experiments, 1 pL droplets for each protein solution were deposited onto a 1" 0 circular coverslip along with 8 2pL droplets of saturated NaCI solution. By mounting the 1" 0 circular coverslip in a 1" 0 lens mount followed by a 1 / 4" thick 1" 0 rubber 0-ring and a 2 mm thick 1" 0 CaF2 window, a sealed volume could be created between the coverslip and the CaF2 window with a controlled humidity of 75.5% RH, mediated by the saturated NaCI solution spots. The protein droplets dried slowly under these humidity-controlled conditions to form flat, hydrated gel-phase structures, similar to the coffee-ring structures previously reported to be formed by a similar approach.1The height of each gel spot was measured to be approximately 10 pm using a digital micrometer.

[0177] EEV 2DIR Spectroscopy

[0178] As EVV 2DIR spectroscopy is relatively unfamiliar, a brief description of its workings is provided here. The technique measures the coupling of a 1 quantum (fundamental vibration) transition to a 2 quantum (combination band or overtone) transition. If the combination band which is excited contains the 1 quantum transition which is also excited, then Raman scattering between the two levels is an allowed transition. The visible beam, therefore, effectively probes the coupling between the two excited vibrations, one of which is a vibrational fundamental frequency and one of which is a combination band or overtone. The spectra produced have columns of features corresponding to fundamental transitions which couple to many other vibrations. They also have diagonal features which are in effect lines, where the second 1 quantum transition, which is contained within the combination band, is constant. The frequency of the second vibration within the combination band is easily estimated by subtracting the frequency of the first excitation beam, coa, from that of the second excitation beam, cop. This assumes that the frequency shift in the combination band due to coupling itself is small or negligible, which is not always the case.

[0179] EVV 2DIR Spectrometer

[0180] The experimental setup of EW 2DIR spectroscopy, as performed in this example, has been described in detail elsewhere.2In brief, an ultrafast Thsapphire laser system (Newport Spectra-Physics) is used to generate a near-IR laser beam at 790 nm with a pulse duration of 1 ps and a repetition rate of 1 kHz. The generated beam is split into three arms, two of which drive optical parametric amplifiers (Newport Spectra-Physics) to produce tunable mid-IR 1 ps outputs calibrated using a 38 pm thick polystyrene calibration film traceable to NIST 1921 b frequencies. All three beams are linearly polarized parallel to the plane of propagation, spatially overlapped, and focused at the sample. The pulse energy of the 790 nm beam is set to 300 nJ at the sample, and the mid-IR beams display near symmetric energy profiles across an coa / 2Kc range from 1250 to 1750 cm1(Erna . = 800 nJ at 1500 cm1) and an cofi / 2Nc range from 2600 to 3400 cm-1(Em = 2.0 Z4 at 3000 cm'1). These pulse energies are measured through a 100 ym diameter pinhole as it closely matches the 1 / e2 diameter of the near-IR beam, which defines the sample area from which the four-wave mixing (FWM) signal is generated. Delay stages are used to achieve temporal control of the laser pulses. The temporal order of the pulses used was coo, — > cop — > coy, leading to the coherence pathway described in Figure 1 . The emitted FWM signal is detected using a photo-multiplier detector (Hamamatsu H7422-50), with the input beams being filtered out. The experiment is enclosed and purged with N2 to achieve a relative humidity below 3%. The EVV 2DIR spectra reported here were all obtained using pulse time delays of zoo = 1 .75 ps and TN = 1 ps. These time delays were chosen so as to provide maximum contrast for the most important features of the spectrum, whilst still keeping the nonresonant background low enough to be able to see the smaller, less-intense features clearly.

[0181] Spectra were collected with a mid-IR frequency step size of 5 cm-1and an acquisition rate of 100 laser shots per data point. The spectra had any mid-IR-independent background subtracted, were smoothed with a nearest neighbor average smoothing filter, and the spectral intensities were normalized by the intensity of the methylene peak. All spectra shown here are plotted on a logarithmic intensity scale. When producing difference spectra, the spectra were first square rooted, then subtracted, as the signal intensity is proportional to the square of the sample concentration.

[0182] 2D Fitting of the Spectrum

[0183] Fitting was performed in MATLAB (Registered Trade Mark) using a nonlinear least squares routine to fit the data with a sum of 2D Gaussians, restricted to having equal x and y variances. The fitting requires an initial input of the Gaussian parameters. Initial locations for the peaks were chosen manually. Initial signal intensities were taken to be the spectral intensity at the initial locations and the initial widths were taken to be the spectral width of the mid IR laser pulses (25 cm1).

[0184] DFT Calculations and EVV 2DIR Simulations

[0185] DFT calculations were performed using the 6-31 G(d,p) basis set using GAUSSIAN 09. The calculations of the EVV 2DIR peaks from the vibrations of the model were made in accordance with the method reported by Kwak et al.16 and applied in our previous work on model systems.3’45A frequency factor of 0.96 was used to produce the frequencies seen in the calculated EVV 2DIR spectra. The calculated peaks were dressed with a width of 25 cm1full width at half-maximum to correspond with the typical widths seen experimentally. The spectrum was calculated for the FGFR1 / SU5402 system using a model based on the cocrystal structure reported by Mohammadi et a1.33 and contained only the protein residues within 5 A of the SU5402 molecule. Methyl groups were added to terminate unsaturated parts of the model, and a water molecule was included which was within the 5 A range.

[0186] Results and Discussion

[0187] This example demonstrates the development of a set of protocols that enabled the identification of one or more EW 2DIR peaks that are definitively correlated with binding of a drug to a protein-active site. This requires identification of features that are not present when the protein or drug is measured independently or when the drug is mixed with a protein for which no specific binding site exists. As it proved impossible to generate a spectrum of SU-5402 in pure water because of aggregation of the drug, a background of a nonbinding control protein is used at a 1 :1 stoichiometric ratio as with the target protein. This allowed the spectrum of unbound SU5402 to be deduced.

[0188] In this example, the protein BSA was used as the control protein and difference spectra were created by subtracting the protein-only spectrum from that of the drug mixed with the target protein. This enabled the identification of peaks that are only present in the case of specific binding, as a means of identifying binding specific peaks. An alternative may be to compare the spectrum when it is in the presence of an organic solvent. Whilst there are drugs that can bind specifically to BSA, most do not, but instead attach themselves to multiple locations on the surface with low binding affinity. Heuristically, the finding here is that SU5402 falls into the category of nonspecific binding. An additional and independent method for assigning binding-specific peaks is to titrate the drug into the FGFR1 protein binding site and demonstrate the expected difference in titration curves between binding-dependent peaks and peaks that are independent of binding. This provides two independent methods for assigning peaks to the interaction between the drug and the target protein. It is noteworthy that both independent methods identify the same six peaks as being because of specific binding of SU5402 to FGFR, whilst the other 57 peaks due to the presence of the drug are presumably dominated by intramolecular couplings.

[0189] Finally, an interpretation of the binding-specific features is provided by using quantum mechanical calculations of the vibrational couplings in the binding site and showing how these correlate with the binding-dependent observations.

[0190] EVV 2DIR spectra were measured both for pure FGFR1 and containing 1 mM 5U5402.6These can be seen in Figure 1A,B, respectively, along with their difference spectra which were produced by subtracting the square root of the protein-only spectrum from the square root of the protein with the inhibitor spectrum. Roots have to be taken to make the spectra linear in concentration, as the signals are homodyne and are proportional to the square of the molecule number. The EVV 2DIR data shown here typically have a dynamic range of -450.

[0191] A combination of inspection and fitting reveals that the FGFR1 protein spectrum (Figure 1 A) contains - 200 peaks which have maximum amplitude at least 36 above the noise floor, where a is the standard deviation of the feature-free regions of the spectrum. The BSA protein spectrum contains a similar number of identifiable peaks. It is worth noting that this specification of statistical validity is reasonably stringent. Signal averaging within the area of a peak provides greater signal averaging than the peak maximum alone, therefore, it is likely that all of the peaks identified are real.

[0192] The protein spectra contain some features which have already been assigned in previous work.578For example, the two largest spectral features seen at coo, = 1470 cm Vcofi = 2945 cm-1and coo, = 1640 cm Vcofi = 3300 cm-1correspond to methyl / methylene and amide I bands, respectively. Features arising from phenylalanine side chains can be seen at coo, = 1470 cm Vcofi = 3050 cm-1and coo, = 1480 cm Vcofi = 3070 cm-1along with a tyrosine peak at coo, = 1530 cm Vcofi = 3130 cm1. The most relevant features for the purposes of this example, however, are the peaks that are visible in the SU5402 + FGFR1 spectrum but not detectable in the BSA + FGFR1 spectrum and the amide-associated features centered at -4650 / 3300.

[0193] The amide features provide an immediate indication of the action of the drug. The change in the amide region of the FGFR1 spectrum suggests that FGFR1 undergoes a degree of structural rearrangement which affects the amide features, and therefore, presumably elements of the protein backbone. In contrast, the absence of any changes to the BSA amide region upon addition of the drug indicates that the presence of SU5402 does not affect the BSA backbone structure at all, presumably because SU5402 does not bind to BSA at any specific site.

[0194] In general, addition of SU5402 to the proteins increases the number of detectable peaks substantially. After subtracting the spectrum 1 A (FGFR1 only) from 1 B (FGFR1 + 5U5402), a difference spectrum was produced which was fitted to a sum of 2D Gaussian features. Aside from the amide features, this revealed 63 features in the difference spectrum. Similarly, the spectra of BSA with and without SU5402 (1 D,E) were subtracted which produced a difference spectrum containing 57 Gaussian features. The crucial point here is that all 57 features in the BSA difference spectrum are also found in the FGFR1 difference spectrum, but there are an additional six peaks clearly seen in the case of the FGFR1 spectrum. The location of these features are marked on the difference spectra (Figure 1 C,F) with pink crosses.

[0195] The difference spectra can contain features from three sources. First, those spectral features arising from intramolecular vibrational coupling within SU5402; second, those arising from changes in the protein spectrum as a result of inhibitor binding-induced protein structural changes, such as those observed in the amide bands; and finally, as a result coupling between the protein and the drug.

[0196] The 57 peaks which are common to both FGFR1 and BSA difference spectra were assigned to intramolecular coupling within SU5402. Because BSA has no binding site for SU5402, this was a reasonable assumption. The changes in the broad amide region for FGFR1 are assigned to changes in the protein structure caused by drug binding while the additional six peaks in the FGFR1 spectrum (Figure 9) were tentatively assigned to changes caused by drug — protein interactions. It is not, in general, possible to definitively tell whether such new crosspeaks are due to coupling between drug modes and protein modes, new collective modes of the bound drug — protein complex, or changes in the protein structure caused by drug binding. Distinguishing these possibilities was the motivation for the QM calculations which are described later.

[0197] It was clear from inspecting the experimental spectra that five of the six peaks came from coupling of one fundamental vibration at 1400 cm-1to a range of different vibrations while the sixth was in the vicinity of the amide vibrations. To confirm the identification of these six peaks as binding-dependent, an additional independent measure of their binding-dependent behavior was performed. EVV 2DIR spectra of the FGFR1 / SU5402 mixture were measured across a range of molar ratios. At the high concentrations of the protein and drug used in these experiments, the expectation was that the intensity of truly binding-dependent peaks should not continue to increase linearly beyond a 1 :1 molar ratio, as the binding sites become saturated. Conversely, features that are caused simply by intramolecular couplings within SU5402 should continue to increase in intensity linearly with the concentration of the drug.

[0198] EVV 2DIR spectra of 1 mM FGFR1 and 0, 0.5, 0.75, 1 , 1.5, and 2 mM SU5402 were recorded, their difference spectra calculated, and the difference spectra fitted to extract peak intensities as described above. All six of the peaks present exclusively in the FGFR1 difference spectrum show a plateau in signal strength. In contrast, the 65 difference spectrum peaks common to both the FGFR1 / SU5402 and BSA / SU5402 difference spectra all displayed linear increases in intensity with the inhibitor: the protein ratio, as would be expected for features that are not associated with binding. This confirmed the nonbinding peaks as intramolecular modes of the drug and therefore proportional to the drug concentration only. The average fitted peak intensities plotted as a function of the inhibitor to protein molar ratio can be seen in Figure 2. The data are normalization to the 1 :1 fitted intensity for ease of comparison. The total number of proteins bound continues to increase beyond the 1 :1 molar ratio as is normal in drug — protein binding curves, as not all ligands are bound when the sample ratio is 1 :1. It is very clear from these data that the six peaks are not present in the BSA difference spectrum, and therefore, because of the presence of both FGFR and SU5402, all have distinctly different concentrationdependent behavior from the other 57 peaks, which provides a second form of verification, over and above their absence in the control measurements with BSA.

[0199] In order to assign the six binding-dependent cross-peaks to the coupled pairs of vibrations which gave rise to them, ab initio calculations were performed of the vibrational couplings of the drug in the binding site and compared with the isolated drug.

[0200] The structure of the protein — drug complex used for calculations is shown in Figure 3 below, the spectrum of SU5402 in free space was also calculated as a guide to the spectrum of SU5402 not bound into the protein-active site. Although this free-space calculation does not contain water molecules, the difference between the calculated spectra for bound and unbound species was expected to be instructive, which proved to be the case as can be seen below.

[0201] As EW 2DIR probes vibrational couplings between fundamental modes (mode #1 at coa) and their combinations (modes #1 + #2 at Wp), peaks with common mode #1 fall in columns of equal co, whilst peaks with common mode #2 fall in diagonal lines of gradient 1 . By following the location of these lines of mode commonality and by inspecting the motions of the calculated modes, corresponding columns of equivalent mode #1 and diagonals of equivalent mode #2 can be found between the two calculated spectra and the experimental difference spectrum of the FGFR1-SU5402 complex. Key lines of mode equivalence have been marked on Figure 4. Each vibrational mode can give rise to peaks in both a column and a diagonal, as it can participate as both the #1 and the #2 modes. An example of this is mode C marked with the magenta lines in Figure 4. The overtone peaks, in which a given mode forms both the #1 and #2 modes, fall on the intersection of the column and diagonal of a given mode and form a new diagonal with a 2:1 gradient and has been marked on Figure 4 as a dotted black line.

[0202] Calculated modes that contribute to the binding-dependent peaks are assigned numbers 1-7 to distinguish them from labelled experimental modes identified as A — E in Figure 4. Images of the modes showing the major atomic motions are found in Figures 5-11 .

[0203] It can also be seen from Figure 4, the column of equivalent mode D which appears in both the calculated spectra, cannot be seen in either of the two experimental spectra. To understand this, it is important to note that the method used to produce the calculated spectra does not take account of lifetimes of the states involved. In fact, with EW 2DIR spectroscopy, it is not unusual to find that some of the expected cross-couplings between vibrations cannot be seen experimentally. This is essentially because some of the state lifetimes are too short to be observed using a method with an effective time resolution of ps. In practice, this means that experimental spectra are usually slightly sparser than those calculated.

[0204] As can be seen from Figure a,c, the frequencies of several equivalent modes are slightly shifted between the two calculated spectra. These shifts range from 3 to 43 cm1such as mode A which appears at 1330.03 cnr1in the binding site calculation and at 1322.41 cm1in the free-space calculation. In contrast, no shift in frequency is observed between the bound- and non-specifically-bound experimental spectra. This can be attributed to weak interactions with the protein binding site slightly shifting the frequency of inhibitor modes compared to those of the free-space inhibitor molecule. The lack of shift between the bound- and non-specifically-bound experimental spectra can be attributed to the chemical environment of the inhibitor is more similar between the two than to free space.

[0205] The calculated binding site inhibitor complex spectra also includes peaks arising from protein modes, such as the amide 1 overtone peak seen at coo, = 1702.5 cm Vcofi = 3405.0 cm-1. The line of equivalent coo, for the amide 1 mode has been marked on Figure 4a, b as mode E. This amide 1 mode can be seen to be coupled to an inhibitor mode of calculated frequency 1558.62 cm-1as a peak at coo, = 1702.5 cm Vcofi = 3261.15 cm-1in Figure 4a. This inhibitor mode gives rise to a diagonal of peaks which has been marked on Figure 4 as mode C. Mode C corresponds to an amide-associated motion of the 5- membered lactam ring with some motion on the adjoined benzene ring. The experimental bindingdependent peak at coo, = 1660 cm Vcofi = 3240 cm1can be assigned to the coupling of this lactam ring mode, mode C, to the protein amide 1 mode, mode E, and is the equivalent of the calculated peak at 1702.5 cm Vcofi = 3261.15 cm'1. This vibrational coupling is mediated by a hydrogen bond formed between the NH group of the SU5402 lactam ring and the amide carbonyl of the proximal amino acid of the binding site. Calculated equivalents of the five experimental binding sensitive peaks at coo, = 1400 cm-1were observed in both the calculated spectra, with and without the inclusion of the protein binding site. These are marked as mode B in Figure 4. The vibrational modes calculated to give rise to them exhibited motion exclusively on the SU5402 molecule. The apparent experimental binding dependence of the intensity of these peaks must therefore be attributed to lifetime effects. In the non-bound case, the SU5402 molecules are in a different chemical environment to those bound in the kinase binding site, and it is perfectly reasonable that they are exposed to different environmental fluctuations. This can easily result in increased dephasing and decreased lifetimes of these vibrational coherences in the non- specifically bound case, compared to that of the protein inhibitor complex. Mode B corresponds to a complex motion involving atom displacements across the entire molecule from the fused benzene ring to the methyl substituent of the pyrrole ring. It was speculated that the extensive nature of the mode which runs across many chemical groups provides more opportunities for "environmental" coupling (e.g., coupling to water) to influence coherence lifetimes through either pure or population dephasing. However, given that most of the modes here are spread over many chemical groups, it was considered somewhat surprising that more modes are not affected in a similar way.

[0206] Figures 5-11 show the main atomic displacements. In summary, the 1400 cm-1mode labeled B in Figure 4 and corresponding to mode 1 shown in Figure 5 is entirely located on SU5402, supporting the idea that the appearance of this column of cross-peaks upon binding is due to a change of the dephasing rate for this mode because of sequestration of SU5402 in the binding site. This was presumed to be due to exclusion of a solvent interaction which reduces the state lifetime. All other modes are collective motions of the drug and the protein binding site and show substantial motion of amide groups.

[0207] In some ways, the most striking feature of the data is that none of the intramolecular SU5402 modes, with the exception of the amide-associated cross-peaks, are lost or show substantial frequency shifts when the drug is bound to FGFR1 , compared with the nonbound situation with BSA. This suggested that the average structure of the solvated drug is very similar, if not more or less identical, to the structure of the drug in the binding site. This would certainly reduce the conformational entropy cost of binding and make the drug a more potent inhibitor and is thought to be a fairly common feature of heterocyclic aromatic drugs which tend to have relatively rigid planar components to their structures.

[0208] There was a substantial change in the amide region of the FGFR spectrum when SU5402 was added but no change at all to the amide region of BSA upon addition of the drug. This further highlighted the utility of BSA as a nonbinding control and suggested that a subset of FGFR structures is stabilized by the binding of SU5402. There is a good precedent for this. Many kinases contain a regulatory loop comprising a Asp-Phe-Gly sequence known as the DFG loop. This loop can either be in the "in" form, within the protein where the phenylalanine resides deep within a hydrophobic pocket, or the "out" form whereby it protrudes into the surrounding environment. SU5402 is a type 1 kinase inhibitor which binds to and stabilizes only to the DFG in form, and this could indeed lead to a net change in the protein structure upon binding. It is not possible to determine from the data whether it is in fact the DFG transition that is being observed or whether some other structural variation is stabilized by SU5402 binding. As proteins do explore a range of structural configurations at equilibrium, observing a change in this structural equilibrium upon drug binding is expected, but these observations do highlight the sensitivity of EW 2DIR to global structural changes in a protein as well as local variations in coupling because of the drug — protein contacts.

[0209] One of the binding-dependent peaks was more or less unequivocally assigned to a drug — protein contact. This is the feature found experimentally at coo, = 1660 cm Vcofi = 3240 cm'1. Its proximity to the amide 1 mode suggested that it is likely to be due to the known hydrogen bonding between SU5402 and a carbonyl of the protein backbone. This intuitive assignment was supported by quantum mechanical calculations which predicted the presence of this feature. It is caused by coupling between a mode of the drug and a protein mode which contains considerable amplitude of motion of the carbonyl in the binding site (see Figure 12), a chemical group which is known to hydrogen bond to the drug molecule. As this is a feature common to many kinase — drug interactions where the drug targets the ATP binding site, it was anticipated that this feature would be a general and generic means to identify and monitor the details of drug binding to these kinase-active sites.

[0210] In picking the most congested spectral region available to EVV 2DIR spectroscopy, namely, the region containing the overtone couplings, a particularly difficult challenge was set to the separation of bindingdependent peaks from binding-independent features. The reasoning here is that if it can be done for such congested spectra, it can certainly be done for less-congested spectral regions. Despite the relatively high levels of congestion from — 200 protein peaks and — 60 drug peaks, the spectra are both sufficiently decongested and sufficiently precise to allow six binding-dependent features to be identified. The spectral region measured here is approximately 1 / 20th of the total spectral region accessible by the EVV 2DIR apparatus used in this instance. It is, therefore, extremely likely that a wider survey of the total accessible spectral space will yield other binding-dependent features too. Indeed, calculations suggest that there are at least 12 more binding-dependent peaks present in the wider accessible spectral space, arising from vibrational couplings between SU5402 and FGFR1 binding site modes. There will presumably also be a significant number of features which experimentally display binding dependence, such as changes in the global protein structure in response to drug binding. As only the vibrational coupling in the drug binding site were calculated, these were not predicable using the approach employed in this instance.

[0211] The ability of EW 2DIR demonstrated here to detect the binding of a drug to a binding site of a target protein of commercial interest bodes well for the utility of the technique in drug discovery in future. The spectra shown here were the result of signal averaging for 90 min. The spectrometer was driven by a 1 kHz laser engine and data were collected at this 1 kHz rate. Laser engines capable of driving an EVV 2DIR spectrometer at 100 kHz now exist which would potentially reduce the collection time for an entire spectrum to approximately 1 min. Moreover, if only a part of the spectrum is required for a binding study, such as the amide region, the speed of collection for a particular drug — protein combination could be reduced to a few seconds. High throughput versions of this experiment would be entirely possible to establish and the high data density produced would enable mining of the data into "contact" modes able to be assigned to classes of protein — drug contact, even without the requirement for detailed structural analysis.

[0212] The prognosis for direct structural analysis is also good. The inventors have previously shown that collecting data at two different polarizations of the beams allows the geometry of the mode interaction to be determined.59As each binding-dependent peak provides one angular constraint, it would only require a few such features to provide quite tight structural constraints for the drug — protein interaction geometry. Similarly, it was shown that the intensity of these peaks can be used to determine the distances between atomic contacts, providing further structural detail.

[0213] Conclusions

[0214] The inventors have shown that it is possible to detect the binding of a drug to a clinically important drug target, namely, the protein kinase FGFR1. Of the six binding-dependent peaks, one can be assigned directly to coupling caused by the drug binding to a carbonyl in the protein-active site. This assignment is supported by ab initio calculations of vibrational couplings in the enzyme-active site when occupied by the drug molecule.

[0215] Example 2: Use of Electron-Vibration-Vibration Two-Dimensional Infrared (EVV 2DIR) Spectroscopy to Analyse and Interpret the Transition from Unfolded to Folded G-Quadruplex DNA.

[0216] In this example we show that it is possible to distinguish between two different structural hypotheses for the geometry of interacting chemical entities, using a combination of EW 2DIR spectroscopy and calculation of the EVV 2DIR spectra for the two hypotheses. The two structural hypotheses being tested against each other here are in one case that guanine nucleotides are bound together and in the other that the guanine nucleotides are not bound to each other. As the presence of potassium ions can cause a structural change in this sample, these two hypotheses are tested in samples containing a range of different ion concentrations in order to determine at what potassium concentration guanine complexes are formed.

[0217] G-quadruplexes (G4) are non-canonical DNA secondary structures which can form in guanine-rich DNA sequences.10 11G4s can be unimolecular, bimolecular or tetramolecular, and structures vary in relative strand direction, loop length and composition, and glycosidic bond rotation.11Over 700,000 G4 sequences have been identified in the human genome using next-generation genome sequencing.12Both experimental methods12and computational algorithms13 14suggest a high percentage of G4- forming sequences are found in gene promoter regions and telomeres. The formation of G4 structures in these regions is thought to inhibit the growth of cancer cells by inhibiting telomere maintenance15 16 and impeding the transcription of oncogenes.17 18The use of G4-stabilising ligands is therefore being explored as a potential molecular therapeutic approach to the suppression and elimination of cancerous cells. As such, structural insight into G4s and G4-ligand complexes is useful to assist in ligand design.

[0218] In this example, EVV 2DIR spectroscopy was used in combination with quantum chemistry calculations based on density functional theory, to study the structural changes that occur in Myc2345 due to the addition of potassium ions, which shifts the position of thermodynamic equilibrium from an unfolded to a folded G4. The 22-mer DNA sequence Myc2345 forms an intramolecular, propeller-type, parallel- stranded G4 in the presence of potassium ions.19

[0219] Methods

[0220] EVV 2DIR Spectroscopy

[0221] The setup comprised a dual Thsapphire amplifier (Thales Laser), synchronised by a single oscillator seed source (Femto Laser), producing two 800 nm, 10 kHz, 0.8 mJ beams with FWHM pulse durations and bandwidths of 1 ps, 20 cm1(a ) and 50 fs, 300 cm1(ft ). Optical parametric amplifiers (4W pump power, Light Conversion TOPAS) were used to tune the central frequency of the beams: )awas varied from 1250-1750 cm-1in step sizes of 5 cm-1, and a>p was set to 3200 cm-1. The y beam was produced by running an arm of the fs beam through an etalon, producing 800 nm, 1 ps pulses with a bandwidth of 12 cm1. All three beams were focused on the sample in a geometry that satisfied phase matching conditions. Time delays between each pulse were created using delay stages, and were set to Tap = Tpy= 0.4 ps. The FWM signal was detected by a combination of a spectrograph and a charge coupled device camera.

[0222] Sample Preparation

[0223] Myc2345, d(5'-TGAGGGTGGGGAGGGTGGGGAA-3' - SEQ ID NO: 2), was purchased from Eurogentec and used without further purification. Myc2345 was dissolved in a Bis-Tris buffer solution in Nanopure water (10 mM, pH 7.1) to a concentration of 1 mM. KCI was added to the sample at concentrations of 0, 5, 10 and 100 mM. The sample solution was vortexed for 10 s, annealed at 95 °C for 2 m and cooled to room temperature. The sample solution was formed into gel spots by depositing 1 pL on a glass coverslip and allowing it to dry in a sealed sample cell of 85% relative humidity, created by depositing 6 x 2 pL spots of a saturated aqueous solution of KCI. The sample cell comprised of the glass coverslip, a %” thick spacer and a 2 mm thick CaF2 disc as the front window. As the spots shrink when they form the gel phase, the final potassium concentration is higher than that of the initial solution.

[0224] Data Processing

[0225] Glass: A spectrum of the glass coverslip measured at time delays of Tap = Tpy= 0 was used to normalise the sample spectra, to account for varying OPA output energies across the wavelength ranges. The glass spectrum was smoothed by averaging over the nearest neighbours. Interference fringes were observed in the a)p - a>aaxis of the spectrum. To remove the fringes, linear interpolation was applied along the a>p - a>aaxis, maintaining the same number of data points. A low-pass brick wall filter was applied to the Fourier transform of the subsequent glass spectrum, using a cut-off frequency of 0.05. Line cuts of the glass spectrum after each step are shown in Figure 31.

[0226] DNA sample spectra: The sample spectra were smoothed by averaging over the nearest neighbours. Linear interpolation was applied along the a>p - a>aaxis to match the glass spectrum. A background subtraction was applied to the spectra by subtracting the mean intensity in the region a>a= 1325-1375 cm'1, a>p - a>a= 1190-1210 cm1. The spectra were then background subtracted and normalised on the glass spectrum. As the FWM signal is proportional to the square of the analyte, the intensity of the spectra was square-rooted. The spectra were also normalised at the highest intensity peak at a>a= 1525 cm1, a>p - a>a= 1450 cm1. Each sample spectrum is the average of three measurements.

[0227] 2D Gaussian Fitting

[0228] To estimate the number of peaks in the experimental 2D spectra, line-cuts at different frequencies were analysed. This provided a reasonable estimate for the number of Gaussian functions (102) to be used for the 2D Gaussian fitting. Each of the spectra were fitted with a sum of 102 2D Gaussian functions using the non-linear least squares fitting function scipy.optimise.curve_fit on Python. The Trust Region Reflective algorithm was used to perform the minimisation of the Chi-squared value. Errors of the parameters of the fit were propagated from the residuals of the best fit. The 100 mM spectrum was first fitted with no fixed parameters. The initial guess locations were based on visual inspection of line cuts of the spectrum, and these were allowed to vary by ±15 cm1. The initial guess intensities were randomly generated numbers between 0-1 and were allowed to vary within the same range. The initial guess widths were 25 cm-1for all peaks in both dimensions, and these were allowed to vary between 5-40 cm-1. Each of the sample spectra were then fitted using the optimised parameters from the initial 100 mM fit as the initial guesses. All parameters were fixed, apart from intensities, which were allowed to vary from 0-1 , and the location of peaks which were observed to shift in frequency with respect to potassium ion concentration in the line cuts of the spectra. These shifting peaks were allowed to vary by ±30 cm-1in the dimension the shift is observed.

[0229] EVV 2DIR Spectra Calculations

[0230] The computational EVV 2DIR spectra were obtained via a method based on the Gaussian software (Gaussian 16, Revision B.01).

[0231] The relevant contributions to the third-order nonlinear susceptibility(3), which is the central quantity for the EVV 2DIR spectral intensities, can by a perturbation-theory development truncated at the first order of anharmonicity in the involved polarization properties (called electrical anharmonicity) and vibrational potential (called mechanical anharmonicity), be expressed as the sum(3)= x^ °f electrical and mechanical anharmonicity contributions ®and respectively. In the present work, the contribution from was ignored and only x^ was considered, which is believed to still be sufficient for observations of a qualitative nature. The evaluation of involves calculating the relevant vibrational energy levels (of singly and double excited states) and, under the excitation pattern considered, the normal-mode geometric first- and second-order derivatives of the molecular dipole moment / z, and first-order derivatives of the molecular polarizability, a (Q, and Qj refer to normal modes).

[0232] In vacuo G-quartet and G-quadruplex molecular structures were first prepared using PyMol and then optimized with Gaussian at the density functional theory (DFT) level of theory using the B3LYP functional, and using the 6-31 G(d,p) basis set, employing the I nt=UltraFine grid granularity parameter and SCF=Very Tight convergence criterion. The optimized G-quartet and G-quadruplex structures were subsequently used with Gaussian, employing the same level of theory and basis set as for the geometry optimization, to calculate harmonic vibrational normal mode frequencies ec^and the dipole moment and polarizability and the first-order normal-mode derivative of the dipole moment, using numerical differentiation with in-house scripts of the dipole moment, dipole moment first derivative, and polarizability calculated at normal-mode displaced geometries to obtain the derivatives described above for the EVV 2DIR spectral intensities. A harmonic regime was assumed for the vibrational energy levels with a uniform scaling parameter of 0.96 , thus resulting in no difference between the sum + a>j of singly excited energy levels and the combination / overtone energy level OJL +]involving the same normal mode indices. The resulting data was then processed using in-house scripts into two-dimensional spectral intensities, each peak being dressed with 2D Gaussian lineshapes with a linewidth of 20 cm1(FWHM), and subsequently rendered using the Matplotlib Python library.

[0233] Results

[0234] The experimental 2DIR spectra of 1 mM Myc2345 in the presence of 0, 5,10 and 100 mM K+are presented in Figure 16. In EVV 2DIR spectroscopy, one infrared pulse excites one vibrational band while the other excites another. The only situation considered was where, in a normal-mode factorization of the vibrational wavefunction, the first pulse induces a one-quantum excitation to a singly excited state and the second induces a two-quantum excitation to a doubly excited state in which at least one of the mode indices corresponds to that of the singly excited state. The coupling was read-out by a visible pulse which Raman-scatters from the coherence if the two excited bands are coupled and does not scatter if they are not. The measured four-wave mixing signals were plotted with the Raman excitation frequency subtracted, yielding a difference frequency axis a>p - o>a(Y-axis) as a function of a>a(X- axis). As a>p excites a combination band, the two axes approximately correspond to the frequencies of the two fundamental vibrations which contribute to the combination band. This leads to rows and columns of features being observable in the data. These correspond to a particular vibrational mode, coupling to multiple other modes. Because the coupling density is high, multiple coupling peaks overlap in these spectra which means that peak positions do not necessarily represent the positions of couplings but instead are sometimes the peaks caused by the presence of multiple couplings. In order to disentangle these complex spectra as much as possible, the 2DIR spectra were fitted with two-dimensional Gaussian functions, which provide a good approximation to the shape of typical EW 2DIR coupling feature. In this case, a minimum of 102 2D Gaussian functions were required, revealing the presence of at least 102 distinct couplings. The fitted 2DIR spectra are presented in Figure 17, with superimposed markers denoting cross-peak central frequencies from the fit. As expected, these did not always perfectly coincide with the peaks of the raw, unfitted data. Vertical line-cuts comparing raw and fitted spectra are presented in Figure 18 which demonstrates the high signal to noise ratio of this data set and the goodness of the fit. Comparisons of raw and fitted data at other frequencies are shown in Figure 25 along with Chi-squared values. In this section the focus is on the frequency shifts, intensity changes and new vibrational couplings that occur as a result of structural changes on K+ion addition.

[0235] It was observed that multiple peaks in the region of in-plane base vibrations at around ~ 1700 cm1redshift with increasing K+concentration, as shown in Figures 19a, b. Peaks 12, 30, and 58 red-shift along the a>p - o)aaxis; peaks 74, 77, 85 and 87 red-shift along the u>aaxis; and peak 80 red-shifts along both axes. The observed red shifts are consistent with the weakening of the guanine carbonyl bonds as Hoogsteen hydrogen bond formation and K+coordination occurs. Interestingly, the - a>afrequency representing guanine carbonyl stretching appeared at a higher frequency than the a>afrequency. In the absence of potassium ions, the a>p - a>afrequency associated with guanine carbonyl stretching was in the range 1694-1707 cm1, whereas the a>aguanine carbonyl stretching frequency was observed between 1671-1688 cm1. Interpretation of this highly unusual observation can be found in the Discussion section. Additionally, the guanine carbonyl stretching frequencies observed in the absence of potassium ions was higher than the single strand guanine frequency typically observed between 1660-1673 cm'1 20, and this observation hinted at the possibility of some pre-existing structure in the absence of cations.

[0236] It is well known that nucleic acids exhibit changes in absorbance (hypochromism and hyperchromism) in the UVA / is region of the spectrum, which is thought to be caused by base stacking interactions.21 23Hypochromism and hyperchromism are reflected in the Raman line intensities of ring modes, which has previously been used to monitor conformational changes in nucleic acids.24 30In the EW 2DIR spectra, peaks which exhibit a monotonic increase / decrease in intensity and those which also show a minimum intensity change of 20% between the 0 K+and 100 mM K+spectra were focused on. 29 peaks which decrease in intensity as K+concentration increases were identified (Figure 19d-i). The majority of these peaks appeared at frequencies characteristic of guanine ring vibrations (1476-1495 cm4, 1564-1568 cm4and 1575-1590 cm4).33Additionally, five peaks were observed to increase in intensity as K+concentration increases (Figure 19c). Assignment of these features and a description of the mechanisms for coupling increases and decreases are given in the discussion. The intensity traces of the remaining peaks as a function of K+ion concentration are shown in Figure 26.

[0237] The folding of G4 from a relatively disordered to a relatively ordered structure would be expected to produce a few new coupling peaks as chemical groups previously separated are brought closer together. Figures 20c and 20d present the raw and fitted 2DIR spectra in the region of the observed new peaks 101 (o)a= 1705 cm1, - o)a= 1675 cm1) and 92 ( ua= 1691 cm1, a>p - )a= 1132 cm1), respectively. Figure 20b presents the intensity - ion concentration dependence of both peaks. New vibrational modes extending over multiple G-quartets, and therefore having greater levels of delocalisation, can also result in the appearance of new peaks. These peaks may serve as markers of G4 formation. Assignments of these peaks are discussed in the Discussion section.

[0238] Discussion

[0239] One of the initially surprising aspects of the experimental data was that relatively few completely new couplings were created by the folding of the G4. Although behaviour like this may be observed in the calculated spectra, such as in Figures 21 a and 21 b, it is useful to reflect on the reasons for this. Chemical moieties such as DNA bases, contain many chemical ‘groups’ which couple to each other in a diverse set of ways. The formation of a quartet of bases coupled by Hoogsteen hydrogen bonding might therefore primarily be characterised by the reorganisation / recombination of the vibrational states of the individual bases (and the couplings between them) rather than the emergence of altogether new couplings not characterisable in this way. This is loosely analogous to the situation in molecular exciton coupling, where electronic states are reorganised by the new couplings, rather than the number of electronic states being increased.

[0240] One of the models for the formation of the G4 is the transition between two states: a single strand in the absence of ions and a fully folded G4 in the presence of an excess concentration of potassium ions. However, the frequencies associated with guanine carbonyl stretching in the EVV 2DIR spectrum of Myc2345 in the absence of ions are not characteristic of that of guanine in a single strand.28Nonetheless, the formation of a higher order structure in response to the addition of potassium ions is indicated by a number of observations, namely, the red-shift of the guanine carbonyl stretch frequencies, the observed changes in cross-peak intensities and the appearance of two new couplings. To gain a better understanding of the structure observed in the absence of ions and the observed spectral changes upon increasing ion concentration, EW 2DIR spectra were calculated for the G4 and its possible precursors. Furthermore, a uniform scaling parameter was employed for the vibrational energy levels, which were calculated harmonic vibrational potential approximation, and thus any difference between overtone / combination band energy levels and the corresponding sum of one-quantum energy levels were not considered. Finally, the calculations were carried out in vacuo, not involving any implicit or explicit solvation while the experimental sample existed in the presence of abundant water. Altogether, the level of confidence in the findings associated with the analysis of the quantum-chemical calculations in this work should therefore be considered with these limitations in mind. The overall strategy in this work was to identify and distinguish three possible structure types. These were isolated guanines which were not part of a quartet; guanines arranged in a quartet and guanine quartets stacked to form a G- quadruplex. The Initial DNA Structure Contains Guanines in a Structure Resembling a G-Quartet in the Absence of Potassium Ions

[0241] Multiple pieces of evidence pointed towards the presence of pre-existing DNA structure in the absence of ions. The first evidence came from the observation of the frequency difference between the row and corresponding column of peaks arising from the coupling of a guanine carbonyl stretch mode with various other vibrational modes (Figure 16a). In general, for EVV 2DIR spectroscopy, the rows and columns of cross-peaks relating to the same mode, have an approximate reflective symmetry about the overtone line = a>p - a>aHowever, this symmetry is frequently broken by a negative frequency shift along the a>p - a>aaxis. This is because a combination band is usually shifted to a frequency lower than the sum of the fundamental frequencies by the process of mode-coupling. In the experimental spectrum (Figure 16a), the row of cross-peaks with frequencies which depend on potassium ion concentration were assigned to the guanine carbonyl stretch with a frequency between 1694 - 1707 cm1. This frequency was unusually higher than the corresponding column of peaks with frequencies that are also dependent on potassium ion concentration and appeared between 1671-1688 cm1. Assuming that these were not combination bands that were shifted to a higher frequency than the sum of the constituent fundamentals, this indicated that the guanine carbonyl stretch mode giving rise to the row, referred to as GCO I, was not the same one which gives rise to the column, referred to as GCO II.

[0242] An occurrence of two different guanine carbonyl modes giving rise to what is normally a corresponding row and column of cross-peaks is impossible to reconcile with the spectroscopy of an isolated guanine base (Figure 27). The calculated EW 2DIR spectra and the associated vibrational analysis do, however, exhibit a comparable behaviour for the structures involving coupled guanine bases as in a G-quartet (Figure 21 a) or a G4 (Figure 21 b). The reason for this is that the roughly symmetric G-quartet structure gives rise to multiple combinations of guanine stretch modes in a manner analogous to simple exciton coupling. Although many different guanine carbonyl centric modes contribute to the column of crosspeaks, four in the case of the G-quartet and 12 in the case of the G4 pictured in Figure 21 e, the coupling strength of the cross-peaks associated with each carbonyl mode differ and one mode has much stronger couplings than the others. Interestingly, the dominating carbonyl stretch mode in the row of cross-peaks, is not the same mode that dominates the corresponding column of cross-peaks. Thus, in these calculated spectra, the rows and columns actually originate from different combinations of guanine motions. While the relative difference between the central frequency of the dominating row and column are different and opposite in sign between the experimental and our calculated spectra, this observation can be taken as support for the notion that the observation of differing guanine stretch modes in the rows and columns of the EW 2DIR spectrum of Myc2345 in the absence of ions, can be due to the presence of a pre-existing structure involving coupled guanine bases. In effect, EVV 2DIR picks out differences in oscillator strength between the one-quantum states of the G-quartet, with one being excited first in the row and another in the corresponding column. As the ‘excitonic’ interactions of the original isolated guanine motions, can result in one quantum transitions of different oscillator strength when they combine ‘excitonically’, then one can have the situation where one state dominates the column whilst another dominates the row. Further evidence for the presence of pre-existing structure in the absence of ions came from the absolute frequencies of the guanine carbonyl stretches. The frequency of nitrogenous base carbonyl stretches is sensitive to DNA conformation and is often used to determine DNA secondary structure. The guanine carbonyl stretching frequency was observed at u>a= 1671 - 1688 cm-1(GCO II), which is higher than a single strand guanine typically observed between 1660-1673 cm1,26which could suggest the presence of a pre-existing DNA Structure. In fact, circular dichroism measurements of MycL1 , a similar G4-forming sequence to Myc2345 (see Figure 28 for sequence comparison), also suggest that some structure can be present in a G4 DNA sequence, even in the absence of potassium ions.27

[0243] Given the above indications that some structure resembling a G-quartet was present, even in the absence of potassium ions, the obvious question was whether the entire G4 structure may be preformed. Evidence for the absence of a fully stacked G4 complex in the absence of potassium ions came from four observations. Firstly, in the geometry optimization, a stacked quartet without potassium ions was not found to be energetically stable. This structure is unfavourable due to the repulsive forces between the 12 carbonyl oxygens. As this is a purely enthalpic effect and as entropic forces will only drive the structure apart in a real solvated system, the presence of a G4-like structure in the absence of potassium ions was therefore considered to be unlikely.31The three other pieces of evidence for the absence of stacked G-quartets (the entire G4 structure) in the absence of potassium ions were the redshifts of particular frequencies observed when the potassium ions were added, the intensity changes for multiple peaks upon the addition of potassium and the formation of new spectral features already known to be associated only with fully stacked G4. These are discussed in more detail below.

[0244] The Red Shift of Guanine Carbonyl Stretches Indicate Ion Coordination

[0245] G4 formation has previously been monitored by noting a frequency shift in the guanine carbonyl stretch frequency.32In this example, a red shift was observed in the frequency of both guanine carbonyl stretch modes, GCO I (Figure 19a) and GCO II (Figure 19b), upon the addition of potassium ions. Peak 80, which is a coupling between GCO I and GCO II, shifted diagonally which supported these assignments. According to the existing literature however, the folding of a single strand to a G4 induces a frequency shift of the guanine carbonyl stretch in the opposite direction.3233In order to reconcile this discrepancy and in order to understand the experimental observations, calculated EW 2DIR spectra of the proposed initial structure of a G-quartet (Figure 21 a) and a G4 (Figure 21 b) were compared to experimental spectra of Myc2345 in the absence and presence of potassium ions.

[0246] The experimentally determined red-shifts in the guanine carbonyl stretch frequencies caused by the addition of potassium ions are shown in Figures 19a,b. The red-shift in the GCO I and GCO II frequencies were approximately reproduced in the calculated difference spectra between a G-quartet (Figure 21 a) and a G4 (Figure 21 b). Experimentally, the average frequency shifts in GCO I and GCO II were -8 cm1and -10 cm1, whilst the calculated frequency shifts were -2 cm1and -25 cm1, respectively. Therefore, the calculations, keeping in mind that they were based on harmonic vibrational wavefunctions, qualitatively agreed with experimental observations, and suggested that potassium ion complexation caused a red shift overall. The mechanism of the red-shift of the guanine carbonyl stretch frequencies can be attributed to the ion coordination to the carbonyl oxygen and the stacking of G- quartets. The ion coordination results in the lengthening of the carbonyl bonds from 1.23 A to 1.25 A. Moreover, the resultant stacking of G-quartets, results in the delocalisation of the guanine carbonyl modes found in the G-quartet at 1704 cm1(Figure 21 g,h) across all three quartets in the equivalent mode in the G4 found at 1679 cm'1.

[0247] Changes in Cross-Peak Intensity Indicate Preformed G-Quartet without Full G4 Structure

[0248] It was experimentally found that a large number of cross-peaks decrease in intensity upon the addition of potassium ions. This observation can also be used to test structural models for the folding as these reductions in intensities would be manifested in calculations. The measured ratio spectrum caused by the addition of potassium ions and the calculated G-quadruplex / G-quartet ratio spectrum are shown in Figure 22 and have comparable features. Excluding the difference peaks due to frequency shifts and new couplings, both spectra show a majority of negative differences, with some positive difference peaks near or below u>a= 1600 cm1. This overall appearance of reduction in cross-peak intensities, although involving a comparison to calculated spectra based only on electrical anharmonicity, is thus consistent with the G-quartet to G-quadruplex model.

[0249] The changes in cross-peak intensity observed going from G-quartet (Figure 21 c,d) to G4 (Figure 21 e) can be explained by the change in relative orientation of the guanine bases of a quartet. In the G-quartet, molecular planes of adjacent guanine bases intersect at an angle of 21°. In the G4, this angle is altered to 19° in the terminal quartets and 0° in the central, planar quartet (Figure 30). These structural changes result in changes in the orientation of transition dipole moments, which account for the changes in intensity observed in cross-peaks across the spectrum.

[0250] Peaks 92 and 101 are Markers of a Fully Folded G4 in Myc2345

[0251] New cross peaks which appear as the potassium ion concentration is increased are potential markers of a fully folded G4. Two new peaks formed on potassium addition were identified: Peak 92 = 1691 cm'1, c p - o)a= 1132 cm'1) and Peak 101 (a>u= 1705 cm'1, a)p - o)a= 1675 cm'1). Neither of the peaks appeared in the calculated G4 spectrum, which suggested that both peaks could involve vibrational modes outside of the core G4 structural elements included in the calculation (Figure 21 e). Using the calculated vibrational frequencies of the G4 structure (Figure 21 e), the a>p - a>afrequencies of both peaks were tentatively assigned to vibrations of the core G4 structure: Namely, guanine ring deformations of terminal quartets in peak 92 and NH bends and C=O stretches of all quartets in peak 101. The d)afrequency of both peaks was characteristic of a G Ce=O6 stretch or a T C2=O2 stretch, which could be located within G / T in the end caps of Myc2345 (Figure 23) that were not included in the calculation. It is already proposed by other work that bases in the end caps of G4s interact with the core G4 structure, further stabilising it. The NMR structure of Myc2345 (PDB: 7KBV) is consistent with this idea.34 Summary of the comparisons between measured and computed spectra

[0252] Figure 22b shows the calculated ratio spectrum for G4 / G-quartet. This was compared with the experimentally measured ratio spectrum induced by the addition of 100 mM potassium ions (Figure 22a). Qualitatively, a number of key elements stood out as present in both calculated and measured spectra. Firstly, there was the pattern of a negative column at -1600 cm-1 next to a lower-frequency positive column. Secondly there was the absence of such a pattern in the corresponding rows, due to the mode combination effects outlined in Figure 21. Thirdly, there was the general reduction in crosspeak intensities in both calculated and measured spectra. Based on these observations, it was concluded that there is evidence to suggest that Myc2345 forms unstacked G-quartets in the absence of ions, and Figure 23 depicts how such a structure may look. In this model, on the addition of potassium ions, the G-quartets stack together to form a different structure presumed to be a G-quadruplex structure and based on the appearance of the two new coupling peaks, it is furthermore plausible that in such a structure, the guanine and / or thymine bases in the end caps of the DNA sequence couple to the stacked quartets.

[0253] Conclusions

[0254] The behaviour of 102 distinct and identifiable vibrational couplings was tracked during the process of Myc2345 G-quadruplex folding, using EVV 2DIR spectroscopy, by titrating these features as a function of potassium ion concentration. The initial 102 peaks were modulated during folding in a number of ways including changes in intensity and frequency shifts plus the formation of two new and additional coupling peaks. Assignments of spectral features and explanations for the cause of changes to spectral features were made with the aid of quantum-chemical (density functional theory) vibrational analysis and calculations of 2DIR spectra. Altogether, it was found that there is evidence to suggest that Myc2345 forms unstacked G-quartet-resembling structures in the absence of ions, and that a different structure, presumed to be a G-quadruplex, forms upon potassium ion addition. More generally, it was demonstrated that it is possible to obtain and interpret details about short range couplings even in the absence of long-range order. Secondly, it was shown that a high density of interpretable features can be obtained using EW 2DIR spectroscopy with high fidelity.

[0255] References

[0256] 1. Mampallil, D.; et al. Acoustic suppression of the coffee-ring effect. Soft Matter 2015, 11 , 7207- 7213.

[0257] 2. Petti, M. K; Lomont, J. P.; Maj, M.; Zanni, M. T. Two-Dimensional Spectroscopy Is Being Used to Address Core Scientific Questions in Biology and Materials Science. J. Phys. Chem. B 2018, 122, 1771-1780.

[0258] 3. Donaldson, P. M.; et aL Direct identification and decongestion of Fermi resonances by control of pulse time ordering in twodimensional I R spectroscopy. J. Chem. Phys. 2007, 127, 114513.

[0259] 4. Fournier, F.; et al. Biological and Biomedical Applications of Two-Dimensional Vibrational Spectroscopy: Proteomics, Imaging, and Structural Analysis. Acc. Chem. Res. 2009, 42, 1322- 1331 .

[0260] 5. Guo, it; et al. Detection of complex formation and determination of intermolecular geometry through electrical anharmonic coupling of molecular vibrations using electron-vibration-vibration two-dimensional infrared spectroscopy. Phys. Chem. Chem. Phys. 2009, 11 , 8417-8421 .

[0261] 6. Klein, T.; et aL Structural and dynamic insights into the energetics of activation loop rearrangement in FGFR1 kinase. Nat. Commun. 2015, 6, 7877.

[0262] 7. Fournier, F.; et al. Optical fingerprinting of peptides using twodimensional infrared spectroscopy: Proof of principle. Anal. Biochem. 2008, 374, 358-365.

[0263] 8. Fournier, F.; et al. Protein identification and quantification by two-dimensional infrared spectroscopy: Implications for an all-optical proteomic platform. Proc. Natl. Acad. Sci. U.S.A. 2008, 105, 15352-15357.

[0264] 9. Guo, R; Mukamel, S.; Klug, D. R. Geometry determination of complexes in a molecular liquid mixture using electron-vibration-vibration two-dimensional infrared spectroscopy with a vibrational transition density cube method. Phys. Chem. Chem. Phys. 2012, 14, 14023-14033.

[0265] 10. Spiegel, J., Adhikari, S. & Balasubramanian, S. The Structure and Function of DNA G- Quadruplexes. Trends Chem. 2, 123-136 (2020).

[0266] 11. Ma, Y., lida, K. & Nagasawa, K. Topologies of G-quadruplex: Biological functions and regulation by ligands. Biochem. Biophys. Res. Commun. 531 , 3-17 (2020).

[0267] 12. Chambers, V. S. et al. High-throughput sequencing of DNA G-quadruplex structures in the human genome. Nat. Biotechnol. 33, 877-881 (2015).

[0268] 13. Huppert, J. L. & Balasubramanian, S. Prevalence of quadruplexes in the human genome. Nucleic Acids Res. 33, 2908-2916 (2005).

[0269] 14. Huppert, J. L. & Balasubramanian, S. G-quadruplexes in promoters throughout the human genome. Nucleic Acids Res. 35, 406-413 (2007).

[0270] 15. Sen, D. & Gilbert, W. Formation of parallel four-stranded complexes by guanine-rich motifs in DNA and its implications for meiosis. Nature 334, 364-366 (1988).

[0271] 16. Fouquerel, E., Parikh, D. & Opresko, P. DNA damage processing at telomeres: The ends justify the means. DNA Repair (Amst) 44, 159-168 (2016). Siddiqui-Jain, A., Grand, C. L., Bearss, D. J. & Hurley, L. H. Direct evidence for a G-quadruplex in a promoter region and its targeting with a small molecule to repress c-MYC transcription. Proc. Natl. Acad. Sci. USA 99, 11593-11598 (2002). Cogoi, S. & Xodo, L. E. G-quadruplex formation within the promoter of the KRAS proto-oncogene and its effect on transcription. Nucleic Acids Res. 34, 2536-2549 (2006). Phan, A. T., Modi, Y. S. & Patel, D. J. Propeller-Type Parallel-Stranded G-Quadruplexes in the Human c-myc Promoter. J. Am. Chem. Soc. 126, 8710-8716 (2004). Banyay, M., Sarkar, M. & Graslund, A. A library of IR bands of nucleic acids in solution. Biophys. Chem. 104, 477-488 (2003). Zhang, Y., Chen, J., Ju, H. & Zhou, J. Thermal denaturation profile: A straightforward signature to characterize parallel G-quadruplexes. Biochimie. 157, 22-25 (2019). Mergny, J.-L., Phan, A.-T. & Lacroix, L. Following G-quartet formation by UV-spectroscopy. FEBS Lett. 435, 74-78 (1998). Weissbluth, M. Hypochromism. Q Rev. Biophys. 4, 1-34 (1971). Chinsky, L. & Turpin, P. Y. Ultraviolet Raman Spectroscopy of Polyribouridylic Acid: Excitation Profile of the Hypochromism Induced by Order-Disorder Transition. Biopolymers 19, 1507-1515 (1980). Tomlinson, B. L. & Peticolas, W. L. Conformational Dependence of Raman Scattering Intensities in Polyadenylic Acid. J. Chem. Phys. 52, 2154-2156 (1970). Turpin, P. Y., Chinsky, L., Laigle, A. & Jolies, B. DNA Structure Studies by Resonance Raman Spectroscopy. J. Mol. Struct. 214, 43-70 (1989). Reipa, V., Niaura, G. & Atha, D. H. Conformational analysis of the telomerase RNA pseudoknot hairpin by Raman spectroscopy. RNA 13, 108-115 (2007). Peticolas, W. L. Raman Spectroscopy of DNA and Proteins. Methods Enzymol. 246, 389-416 (1995). Friedman, S. J. & Terentis, A. C. Analysis of G-quadruplex conformations using Raman and polarized Raman spectroscopy. J. Raman Spectrosc. 47, 259-268 (2016). Miura, T. & Thomas, G. J. Structural Polymorphism of Telomere DNA: Interquadruplex and Duplex -Quadruplex Conversions Probed by Raman Spectroscopy. Biochemistry 33, 7848-7856 (1994). Lane, A. N., Chaires, J. B., Gray, R. D. & Trent, J. O. Stability and kinetics of G-quadruplex structures. Nucleic Acids Res. 36, 5482-5515 (2008). Guzman, M. R., Liquier, J., Brahmachari, S. K. & Taillandier, E. Characterization of parallel and antiparallel G-tetraplex structures by vibrational spectroscopy. Spectrochim. Acta A Mol. Biomol. Spectrosc. 64, 495-503 (2006). Price, D. A. et al. Infrared Spectroscopy Reveals the Preferred Motif Size and Local Disorder in Parallel Stranded DNA G-Quadruplexes. ChemBioChem 21 , 1-14 (2020). 34. Dickerhoff, J., Dai, J. & Yang, D. Structural recognition of the MYC promoter G-quadruplex by a quinoline derivative: insights into molecular targeting of parallel G-quadruplexes. Nucleic Acids Res. 49, 5905-5915 (2021).

[0272] Also disclosed herein are the following numbered clauses:

[0273] A1 . A method of determining the geometry of interaction between a protein or nucleic acid and one or more chemical entities, the method comprising the steps of:

[0274] (a) providing an experimentally obtained multi-dimensional spectrum for a sample comprising the protein or nucleic acid and the one or more chemical entities;

[0275] (b) providing two or more structural hypotheses;

[0276] (c) obtaining a computationally generated multi-dimensional spectrum of the same type as the experimentally obtained multi-dimensional spectrum for each of the two or more structural hypotheses;

[0277] (d) determining which of the computationally generated multi-dimensional spectra best fits the experimentally obtained multi-dimensional spectrum; and

[0278] (e) based on step (d), identifying the geometry of interaction between the protein or nucleic acid and one or more chemical entities.

[0279] A2. The method according to clause A1 , wherein the experimentally obtained multi-dimensional spectrum is obtained using two-dimensional infrared ‘2DIR’ spectroscopy.

[0280] A3. The method according to clause A2, wherein the experimentally obtained multi-dimensional spectrum is obtained using electron-vibration-vibration two-dimensional infrared ‘EVV 2DIR’ spectroscopy.

[0281] A4. The method according to any preceding clause, wherein the one or more chemical entities is a drug molecule.

[0282] A5. The method according to clause A4, wherein the geometry of interaction is located at a target binding site.

[0283] A6. The method according to any preceding clause, wherein the computationally generated multidimensional spectra are calculated using density functional theory ‘DFT’.

[0284] A7. The method according to any preceding clause, wherein step (d) comprises ranking the two or more structural hypotheses.

[0285] A8. The method according to any preceding clause, wherein step (d) comprises determining a correlation. A9. The method according to any preceding clause, wherein the experimentally obtained multidimensional spectrum is pre-processed.

[0286] A10. The method according to any preceding clause, wherein the experimentally obtained multidimensional spectrum is cleaned to remove artefacts.

[0287] A11. The method according to any preceding clause, wherein the computationally generated multidimensional spectra are each pre-processed.

[0288] A12. The method according to any preceding clause, wherein the computationally generated multidimensional spectra are each cleaned to remove artefacts.

[0289] A13. The method according to any preceding clause wherein step (a) comprises measuring the multidimensional spectrum for the sample.

[0290] A14. The method according to clause A13, wherein the sample is deposited as a substrate in the form of a microarray comprising at least one spot of deposited material.

[0291] A15. The method according to clause A14, wherein the substrate is transparent to visible light; or reflective of visible light.

[0292] A16. The method according to clause A14 or clause A15, wherein the microarray is maintained within a humidity range before and during measurement of the sample.

[0293] A17. The method according to any of clauses A13 to A16, wherein the sample scatters a substantial amount of the incident visible light.

[0294] A18. The method according to claim A17, wherein the scattered light is collected by an optical component or components forming a lens or mirror before being detected.

[0295] A19. A computer-readable medium comprising computer-executable instructions which, when executed by a processor, cause the processor to perform the method according to any of clauses A1 to A18.

[0296] B1 . A system for determining the geometry of interaction between a protein or nucleic acid and one or more chemical entities, the system comprising:

[0297] (a) a spectrometer arranged to experimentally obtain a multi-dimensional spectrum for a sample comprising the protein or nucleic acid and the one or more chemical entities; and

[0298] (b) one or more processors arranged to: generate or receive two or more structural hypotheses; obtain a computationally generated multi-dimensional spectrum from each of the two or more structural hypotheses; determine which of the computationally generated multi-dimensional spectra best fits the experimentally obtained multi-dimensional spectrum; and identify the geometry of interaction between the protein or nucleic acid and one or more chemical entities based on the determination of which of the computationally generated multi-dimensional spectra best fits the experimentally obtained multidimensional spectrum.

[0299] B2. The system according to clause B1 , wherein the spectrometer is a two-dimensional infrared ‘2DIR’ spectrometer.

[0300] B3. The system according to clause B2, wherein the spectrometer is a two-dimensional infrared ‘EVV 2DIR’ spectrometer.

[0301] B4. The system according to any preceding clause, wherein the one or more chemical entities is a drug molecule.

[0302] B5. The system according to clause B4, wherein the geometry of interaction is located at a target binding site.

[0303] B6. The system according to any preceding clause, wherein the computationally generated multidimensional spectra are calculated using density functional theory ‘DFT’.

[0304] B7. The system according to any preceding clause, wherein the one or more processors is arranged to rank the two or more structural hypotheses as part of determining which of the computationally generated multi-dimensional spectra best fits the experimentally obtained multidimensional spectrum.

[0305] B8. The system according to any preceding clause, wherein the one or more processors is arranged to receive a user input in order to determine which of the computationally generated multidimensional spectra best fits the experimentally obtained multi-dimensional spectrum.

[0306] B9. The system according to any preceding clause, wherein the one or more processors is arranged to pre-process the experimentally obtained multi-dimensional spectrum

[0307] B10. The system according to any preceding clause, wherein the one or more processors is arranged to clean the experimentally obtained multi-dimensional spectrum to remove artefacts. B11 . The system according to any preceding clause, wherein the one or more processors is arranged to pre-process each of the computationally generated multi-dimensional spectra.

[0308] B12. The system according to any preceding clause, wherein the one or more processors is arranged to clean each of the computationally generated multi-dimensional spectra to remove artefacts.

[0309] B13. The system according to any preceding clause, further comprising the sample, wherein the sample comprises a substrate in the form of a microarray comprising at least one spot of deposited material.

[0310] B14. The system according to clause B13, wherein the substrate is transparent to visible light; or reflective of visible light.

[0311] B15. The system according to clause B13 or clause B14, further comprising a humidity control device arranged to maintain the humidity of the microarray within a humidity range before and / or during measurement of the sample with the spectrometer.

[0312] B16. The system according to any of clauses B13 to B15, wherein the sample is arranged to scatter a substantial amount of incident visible light.

[0313] B17. The system according to clause B16, wherein the system further comprises at least one optical component forming a lens or mirror, wherein the at least one optical component is arranged collect the scattered light.

Claims

Claims1. A method of determining the geometry of interaction between a protein or nucleic acid and one or more chemical entities, the method comprising the steps of:(a) providing an experimentally obtained multi-dimensional spectrum for a sample comprising the protein or nucleic acid and the one or more chemical entities;(b) providing two or more structural hypotheses;(c) obtaining a computationally generated multi-dimensional spectrum of the same type as the experimentally obtained multi-dimensional spectrum for each of the two or more structural hypotheses;(d) determining which of the computationally generated multi-dimensional spectra best fits the experimentally obtained multi-dimensional spectrum; and(e) based on step (d), identifying the geometry of interaction between the protein or nucleic acid and one or more chemical entities.

2. The method according to claim 1 , wherein the experimentally obtained multi-dimensional spectrum is obtained using two-dimensional infrared ‘2DIR’ spectroscopy.

3. The method according to claim 2, wherein the experimentally obtained multi-dimensional spectrum is obtained using electron-vibration-vibration two-dimensional infrared ‘EVV 2DIR’ spectroscopy.

4. The method according to any preceding claim, wherein the one or more chemical entities is a drug molecule.

5. The method according to claim 4, wherein the geometry of interaction is located at a target binding site.

6. The method according to any preceding claim, wherein the computationally generated multidimensional spectra are calculated using density functional theory ‘DFT’.

7. The method according to any preceding claim, wherein step (d) comprises ranking the two or more structural hypotheses.

8. The method according to any preceding claim, wherein step (d) comprises determining a correlation.

9. The method according to any preceding claim, wherein the experimentally obtained multidimensional spectrum is pre-processed.

10. The method according to any preceding claim, wherein the experimentally obtained multidimensional spectrum is cleaned to remove artefacts.

11. The method according to any preceding claim, wherein the computationally generated multidimensional spectra are each pre-processed.

12. The method according to any preceding claim, wherein the computationally generated multidimensional spectra are each cleaned to remove artefacts.

13. The method according to any preceding claim wherein step (a) comprises measuring the multidimensional spectrum for the sample.

14. The method according to claim 13, wherein the sample is deposited as a substrate in the form of a microarray comprising at least one spot of deposited material.

15. The method according to claim 14, wherein the substrate is transparent to visible light; or reflective of visible light.

16. The method according to claim 14 or claim 15, wherein the microarray is maintained within a humidity range before and during measurement of the sample.

17. The method according to any of claims 13 to 16, wherein the sample scatters a substantial amount of the incident visible light.

18. The method according to claim 17, wherein the scattered light is collected by an optical component or components forming a lens or mirror before being detected.

19. A system for determining the geometry of interaction between a protein or nucleic acid and one or more chemical entities, the system comprising:(a) a spectrometer arranged to experimentally obtain a multi-dimensional spectrum for a sample comprising the protein or nucleic acid and the one or more chemical entities; and(b) one or more processors arranged to: generate or receive two or more structural hypotheses each comprising a computationally generated multi-dimensional spectrum; determine which of the computationally generated multi-dimensional spectra best fits the experimentally obtained multi-dimensional spectrum; and identify the geometry of interaction between the protein or nucleic acid and one or more chemical entities based on the determination of which of the computationallygenerated multi-dimensional spectra best fits the experimentally obtained multidimensional spectrum.

20. A computer-readable medium comprising computer-executable instructions which, when executed by a processor, cause the processor to perform the method according to any of claims 1 to 19.