Method for determining interaction geometry
Patent Information
- Application Number
- JP2026501228
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-07-13
- Filing Date
- 2024-07-12
- Publication Date
- 2026-09-04
Smart Images

Figure 2026530132000010 
Figure 2026530132000011 
Figure 2026530132000012
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for determining the interaction geometry between a target macromolecule, such as a protein or nucleic acid, and one or more chemical entities. [Background technology]
[0002] A wide range of techniques are available to obtain structural insights into molecular structure and the geometry of intermolecular interactions. These techniques include X-ray crystallography, cryo-electron microscopy, and two-dimensional NMR spectroscopy. Molecular recognition, where one molecule recognizes another, is a crucial component of biological mechanisms. Molecular recognition facilitates selective and specific interactions and bindings while simultaneously suppressing interactions with other molecular species. The binding of enzyme substrates or cofactors to the active site of a protein is a good example. Other examples include protein-protein binding for regulatory purposes and the binding of proteins to nucleic acids for gene expression regulation. The stability and specificity of the interactions underlying molecular recognition are determined by the specific spatial arrangement of atoms and functional groups. Sometimes dynamic elements are involved, and potential binding motifs or structures may temporarily appear until they are stabilized by the interaction.
[0003] One specific application is the drug discovery and development process for small molecule drugs. The early stages of small molecule drug discovery are highly iterative, involving the synthesis of numerous potential drug molecules to improve the properties of candidate drug molecules. This can occur at any point in the process, whether it's hit-to-lead development, lead optimization, or simply to enhance drug performance. As part of these iterative processes, the ability to detect and characterize drug-target interactions (e.g., protein-ligand interactions) is essential. Furthermore, it's crucial to determine the drug's position, detect contact between the drug and target, and determine the geometry of the entire complex formed by the drug and target. This data allows chemists to compare candidate drugs and identify which molecular elements should be modified to improve binding properties. Therefore, obtaining structural information about drug-protein interactions is an essential element in most drug discovery and development processes.
[0004] Despite the excellent performance and usefulness of the current set of methods used for structural analysis, practical limitations mean that in many cases, even with existing technologies, it is not possible to obtain desired structural information in sufficient quantity with sufficient accuracy, or even to obtain any desired structural information at all. These include, for example, cases where protein crystallization is impossible, cases where there are disordered regions in the crystal structure that cannot be elucidated, cases where the structural accuracy of the ligand contact site is insufficient to advance the drug discovery process, or cases where the target region within the protein is relatively unstructured in its native state.
[0005] Currently, drug development activities may involve the synthesis of thousands of molecular variants and subsequent synthesis steps based on the accidental crystal structures of lead series candidate molecules. While electron microscopy is beginning to replace X-ray crystallography in determining novel structures, it remains slow and costly for use in drug discovery activities. Furthermore, it is unclear whether electron microscopy is applicable to inherently disordered targets. Currently, approximately 30% of all proteins are thought to fall under the category of inherently disordered targets. Overall, the cost, time, and other practical limitations of current methods necessitate improved methods for obtaining structural insights into intermolecular interactions.
[0006] Spectroscopy can be used to develop versatile assay methods with sufficient capabilities for binding detection, binding quantification, binding site confirmation, and structural analysis. In particular, modified two-dimensional nuclear magnetic resonance (2D NMR) spectroscopy is used for structural analysis, and various optical spectroscopy methods are used for binding detection (often in the form of fluorescence assays). However, 2D NMR generally lacks sufficient sensitivity and throughput for use as a screening method for drug binding analysis, and optical spectroscopy generally does not have sufficient structural analysis capabilities for use in structural analysis. [Overview of the project]
[0007] Despite the relatively limited use of optical methods in structural analysis, the inventors employed a modified two-dimensional infrared (2DIR) spectroscopy technique called electron-vibration-vibration (EVV) to determine the geometry of intermolecular interactions. The inventors discovered that by using EEV 2DIR spectroscopy in combination with quantum chemical calculations based on density functional theory, a series of structural hypotheses regarding the geometry of molecular complexes can be ranked by comparing experimental 2DIR spectra with predicted 2DIR spectra calculated from structural hypotheses. Ranking of structural hypotheses may involve determining the geometry of the molecular complex from a series of structural hypotheses, including the actual geometry. The method developed by the inventors is particularly advantageous in that it can significantly reduce the cost and time of the drug discovery process. This significant reduction in cost and time is possible because the invention allows for the determination of many structures per day and can be performed using relatively small amounts of target molecules (typically proteins, but not necessarily proteins).
[0008] In this context, geometry is synonymous with the molecular pose at the drug binding site, as is generally known. In drug-target interactions and other forms of biomolecular recognition, as well as non-biological interactions, geometry involves a series of contacts between chemical entities. These contacts include hydrogen bonds, salt bridges, dipole interactions, and dispersion forces, including π-π interactions.
[0009] A method for determining the geometry of an interaction between a protein or nucleic acid and one or more chemical entities, comprising the following steps: (a) Prepare experimentally obtained multidimensional spectra of a sample containing a protein or nucleic acid and one or more chemical entities; (b) Prepare two or more structural hypotheses; (c) Obtain computationally generated multidimensional spectra of the same type as experimentally obtained multidimensional spectra for each of two or more structural hypotheses; (d) Determining which of the computationally generated two-dimensional spectra best fits the experimentally obtained multidimensional spectrum; and A method is provided which includes (e) identifying the geometry of the interaction between a protein or nucleic acid and one or more chemical entities based on step (d).
[0010] Furthermore, a system for determining the geometry of interactions between a protein or nucleic acid and one or more chemical entities, (a) A spectrometer configured to experimentally obtain a multidimensional spectrum of a sample comprising a protein or nucleic acid and one or more chemical entities; and (b) One or more processors, Generate or receive two or more structural hypotheses, each containing a computationally generated multidimensional spectrum; Determine which of the computationally generated multidimensional spectra best fits the experimentally obtained multidimensional spectrum; and Also provided is a system including a processor configured to identify the geometry of interactions between a protein or nucleic acid and one or more chemical entities, based on the determination of the computationally generated multidimensional spectrum that best fits the experimentally obtained multidimensional spectrum.
[0011] Also provided is a computer-readable medium containing computer-executable instructions that, when executed by a processor, cause the processor to perform any of the methods described herein.
[0012] The inventors have advantageously discovered that it is possible to determine the interaction geometry by comparing two or more structural hypotheses and determining which best fits the experimental data. This approach makes it possible to investigate the geometry of the interaction in question, which would not be possible from experimental data alone. [Brief explanation of the drawing]
[0013] [Figure 1] (A) FGFR1, (B) FGFR1 + SU5402, (C) difference spectrum between FGFR1 + SU5402 and FGFR1, (D) BSA, (E) BSA + SU5402, (F) experimental EVV 2DIR spectra showing the difference spectrum between BSA + SU5402 and BSA. The difference spectrum from BSA contains spectral features of SU5402 in a nonspecifically bound state, while the difference spectrum from FGFR1 contains features attributable to SU5402 in a protein-bound state. The difference spectrum from FGFR1 contains six additional peaks in addition to all the drug peaks present in the difference spectrum from BSA. These may be binding-dependent. The frequency positions of these six binding-related peaks are shown on both difference spectra. [Figure 2] This figure shows the frequencies of the peaks that appear in the FGFR1 / SU5402 difference spectrum due to coupling, rounded to the nearest 5cm⁻¹ value. [Figure 3] This figure shows the average peak response as a function of the SU5402 / FGFR molar ratio for the difference spectral peaks that appear with and without binding. The tendency for the signals that appear with the six bindings to flatten as a function of increasing drug concentration supports the idea that these appear with binding due to the saturation of available protein binding sites. Since the signal is determined by the square of the number of molecules, the square root of the fitted peak intensity was plotted against the SU5402 / FGFR molar ratio. [Figure 4] This figure shows the geometry of the SU5402 molecule, including five partial amino acid residues and one water molecule, used for EVV 2DIR spectral calculations of SU5402 bound to FGFR1. The carbon atoms of the SU5402 molecule are shown in orange, and the carbon atoms of the contained protein are shown in pink. [Figure 5]This figure shows (a) the calculated spectrum of the SU5402 and FGFR1 binding site, (b) the experimentally bound FGFR1 / SU5402 difference spectrum, (c) the calculated SU5402 in free space spectrum, and (d) the experimentally nonspecific binding BSA / SU5402 difference spectrum. From the comparison, the peak appearing with binding at 1660 / 3240 can be attributed to the coupling of mode E, the amide mode of the protein, and mode C of SU5402. Since mode B is present in the calculated spectra with and without the binding site, the experimental binding dependence of its peak is due to the lifetime effect. The black dotted line is the harmonic diagonal, and all features located on this line are harmonics. Therefore, this line passes through the intersection of lines of the same color, because the same mode is represented as a harmonic coupled with itself. [Figure 6] The figure shows (a) the calculated EVV spectra of SU-5402, its neighboring amino acid residues, and water molecules, and (b) the experimental 1:1 SU-5402 / FGFR1 difference spectrum for comparison. The letter labels indicate how the experimental peaks are assigned to the calculated peaks. Binding-sensitive peaks are marked with pink circles. [Figure 7] This figure shows the modes assigned to the α and β modes of the cross-peaks that appear with respect to the six bonds. [Figure 8] This figure shows the structures of SU-5402 and FGFR1 binding site atoms used in the EVV spectral calculation. Atoms exhibiting motion in Mode 1 are highlighted in pink. As can be seen from the figure, all vibrational motion in Mode 1 is present on atoms of the SU-5402 molecule. Motion is observed on all atoms of SU-5402 except for the atoms of the propionic acid group. The contribution of the maximum amplitude in this mode is the asymmetric ring deformation of the pyrrole group, which is composed of atoms labeled 38-42. This mode forms the α modes of peaks D, E, F, G, and H. [Figure 9]This figure shows the structures of the SU-5402 and FGFR1 binding site atoms used in the EVV spectral calculation. Atoms exhibiting motion in mode 2 are highlighted in pink. As can be seen from the figure, the vibrational motion in mode 2 is divided between the SU-5402 molecule and the protein binding site. In the SU-5402 molecule, motion is shown in the pyrrole group, the amide portion of the oxiindole group, and the alkene carbon linking these two. In the protein, motion is shown in the amide backbone groups formed by atoms labeled 14, 13, 16, 66 and 9, 8, 11, 61. This vibrational mode is likely coupled between the SU-5402 molecule and the protein residues via hydrogen bonding interactions between the oxiindole amide and the protein amide backbone. The maximum amplitude contribution to this mode is due to the oxiindole amide motion of the SU-5402 molecule. This mode forms the α mode of peak R. [Figure 10] This figure shows the structures of the SU-5402 and FGFR1 binding site atoms used in the EVV spectral calculation. Atoms exhibiting motion in mode 3 are highlighted in pink. As can be seen from the figure, the vibrational motion in mode 3 is divided between the SU-5402 molecule and the protein binding site. In the SU-5402 molecule, motion is shown in the pyrrole group, the amide portion of the oxiindole group, and the alkene carbon linking these two. In the protein, motion is shown in the amide skeleton group formed by atoms labeled 14, 13, 16, 66 and 9, 8, 11, 61. This vibrational mode is likely coupled between the SU-5402 molecule and the protein residues via hydrogen bonding interactions between the oxiindoleamide and the protein amide skeleton. The maximum amplitude contribution to this mode is due to the oxiindoleamide motion of the SU-5402 molecule. This mode forms the β mode of peak H. [Figure 11]This figure shows the structures of the SU-5402 and FGFR1 binding site atoms used in the EVV spectral calculation. Atoms exhibiting motion in mode 4 are highlighted in pink. As can be seen from the figure, the vibrational motion in mode 4 is separated between the SU-5402 molecule and the protein binding site. In the SU-5402 molecule, motion is shown in the pyrrole group and its methyl group, the 5-membered ring-containing amide of the oxindole group, and the alkene carbon linking these two. In the protein, motion is shown in the amide skeleton group formed by atoms labeled 14, 13, 16, 66 and 18, 19, 21, 68, as well as the methyl group at carbon 20. This vibrational mode is likely coupled between the SU-5402 molecule and the protein residue via hydrogen bonding interactions between the oxindole amide and the protein amide skeleton. The maximum amplitude contribution to this mode is due to the oxindole amide motion of the SU-5402 molecule. This mode forms the β mode with peaks G and R. [Figure 12] This figure shows the structures of SU-5402 and FGFR1 binding site atoms used in the EVV spectral calculation. Atoms exhibiting motion in mode S are highlighted in pink. As can be seen from the figure, the vibrational motion in mode 5 is separated between the SU-5402 molecule and the protein binding site. Motion is shown in all atoms of the SU-5402 molecule except for the atoms of the propionic acid group. In the protein, motion is shown in some of the amide skeleton groups formed by atoms labeled 14, 13, 16, 66 and 18, 19, 21, 68, as well as in the three methyl groups on carbons 15, 20, and 22. This vibrational mode is likely coupled between the SU-5402 molecule and the protein residue via hydrogen bonding interactions between the oxiindoleamide and the protein amide skeleton. The maximum amplitude contribution to this mode is due to the oxiindoleamide motion of the SU-5402 molecule. This mode forms the β mode of peak F. [Figure 13]This figure shows the structures of SU-5402 and FGFR1 binding site atoms used in the EVV spectral calculation. Atoms exhibiting motion in mode 6 are highlighted in pink. As can be seen from the figure, the vibrational motion in mode 6 is separated between the SU-5402 molecule and the protein binding site. Motion is shown in all atoms of the SU-5402 molecule except for the atoms of the propionic acid group. In the protein, motion is shown in the amide skeleton group formed by atoms labeled 5, 6, 7, 8, and 9, as well as in the methyl group at carbon 2. Motion is also observed in hydrogen atoms 66 and 60. This vibrational mode is likely coupled between the SU-5402 molecule and protein residues via hydrogen bonding interactions between the oxiindoleamide and the protein amide skeleton. The maximum amplitude contribution to this mode is due to the oxiindoleamide motion of the SU-5402 molecule. This mode forms the β mode of peak E. [Figure 14] This figure shows the structures of the SU-5402 and FGFR1 binding site atoms used in the EVV spectral calculation. Atoms exhibiting motion in mode 7 are highlighted in pink. As can be seen from the figure, the vibrational motion in mode 7 occurs mainly on the SU-5402 molecule, with slight motion observed in some protein residues. In the SU-5402 molecule, motion is shown in all atoms including the propionic acid group. In the protein, motion is shown only in the hydrogen atoms of the amide skeleton labeled 60 and 61. This vibrational mode is likely coupled between the SU-5402 molecule and the protein residues via hydrogen bonding interactions between the oxyindoleamide and the protein amide skeleton. The maximum amplitude contribution in this mode is the ring deformation of the pyrrole group on the SU-5402 molecule. This mode forms the β mode of peak D. [Figure 15]a) This figure shows the chemical structure of a G-quartet with metal ion coordination to oxygen O6. This structure shows a Hoogstin-type hydrogen bonding motif present between the four guanine bases of the planar G-quartet. b) This figure shows the structure of an intramolecular parallel propeller-type G-quadruplicate formed by Myc2345 in the presence of K+. K+ is located between two G-quartets (not shown here to clarify the nucleotide folding structure). The yellow rectangles indicate guanine bases in C2-end / anticonformation. c) This is a schematic diagram of the laser pulses used in the EVV 2DIR experiment. The third probe beam (ωγ) is non-resonant in the electronic region and is used to observe the polarization produced by the two resonant IR beams. Meanwhile, the two infrared laser beams ωα and ωβ induce resonant oscillations in the system. The signal is anti-Stokes with respect to the probe beam photons, and ωδ = ωγ + ωβ - ωα. d) This figure shows a pulse sequence illustrating the delay time between different pulses. The delay time used in all experimental measurements described herein was Tαβ = Tβγ = 0.4 ps. [Figure 16]This figure shows experimental EVV 2DIR spectra of the G4-forming DNA sequence Myc2345 in various [DNA]:[K+] concentrations. a-d are processed EVV 2DIR spectra of 1 mM Myc2345 in the presence of 0 mM (a), 5 mM (b), 10 mM (c), and 100 mM (d) K+. All spectra are normalized to the strongest peak at ωα = 1525 cm⁻¹ and ωβ-ωα = 1450 cm⁻¹. In EVV 2DIR spectroscopy, if one IR pulse (frequency ωβ) excites a coupling band and harmonics, and a second IR pulse (frequency ωα) excites a single quantum transition, the coupling of the single quantum transition and its harmonics is expected to appear on the "harmonic diagonal" line, indicated by the dotted line. However, due to other factors such as anharmonics, it may actually appear slightly below this line. The features outside the diagonal are due to the coupling of vibrational modes at frequencies ωα and ωβ-ωα, where ωβ excites a combination band with a frequency approximately equal to the sum of the frequencies of the two coupled modes. The horizontal and vertical dashed lines in (a) highlight rows and columns of cross-peaks with frequencies dependent on potassium ion concentration, which are attributed to the guanine carbonyl stretching mode coupled with various other modes. The guanine carbonyl stretching mode (GCO I) giving rise to the row of cross-peaks has a higher frequency than the guanine carbonyl stretching mode (GCO II) giving rise to the column of cross-peaks. [Figure 17]This figure shows fitted EVV 2DIR spectra of the G4-forming DNA sequence Myc2345 in various [DNA]:[K+] states. a-d are processed EVV 2DIR spectra of 1mM Myc2345 in the presence of 0mM(a), 5mM(b), 10mM(c), and 100mM(d) K+. They were fitted with a sum of 102 Gaussian functions. Markers indicate the median frequencies of the 102 Gaussian functions used to fit each spectrum. The diagonal dotted lines indicate the harmonic diagonals. Features on the diagonals are due to the coupling of the fundamental mode ωα and its harmonic ωβ. Features outside the diagonals are due to the coupling of the vibrational mode frequencies ωα and ωβ-ωα. Spectral changes observed as a result of DNA structural changes induced by K+ include a decrease in cross-peak intensity (blue circles), an increase in cross-peak intensity (red circles), a redshift in the frequency of guanine carbonyl stretching vibrations at approximately 1700 cm⁻¹ (green diamonds), and the appearance of new vibrational couplings (purple circles). [Figure 18] This figure shows a line-cut comparison of raw data and fitted results of the EVV 2DIR spectrum of Myc2345. The raw data (dotted line) and fitted (solid line) EVV 2DIR spectra at ωα=1376cm⁻¹ (purple) and ωα=1482cm⁻¹ measured with 1mM Myc2345 in the presence of 100mM K+ are shown. Comparisons of raw and fitted data at other ωα and ωβ-ωα frequencies can be found in Figure 26. [Figure 19] This figure shows the frequency shift and intensity changes of cross-peaks observed as a result of the formation of a complete G4 structure by K+. a-b are the frequency-potassium ion concentration dependence of the guanine carbonyl stretching modes GCO I (a) and GCO II (b). c-i are the relative intensities of the peaks, showing the increase (c) and decrease (d-i) of the cross-peak intensity as a function of potassium ion concentration. The intensity is shown as a ratio to the intensity in the absence of K+ ions. The peaks shown here show an intensity difference of ≥20% overall. The intensity-ion concentration dependence of the remaining peaks is shown in Figure 26. [Figure 20]This figure shows the new coupling peaks 92 (ωα=1691cm-1, ωβ-ωα=1132cm-1) and 101 (ωα=1705cm-1, ωβ-ωα=1675cm-1) resulting from K+-induced G4 formation. a shows the experimental spectrum of 1mM Myc2345 in the presence of 100mM K+ as a ratio to the Myc2345 spectrum in the absence of the ion. The dotted black boxes indicate the regions where the new coupling peaks 92 (bottom) and 101 (top) appear. b shows the relative intensities of the new peaks 92 and 101 as a function of [K+] concentration. The intensities are shown as a ratio to the intensities in the absence of K+ ions. c shows the raw (top) and fitted (bottom) spectra of 1mM Myc2345 with the addition of 5 (left), 10 (center), and 100mM K+ (right). The ratios are shown to the 1 mM Myc2345 spectrum in the absence of ions, showing the gradual appearance of peak 101 with increasing [K+]. d shows the raw spectrum (top) and fitted spectrum (bottom) of 1 mM Myc2345 with the addition of 5 (left), 10 (center), and 100 mM K+ (right). The ratios are shown to the 1 mM Myc2345 spectrum in the absence of ions, showing the gradual appearance of peak 92 with increasing [K+]. The vertical and horizontal dashed lines indicate the peak positions by Gaussian fitting. [Figure 21]This figure shows the calculated EVV 2DIR spectra of the G-quartet and G-quadruplicate. a is the calculated EVV 2DIR spectrum of the non-planar G-quartet. There are four combinations of guanine carbonyl stretching modes in the G-quartet at approximately 1700 cm⁻¹. The guanine carbonyl stretching mode that yields the highest intensity coupling along the row of cross peaks is the in-phase guanine carbonyl stretching mode at 1694 cm⁻¹, as shown in (f). Two degenerate, out-of-phase guanine carbonyl stretching modes at 1704 cm⁻¹ yield the highest intensity coupling along the column of cross peaks, as shown in (g) and (h). b is the calculated EVV 2DIR spectrum of the G-quadruplicate. There are 12 combinations of guanine carbonyl stretching modes in G₄. The guanine carbonyl stretching modes that result in the highest intensity coupling along the row of cross peaks at 1692 cm⁻¹ exhibit the same motion as the G-quartet modes at 1694 cm⁻¹ in all three quartets of G4, as shown in (f). Similarly, two degenerate guanine carbonyl stretching modes produce the highest intensity coupling along the row of cross peaks at 1679 cm⁻¹. These exhibit the same motion as the G-quartet modes at 1704 cm⁻¹, as shown in (g) and (h), but span all three quartets within G4. Animations of the G4 vibrational modes at 1692 cm⁻¹ and 1679 cm⁻¹ can be found in the supplementary information. c-d are overhead (c) and side view (d) of the G-quartet structure used for the EVV 2DIR spectral calculation in (a). e is the structure of G4 used for the EVV 2DIR spectral calculation in (b). f is the in-phase guanine carbonyl stretching mode of the G-quartet at 1694 cm⁻¹. g~h are the two degenerate out-of-phase guanine carbonyl stretching modes of the G-quartet at 1704 cm⁻¹. [Figure 22]This figure shows the calculated and experimental EVV 2DIR ratio spectra (G-quadruplicate / G-quartet). 'a' is the experimental EVV 2DIR ratio spectrum of 1 mM Myc2345 in the presence of 100 mM K+ divided by 1 mM Myc2345 in the absence of ions. 'b' is the calculated EVV 2DIR ratio spectrum of the G-quadruplicate / G-quartet. While the calculated spectrum includes only electrical coupling, similar overall changes are observed when mechanical coupling is included. The increase or decrease in coupling observed around 1700 cm⁻¹ (dashed line) is due to the redshift of the frequency of the guanine carbonyl stretching vibration. In the experimental difference spectra, the positive features (black crosses) at ωα=1691cm⁻¹, ωβ-ωα=1132cm⁻¹ and ωα=1705cm⁻¹, ωβ-ωα=1675cm⁻¹ are attributed to the appearance of new peaks 92 and 101, respectively. The remaining difference features are attributed to changes in the intensity of the cross-peaks as a result of DNA structural changes associated with potassium ion introduction. The experimental ratio spectra are consistent with the structural differences in the G-quartet and G4. The alternating red and blue columns on the right side of both spectra qualitatively indicate that the essential features associated with those columns are reflected in both computational and experimental results, and the absence of this pattern in the corresponding rows similarly supports this. [Figure 23] This is a schematic diagram of the proposed equilibrium state between the non-stacked G-quartet formed on Myc2345 in the absence of potassium ions and the G4 formed in the presence of potassium ions. Left: A depiction of possible pathways through which Myc2345 can form a non-stacked G-quartet in the absence of potassium ions. Right: Schematic diagram of the G4 structure of Myc2345 in the presence of potassium ions. For clarity, only the guanine bases and end cap bases involved in the quartet are included. Novel coupling peaks 92 and 101 suggest that the G / T of the end cap is coupled with the core G4 structure. [Figure 24]This figure compares the raw data and fitted experimental EVV 2DIR spectra of 1mM Myc2345 in the presence of 0mM and 100mM K+. The line cuts show a comparison of the raw data (dotted lines) and fitted results (solid lines) of the EVV 2DIR spectra of 1mM Myc2345 in the presence of 0 K+ (orange) and 100mM K+ (purple) at frequencies ωα=1385cm⁻¹ (top left), ωα=1525cm⁻¹ (bottom left), ωβ-ωα=1560cm⁻¹ (top right), and ωβ-ωα=1560cm⁻¹ (bottom right). [Figure 25] This figure shows the chi-squared values of the two-dimensional Gaussian function fitting of the experimental EVV 2DIR spectrum of 1 mM Myc2345 in the presence of various potassium ion concentrations. [Figure 26] This figure shows the relative intensity of the cross peaks as a function of potassium ion concentration. The intensity is shown as a ratio to the intensity in the absence of K+ ions. The experimental EVV 2DIR spectrum of Myc2345-K+ was fitted with a total of 102 two-dimensional Gaussian functions. The intensity traces of the cross peaks that increased or decreased with [K+], and the intensity traces of new vibrational couplings, can be found in Figures 19 and 20b, respectively. [Figure 27] This figure shows the calculated EVV 2DIR spectrum of an isolated guanine base. a is the calculated EVV 2DIR spectrum of the guanine base. The isolated guanine base has one carbonyl stretch mode at 1758 cm⁻¹, which causes the rows and columns of cross-peaks. b is a depiction of the guanine carbonyl stretch mode in the isolated guanine base. [Figure 28]This figure shows a comparison of G4-forming sequences in the human c-myc promoter region from different literature. a is the purine-rich wild-type sequence with 27-nt present in the NHE of the Pu27:c-myc promoter (SEQ ID NO:1). b is the Myc2345 sequence used in the EVV 2DIR measurement described herein (SEQ ID NO:2). c is the wild-type Myc2345 sequence used in the NMR structure PDB:7KBV (SEQ ID NO:3). d is the modified Myc2345 sequence MycL1 used in the CD measurement in reference 11 (SEQ ID NO:4). Bases shown in red indicate differences from the Myc2345 sequence used in the EVV 2DIR measurement. [Figure 29] This figure shows the G-quadruplicate and G-quartet structures used in the calculation of the EVV 2DIR spectrum. Figures a-e show the structure of the G-quadruplicate. Figure (b) shows a top view of the G-quadruplicate. The terminal quartets (purple) are aligned with each other, and the central quartet (green) is rotated approximately 45° relative to the terminal quartets. Figures (d) and (e) show the terminal quartets and central quartets of the G-quadruplicate, respectively. Figures f-g show (f) a top view and (g) a side view of the structure of the G-quartet. [Figure 30] This figure shows a comparison of the structures of the G-quartet and the terminal quartet in G4. (a-b) shows the structure of the G-quartet and (c) the terminal quartet of G4, illustrating the intersection of the molecular planes of adjacent guanine bases. The guanine molecular planes intersect at an angle of 21 degrees in the G-quartet and at 19 degrees in the terminal quartet of G4. [Figure 31] This figure shows the processing of the EVV 2DIR spectrum of the coverslip used to normalize the EVV 2DIR spectrum of Myc2345. Coverslip EVV 2DIR spectrum: (a) raw data, (b) after smoothing by nearest neighbor averaging, (c) after interpolation in the vertical axis direction, (d) after applying a low-pass filter for Fourier transform. [Modes for carrying out the invention]
[0014] In the first aspect, a method for determining the geometry of an interaction between a protein or nucleic acid and one or more chemical entities, comprising the following steps: (a) Prepare experimentally obtained multidimensional spectra of a sample containing a protein or nucleic acid and one or more chemical entities; (b) Prepare two or more structural hypotheses; (c) For each of two or more structural hypotheses, determine a computationally generated multidimensional spectrum of the same type as the experimentally obtained multidimensional spectrum; (d) Determining which of two or more structural hypotheses best fits the experimentally obtained multidimensional spectrum; and A method is provided which includes (e) identifying the geometry of the interaction between a protein or nucleic acid and one or more chemical entities based on step (d).
[0015] This method can be carried out in any appropriate order. This method can be carried out in the order of (a), (b), (c), and (d). This method can be carried out in the order of (b), (a), (c), and (d). Step (b), or steps (b) and (c), can be carried out before step (a).
[0016] As used herein, the term “interaction geometry” refers to the arrangement or spatial configuration of two molecules, or parts of molecules, or within a molecule. A particular configuration of two molecules (or parts of molecules) may arise from interactions such as hydrogen bonding, electrostatic interactions, and / or dispersion forces. This particular configuration may arise from the cumulative sum of forces from these interactions, and other effects such as thermal fluctuations. This configuration is typically considered a geometry that minimizes the free energy of the interacting chemical entities.
[0017] The interaction geometry is, appropriately speaking, the relative orientation and / or distance between two molecules or molecular parts that are not covalently bonded, for example. The interaction geometry may include, or consist of, the distance between interacting chemical groups. As used herein, “interacting chemical groups” refers to substructures that give character to the spectrum (for example, as a result of proximity).
[0018] As used herein, “chemical entity” refers to any molecule that can form an interaction with the protein or molecule of interest (as described above). This interaction includes interactions between chemical groups that interact with each other without direct covalent bonding (such as amino acid side chains or nucleotide bases of nucleic acids). A chemical entity may be another part of the protein or molecule of interest, for example, a DNA molecule or a part of it (such as a nucleotide base). Thus, the interacting chemical groups and / or chemical entities may be distinct molecular species, and the interaction may occur within a single chemical entity. This interaction may include, but is not limited to, interactions within a protein, within a nucleic acid, between a protein and another chemical entity, or between a nucleic acid and another chemical entity. Therefore, it will be understood that, alternatively or additionally, this disclosure describes a method for determining the geometry of interactions between a molecule or part of a molecule of interest and one or more chemical entities. The method involves the following steps: (a) Prepare experimentally obtained multidimensional spectra of a sample containing the target molecule or a part of a molecule and one or more chemical entities; (b) Prepare two or more structural hypotheses; (c) Obtain computationally generated multidimensional spectra of the same type as experimentally obtained multidimensional spectra for each of two or more structural hypotheses; (d) Determine which of the structural hypotheses best fits the experimentally obtained multidimensional spectrum; and (e) Based on step (d), the process includes identifying the geometry of the interaction between the target molecule or part of a molecule and one or more chemical entities.
[0019] The target molecule may be any suitable molecule. The target molecule is typically a macromolecule. The target molecule may be a polymer or biopolymer. The target molecule may be part of another molecule. The target molecule may be a protein. The target molecule may be part of a protein. The target molecule may be a nucleic acid, such as RNA or DNA. Therefore, the methods described herein can be used to determine the geometry of interactions between nucleic acids, such as pre-microRNA, and one or more chemical entities.
[0020] The methods described herein can be used to determine the geometry of the interaction between a ligand and a protein at a cryptic or allosteric site.
[0021] The methods described herein can be used to determine the geometry of the interaction between two or more polymer chains. For example, they can be used to determine the relative orientation of the chains.
[0022] One or more chemical entities may be one or more drug molecules (or parts thereof). As used herein, drug molecules may be candidate drugs in drug development, or chemical entities suitable for use as drugs or in drug development.
[0023] Preferably, there is only one drug molecule, and the method is a method for determining the geometry of the interaction between the protein and the drug. The geometry of the interaction between the protein and the drug (ligand) is sometimes called the ligand pose.
[0024] Therefore, the method of the first embodiment may be a method for determining the geometry of the interaction between a protein (or a part thereof) and a drug molecule (or a part thereof).
[0025] As is well known in the art, drugs are thought to bind to receptors and are therefore sometimes referred to as ligands, either surrogately or additionally.
[0026] The interaction geometry may be located at the target binding site. This disclosure may describe a method for determining the interaction geometry of a ligand-receptor complex structure.
[0027] The method of the first embodiment includes (a) preparing a multidimensional spectrum experimentally obtained for a sample.
[0028] As used herein, "multidimensional spectrum" refers to a spectrum obtained by a spectroscopic method that is not one-dimensional, i.e., a spectrum obtained by a spectroscopic method with two or more dimensions.
[0029] Conventional spectroscopy involves exposing a sample to various frequencies or wavelengths and measuring the sample's response as a function of a single variable, such as time or frequency. This measurement provides insights into the absorption, emission, and scattering properties of the sample at specific energies or frequencies.
[0030] In contrast, multidimensional spectroscopy introduces additional dimensions, often incorporating multiple variables such as delay time, frequency, and pulse sequence parameters. By manipulating these parameters, multidimensional spectra dependent on multiple variables are obtained, allowing for the exploration of correlations between different variables. This enables a more thorough investigation of the sample's behavior. Correlations can be generated through various mechanisms, such as coupling between molecular vibrations, coupling between electronic states, and coupling between vibrations and electronic states.
[0031] Two-dimensional (2D) spectroscopy obtains a two-dimensional spectrum by exciting a sample with a series of pulses and measuring the response as a function of at least two excitation frequencies and a detection frequency. When short light pulses are used, multiple correlations can be excited simultaneously due to the Fourier relations in the time and frequency domains. This Fourier relation can be used to measure the response in the time domain. Therefore, the sample response and the resulting correlations can be measured in the time domain, the frequency domain, or both. Cross-peaks appear in the resulting spectrum, providing information about the coupling and interaction between different electronic states and / or vibrational modes within the sample.
[0032] The methods described herein may involve the measurement of multidimensional spectra. Alternatively or additionally, these methods may involve the input of already measured experimental multidimensional spectra.
[0033] Methods for measuring multidimensional spectra are known in the art.
[0034] Multidimensional spectra can be obtained using two-dimensional infrared (2DIR) spectroscopy.
[0035] Multidimensional spectra can be obtained using electron vibration (EVV) two-dimensional infrared (EVV 2DIR) spectroscopy, also known as double vibrationally enhanced (DOVE) spectroscopy.
[0036] EVV 2DIR spectroscopy measures the coupling between single-quantum (fundamental vibration) transitions and double-quantum (combination band or harmonic) transitions. When an excited combination band contains an equally excited single-quantum transition, the Raman scattering between the two levels is an acceptable transition. Therefore, the visible light beam effectively probes the coupling between two excited vibrations (one at the fundamental frequency, the other at the combination band or harmonic). The resulting measured correlation can be plotted in various ways. Using the frequencies of the two excitation beams as axes, the resulting spectrum shows a series of features corresponding to the ground transition, which couples with many other vibrations. Diagonal features are also observed, appearing as virtually straight lines. This is because the second single-quantum transition contained in the combination band is constant. The frequency of the second vibration in the combination band corresponds to the frequency ω of the second excitation beam. β The frequency ω of the first excitation beam α This can be easily estimated by subtracting the following. This assumes that the frequency shift within the combination band due to the coupling itself is small or negligible, but this is not always the case. Plotting the spectrum by plotting the estimated second one-quantum transition excitation frequency against the first one-quantum transition results in columns and rows of features in the generated spectrum, each corresponding to a one-quantum transition coupled with multiple other one-quantum transitions. Regardless of how the spectrum is plotted, the features on the rows, columns, and diagonals in the two-dimensional spectrum are sometimes approximate due to factors such as interference between spectral features and frequency shifts due to additional coupling or processes. These rows and columns can be useful for confirming cross-peaks and assigning them to specific interactions by increasing confidence in multiple couplings and improving the reliability of comparisons between the computational spectrum and the experimental spectrum. In some cases, only a single peak may be present in a row or column due to various factors such as the short lifetimes of some couplings or their weakness making detection difficult.
[0037] Spectrometers suitable for EVV 2DIR spectral measurement are known in the art.
[0038] A suitable spectrometer includes an ultrafast laser system used to generate one or more high-intensity laser beams. The generated laser beam(s) are split into three arms, two of which can drive optical parametric amplifiers to generate tunable mid-infrared forces. The third beam is used to generate Raman scattering within the sample and therefore must not resonate with the sample's vibrational absorption. This third beam typically has the frequency of the first generated high-intensity laser beam. All three beams may be linearly polarized parallel to the propagation plane or have different polarization states from each other, and are spatially superimposed within the sample with beam diameters selected to maximize signal intensity without damaging the sample.
[0039] As is known in the art, different pulse sequences or different polarization combinations may be used.
[0040] Advantageously, the inventors discovered that by measuring EVV 2DIR spectra, the sequence can be adjusted over a wide spectral range, allowing for the utilization of more coupling spectra than with two-dimensional infrared spectrometers (designed to measure vibrational coupling by methods other than the Raman scattering process as part of the two-dimensional spectrum generation mechanism). Furthermore, a wider spectral range can be measured in a given experiment, and for large biomolecules such as proteins and nucleic acids, typically about 100 spectral features can be measured. These multiple features are extremely advantageous in assigning experimental features by spectral calculation or other means. They also provide numerous spectral features (often aligned in rows or columns) that can be used to compare structural hypotheses, enhancing the ability of this technique to distinguish structural hypotheses.
[0041] The delay time between the three optical pulses must be selected to maximize the signal resulting from the vibrational coupling while minimizing the non-resonant background caused by the non-linear frequency mixing within the sample. The chosen timing is usually a compromise between these two conflicting factors, and this compromise is generally determined empirically and is a function of pulse duration, pulse frequency, and sometimes the lifetime of the coupling in question. Typical timing delays range from 0.5 ps to 2 ps, but can be shortened or extended to 0.1 ps to 10 ps under specific circumstances.
[0042] Depending on the design of the EVV 2DIR spectrometer, it is possible to excite multiple vibrational states simultaneously using shorter infrared pulses, as is known in the art.
[0043] According to a second aspect, a system for determining the geometry of interactions between a protein or nucleic acid and one or more chemical entities, (a) A spectrometer configured to experimentally obtain a multidimensional spectrum of a sample comprising a protein or nucleic acid and one or more chemical entities; and (b) A processor, To generate or receive two or more structural hypotheses; For each of two or more structural hypotheses, determine the computationally generated multidimensional spectrum; Determine which of two or more structural hypotheses best fits the experimentally obtained multidimensional spectrum; and A system is provided which includes a processor configured to identify the geometry of the interaction between a protein or nucleic acid and one or more chemical entities, based on the determination of the structural hypothesis that best fits the experimentally obtained multidimensional spectrum.
[0044] The features according to the first and second embodiments are described below.
[0045] To avoid misunderstanding, the sample contains a target molecule (e.g., a protein or nucleic acid) and a chemical entity whose interaction geometry is determined. The interaction may occur between a first part of the target molecule and a second part of the target molecule. The interaction may be one or more selected from nucleic acid-drug, protein-nucleic acid, protein-protein, nucleic acid-nucleic acid, intraprotein interactions, intranucleic acid interactions, and non-biological chemical interactions within the same molecule or between molecules (e.g., interactions between artificial polymer chains).
[0046] If the method according to the first embodiment is a method for determining the geometry of the interaction between a protein and one or more drug molecules, the sample includes a protein and one or more drug molecules.
[0047] If the method of the first embodiment is a method for determining the geometry of the interaction between a protein and a drug molecule, the sample includes a protein and a drug molecule. Therefore, the sample may be referred to as a "protein-drug" sample.
[0048] The sample may contain some water. The sample may contain a buffer suitable for the sample (e.g., phosphate buffer). The sample may contain a salt (e.g., NaCl). The sample may also contain a solvent such as dimethyl sulfoxide (DMSO), or other substances necessary to place the sample in a particular state or condition.
[0049] The sample may be deposited on the substrate in the form of a microarray containing at least one spot-like deposit. The microarray may be maintained at a controlled humidity (e.g., within a predetermined humidity range) before and during the measurement of the sample.
[0050] As is well known in the art, conventional microarrays consist of regularly arranged, minute spots (also called features). These are deposited to adhere to a solid surface, i.e., a substrate. Spot sizes can range from 1 mm to 1 μm. Typically, spots are 200 to 500 μm in size.
[0051] The substrate may be transparent or reflective to visible light. The substrate may be glass. The substrate may be plastic. If infrared (IR) transmittance measurement is also useful, the substrate may further be transparent to IR light. In this case, the substrate may be calcium fluoride or another IR-transmitting material.
[0052] The sample may scatter a considerable amount of incident visible light. The source of the incident visible light is a third laser beam. This is the beam that is measured after passing through the sample. If there were no Mie scattering or Rayleigh scattering, all the light would be visible, but some undergoes inelastic (Raman) scattering. Raman scattering is generated by a coherent process and appears at a wavelength slightly different from the incident visible light beam, so the Raman scattering signal (beam) is generated at a slightly different angle from the input beam. However, because this is coherent Raman scattering, it is still a beam and not 360-degree scattering.
[0053] The Raman scattering signal may be transmitted through the substrate before detection, or reflected by the substrate before detection. If the sample exhibits a high degree of elastic scattering (i.e., a powder that generates a lot of Mie scattering), additional lenses or mirrors may be used to collect as many scattered Raman signals as possible.
[0054] In EVV 2DIR spectroscopy, the Raman scattering wavelength is different from any of the input beams, allowing for easy separation of the Raman signal from any scattered light using filters, a spectrometer, or both. This makes the technique more robust compared to other 2DIR spectroscopy methods where the output wavelength overlaps with one or more input wavelengths, especially in the presence of Mie or Rayleigh scattering.
[0055] Scattered light may be collected before detection by an optical component(s) that forms a lens or mirror.
[0056] The method of the first embodiment includes (b) preparing two or more structural hypotheses and obtaining a computationally generated multidimensional spectrum for each structural hypothesis that is of the same type as the experimentally obtained multidimensional spectrum.
[0057] For example, in the context of drug discovery, where the methods described herein may be particularly useful, a structural hypothesis is the pose that a ligand (of a drug molecule) takes relative to a protein, more specifically, within a known or presumed binding site. As is known in the art, poses are used to indicate the relative orientation of a ligand to a protein structure.
[0058] Structural hypotheses may be formulated using docking models.
[0059] In this type of model, multiple structural substitutions are validated according to specific criteria. The ranking process of a typical docking model usually involves two main components: a search algorithm that explores the conformational space and a scoring algorithm that ranks the likelihood of each possible pose.
[0060] Alternatively or additionally, structural hypotheses may be generated using ab initio protein structure prediction algorithms such as alphafold coding, and then the docking model may be used to generate the structure.
[0061] Alternatively or additionally, the structural hypothesis may be based on spectra obtained by X-ray crystallography, NMR, electron microscopy, or any other method that can provide any indication of ligand pose.
[0062] Therefore, the structural hypothesis provides possible relative orientations between the target molecule or a part thereof and one or more chemical entities.
[0063] Therefore, computational spectra of the same type as experimentally obtained spectra can be generated or calculated for each structural hypothesis. Accordingly, "computationally generated spectrum" in this specification refers to the spectrum generated based on the structural hypothesis.
[0064] Therefore, process (c) may be additional or alternative, (c) For each structural hypothesis, this may include calculating a computationally generated multidimensional spectrum of the same type as the experimentally obtained multidimensional spectrum.
[0065] Therefore, when the term "spectrum" is used in general in this specification, it is intended to refer to experimentally obtained multidimensional spectra and computationally generated multidimensional spectra, respectively.
[0066] The computational spectrum may be calculated using density functional theory (DFT). However, it will be understood that other methods for calculating the computational spectrum are known in the art. Appropriate electronic structure calculations should include calculations of the normal modes of vibration, differentiation of the potential energy surface (leading to mechanical anharmonics), and differentiation of the polarizability (leading to electrical anharmonics). Computational cost is an important factor in determining which type of quantum mechanical calculation to use. It is possible to reduce computational cost by making further approximations. For example, in many cases, electronic coupling is far more important than mechanical coupling, so it is sufficient to calculate only the electronic coupling. It is also possible to calculate coupling using the dipole approximation.
[0067] Step (d) may include ranking two or more structural hypotheses.
[0068] This method can be alternatively expressed as a way to rank two or more structural hypotheses in order to determine the interaction geometry between a target molecule, such as a protein, and one or more chemical entities.
[0069] Step (d) may include comparing the goodness of fit (correlation) between experimentally obtained multidimensional spectra and computationally generated multidimensional spectra determined from the first structural hypothesis (and so on). This comparison is continued until all comparisons of the goodness of fit between experimentally obtained spectra and each structural hypothesis are completed.
[0070] Experimentally obtained multidimensional spectra may contain artifacts. Computationally generated multidimensional spectra may also contain artifacts. Such artifacts may alter and / or distort the spectrum.
[0071] Experimentally obtained multidimensional spectra may be cleaned to remove artifacts.
[0072] Therefore, the method of the first embodiment may include cleaning the experimentally obtained multidimensional spectrum to remove artifacts.
[0073] For example, experimentally obtained spectra may contain artifacts in regions with a low signal-to-noise ratio. Experimentally obtained spectra may also contain removable non-resonant background. Non-resonant background is identified in non-overlapping frequency regions between spectral peaks. At these frequencies, non-zero signals are caused by non-resonant background. These regions can be used to determine the non-resonant background across the entire sample.
[0074] Since the non-resonant background is the same for identical pulse sequences and sample compositions, it only needs to be determined once in each case. Experimentally obtained spectra may be normalized at specific frequencies. The intensity of cross-peaks measured in experimental spectra depends on the intensity of the infrared laser beam at each frequency combination. Because infrared laser beam generation methods produce beams of different intensities at different frequencies, these variations must be normalized before comparison with calculated spectra. To achieve this, the beam intensity must be measured across the entire spectral range used so that normalization can be performed. One way to measure the beam intensity and also consider the possibility of beam overlap variations at different infrared frequencies is to first measure the 2DIR spectrum of the non-resonant substrate. The intensity of this non-resonant spectrum then provides the relative intensity of the product of the laser beams at each frequency pair in the 2D spectrum and also explains any changes in overlap due to variations in beam characteristics associated with changes in the output wavelength of the optical parametric amplifier.
[0075] Experimentally determined spectra may be pre-processed to improve the accuracy of comparison with calculated spectra. One form of pre-processing involves measuring experimental spectra of ligand-binding targets, such as proteins and nucleic acids, both in the presence and absence of the drug or ligand. These two spectra are then combined to generate a difference spectrum or ratio spectrum. This difference spectrum or ratio spectrum is then compared to a calculated difference spectrum or ratio spectrum to rank the structural hypotheses.
[0076] Computationally generated multidimensional spectra may be cleaned and artifacts removed.
[0077] Therefore, the method of the first embodiment may include cleaning the computationally generated multidimensional spectrum to remove artifacts.
[0078] For example, computationally generated spectra may contain Fermi resonances that have become high in intensity due to artifacts, which can result in coupling features of unusual intensity within the spectrum. These resonances can be ignored or their intensity can be limited using standard methods known in the art. One way to remove high-intensity Fermi resonances due to artifacts is to decompose the spectral calculation result into a numerator and a denominator. Since high-intensity Fermi resonances due to artifacts result in low values in the denominator approaching zero, these resonances can be removed from the calculation by setting a denominator threshold. Alternatively or additionally, computational features exceeding a certain threshold can be removed. As a more advanced method for cleaning computational data, setting a maximum value for Fermi resonances is known in the art.
[0079] To support an effective comparison between calculated and experimental spectra, other forms of preprocessing of the calculated spectrum may also be used. Preprocessing includes considering scaling factors between calculated and measured frequencies. This is a known limitation of vibration frequency calculations and is frequently used in comparing calculated linear one-dimensional vibration spectra with corresponding experimental spectra.
[0080] Computationally generated spectra may be preprocessed to improve the accuracy of comparison with experimental spectra. One form of preprocessing involves calculating spectra for subelements, such as ligand-binding targets or ligand-binding sites, both in the presence and absence of the ligand. These two spectra are then combined to generate a difference spectrum or ratio spectrum. This difference spectrum or ratio spectrum is then compared to experimentally determined difference or ratio spectra to rank the structural hypotheses.
[0081] The experimentally obtained spectrum can be decomposed into individual feature components through a fitting process. In this process, the spectrum is represented by two-dimensional Gaussian spectral features, each with width and amplitude. Fitting is performed using the Levenberg-Marquardt method, the simplex method, or a similar nonlinear fitting algorithm. Through trial and error and iteration, the minimum number of Gaussian features required to fit the entire experimental spectrum is determined. Fitting can be performed on both target spectra in the presence and absence of the ligand, thereby obtaining fitted difference spectra or ratio spectra for comparative purposes.
[0082] Computationally generated spectra can be decomposed into identical individual features by fitting them in the same way as experimentally obtained spectra, or simply by assigning each calculated cross-peak to a Gaussian feature with known intensity and fixed spectral width. Calculated cross-peaks are often a mixture of multiple couplings present in the same region of the spectrum, further summed and smoothed by the finite spectral resolution of the spectroscopic method. This decomposition of the calculated spectrum into a series of Gaussian features is possible for both target spectra in and without ligands, yielding decomposed difference or ratio spectra for comparative purposes.
[0083] The experimentally obtained spectra can be fitted with two-dimensional Gaussian functions. The number of Gaussian functions required corresponds to the number of individual couplings. The computationally generated spectra can each be fitted with two-dimensional Gaussian functions.
[0084] By comparing the fitting results, we can verify how closely the computationally generated spectrum is to the experimentally obtained spectrum.
[0085] Therefore, the method of the first embodiment is (d1a) This may include decomposing experimentally obtained spectra and two or more computationally generated spectra into individual features.
[0086] In the method of the first embodiment, (d) may include a reference or control. This may be, for example, a state in which the target molecule itself, i.e., the chemical entity, is absent. This reference or control may be a state in which the protein itself, i.e., the chemical entity, is absent, for example, a state in which one or more drug molecules are absent. This reference or control may be a state in which the nucleic acid itself, i.e., the chemical entity, is absent. The reference or control may be a chemical entity such as a drug molecule, and a macromolecule such as a protein is present, but does not form a specific interaction with the macromolecule. For example, in the case of a protein-drug complex, the reference may be the drug in the presence of bovine serum albumin (BSA), which is a protein. The reference or control may be the target molecule (such as DNA) in a different chemical state, for example, a different pH, in the presence or absence of ions (e.g., in the absence of potassium ions).
[0087] Therefore, in the method of the first embodiment, (d) is, (d2a) Comparing the experimentally obtained spectrum of a sample containing a protein or nucleic acid without one or more chemical entities with the experimentally generated spectrum of a sample containing a protein or nucleic acid with one or more chemical entities, and / or the reverse; and / or (d2b) This may include comparing a computationally generated spectrum of a protein or nucleic acid without one or more chemical entities with a computationally generated spectrum of a protein or nucleic acid with one or more chemical entities, and / or the reverse.
[0088] The method of the first embodiment may include (d2a) and (d2b).
[0089] The method of the first embodiment is (d2a') Analyzing the difference (or ratio) between the experimentally obtained spectrum of a sample containing a protein without one or more chemical entities and the experimentally generated spectrum of a sample containing a protein or nucleic acid with one or more chemical entities, or vice versa; (d2b') Analyzing the difference (or ratio) between the computationally generated spectrum of a protein without one or more chemical entities and the computationally generated spectrum of a protein or nucleic acid with one or more chemical entities, or vice versa; and This may involve determining the correlation between the differences between (d2c'), (d2a'), and (d2b').
[0090] The higher the correlation found in (d2c'), the higher the rank of the structural hypothesis associated with a particular computationally generated spectrum.
[0091] The method of the first embodiment is (1) Multidimensional spectra of a protein or nucleic acid and one or more chemical entities, and A multidimensional spectrum of a protein or nucleic acid without one or more chemical entities, Calculating the multidimensional spectrum of one or more chemical entities that do not involve proteins or nucleic acids; (2) Identify the changes in the multidimensional spectrum calculated in (1) that result from the interaction between the modes of one or more chemical entities and the modes of proteins or nucleic acids; (3) Repeat (1) and (2) for each structural hypothesis; (4) Multidimensional spectra of proteins or nucleic acids and one or more chemical entities, A multidimensional spectrum of a protein or nucleic acid without one or more chemical entities, Measuring the multidimensional spectrum of one or more chemical entities that do not involve proteins or nucleic acids; (5) For each structural hypothesis, compare the experimentally measured changes obtained in (4) with the changes calculated in (2); (6) This may include ranking each structural hypothesis.
[0092] This method can be performed in any appropriate order.
[0093] This method can be performed in the order of (1), (2), (3), (4), (5), (6), or in the order of (4), (1), (2), (3), (5), (6).
[0094] The method of the first embodiment is (e) Based on step (d), the process includes identifying the geometry of the interaction between a protein or nucleic acid and one or more chemical entities.
[0095] The various methods described above can be executed by computer programs.
[0096] Accordingly, according to the third aspect, a computer-readable medium is provided which, when executed by a processor, includes computer-executable instructions causing the processor to perform the method according to the first aspect.
[0097] A computer program may include computer code configured to cause a computer to perform one or more of the functions described above in various ways. Computer programs and / or code for performing such functions may be provided to a computer or other device as one or more computer-readable media, or more generally as a computer program product. Computer-readable media may be temporary or non-temporary. One or more computer-readable media may be, for example, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, or a propagation medium for data transmission (e.g., for downloading code over the Internet). Alternatively, one or more computer-readable media may take the form of one or more physical computer-readable media, such as semiconductor or solid-state memory, magnetic tape, removable computer disks, random-access memory (RAM), read-only memory (ROM), rigid magnetic disks, or optical discs such as CD-ROMs, CD-R / Ws, and DVDs.
[0098] In some implementation embodiments, the modules, components, and other features described herein may be implemented as individual components or integrated into the functionality of a hardware component such as an ASIC, FPGA, DSP, or similar device.
[0099] A “hardware component” is a tangible (e.g., non-temporary) physical component (e.g., a set of one or more processors) capable of performing a specific operation and that can be in a specific physical configuration or arrangement. A hardware component may include dedicated circuitry or logic permanently configured to perform a specific operation. A hardware component may be or include dedicated processors such as field-programmable gate arrays (FPGAs) or ASICs. A hardware component may also include programmable logic or circuitry temporarily configured by software to perform a specific operation.
[0100] Therefore, the term “hardware component” should be understood to include tangible entities that are physically constructed and permanently configured (e.g., hardwired) or temporarily configured (e.g., programmed) in order to operate in a particular way or to perform the particular operations described herein.
[0101] Furthermore, modules and components may be implemented as firmware or functional circuits within hardware devices. Additionally, modules and components may be implemented in any combination of hardware devices and software components, or solely in software (e.g., code stored or otherwise embodied in a machine-readable medium or transmission medium).
[0102] Unless otherwise stated, as will be apparent from the following description, throughout this specification, any discussion using terms such as “receive,” “determine,” “compare,” “enable,” “maintain,” and “identify” is understood to refer to the actions and processes by which a computer system or similar electronic computing device manipulates data represented as physical (electronic) quantities in its registers and memory and converts it into other data similarly represented as physical quantities in the computer system’s memory, registers, or other information storage, transmission, or display devices.
[0103] The phrase "at least two" is synonymous with "two or more," i.e., two, three, four, five, six, or more.
[0104] As used herein, the term “approximately” generally includes or refers to a range of values that a person skilled in the art would consider equivalent to (i.e., having the same function or result as) the stated value. When the term “approximately” is used in relation to a numerical value, it may represent a deviation of 10%, 5%, 2%, 1%, or 0% from that value (in order of increasing priority).
[0105] It will be understood that all embodiments described herein are considered to be broadly applicable and combinable with any other consistent embodiments as needed. Such combinature is considered to fall within the scope of the appended claims.
[0106] Here, embodiments will be further described with reference to the following examples, which are presented for illustrative purposes only. The examples will refer to the following drawings. [Examples]
[0107] Use of Electron Vibration Two-Dimensional Infrared Spectroscopy (EVV 2DIR) for detecting drug binding to target proteins This example demonstrates that spectral features appearing only when SU5402 specifically binds to a target protein can be identified, even in the harmonic coupling region, which is the most crowded region in the EVV 2DIR spectral space. Two independent methods for verifying binding specificity are also presented. The stoichiometry of binding is shown to be largely as expected. Bovine serum albumin (BSA), which does not have a binding site for SU5402, is used as a counterexample. Finally, quantum mechanical calculations of vibrational coupling at the binding site were used in the presence of the bound drug to identify features that appear with binding. This makes it possible to assign at least one contact, demonstrating the potential for calculations to be used as an assignment tool even in complex chemical systems such as protein-drug binding.
[0108] method Protein sample preparation The kinase domain construct of FGFR1 (100 μM) was concentrated using a centrifuge (Corning Spin-X OF 30k MWCO) by buffer exchange from a storage buffer containing 40 mM N-(2-hydroxyethyl)piperazine-N'-ethanesulfonic acid (pH 7.5), 200 mM NaCl, and 1 mM tris(2-carboxyethyl)phosphine to a final solution containing 1.5 mM phosphate buffer (pH 8.0) and 5 mM NaCl. Lyophilized powder of BSA (heat shock fraction, Sigma-Aldrich) was dissolved in 5 mM phosphate buffer (pH 8.0) and 5 mM NaCl to a concentration of 1 mM. To a 1 mM protein solution sample, 2 vol% of SU-5402 solution in dimethyl sulfoxide (DMSO) (or pure DMSO) was added to obtain FGFR1 solutions containing 0, 0.5, 0.75, 1, 1.5, and 2 mM SU-5402, as well as BSA solutions containing 0 and 1 mM SU-5402. In the EVV 2DIR experiment, 1 μL droplets of each protein solution were dropped onto a 1-inch diameter circular coverslip, and eight 2 μL droplets of saturated NaCl solution were placed in addition. By mounting the 1-inch diameter circular coverslip on a 1-inch diameter lens mount and sequentially stacking a 1 / 4-inch thick 1-inch diameter rubber O-ring and a 2 mm thick 1-inch diameter CaF2 window, a sealed space with controlled humidity of 75.5% RH was formed between the coverslip and the CaF2 window via the saturated NaCl solution spot. Under these humidity-controlled conditions, the protein droplets slowly dried, forming a flat hydrated gel phase structure. This is similar to the coffee ring structure reported using a similar approach. The height of each gel spot was measured to approximately 10 μm using a digital micrometer.
[0109] EEV 2DIR spectroscopy Since EEV 2DIR spectroscopy is relatively unknown, its operating principle will be briefly explained. This technique measures the coupling between a single-quantum (fundamental vibration) transition and a double-quantum (combination band or harmonic) transition. When an excited combination band contains an equally excited single-quantum transition, the Raman scattering between the two levels is an acceptable transition. Therefore, the visible light beam effectively probes the coupling between two excited vibrations (one at the fundamental frequency and the other at the combination band or harmonic). The resulting spectrum shows a series of features corresponding to the ground transition, which couples with many other vibrations. Diagonal features are also observed, which appear as virtually straight lines. This is because the second single-quantum transition contained in the combination band is constant. The frequency of the second vibration in the combination band can be easily estimated by subtracting the frequency coa of the first excitation beam from the frequency cop of the second excitation beam. This assumes that the frequency shift in the combination band due to the coupling itself is small or negligible, but this is not always the case.
[0110] EVV 2DIR spectrometer The experimental setup for the EVV 2DIR spectroscopy method used in this embodiment is described in detail in other literature. Briefly, a 790 nm near-infrared laser beam is generated using an ultrafast Ti:sapphire laser system (Newport Spectra-Physics). The pulse width is 1 ps and the repetition rate is 1 kHz. The generated beam is split into three arms, two of which drive optical parametric amplifiers (Newport Spectra-Physics) to generate a tunable mid-infrared 1 ps output calibrated using a 38 μm thick polystyrene calibration film traceable to the NIST 1921b frequency. All three beams are linearly polarized parallel to the propagation plane, spatially superimposed, and focused onto the sample. The pulse energy of the 790 nm beam is set to 300 nJ at the sample, and the mid-infrared beam is 1250-1750 cm⁻¹. -1 COA / 2KC range (1500cm) -1(Erna = 800 nJ) and cofi / 2Nc range 2600~3400 cm -1 (3000cm -1 The sample exhibits a nearly symmetrical energy profile over Em = 2.0 / 4. These pulse energies were measured through a 100 μm diameter pinhole. This pinhole diameter closely matches the 1 / e2 diameter of the near-infrared beam, defining the sample region where the four-wave mixing (FWM) signal is generated. A delay stage is used for time control of the laser pulses. The pulse sequence used was coo → cop → coy, leading to the coherence path shown in Figure 1. The emitted FWM signal is detected using a photomultiplier tube detector (Hamamatsu H7422-50) with the input beam removed. The experiment is sealed and purged with N2 to maintain a relative humidity of less than 3%. All EVV 2DIR spectra reported herein were obtained using pulse delay times of zoo = 1.75 ps and TN = 1 ps. These delay times were chosen to provide maximum contrast for the most important features of the spectrum while keeping the non-resonant background sufficiently low and allowing for clear observation of smaller, lower-intensity features.
[0111] The spectrum is in the mid-infrared region at 5 cm. -1 Spectra were collected with a frequency step size of 100 laser irradiations per data point. Spectra were smoothed with a nearest neighbor mean smoothing filter, with arbitrary backgrounds independent of the mid-infrared region subtracted, and spectral intensities normalized by the methylene peak intensity. All spectra shown herein are plotted on a logarithmic intensity scale. When generating difference spectra, since signal intensity is proportional to the square of the sample concentration, the spectra were first squared and then finitely processed.
[0112] Two-dimensional fitting of the spectrum Fitting was performed using MATLAB (registered trademark). Data was fitted using the sum of two-dimensional Gaussian functions by the nonlinear least squares method under the constraint that the variances in the x-direction and y-direction are equal. Initial input of Gaussian function parameters is required for fitting. The initial positions of peaks were selected manually. The initial signal intensity was defined as the spectral intensity at the initial position, and the initial width was the spectral width of the mid-infrared laser pulse (25 cm -1 )
[0113] DFT Calculation and EVV 2DIR Simulation DFT calculation was performed using GAUSSIAN 09 with the 6-31G(d,p) basis set. EVV 2DIR peak calculation from model vibrations was carried out according to the method reported by Kwak et al. (16), which was applied in previous studies on the model system of the present inventors. A frequency scaling factor of 0.96 was used to obtain the frequencies observed in the calculated EVV 2DIR spectra. The calculated peaks are expressed with a full width at half maximum of 25 cm -1 to correspond to the typical width experimentally observed. Spectra were calculated for the FGFR1 / SU5402 system using a model based on the co-crystal structure reported by Mohammadi et al. (33), and contained only protein residues within 5 Å from the SU5402 molecule. A methyl group was added for terminal capping at the unsaturated portions of the model, and water molecules were included within the 5 Å range.
[0114] Results and Discussion This example shows the development of a series of protocols that enable the identification of one or more EVV 2DIR peaks that are clearly associated with the binding of a drug to the active site of a protein. This requires identification of features that are not present when proteins or drugs are measured individually, or when a protein that does not have a specific binding site is mixed with the drug. Since it was found impossible to generate a spectrum of SU-5402 in pure water due to drug aggregation, the background of non-binding control protein at a 1:1 stoichiometric ratio with the target protein was used, which enabled estimation of the spectrum of unbound SU5402.
[0115] In this example, protein BSA was used as a control protein, and a difference spectrum was created by subtracting the spectrum of the protein alone from the spectrum of the drug mixed with the target protein. This made it possible to identify peaks that are present only in the case of specific binding, serving as a means of identifying binding-specific peaks. Another method would be to compare spectra in the presence of organic solvents. While some drugs can bind specifically to BSA, most do not, and instead adhere to multiple sites on the surface with low binding affinity. Empirically, the findings here suggest that SU5402 belongs to the category of non-specific binding. An additional independent method for assigning binding-specific peaks is to titrate the drug to the FGFR1 protein binding site and demonstrate the difference in the predicted titration curve between the peak that appears with binding and the peak derived from non-binding. This provides two independent methods for assigning peaks to the interaction between the drug and the target protein. Notably, both independent methods identified the same six peaks as being due to the specific binding of SU5402 to FGFR, while the remaining 57 peaks were due to the presence of the drug, suggesting that intramolecular coupling is likely dominant.
[0116] Finally, by using quantum mechanical calculations of vibrational coupling at the coupling site and showing the correlation with the observed results associated with these couplings, we provide an interpretation of coupling-specific features.
[0117] EVV 2DIR spectra were measured for both pure FGFR1 and 1 mM 5U5402-containing samples. These can be seen in Figures 1A and 1B, respectively, and are shown along with the difference spectra generated by subtracting the square root of the protein-only spectrum from the square root of the inhibitor-containing protein spectrum. Since the signal is homodyne and proportional to the square of the number of molecules, it is necessary to take the square root to linearize the spectrum with respect to concentration. The dynamic range of the EVV 2DIR data shown here is typically -450.
[0118] A combination of observation and fitting revealed that the FGFR1 protein spectrum (Figure 1A) contains approximately 200 peaks, with their maximum amplitudes at least 36 above the noise floor. Here, 'a' is the standard deviation of the featureless region of the spectrum. The BSA protein spectrum also contains a similar number of identifiable peaks. It is noteworthy that this definition of statistical validity is considerably strict. Signal averaging within the peak region yields a higher signal average than peak maximum value alone, making it highly probable that all identified peaks are real.
[0119] The protein spectrum contains several features already identified in previous studies. For example, coo = 1470 cm². -1 / cofi=2945cm -1 and coo=1640cm -1 / cofi=3300cm -1 The two maximum spectral features observed correspond to the methyl / methylene group and the amide I band, respectively. The feature originating from the phenylalanine side chain is coo = 1470 cm⁻¹. -1 / cofi=3050cm -1 and coo=1480cm -1 / cofi=3070cm -1 This can be confirmed, and the tyrosine peak is at coo = 1530 cm. -1 / cofi=3130cm -1 This can be seen. However, the features most relevant to the purpose of this embodiment are the peaks detected in the SU5402+FGFR1 spectrum but not in the BSA+FGFR1 spectrum, and the amide-related features that appear around -4650 / 3300.
[0120] The characteristics of the amide immediately suggest the action of the drug. Changes in the amide region of the FGFR1 spectrum suggest that FGFR1 undergoes structural rearrangement to the extent that it affects the amide structure, and possibly also affects elements of the protein backbone. On the other hand, the absence of changes in the BSA amide region upon drug addition suggests that the presence of SU5402 has no effect whatsoever on the BSA backbone structure. This is probably because SU5402 does not bind to any specific site on BSA.
[0121] Generally, adding SU5402 to a protein significantly increases the number of detectable peaks. A difference spectrum was created by subtracting spectrum 1A (FGFR1 only) from spectrum 1B (FGFR1 + SU5402) and then fitting it using a two-dimensional Gaussian sum. Excluding amide features, 63 features were observed in this difference spectrum. Similarly, the BSA spectra (1D,E) with and without SU5402 were subtracted to generate a difference spectrum containing 57 Gaussian features. Importantly, all 57 features in the BSA difference spectrum are also present in the FGFR1 difference spectrum, but in the FGFR1 spectrum, six additional peaks are clearly visible. The locations of these features are indicated by pink crosses in the difference spectra (Figures 1C, F).
[0122] The difference spectrum may contain features from three sources: firstly, spectral features resulting from intramolecular vibrational coupling within the SU5402 molecule; secondly, features resulting from changes in the protein spectrum as a result of protein structural changes due to inhibitor binding, such as those observed in the amide band; and finally, features resulting from the coupling of the protein with the drug.
[0123] The 57 peaks common to both the difference spectra of FGFR1 and BSA were attributed to intramolecular coupling within SU5402. This is a reasonable assumption, as BSA does not have a binding site for SU5402. The changes in the broadband amide region of FGFR1 were attributed to protein structural changes due to drug binding. On the other hand, the additional six peaks in the FGFR1 spectrum (Figure 9) were tentatively attributed to changes due to drug-protein interactions. In general, it is impossible to definitively determine whether these new cross-peaks are due to drug-mode and protein-mode coupling, a new collective mode of the bound drug-protein complex, or protein structural changes due to drug binding. Distinguishing between these possibilities motivated the quantum mechanical calculations described later.
[0124] Analysis of the experimental spectrum revealed that five of the six peaks were at 1400 cm⁻¹. -1 It was revealed that the sixth vibration was located near the amide vibration, due to the coupling of one fundamental vibration with various other vibrations.
[0125] To confirm that these six peaks appear in conjunction with binding, we performed verification using additional independent methods to evaluate binding behavior. EVV 2DIR spectra of the FGFR1 / SU5402 mixture were measured at various molar ratios. With the high concentrations of protein and drug used in these experiments, it was expected that the intensity of the peaks appearing in conjunction with true binding would not continue to increase linearly beyond a 1:1 molar ratio, as the binding site would reach saturation. Conversely, features resulting simply from intramolecular coupling within SU5402 should continue to increase linearly in intensity with drug concentration.
[0126] EVV 2DIR spectra were recorded for 1 mM FGFR1 and SU5402 at concentrations of 0 mM, 0.5 mM, 0.75 mM, 1 mM, 1.5 mM, and 2 mM. The difference spectra were calculated, and peak intensities were extracted by fitting the difference spectra as described above. All six peaks present only in the FGFR1 difference spectrum showed a plateau in signal intensity. In contrast, all 65 difference spectral peaks common to both the FGFR1 / SU5402 and BSA / SU5402 difference spectra showed a linear increase in intensity proportional to the inhibitor:protein ratio. This is expected behavior for features unrelated to binding. This confirms that the unbound peaks are intramolecular modes of the drug and are therefore proportional only to the drug concentration. The average fitted peak intensities plotted as a function of the inhibitor-to-protein molar ratio can be seen in Figure 2. For ease of comparison, the data are normalized to 1:1 fitted intensities. The total number of bound proteins continues to increase even beyond a 1:1 molar ratio. This is typical in drug-protein binding curves, as not all ligands are bound at a 1:1 sample ratio. From these data, it is very clear that six peaks are absent in the BSA difference spectrum. Therefore, since both FGFR and SU5402 are present, all of these peaks exhibit clearly different concentration-dependent behavior from the other 57 peaks, which, in addition to their absence in the control measurement using BSA, provides a second means of verification.
[0127] To attribute the cross-peaks associated with these six bonds to the vibrational coupling pairs that produced them, first-principles calculations were performed on the vibrational coupling of the drugs at the binding sites and compared with the case of the individual drugs.
[0128] The structure of the protein-drug complex used in the calculation is shown in Figure 3. To estimate the spectrum of SU5402 not bound to the protein active site, the spectrum of SU5402 in free space was also calculated. Although this free-space calculation does not include water molecules, it was expected that useful insights could be obtained from the difference between the calculated spectra of the bound and unbound species, and this proved to be the case, as can be seen below.
[0129] EVV 2DIR probes the vibrational coupling between the fundamental mode (mode #1 in coa) and its combined vibration (modes #1+#2 in Wp). Therefore, peaks with a common mode #1 appear on the column corresponding to the same co, while peaks with a common mode #2 lie on a diagonal with a slope of 1. By tracking the positions of these lines indicating mode commonality and further examining the motion of the calculated modes, we can find the column corresponding to equivalent mode #1 and the diagonal corresponding to equivalent mode #2 between the two calculated spectra and the experimental difference spectrum of the FGFR1-SU5402 complex. Key lines of mode equivalence are shown in Figure 4. Since each vibrational mode can be involved as both a #1 mode and a #2 mode, peaks can appear on both the column and the diagonal. An example of this is mode C, shown as a magenta line in Figure 4. Harmonic peaks where a given mode forms both a #1 and a #2 mode are located at the intersection of the mode's column and diagonal, forming a new diagonal with a 2:1 slope. This diagonal line is shown as a black dotted line in Figure 4.
[0130] The computational modes that contribute to the peaks that appear with bonding are assigned numbers 1 to 7 to distinguish them from the experimental modes labeled A to E in Figure 4. Images of the modes that show the main atomic motions are shown in Figures 5 to 11.
[0131] As can also be seen in Figure 4, the column of equivalent mode D that appears in both calculated spectra is not found in either of the two experimental spectra. To understand this, it is important to note that the lifetimes of the states involved are not considered in the method used to generate the calculated spectra. In fact, in EVV 2DIR spectroscopy, it is not uncommon for some of the expected cross-coupling between vibrations to be not experimentally observed. This is essentially because some state lifetimes are too short to be observed using a method with an effective time resolution in the ps range. In practice, this means that the experimental spectra are usually slightly sparser than the calculated spectra.
[0132] As is evident from Figures a and c, the frequencies of several equivalent modes are slightly shifted between the two calculated spectra. These shifts are 3–43 cm⁻¹. -1 This range includes, for example, mode A is 1330.03 cm in the joint site calculation. -1 In free-space calculations, 1322.41 cm -1 This appears in the experimental spectra of the bound state and the non-specific bound state. In contrast, no frequency shift is observed between the experimental spectra of the bound state and the non-specific bound state. This may be due to the weak interaction with the protein binding site causing a slight shift in the frequency of the inhibitor mode compared to the inhibitor molecule in free space. The absence of a shift between the experimental spectra of the bound state and the non-specific bound state may be due to the fact that the chemical environment of the inhibitor is more similar between the two states than in free space.
[0133] The calculated binding site inhibitor complex spectrum shows coo = 170 2.5 cm². -1 / cofi=3405.0cm -1 The peaks observed include those attributable to protein modes, such as the amide 1 harmonic peak. The equivalent coo line for the amide 1 mode is shown as mode E in Figures 4a and 4b. This amide 1 mode corresponds to coo = 1702.5 cm in Figure 4a. -1 / cofi=3261.15cm -1 The peak of this calculation was 1558.62 cm⁻¹. -1It can be confirmed that it is coupled with the inhibitor mode. This inhibitor mode produces the diagonal of the peak shown as mode C in Figure 4. Mode C corresponds to the amide-related motion of the 5-membered lactam ring and the motion on the adjacent benzene ring. The experimental peak that appears with bonding (coo = 1660 cm) -1 / cofi=3240cm -1 This can be attributed to the coupling between the lactam ring mode (mode C) and the protein amide 1 mode (mode E), with a calculated peak of 1702.5 cm⁻¹. -1 / cofi=3261.15cm -1 This is equivalent to [the above]. This vibrational coupling is mediated by a hydrogen bond formed between the NH group of the SU5402 lactam ring and the amide carbonyl group of the amino acid near the binding site.
[0134] coo=1400cm -1 The computational equivalent peaks corresponding to the five experimentally observed binding sensitivity peaks were observed in the computational spectra both with and without the protein binding site. These are shown as Mode B in Figure 4. The vibrational modes calculated to produce these showed motion only on the SU5402 molecule. Therefore, the experimental binding dependence clearly seen in the intensity of these peaks is thought to be due to lifetime effects. In the unbound case, the SU5402 molecule is in a different chemical environment than the molecule bound to the kinase binding site and is naturally exposed to different environmental fluctuations. For this reason, in the case of nonspecific binding, the phase relaxation of these vibrational coherences is increased and the lifetime is shortened compared to the protein inhibitor complex. Mode B corresponds to a complex motion with atomic displacements across the entire molecule, from the fused benzene ring to the methyl substituent of the pyrrole ring. The broad nature of this mode, spanning many chemical groups, was hypothesized to increase the chances that "environmental" coupling (e.g., coupling with water) influences the coherence lifetime through pure or collective phase relaxation. However, considering that most of the modes here are dispersed across many chemical groups, it was somewhat surprising to find that there were even more modes that were not affected in the same way.
[0135] Figures 5 to 11 show the major atomic displacements. In summary, the displacement labeled B in Figure 4 corresponds to mode 1 in Figure 5, which is 1400 cm². -1 The mode is entirely located on SU5402, and the appearance of this row of cross peaks upon binding supports the idea that the isolation of SU5402 at the binding site alters the phase relaxation rate of this mode. This is presumed to be due to the elimination of solvent interactions that shorten the state lifetime. All other modes are collective motions of the drug and protein binding site, indicating significant motion of the amide group.
[0136] In a sense, the most striking feature of this data is that the intramolecular modes of SU5402, excluding the amide-related cross-peaks, do not disappear when bound to FGFR1 compared to the unbound state using BSA, nor do they show a significant frequency shift. This suggests that the average structure of the solvated drug is very similar to the drug structure at the binding site, to the point of being almost identical. This certainly reduces the conformational entropy cost associated with binding, making the drug a more potent inhibitor, and is considered a relatively common feature of heterocyclic aromatic drugs that tend to have relatively rigid planar structural components.
[0137] When SU5402 was added, a significant change occurred in the amide region of the FGFR spectrum, but no change was observed in the amide region when the drug was added to BSA. This further emphasizes the usefulness of BSA as an unbound control and suggests that a portion of the FGFR structure is stabilized by SU5402 binding. There is excellent precedent for this. Many kinases contain a regulatory loop containing an Asp-Phe-Gly sequence known as the DFG loop. This loop can exist in either an "in" form within the protein, where phenylalanine is located deep within a hydrophobic pocket, or an "out" form, where phenylalanine protrudes into the surrounding environment. SU5402 is a type 1 kinase inhibitor that binds and stabilizes DFG only when it is in the "in" state, and it is possible that a net change in protein structure actually occurs upon binding. From the data, it is impossible to determine whether what is observed is actually a transfer of DFG or whether other structural changes are being stabilized by SU5402 binding. Since proteins explore various structural configurations in equilibrium, it is expected that changes in this structural equilibrium will be observed upon drug binding. These observations highlight that EVV 2DIR is sensitive to both overall structural changes in the protein and local coupling changes due to drug-protein contact.
[0138] One of the peaks that appeared upon binding was almost certainly attributed to drug-protein contact. This was experimentally determined to be coo = 1660 cm². -1 / cofi=3240cm -1This feature was observed. Its proximity to the amide 1 mode suggests it is likely due to a known hydrogen bond between SU5402 and the carbonyl group of the protein backbone. This intuitive attribution was supported by quantum mechanical calculations that predicted the presence of this feature. It arises from the coupling between a drug mode and a protein mode that contains a significant amount of the carbonyl group's motion amplitude at the binding site (see Figure 12). This carbonyl group is a chemical group known to hydrogen bond with drug molecules. Since this is a common feature in many kinase-drug interactions with drugs targeting the ATP binding site, this feature was expected to be a general and versatile means of identifying and monitoring the details of drug binding to the kinase active site.
[0139] In selecting the most crowded spectral region available in EVV 2DIR spectroscopy—that is, the region containing harmonic coupling—separating peaks that appear with binding from features derived from non-binding was a particularly challenging task. The reasoning is that if it's possible in such a crowded spectrum, it should certainly be possible in less crowded spectral regions. Despite the relatively high level of crowding with over 200 protein peaks and over 60 drug peaks, both spectra are well-organized and possess the precision to identify features associated with six bindings. The spectral region measured here represents approximately 1 / 20 of the entire spectral region accessible by the EVV 2DIR instrument used in this embodiment. Therefore, a broader investigation of the entire accessible spectral space would likely yield features associated with other bindings. Indeed, calculations suggest that in the broader accessible spectral space, there are at least 12 additional peaks associated with binding resulting from vibrational coupling between SU5402 and FGFR1 binding site modes. A considerable number of experimentally binding-dependent features, such as changes in the entire protein structure in response to drug binding, are also likely to exist. Since this study only calculated vibrational coupling at the drug binding site, these predictions were not possible with the approach adopted in this embodiment.
[0140] The ability of EVV 2DIR to detect drug binding to commercially relevant target protein binding sites, as demonstrated herein, suggests the potential usefulness of this technology in future drug discovery. The spectra shown here are the result of signal averaging over 90 minutes. The spectrometer was driven by a 1 kHz laser engine, and data were acquired at a rate of 1 kHz. Currently, there are laser engines capable of driving EVV 2DIR spectrometers at 100 kHz, which could potentially reduce the total spectral acquisition time to approximately 1 minute. Furthermore, if only a portion of the spectrum (such as the amide region) is needed in binding studies, the acquisition rate for specific drug-protein combinations could be reduced to a few seconds. It is entirely possible to establish a high-throughput version of this experiment, and the high data density generated would allow for data mining into "contact" modes that can be attributed to a class of protein-drug contacts without requiring detailed structural analysis.
[0141] Direct structural analysis also shows promising prospects. The inventors have previously shown that the geometry of mode interactions can be determined by collecting data under two different polarization states of the beam. Since each peak appearing with binding provides a single angular constraint, only a few such features are sufficient to obtain very strict structural constraints for the drug-protein interaction geometry. Similarly, it has been shown that the intensities of these peaks can be used to determine interatomic contact distances, providing further structural detail.
[0142] conclusion The inventors demonstrated the ability to detect drug binding to a clinically important drug target, namely the protein kinase FGFR1. One of the six peaks associated with binding can be directly attributed to the coupling that occurs when the drug binds to the carbonyl group of the protein's active site. This attribution is supported by first-principles calculations of vibrational coupling at the enzyme's active site occupied by the drug molecule. [Examples]
[0143] Analysis and interpretation of the transition of G-quadruplicate DNA from unfolded to folded state using electron vibration two-dimensional infrared spectroscopy (EVV 2DIR). In this embodiment, the inventors demonstrate that two different structural hypotheses regarding the geometry of interacting chemical entities can be distinguished by combining EVV 2DIR spectroscopy with calculations of EVV 2DIR spectra for two hypotheses. The two structural hypotheses cross-validated here are one that guanine nucleotides are bound and the other that guanine nucleotides are not bound to each other. Since the presence of potassium ions can cause structural changes in this sample, these two hypotheses were tested in samples containing various different ion concentrations to determine the potassium concentration at which the guanine complex is formed.
[0144] G-quadrivalent (G4) structures are non-standard DNA secondary structures formed in guanine-rich DNA sequences. G4s can be monolayer, bilayer, or tetralayer, and their structure varies depending on relative strand orientation, loop length and composition, and glycosidic bond rotation. Using next-generation genome sequencing, more than 700,000 G4 sequences have been identified in the human genome. Both experimental methods and computational algorithms suggest that a high percentage of G4-forming sequences are found in gene promoter regions and telomeres. The formation of G4 structures in these regions is thought to suppress cancer cell proliferation through inhibition of telomere maintenance and transcriptional inhibition of oncogenes. Therefore, the use of G4-stabilizing ligands is being investigated as a potential molecular therapeutic approach for suppressing and eliminating cancer cells. For this reason, structural knowledge of G4 and G4-ligand complexes is useful as an aid in ligand design.
[0145] In this example, we studied the structural changes occurring in Myc2345 upon potassium ion addition by combining EVV 2DIR spectroscopy and quantum chemical calculations based on density functional theory. This addition causes the thermodynamic equilibrium position to shift from the unfolded state to the folded state G4. The 22-mer DNA sequence Myc2345 forms an intramolecular propeller-type parallel strand G4 in the presence of potassium ions.
[0146] method EVV 2DIR spectroscopy The apparatus featured a dual Ti:sapphire amplifier (Thales Laser) synchronized by a single oscillator seed light source (Femto Laser), generating two beams of 800 nm, 10 kHz, and 0.8 mJ. These beams had FWHM pulse widths and bandwidths of 1 ps and 20 cm, respectively. -1 (α) and 50fs·300cm -1 (β) was obtained. The beam center frequency was adjusted using an optical parametric amplifier (pump output 4W, TOPAS manufactured by Light Conversion), and ω α 1250~1750cm -1 Step size 5cm within the range -1 Change it with ω β It is 3200cm -1 The settings were configured as follows: The gamma beam was generated by passing one arm of the fs beam through an etalon, with a bandwidth of 12 cm. -1 The beam was set to 800 nm, 1 ps pulse. All three beams were focused onto the sample in an arrangement that satisfied the phase matching condition. The delay time between each pulse was set using a delay stage, T αβ =T βγ The value was set to 0.4 ps. The FWM signal was detected by combining a spectral graph and a charge-coupled device (CCD) camera.
[0147] Sample preparation Myc2345, d(5'-TGAGGGTGGGGAGGGTGGGGAA-3' (SEQ ID NO:2) was purchased from Eurogentec and used without further purification. Myc2345 was dissolved at a concentration of 1 mM in bis-tris buffer solution (10 mM, pH 7.1) in nanopure water. KCl was added to the sample at concentrations of 0, 5, 10, and 100 mM. The sample solution was vortexed for 10 seconds, annealed at 95°C for 2 minutes, and then cooled to room temperature. A gel spot was formed by dropping 1 μL of the sample solution onto a coverslip and drying it in a sealed sample cell at 85% relative humidity, which was prepared by dropping a 6 × 2 μL spot of saturated KCl aqueous solution onto the coverslip. The sample cell consisted of a coverslip, a 1 / 4-inch thick spacer, and a 2 mm thick CaF2 disk as a front window. Due to the shrinkage of the spot during gel phase formation, the final potassium concentration was higher than that of the initial solution.
[0148] Data processing Glass: To explain the different OPA output energies for each wavelength range, T αβ =T βγ The sample spectrum was normalized using the cover glass spectrum measured with a delay time of =0. The glass spectrum was smoothed by nearest neighbor averaging. β -ω α Interference fringes were observed along the axis. To remove the fringes, the number of data points was maintained while ω β -ω α Linear interpolation was applied along the axis. A low-pass brick-wall filter with a cutoff frequency of 0.05 was then applied to the Fourier transform of the glass spectrum. The line cuts of the glass spectrum after each step are shown in Figure 31.
[0149] DNA sample spectrum: The sample spectrum was smoothed by nearest neighbor averaging. To match the glass spectrum, ω β -ω α Linear interpolation was applied along the axis. For the spectrum, ω α = 1325~1375cm -1 , ω β -ω α= 1190~1210cm -1 Background removal was performed by subtracting the average intensity in the region. Subsequently, the spectrum was background-removed and normalized based on the glass spectrum. Since the FWM signal is proportional to the square of the analyte, the spectral intensity was squared. The spectrum also had the highest intensity peak (ω α = 1525cm -1 , ω β -ω α = 1450cm -1 The spectra are normalized by ). Each sample spectrum is the average of three measurements.
[0150] Fitting using a two-dimensional Gaussian function To estimate the number of peaks in the experimental two-dimensional spectrum, line cuts at different frequencies were analyzed. This yielded reasonable estimates for the number of Gaussian functions (102) to be used for fitting with two-dimensional Gaussian functions. Each spectrum was fitted with the sum of 102 two-dimensional Gaussian functions using the Python nonlinear least-squares fitting function scipy.optimise.curve_fit. The Trust Region Reflective algorithm was used to minimize the chi-squared value. Errors in the fitting parameters were propagated from the residuals of the best fit. The 100mM spectrum was initially fitted without fixed parameters. The initial estimated positions were based on a visual inspection of the spectrum line cuts, and these were within ±15cm. -1 A range of variation was allowed. The initial estimated intensity was a randomly generated value between 0 and 1, and variation within the same range was allowed. The initial estimated width was 25 cm for all peaks in both dimensions. -1 5-40cm -1A range of variation was allowed. Subsequently, each sample spectrum was fitted using the parameters optimized with an initial 100 mM fitting as initial estimates. All parameters were fixed except for the intensity (variability allowed in the range of 0 to 1) and the position of peaks where frequency shifts were observed depending on the potassium ion concentration at the spectral line cut. These shifting peaks were fixed at ±30 cm in the dimension where the shift was observed. -1 A range of variation was permitted.
[0151] EVV 2DIR spectrum calculation The calculated EVV 2DIR spectra were obtained using a method based on Gaussian software (Gaussian 16, Revision B.01).
[0152] The third-order nonlinear susceptibility χ is the central quantity of the EVV 2DIR spectral intensity. (3) The relevant contributions are related to the anharmonicity in the polarization characteristics (called electrical anharmonicity) and the anharmonicity in the oscillation potential (called mechanical anharmonicity), and the contributions due to electrical anharmonic are determined by a perturbation theory development that is truncated up to the first order of anharmonicity.
[0153]
number
[0154]
number
[0155]
number
[0156]
number
[0157]
number
[0158]
number
[0159]
number
[0160]
number
[0161]
number
[0162] The G-quartet and G-quadruplicate molecular structures in a vacuum were first constructed using PyMol and then optimized at the density functional theory (DFT) level using Gaussian. The theoretical level used was the B3LYP functional, the basis set was 6-31G(d,p), the lattice refinement parameter was Int=UltraFine, and the convergence criterion was SCF=VeryTight. The optimized G-quartet and G-quadruplicate structures were then optimized using Gaussian with the same theoretical level and basis set as the geometry optimization, determining the harmonic oscillation normal mode frequency ω i The dipole moment, polarizability, and the first derivative of the dipole moment in normal mode were calculated. These first derivatives were obtained by numerically differentiating the dipole moment, the first derivative of the dipole moment, and the polarizability calculated in the geometry with normal mode displacement using a custom script, thereby obtaining the derivative quantities necessary for calculating the EVV 2DIR spectral intensity. For the vibrational energy levels, a harmonic region was assumed with a uniform scaling parameter of 0.96, so the sum of the single excitation energy levels ω i +ω j Combinations including the same normal mode index / harmonic energy levels ω (i+j) No difference was found between the two. The obtained data was then converted to two-dimensional spectral intensity using a custom script, and each peak was represented by a linewidth of 20 cm. -1 The (FWHM) was represented as a two-dimensional Gaussian linear model and then visualized using the Matplotlib Python library.
[0163] result K at 0mM, 5mM, 10mM, and 100mM +Figure 16 shows the experimental 2DIR spectrum of 1 mM Myc2345 in the presence of EVV. In EVV 2DIR spectroscopy, one infrared pulse excites one vibrational band, and the other excites another band. The only situation considered was the case in the normal mode factorization of the vibrational wavefunction where the first pulse induces a single quantum excitation to a single excited state, and the second pulse induces two quantum excitations to a double excited state, with at least one mode exponent corresponding to the mode exponent of the single excited state. Coupling is read out by the visible pulse; coherent Raman scattering occurs when the two excitation bands are coupled, but not when they are not coupled. The measured four-wave mixed signal is the Raman excitation frequency ω γ Subtracting ω, the plot is generated. α The difference frequency axis ω as a function of the (X axis) β -ω α (Y-axis) was obtained. ω β Because the combination band is excited, the two axes roughly correspond to the frequencies of the two fundamental vibrations contributing to the combination band. This is why feature rows and columns are observed in the data. These correspond to a particular vibration mode being coupled with several other modes. Due to the high coupling density, multiple coupling peaks overlap in these spectra. In other words, the peak position does not necessarily indicate the location of the coupling, but may be a peak resulting from the presence of multiple couplings.
[0164] In order to organize these complex spectra as much as possible, the 2DIR spectra were fitted with a two-dimensional Gaussian function. The two-dimensional Gaussian function well approximates the shape of typical EVV 2DIR coupling features. In this case, at least 102 two-dimensional Gaussian functions are required, revealing that there are at least 102 distinct couplings. The fitted 2DIR spectrum is shown in Figure 17. Markers indicating the center frequencies of cross-peaks obtained from the fitting are overlaid on the spectrum. As expected, these do not always perfectly align with the peaks in the unprocessed raw data. A vertical line-cut plot showing a comparison between the raw data and the fitted spectrum is provided in Figure 18, which demonstrates the high signal-to-noise ratio of the dataset and the good quality of the fitting. A comparison between the raw data and the fitted data at other frequencies, along with the chi-square values, is shown in Figure 25. In this section, we focus on K + + frequency shifts, intensity changes, and new vibrational couplings resulting from structural changes induced by ion addition.
[0165] As shown in Figures 19a and 19b, with an increase in K + + concentration, it was observed that multiple peaks in the in-plane fundamental vibrational region (around 1700 cm -1 -1) undergo a red shift. Peaks 12, 30, and 58 red shift along the ω β 1-ω α axis, peaks 74, 77, 85, and 87 red shift along the ω α 3 axis, and peak 80 red shifts along both axes. The observed red shifts are consistent with the weakening of guanine carbonyl bonds as Hoogsteen hydrogen bonds form and K + + coordination occurs. Interestingly, the ω β 1-ω α 3 frequency corresponding to guanine carbonyl stretching appears at a higher frequency than the ω α 3 frequency. In the absence of potassium ions, the ω β 1-ω α 3 frequency associated with guanine carbonyl stretching ranges from 1694 to 1707 cm -1 -1, while the ω αThe frequency was observed at 1671~1688 cm -1 Interpretation of this extremely unusual observation result can be found in the discussion section. Furthermore, the frequency of guanine carbonyl stretching observed in the absence of potassium ions is normally 1660~1673 cm -1 which is higher than the frequency of single-stranded guanine observed at . This observation suggests that some structure may pre-exist even in the absence of cations.
[0166] It is well known that nucleic acids exhibit absorbance changes (hypochromism and hyperchromism) in the UV / Vis region of the spectrum, which is considered to result from base stacking interactions. Hypochromism and hyperchromism are also reflected in the intensity of Raman lines for ring modes, and have conventionally been used to monitor conformational changes in nucleic acids. In the EVV 2DIR spectrum, among peaks whose intensity monotonically increases or decreases, we further focused on + peaks exhibiting an intensity change of at least 20% between the 0K + spectrum and the 100 mM K + spectrum. Twenty-nine peaks whose intensity decreases as the K concentration increases were identified (Figs. 19d-i). Most of these peaks appear in the frequency range characteristic of guanine ring vibration (1476~1495 cm -1 , 1564~1568 cm -1 , 1575~1590 cm -1 ). Furthermore, five peaks were observed to increase in intensity as the K + concentration increases (Fig. 19c). Assignment of these features and explanation of the mechanism of increase / decrease in coupling are provided in the discussion. The intensity traces of the remaining peaks are shown in Fig. 26 as a function of K + ion concentration.
[0167] When G4 folds from a relatively disordered structure to a relatively ordered structure, it is expected that several new coupling peaks will be generated as chemical groups that were previously separated approach each other. Figs. 20c and 20d show the newly observed peak 101 (ω α = 1705 cm -1 , ω β -ωα = 1675cm -1 ) and 92(ω α = 1691cm -1 , ω β -ω α = 1132cm -1 The raw data and fitted 2DIR spectra for each region are shown. Figure 20b shows the dependence of both peaks on intensity and ion concentration. Novel vibrational modes spanning multiple G-quartets may lead to the emergence of new peaks due to their higher delocalization levels. These peaks may function as markers of G4 formation. The assignment of these peaks will be discussed in the discussion section.
[0168] Consideration One of the initial surprising aspects of the experimental data was the relatively small number of entirely new couplings resulting from G4 folding. While this behavior can sometimes be observed in computational spectra such as those in Figures 21a and 21b, it is worthwhile to consider the reasons for this. Chemical parts such as DNA bases contain many chemical "groups" that link to each other in various ways. Therefore, the formation of a quartet of four bases linked by Hougstin-type hydrogen bonds can be characterized primarily by the rearrangement / recombination of the vibrational states of the individual bases (and the couplings between them) rather than the emergence of entirely new couplings not characterized by this method. This is roughly analogous to the situation of molecular exciton coupling, where the number of electronic states does not increase, but rather the electronic states are rearranged by new couplings.
[0169] One model for G4 formation involves a transition between two states: a single-chain state in the absence of ions and a fully folded G4 state in the presence of excess potassium ions. However, the frequencies associated with guanine carbonyl stretching in the EVV 2DIR spectrum of Myc2345 in the absence of ions do not match the characteristics of single-chain guanine. Nevertheless, several observations—redshift of the frequency of guanine carbonyl stretching vibrations, observed changes in cross-peak intensity, and the appearance of two new couplings—suggest the formation of higher-order structures in response to potassium ion addition. To gain a deeper understanding of the structures observed in the absence of ions and the spectral changes observed with increasing ion concentration, EVV 2DIR spectra were calculated for G4 and its candidate precursors. Furthermore, a uniform scaling parameter was applied to the vibrational energy levels, and calculations were performed using the harmonic vibrational potential approximation; therefore, the difference between the energy levels of the overtone / combination bands and the sum of the corresponding single quantum energy levels was not considered. Finally, the calculations were performed in a vacuum, and the experimental samples were in the presence of abundant water, but no implicit or explicit solvation was involved. Overall, the reliability of the findings related to the analysis of quantum chemical calculations in this study should therefore be considered with these limitations in mind. The overall strategy of this study was to identify and distinguish three possible structural types: isolated guanine not part of a quartet, guanine arranged in a quartet, and a guanine quartet stacked to form a G-quadrichain.
[0170] In the absence of potassium ions, the initial DNA structure resembles a G-quartet and contains guanine. Multiple pieces of evidence suggested that the DNA structure was present in the absence of ions. The first piece of evidence came from observing the frequency difference between rows and corresponding columns of peaks resulting from the coupling of the guanine carbonyl stretching mode with various other vibrational modes (Figure 16a). In general, in EVV 2DIR spectroscopy, rows and columns of cross-peaks associated with the same mode are harmonic lines (ω α =ω β -ω αIt has approximate reflection symmetry with respect to ω. However, this symmetry is ω β -ω α It is frequently lost due to a negative frequency shift along the axis. This is because, due to the process of mode coupling, the combined bands are usually shifted to frequencies lower than the sum of the fundamental frequencies. In the experimental spectrum (Figure 16a), the row of cross peaks with frequencies dependent on potassium ion concentration is 1694–1707 cm⁻¹. -1 This frequency range was attributed to guanine carbonyl stretching vibrations. This frequency range, which also depends on potassium ion concentration, is 1671–1688 cm⁻¹. -1 These peaks were unusually high compared to the corresponding peaks that appeared. Assuming these are not combinational bands shifted to frequencies higher than the sum of the constituent fundamentals, it is suggested that the guanine carbonyl stretching mode (GCO I) that gives rise to this row is not identical to the mode (GCO II) that gives rise to the row.
[0171] The appearance of two different guanine carbonyl modes, which normally give rise to corresponding rows and columns of cross-peaks, is inconsistent with the spectroscopy of isolated guanine bases (Figure 27). However, the calculated EVV 2DIR spectra and associated vibrational analyses show similar behavior in structures containing coupled guanine bases, such as the G-quartet (Figure 21a) and G4 (Figure 21b). This is because the nearly symmetrical G-quartet structure gives rise to multiple combinations of guanine stretching modes in a manner similar to simple exciton coupling. Many different guanine carbonyl center modes contribute to the columns of cross-peaks (4 in the G-quartet in Figure 21e, and 12 in G4). However, the coupling intensities of the cross-peaks associated with each carbonyl mode differ, with one mode showing much stronger coupling than the others. Interestingly, the carbonyl stretching mode dominant in the rows of cross-peaks is not identical to the mode dominant in the corresponding columns of cross-peaks. Therefore, in these calculated spectra, the rows and columns actually originate from different combinations of guanine motion. The relative difference in the dominant row and column center frequencies shows a sign difference between the experimental spectrum and our calculated spectrum. This observation can be interpreted as supporting the idea that the observation of different guanine stretching modes in the rows and columns of the EVV 2DIR spectrum of Myc2345 in the absence of ions may be due to the presence of a pre-existing structure containing coupled guanine bases. In effect, EVV 2DIR is a technique for extracting the difference in oscillator intensity between one quantum state of a G-quartet, where one vibrational mode is excited first in the row and another vibrational mode is excited first in the corresponding column. The “excitonic” interactions of the original isolated guanine motions can lead to one-quantum transitions of different oscillator intensities when they combine “excitonically,” so situations can arise where one state dominates a column and another state dominates a row.
[0172] Further evidence indicating the existence of pre-existing structures even in the absence of ions was obtained from the absolute frequency of the guanine carbonyl stretching vibration. The frequency of the carbonyl stretching vibration of nitrogen-containing bases is sensitive to the conformation of DNA and is often used to determine the secondary structure of DNA. The frequency of the guanine carbonyl stretching vibration is ω α = 1671~1688cm -1 Observed in (GCO II), which is typically observed in single-chain guanine at 1660–1673 cm². -1 This value is higher than that. This could suggest the presence of an existing DNA structure. In fact, circular dichroism measurements of MycL1 (see Figure 28 for sequence comparison), a G4-forming sequence similar to Myc2345, also suggest that some structure may exist in the G4 DNA sequence even in the absence of potassium ions.
[0173] Given the above suggestion that some structure similar to the G-quartet existed even in the absence of potassium ions, the obvious question arose as to whether the entire G4 structure could be pre-formed. Evidence that a fully stacked G4 complex does not exist in the absence of potassium ions was obtained from four observations. First, in geometry optimization, the stacked quartet without potassium ions was found to be energetically unstable. This structure is undesirable due to the repulsive forces between the 12 carbonyl oxygens. This is a pure enthalpy effect, and in a real solvent system, entropy forces only decompose the structure, so the existence of a G4-like structure in the absence of potassium ions was considered unlikely. Three other pieces of evidence indicating the absence of a stacked G-quartet (complete G4 structure) in the absence of potassium ions were the redshift of specific frequencies observed upon potassium ion addition, the intensity changes of multiple peaks associated with potassium addition, and the formation of new spectral features known to be associated only with fully stacked G4. These will be discussed in more detail below.
[0174] The redshift of the guanine carbonyl stretching vibration indicates ionic coordination. G4 formation has traditionally been monitored by observing frequency shifts in the guanine carbonyl stretching vibration frequency. In this example, redshift was observed in both the guanine carbonyl stretching modes GCO I (Figure 19a) and GCO II (Figure 19b) upon potassium ion addition. The oblique shift of peak 80, which is the coupling between GCO I and GCO II, also supported this attribution. However, according to existing literature, the frequency shift of the guanine carbonyl stretching vibration due to folding from a single strand to G4 is in the opposite direction. To resolve this contradiction and understand the experimental observations, the calculated EVV 2DIR spectra of the proposed G-quartet (Figure 21a) and G4 (Figure 21b) initial structures were compared with the Myc2345 experimental spectra in and without potassium ions.
[0175] The experimentally determined redshifts of the frequencies of guanine carbonyl stretching vibrations due to potassium ion doping are shown in Figures 19a and 19b. The frequency redshifts of GCO I and GCO II were nearly reproduced in the calculated difference spectra of the G-quartet (Figure 21a) and G4 (Figure 21b). Experimentally, the mean frequency shifts of GCO I and GCO II were -8 cm, respectively. -1 and -10cm -1 In contrast, the calculation frequency shift was -2cm each. -1 and -25cm -1 Therefore, the calculations qualitatively agree with experimental observations, even considering that they are based on harmonic vibrational wave functions, suggesting that potassium ion complexing causes a redshift overall. The mechanism of the redshift in the frequency of the guanine carbonyl stretching vibration can be attributed to ionic coordination to the carbonyl oxygen and the stacking of G-quartets. Ionic coordination elongates the carbonyl bond from 1.23 Å to 1.25 Å. Furthermore, the resulting stacking of G-quartets causes a shift of 1704 cm⁻¹. -1 The guanine carbonyl mode observed in the G-quartet (Figure 21g, h) is 1679 cm⁻¹ in G4. -1 The equivalent mode observed is delocalized across all three quartets.
[0176] The change in cross-peak intensity suggests a pre-existing G-quartet without a complete G4 structure. Experimentally, it was confirmed that the intensity of numerous cross-peaks decreases with the addition of potassium ions. Since this intensity decrease also appears in calculations, this observation can be used to validate folding structure models. Figure 22 shows the measured ratio spectrum with potassium ion addition and the calculated G-quadruplicate / G-quartet ratio spectrum. Both show similar characteristics. Excluding difference peaks due to frequency shifts and new couplings, negative differences predominate in both spectra, and ω α = 1600cm -1 A positive difference peak exists in or below this range. The overall decreasing trend in the intensity of this cross-peak is consistent with the G-quartet to G-quadruplicate model, although this is accompanied by a comparison with a calculated spectrum based solely on electrical anharmonics.
[0177] The change in cross-peak intensity observed during the transition from the G-quartet (Figure 21c, d) to G4 (Figure 21e) can be explained by a change in the relative orientation of guanine bases within the quartet. In the G-quartet, the molecular planes of adjacent guanine bases intersect at an angle of 21 degrees. In G4, this angle changes to 19 degrees in the terminal quartet and to 0 degrees in the central planar quartet (Figure 30). These structural changes result in a change in the orientation of the transition dipole moment, which explains the change in cross-peak intensity observed across the entire spectrum.
[0178] Peaks 92 and 101 are markers for the fully folded G4 in Myc2345. The new cross-peaks that appear with increasing potassium ion concentration are potential markers of fully folded G4. Two new peaks formed by potassium addition were identified. Peak 92(ω α = 1691cm -1 , ω β -ω α = 1132cm -1 ) and peak 101(ω α = 1705cm -1 , ωβ -ω α = 1675cm -1 Neither peak appeared in the calculated G4 spectrum, suggesting that both peaks may be involved in vibration modes outside the core G4 structural elements included in the calculation (Figure 21e). Using the calculated vibration frequencies of the G4 structure (Figure 21e), the ω of both peaks was calculated. β -ω α The frequencies were tentatively attributed to vibrations of the core G4 structure. Specifically, peak 92 represents guanine ring deformation of the terminal quartet, and peak 101 represents NH bending vibrations and C=O stretching vibrations of all quartets. The ω of both peaks α The frequency is characteristic of G C6=O6 stretching or T C2=O2 stretching, which may be located within the G / T of the Myc2345 end, which was not included in the calculations (Figure 23). It has already been suggested in other studies that the base at the G4 end interacts with and further stabilizes the core G4 structure. The NMR structure of Myc2345 (PDB: 7KBV) is consistent with this idea.
[0179] Overview of the Comparison of Measured and Calculated Spectra Figure 22b shows the calculated ratio spectrum of the G4 / G-quartet. This spectrum was compared with the experimentally measured ratio spectrum (Figure 22a) obtained by adding 100 mM potassium ions. Qualitatively, several important elements present in both the calculated and measured spectra stood out. Firstly, adjacent to the low-frequency positive column, approximately 1600 cm⁻¹ -1A pattern of negative columns was observed. Secondly, due to the mode combination effect outlined in Figure 21, such a pattern is absent in the corresponding row. Thirdly, the cross-peak intensity was generally reduced in both the calculated and measured spectra. Based on these observations, it was concluded that there is evidence suggesting that Myc2345 forms a non-stacked G-quartet in the absence of ions, and Figure 23 shows what such a structure might look like. In this model, upon potassium ion addition, the G-quartet stacks to form another structure, which is presumed to be a G-quadruplicate structure. Furthermore, the appearance of two new coupling peaks suggests that guanine and / or thymine bases at the ends of the DNA sequence are likely coupling with the stacked quartet in this structure.
[0180] conclusion During the folding process of the Myc2345 G-quadruplicate, the distinct and recognizable vibrational coupling behavior of 102 was tracked using EVV 2DIR spectroscopy. These features were titrated as potassium ion concentration functions. During the folding process, the initial 102 peak was modulated in several ways, including changes in intensity and frequency shifts, as well as the formation of two new coupling peaks. Vibrational analysis using quantum chemistry (density functional theory) and calculations of 2DIR spectra were utilized to elucidate the attribution of spectral features and the causes of their changes. Overall, evidence was found suggesting that Myc2345 forms a structure similar to a non-stacked G-quartet in the absence of ions, and that upon potassium ion addition, it forms another structure presumed to be a G-quadruplicate. More generally, it was demonstrated that detailed information regarding short-range coupling can be obtained and interpreted even in the absence of long-range order. Secondly, it was shown that high-fidelity EVV 2DIR spectroscopy enables high-density acquisition of interpretable features.
[0181] (References) 1. Mampallil, D.; et al. Acoustic suppression of the coffee-ring effect. Soft Matter 2015, 11, 7207-7213. 2. Petti, M. K; Lomont, J. P.; Maj, M.; Zanni, M. T. Two-Dimensional Spectroscopy Is Being Used to Address Core Scientific Questions in Biology and Materials Science. J. Phys. Chem. B 2018, 122, 1771-1780. 3. Donaldson, P. M.; et aL Direct identification and decongestion of Fermi resonances by control of pulse time ordering in twodimensional IR spectroscopy. J. Chem. Phys. 2007, 127, 114513. 4. Fournier, F.; et al. Biological and Biomedical Applications of Two-Dimensional Vibrational Spectroscopy: Proteomics, Imaging, and Structural Analysis. Acc. Chem. Res. 2009, 42, 1322-1331. 5. Guo, it; et al. Detection of complex formation and determination of intermolecular geometry through electrical anharmonic coupling of molecular vibrations using electron-vibration-vibration two-dimensional infrared spectroscopy. Phys. Chem. Chem.Phys. 2009, 11, 8417-8421. 6. Klein, T.; et aL Structural and dynamic insights into the energetics of activation loop rearrangement in FGFR1 kinase. Nat. Commun. 2015, 6, 7877. 7. Fournier, F.; et al. Optical fingerprinting of peptides using twodimensional infrared spectroscopy: Proof of principle. Anal. Biochem. 2008, 374, 358-365. 8. Fournier, F.; et al. Protein identification and quantification by two-dimensional infrared spectroscopy: Implications for an all-optical proteomic platform. Proc. Natl. Acad. Sci. U.S.A. 2008, 105, 15352-15357. 9. Guo, R; Mukamel, S.; Klug, D. R. Geometry determination of complexes in a molecular liquid mixture using electron-vibration-vibration two-dimensional infrared spectroscopy with a vibrational transition density cube method. Phys. Chem. Chem. Phys. 2012, 14, 14023-14033. 10. Spiegel, J., Adhikari, S. & Balasubramanian, S. The Structure and Function of DNA G-Quadruplexes. Trends Chem. 2, 123-136 (2020). 11. Ma, Y., Iida, K. & Nagasawa, K. Topologies of G-quadruplex: Biological functions and regulation by ligands. Biochem. Biophys. Res. Commun. 531, 3-17 (2020). 12. Chambers, V. S. et al. High-throughput sequencing of DNA G-quadruplex structures in the human genome. Nat. Biotechnol. 33, 877-881 (2015). 13. Huppert, J. L. & Balasubramanian, S. Prevalence of quadruplexes in the human genome. Nucleic Acids Res. 33, 2908-2916 (2005). 14. Huppert, J. L. & Balasubramanian, S. G-quadruplexes in promoters throughout the human genome. Nucleic Acids Res. 35, 406-413 (2007). 15. Sen, D. & Gilbert, W. Formation of parallel four-stranded complexes by guanine-rich motifs in DNA and its implications for meiosis. Nature 334, 364-366 (1988). 16. Fouquerel, E., Parikh, D. & Opresko, P. DNA damage processing at telomeres: The ends justify the means. DNA Repair (Amst) 44, 159-168 (2016). 17. Siddiqui-Jain, A., Grand, C. L., Bearss, D. J. & Hurley, L. H. Direct evidence for a G-quadruplex in a promoter region and its targeting with a small molecule to repress c-MYC transcription. Proc. Natl. Acad. Sci. U S A 99, 11593-11598 (2002). 18. Cogoi, S. & Xodo, L. E. G-quadruplex formation within the promoter of the KRAS proto-oncogene and its effect on transcription. Nucleic Acids Res. 34, 2536-2549 (2006). 19. Phan, A. T., Modi, Y. S. & Patel, D. J. Propeller-Type Parallel-Stranded G-Quadruplexes in the Human c-myc Promoter. J. Am. Chem. Soc. 126, 8710-8716 (2004). 20. Banyay, M., Sarkar, M. & Graslund, A. A library of IR bands of nucleic acids in solution. Biophys. Chem. 104, 477-488 (2003). 21. Zhang, Y., Chen, J., Ju, H. & Zhou, J. Thermal denaturation profile: A straightforward signature to characterize parallel G-quadruplexes. Biochimie. 157, 22-25 (2019). 22. Mergny, J.-L., Phan, A.-T. & Lacroix, L. Following G-quartet formation by UV-spectroscopy. FEBS Lett. 435, 74-78 (1998). 23. Weissbluth, M. Hypochromism. Q Rev. Biophys. 4, 1-34 (1971). 24. Chinsky, L. & Turpin, P. Y. Ultraviolet Raman Spectroscopy of Polyribouridylic Acid: Excitation Profile of the Hypochromism Induced by Order-Disorder Transition. Biopolymers 19, 1507-1515 (1980). 25. Tomlinson, B. L. & Peticolas, W. L. Conformational Dependence of Raman Scattering Intensities in Polyadenylic Acid. J. Chem. Phys. 52, 2154-2156 (1970). 26. Turpin, P. Y., Chinsky, L., Laigle, A. & Jolles, B. DNA Structure Studies by Resonance Raman Spectroscopy. J. Mol. Struct. 214, 43-70 (1989). 27. Reipa, V., Niaura, G. & Atha, D. H. Conformational analysis of the telomerase RNA pseudoknot hairpin by Raman spectroscopy. RNA 13, 108-115 (2007). 28. Peticolas, W. L. Raman Spectroscopy of DNA and Proteins. Methods Enzymol. 246, 389-416 (1995). 29. Friedman, S. J. & Terentis, A. C. Analysis of G-quadruplex conformations using Raman and polarized Raman spectroscopy. J. Raman Spectrosc. 47, 259-268 (2016). 30. Miura, T. & Thomas, G. J. Structural Polymorphism of Telomere DNA: Interquadruplex and Duplex-Quadruplex Conversions Probed by Raman Spectroscopy. Biochemistry 33, 7848-7856 (1994). 31. Lane, AN, Chaires, JB, Gray, RD & Trent, JO Stability and kinetics of G-quadruplex structures. Nucleic Acids Res. 36, 5482-5515 (2008). 32. Guzman, MR, Liquier, J., Brahmachari, SK & Taillandier, E. Characterization of parallel and antiparallel G-tetraplex structures by vibrational spectroscopy. Spectrochim. Acta A Mol. Biomol. Spectrosc. 64, 495-503 (2006). 33. Price, DA et al. Infrared Spectroscopy Reveals the Preferred Motif Size and Local Disorder in Parallel Stranded DNA G-Quadruplexes. ChemBioChem 21, 1-14 (2020). 34. Dickerhoff, J., Dai, J. & Yang, D. Structural recognition of the MYC promoter G-quadruplex by a quinoline derivative: insights into molecular targeting of parallel G-quadruplexes. Nucleic Acids Res. 49, 5905-5915 (2021).
[0182] The following numbered clauses are also disclosed herein: A1. A method for determining the geometry of an interaction between a protein or nucleic acid and one or more chemical entities, comprising the following steps: (a) Prepare experimentally obtained multidimensional spectra of a sample comprising the protein or nucleic acid and the one or more chemical entities; (b) Prepare two or more structural hypotheses; (c) For each of the two or more structural hypotheses, obtain a computationally generated multidimensional spectrum of the same type as the experimentally obtained multidimensional spectrum; (d) Determining which of the computationally generated multidimensional spectra best fits the experimentally obtained multidimensional spectrum; and (e) A method comprising identifying the geometry of the interaction between the protein or nucleic acid and the one or more chemical entities based on step (d); A2. The method according to clause A1, wherein the experimentally obtained multidimensional spectrum is obtained using two-dimensional infrared spectroscopy (2DIR); A3. The method described in clause A2, wherein the experimentally obtained multidimensional spectrum is obtained using electron vibration two-dimensional infrared spectroscopy (EVV 2DIR); A4. The method according to any one of the provisions A1 to A3, wherein the one or more chemical entities are drug molecules; A5. The method according to clause A4, wherein the geometry of the interaction is located at the target binding site; A6. The method described in any one of clauses A1 to A5, wherein the computationally generated multidimensional spectrum is calculated using density functional theory "DFT"; A7. The method according to any one of the clauses A1 to A6, wherein step (d) includes ranking the two or more structural hypotheses; A8. The method according to any one of the clauses A1 to A7, wherein step (d) includes determining the correlation; A9. The method according to any one of the provisions A1 to A8, wherein the experimentally obtained multidimensional spectrum is pretreated; A10. The method according to any one of the clauses A1 to A9, wherein the experimentally obtained multidimensional spectrum is cleaned and artifacts are removed; A11. The method according to any one of the provisions A1 to A10, wherein the computationally generated multidimensional spectra are each preprocessed; A12. The method according to any one of the clauses A1 to A11, wherein the computationally generated multidimensional spectra are each cleaned and artifacts removed; A13. The method according to any one of the clauses A1 to A12, wherein step (a) includes measuring the multidimensional spectrum of the sample; A14. The method according to clause A13, wherein the sample is deposited as a substrate in the form of a microarray containing at least one spot-like deposit; A15. The aforementioned substrate It is transparent to visible light, or The method described in clause A14, which is reflective to visible light; A16. The method according to clause A14 or A15, wherein the microarray is maintained within a predetermined humidity range before and during the measurement of the sample; A17. The method according to any one of the provisions A13 to A16, wherein the sample scatters a considerable amount of incident visible light; A18. The method according to claim A17, wherein the scattered light is collected by an optical component(s) forming a lens or mirror before detection; A19. A computer-readable medium containing computer-executable instructions, which, when executed by a processor, cause the processor to perform the method described in any one of the clauses A1 to A18.
[0183] B1. A system for determining the geometry of interactions between a protein or nucleic acid and one or more chemical entities, (a) A spectrometer configured to experimentally obtain a multidimensional spectrum of a sample comprising the protein or nucleic acid and the one or more chemical entities; and (b) One or more processors, To generate or receive two or more structural hypotheses; For each of the two or more structural hypotheses mentioned above, a computationally generated multidimensional spectrum is obtained; From the computationally generated multidimensional spectra, determine the one that best fits the experimentally obtained multidimensional spectrum; and A system comprising a processor configured to identify the geometry of the interaction between the protein or nucleic acid and one or more chemical entities, based on the determination of the computationally generated multidimensional spectrum that best fits the experimentally obtained multidimensional spectrum; B2. The system described in clause B1, wherein the spectrometer is a two-dimensional infrared (2DIR) spectrometer; B3. The system described in Clause B2, wherein the spectrometer is a two-dimensional infrared (EVV 2DIR) spectrometer; B4. The system described in any one of the clauses B1 to B3, wherein one or more of the chemical entities are drug molecules; B5. The system according to clause B4, wherein the geometry of the interaction is located at the target binding site; B6. The system described in any one of clauses B1 to B5, wherein the computationally generated multidimensional spectrum is calculated using density functional theory "DFT"; B7. The system according to any one of clauses B1 to B6, wherein one or more processors are configured to rank the two or more structural hypotheses as part of determining which of the computationally generated multidimensional spectra best fits the experimentally obtained multidimensional spectrum; B8. The system according to any one of the clauses B1 to B7, wherein one or more processors are configured to receive user input in order to determine which of the computationally generated multidimensional spectra best fits the experimentally obtained multidimensional spectrum; B9. The system according to any one of the clauses B1 to B8, wherein one or more processors are configured to preprocess the experimentally obtained multidimensional spectrum; B10. The system according to any one of clauses B1 to B9, wherein the one or more processors are configured to clean the experimentally obtained multidimensional spectrum to remove artifacts; B11. The system according to any one of clauses B1 to B10, wherein the one or more processors are configured to perform preprocessing on each of the computationally generated multidimensional spectra; B12. The system according to any one of clauses B1 to B11, wherein the one or more processors are configured to clean each of the computationally generated multidimensional spectra to remove artifacts; B13. The system according to any one of clauses B1 to B12, further comprising a sample including a substrate in the form of a microarray comprising at least one spot-shaped deposit; B14. The substrate is transparent to visible light, or reflective to visible light, the system according to clause B13; B15. The system according to clause B13 or B14, further comprising a humidity control device configured to maintain the humidity of the microarray within a predetermined humidity range before and during measurement of the sample by the spectrometer; B16. The system according to any one of clauses B13 to B15, wherein the sample is configured to scatter a substantial amount of incident visible light; B17. The system according to clause B16, further comprising at least one optical component forming a lens or a mirror, wherein the at least one optical component is configured to collect scattered light.
Claims
1. A method for determining the geometry of an interaction between a protein or nucleic acid and one or more chemical entities, comprising the following steps: (a) Prepare experimentally obtained multidimensional spectra of a sample comprising the protein or nucleic acid and one or more chemical entities; (b) Prepare two or more structural hypotheses; (c) For each of the two or more structural hypotheses, obtain a computationally generated multidimensional spectrum of the same type as the experimentally obtained multidimensional spectrum; (d) Determining which of the computationally generated multidimensional spectra best fits the experimentally obtained multidimensional spectrum; and (e) A method comprising identifying the geometry of the interaction between the protein or nucleic acid and the one or more chemical entities based on step (d).
2. The method according to claim 1, wherein the experimentally obtained multidimensional spectrum is obtained using two-dimensional infrared spectroscopy (2DIR).
3. The method according to claim 2, wherein the experimentally obtained multidimensional spectrum is obtained using electron vibration two-dimensional infrared spectroscopy (EVV 2DIR).
4. The method according to any one of claims 1 to 3, wherein the one or more chemical entities are drug molecules.
5. The method according to claim 4, wherein the geometry of the interaction is located at a target binding site.
6. The method according to any one of claims 1 to 5, wherein the computationally generated multidimensional spectrum is calculated using density functional theory "DFT".
7. The method according to any one of claims 1 to 6, wherein step (d) includes ranking the two or more structural hypotheses.
8. The method according to any one of claims 1 to 7, wherein step (d) includes determining the correlation.
9. The method according to any one of claims 1 to 8, wherein the experimentally obtained multidimensional spectrum is subjected to pretreatment.
10. The method according to any one of claims 1 to 9, wherein the experimentally obtained multidimensional spectrum is cleaned to remove artifacts.
11. The method according to any one of claims 1 to 10, wherein the computationally generated multidimensional spectra are each subjected to preprocessing.
12. The method according to any one of claims 1 to 11, wherein the computationally generated multidimensional spectra are each cleaned to remove artifacts.
13. The method according to any one of claims 1 to 12, wherein step (a) includes measuring the multidimensional spectrum of the sample.
14. The method according to claim 13, wherein the sample is deposited as a substrate in the form of a microarray containing at least one spot-like deposit.
15. The aforementioned substrate, It is transparent to visible light, or The method according to claim 14, wherein it is reflective to visible light.
16. The method according to claim 14 or 15, wherein the microarray is maintained within a predetermined humidity range before and during the measurement of the sample.
17. The method according to any one of claims 13 to 16, wherein the sample scatters a considerable amount of incident visible light.
18. The method according to claim 17, wherein the scattered light is collected by one or more optical components forming a lens or mirror before detection.
19. A system for determining the geometry of interactions between a protein or nucleic acid and one or more chemical entities, (a) A spectrometer configured to experimentally obtain a multidimensional spectrum of a sample comprising the protein or nucleic acid and one or more chemical entities; and (b) One or more processors, Generate or receive two or more structural hypotheses, each containing a computationally generated multidimensional spectrum; Of the computationally generated multidimensional spectra, determine the one that best fits the experimentally obtained multidimensional spectrum; and A system comprising a processor configured to identify the geometry of the interaction between the protein or nucleic acid and one or more chemical entities, based on the determination of the computationally generated multidimensional spectrum that best fits the experimentally obtained multidimensional spectrum.
20. A computer-readable medium comprising a computer-executable instruction, when executed by a processor, causing the processor to perform the method according to any one of claims 1 to 19.