Methods and tools for determining conformation and conformational changes of proteins and their derivatives - Patents.com

JP2024545439A5Pending Publication Date: 2025-12-03ビオグノシス アーゲー +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024533240
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-03
Filing Date
2022-11-25
Publication Date
2025-12-03

AI Technical Summary

Technical Problem

Current methods for studying protein conformational changes in their biological context are limited by the inability to handle complex biological samples and provide multiplexed, scalable, and easy-to-use analysis, especially for clinical applications.

Method used

A method combining limited proteolysis with mass spectrometry techniques like DIA (SWATH) and non-specific proteases under native conditions, followed by filtration to enrich for non-tryptic peptides, which are then analyzed by LC-MS/MS to identify conformational states of proteins.

Benefits of technology

This approach enhances the sensitivity, coverage, and throughput of protein conformational analysis, allowing for the detection of structural changes and potential drug targets, and providing insights into disease mechanisms and biomarkers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for detecting the conformational state of a protein, said protein being in a complex mixture of further proteins and other biomolecules, wherein the protein in said complex mixture is exposed to conditions inducing a structural change in said protein, comprising the following sequence of steps: 1. obtaining a first fragment sample by performing limited proteolysis of an extraction mixture under conditions in which the protein is in an initial conformational state to be detected, immediately followed by: 2. removing large peptides and proteins or other biomolecules from said first fragment sample to form an enriched fragment sample; 3. performing an analytical analysis of the enriched fragment sample to determine the fragments that are characteristic when the limited proteolysis of step 1 is brought about, but also remain after the removal step 2, thereby determining the conformational state of said at least one protein.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to methods for determining the conformation and conformational changes of proteins and their derivatives, optionally in their native biological context, in particular using limited proteolysis in combination with e.g. selected reaction monitoring, data-dependent acquisition (DDA), data-independent acquisition (DIA) including the Sequential Windowed Acquisition of All Theoretical Fragment Ion Mass Spectra (SWATH) method. [Background technology]

[0002] Proteins are crucial effectors and regulators of a wide variety of cellular processes. In response to perturbations (e.g., in disease), proteins can change their cellular concentration, activity, location, and their structure. Being able to capture such transitions is a vital challenge in life sciences to understand the functioning of fundamental cellular processes in health and disease, and to identify new options for disease diagnosis and treatment. Changes in cellular protein concentration in response to perturbations can be routinely investigated by mass spectrometry (MS)-based proteomic techniques. Little is known about cellular protein conformational switches, mainly due to the lack of suitable methods to study protein folding in cells. This switch represents a major limitation for biological and clinical applications, since conformational changes can strongly affect protein activity, location, and stability, thus profoundly affecting cellular physiology.

[0003] Proteins can change their conformation due to binding to lipids, ions, small molecules, or nucleic acids, interactions with other proteins, chemical modifications (e.g., phosphorylation), or environmental changes such as fluctuations in pH, ionic strength, or temperature. The extent of conformational changes ranges from small localized movements such as allosteric / local rearrangements to larger scale fluctuations such as domain movements, and even dramatic switches between folded and unfolded states or between monomeric and multimeric states. In particular, the transition of monomeric proteins to higher aggregated structures has recently attracted increasing attention in both biology and biomedicine. Over the past two decades, various human diseases, called protein aggregation diseases (more than 20 different pathologies), have been found to be associated with the intracellular or extracellular accumulation of aggregates of specific misfolded proteins. Many neurodegenerative diseases with previously unknown etiology, such as Parkinson's disease or Alzheimer's disease, are now classified in this category. It is even possible to classify various diseases according to the predominant protein component of the protein aggregates, which also differentiates the clinical manifestations of the disease. For example, α-synuclein (αSyn)-containing Lewy bodies are typical for most types of Parkinson's disease (PD), while amyloid- β Peptide inclusions are generated in Alzheimer's disease (see, for example, Non-Patent Document 1). The ability to monitor such protein conformational transitions in biological specimens will open new possibilities for the diagnosis and therapy of these protein-centric pathologies and shed light on their pathogenesis.

[0004] Many biophysical techniques have been applied to monitor the conformational features of proteins, such as nuclear magnetic resonance (NMR), X-ray crystallography, infrared and Raman spectroscopy, circular dichroism, atomic force microscopy, or fluorescence spectroscopy. These techniques are mainly used to analyze (purified) proteins in vitro, since they cannot deal with the complex biological background. This is a substantial limitation, since the conformation adopted by proteins is controlled in cells by a large number of simultaneous events specific to the cellular context, such as environmental factors, binding events, or post-translational modifications, which cannot be reproduced by in vitro systems. Techniques based on Förster resonance energy transfer (FRET) have the advantage that conformational changes of proteins are monitored in their native cellular environment, but they require the introduction of fluorescent probes at appropriate sites of each target protein, making them inapplicable on a large scale or for clinical samples.

[0005] In light of the above considerations, there is an urgent need to have available methods to follow protein conformational changes in their biological environment in a multiplexed manner (many proteins at once). Additional features of an ideal method would be i) amenable to scale-up (rapid analysis of many samples), and ii) easily adaptable to a variety of applications (clinical or biotechnological applications or basic biological research).

[0006] Gupta, Lapadula, and Abou-Donia reported the purification and characterization of pure cytochrome P450 isozymes from β-Naphthoflavone-Induced Adult Hen Liver (Non-Patent Document 2) in their paper entitled "Purification and Characterization of Cytochrome P450 Isozymes from β-Naphthoflavone-Induced Adult Hen Liver." Characterization is performed by protease treatment using chymotrypsin under denaturing conditions.

[0007] Cohen, Ferre-D'Amare, Burley, and Chait, in their paper entitled "Probing the solution structure of the DNA-binding protein Max by a combination of proteolysis and mass spectrometry" (Non-Patent Document 3), propose a simple biochemical method to probe the solution structure of DNA-binding proteins by combining enzymatic proteolysis with matrix-assisted laser desorption / ionization mass spectrometry. The method is based on inferring structural information from the determination of protection against enzymatic proteolysis, which depends on the solvent accessibility and flexibility of the protein.

[0008] US Patent No. 5,399,633 discloses a limited proteolysis (LiP) protocol, i.e. a method for detecting the conformational state of a protein contained in a complex mixture of further proteins and / or other biomolecules, in particular in a complex natural biological matrix, as well as an assay for such a method. The method comprises, optionally after an extraction and / or lysis step, the following steps: 1. limited proteolysis of the complex mixture under conditions in which the protein is in the conformational state to be detected to obtain a first fragment sample, 2. denaturing the first fragment sample to obtain a denatured first fragment sample, 3. completely fragmenting the denatured first fragment sample in a digestion step to obtain a completely fragmented sample, 4. analytical analysis of the completely fragmented sample to determine the characteristic fragments when both the limited proteolysis of step 1 as well as the complete fragmentation of digestion step 3 have been obtained, thereby determining the conformational state.

[0009] In a paper entitled "Measuring protein structural changes on a proteome-wide scale using limited proteolysis-coupled mass spectrometry" (Non-Patent Document 4), Schopper et al. report on protein structural changes caused by external perturbations or internal factors that regulate cell physiology to profoundly affect protein activity. Limited proteolysis-coupled mass spectrometry (LiP-MS) is reported to be a recently developed proteomic approach that allows protein structural changes to be directly identified in their complex biological context on a proteome-wide scale. After the perturbation of interest, the proteome extract is subjected to a double protease digestion step in which a non-specific protease is applied under native conditions, followed by complete digestion with the sequence-specific protease trypsin under denaturing conditions. This sequential process generates structure-specific peptides suitable for bottom-up MS analysis. Proteomic workflows including shotgun or targeted MS and label-free quantification are then applied to directly measure structure-dependent protein degradation patterns in proteomic extracts. Potential applications of LiP-MS include discovery of perturbation-induced protein structural changes, identification of drug targets, detection of disease-related protein structural states, and direct analysis of protein aggregates in biological samples. This approach also allows identification of specific protein regions involved in structural transitions or affected by binding events.

[0010] In their paper entitled "The cellular thermal shift assay for evaluating drug target interactions in cells" (Non-Patent Document 5), Jafari et al. report on a thermal shift assay used to study the thermal stabilization of proteins upon ligand binding. Such assays have been used in the drug discovery industry and academia to detect interactions on purified proteins. Proof-of-principle studies have been published describing the implementation of thermal shift assays in cellular formats, which have been termed cellular thermal shift assays (CETSA). This method allows the study of target engagement of drug candidates in a cellular context, as exemplified by experimental data on the human kinases p38a and ERK1 / 2. The assay involves treating cells with the compound of interest, heating to denature and precipitate proteins, lysing the cells, and separating cell debris and aggregates from the soluble protein fraction. Unbound proteins denature and precipitate at high temperature, while ligand-bound proteins remain in solution. They describe two procedures to detect stabilized proteins in the soluble fraction of the sample. The first approach involves sample post-processing and detection using quantitative Western blotting, whereas the second approach is performed directly in solution and relies on the induction of proximity of two targeting antibodies upon binding to soluble proteins. The latter protocol has been optimized to allow for increased throughput, as potential applications require large numbers of samples.

[0011] In a paper entitled "Dynamic 3D proteomes reveal protein functional alterations at high resolution in situ" (Non-Patent Document 6), Cappelletti et al. report that it is possible to read out the overall protein structure based on limited proteolysis-mass spectrometry (LiP-MS) that detects many functional changes simultaneously in situ in bacteria undergoing nutritional adaptation and yeast responding to acute stress. The readout of the structure visualized as a structural barcode captured changes in enzyme activity, phosphorylation, protein aggregation, and complex formation, along with the elucidation of individual regulated functional sites such as binding sites and active sites. By comparing with existing knowledge including other omics data, it was revealed that LiP-MS detects many known functional changes in well-studied pathways. This suggested distinct metabolite-protein interactions and allowed the identification of the control mechanism of glucose uptake based on fructose-1,6-bisphosphate in E. coli. Structural readout dramatically expands the coverage of classical proteomics, informs mechanistic hypotheses, and paves the way for in situ structural systems biology.

[0012] In a paper titled "Tracking cancer drugs in living cells by thermal profiling of the proteome" (Non-Patent Document 7), Savitzki et al. reported that thermal proteome profiling (TPP) was performed on human K562 cells by heating intact cells or cell extracts, and that there was a marked difference in melting characteristics between the two settings, as well as a trend toward increased protein stability in cell extracts. It has been reported that thermal profiling of cellular proteomes allows differential evaluation of protein ligand binding and other protein modifications, provides an unbiased measurement of drug target occupancy for multiple targets, and facilitates the identification of markers for drug efficacy and toxicity.

[0013] US Patent No. 5,399,633 discloses a method for detecting the conformational state of a protein contained in a complex mixture of further proteins and / or other biomolecules, in particular in a complex natural biological matrix, as well as an assay for such a method. The method comprises, optionally after an extraction and / or dissolution step, the following steps: 1. performing limited proteolysis of the complex mixture under conditions in which the protein is in the conformational state to be detected to obtain a first fragment sample, 2. denaturing the first fragment sample to obtain a denatured first fragment sample, 3. completely fragmenting the denatured first fragment sample in a digestion step to obtain a completely fragmented sample, and 4. performing an analytical analysis of the completely fragmented sample to determine the characteristic fragments when both the limited proteolysis of step 1 as well as the complete fragmentation of digestion step 3 are obtained, thereby determining the conformational state.

[0014] Schopper et al., "Measuring protein structural changes on a proteomewide scale using limited proteolysis-coupled mass spectrometry" (Non-Patent Document 4), reports on protein structural changes caused by external perturbations or internal factors, and how these can profoundly affect protein activity and therefore regulate cell physiology. Limited proteolysis-coupled mass spectrometry (LiP-MS) is reported as a technique that allows protein structural changes to be directly identified in their complex biological context on a proteome-wide scale. After the perturbation of interest, the proteome extract is subjected to a double protease digestion step in which a non-specific protease is applied under native conditions, followed by complete digestion with the sequence-specific protease trypsin under denaturing conditions. This sequential process generates structure-specific peptides suitable for bottom-up MS analysis. A proteomic workflow including shotgun or targeted MS and label-free quantification is then applied to directly measure structure-dependent protein degradation patterns in proteomic extracts. Potential applications of LiP-MS have been reported to include discovery of perturbation-induced protein structural changes, identification of drug targets, detection of disease-related protein structural states, and direct analysis of protein aggregates in biological samples. This approach also allows identification of specific protein regions involved in structural transitions or affected by binding events. Sample preparation takes approximately 2 days, followed by MS and data analysis times ranging from 1 to several days depending on the number of samples analyzed.

[0015] (2009) describes a chemoselective enrichment strategy called semitryptic peptide enrichment strategy for proteolysis procedures (STEPP) to isolate semitryptic peptides generated in mass spectrometry-based proteome-wide applications of limited proteolysis. The strategy involves reacting isobaric mass tags with the ε-amino groups of lysine side chains and any N-termini generated in the limited proteolysis reaction. Subsequent digestion of the sample with trypsin and chemoselective reaction of the newly exposed N-termini of the tryptic peptides with N-hydroxysuccinimide (NHS)-activated agarose resin removes the tryptic peptides from the solution, leaving only the semitryptic peptides with one nontryptic cleavage site generated in the limited proteolysis reaction for subsequent LC-MS / MS analysis. As part of this work, the STEPP technique is coupled with two different proteolysis methods, including pulsed proteolysis (PP) and limited proteolysis (LiP). The STEPP-PP workflow is evaluated in two proof-of-principle experiments targeting proteins in yeast cell lysates as well as two well-studied drugs, namely cyclosporine A and geldanamycin. The STEPP-LiP workflow is evaluated in two proof-of-principle experiments targeting proteins in two cell culture models of human breast cancer, the MCF-7 and MCF-10A cell lines. The STEPP protocol increased the number of semitryptic peptides detected in the LiP and PP experiments by 5- to 10-fold. The STEPP protocol not only expands the coverage of the proteome but also increases the amount of structural information that can be gleaned from limited protein digestion experiments. Moreover, the protocol also enables quantitative determination of ligand binding affinities.

[0016] Non-Patent Document 9 describes an integrated experimental and computational technique to quantify hundreds of protein complexes in a single run. The method consists of size-exclusion chromatography (SEC) to fractionate native protein complexes, SWATH / DIA mass spectrometry to accurately quantify proteins in each SEC fraction, and the computational framework CCprofiler to detect and quantify protein complexes by error-controlled complex-centric analysis using prior information from a comprehensive protein interaction map. Analysis of the HEK293 cell line proteome revealed 462 complexes composed of 2127 protein subunits. The technique identifies novel subcomplexes and assembly intermediates of central regulatory complexes while assessing quantitative subunit distributions across them. The toolset CCprofiler is freely available and provides a web platform SECexplorer for custom exploration of the modularity of the HEK293 proteome. [Prior art documents] [Patent documents]

[0017] [Patent Document 1] International Publication No. 2014 / 082733 [Non-patent literature]

[0018] [Non-Patent Document 1] A Aguzzi & T O'Connor, Nat Rev Drug Discov 9 (3), 237 [Non-Patent Document 2] Archives of Biochemistry and Biophysics, 282 (1) 170-182 (1990) [Non-Patent Document 3] Protein Science (1995), 4:1088-1099 [Non-Patent Document 4] nature protocols, VOL.12 NO.11, 2017, 2391ff

Non-Patent Document 5

Non-Patent Document 6

Non-Patent Document 7

Non-Patent Document 8

Non-Patent Document 9

Summary of the Invention

[0019] Current LiP-MS approaches, for example those disclosed in US Pat. No. 5,399,633 (e.g., Schopper et al. or Cappelletti et al.), specifically rely on a double digestion workflow that involves extensive non-specific enzymatic digestion of native proteins followed by a complete tryptic digestion of the denatured proteome. Subsequent analysis is then focused on tryptic and / or semi-tryptic peptides. The approach proposed here, on the other hand, is designed to focus on peptides that are generated only from native, non-denatured proteins and are largely non-tryptic and therefore largely absent in a typical LiP-MS experiment. It is important to note that attempts to utilize completely non-tryptic peptides in the analysis of standard LiP-MS experiments would provide little useful information for two reasons. First, as a result of the complete tryptic digestion performed under denaturing conditions that are standard in LiP-MS protocols, very few completely non-tryptic peptides remain. Second, any such peptides generated are difficult to accurately detect and quantitate via mass spectrometry because their signals are weak and / or masked by the sheer amount and volume of tryptic and / or semi-tryptic peptides present. Analysis of a typical LiP-MS experiment reveals that less than 1% of the total peptides are completely non-tryptic and do not contribute to target identification in positive control experiments.

[0020] The proposed workflow uses LC-MS, DIA (data-independent acquisition of product ion spectra) mass spectrometry, and non-specific database searching to enable MS-based identification of unique structural / conformational states of proteins and / or structural / conformational protein changes in complex biological or clinical specimens with high sensitivity, coverage, and throughput.

[0021] The structural / conformational states of proteins sampled by the proposed approach can be native, non-native, or a mixture of both, depending on the application. Implementations of this technology can be used to investigate protein structures under standard conditions, during protein binding to drugs / small molecules or metabolites, when bound to other proteins (protein-protein interactions, i.e., protein complexes), when bound to various other molecules (e.g., lipids or DNA) as a result of chemical modifications (e.g., PTMs such as protein phosphorylation), or upon changes in the local environment (e.g., increased temperature, changes in ionic strength, or the presence of chaotropes). The proposed technology allows for hypothesis-free detection and identification of proteins that undergo structural changes when perturbations are induced in the investigated system (e.g., immune signaling or disease initiation). Identifying such unique protein conformational changes and states provides valuable information about protein structure and function, allows identification of drug or metabolite targets of interest, characterizes biochemical and signaling pathways involved in the response to perturbations, and provides information about disease mechanisms. Furthermore, altered protein structures can be used as surrogate indicators (i.e., structural biomarkers) for disease detection. Insights into the structural state and dynamics of the structural proteome provide a deeper understanding of both physiological and non-physiological mechanisms of action that can be a major obstacle in advancing our understanding of disease and supporting drug design and improvement. Thus, the proposed technology allows the exploitation of structural proteomics for a variety of applications ranging from basic biology to target deconvolution and biomarker discovery.

[0022] More specifically, the proposed technique uses a novel approach that addresses the problems / disadvantages inherent in existing mass spectrometry techniques (e.g., the above-mentioned approaches by Schopper et al., Cappelletti et al., Jafari et al., and Savitzki et al.) that aim to address the aforementioned challenges. These techniques focus on maximizing the depth (coverage, sensitivity) of proteome discovery by relying on the generation and identification of tryptic and / or semi-tryptic digestion mixtures (i.e., one tryptic digestion cleavage per peptide). These peptides are typically suitable for mass spectrometry in terms of their size and identifiability. However, in these experiments, many peptides are uninformative in terms of structural / conformational information, while significantly increasing the complexity, dynamic range, and noise of the sample. Furthermore, in current versions of LiP-MS, after tryptic digestion, many peptides are lost in the analysis because they are too short for reliable MS-based identification. The proposed technique is a new variant of the LiP-MS approach, which sacrifices the identification of peptides (and therefore proteins) that do not necessarily report structural information, and instead focuses on increasing the relative ability to identify peptides that convey structural / conformational information from proteins. By focusing on enrichment of structurally informative peptides, the proposed technique increases the signal-to-noise ratio, making the identification of protein structural changes more robust.

[0023] A key feature of the proposed technique is that it increases the number and abundance of truly informative peptides relative to the total number of peptides contained in a sample, thus significantly reducing the dynamic range challenge inherent in proteomic samples.

[0024] This problem is most pronounced in human body fluids such as plasma, but is also seen in samples with low abundance of the protein(s) of interest, as is common in the study of drugs with a single (or few) protein targets. This wide dynamic range is due to, among other factors, a combination of the size / concentration distribution of proteins and the number and distribution of peptide responses in the mass spectrometer. In classical proteomic approaches, including LiP-MS approaches, the larger the protein, the more tryptic peptides will be generated. Thus, there is a strong correlation between the size and / or abundance of a protein and the likelihood that such a protein will generate at least some peptides that show a strong response in the mass spectrometer. In contrast, in the proposed technology approach, peptides are generated primarily from protease-accessible regions of proteins in their native or near-native conformation (i.e., not denatured). These peptides tend to be present on the solvent-exposed surface of the protein. This greatly reduces the proportional number of peptides derived from large proteins, since the relationship between protein surface area and volume is not a fixed ratio but, on average, decreases as the size of the protein increases. By exploiting this bias in surface area to size ratio, the proposed technique aims to reduce the dynamic range inherent in proteomics samples, which is at least in part due to the natural size distribution of proteins. Furthermore, the proposed technique focuses on information-rich accessible peptides that report on protein structure and protein structural changes.

[0025] The proposed technique exploits two main processes: enrichment of short MS-compatible peptides coupled with limited digestion. The overall goal of this procedure is to increase the number of peptides that carry conformational / structural information (signal) while reducing the amount of peptides that do not contain such information (noise) in the mass spectrometer-ready sample.

[0026] This is achieved in a surprisingly simple and efficient manner by using a limited digestion step under non-denaturing (i.e., protein structure-preserving) conditions, using specific or non-specific proteases. The limited digestion can be varied by changing the enzyme to substrate ratio, performing the digestion at lower temperatures, or performing the digestion on a relatively short time scale. The resulting peptide mixture is then not fully denatured, but fully fragmented, in contrast to the LiP-MS protocol, and this is filtered or processed in other ways (see below) to remove large peptide / protein fragments that represent the majority of the mixture and contain little information about the protein structure / conformation and / or structural / conformational changes. By removing large peptide / protein pieces from the digest prior to mass spectrometry, the proposed technique also reduces the number of non-informative peptides that may introduce artifacts and / or noise into the experiment, even during downstream data analysis in terms of peptide identification on the mass spectrometer. This removal can be achieved by processes that separate large peptide / protein fragments to enrich for suitable peptides, including size filtration (e.g., using a filtration device with a MWCO of 10k), chromatography (e.g., size exclusion, hydrophobic, or anion exchange), or physical processes (e.g., phase separation, absorption, or precipitation). Compared to classical LiP-MS, in the proposed technique the full trypsinization step under denaturing conditions after the limited proteolysis step is omitted.

[0027] In addition to using a single protease for the limited digestion step, the proposed technique can be enhanced by the use of a protease mixture and / or a sensitizing agent (e.g., heat, urea). Instead of a protease mixture, two different proteases (or a set of proteases) can be used on an aliquot of the sample, followed by pooling of the samples. Pooling can be performed after peptides have been digested or separated from the rest of the sample. Instead of pooling, samples can also be processed and measured completely separately. According to the proposed technique, the protease mixture and the sensitizing agent act in two ways to improve the identification of the protein of interest. The protease mixture contains proteases that target different amino acids for cleavage, making it possible to generate more and more unique peptides during the digestion step very easily. The sensitizing agent works by slightly disrupting the native state of the protein, allowing new protease cleavage (cleavage sites). Sensitizing agents are particularly useful for the proposed technology when a particular condition is being investigated for changes in protein structure (e.g., both in the presence and absence of a drug or metabolite). In such cases, the effect of the sensitizing agent is enhanced when an event such as drug binding alters the structure (i.e., stabilizes or destabilizes) rendering the protein of interest more or less sensitive to the particular sensitizing agent.

[0028] Thus, the proposed technique is able to identify the conformation or structure of a particular protein by exploiting previously ignored peptides (i.e., peptides with two non-tryptic digestible ends located on the surface of the protein). The step of removing large peptides and proteins included in the proposed technique workflow (e.g., filtration) led to a relative enrichment of structurally informative peptides, although it altered the overall peptide / protein identification counts.

[0029] The proposed technical workflow includes or consists of the following steps: 1. Protein or proteome-containing samples from cells, tissues, or bodily fluids are interrogated for protein structure and / or conformational state. This includes, but is not limited to, incubation with ligands (e.g., small molecules, metabolites, etc.) at defined concentrations (treatment mixtures) including control (vehicle) samples, or exposure of lysates to conditions that induce changes in structural state such as temperature, metabolic activators, etc. including unstimulated samples. Whatever conditions are utilized, native protein structure (primary, secondary, tertiary, and preferably / potentially also quaternary) is preserved throughout. 2. Each sample is subjected to a limited digestion step using a specific or non-specific protease (or a combination thereof). This digestion should be relatively short, typically 1-5 minutes, and immediately quenched. 3. Following this digestion step, larger peptides and protein fragments are removed, for example via a filtration device, separation device, or concentration device (see methods above and below). 4. The remaining peptides can then be processed using standard proteomics workflows (e.g., denaturation, C18 cleanup, etc.) and analyzed via LC-MS / MS, thus identifying and quantifying the peptides.

[0030] An illustrative schematic of how this technology supports the detection of structural / conformational changes is shown in Figure 1, using the investigation of ligand binding events as an example.

[0031] Thus, the proposed approach is a novel complementary approach to LiP-MS that can yield completely unique information from conventional LiP-MS experiments due to its focus on generating and analyzing a set of unique peptides. By digesting only native proteins for a limited time and introducing subsequent filtration / concentration steps, the proposed technique reduces the number of informative peptides and / or protein pieces that are problematic during sample preparation, data acquisition, and analysis. This increased signal-to-noise ratio can equate to a significant improvement in data quality and thus improved biological insights.

[0032] More generally, the invention relates to a method for detecting the conformational state of at least one protein, said at least one protein being contained in a complex mixture of further proteins and / or other biomolecules (such a complex mixture can be, for example, a complex cell extract mixture or, in general, a biological probe such as a tissue-based body fluid (plasma, cerebrospinal fluid, urine, etc.), an environmental sample (e.g. seawater, etc.), e.g. a biological secretion obtained by cells lysed according to the published LiP protocol (see above), and the protein concentration can be determined using an assay kit), said at least one protein in said complex (e.g. cell extract) mixture being exposed to conditions inducing a conformational change in said at least one protein. The method comprises, optionally after an extraction and / or lysis step, the following sequence of steps: 1. Obtaining a first fragment sample by limited proteolysis of a complex (e.g., cell extract) mixture under conditions in which at least one protein is in an original conformational state to be detected, followed immediately by 2. Removing large peptides and proteins or other biomolecules from the first fragment sample to form an enriched fragment sample; 3. Analytical analysis of the enriched fragment sample to determine the fragments that are characteristic of the limited proteolysis of step 1 and that remain after removal step 2 to determine the conformational state of said at least one protein; Includes.

[0033] When "immediately after" is mentioned at the end of step 1, this excludes further denaturation and / or proteolysis steps, but does not exclude further steps involved in stopping the limited proteolysis of step 1, i.e. quenching the limited proteolysis by adding a corresponding reagent (e.g. sodium deoxycholate solution), increasing the temperature, washing, filtration, precipitation, adjusting the solvent, ionic strength and / or pH, or a combination thereof, etc. In other words, step 1 is "completed" from the point of view of digestion, since there are no other steps during the limited proteolytic digestion that contribute to peptide generation (e.g. denaturation or addition of other proteases that increase access to the cleavage site, etc.). However, deoxycholate can be added and the temperature can be, for example, up to 98° C. before step 2, which stops the protease activity from step 1.

[0034] The term "conformational state" of at least one protein should be understood broadly as generally accepted in the field and means the spatial arrangement of the constituent atoms of the protein that determines the overall shape of the molecule. In other words, the expression "conformational state" includes any kind of information that goes beyond the primary structure information, i.e. the linear sequence of amino acids in a peptide or protein and its potential chemical modifications. The expression "conformational state" therefore includes secondary structure (three-dimensional arrangement of local segments of a protein, the two most common secondary structure elements are alpha helices and beta sheets, but also beta turns and omega loops), supersecondary structure (a compact three-dimensional protein structure of several adjacent elements of secondary structure smaller than a motif, i.e. a protein domain or protein subunit), tertiary structure (domain, i.e. the three-dimensional shape of a protein), and quaternary structure (the structure of a protein that is itself composed of two or more smaller protein chains, also called subunits).

[0035] The conformational changes of proteins that can be identified using the proposed method are made possible by their inherent flexibility. These changes can occur with relatively little energy expenditure. At the molecular structure level, conformational changes in single polypeptides are the result of changes in main-chain torsion angles and side-chain orientations. The overall effect of such changes can be localized by reorientations of a few residues and small local main-chain torsional changes. On the other hand, torsional changes localized to very few residues in key positions can also lead to large changes in the tertiary structure. The latter type of conformational changes is described as domain motions. In their paper entitled “Structural changes involved in protein binding correlate with intrinsic motions of proteins in the unbound state” (PNAS December 27, 2005vol. 102 no. 52), Tobi et al. report that conformational changes associated with protein-protein interactions range from local changes in the rotamer states of side chains to global changes in structure, including collective domain motions.

[0036] The proposed approach is a structural proteomics approach, which is defined here as the use of proteomic approaches to determine the structural features / structures of proteins and structural changes in proteins, either individually or proteome-wide.

[0037] The proposed method also enables what is called target deconvolution, which is the identification of direct interaction partners of a drug / small molecule or protein. Target deconvolution can be achieved by a number of methods, including affinity chromatography, expression cloning, protein microarrays, "reverse transfection" cell microarrays, and biochemical inhibition.

[0038] Preferably, the analysis in step 3 is performed by quantitative mass spectrometry.

[0039] The proposed method is based on coupling a biochemical technique called limited protein proteolysis (LiP) with other mass spectrometry techniques such as DIA (including SWATH-MS) or selected reaction monitoring (SRM), or advanced targeted mass spectrometry workflows including SRM-like techniques including parallel reaction monitoring (PRM), data-dependent acquisition (DDA) and isobaric labeling quantification.

[0040] Liquid chromatography coupled to mass spectrometry (LC-MS) has been used for many years in the proteomic community for the identification and quantification of peptides (and therefore proteins) from complex sample mixtures. The most commonly used approaches are variants of the so-called LC-MS / MS or "shotgun" MS approaches, which are based on the generation of fragment ions from precursor ions that are automatically selected based on the precursor ion profile (data-dependent analysis, DDA). The most mature technique is called selected reaction monitoring (SRM), also often called multiple reaction monitoring (MRM). Targets for MRM experiments are defined on a rational basis and depend on the hypothesis to be tested in the experiment. Combinations of precursor and fragment ions selected for these targets (so-called transitions, a set of transitions for one target precursor is called an MRM assay) are programmed into the mass spectrometer, after which measurement data are generated only for the defined targets. A variant of targeted proteomics is data-independent acquisition, a more recently published variant commonly called SWATH-MS approach or data-independent acquisition (DIA). Here, the targeted aspect is introduced only at the data analysis level. In contrast to MRM, this approach does not require any preliminary peptide-specific method design prior to sample injection. The LC-MS acquisition covers the entire analyte content of the sample over the entire mass and retention time (RT) range, so that the data can be mined a posteriori for any peptide / precursor of interest. Data is acquired in a data-independent manner across the entire chromatography in the complete mass range (e.g., 200 Thomsons to 2000 Thomsons) regardless of the sample content. This is typically achieved by stepping a peptide precursor selection window across the complete mass range. In effect, this data acquisition method creates a complete fragment ion map for all analytes present in the sample and links the fragment ion spectra back to the precursor ion selection window in which the fragment ion spectra were acquired.This is achieved by extending the precursor isolation window and thus taking into account in advance a large number of precursors that co-elute and thus participate simultaneously in the fragmentation pattern recorded during the analysis. Such a precursor window is called swath. As a consequence, complex fragment ion spectra arise from the fragmentation of a large number of precursors, which requires more challenging data analysis. Unlike shotgun proteomics, in MRM and SWATH or DIA techniques, spectra for the same analyte are repeatedly recorded with high time resolution (LC retention time resolution). The limited fragment ion information in the case of MRM and the limited association of fragment ions with precursor ions in the case of SWATH / DIA, together with the (high) time resolution compared to shotgun proteomics, both require and enable entirely new types of data analysis. Since only a limited number of predefined analytes are monitored, there is no need to perform a shotgun proteomics type database search by comparing the spectrum to a complete theoretical proteome. Instead, a number of scores based on signal features such as shape, co-elution of transitions, and similarity of transition intensities with the assay library are described. Entirely new types of data analysis are targeted (peptide-centric) analysis and untargeted data (spectrum-centric) analysis.

[0041] Spectral-centric analysis can be defined as follows: data analysis of data, which may be DDA or DIA data acquired in an LC-MS / MS experiment, where the search is spectral-centric. This means scanning the spectrum in the MS2 dimension for possible matches with all theoretical peptides and their fragments, typically derived from a protein database with no or limited prior spectral information. Typically, the parent precursor ions for the MS2 spectrum are matched with a certain m / z tolerance against the theoretical m / z for all precursors in the search space, resulting in a set of candidate peptides. The candidate peptide that best explains the spectrum in terms of theoretical fragment ions is then considered as the peptide spectral match (PSM). No further prior information on the fragments is required.

[0042] Peptide-centric analysis / peptide-centric search can be defined as follows: data analysis of data, which may be DDA or DIA data acquired in an LC-MS / MS experiment, where the search is precursor-centric. Possible predicted peptides and their fragments derived from predicted or empirical spectral libraries are queried against spectra in the MS1 and MS2 dimensions. In this analysis, the peptide's spectral information is required, especially the possible observed fragment ions with retention time, ion mobility, and relative fragment intensity. This information is used to narrow down the peptide search space by querying only those spectra that are within a certain m / z, iRT, or IM tolerance range and scoring the matches. Having this additional information significantly improves the sensitivity of the analysis by leading to a more robust score.

[0043] Moreover, a specific confidence estimate in MRM by false discovery rate cannot be made as in classical shotgun proteomics. Therefore, a new approach was developed by measuring transitions for non-existent peptides (decoy transitions) (Reiter L, Rinner O, Picotti P, Huttenhain R, Beck M, Brusniak MY, Hengartner MO, Aebersold R. "mProphet: automated data processing and statistical validation for large-scale SRM experiments." Nature methods 2011, 8(5):430-435). Data from these decoy transitions can be used to derive a false discovery rate, as is done in shotgun proteomics. This confidence estimate by false discovery rate is necessary to determine the significance level of the data and to allow user-defined quality filtering of the data. SWATH / DIA data is different from MRM data. In contrast to MRM, complete fragment ion spectra are recorded using the SWATH / DIA method. The time resolution is usually selected similarly to MRM. Comparing SWATH / DIA with shotgun proteomics, the difference is that in SWATH / DIA the window for precursor selection is usually chosen wider (e.g. up to 25Th or around 25Th or up to 32Th instead of approximately 1Th as in shotgun proteomics), so that fragment ion spectra are derived from a much larger number of precursors. This high complexity of the fragment ion spectra makes it impractical / inefficient to analyze the data as in shotgun proteomics using database searches. However, the data can be analyzed similarly to MRM data, with the added advantage that the time resolution of the data is higher. This can be done by extracting the ion currents corresponding to the transitions in MRM.The data obtained can then be analyzed in a manner very similar to MRM. In all variants of mass spectrometry coupled with LC, the analysis is performed after digesting the proteins in the sample for MRM experiments into smaller peptides. The resulting peptide mixture is usually separated by chromatography to reduce the complexity of the sample. Chromatographic separation adds a time dimension, i.e. retention time (RT), to the data recorded by the mass spectrometer. Data-independent acquired data can also be analyzed in a spectrally centric or non-targeted manner, where a query is performed in the search space based on the data, for example based on the precursor ion (MS1) signal in the data (different from SWATH). This is the same analysis type typically used for DDA. In addition, various quantification techniques can be used, such as isobaric labeling, such as TMT or iTRAQ, or stable isotope labeling with amino acids in cell culture (SILAC), or heavy isotopically labeled peptides can be added to perform absolute quantification of peptides and proteins. Parallel reaction monitoring (PRM) can also be used, which is similar to MRM but is performed on a high resolution instrument in which fragment ion scans (MS2 scans) are acquired over the full range of the targeted analyte(s) in the analysis.

[0044] SRM assays are quantitative mass spectrometry-based assays specific for peptides or proteins of interest similar to antibodies for Western blotting, but with higher multiplexing capabilities and faster development times (assays for 100 peptides can be developed in 1 hour). We have previously demonstrated that SRM can quantify proteins with a wide range of cellular abundances, down to less than 50 copies per cell in whole cell lysates (see P Picotti et al., Cell 138 (4), 795 (2009) and Picotti at al. Nature Methods, VOL.9 NO.6, JUNE 2012, which references relate to SRM technology specifically included in this disclosure), resolve proteins with high (>95%) sequence overlap, and measure target peptides across multiple samples. Thus, this technology allows for quantitative measurement of specific peptides in highly complex samples. Recently, further developments in SRM methods include SRM-like methods based on data-independent acquisition of product ion spectra and their targeted analysis (SWATH methods, see LC Gillet et al., Mol Cell Proteomics 11 (6), O111 016717 (2012), the disclosure of which is incorporated herein as it relates to SWATH methods and data extraction).

[0045] According to a first preferred embodiment, for detection of the conformational state itself, in parallel to steps 1 to 3, an original complex mixture (cell extract) comprising said at least one protein and not exposed to said conditions inducing structural changes is subjected to steps 1 to 3 to generate an enriched fragment control sample, and the determination of the conformational state of the at least one protein is based on a quantitative comparison of the analytical analysis of the enriched fragment sample with the analytical analysis of the enriched fragment control sample.

[0046] According to another preferred embodiment, for detecting changes in conformational states depending on different conditions in a complex mixture, two complex mixtures are generated by subjecting a first complex mixture and a second complex mixture individually to steps 1 to 3, thereby exposing them to different conditions inducing a conformational change in said at least one protein, and the determination of the conformational change of the at least one protein is based on a comparison between an analytical analysis of the first enriched fragment sample and an analytical analysis of the second enriched fragment sample.

[0047] Typically, the conditions inducing a structural change in said at least one protein in said complex (cell extract) mixture are preferably selected from the group consisting of temperature change, pressure change, ionic strength change, pH change, metabolic enhancer change, addition of a drug / small molecule, addition of a metabolite, addition of a ligand including addition of a protein, addition of a peptide, addition of a lipid, addition of DNA, addition of RNA, disease / health state or condition and genetic variation (e.g. mutation, etc.), or a combination thereof, addition of a chaotrope, chemical modification including post-translational modification, in particular phosphorylation, disulfide bridge formation, ADP-ribosylation, ubiquitination, sumoylation, acetylation, methylation, oxidation, glycosylation, or a combination thereof.

[0048] Preferably, in step 2, peptides and proteins are removed by filtration, separation, or any other concentration step.

[0049] For the purposes of the present invention, the term "enrichment" is defined as follows: a first fragment sample is obtained by limited proteolysis of a complex (cell extract) mixture under conditions in which at least one protein is in the original conformational state of interest, and the enrichment step corresponds to the removal of larger peptides and proteins or other biomolecules not of interest from said first fragment sample, for example via a filtration, separation or concentration device, thus obtaining an enriched fragment sample.

[0050] Preferred methods include size filtering (e.g., using a 10k MWCO filtration device), chromatography including size exclusion, hydrophobic, or anion exchange chromatography, physical removal including phase separation, absorption, precipitation, filtration, separation, or concentration based on hydrophilic / hydrophobic properties, filtration, separation, or concentration based on electric / magnetic fields, or combinations thereof. Filtration can also be performed using two or more filters, thus enriching for a particular peptide size range. That is, not only are large peptides removed, but also very small peptides whose amino acid sequences are too short to be specific enough or suitable for mass spectrometry are removed.

[0051] Preferably, in step 2, particularly good results can be obtained by removing peptides, proteins, or other biomolecules, or both, having a molecular weight greater than 20 kDa, preferably greater than 15 kDa, and most preferably greater than 10 kDa, from the first fragment sample.

[0052] According to the above, preferably in step 2, peptides having a molecular weight of less than 0.1 kDa, or less than 0.2 kDa, or less than 0.4 kDa can also be removed from the first fragment sample with good results.

[0053] According to another preferred embodiment, step 3 comprises a proteomics workflow including, inter alia, denaturation, C18 cleanup, or a combination thereof, prior to the actual analysis.

[0054] According to a preferred protocol, in the limited proteolysis step 1, a proteolytic system is used selected from the group consisting of protease K, thermolysin, subtilisin, pepsin, papain, α-chymotrypsin, elastase, and mixtures thereof.

[0055] In step 1, the proteolytic system is preferably used at a concentration in the range of 1 / 50 to 1 / 10,000, preferably 1 / 100 to 1 / 1000, on a weight basis expressed as the ratio of enzyme content to biomolecule content relative to the total biomolecule content in the sample.

[0056] Step 1 may be carried out for a period of 1 to 60 minutes, preferably 2 to 30 minutes, or 2 to 10 minutes, or 2 to 5 minutes, more preferably at a temperature of 20°C to 40°C.

[0057] Preferably, the temperature in the limited proteolysis step 1 is in the range of 20° C. to 40° C. or 4° C. to 90° C. The temperature range is usually around room temperature (20° C. to 25° C.) or 37° C., while thermolysin is active up to 80° C., and 4° C. is also applicable for slowing down the proteolysis reaction.

[0058] The characteristics of the non-specific proteases used are summarized in the following table:

[0059] [Table 1]

[0060] To perform quantitative determination, the heavily labeled fragments that are characteristic when the limited proteolysis of step 1 is brought about but remain after the removal step 2 can be spiked into the original complex mixture and / or the first fragment sample and / or the enriched fragment sample. Thus, if desired, absolute quantification can be achieved using heavily labeled synthetic internal standard peptides. This approach can be applied directly to unfractionated proteomic extracts or can be coupled with various isotope labeling and sample fractionation techniques previously used in proteomic experiments (e.g., iTRAQ and TMT labeling, as well as the TAILS workflow, O Kleifeld et al., Nature biotechnology 28 (3), 281 (2010)).

[0061] The analytical analysis in step 3 is preferably carried out using specific quantitative mass spectrometry based assays in the form of selected reaction monitoring (SRM) and / or data independent acquisition (DIA) of product ion spectra.

[0062] Typically, a complex (eg, cell extract) mixture of additional proteins and / or other biomolecules is a complex natural biological matrix.

[0063] At least one protein is usually a protein based solely on proteinogenic amino acids or based on proteinogenic amino acids and having post-translational modifications.

[0064] Furthermore, the present invention relates to the use of the methods detailed above for determining medically relevant conformations of proteins, for determining protein-based drugs, for determining the effect of drugs or other ligands on proteins, or for the quality control of protein-based pharmaceutical formulations.

[0065] The present invention also relates to the use of the methods detailed above in combination with peptide fragment enrichment techniques such as TAILS for the peptides generated by step 1.

[0066] The proposed method opens numerous possibilities in biomedical, biotechnological and pharmaceutical applications as well as in fundamental and biological research, only a few of which will be presented below. It provides a novel platform to measure protein conformational changes in addition to the traditional changes in protein abundance or protein modifications currently measured by MS.

[0067] The method is particularly promising for the detection and treatment (testing of new drugs) of diseases caused by protein misfolding and aggregation, such as Alzheimer's or Parkinson's disease. Peptides carrying structural information can be used to investigate the structure of disease-related proteins in clinical samples and have potential as disease biomarkers. Furthermore, they can be used in drug screening to test the ability of chemical modulators (drugs) to directly affect the aggregation process in cell extracts.

[0068] The technique can also be applied to monitor the stability and proper protein folding of protein-based drugs, a critical quality control step for pharmaceutical companies manufacturing drugs.

[0069] Because proteins can change conformation upon binding to a drug or other ligand, the method can also be used to identify receptors for the drug or ligand based on the conformational change detected.

[0070] This method can be used to study the structure of protein receptors of interest directly within a cellular matrix, aiding in the design of molecules that target those protein receptors.

[0071] Marker assays can be translated into kits for the diagnosis of human diseases (disease biomarkers).

[0072] The method can be used for quality control of protein-based pharmaceutical formulations in the drug development pipeline.

[0073] The method can be used in drug development for protein conformational diseases.

[0074] This method may help identify receptors for existing drugs and understand their mechanisms of action.

[0075] Further embodiments of the invention are defined in the dependent claims.

[0076] Preferred embodiments of the present invention are described below with reference to the drawings, which are intended to illustrate the present preferred embodiments of the present invention and are not intended to limit the invention. [Brief description of the drawings]

[0077] [Figure 1] Figure 1 shows a schematic diagram of the proposed approach. Native proteins are incubated with defined concentrations of ligands, including a control (vehicle) sample. Each sample is subjected to a short, rapidly quenched limited digestion step. Larger peptides and protein pieces are then removed, resulting in unique peptide populations depending on the protein conformation during the limited digestion. The remaining peptides can be used to implement standard proteomics workflows. [Diagram 2] FIG. 2 shows a comparison of the LiP-MS protocol with two different variants of the proposed Dark-LiP methodology using the drug rapamycin. [Diagram 3] FIG. 3 shows the identification of peptides from FKBP1 by LiP-MS and comparison with Dark-LiP. [Figure 4] FIG. 4 shows a comparison of the LiP-MS protocol with two different variants of the Dark-LiP methodology using the general kinase inhibitor staurosporine. [Figure 5A] Figure 5A shows a comparison of the LiP-MS and Dark-LiP protocols for the identification of staurosporine protein targets, where A shows the total number of proteins and peptides identified by LiP-MS and Dark-LiP. [Figure 5B] Figure 5B shows a comparison of LiP-MS and Dark-LiP protocols for the identification of staurosporine protein targets, where B shows the peptide length distribution in LiP-MS and Dark-LiP. [Figure 5C]Figure 5C shows a comparison of the LiP-MS protocol with the Dark-LiP protocol for the identification of staurosporine protein targets, where C shows a graph showing the number of kinases identified as potential drug targets (true positives) versus the total number of potential drug targets (true positives + false positives). [Figure 6] FIG. 6 illustrates the experimental design for identifying differences in protein structure by Dark-LiP. [Figure 7] FIG. 7 shows a comparative analysis of limited proteolysis of tau monomers and tau fibrils. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0078] Figure 1 shows a schematic of the proposed approach, where the condition that distinguishes conformational differences between a reference sample and an altered sample is the addition of a ligand. Native protein 100 is incubated with a defined concentration of ligand 101a in the altered sample (upper path), whereas no ligand is added to the control (vehicle) sample 101b (lower path). Each sample is subjected to a short (1-5 min) rapidly quenched limited digestion step 102, typically with a non-specific protease. Then, in step 103, larger peptides and protein pieces are removed (e.g., via filtration), thus obtaining a unique peptide population 104a / b according to the conformation of the protein during the limited digestion 102. The remaining peptides can then be used to perform a standard proteomics workflow (e.g., denaturation, C18 cleanup, LC-MS, and analysis). EXAMPLES

[0079] Example 1: Comparison of LiP-MS with the proposed so-called Dark-LiP technique The results of Example 1 are shown in Figure 2, which shows a comparison of the LiP-MS protocol with two different variants of the proposed approach, the Dark-LiP methodology, using the drug rapamycin. (FA: formic acid, TCEP: reducing agent tris(2-carboxyethyl)phosphine, CAA: alkylating agent chloroacetamide, DOC: deoxycholate, ABC: ammonium bicarbonate, LysC: endoproteinase LysC).

[0080] To compare the current LiP-MS protocol with the Dark-LiP configuration, four aliquots of 300 μg HeLa lysate were incubated with 2 μM rapamycin and four aliquots were incubated with carrier (DMSO). Each aliquot was then treated with proteinase K at a protein ratio of 1:100 (mass / mass). After stopping the reaction with sodium deoxycholate (DOC) and boiling, the sample was divided into three aliquots.

[0081] One aliquot is used to evaluate downstream processing strategies for LiP-MS.

[0082] Two aliquots are used to evaluate downstream processing strategies of Dark-LiP.

[0083] For the LiP-MS procedure, 50 μg of protein extract was reduced, alkylated, diluted with ammonium bicarbonate buffer, and digested with LysC and trypsin. Finally, the generated peptides were acidified with formic acid and cleaned with C18 resin.

[0084] In the Dark-LiP methodology, the PK-digested cellular proteome was diluted with ammonium bicarbonate and acidified with formic acid, thereby omitting sample reduction, alkylation, and LysC / trypsin digestion.

[0085] The first variant of Dark-LiP consisted in filtering the acidified sample through a 10 MWCO cutoff spin filter immediately after limited proteolysis, followed by peptide cleanup with C18.

[0086] In the second variant of Dark-LiP, the filtration step was omitted and PK-generated peptides were directly acidified and cleaned with C18 resin immediately after limited proteolysis.

[0087] The purified peptides obtained with both LiP-MS and Dark-LiP methodologies after C18 cleanup were analyzed by mass spectrometry.

[0088] Identification of peptides from FKBP1 by LiP-MS and Dark-LiP FIG. 3 shows the identification of peptides from FKBP1 by LiP-MS and Dark-LiP.

[0089] Figure 3A.1 shows the total number of proteins identified by LiP-MS and Dark-LiP. A similar number of proteins were identified by the two variants of Dark-LiP. LiP-MS identified the largest number of proteins.

[0090] Figure 3A.2 shows the total number of peptides identified by LiP-MS and Dark-LiP. A comparable number of peptides were identified by the two variants of Dark-LiP. LiP-MS identified the largest number of peptides.

[0091] The first variant of Dark-LiP (consisting in filtration of the acidified sample through a 10 MWCO cutoff spin filter immediately after limited proteolysis, followed by peptide cleanup with C18) is designated "Dark-Lip 10K" in Figures 3A.1 and 3A.2 and 3B.

[0092] The second variant of Dark-LiP (which consists in omitting the filtration step immediately after limited proteolysis and directly acidifying the PK-generated peptides and cleaning with C18 resin) is designated as "Dark-Lip C18" in Figures 3A.1 and 3A.2 and 3B.

[0093] FIG. 3B shows the number of FKBP1 peptides identified as potential drug targets (true positives) versus the total number of potential drug targets (true positives + false positives).

[0094] The analysis by LiP-MS identified a larger number of peptides and proteins than Dark-LiP (Figure 3), because these samples are inherently more complex without filtering out large proteins and peptides. Since peptides are obtained after protein fragmentation, it is also expected that the removal of the large protein fraction in Dark-LiP would result in the identification of fewer peptides than LiP-MS.

[0095] The number of proteins and peptides identified by the two variants of Dark-LiP is comparable, as is the number of FKBP1 peptides (true positives). Therefore, the "simpler" Dark-LiP variant, i.e. the one omitting the spin filter with a cutoff of 10 MWCO, is selected for further experiments.

[0096] To test the performance of Dark-LiP and LiP-MS in identifying FKBP proteins (targets of rapamycin) as potential drug targets, the machine learning pipeline LiP-quant was used to score peptides identified in several analyses and obtain a list of potential peptide drug targets ranked by LiP score.

[0097] The number of peptides derived from known described rapamycin targets (FKBP proteins) was plotted against the total number of peptide candidates (true positives + false positives) (Figure 3). Thus, a steeper line implies a higher number of peptides derived from FKBP than a flatter line. These results show that both variants of the Dark-LiP methodology outperform standard LiP-MS by identifying more FKBP peptides with higher scores than LiP-MS.

[0098] LiP-Quant is a drug target deconvolution data analysis pipeline based on mass spectrometry coupled with limited protein digestion that works across species including human cells. Machine learning is used to identify features indicative of drug binding and combine them into a single score to identify protein targets of small molecules and estimate their binding sites.

[0099] Example 2: Comparison of the LiP-MS protocol with one variant of the Dark-LiP methodology (C18 variant) using the general kinase inhibitor staurosporine Figure 4 shows a comparison of the LiP-MS protocol with one variant (C18 variant) of the DARK-LiP methodology using the drug staurosporine. (FA: formic acid, TCEP: reducing agent, CAA: alkylating agent, DOC: deoxycholate, ABC: ammonium bicarbonate, LysC: endoproteinase LysC).

[0100] The general kinase inhibitor staurosporine was used as a model system to compare the performance of LiP-MS and Dark-LiP in a drug-dose response experimental setup (Figure 4).

[0101] Kinases are a large family of enzymes that phosphorylate other proteins and are essential for normal cellular function. Due to the large number of kinases, the staurosporine assay is commonly used to compare the efficiency of different methodologies to identify drug targets.

[0102] 175 μg of native HeLa protein lysate was treated in duplicate with seven different concentrations of the drugs staurosporine and DMSO and samples were processed according to the LiP-MS protocol or the C18 variant of the Dark-LiP methodology as described above.

[0103] FIG. 5 shows a comparison of LiP-MS and Dark-LiP for the identification of staurosporine protein targets.

[0104] FIG. 5A shows the total number of proteins and peptides identified by LiP-MS and Dark-LiP.

[0105] FIG. 5B shows the peptide length distribution in LiP-MS and Dark-LiP.

[0106] FIG. 5C shows the number of kinases identified as potential drug targets (true positives) versus the total number of potential drug targets (true positives + false positives).

[0107] Comparable to the rapamycin experiments, LiP-MS analysis identified a larger number of peptides and proteins than Dark-LiP, while Dark-LiP even identified more semitryptic peptides than LiP-MS (Figure 5A). Furthermore, the peptides identified by Dark-LiP were also longer in length due to the absence of a second fragmentation step by trypsin / LysC (Figure 5B).

[0108] LiP-quant was used to score the identified peptides based on the correlation between peptide abundance and drug concentration, thereby generating a list of the most likely staurosporine targets. To visually compare the performance of Dark-LiP and LiP-MS in identifying kinases as drug targets, the number of kinases identified as target candidates was plotted against the total number of candidates (Figure 5C). As above, a steeper line means a larger number of kinases were identified as top drug target candidates, representing a more efficient method. These results indicate that Dark-LiP outperformed LiP-MS by identifying more kinases as targets of staurosporine.

[0109] Example 3: Dark-Lip technology enables detection of conformational changes in proteins It is estimated that over 45 million people worldwide suffer from dementia, the most common form of which is Alzheimer's disease, which is characterized by the accumulation of abnormal fibrillar tangles of the microtubule-associated protein tau, a natively misfolded protein that has a highly flexible conformation under physiological conditions.

[0110] The references illustrate clinically relevant conformational changes of tau protein in health and disease that can be discriminated by Dark-LiP.

[0111] In the Dark-LiP procedure, during the limited proteolysis step, proteinase K fragments tau protein in solution. The rate of fragmentation of a particular region of tau depends on the accessibility of that region to proteinase K. Monomeric tau is described in the literature as a highly disordered protein composed of protein segments with high flexibility. As a result, regions of monomeric tau are more accessible to proteinase K, which leads to fast kinetics of fragmentation and ultimately to a higher abundance of peptides. In contrast, fibrillar tau is described in the literature as a large structure of aggregated molecules of monomeric tau. This means that some regions and molecules of fibrillar tau are protected from proteinase K by other regions and molecules of tau aggregated together. Overall, this lower accessibility leads to a slower fragmentation rate and, consequently, a lower abundance of peptides.

[0112] To demonstrate this, 6 μg of monomeric tau was diluted using three aliquots of 100 μL LiP-MS buffer, and 6 μg of fibrillar tau was diluted using three aliquots of 100 μL LiP-MS. Each aliquot was then treated with proteinase K at a 1:50 molecular ratio (50 molecules of tau for each PK molecule) and incubated at room temperature for 2 min. The reaction was stopped by adding sodium deoxycholate (DOC) and boiling.

[0113] The PK-fragmented tau sample was then diluted with ammonium bicarbonate and acidified with formic acid (thereby omitting sample reduction, alkylation, and LysC / trypsin digestion as a feature from the Dark-LiP methodology). After acidification, the PK-generated peptides were cleaned with C18 resin. This step removes large peptides and proteins that are retained on the C18 resin.

[0114] Figure 6 illustrates the experimental design to identify differences in protein structure by Dark-LiP.

[0115] Fragmented tau samples generated upon processing with the Dark-LiP workflow will be analyzed by mass spectrometry and peptides identified in both conditions will be compared using a statistical Student's t-test.

[0116] The two-sample Student's t-test is used to test whether two sample groups (two populations) differ in terms of a quantitative variable based on a comparison of two samples drawn from the two groups (Equation 1). In other words, the two-sample Student's t-test allows testing the null hypothesis of whether the means of two populations are equal (the samples are measured on a quantitative continuous variable). For the Student's t-test, the mean and standard error of the mean are used to compare the two samples and assume a normal distribution of the data.

[0117] The comparison of the means is carried out according to Equation 1, and the output is a value t. This t-value is a measure of the magnitude of the difference between the means of the two populations relative to the variability of the measurements. The higher the value of t, the larger and more significant the difference between the two populations and the smaller the variability between the measurements. A particular t-value can then be converted into the probability of obtaining that t-value (p-value) via a two-sample t-distribution value. The relationship between t-values ​​and p-values ​​has already been established for a particular probability distribution of the defined test, so that the p-value can be calculated directly (in this case, a two-sample t-test and a normal distribution are assumed). The p-value (p stands for probability) is frequently used to measure statistical significance, describing the likelihood that the observed difference in the means is explained by chance. The p-value represents a probability between 0% and 100%, so for example, a p-value of 0.01 corresponds to a probability of 1%. Thus, the lower the p-value, the lower the probability that the observed value is explained by chance, and consequently the higher the statistical significance (in the example described, the observed value would occur by chance in 1% of a population of random measurements).

[0118]

number

[0119] Formula 1: Formula for performing a statistical t-test between two independent populations t: A measure of the magnitude of the difference between populations in the variability of measurements. The t-value ranges from 0 (no difference in the mean) to infinity. x1: The mean of the measurements for population 1 x2: The mean of the measurements for population 2 s1: Standard deviation of the measurements for population 1 s2: Standard deviation of the measurements for population 2 n1: Number of measurements in population 1 n2: The number of measurements in population 2.

[0120] In our example with tau protein, a t-test is used to calculate whether the difference in abundance of a particular peptide generated by Dark-LiP fragmentation of monomeric and fibrillar tau is statistically significant. If the difference in average abundance between the three measurements made on monomeric tau is different from the average abundance between the three measurements made on fibrillar tau, it is likely that the fragmentation rate and, consequently, the accessibility of that particular peptide to proteinase K was also different between the two tau isoforms. Thus, peptides with statistically significant differences in abundance are directly linked to structural differences in tau. To globally visualize the results of the statistical analysis, the inverse logarithm of the p-value (derived from the t-value obtained in Equation 1) was plotted against the logarithm of the fold change in peptide abundance (the ratio between x1 and x2 described in Equation 1) for each peptide (Figure 7). Because p-values ​​can vary widely between peptides, a logarithmic transformation is used to improve visualization. Otherwise, the peptides with the lowest p-values ​​would be located far away from the peptides with the highest p-values, making it difficult to visualize the data efficiently. Because the logarithm of numbers lower than 1 is negative, the inverse logarithm is used to convert the data to positive values ​​that are easier to visualize and interpret. By applying a logarithmic transformation to the fold changes, we can also efficiently visualize fold changes with widely varying values ​​and distinguish between peptides where x1 is greater than x2 (positive AVG Log2 ratio) and peptides where x1 is less than x2 (negative AVG Log2 ratio).

[0121] Furthermore, from Equation 1, we observe that the p-value depends not only on the difference in abundance between the two populations, but also on the number of measurements and the standard deviation of those measurements (which correlates with the variability of observations between replicates). Thus, it is possible that the p-value derives primarily from a large difference in the means of the sample groups (the fold change given by x1-x2 in the numerator of Equation 1) or that the p-value is primarily due to a low standard error of the mean (the low variability between replicates, given by (s1 / n1) in the denominator of Equation 1). 2 +(s2 / n2) 2 ), since n1 = n2 = 3 replicates). As a result, the p-value was plotted against the fold change (ratio between x1 and x2), allowing a simultaneous visualization of the difference in the means of the two sample groups (fold change) and an estimate of the variability of the measurements.

[0122] In proteomics studies, values ​​of 1% for significance and a fold change of 2 are commonly accepted in the industry. Therefore, peptides with p-values ​​greater than 0.01 (equivalent to a -log2P value of 6.64, represented by the horizontal dashed line) and fold changes lower than 2 (log2 ratios lower than -1 and higher than 1, represented by the vertical dashed lines) were considered statistically non-significant. These non-significant peptides are represented by small dots and are located in the area below and between the dashed lines.

[0123] Peptides with p-values ​​lower than 0.01 and fold changes higher than 2 were considered statistically significant. Statistically significant peptides from tau are represented as crosses, while "Y" shapes represent statistically significant peptides from bacterial proteins co-purified during the preparation of tau protein. If x1 is greater than x2, the fold change (ratio between x1 and x2) is greater than 1, and this peptide is displayed in the right part of the graph. If x2 is greater than x1, the fold change is less than 1, and this peptide is displayed in the left part of the graph. Overall, peptides with high values ​​on the y-axis correspond to peptides with low p-values, and thus correspond to high statistical significance (which may correlate with low standard deviations and, therefore, high reproducibility between replicates). Symbols with high module values ​​on the x-axis correspond to large abundance differences, which may correlate with larger structural changes.

[0124] Overall, limited proteolysis of the two different forms of tau using the Dark-LiP methodology generated 412 peptides with significant differences in abundance.

[0125] According to the above explanation, peptides represented within the triangular region in Figure 7 are high on the y-axis and low on the x-axis, and therefore have high -LogP values ​​and low fold changes, which means that the differences in abundance of these peptides are small between monomeric and fibrillar tau (which may be related to mild structural changes), but the quantification is highly reproducible (low standard deviation) between three replicates of the same conditions.

[0126] The peptide within the circle in Figure 7 has a high value on the x-axis and a low value on the y-axis (hence a high fold change and a relatively low -LogP value), so that this peptide may be associated with a strong structural difference between the two tau variants, although the variability in the abundance of the peptide was large between measurements (leading to relatively low statistical significance).

[0127] The peptides within the rectangle in Figure 7 have moderate x and y values, and therefore moderate fold changes and moderate -LogP values. Because they can be reproducibly measured between replicates and have relatively large differences in abundance, these peptides can be used to robustly identify structural changes between tau variants.

[0128] In our analysis, monomeric tau was considered as the first variable x1 in Equation 1, and fibrillar tau was considered as the second variable x2 in Equation 1. Figure 7 shows that the majority of tau peptides with high statistical significance and high fold change are located in the positive region of the x-axis, which means that x1 is greater than x2, and consequently the abundance of peptides is higher in monomeric tau than in fibrillar tau. This indicates that, as expected (as described above), monomeric tau was more susceptible to proteolytic fragmentation by proteinase K than fibrillar tau.

[0129] Overall, Example 3 demonstrates that the Dark-LiP technique can be used to reveal differences in the proteolytic susceptibility of several regions of monomeric and fibrillar tau, thereby highlighting the distinct structures of the two tau variants. [Explanation of symbols]

[0130] 100% Natural Protein 101a Ligand 101b Control (vehicle) sample 102 Limited digestion process 103 Filtration 104a / b Unique peptide population DDA Data-Dependent Acquisition DIA Data Independent Acquisition LC Liquid Chromatography LC-MS Liquid Chromatography Coupled with Mass Spectrometry LiP Limited protein digestion MRM Multiple Reaction Monitoring MS mass spectrometry RT retention time SILAC: Stable isotope labelling of amino acids in cell cultures SRM Selected Reaction Monitoring SWATH: Sequential Windowed Acquisition of All Theoretical Fragment Ion Mass Spectra.

Claims

1. 1. A method for detecting the conformational state of at least one protein, said at least one protein being contained in a complex mixture of further proteins and / or other biomolecules, including a complex cell extract mixture, wherein said at least one protein in said complex mixture is exposed to conditions which induce a conformational change in said at least one protein, optionally after an extraction and / or lysis step, comprising the following series of steps:

1. Obtaining a first fragment sample by limited proteolysis of said complex mixture under conditions in which said at least one protein is in its original conformational state to be detected, followed immediately by:

2. Removing large peptides and proteins or other biomolecules from the first fragment sample to form an enriched fragment sample; 3. Analyzing the enriched fragment sample to determine the fragments that are characteristic of the limited proteolysis of step 1 and that remain after removal step 2, thereby determining the conformational state of the at least one protein; The method comprising:

2. To detect the conformational state itself, in parallel with steps 1 to 3, the original complex mixture comprising the at least one protein and not exposed to the conditions that induce the structural change is subjected to steps 1 to 3 to generate an enriched fragment control sample, and determining the conformational state of the at least one protein is based on a quantitative comparison of the analytical analysis of the enriched fragment sample with the analytical analysis of the enriched fragment control sample; or 2. The method of claim 1, wherein detecting the change in conformational state depending on different conditions in the complex mixture comprises generating two complex mixtures by subjecting a first complex mixture and a second complex mixture to different conditions that induce a structural change in the at least one protein by subjecting them separately to steps 1 to 3, and determining the conformational change of the at least one protein based on a comparison of the analytical analysis of the first enriched fragment sample with the analytical analysis of the second enriched fragment sample.

3. 3. The method of claim 1 or 2, wherein the conditions that induce a structural change in the at least one protein in the complex mixture are selected from the group consisting of: a temperature change; a pressure change; a change in ionic strength; a pH change; a metabolic enhancer change; addition of a ligand, including addition of a drug / small molecule, addition of a metabolite, addition of a protein, addition of a peptide, addition of a lipid, addition of DNA, addition of RNA; a disease / health state or condition, including a mutation, and a genetic variation, or a combination thereof; addition of a chaotrope; a chemical modification, including a post-translational modification, including phosphorylation, disulfide bridge formation, ADP-ribosylation, ubiquitination, sumoylation, acetylation, methylation, oxidation, glycosylation, or a combination thereof.

4. A method as described in claim 1 or 2, wherein in step 2, peptides and proteins are removed in a filtration process, separation process, or another concentration process, including size filtration, chromatography, including size exclusion chromatography, hydrophobic chromatography, or anion exchange chromatography, physical removal, including phase separation, absorption, precipitation, filtration, separation, or concentration based on hydrophilic / hydrophobic properties, filtration, separation, or concentration based on electric / magnetic fields, or combinations thereof.

5. A method according to claim 1 or 2, wherein in step 2, peptides, proteins, and / or other biological molecules having a molecular weight greater than 20 kDa, or greater than 15 kDa, or greater than 10 kDa are removed from the first fragment sample.

6. The method described in claim 1 or 2, wherein step 3 includes a proteomics workflow prior to actual analysis.

7. 3. The method according to claim 1, wherein in step 1, a proteolytic system selected from the group consisting of protease K, thermolysin, subtilisin, pepsin, papain, α-chymotrypsin, elastase, and mixtures thereof is used.

8. 3. The method according to claim 1, wherein in step 1, the proteolytic system is used at a concentration in the range of 1 / 50 to 1 / 10,000, or 1 / 100 to 1 / 1,000, by weight, expressed as a ratio of enzyme content to biomolecule content, relative to the total biomolecule content in the sample.

9. 3. The method of claim 1, wherein step 1 is carried out for a period of from 1 minute to 60 minutes, or from 2 minutes to 30 minutes, or from 2 minutes to 10 minutes, or from 2 minutes to 5 minutes.

10. 3. The method according to claim 1 or 2, wherein for quantitative determination, heavily labeled fragments that are characteristic when the limited proteolysis of step 1 is brought about but remain after removal step 2 are spiked into the original complex mixture and / or the first fragment sample and / or the enriched fragment sample.

11. A method as described in claim 1 or 2, wherein the analytical analysis in step 3 is performed using an assay based on a specific quantitative mass spectrometry method in the form of selected reaction monitoring (SRM) and / or data-independent acquisition of product ion spectra.

12. 3. The method of claim 1 or 2, wherein the complex mixture of further proteins and / or other biomolecules is a complex natural biological matrix.

13. 3. The method of claim 1, wherein the at least one protein is a protein based solely on proteinogenic amino acids or a protein based on proteinogenic amino acids and having post-translational modifications.

14. 3. Use of the method according to claim 1 or 2 for the hypothesis-free determination of the conformation of said at least one protein that has undergone a conformational change after an induced perturbation in a complex mixture under investigation or a medically relevant conformation of said protein, for the determination of protein-based drugs, for the effect of drugs or other ligands on proteins, or for the quality control of protein-based pharmaceutical formulations.

15. 3. Use of the method according to claim 1 or 2 in combination with a peptide fragment enrichment technique such as TAILS for the peptides generated by step 1.

16. 3. Use of conformationally altered peptides / proteins contained in the enriched fragment sample obtained in step 2 of the method according to claim 1 or 2 as biomarkers.

17. The method described in claim 1 or 2, wherein step 3 includes a proteomics workflow including denaturation, C18 cleanup, or a combination thereof prior to actual analysis.

18. The method described in claim 9, wherein step 1 is carried out at a temperature of 20°C to 40°C.