A pressure-based method for detecting protein-molecular affinity
By cleavage and size exclusion chromatography of biological tissue samples under pressure cycler and data-independent mass spectrometry detection, the problems of limited flux and insufficient accuracy of interprotein affinity detection in the prior art are solved, and efficient evaluation of intra-protein affinity in mixed samples is achieved.
Patent Information
- Application Number
- CN202310109993.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-14
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-02-14
AI Technical Summary
Existing protein interaction detection methods have limited throughput, inaccurate results and inability to apply to mixed samples in evaluating interprotein affinity, especially inadequate detection capabilities for clinical samples such as serum, feces and tissue samples.
The pressure cycler is used to provide a stable and uniform pressure environment, cleavage biological tissue samples, and the protein complex is isolated by size exclusion chromatography, and data-independent mass spectrometry detection is used to evaluate the affinity between protein molecules by calculating the peak area changes of proteins under different pressures.
High-throughput and accurate detection of protein affinity in mixed samples is achieved, and it can comprehensively scan multiple intermolecular interaction forces without sample purification, providing a more stable and efficient evaluation of protein interactions.
Smart Images

Figure CN116699043B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of proteomics, and in particular relates to a method for detecting affinity between protein molecules based on pressure. Background Art
[0002] Proteins typically bind to other proteins or molecules through non-covalent protein-protein interactions to form protein complexes, which carry out vital functions in biological systems. Disease mechanisms are often mediated by interactions between proteins. Quantifying the stability of protein-protein interactions (PPIs) and the strength of drug-protein binding can advance our understanding of disease etiology, progression, and pathogenesis, aiding our understanding of vital functions, helping us more quickly identify key disease mechanisms and targets, and promoting clinical translation and precision medicine.
[0003] Existing methods for assessing protein interactions include affinity purification mass spectrometry (AP-MS), proximity labeling, cross-linking mass spectrometry (XL-MS), protein thermal shift, and size exclusion chromatography-SWATH (SEC-SWATH). Labeling-based methods such as AP-MS and proximity labeling require the additional expression of bait proteins in cells and cannot be applied to clinical samples that do not express proteins independently, such as serum, feces, and tissues. Although XL-MS can measure protein interactions in their native state, the depth of identification is limited by the need for cross-linking agents and complex computational identification methods to assess protein interactions. Thermal shift is a relatively high-throughput method for identifying protein interactions. It exploits the property of proteins to denature at specific temperatures and measures the stability of proteins by measuring the number of hydrophobic groups exposed at different temperatures. For example, when a protein binds to other proteins or molecules, its stability increases, and the number of hydrophobic groups exposed at the same temperature decreases. This allows different protein solubility defects to be identified to screen protein ligands or evaluate protein interactions. Thermal drift technology measures protein complexes in different states by regulating temperature. However, its reliance on PCR thermal cyclers for temperature control results in low heat conduction efficiency and hysteresis. The heating method is uneven, leading to differences in internal and external temperatures of the sample, potentially leading to errors in the evaluation results. Furthermore, thermal drift technology cannot measure the affinity between multiple molecules in a mixed sample, resulting in certain throughput limitations.
[0004] In recent years, for example, the Aebersold laboratory has developed a method for systematically detecting protein complexes using complex-centric analysis, which allows for the fractionation of protein complexes by size exclusion chromatography (SEC). This technique allows for high-throughput identification of complexes without the need for epitope tags. However, this method can only assess the relevance of protein interactions within a protein complex, but cannot quantitatively assess the affinity between proteins.
[0005] Therefore, there is still a need to develop an effective method to evaluate the affinity between proteins to overcome the above defects. Summary of the Invention
[0006] The present invention is accomplished based on the following principles:
[0007] Pressure is used to provide stable and uniform energy to protein-containing samples. As pressure increases, the molecular interactions are disrupted, leading to the separation of bound molecules. The size exclusion column reveals differences in chromatographic peaks before and after the disruption of these interactions. Each chromatographic fraction is collected and analyzed by mass spectrometry to obtain qualitative and quantitative information about the molecules in each fraction. This allows calculation of the distribution differences of each molecule in the mixed sample at different pressures, thereby inferring the strength of different intermolecular interactions.
[0008] In one aspect, the present invention provides a method for detecting affinity between protein molecules based on pressure, comprising the following steps:
[0009] 1) Place the biological tissue sample at P1 to P n The lysis was performed under a series of pressures to obtain a series of lysates, where P1 represents the first pressure, P n represents the nth pressure, where n is an integer greater than 3, and the difference Δp between two adjacent pressures is a constant;
[0010] 2) centrifuging the series of lysates obtained at various pressures in step 1), taking the supernatants, performing high performance liquid chromatography (HPLC) separation based on size exclusion chromatography (SEC), and collecting fractions;
[0011] 3) proteolyzing the fractions collected in step 2) into peptide fragments using trypsin;
[0012] 4) performing mass spectrometry data acquisition on the different fractions after proteolysis in step 3) using a data-independent acquisition (DIA) method;
[0013] 5) performing protein library analysis on the DIA mass spectrometry data of the different fractions collected in step 4) to obtain a protein quantitative matrix;
[0014] 6) Using a protein quantification matrix, the abundance distribution of protein molecules in different fractions of biological tissue samples can be inferred;
[0015] 7) Grouping the fractions, and performing correlation analysis based on the abundance distribution of the target protein pairs for analyzing intermolecular forces in each group of fractions to calculate the correlation coefficients of the target protein pairs corresponding to each group of fractions;
[0016] 8) Based on the correlation coefficients of the target protein pairs corresponding to the fractions of each group, identifying whether there are interaction peaks between the target protein pairs;
[0017] 9) In the case where there is an interaction peak between the target protein molecule pair, the peak area of the corresponding interaction peak is calculated using the protein quantitative matrix at each pressure, and the affinity coefficient AF of the protein molecule pair is calculated using the following formula, wherein AF is used to characterize the change in the peak area of the interaction peak between the protein molecule pair as the pressure increases:
[0018]
[0019] in
[0020] In the above formula, A1 represents the peak area of the interaction peak at the first pressure, A n-1 represents the peak area of the interaction peak at the n-1th pressure, A n is the peak area of the interaction peak at the nth pressure, A n-j-1 represents the peak area of the interaction peak at the nj-1th pressure, a n-j represents the peak area of the interaction peak at the njth pressure;
[0021] k j For the adjacent pressure interval P n-j-1 and P n-j The weight of the interaction peaks in the corresponding adjacent pressure intervals represents the degree of change A in the peak area. n-j-1 -A n-j Regarding the degree of influence of the affinity coefficient AF, the constant k is the pressure interval coefficient, which is a constant greater than 1.
[0022] The above formula is explained in detail below.
[0023] From D j It can be seen from the expression that from j = 1 to j = n-2, D j They are And considering that k>1, the adjacent pressure interval sequence [(A1-A2), (A2-A3),…, (A n-2 -An-1 )] the corresponding weight coefficient is from k n-2 to k in descending order, that is, through k j The setting can make the adjacent pressure intervals closer to the front have a greater impact on the affinity coefficient AF, and conversely, the adjacent pressure intervals closer to the back have a smaller impact on the affinity coefficient AF. Specifically, combined with the calculation formula of AF, it can be seen that the change in the peak area of the interaction peak corresponding to the adjacent pressure interval with a closer degree of decline can make the AF value decrease with a steeper trend.
[0024] As for the AF calculation formula, and The setting of subtracting and dividing the difference by Δp is to control the data scale of the exponent in the AF formula within a more reasonable range, preventing the data corresponding to each adjacent pressure interval from being aggregated in a small interval and failing to fully reflect the impact of the peak area change of the corresponding interaction peak on the AF value.
[0025] Hereinafter, the method of the present application is described in detail step by step.
[0026] Step 1)
[0027] In step 1), the biological tissue sample is subjected to cyclic pressurization treatment for a period of time under a series of pressures to lyse and obtain a series of lysates (if the sample is cell-free, lysis is not required and the pressurization time can be appropriately shortened. For the sake of consistency in description, the sample after pressurization is still referred to as lysate in the following description).
[0028] The source of the biological tissue sample is not particularly limited. For example, it can be derived from: animals, such as animal organs, tissues, cells, body fluids, secretions, excretions, etc.; plants, such as roots, stems, leaves, flowers, fruits, etc.; such samples require the removal of plant cell walls using cellulase; or microorganisms, such as bacteria and fungi, require the removal of bacterial cell walls using lysozyme. In one embodiment, the biological tissue sample is mouse liver.
[0029] The purpose of lysis is to destroy the cell membrane to release proteins. Protein complexes mainly exist inside cells, and there may also be protein complexes on the cell membrane, which need to be fully released. Lysis under pressurized conditions can promote the separation of protein complexes, thereby releasing proteins faster and more fully. In addition, pressurization can make the method of the present invention compatible with tissue samples. There are connections between cells in the tissue sample, which hinder the release of intracellular proteins. This obstacle can be broken under the action of pressure. After the pressure is removed, the separated protein complex molecules cannot be recombined. In addition, if the biological tissue sample to be processed does not have a cell membrane, the pressurization time can be appropriately shortened.
[0030] A weaker tissue lysis buffer should be used for lysis. This buffer should disrupt the cell membrane and release proteins while minimizing the risk of disrupting protein-protein interactions. This buffer can also maintain the stability of protein complexes to a certain extent. For example, NP-40 can be used as the detergent component in tissue lysis buffers.
[0031] Specifically, the lysis buffer includes 0.3-0.8 vol% NP-40, 35-60 mM HEPES (pH 7.5), 130-160 mM NaCl, 35-60 mM NaF, 190-215 μM Na 3 VO 4 , 0.7-1.2 mM PMSF, and 1× protease inhibitor cocktail. At this time, 40-50 μl of lysis buffer can be added to 1-5 mg of biological tissue sample.
[0032] As mentioned above, the purpose of applying pressure during lysis is to use pressure to provide stable and uniform energy to the sample containing protein, and to destroy the intermolecular interaction force under pressure, thereby separating the protein molecules (i.e., protein complexes) in a bound state. When lysis is performed under different pressures, the degree of separation of the protein complex under different pressures is different, so that the existence state of the molecules in the obtained series of lysates is different, and the chromatographic peaks after separation by high-performance liquid chromatography based on size exclusion chromatography SEC are also different ( Figure 2 ).
[0033] There are no specific restrictions on the equipment used for pressurization, as long as it can provide the appropriate pressure environment. For example, a pressure cycler (PCT) can be used to perform pressurization on biological tissue samples. Using a PCT, temperature and pressure can be precisely controlled over a short period of time, providing a stable, controllable pressure environment.
[0034] In step 1), a series of pressures P1 to P n The number of n can be 3 or more, that is, n is an integer greater than 3, such as 3, 4, 5, 6, 7, 8, 9, 10, etc. There is no upper limit for n. More n is conducive to more accurate results, but it also increases the difficulty of calculation. Therefore, n is preferably less than 10.
[0035] The set pressure parameter can be in the range of 5-50kpsi, that is, P1 ≥ 5kpsi and P n≤50kpsi. Within this pressure range, intermolecular affinity can be effectively measured. If the pressure exceeds 50kpsi, single protein molecules may be further broken into protein fragments, affecting HPLC separation. Pressures below 5kpsi may affect sample lysis (especially for samples containing cells), resulting in inadequate protein release.
[0036] The pressure difference Δp between adjacent cells is a constant value to simplify calculations. There are no specific limitations on Δp, as long as the pressure difference is sufficient to achieve significantly different degrees of separation. For example, Δp can be any value between 5 and 20 kpsi, such as 8, 10, 12, 15, 18, etc., but is not limited thereto.
[0037] In one embodiment, three pressures can be selected. In one embodiment, the first pressure = 15 kpsi, the second pressure = 25 kpsi, and the third pressure = 35 kpsi, corresponding to obtaining three lysates L1, L2, and L3.
[0038] In summary, by placing biological tissue samples in lysis buffer, n The lysates were lysed under a series of pressures to obtain a series of lysates L1 to L2 with different degrees of protein molecule separation. n , that is, the lysate L1 is obtained by cleavage at the pressure of P1, the lysate L2 is obtained by cleavage at the pressure of P2, ..., n The lysate L was obtained by pyrolysis under pressure. n .
[0039] Step 2)
[0040] In step 2), the series of lysates obtained in step 1) are centrifuged, and the supernatants are collected and subjected to high performance liquid chromatography separation based on size exclusion chromatography (SEC), and fractions are collected.
[0041] The purpose of centrifugation is to remove insoluble matter from the lysate to prevent clogging of the SEC column. The centrifugation conditions are not particularly limited. For example, the centrifugation can be performed at a speed greater than 18,000 rcf for 5-15 minutes.
[0042] In HPLC separations based on size exclusion chromatography (SEC), proteins of varying molecular sizes within a lysate can be separated into fractions based on their molecular weight. This allows protein complexes of varying molecular weights to be separated and distinguished. Proteins comprising the same complex will appear in the same fractions, allowing interaction peaks to be detected in subsequent steps.
[0043] For example, lysate L1 is subjected to high performance liquid phase separation to obtain fraction F. 1-1 to F1-m m fractions; HPLC separation of lysate L2 can obtain fraction F 2-1 to F 2-m m fractions; . . . . , and the lysate L n By high performance liquid separation, fraction F can be obtained n-1 to F n-m’ m ’ Here m refers to the number of fractions obtained from each lysate, which is the same.
[0044] Since lysate L1 to L n The degree of separation of protein molecules in the medium is different, so the fractions F obtained by each 1-1 to F 1-m With F 2-1 to F 2-m ,……,F n-1 to F n-m’ Each one is different.
[0045] For example, in step 1), the first pressure = 15 kpsi, the second pressure = 25 kpsi, and the third pressure = 35 kpsi, corresponding to obtaining three lysates L1, L2, and L3, fraction F is obtained from L1. 1-1 to F 1-90 , fraction F is obtained from L2 2-1 to F 2-90 , and fraction F is obtained from L3 3-1 to F 3-90 , a total of 270 fractions.
[0046] Step 3)
[0047] In step 3), the fractions collected in step 2) were respectively proteolyzed using trypsin.
[0048] The purpose of proteolysis with trypsin is to digest proteins into peptide fragments that can be detected by mass spectrometry.
[0049] In an embodiment, conventional trypsin proteolysis conditions can be used. For example, the enzymatic digestion can be performed at 37° C. overnight. The amount of trypsin used can be 1 (trypsin): 50 (total protein) by weight.
[0050] Alternatively, trypsin digestion can be followed by intracellular protease Lys-C digestion at 37°C for 4 hours. The amount of intracellular protease Lys-C can be used in a ratio of 1 (Lys-C enzyme): 100 (total protein) by weight. Digestion with intracellular protease Lys-C allows for more complete digestion, yielding more cleaved peptides, which facilitates mass spectrometry detection.
[0051] Step 4)
[0052] In step 4), the different fractions following the proteolysis in step 3) are analyzed by liquid chromatography-tandem mass spectrometry. The liquid chromatography utilizes a reversed-phase column to separate the peptides. For example, a 20-minute effective gradient at 5 μl / min can be used, from 5% to 32% buffer B (0.1% formic acid, 98% acetonitrile), i.e., from 95% to 68% buffer A (0.1% formic acid, 2% acetonitrile). Data-independent acquisition (DIA) is used for mass spectrometry data acquisition.
[0053] The purpose of mass spectrometry data acquisition using data-independent acquisition (DIA) is to collect all fragment ions after protein digestion into peptides in the sample. Library construction enables short gradients and rapid analysis, ensuring high throughput and high reproducibility, making it more suitable for samples with a large number of fractions. In contrast, data-dependent acquisition (DDA) methods selectively scan highly abundant peptides, resulting in poor reproducibility and making them unsuitable for this method.
[0054] The method for collecting mass spectrometry data using data-independent acquisition (DIA) can be a conventional process. For example, the following operation can be performed: the mass spectrometer scanning range is divided into a plurality of windows, and all ions in each window are scanned, thereby obtaining all ion information in the sample without bias.
[0055] The DIA mass spectrometric data of different fractions were acquired by performing mass spectrometric data acquisition in a data-independent acquisition (DIA) manner, which constituted characteristic tandem mass spectra of proteins in different fractions.
[0056] Step 5)
[0057] In step 5), protein library analysis is performed on the DIA mass spectrometry data of the different fractions collected in step 4) to obtain a protein quantitative matrix.
[0058] Here, protein library search analysis refers to comparing the obtained mass spectra with the theoretical spectra in the protein database by using an automatic data alignment program, first identifying the peptide sequences in the DIA mass spectrometry data, and then inferring the protein identification and quantification matrix based on the peptide information.
[0059] Here, the automatic data comparison program may adopt any suitable database search software, such as DIA-NN, Spectronaut, and OpenSWATH.
[0060] The analysis can be performed using DIA-NN or OpenSWATH software to perform spectrum search and then quantify the protein.
[0061] The protein database used may be, for example, a pre-built spectral library, or a library simulated from a fasta file, for example, a spectral library may be used.
[0062] In some embodiments, proteins with a high number of missing values need to be removed. For example, proteins with a missing rate greater than or equal to 7 / 9 can be removed from the protein quantification matrix to obtain a protein quantification matrix with a low missing rate. In addition, the data baseline can be controlled (optional step). For example, a protein with a maximum abundance of no more than 100 is considered noise (this protein is not included in subsequent analysis).
[0063] Here, the protein quantitative matrix refers to the expression information of each protein in different fractions, with the protein name as the horizontal direction and the fraction sample name as the vertical direction.
[0064] Step 6)
[0065] In step 6), the abundance distribution of protein molecules in the biological tissue sample in different fractions is obtained through mass spectrometry identification and quantification results.
[0066] Based on the results of step 5), the abundance distribution of protein molecules in the biological tissue sample in different fractions (as above, 90) is obtained.
[0067] Step 7)
[0068] In step 7), the fractions are grouped, and a correlation analysis is performed based on the abundance distribution of the target protein pairs whose intermolecular forces are to be analyzed in each group of fractions to calculate the correlation coefficients of the target protein pairs corresponding to each group of fractions.
[0069] The correlation analysis of abundance distribution can adopt the Pearson correlation analysis method, the Spearman correlation analysis method, etc., but is not limited thereto.
[0070] For example, the correlation analysis of abundance distribution can be performed as follows: the abundance information of two proteins is divided into several windows (parts), such as 9 fractions as one window (part), but the present invention is not limited thereto.
[0071] Specifically, for each pair of protein molecules, the abundance of the protein molecules in 90 fractions is divided into 10 windows (parts), and the specific allocation is as follows: window 1 is fraction 1 to fraction 9, window 2 is fraction 10 to fraction 18, window 3 is fraction 19 to fraction 27, window 4 is fraction 28 to fraction 36, window 5 is fraction 37 to fraction 45, window 6 is fraction 46 to fraction 54, window 7 is fraction 55 to fraction 63, window 8 is fraction 64 to fraction 72, window 9 is fraction 73 to fraction 81, and window 10 is fraction 82 to fraction 90. Each window corresponds to 9 fractions.
[0072] Then, within each window, the correlation of the target protein molecule pair corresponding to the window is calculated using the information on the abundance distribution of the protein molecules in the target protein molecule pair in each fraction within the window. Specifically, the Pearson correlation analysis method can be used to calculate the correlation coefficient of the protein molecule pair to represent the correlation of the target protein molecule pair. Alternatively, any other applicable correlation analysis method such as Spearman rank correlation analysis can be considered to calculate the correlation of the target protein molecule pair, and this application does not limit this. In addition, the specific calculation formulas of the Pearson correlation analysis method and the Spearman rank correlation analysis method are prior art and can be calculated in a manner known to those skilled in the art and will not be described in detail here.
[0073] Step 8)
[0074] Based on the correlation coefficients of the target protein pairs corresponding to each group of fractions, it is identified whether there are interaction peaks between the target protein pairs.
[0075] Identifying molecular interaction peaks can be performed by statistically analyzing the correlation coefficients of each pair of molecules to determine whether windows with high correlation coefficients appear consecutively. If there are only two consecutive windows with correlations greater than 0.5, and at least one window with correlations greater than 0.7, then the pair is considered to have an interaction peak. This approach can be used to analyze multiple protein pairs, for example, to determine whether each pair of protein molecules has an interaction peak.
[0076] Step 9)
[0077] In the case of interaction peaks between the target protein pairs, the protein quantitative matrix at each pressure is used to calculate the peak area of the corresponding interaction peak, and the affinity coefficient AF of the protein pair is calculated using the formula.
[0078] Based on the molecular interaction peaks identified in step 8), if the correlation coefficient of the second and third windows is greater than 0.7, the first window is taken into account to calculate the peak area and peak top position of the interaction peak within the window, where the peak top position is used to match the interaction peaks under different pressures. Considering that a certain error range is allowed, in some embodiments, ±2 fractions can be regarded as the same peak.
[0079] The peak area of the interaction peak can be calculated as follows: for each molecule in each pair of molecules, the peak area in the window corresponding to the interaction peak identified by the correlation coefficient is calculated. The peak area of the interaction peak is then averaged between the two corresponding peak areas. In this way, the peak areas of the interaction peaks for many pairs of molecules (i.e., the peak area of the protein complex) are obtained.
[0080] Then, the interacting molecular pairs whose peak area of the interaction peak decreases with increasing pressure are screened out, and the affinity coefficient AF between protein molecules is evaluated by the following formula, wherein the AF is used to characterize the change in the peak area of the interaction peak between protein molecular pairs with increasing pressure, and the larger the AF, the stronger the affinity between the target protein molecular pairs. Conversely, the smaller the AF, indicating that the peak area of the interaction peak of the protein molecular pair decreases significantly with increasing pressure, that is, its affinity is weaker.
[0081]
[0082] in
[0083] In the above formula, A1 represents the peak area of the interaction peak at the first pressure, A n-1 represents the peak area of the interaction peak at the n-1th pressure, A n is the peak area of the interaction peak at the nth pressure, A n-j-1 represents the peak area of the interaction peak at the nj-1th pressure, A n-j represents the peak area of the interaction peak at the njth pressure;
[0084] P1, P n and Δp have the same meanings as above;
[0085] k j For the adjacent pressure interval P n-j-1 and P n-j The weight of the interaction peaks in adjacent pressure intervals is represented by the degree of change in peak area A. n-j-1 -A n-j The degree of influence on the affinity coefficient AF, where k is the pressure interval coefficient and is a constant greater than 1.
[0086] In some embodiments, the value of k can be set to any value between 1 and 2. In addition, the constant k can be set in association with the number of pressure intervals and the difference Δp between two adjacent pressures, so that when the product of the number of pressure intervals and Δp is constant, that is, when the total pressure difference is constant, the smaller Δp is, the larger the value of k is. As an example only, in some cases, the value of k can be set to 1.2. While keeping the product of the number of pressure intervals and Δp unchanged, if more pressure intervals are set, that is, the Δp between the corresponding pressures becomes smaller, in this case, in order to reflect the effect of the unit pressure difference on the affinity coefficient AF, the value of k can be adjusted accordingly. Its specific value can be set based on the results of a series of experiments with the span of the pressure interval and the division method of the pressure interval as variables, and this application does not limit this. Setting the value of k according to the above method can more reasonably reflect the influence of pressure on the interaction force between protein molecule pairs. In addition, by setting the number of mutually adapted pressure intervals, the difference Δp between two adjacent pressures, and the pressure interval coefficient k in the experiment, the interaction force between the target protein molecule pairs can be more reliably determined. When calculating the AF between multiple pairs of protein molecules at the same time, the difference in interaction force between different pairs of protein molecules can be more significantly identified.
[0087] In one embodiment, when three pressures P1 to P3 are selected in step 1), the affinity AF between protein molecules can be evaluated by the following formula:
[0088]
[0089] Wherein, A1 represents the peak area of the interaction peak at the first pressure, A2 represents the peak area of the interaction peak at the second pressure, A3 represents the peak area of the interaction peak at the third pressure, P1 represents the first pressure, P3 represents the third pressure, wherein the first pressure (P1) < the second pressure (P2) < the third pressure (P3), and the adjacent pressure difference is a constant value;
[0090] Δp represents the adjacent pressure difference, k 1 The weight coefficient representing the pressure range from the second pressure to the third pressure may be 1.2.
[0091] In another embodiment, when four pressures P1 to P4 are selected in step 1), the affinity coefficient AF between protein molecules can be estimated by the following formula:
[0092]
[0093] The affinity coefficient AF calculated by the above formula in the present application characterizes the rate at which the abundance of an interacting molecule pair (protein complex) changes with pressure, thereby characterizing the interaction force between protein molecules.
[0094] Those skilled in the art will understand that, in the method of the present application, the more pressure conditions are set, the more accurate the affinity assessment will be.
[0095] Beneficial effects
[0096] 1. The present invention evaluates the affinity between protein molecules by comprehensively scanning the protein expression intensity of protein complexes under different pressures, called Pressure-based Molecular Affinity Full Scan (P-MAFS), which provides a new technology for measuring the affinity between molecules.
[0097] 2. The method of the present application uses a pressure cycler to adjust the pressure in order to provide a stable and controllable pressure environment of different pressures for the protein complex, thereby processing the sample more evenly and fully.
[0098] 3. Combining PCT, HPLC, and DIA mass spectrometry methods can detect the affinity between protein molecules or drug molecules in mixed sample types, and comprehensively scan all intermolecular interactions at one time without the need for sample purification.
[0099] 4. Through computational analysis model methods, the affinity between multiple molecules can be quantitatively analyzed simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0100] Figure 1 A schematic flow chart showing the method for detecting pressure-based intermolecular affinity according to the present application.
[0101] Figure 2 The figure shows the HPLC chromatograms of fresh mouse liver tissue lysate separated by SEC column under different pressures.
[0102] Figure 3 A distribution diagram showing the affinity between 2384 pairs of protein molecules.
[0103] Figure 4 The relative abundance of the interacting molecular pair (Q922E4_Q99JY9) at different pressures (15 kpsi, 25 kpsi, 35 kpsi) is shown when the intermolecular affinity is 20.
[0104] Figure 5The relative abundance of the interacting molecule pair (Q8R086_Q9ET01) at different pressures (15 kpsi, 25 kpsi, 35 kpsi) is shown when the intermolecular affinity is 2.
[0105] Figure 6 The relative abundance of the interacting molecule pair (P50396_P60843) at different pressures (15 kpsi, 25 kpsi, 35 kpsi) is shown when the intermolecular affinity is 0.2. DETAILED DESCRIPTION
[0106] The technical solutions of the present invention are described in detail below through specific embodiments to enable those skilled in the art to better understand the present invention. However, these embodiments are not intended to limit the scope of the present invention.
[0107] the term:
[0108] P-MAFS: Pressure-based molecular affinity full scan
[0109] PPI: protein-protein interaction
[0110] DIA: Data Independent Acquisition
[0111] instrument:
[0112] Pressure cycler (PCT): Barocycler 2320EXT, Pressure BioSciences Inc.
[0113] SEC column: Biosep SEC-s4000, Agel-Phenome
[0114] Reagents:
[0115] 20% NP40: NP-40 diluted with ultrapure water.
[0116] 500 mM HEPES (pH 7.5): Dissolve 5.96 g HEPES in ultrapure water to a volume of 50 ml.
[0117] 1500mM NaCl: Dissolve 4.38g NaCl in ultrapure water to a volume of 50ml.
[0118] 200mM NaF: Dissolve 419.9mg NaF in ultrapure water to a volume of 50ml.
[0119] 200mM Na3VO4: 36.78mg NaF was dissolved in ultrapure water to a volume of 1ml.
[0120] 100mM PMSF: Dissolve 17.42mg NaF in ultrapure water to a volume of 1ml.
[0121] 50× protease inhibitor cocktail: Dissolve 1 tablet of protease inhibitor cocktail (Roche) in ultrapure water to a volume of 1 ml.
[0122] Lysis buffer was prepared from the above reagents. Specifically, 1 ml of lysis buffer contained: 25 μl 20% NP-40, 100 μl 500 mM HEPES, 100 ul 1500 mM NaCl, 250 μl 200 mM NaF, 1 μl 200 mM Na3VO4, 10 μl 100 mM PMSF, 20 μl 50× protease inhibitor cocktail, and 494 μl ultrapure water.
[0123] In this application, if certain operations are not described in detail, these operations and their operating conditions are conventional in the art or well known to those skilled in the art.
[0124] Hereinafter, mouse liver is used as a biological tissue sample to describe in detail the method for detecting the affinity between protein molecules based on pressure according to the present invention.
[0125] 1) Place 4 mg of mouse liver in a pressure cycler (PCT) tube and add 50 μl of lysis buffer, which consists of 0.5% NP40 by volume, 50 mM HEPES (pH 7.5), 150 mM NaCl, 50 mM NaF, 200 μM Na3VO4, 1 mM PMSF, and 1× protease inhibitor cocktail.
[0126] The PCT pressure parameters (temperature: 4°C; number of cycles: 10; on time: 30 s; off time: 10 s; pressure values: first pressure: 15 kpsi, second pressure: 25 kpsi, third pressure: 35 kpsi) were set for lysis to obtain lysates L1, L2 and L3, respectively.
[0127] 2) Centrifuge lysates L1, L2, and L3 at >18,000 rcf for 10 min, and collect the supernatants. Perform high-performance liquid separation on the supernatants of L1, L2, and L3, monitor protein concentration using UV absorbance at 280 nm, and collect fractions (using an SEC column with a column temperature of 4°C, a flow rate of 0.5 ml / min for phase A, and fractions collected every 12 seconds. Fractions were stored on ice; phase A: 50 mM HEPES (pH 7.5), 150 mM NaCl).
[0128] Figure 2The figure shows the HPLC chromatograms of fresh mouse liver tissue lysates separated by SEC column under different pressures. Figure 2 It can be seen that as the pressure increases, the large molecular peak (from left to right, the molecular weight gradually decreases) gradually decreases, and the small molecular peak gradually increases.
[0129] 3) Perform conventional trypsin digestion steps on different fractions. Specifically, each fraction was subjected to conventional reductive alkylation operation, and then 400 ng of trypsin was added to each fraction and digested at 37°C overnight.
[0130] 4) Mass spectrometry data were collected for the different fractions after proteolysis in the previous step using data-independent acquisition (DIA) mode.
[0131] 5) The mass spectrometry data of different fractions were analyzed using DIA-NN to obtain a protein quantification matrix (using the default parameters of the DIA-NN software).
[0132] 6) Through mass spectrometry identification and quantification results, the abundance distribution of protein molecules in different fractions of biological tissue samples is obtained.
[0133] 7) Proteins with many missing values (missing rate greater than or equal to 7 / 9) were removed from the protein quantitative matrix obtained in step 6). Proteins with a peak height maximum value less than 100 were considered as noise (the protein was deleted). The 90 fractions were divided into 10 windows, and the Spearman correlation was calculated for each window.
[0134] 8) Based on the correlation coefficients of the target protein molecule pairs corresponding to each group of fractions, identify whether there is an interaction peak between the target protein molecule pairs.
[0135] The correlation coefficients of each pair of molecules were statistically analyzed to see whether windows with high correlation coefficients appeared continuously. If there were only two consecutive windows with correlations greater than 0.5, and at least one of them was greater than 0.7, the interaction peak was considered to be found.
[0136] 9) The protein quantitative matrix under each pressure is used to calculate the peak area of the corresponding interaction peak and the affinity coefficient AF of the protein molecule pair is calculated by the formula.
[0137] The peak area of the interaction peak can be calculated as follows: for each pair of molecules, the peak area of the window corresponding to the interaction peak identified in the portion with the highest correlation coefficient is calculated. The peak area of the interaction peak is then averaged over the two corresponding peak areas. In this way, the peak areas of the interaction peaks for many pairs of molecules (i.e., the peak area of the protein complex) are obtained.
[0138] Then, the interacting molecular pairs whose peak area decreases with increasing pressure were screened out, and the affinity coefficient AF between the protein molecules was evaluated using the following formula:
[0139]
[0140] In the above formula, A1 represents the peak area of the interaction peak at the first pressure, A2 represents the peak area of the interaction peak at the second pressure, and A3 represents the peak area of the interaction peak at the third pressure.
[0141] From step 9), we get the affinity of 2384 pairs of protein molecules (AF) (see Figure 3 ).from Figure 3 It can be seen that a total of 2,384 protein-molecular interactions were scanned in this experiment, and their relative affinities were mainly concentrated between 0 and 5.
[0142] Figure 4-6 The relative abundances of different interacting molecular pairs under different pressures are shown respectively when the intermolecular affinities are different. Figure 4-6 , it can be seen that the peak height of the molecular peak of two different molecules in the interaction state decreases with increasing pressure, and the faster it decreases (for example Figure 6 , it decreases significantly at 25kpsi), and the smaller the relative intermolecular affinity is, the slower it decreases (e.g. Figure 4 , only a small degree of decrease at 25kpsi), the greater the relative intermolecular affinity is evaluated.
Claims
1. A method for detecting affinity between protein molecules based on pressure, comprising the following steps: 1) Place the biological tissue sample in P 1 to P n The pyrolysis was carried out under a series of pressures to obtain a series of pyrolysis products, among which P 1 represents the first pressure, P n Indicates the nth pressure, where n is an integer greater than 3, and the difference between two adjacent pressures is a fixed value; 2) centrifuging a series of lysates obtained at various pressures in step 1), taking the supernatant, performing high performance liquid chromatography (HPLC) separation based on size exclusion chromatography (SEC), and collecting fractions; 3) The fractions collected in step 2) were proteolyzed into peptide fragments using trypsin; 4) performing mass spectrometry data acquisition on the different fractions after proteolysis in step 3) using data-independent acquisition (DIA) mode; 5) Perform protein library analysis on the DIA mass spectrometry data of different fractions collected in step 4) to obtain a protein quantitative matrix; 6) Infer the abundance distribution of protein molecules in different fractions of biological tissue samples through protein quantitative matrix; 7) Grouping the fractions to generate different windows, performing correlation analysis based on the abundance distribution of the target protein pairs in each group of fractions to analyze intermolecular forces, and calculating the correlation coefficients of the target protein pairs corresponding to each group of fractions; 8) Based on the correlation coefficients of the target protein pairs corresponding to each group of fractions, identify whether there are interaction peaks between the target protein pairs; 9) In the case of interaction peaks between the target protein pairs, the peak area of the corresponding interaction peak is calculated using the protein quantitative matrix at each pressure, and the affinity coefficient AF of the protein pair is calculated using the following formula, wherein the larger the AF, the stronger the affinity between the target protein pairs: In the above formula, represents the peak area of the interaction peak at the first pressure, represents the peak area of the interaction peak at the n-1th pressure, is the peak area of the interaction peak at the nth pressure, represents the peak area of the interaction peak at the nj-1th pressure, represents the peak area of the interaction peak at the njth pressure; Adjacent pressure intervals P n-j-1 and P n-j The weight of the interaction peaks in the corresponding adjacent pressure intervals indicates the degree of change in the peak area. - The influence degree of affinity coefficient AF is as follows: is the pressure interval coefficient, which is a constant greater than 1.
2. The method according to claim 1, wherein In step 1), the biological tissue sample is derived from: animals or plants, and / or In step 1), use a pressure cycler to pressurize the biological tissue sample, and select more than 3 pressures as P 1 to P n , and / or In step 1), the lysis buffer used includes 0.3-0.8 vol% NP-40, 35-60 mM HEPES pH 7.5, 130-160 mM NaCl, 35-60 mM NaF, 190-215 μM Na3VO4, 0.7-1.2 mM PMSF and 1× protease inhibitor cocktail.
3. The method according to claim 2, wherein: The animal biological tissue samples are derived from: Animal organs.
4. The method according to claim 2, wherein: The plant biological tissue samples are derived from roots, stems, leaves, flowers and fruits.
5. The method according to claim 3, wherein The animal biological tissue sample comes from: mouse liver.
6. The method according to claim 1, wherein In step 2), centrifuge at >18,000 rcf for 5-15 minutes.
7. The method according to claim 1, wherein In step 3), trypsin is used for proteolysis, or, based on trypsinization, intracellular protease Lys-C is used for enzymatic cleavage.
8. The method according to claim 1, wherein In step 5), peptide sequences are first identified in the DIA mass spectrometry data, and then the protein identification and quantification matrix is inferred based on the peptide information.
9. The method according to claim 1, wherein In step 5), use DIA-NN, Spectronaut, or OpenSWATH to search the database.
10. The method according to claim 1, wherein In step 7), the Pearson correlation analysis method or the Spearman correlation analysis method is used to perform correlation analysis on the abundance distribution.
11. The method according to claim 1, wherein In step 8), the molecular interaction peaks are identified by performing the following operations: the correlation coefficients of each pair of molecules are statistically analyzed. When the Spearman correlation of two consecutive windows is greater than 0.5, and the Spearman correlation of at least one window is greater than 0.7, the pair of molecules is considered to have an interaction peak.
12. The method according to claim 1, wherein In step 9), the pressure interval coefficient k is related to the number of pressure intervals and the difference between two adjacent pressures. The number of pressure intervals is set in relation to When the product of is constant, The smaller it is, the larger the value of k is.
13. The method according to claim 1, wherein In step 1), three pressures P1 to P3 are selected, and in step 9), the affinity coefficient AF between protein molecules is estimated by the following formula: in, represents the peak area of the interaction peak at the first pressure, is the peak area of the interaction peak at the second pressure, is the peak area of the interaction peak at the third pressure, P 1 represents the first pressure, P 3 represents the third pressure, where the first pressure ( P 1)<Second pressure( P 2)<3rd pressure( P 3), and the adjacent pressure difference is a constant value; Indicates the adjacent pressure difference, The weight coefficient representing the pressure range from the second pressure to the third pressure is set to 1.2.
Citation Information
Patent Citations
Novel method for obtaining high-credibility phosphorylation site occupancy rate in paired samples on large scale
CN113884583A
Methods for systematic identification of protein - protein interactions
US20020102741A1