Mass spectrometry method for large-scale discovery of D-type proteins
Through the conformational resolution of proteinomics mass spectrometry method, five-dimensional protein data were collected and combined with signal amplification strategy, the problem of D-type protein recognition in massive data was solved, and fast and accurate protein stereoisomeric modification screening was achieved, improving the reliability and operation efficiency of recognition.
Patent Information
- Application Number
- CN202410662451.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-05-27
AI Technical Summary
The prior art is difficult to quickly and accurately identify and screen D-type proteins, especially in massive omics data. The identification specificity and sensitivity of traditional methods are low, and it is impossible to effectively identify protein stereoisomerical modifications on a large scale.
Through the conformational resolution proteomics mass spectrometry method, five-dimensional omics data were collected, combined with signal amplification strategy and Gaussian fitting algorithm, stereoisomerized polypeptides were screened, and the isomerase was used to verify it, and D-type protein data was finally generated.
It realizes the rapid and accurate identification of protein stereoisomerial modifications from massive omics data, improves the reliability and operational efficiency of recognition, simplifies the data processing process, and has universality and visualization capabilities.
Smart Images

Figure CN118711659B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of proteomics, computational biology, mass spectrometry chemistry, and clinical medicine. It involves a high-throughput and large-scale screening method for D-type proteins from a vast amount of proteomics data through advanced mass spectrometry experimental methods and computer data processing algorithms. Specifically, it relates to a mass spectrometry method for large-scale discovery of D-type proteins. Background Art
[0002] The central dogma tells us that the vast majority of proteins in nature have an obvious chiral preference, and proteins synthesized from L-amino acids as raw materials account for the vast majority of humans. However, under certain disease states or special pathological environmental conditions, there are also a small number of proteins composed of D-amino acids. This type of D-type protein is closely related to a variety of major human diseases. Currently, a variety of bioenzymes that regulate protein chirality have been reported. Therefore, it is considered that the generation of D-type proteins mainly comes from post-translational modification of proteins, that is, protein stereoisomerism. Protein stereoisomerism refers to a new type of post-translational modification in which the backbone amino acids of a protein undergo a stereoisomeric inversion, suddenly changing from the traditional L-configuration to the D-configuration or other isomeric forms without changing its molecular composition. Stereoisomeric modification is a naturally occurring low-abundance post-translational modification related to neurological diseases. Similar to other common protein post-translational modifications, protein stereoisomerism regulates the structure and function of a series of important proteins. However, previous studies on stereoisomeric modification have mostly focused on the amino acid and polypeptide levels, and there is still relatively little research on stereoisomeric modification at the protein level. The main reason is that since protein stereoisomeric modification does not change the overall amino acid sequence and molecular weight, and the traditional immunohistochemistry has low recognition specificity and sensitivity for isomeric proteins, its large-scale identification is still full of challenges. Therefore, there is an urgent need for a technology to achieve rapid identification and large-scale screening of D-type proteins to help discover potential new disease-associated markers and precision drug targets.
[0003] Developing a large-scale identification algorithm for protein stereoisomeric modification can quickly and effectively extract D-type protein information from a vast amount of omics data sets, and assist in identifying potential isomerization modification domains and their sites. By developing such an experimental process and data processing algorithm for large-scale D-type protein identification, it will help to deeply and comprehensively analyze post-translational modifications with unchanged molecular weight, and provide new ideas for aspects such as early precise diagnosis, molecular typing, drug target discovery, and pathological mechanism analysis of major human diseases. Summary of the Invention
[0004] An object of the present invention is to solve at least the above problems and / or deficiencies, and provide at least the advantages described later.
[0005] Another object of the present invention is to provide a mass spectrometry method for large-scale discovery of D-type proteins, which can be used to screen clinical markers of D-type proteins associated with major human diseases and solve the technical difficulties of manually screening and identifying protein stereoisomeric modifications in massive data.
[0006] To this end, the technical solution provided by the present invention is as follows:
[0007] A mass spectrometry method for large-scale discovery of D-type proteins, comprising:
[0008] Obtaining conformation-resolved proteomic data of a biological sample, including retention time information, tandem mass spectrometry information, collision cross-sectional area, and stereoisomer signal response intensity of stereoisomeric polypeptides with the same amino acid sequence;
[0009] Performing feature extraction on the amino acid sequences in the conformation-resolved proteomic data through a dataset generated by searching a protein database to obtain multiple feature data of a protein sample, so as to establish a feature set for isomeric modification retrieval of the protein sample. Each piece of feature data includes protein name, amino acid sequence, chromatographic retention time, charge, signal response intensity, mass-to-charge ratio, and secondary mass spectrometry map;
[0010] For the feature data under the same amino acid sequence, deleting the data with a retention time difference less than 2 minutes and / or a collision cross-sectional area difference less than 1.5% to screen out proteins or polypeptides with stereoisomers, and obtaining a preliminary screening result by calculating the isomerization ratio and quantity of each protein;
[0011] Performing isomeric verification on the preliminary screening result to generate final D-type protein data.
[0012] Preferably, in the mass spectrometry method for large-scale discovery of D-type proteins, the method for obtaining the conformation-resolved proteomic data is:
[0013] Establishing a separation method applicable to chromatographic-level isomeric polypeptides, and identifying stereoisomeric polypeptide information through a signal amplification strategy algorithm and a Gaussian fitting algorithm;
[0014] Performing a conformation-resolved proteomic experiment on the biological sample to collect five-dimensional omics data including the retention time information, tandem mass spectrometry information, collision cross-sectional area, and the distribution of stereoisomer signal response intensity of the stereoisomeric polypeptides based on reversed-phase liquid chromatography under the same amino acid sequence.
[0015] Preferably, the mass spectrometry method for large-scale discovery of D-type proteins further includes:
[0016] Based on the behavioral characteristics of proteomic analysis of heterogeneous polypeptide standards, screening conditions for retrieving polypeptide heterogeneous modifications at the proteomic level are established. The screening conditions include the threshold of the retention time difference and / or the threshold of the collision cross-sectional area difference between the characteristic data under the same amino acid sequence.
[0017] Preferably, in the mass spectrometry method for large-scale discovery of D-type proteins, for the initial screening results, for the chromatographic peak signals obtained from the corresponding original data, the chromatographic peak signals of the original data are processed by a signal amplification strategy algorithm and a Gaussian fitting algorithm, and the original data chromatographic peak signals are iteratively processed to verify the quality and quantity of the heterogeneous signal peaks under the same amino acid sequence.
[0018] Using means such as heterogeneous signal amplification algorithms and isomerases, combined with biological experiments to verify the D-type mutations on the initial screening results, the final D-type protein data is generated.
[0019] Preferably, in the mass spectrometry method for large-scale discovery of D-type proteins, a five-dimensional proteomic dataset is integrated to efficiently separate and identify polypeptide isomers from the perspectives of amplification of collision cross-sectional area differences and chiral isomer differences.
[0020] Preferably, in the mass spectrometry method for large-scale discovery of D-type proteins, the protein isomerization modification data includes protein name, amino acid sequence, charge, mass-to-charge ratio, retention time, collision cross-sectional area, response intensity, and isomerization ratio.
[0021] Preferably, in the mass spectrometry method for large-scale discovery of D-type proteins, both the characteristic set of the protein sample heterogeneous retrieval modification and the output protein isomerization modification data are displayed in the form of rows and columns.
[0022] Preferably, in the mass spectrometry method for large-scale discovery of D-type proteins, when calculating the relative ratio of the target peaks under the characteristic data of the same amino acid sequence, false positive noise data is excluded.
[0023] Optionally, first group the feature set retrieved from the isomeric modification of the protein sample according to the amino acid sequence and charge. For the feature data in the same group, according to the signal response intensity of their different charges, delete the charge information with a relatively low signal response intensity value (lower than 1%, where the relative response intensity refers to the absolute intensity of each isomer under the same sequence divided by the sum of all intensities under that sequence (the sum of all isomers is 1)), so that each amino acid sequence of the feature data has a unique main charge peak to obtain its target peak; for the feature data in the same group, delete the feature data with an isotope mass interpolation of 1 or 2 and a relatively small signal response intensity (lower than 1%), and delete the feature data with a relative retention time value ΔRT% < 0.5% and a relatively small signal response intensity (lower than 1%). Calculate the retention time distribution between adjacent target peaks of the feature data in the same group, and screen out the feature data with a retention time difference of less than 2 minutes to screen out proteins or polypeptides with isomeric modifications; calculate the relative proportion of the target peaks under the feature data of the same amino acid sequence to obtain the isomerization ratio and the number of isomers of each protein. It should be noted that each amino acid sequence in the feature set retrieved from the isomeric modification of the protein sample has only a unique mass-to-charge ratio value.
[0024] Optionally, the method for feature extraction from the omics mass spectrometry data of a protein sample includes the following steps:
[0025] Obtain the proteomics raw data for protein database retrieval,
[0026] Import the proteomics raw data into the MaxQuant database for database retrieval analysis to obtain the data for isomeric retrieval,
[0027] Perform feature extraction on the data for isomeric retrieval to obtain multiple pieces of feature data of the protein sample.
[0028] An electronic device includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any one of the methods.
[0029] A storage medium stores a computer program, and when the program is executed by a processor, it implements any one of the methods.
[0030] The present invention has at least the following beneficial effects:
[0031] The present invention provides a high-performance mass spectrometry experimental method and a large-scale data processing algorithm for quickly screening and identifying stereoisomeric proteins from a large amount of omics data for the first time, including the following beneficial effects:
[0032] 1) High efficiency and speed: Through this algorithm, large-scale omics data can be processed quickly, and proteins with potential isomeric modifications can be accurately screened out, saving time and labor costs.
[0033] 2) Accuracy: The algorithm design combines feature extraction and optimization steps, and can accurately identify protein isomeric modifications, improving the reliability and accuracy of identification.
[0034] 3) Universality: The algorithm has been verified and applied with actual samples, and can be applied to different types of omics sample data, with certain universality and scope of application.
[0035] 4) Simple and efficient: Compared with traditional methods, this algorithm reduces the complexity of the preprocessing process, simplifies the data post-processing flow, and improves the operation efficiency of researchers.
[0036] 5) Visualized results: Through data visualization methods, the screened target proteins and isomeric modification information can be intuitively displayed, facilitating researchers to interpret and analyze the identification results.
[0037] Other advantages, objectives, and features of the present invention will be partially reflected by the following description, and partially will also be understood by those skilled in the art through the research and practice of the present invention. Description of the Drawings
[0038] Figure 1 This is the general workflow of the stereoisomer screening algorithm at the proteome level in the embodiments of the present invention.
[0039] Figure 2 This is the analysis chart of the retention time difference of stereoisomeric polypeptides at the standard level (A) and the multi-dimensional stereoisomer distribution chart based on m / z, RT, and CCS (B) in the embodiments of the present invention.
[0040] Figure 3 This is the large-scale stereoisomeric protein identification chart on the N2a cell line through mass spectrometry experiments and data processing screening algorithms in the embodiments of the present invention. 3A, 3E) Distribution of the proportion and number of stereoisomeric polypeptides in the control group; 3B, 3F) Distribution of the proportion and number of stereoisomeric polypeptides in the low-dose group; 3C, 3G) Distribution of the proportion of stereoisomeric polypeptides in the medium-dose group and the number of isomers; 3D, 3H) Distribution of the proportion and number of stereoisomeric polypeptides in the high-dose group;
[0041] Figure 4 This is the representative chromatogram (A), first-order mass spectrum (B), and second-order mass spectrum (C-D) of the endogenous stereoisomeric polypeptide IGGIGTVPVGR in N2a cells in the embodiments of the present invention.
[0042] Figure 5This is a diagram showing the distribution of stereoisomeric proteins and related pathways in the N2a cell line in the embodiments of the present invention. A) The number of N2a proteins identified at the traditional omics level; B) The number of stereoisomeric proteins in each group after algorithm screening; C) Hierarchical clustering of 166 quantified stereoisomeric proteins. Based on four different intensity distributions, they are divided into two clusters. D) Biological processes and KEGG pathways enriched by the quantified proteins in Cluster 1, sorted according to the P value. E) Biological processes and KEGG pathways enriched by the quantified proteins in Cluster 2, sorted according to the P value. Detailed implementation manners
[0043] The following provides a further detailed description of the present invention so that those skilled in the art can implement it with reference to the text of the specification.
[0044] The present invention is described below with reference to specific embodiments. Those skilled in the art can understand that the experimental methods in the following embodiments are all conventional methods unless otherwise specified; the raw materials, reagents, materials, etc. used in the following embodiments are all commercially purchased products unless otherwise specified.
[0045] The object of the present invention is to provide a large-scale algorithm and workflow for quickly screening and identifying protein stereoisomeric modifications and D-type protein variants from massive proteomics data. The molecular weight changes caused by traditional post-translational modifications can be directly identified by omics mass spectrometry, while isomeric modifications cannot be directly identified by traditional omics mass spectrometry techniques because they do not change the molecular weight. The present invention aims to develop an isomeric algorithm to achieve large-scale stereoisomeric retrieval of proteomics data and screen and extract information on protein variants with stereoisomeric modifications from massive data.
[0046] The specific design ideas are as follows:
[0047] 1) Conduct conformational resolution proteomics mass spectrometry experiments on complex biological samples to collect five-dimensional proteomics datasets;
[0048] 2) Use a protein database for sequence retrieval and protein identification to obtain characteristic information such as the sequence, abundance, retention time, and biological function of the identified proteins;
[0049] 3) Develop a chromatographic separation method suitable for a large number of isomeric polypeptides and a high-throughput data processing algorithm for signal recognition;
[0050] 4) Optimization of stereoisomeric modification conditions: By optimizing stereoisomeric standards, obtain conditions that can be used for stereoisomeric screening algorithms for large-scale omics samples;
[0051] 5) Extraction of stereoisomeric modification features: By extracting features from the omics mass spectrometry data of protein samples, such as amino acid sequences, chromatographic retention times, ion mobility values, secondary mass spectrometry spectra, and signal intensities, etc., a feature set capable of accurately identifying isomeric modifications is established;
[0052] 6) Algorithm design and optimization: Design a specialized algorithm based on these features that can efficiently and quickly screen out target proteins that may have isomeric modifications from a large amount of omics datasets;
[0053] 7) Model training and verification: Use known isomeric modification samples to verify the accuracy and universality of the algorithm;
[0054] 8) Big data processing and application: Develop an effective data processing and analysis process to adapt to the rapid processing and extraction of a large amount of omics data, and ensure the large-scale screening of isomeric proteins;
[0055] 9) Application to actual samples: Apply the algorithm to real omics sample datasets, test its identification effect, and verify the feasibility and accuracy of the algorithm in actual samples;
[0056] 10) Data presentation and interpretation: Design an intuitive and easy-to-understand data visualization method to display the screened target proteins and possible isomeric modification information.
[0057] Based on the mass spectrometry experimental process for large-scale discovery of protein stereoisomeric modifications and the isomeric data retrieval algorithm, the present invention establishes a mass spectrometry method for large-scale discovery of protein stereoisomeric modification variants, which has the following differences compared with the existing post-translational modification mass spectrometry techniques:
[0058] 1) This method is the first mass spectrometry method for large-scale screening and identification of stereoisomeric proteins from a large amount of proteomics data.
[0059] 2) This method integrates for the first time the stereoisomeric separation technology driven by a signal amplification algorithm and a data search algorithm for screening potential stereoisomeric proteins from a large amount of proteomics data.
[0060] 3) This method establishes for the first time the molecular map of stereoisomeric proteomics of the N2a cell line.
[0061] 4) Standard sample verification: By using existing standard samples to explore and set conditions and combining a large amount of omics data for verification, the present invention can more accurately screen and identify protein stereoisomeric modifications, reduce the error rate, and improve the result credibility.
[0062] 5) Application of omics data: The present invention makes full use of a large amount of omics sample data, and through in-depth mining and analysis, realizes the large-scale discovery of protein isomeric modifications, providing richer and more comprehensive data support for related research.
[0063] 6) Algorithm efficiency: The algorithm of the present invention is optimized and designed for protein isomeric modification characteristics, with high processing efficiency and accuracy, and can quickly and accurately identify potential modification domains and sites, providing important references for subsequent biological research.
[0064] In summary, the present invention has obvious innovative advantages in aspects such as large-scale isomeric data processing and large-scale identification of protein stereoisomeric variants, breaks through the technical bottleneck that traditional mass spectrometry cannot identify post-translational modifications with unchanged molecular weight, can be applied to the research on the structure-activity relationship of protein stereoisomeric modifications in neurodegenerative diseases and tumor markers, discovers new protein variants associated with diseases, has important scientific research significance for understanding the molecular mechanisms of special protein structure subtypes associated with diseases and their dynamic change processes, and has considerable application prospects in aspects such as pathological mechanism analysis, early molecular typing, personalized prognosis diagnosis and treatment, and precise target discovery of human diseases.
[0065] The present invention provides a mass spectrometry method for large-scale discovery of D-type proteins, including:
[0066] Develop a separation method suitable for chromatographic-level isomeric polypeptides, and identify isomeric polypeptide information through signal amplification strategies (continuous wavelet transform) and Gaussian fitting);
[0067] Conduct conformational resolution proteomics experiments on complex biological samples, collect five-dimensional omics data including retention time information of isomeric polypeptides based on reversed-phase liquid chromatography, tandem mass spectrometry information, high-resolution collision cross-sectional area, and signal response intensity distribution of sequence-matched stereoisomers, so as to achieve precise quantification of stereoisomers;
[0068] Based on the proteomics analysis behavior characteristics of isomeric polypeptide standards, establish screening conditions for polypeptide isomeric modification retrieval at the omics level;
[0069] Extract features from the dataset generated by sequence database searching to obtain multiple feature data of protein samples, so as to establish a feature set for protein sample isomeric modification retrieval. Each feature data includes protein name, amino acid sequence, chromatographic retention time, charge, signal response intensity, mass-to-charge ratio, and secondary mass spectrum;
[0070] For the feature data of the same group, delete the data with a retention time difference less than 2 minutes and / or a collision cross-sectional area difference less than 1.5% to screen out proteins or polypeptides with isomeric modifications, and obtain a preliminary screening result by calculating the isomerization ratio and quantity of each protein;
[0071] For the primary screening results, further utilize the signal amplification algorithm to iteratively process the original data, check the quality and quantity of heterogeneous signal peaks, and combine biological experiments to verify D-type mutations, generating a final list of D-type proteins.
[0072] An efficient liquid-phase separation method for the large-scale discovery of D-type proteins, based on reversed-phase liquid separation, with certain differences in the retention times of stereoisomeric polypeptides. Through the innovative application of the signal amplification algorithms CWT and GA, the chromatographic-based separation and recognition capabilities of stereoisomers are maximized.
[0073] A high-resolution gas-phase separation method for the large-scale discovery of D-type proteins, integrating five-dimensional proteomic datasets, and efficiently separating and recognizing polypeptide isomers from the perspectives of differences in collision cross-sectional area and amplified chiral isomer differences.
[0074] A mass spectrometry method for the large-scale discovery of D-type proteins, where the protein isomerization modification data includes protein name, amino acid sequence, charge, mass-to-charge ratio, retention time, collision cross-sectional area, response intensity, and isomerization ratio.
[0075] The mass spectrometry method for the large-scale discovery of D-type proteins, a method for large-scale automatic feature extraction of omics mass spectrometry data of protein samples, specifically including the following steps:
[0076] Obtain the proteomic raw data for protein database retrieval; comprehensively reveal the differences of stereoisomeric polypeptide standards based on conformation-resolved mass spectrometry and obtain the screening conditions for the isomer retrieval algorithm; increase the differences between stereoisomeric polypeptides through the signal amplification algorithm to improve their resolution and separation capabilities; import the proteomic raw data into the MaxQuant database for database retrieval analysis to obtain the data for isomer retrieval; perform feature extraction on the data for isomer retrieval to obtain multiple feature data of the protein sample.
[0077] Preferably, in the mass spectrometry method for the large-scale discovery of D-type proteins, the isomer algorithm screening conditions are based on stereoisomeric standards and are significantly persuasive.
[0078] A mass spectrometry method for the large-scale discovery of D-type proteins, where both the feature set of protein sample isomer retrieval modification and the output protein isomerization modification data are presented in the form of rows and columns.
[0079] A mass spectrometry method for the large-scale discovery of D-type proteins, when calculating the relative ratio of target peaks under the feature data of the same amino acid sequence, false positive noise data is excluded.
[0080] The specific implementation plan is as follows:
[0081] 1) Establishment of a conformational resolution proteomics mass spectrometry method:
[0082] Step 1: Synthesize 40 polypeptide fragments with the most abundant post-translational modifications, including disease markers such as Alzheimer's and Parkinson's diseases. After mixing the above-mentioned polypeptide data in equal proportions, collect the corresponding omics data based on the LTQ Orbitrap Eclipse MS of Thermo Fisher.
[0083] Step 2: Collect the ion mobility data of the above standard peptides through Synapt XS of Waters and obtain the corresponding CCS values.
[0084] Step 3: Introduce signal amplification strategies including CWT and GA into the liquid chromatography separation level to improve the resolution and separation ability of stereoisomeric polypeptides. Gaussian fitting (GA) and continuous wavelet transform (CWT) are applied to chromatographic separation. Based on the chromatographic peak signals obtained from the original data, after being processed by GA and CWT, the chromatographic curves are fitted to improve the resolution and separation ability.
[0085] Step 4: At the standard level, obtain information such as liquid chromatography retention time, tandem mass spectrometry, collision cross-sectional area, and stereoisomer response intensity, optimize the conditions for large-scale omics isomer screening, and obtain screening conditions with a retention time difference of less than 2 minutes and / or a collision cross-sectional area difference of less than 1.5%.
[0086] 2) Preliminary extraction of proteomics data:
[0087] Step 5: Obtain omics raw data (suffixes such as.raw,.wiff,.mzML, etc.) from different biological samples (tissues, cells, urine, etc.). In this example, the current data used is the omics data collected based on the LTQ Orbitrap Eclipse MS of Thermo Fisher. The omics raw data can be obtained through the online website ProteomeXchange or by self-sampling. The goal of this step is to obtain the original file for protein database retrieval.
[0088] Step 6: Import the target raw data into MaxQuant for database search and analysis. Select the corresponding database based on the type of input raw data, set appropriate parameters to perform database search on the raw data, and a series of txt files will be obtained after the database search. The goal of this step is to obtain the data for isomer retrieval.
[0089] Step 7: Extract the file named evidence and filter out the contamination library and the reverse library. Convert the txt file into an excel spreadsheet with the file extension.xlsx. Divide it into different sheets according to the different numbers of sample collection needles for subsequent analysis, and use this excel spreadsheet (including parameters such as protein sequence, name, charge, m / z, Retention time, Intensity, Ratio_D, etc.) as the final input file for isomeric retrieval. The goal of this step is to obtain a data file that can be recognized by the algorithm.
[0090] 3) Large-scale screening of stereoisomeric proteins:
[0091] Step 8: According to the information in the input file, delete the rows where the response intensity (Intensity) column is 0 or has a null value. The goal of this step is to filter out invalid data.
[0092] Step 9: According to the specific experimental requirements, delete or not delete the rows where the modification column value is "modified". This step is to ensure that each sequence has the same mass-to-charge ratio.
[0093] Step 10: Group the remaining rows by the sequence and charge columns. According to the response intensity of different charges, delete the charge information with a low response value. The purpose of this step is to ensure that each sequence has a unique main charge peak.
[0094] Step 11: Calculate the retention time difference between adjacent peaks. Based on the retention time difference, delete the rows where the retention time difference is less than 2 minutes or the difference in collision cross-sectional area is less than 1.5%. The goal of this step is to screen the list of polypeptides with isomeric modifications according to the retention time.
[0095] Step 12: Calculate the relative ratio of the target peaks under the same sequence, and eliminate the false positive background fluctuation noise signals. Calculate the ratio value of the remaining target peaks to obtain the isomerization ratio and the number of isomers for each sequence. The goal of this step is to ensure the reliability of the screened data.
[0096] Step 13: Run the algorithm and output a simplified list of isomeric modifications including protein name (Proteins), sequence (Sequence), charge (Charge), mass-to-charge ratio (m / z), retention time (RT), collision cross-sectional area (CCS), response intensity (Intensity), isomerization ratio (Ratio, etc.), and a list containing all input information, and generate an excel file. The specific technical route and workflow are as Figure 1 shown.
[0097] As attached Figure 1 As described above, the specific implementation process of this application can be divided into three steps. The first is the establishment of the conformational resolution proteomics mass spectrometry method (Steps 1 to 4); the second is the preliminary extraction of proteomics data (Steps 5 to 7), and the third is the large-scale screening of stereoisomeric proteins (Steps 8 to 13). The following is an example to illustrate this application:
[0098] Example 1: Verification of the isomeric algorithm for protein isomers
[0099] (1) Select isomeric modification sites and target polypeptides
[0100] Based on previous literature research, aspartic acid and serine were used as isomeric modification sites to construct a series of stereoisomeric modified polypeptides including disease markers such as Alzheimer's disease (AD) and Parkinson's disease (PD), among which 40 stereoisomeric polypeptides including disease markers such as Aβ42, Tau, αSynuclein, Parkin, Pink, DJ, and P53 were included.
[0101] (2) C18-based chromatographic separation
[0102] Based on the synthesized stereoisomeric polypeptides, the relative retention times of the polypeptides were obtained by liquid chromatography-mass spectrometry (LC-MS) analysis. Due to the change in the spatial structure between stereoisomeric polypeptides, the interaction force between them and C18 changes, resulting in a change in the retention time. By calculating the retention time difference of each pair of polypeptides, the differences between stereoisomers were identified.
[0103] (3) Liquid-phase separation screening and identification of stereoisomeric standards
[0104] The retention time differences of 40 stereoisomeric polypeptides were analyzed. The results showed that the retention time difference was more than 2 minutes (attached Figure 2 Figure A), so the polypeptides with a retention time difference of more than 2 minutes were used as potential isomeric modified polypeptides, providing guidance for the screening and identification of stereoisomeric polypeptides in large omics samples.
[0105] (4) Gas-phase separation of stereoisomeric standards
[0106] As a key dimension of conformational resolution mass spectrometry, ion mobility mass spectrometry (IM-MS) provides valuable insights into the subtle differences between stereoisomers. We collected the CCS values of 40 stereoisomeric peptides and observed a different distribution from RT, with a CCS difference of less than 4% (attached Figure 2B). The complementarity of gas-phase CCS and liquid-phase RT for stereoisomeric polypeptides indicates that combining these parameters can further enhance the differences between stereoisomers. A three-dimensional vector is assigned to each stereoisomeric polypeptide, with the coordinates (x, y, z) representing RT, CCS, and m / z respectively. Different data points of each isomer set are shown on the spatial vectors of the x and y axes, enabling the discrimination of multidimensional stereoisomers.
[0107] Example 2: Construction of N2a cell model and screening of protein isomerization from omics data
[0108] (1) Construction of Parkinson's disease cell model
[0109] Take N2a cells in the logarithmic growth phase and seed them in 96-well plates at a density of 5×10^3 cells / well, and culture for 24 hours. After the cells adhere, add different concentrations of 6-hydroxydopamine (6-OHDA) to treat the cells. After treatment, use methods such as MTT and Western blot to confirm the successful construction of the model. According to the results of MTT and Western blot, three different concentrations (50 / 150 / 300 μM) of 6-OHDA are selected for subsequent omics experiments.
[0110] (2) Sample preparation and LC-MS omics analysis
[0111] Take cells from the normal group and the disease group, and prepare samples according to the workflow of proteomics, including steps such as protein extraction and enzymatic digestion. Perform omics data collection and analysis, and process the mass spectrometry data through Maxquant to obtain a series of txt files for isomer retrieval.
[0112] (3) Isomer retrieval
[0113] According to the previous isomerization retrieval workflow (step seven above), perform isomer retrieval on the samples of the normal group and the disease group to discover potential isomeric modifications. Obtain the isomeric modification lists of the normal group and the disease group to provide a basis for subsequent data analysis and research.
[0114] (4) Post-processing data analysis
[0115] Compare the omics data of the normal group and the disease group, and analyze the differences caused by isomeric modifications at the polypeptide and protein levels. Analyze the data, further explore the biological significance of protein isomeric modifications, and investigate the association between protein isomeric modifications and the occurrence and development of diseases. Through the analysis of the selected polypeptides, it is found that most isomeric polypeptides have only two isomers (including one wild-type polypeptide without mutation), and a few have three isomers ( Figure 3 ). Further select some of the screened polypeptides for verification. As shown in the appendix Figure 4As shown, after screening by the heterogeneous algorithm, the polypeptide IGGIGTVPVGR identified two isomers. Further, based on the extracted ion m / z 513.31, the original data was extracted to obtain its chromatogram. The results showed that two chromatographic peaks appeared at retention times of 50.78 min and 111.76 min respectively, which was consistent with the results after isomer screening, indicating the accuracy of the algorithm.
[0116] Further analysis was carried out at the protein level, as shown in the appendix Figure 5 As shown, from 5630 proteins of N2a, 182 potential stereoisomerically modified proteins were screened. The results of quantitative hierarchical clustering showed that based on the protein abundances of different groups, a total of two clusters were generated. Among them, the within-group differences in the experimental groups with low and medium concentrations were greater than the between-group differences, which may suggest that these proteins have different expression patterns or functions at different concentration levels. Through gene variant analysis, the changes in genes or gene sets related to the functions of these proteins can be further identified, and in-depth analysis can be carried out on the selected biological processes and pathways in each cluster. The results of GO enrichment analysis and pathway analysis showed that after 6-OHDA administration, cell metabolism was enhanced and ATP production increased to maintain the normal life activities of cells. In addition, pathways related to neurological diseases, such as Parkinson's disease, were significantly upregulated, indicating that stereoisomeric modification may be closely related to neurological diseases and may play a certain role in the pathogenesis of related diseases. These findings may contribute to a deeper understanding of the association between stereoisomerically modified proteins and neurological diseases, providing important insights and research directions for further research and treatment of neurological-related diseases.
[0117] The mass spectrometry method for large-scale discovery of D-type proteins according to the present invention includes: developing a chromatographic separation method applicable to a large number of heterogeneous polypeptides and a high-throughput data processing algorithm for signal recognition; performing a conformational resolution proteomics mass spectrometry experiment on a complex biological sample to collect a five-dimensional proteomics data set; using a protein database for sequence retrieval and protein identification to obtain characteristic information such as the sequence, abundance, retention time, and biological function of the identified proteins; establishing screening conditions for polypeptide isoform modification retrieval at the omics level based on the proteomics analysis behavior characteristics of heterogeneous polypeptide standards; for the data set generated by sequence database searching, using appropriate isoform screening conditions, including data characteristics such as retention time, collision cross-sectional area, and abundance, to initially screen to obtain a list of potential D-type proteins; for the initially screened list, further using an isoform chromatographic signal amplification algorithm to iteratively process the original data, checking the quality and quantity of isoform signal peaks, and combining biological experiments to verify D-type mutations, and finally obtaining a list of D-type proteins. The present invention combines an isomeric polypeptide separation based on reversed-phase liquid chromatography, an isomeric chromatographic signal amplification, and an algorithm for large-scale extraction of high-order isomeric proteomics data, and can accurately identify protein conformational isoform modifications, improve the reliability and accuracy of large-scale identification of D-type proteins; can quickly process large-scale proteomics data; and has universality.
[0118] The large-scale algorithm for protein isoform modification proposed by the present invention provides an innovative solution for the rapid discovery of novel post-translational modifications in large omics samples. The close connection between protein isoform modification and neurological diseases and the association between the conformational isomer ratio and the isomerization ratio of neurological disease markers are of great significance, and these associations may have a profound impact on the pathogenesis, diagnosis, and treatment of neurological diseases. The following are some of the key contributions and impacts that this invention may bring:
[0119] 1) Rapid discovery of novel modifications: Through the new method of conformational resolution omics mass spectrometry and the algorithm for large-scale processing of isomeric data of the present invention, novel protein isoform modifications can be quickly and efficiently discovered in large omics samples, providing more comprehensive modification information for researchers and promoting in-depth exploration in the field of protein modification.
[0120] 2) Breakthrough in neurological disease research: Linking protein isoform modification with neurological diseases helps to deeply understand the pathogenesis of neurological diseases and provides new treatment strategies for the diagnosis, treatment, and prevention of neurological diseases.
[0121] 3) Revelation of the association between the degree of isoform modification and markers: Studying the association between the isoform modification ratio and neurological disease markers can provide important clues for further revealing the molecular mechanism of neurological diseases and help to discover new biomarkers.
[0122] 4) Promotion of personalized medicine: By enabling large-scale discovery of protein isoform modifications, it lays the foundation for the realization of personalized medicine and will contribute to the development of diagnostic and treatment plans tailored to individual differences.
[0123] 5) Advancement of scientific and technological innovation: This invention is innovative and will promote scientific and technological innovation in related fields, driving forward proteomics research and contributing to precision medicine and neurological disease research.
[0124] These impacts will help drive research and applications in related fields, broaden the understanding of the relationship between protein isoform modifications and neurological diseases, and provide new directions and inspiration for the progress and innovation in related fields.
[0125] The number of modules and the scale of processing described here are used to simplify the description of the present invention. Applications, modifications, and variations of the present invention will be apparent to those skilled in the art.
[0126] Although the embodiments of the present invention have been disclosed above, it is not limited to the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to specific details and the examples shown and described here.
Claims
1. A mass spectrometry method for large-scale discovery of D-type proteins, characterized in that, Comprising: Obtaining conformational resolution proteomics data of a biological sample, including retention time information, tandem mass spectrometry information, collision cross-sectional area, and stereoisomer signal response intensity of stereoisomeric polypeptides with the same amino acid sequence; Performing feature extraction on the data set generated by searching a protein database with the amino acid sequences in the conformational resolution proteomics data to obtain multiple pieces of feature data of a protein sample, so as to establish a feature set for isomeric modification retrieval of the protein sample, and each piece of feature data includes a protein name, an amino acid sequence, a retention time of chromatography, a charge, a signal response intensity, a mass-to-charge ratio, and a secondary mass spectrometry map; For the feature data under the same amino acid sequence, deleting the data with a retention time difference less than 2 minutes and / or a collision cross-sectional area difference less than 1.5% to screen out proteins or polypeptides with stereoisomers, and obtaining a preliminary screening result by calculating the isomerization ratio and quantity of each protein; When calculating the relative ratio of target peaks in the feature data of the same amino acid sequence, false positive noise data is excluded; Performing isomeric verification on the preliminary screening result to generate final D-type protein data.
2. The method for large-scale discovery of D-type proteins according to claim 1, wherein The method for obtaining the conformational resolution proteomics data is: Establishing a separation method applicable to chromatographic-level isomeric polypeptides, and identifying stereoisomeric polypeptide information through a signal amplification strategy algorithm and a Gaussian fitting algorithm; Performing a conformational resolution proteomics experiment on the biological sample, and collecting five-dimensional omics data including the retention time information, tandem mass spectrometry information, collision cross-sectional area, and the distribution of stereoisomer signal response intensity of the stereoisomeric polypeptides based on reversed-phase liquid chromatography under the same amino acid sequence.
3. The method for large-scale discovery of D-type proteins according to claim 1, characterized in that, Also comprising: Based on the proteomics analysis behavior characteristics of isomeric polypeptide standards, establishing screening conditions for isomeric modification retrieval of polypeptides at the proteomics level, and the screening conditions include a threshold for the retention time difference and / or a threshold for the collision cross-sectional area difference between the feature data under the same amino acid sequence.
4. The method for large-scale discovery of D-type proteins according to claim 1, characterized in that, For the preliminary screening result, for the chromatographic peak signal obtained from its corresponding original data, processing the chromatographic peak signal of the original data through a signal amplification strategy algorithm and a Gaussian fitting algorithm, and iteratively processing the chromatographic peak signal to verify the quality and quantity of the isomeric signal peaks under the same amino acid sequence; At the same time, combining biological experiments to verify the D-type mutation to verify the preliminary screening result, and generating the final D-type protein data.
5. The method for large-scale discovery of D-type proteins according to claim 2, characterized in that, Integrating the five-dimensional proteomics data set, and efficiently separating and identifying polypeptide isomers from the perspective of amplifying the collision cross-sectional area difference and the chiral isomer difference.
6. The mass spectrometry method for large-scale discovery of D-type proteins according to claim 2, wherein The protein isomerization modification data includes a protein name, an amino acid sequence, a charge, a mass-to-charge ratio, a retention time, a collision cross-sectional area, a response intensity, and an isomerization ratio.
7. The mass spectrometry method for large-scale discovery of D-type proteins according to claim 1, characterized in that, Both the feature set for isomeric retrieval modification of the protein sample and the output protein isomerization modification data are displayed in a row-column form.
8. An electronic device, characterized in that, Comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-7.
9. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, the method described in any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Disease protein biomarker identification method
CN115112778A