Application of aa-score driven single-site polypeptide clustering strategy in peptidomics research

By employing a single-site peptide clustering strategy and the aa-score algorithm, this study reveals the effects of enzyme system activity on changes in protein amino acid sites, resolving redundancy and visualization issues in peptidomics data analysis, providing efficient biomarkers for disease diagnosis, and improving the accuracy and interpretability of peptidomics research.

CN121148475BActive Publication Date: 2026-04-07INSTITUTE OF BIOPHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing peptidomics data analysis methods cannot effectively reveal information about changes in enzyme systems, resulting in high variability in peptide identification, large inter-individual differences, and many missing values. Traditional biomarker methods are not very reliable, and data visualization tools lack rationality and resolution.

Method used

A single-site peptide clustering strategy is proposed. The aa-score algorithm and data visualization tools are used to reveal the changes in protein amino acid sites caused by enzyme system activity. Biomarkers of amino acid site peptide clusters are developed, and data analysis and visualization are performed using R language and SpectroDive software.

Benefits of technology

It improves the efficiency and accuracy of peptidomimetics data analysis, reduces redundancy, provides highly sensitive and specific disease diagnostic biomarkers, enhances the interpretability of results, and supports the understanding of disease pathophysiology and the search for treatment options.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121148475B_ABST
    Figure CN121148475B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of peptidomics research strategy, and in particular to a data analysis ecosystem of single-position peptide clustering based on aa-score method, a construction method thereof and application thereof. The aa-score algorithm provides a new peptidomics research strategy, i.e. single-position peptide clustering, based on the overall change effect of a certain amino acid site of a protein by enzyme system activity, which greatly reduces the redundancy of polypeptide group data, weakens the serious missing value problem of peptidomics, improves the efficiency and accuracy of peptidomics data analysis, and can quickly process large-scale data sets, thereby providing a new research idea for the field of peptidomics research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of peptidomics research strategy technology, and in particular to a data analysis ecosystem, construction method, and application of a single-position peptide clustering strategy based on the aa-score method. Background Technology

[0002] Once considered a subset of the proteome, the peptide genome is now recognized as a distinct molecular entity and a key molecular component in biological systems. Mass spectrometry-based peptidomics research aims to directly measure and structurally characterize peptides in biological systems in a high-throughput manner, providing an important tool for studying disease mechanisms, discovering biomarkers, and developing innovative therapies.

[0003] Endogenous peptides are primarily generated through proteolysis, and their sequence amino acid composition and abundance information represent the protein degradation status and specific enzyme system states and activities. However, specific enzymatic cleavage and uncontrollable exopeptidase activities introduce gradient-like, seemingly redundant, dynamic sequence information into peptidomics. Its inherent characteristics, such as low abundance, broad chemical properties, and limitations of detection methods, lead to high variability in peptide identification, significant inter-individual differences, and a large number of missing values ​​in the peptidomimetry. These unique omics attributes of peptidomics have significantly constrained the development of this field. Current peptidomics data analysis methods are mainly borrowed from other omics fields, especially proteomics, including mass spectrometry qualitative and quantitative analysis, differential expression analysis, and functional enrichment analysis. While these borrowed methods provide important technical support for peptidomics research, these data analysis methods, lacking distinctive features, cannot fully reveal the true information hidden within this omics, especially the changes in enzyme systems. Currently, most peptidomics studies focus only on the most significantly different peptides, ignoring the influence of other peptides or missing peptides; biased data analysis may produce contradictory conclusions. Furthermore, peptide generation is closely related to protein and enzyme systems, and the rich biological information contained in the peptidomme presents challenges for data visualization. Existing peptidomme visualization tools, such as Peptimetric, PepEx, and Peptigram, are based on calculations using mass spectrometry signal response intensity or the accumulation of PSM numbers, which lack reasonable or sufficient data resolution. Therefore, there is an urgent need to develop novel data analysis and effective visualization methods targeting peptidommics characteristics in order to more directly uncover the unique features of clinical peptidommics in diseases.

[0004] Currently, the screening of peptide-based biomarkers mainly relies on the difference in peptide abundance between experimental and control groups. These biomarkers fall into two main categories: classic single-peptide biomarkers and peptide combination biomarkers screened using machine learning or other statistical methods. Both types of biomarkers can be represented by peptide sequences or by m / z values ​​detected by mass spectrometry. Theoretically, classic single-peptide biomarkers reduce or minimize the differences caused by enzyme systems acting on specific protein sites. Furthermore, single peptides are variable under uncontrolled exonuclease activity, thus their reliability is low. Peptide combination biomarkers offer higher specificity and sensitivity, but the introduction of more analytical variables increases data complexity, and their biological interpretability is limited. Summary of the Invention

[0005] To address the aforementioned issues, based on the characteristics of peptide generation, this invention proposes a novel peptidomics research strategy of single-position peptide clustering. The biological significance of this strategy lies in revealing the overall effect of enzyme system activity on a specific amino acid site in a protein. Based on this strategy, this invention introduces for the first time the concepts of amino acid sites (including stable amino acid sites, referred to as stable points or stable aa positions; dynamic amino acid sites, aa positions; changing points / aa positions; transition points / aa positions) and peptide clusters. Specifically, when an enzyme system acts on a specific site in a single protein, a group of peptides with the same terminal sequence but varying lengths is called a peptide cluster. The shared terminal amino acids within a peptide cluster are called stable amino acid sites (stable points or stable aa positions), while the non-shared terminal amino acids are called dynamic amino acid sites (dynamic points / aa positions). Stable and dynamic amino acid sites are collectively referred to as changing points / aa positions. Amino acid sites are represented by gene names or protein names and position numbers or similar formats; amino acid sites are directional, that is, if the site is the N-terminus of a polypeptide, it is represented by N-term or N, and if the site is the C-terminus of a polypeptide, it is represented by C-term or C; directional stable amino acid sites are used to represent polypeptide clusters at that site.

[0006] In the data science ecosystem surrounding the aforementioned single-site peptide clustering strategy, this invention has developed the aa-score (amino acid score) algorithm and its supporting data analysis and visualization tools. Furthermore, this invention proposes a novel category of biomarkers, namely peptide cluster biomarkers based on amino acid sites. The discovery process includes: the discovery and verification of amino acid sites with diagnostic potential, and the mining of their representative peptide clusters.

[0007] The aa-score algorithm is divided into two forms based on different mass spectrometry experimental designs and analytical purposes: grouped aa-score and individual aa-score. Grouped aa-score is used for overall comparison between groups and has no special requirements for mass spectrometry experimental design. The positive and negative aa-score values ​​represent the enzyme activity in an activated or inhibited state at the site, which facilitates data visualization. Individual aa-score is used for calculation of a single sample and requires the accumulation of degradation effects by using reference samples or synthetic peptides. Its calculation method is different from grouped aa-score. It is a direct accumulation of the relative or absolute quantitative values ​​of peptides, thereby generating a diagnostically significant digital feature of peptide cluster biomarkers.

[0008] Schematic diagram of the implementation of the novel research strategy of this invention

[0009]

[0010] The aa-score algorithm and its supporting data analysis and visualization tools, along with the process for discovering and validating peptide cluster biomarkers based on amino acid sites, in the single-site peptide clustering peptide research strategy provided by this invention are as follows:

[0011] This invention provides a method for data analysis and visualization related to the aa-score algorithm based on the aascore package developed in R language. Other programming software can also be used to perform related analysis and visualization according to the relevant definitions.

[0012] The first phase of non-targeted quantitative peptidomics research identified amino acid sites with diagnostic potential and extracted disease characteristics.

[0013] 1. Prepare cohort peptide samples and acquire data from them using conventional liquid chromatography-mass spectrometry (LC-MS / MS).

[0014] Step 1 preferably involves mixing equal volumes of queue samples (e.g., plasma / serum) and extracting peptides to serve as reference peptide samples. When submitting the mass spectrometry sample collection list, the reference samples are evenly interspersed among the queue samples for data collection.

[0015] 2. Data analysis and visualization based on grouped aa-scores were performed to identify amino acid sites with diagnostic potential and their corresponding differentially expressed peptides.

[0016] 2.1 In non-targeted peptidomics research, traditional peptide differential analysis is performed.

[0017] The characteristic of this differential analysis is that at least three quantitative values ​​must exist in each group to be included in the differential analysis. This can be calculated using the ttest() or ttest_function() functions in the aascore package, or other statistical software.

[0018] 2.2 Single-site cluster analysis was performed on differentially expressed peptides and peptides with diagnostic potential.

[0019] Differentially expressed peptides are clustered at single sites to extract stable and dynamic amino acid points. Specifically, protein sites corresponding to the N- and C-termini of all differentially expressed peptides are extracted. Sites with ≥2 peptides are considered stable amino acid sites, while sites with 1 peptide are considered dynamic amino acid sites. Changing points are the combined set of stable and dynamic amino acid sites. This analysis can be performed directly using functions such as `cluster_only_peptides()` and `find_stable_points()` in the `aascore` package, or it can be performed using other software.

[0020] ROC curves were used to analyze the diagnostic potential of differentially expressed peptides, and single peptides with diagnostic potential (e.g., AUC > 90) were screened based on their AUC values. Furthermore, the stability sites contained in the peptides with diagnostic potential were extracted.

[0021] 2.3 Calculate the aa-score value of amino acid sites based on the grouped aa-score method to achieve visual analysis.

[0022] First, the specific method for calculating the grouped aa-score value for each amino acid site is to directly sum the fold change values ​​of all differentially expressed peptides contained at that site. Unlike traditional calculation methods, the specific method for calculating the fold change of a single peptide is as follows: In comparative quantitative analysis, when the ratio (ratio) of the average (or median, or other group-representative figures) of the experimental group peptide quantification to the average (or median, or other group-representative figures) of the control group peptide quantification is >1 (i.e., the peptide is upregulated), fold change = ratio; and when the ratio (ratio) of the average (or median, or other group-representative figures) of the experimental group peptide quantification to the average (or median, or other group-representative figures) of the control group peptide quantification is <1 (i.e., the peptide is downregulated), fold change = -1 / ratio.

[0023] The following diagram illustrates the single-site clustering process of peptides and the calculation method of grouped aa-score:

[0024]

[0025] Grouped aa-scores can be calculated directly using the `grouped_aascore()` function in the `aascore` package, or by using other programming software. Specifically, the function can be input with the gene name of the protein to be analyzed, a differentially expressed polypeptide information table, etc. The function will output: a data frame including the differentially expressed polypeptide sequence and its initiation site information, ID information for visualization, and abundance variation information; and the score (aa-score) and cumulative aa-score for each amino acid site in different comparison groups. In R, the `lapply()` function can be used to execute the `grouped_aascore()` function in batches, ultimately obtaining a data frame containing the amino acid score and cumulative score for each protein site, as well as a data frame for visualization.

[0026] Then, the peptidomics data of a specific protein can be visualized using fragment profiling graphs, Aa-score graphs, and cumulation of Aa-score graphs. The cumulation of Aa-score graph can be further extended into a waterfall map graph. These graphical visualizations can also be designed and implemented using other programming software.

[0027] Fragment profiling graphs display the distribution and abundance changes of all differentially expressed peptide degradation products from the perspective of a single protein. These peptides are arranged according to the position of their N-terminal amino acid residues in the protein and their total amino acid length. The peptide color represents the abundance change of the peptide in the group comparison. The Peptide data frame output by the grouped_aascore() function in the aascore package can be used as the data input for the fragment_profiling() function to achieve visualization of fragment profiling graphs.

[0028] The aa-score graph displays information about the intensity of protein backbone degradation from the perspective of enzyme system activity. The horizontal axis represents the amino acid sites of the protein from the N-terminus to the C-terminus, and the vertical axis represents the aa-score value for each amino acid site. The aa-score graph can be visualized using the cumulative aascore data frame output by the grouped_aascore() function in the aascore package as input to the aa_score_plot() function.

[0029] The cumulation of aa-score graph simplifies data by accumulating aa-score values, making it easier to display the changes in cumulative aa-score values ​​at amino acid sites across multiple groups. This graph can use the cumulative_aascore data frame output by the grouped_aascore() function in the aascore package as data input for the cumulative_aascore_plot() function for visualization.

[0030] The difference between a waterfall map and a cumulation of aa-score (AA-Score) graphs is that a waterfall map allows you to add specific amino acid sites to the data visualization, such as transition points or protein amino acid sites recorded on the UniProt website. You can use the cumulative_aascore data frame output by the grouped_aascore() function in the aascore package as input to the waterfall_map() function to visualize the waterfall map.

[0031] In addition, the `grouped_aascore_plot()` function in the `aascore` package can align and merge the three types of graphs mentioned above: fragment profiling, aa-score, cumulation of aa-score, and waterfall map. The `aascore` package also provides other data analysis or visualization functions, such as the `identified_peptides()` function for visualizing missing data.

[0032] 2.4 Extraction of important amino acid sites related to the disease

[0033] First, transition points are extracted from the changing points of differentially assigned peptides. Transition points are defined as points on each curve in the cumulation of aa-score graph where the slope changes positively, negatively, or zero. These changes reflect the most drastic changes in the grouped aa-score value of the measured variable, i.e., the amino acid site, such as from no peptide coverage to peptide coverage, from peptide coverage to no peptide coverage, or a sudden increase or decrease of peptides with opposite trends at that position. The specific mathematical characteristics are as follows: First, the slope of the line segment formed by a site and its two adjacent sites on the left and right changes to either sign or zero. Then, if the site to the left is a changing site covered by peptides, the site to the left of this site (N-1) is a transition point with a direction sign of C. If the site to the right is a changing site covered by peptides, the site to the right of this site (N+1) is a transition point with a direction sign of N. If the site to the left or right is a non-peptide-covered region or a non-changing site, then the site to the left or right of this site is not a transition point. Observations revealed that transition sites often partially overlap with restriction enzyme sites or peptide endpoints recorded in UniProt. Therefore, transition sites are considered to be disease-related and potentially functional sites, belonging to important amino acid sites. Transition sites can be extracted using the `inflect_point_waterfall()` function in the `aascore` package, or extracted manually using other programming software. Furthermore, if the cohort contains multiple disease subtypes, transition sites present in comparisons between all subgroups and healthy controls can be extracted. These sites exhibit more stable and unique hydrolytic properties and are considered important amino acid sites related to the disease. To reduce the influence of the total protein length on the amount of degradation products, the mean of aa-score of these important amino acid sites was normalized using the full-length protein amino acid sequence. Transition amino acid sites with larger normalized mean of aa-scores or larger mean of aa-scores have higher diagnostic potential.

[0034] 2.5 Identify amino acid sites suitable for subsequent targeted analysis and possessing diagnostic potential, and extract all corresponding peptides and those that can be preferentially included.

[0035] Venn diagram analysis was performed on the important amino acid sites related to the aforementioned diseases and the stability sites contained in peptides with diagnostic potential. The intersecting amino acid sites not only possess diagnostic potential, but some of their corresponding single peptides also exhibit high diagnostic potential. These peptides with high diagnostic potential can be preferentially included in the representative peptide clusters of that site, helping to reduce the number of peptides in the representative peptide clusters. Furthermore, all peptides contained in amino acid sites with diagnostic potential were included in the targeted peptidomics analysis.

[0036] Step 2.5 Preferably, the individual aa-score of amino acid sites is calculated using reference samples to perform differential analysis on the amino acid sites. Then, ROC curve analysis is used to analyze the differentially expressed amino acid sites, and amino acid sites with diagnostic potential are directly identified based on the AUC value. Simultaneously, peptides with diagnostic potential can be preferentially included in the representative peptide cluster of that site.

[0037] The second phase of targeted quantitative peptidomics research verifies the diagnostic potential of amino acid sites, identifies representative polypeptide clusters, and proposes polypeptide cluster biomarkers.

[0038] 3. Preparation of peptide samples

[0039] Peptide extraction from cohort samples is used for targeted analysis; simultaneously, equal volumes of cohort samples (such as plasma / serum) are mixed and peptides are extracted for use in preliminary experiments, the establishment of targeted methods, or as reference samples.

[0040] 4. Preliminary PRM screening of representative polypeptide clusters at amino acid sites

[0041] 4.1 Establishment of PRM Pre-experiment Methods

[0042] The method was established using SpectroDive v12.1 (Biognosys) software. When developing the PRM detection method, the Pulsar engine was used to retrieve raw data from the entire peptide genome, generating a spectral library. Key search parameters were set as follows: non-enzymatic digestion; peptide length 7-40; modification: variable modification: Gln->pyro-Glu, Oxidation (P, M). Amino acid sites with diagnostic potential and their contained peptides were included in the PRM preliminary analysis list as much as possible, with the following parameter settings: precursor ion mass-to-charge ratio range, 350-1500; precursor ion charge number, 2-6; peptide length, 8-40; daughter ion mass-to-charge ratio range, 300-1800; maximum daughter ion charge number, 3; ion type, b, y ions; allowed neutral loss types, H2O, NH3, and no loss; top 6 daughter ions were selected. Unscheduled PRM was selected to establish the preliminary experimental method.

[0043] 4.2 PRM Pre-Experiment Data Acquisition

[0044] The LC-MS / MS instrument used in the parallel reaction monitoring (PRM) targeted quantitative analysis stage was an Orbitrap Eclipse Tribrid mass spectrometer (Thermo Fisher Scientific) equipped with an EASY-nLC1200 HPLC system. The chromatographic elution gradients were set as follows: 5-10% B, 3 min; 10-20% B, 22 min; 20-30% B, 22 min; 30-40% B, 13 min; 40-99% B, 4 min; 95% B, 9 min. The mass spectrometry parameters for PRM data acquisition were set as follows: fullMS mass-to-charge ratio scan range of 350-1200 m / z, resolution of 60,000 (m / z 200), and AGC of 6 × 10⁻⁶. 5 The maximum ion implantation time was 50 ms. The MS / MS resolution was 30,000 (m / z 200), the isolation window was 1.0 Da, the collision energy was 30%, and the AGC was 2 × 10⁻⁶. 5 The maximum ion implantation time is 80 ms.

[0045] 4.3 Analysis of PRM preliminary experimental mass spectrometry data to screen representative polypeptide clusters at amino acid sites

[0046] Data were analyzed using SpectroDive v12.1 (Biognosys) software. Based on the preliminary PRM results, target peptides were screened. The screening criteria were as follows: a. targetable by PRM; b. good peak shape with no interference; c. high abundance. Ultimately, multiple peptides with diagnostic potential at specific amino acid sites were preferentially selected as representative peptide clusters for those sites; when no peptides with diagnostic potential were found, representative peptide clusters for those sites were selected based on peak shape and abundance.

[0047] Step 4.3 Preferably, for representative polypeptide clusters, polypeptides in the form of heavy isotope labels are synthesized.

[0048] 5. Establish a targeted analysis method to calculate the individual aa-score.

[0049] 5.1 Establish a targeted analysis method for formal experiments, and conduct targeted analysis on the cohort samples.

[0050] A targeted analysis method was established using a pooled sample approach. 10×iRT standard peptides (Biognosys) were added to the pooled sample for unscheduled PRM analysis, thereby establishing a scheduled PRM targeted quantitative analysis method that includes the m / z of the target endogenous peptide precursor ion and the retention time of the precursor ion. The retention time window was set to ±2.5 minutes. Based on this method, scheduled PRM data were formally collected from cohort samples and at least three reference samples.

[0051] Step 5.1 Preferably, based on the abundance ratio of the target endogenous peptide in the pre-experimental samples, a relabeled synthetic peptide is mixed as an internal standard. A targeted analysis method is established using mixed samples. The internal standard and 10×iRT standard peptide (Biognosys) are added to the mixed samples for unscheduled PRM analysis, thereby establishing a scheduled PRM targeted quantitative analysis method that includes the m / z of the target relabeled peptide precursor ion, the m / z of the endogenous peptide precursor ion, and the retention time of the precursor ion. The retention time window is set to ±2.5 minutes. Based on this method, scheduled PRM data collection for the cohort samples is formally carried out.

[0052] Step 5.1 can also preferably involve establishing a standard curve for the relabeled peptide for absolute quantification analysis. Based on the expected abundance of the target peptide, serially diluted relabeled peptide standards are prepared. The serially diluted relabeled peptide standards and iRT peptides are added to the mixed sample for scheduled PRM analysis to establish a standard curve.

[0053] 5.2 Scheduled PRM data analysis was conducted to calculate the individual aa-score value of amino acid sites, and peptide cluster biomarkers were proposed.

[0054] Data were processed in SpectroDive software to exclude daughter ions with significant interfering signals. For each precursor ion, at least three daughter ion pairs were selected for quantitative analysis. The data were exported and further processed in R. Based on the average abundance of endogenous peptides in the reference samples, the relative content of peptides in individual samples was calculated, thereby calculating the individual aa-score of amino acid sites. The diagnostic potential of amino acid sites, including AUC, specificity and sensitivity, and optimal cutoff value, was analyzed using ROC curves to propose peptide cluster biomarkers.

[0055] Step 5.2 Preferably, based on the relabeled polypeptide mixture standard added to a single sample, i.e., the internal reference standard, the relative content of the polypeptide in the single sample is calculated, thereby calculating the individual aa-score of the amino acid site.

[0056] Step 5.2 can also preferably be performed by calculating the absolute content of the polypeptide in a single sample based on the established standard curve, thereby calculating the individual aa-score of the amino acid site.

[0057] Beneficial effects: 1. Based on the aa-score algorithm, a novel peptidomics research strategy is provided, namely the single-site peptide clustering strategy, which focuses on the overall change effect of enzyme system activity on a certain amino acid site of protein. This strategy greatly reduces the redundancy of peptidomics data, mitigates the serious missing value problem in peptidomics, improves the efficiency and accuracy of peptidomics data analysis, and can quickly process large-scale datasets, providing new research ideas for the field of peptidomics research.

[0058] 2. A comprehensive data visualization solution developed based on the relationship between proteins, peptides, and enzyme systems. This enhances the interpretability of results and facilitates integration and analysis with information obtained from other technologies, such as proteomics and Western blotting, as well as clinical information. This will contribute to a better understanding of the pathophysiological processes of diseases and the search for new treatment options.

[0059] 3. A new class of biomarkers based on amino acid sites. These biomarkers are biologically significant and can be detected by mass spectrometry. The detection method is stable and efficient, and does not require cumbersome procedures. They have the advantages of high sensitivity and high specificity in disease diagnosis and have important scientific research and clinical application value.

[0060] 4. Based on disease-related and potentially functional sites provided by inflection points, these sites are conserved in the population cohort and partially overlap with the enzyme cleavage or peptide sites recorded in UniProt. Combining them with existing bioactive peptide prediction software can further improve the reliability of bioactive peptide prediction. Attached Figure Description

[0061] Figure 1 These are polypeptides with diagnostic potential from the whole polypeptide group in Example 1 of this invention.

[0062] Figure 2 These are important amino acid sites related to the disease that were found in the whole polypeptide group in Example 1 of this invention.

[0063] Figure 3 This figure shows the overlap between the disease-related important amino acid sites contained in the FGA protein in Example 1 of this invention and the significant amino acid sites recorded in UniProt. The information marked in the figure is the significant sites recorded on the UniProt website, and the numbers above the figure represent the overlapping sites.

[0064] Figure 4 This describes the overlap between the variable amino acid sites and disease-related amino acid sites corresponding to the polypeptides with diagnostic potential discovered in Example 1 of this invention. Six of the overlapping amino acid sites are stable amino acid sites.

[0065] Figure 5 This is a novel disease feature discovered based on grouped aa-score in Example 1 of this invention. (A) Differentially expressed peptides discovered based on traditional single-peptide analysis. (B) Abundance of AHSG[312-339] peptides in different groups. (C) Grouped aa-score analysis of AHSG

[339] sites and their visualization. (D) Abundance of AHSG protein detected by mass spectrometry. (E) Western blot analysis of AHSG protein degradation. (F) The relationship between AHSG protein and HAMP in vivo as reported in the literature.

[0066] Figure 6 Figure A shows the differences in individual aa-scores of amino acid sites in different groups during PRM-targeted analysis in Example 2 of this invention, and the diagnostic performance of representative polypeptide clusters at amino acid sites (Figure B). The individual aa-scores are calculated based on reference samples.

[0067] Figure 7 Figure A shows the differences in individual aa-scores of amino acid sites in different groups during PRM-targeted analysis in Example 3 of this invention, and the diagnostic performance of representative polypeptide clusters at amino acid sites (Figure B). The individual aa-score is calculated based on the absolute concentration of the polypeptide.

[0068] Figure 8 Figure A shows the differences in individual aa-scores of amino acid sites in different groups during PRM-targeted analysis in Example 3 of this invention, and the diagnostic performance of representative polypeptide clusters at amino acid sites (Figure B). The individual aa-score is calculated based on a mixture of relabeled polypeptides, i.e., internal standard.

[0069] Figure 9Example 4 of this invention uses individual aa-score-based amino acid site analysis to discover amino acid sites with diagnostic potential for colorectal cancer. Other commonly used data analysis methods in omics analysis can also be applied to amino acid site analysis. Figure (A) shows UMAP analysis, and Figure B, a volcano plot, illustrates amino acid sites with AUC > 90 and the number of peptides in each peptide cluster at each amino acid site. The individual aa-score is calculated based on reference samples, demonstrating that amino acid site-based analysis provides a new approach for peptidomics research.

[0070] Figure 10 This is Example 4 of the present invention, comparing the missing values ​​of amino acid site analysis based on individual aa-score with conventional peptidomics analysis. Under the criterion of ensuring at least three quantitative values ​​for each group are involved in the differential analysis, the individual aa-score analysis based on the reference sample showed 1766 amino acid sites involved in the differential analysis, containing a total of 2327 peptides; while this cohort only had 1774 single peptides involved in the differential analysis. This indicates that the individual aa-score-based analysis strategy significantly improves peptide utilization, reduces data redundancy, and minimizes the impact of missing values. Detailed Implementation

[0071] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.

[0072] Example 1: Discovery of amino acid sites with potential for screening and diagnosis of β-thalassemia

[0073] Patients and healthy participants involved in this invention were recruited by the First Affiliated Hospital of Guangxi Medical University (Guangxi Zhuang Autonomous Region, China). All participants provided written informed consent; for patients under 18 years of age, consent from their parents or legal guardians was required. Two weeks after their last transfusion, patients provided blood samples and completed a questionnaire including information on their first transfusion, treatment, and other medical conditions. This study recruited 286 patients with β-thalassemia and 51 healthy controls. Patients included in the plasma polypeptide study received both transfusions and iron chelation therapy; none underwent splenectomy. Whole blood was collected in EDTA vacuum aspirators and centrifuged at 3000 × g for 10 minutes. Plasma collected from the supernatant was stored at -80°C until use. This study was designed and conducted in accordance with the Declaration of Helsinki. The Ethics Committee of the First Affiliated Hospital of Guangxi Medical University approved this study.

[0074] The above samples were diagnosed by BGI Genomics Clinical Laboratory (Shenzhen, China) using Gap-PCR and SNP testing to determine the mutation type of the patients. Mild β-thalassemia carriers were classified as β0βN / β+βN, intermediate β-thalassemia patients as β+β+ / β+β0, and severe β-thalassemia patients as β0β0. Gene mutations were categorized as follows: β+ mutations included -28A>G (HBB: c.-78A>G), -29A>G (HBB: c.-79A>G), codon 26G>A (HbE, HBB: c.79G>A), IVS-II-654C654>T (HBB: c.316-197C>T), and IVS-II-5G>C (HBB: c.315+5G>C). β0 mutations include codons 41 / 42-TTCT (HBB: c.126_129delCTTT), codon 17A>T (HBB: c.52A>T), codons 71 / 72 + A (HBB: c.216_217insA), IVS-I-1G>T (HBB: c.92+1G>T), codon 43 G>T (HBB: c.130G>T), IVS-I-130 G>C (HBB: c.93-1G>C), codon37 G>A (HBB: c.114G>A), codons 27 / 28 +C (HBB: c.84_85insC), and codon 30A>G (HBB: c.91A>G).

[0075] Clinical parameters including HbF, SF, HbA2, HGB, MCV, and MCH were measured using standard techniques. These included the CELL-DYN fully automated hematology analyzer (Abbott Diagnostics) for measuring MCV, MCH, and HGB; high-performance liquid chromatography (VARIANT II, ​​Bio-Rad) for measuring HbF and HbA2; and electrochemiluminescence immunoassay (Cobas e601, Roche) for measuring serum ferritin (SF).

[0076] This invention is based on differential expression of whole peptides in plasma. It identifies important amino acid sites of disease-related proteins using differentially expressed peptides and a grouped aa-score algorithm. Then, ROC curve analysis is used to select differentially expressed peptides with diagnostic potential and their associated stable amino acid sites in plasma. The selection of amino acid sites with diagnostic potential and their contained peptide clusters from important disease-related protein amino acid sites and stable amino acid sites of differentially expressed peptides with diagnostic potential includes the following aspects:

[0077] 1. Whole-peptide genome study population cohort

[0078] A total of 54 people were included, including 13 healthy controls (Ctr), 8 mild β-thalassemia carriers (TT), 17 severe β-thalassemia patients (TI), and 16 severe β-thalassemia patients (TM).

[0079] 2. Isolation and enrichment of peptides in plasma samples

[0080] 2.1 Add 250 µL of methyl tert-butyl ether (MTBE), 50 µL of deionized water and 150 µL of methanol to 50 µL of plasma in sequence, mix well and let stand at 4°C for 30 min.

[0081] 2.2 Centrifuge the above sample at 21,000 g, 4℃ for 30 min, and take out the supernatant after centrifugation.

[0082] 2.3 Add 500 µL of MTBE and 100 µL of deionized water to the supernatant above, mix well, and centrifuge at 1,000 g and 4 °C for 10 min.

[0083] 2.4 The above samples were divided into two phases. The upper layer, rich in hydrophobic interfering substances such as lipids, was removed, and the lower clear liquid was retained and dried at 4°C.

[0084] 2.5 The obtained samples are processed by a desalting column and then ready for testing.

[0085] 3. Detection of the whole peptide genome using liquid chromatography-mass spectrometry (LC-MS / MS)

[0086] 3.1 Data acquisition was performed using an Orbitrap Exploris 480 mass spectrometer equipped with FAIMS Pro.

[0087] The LC-MS / MS instrument used was an OrbitrapExploris 480 mass spectrometer (Thermo Fisher Scientific) equipped with an EASY-nLC 1200 HPLC system and FAIMS Pro. The peptide was dissolved in 15 µL of 0.1% FA aqueous solution, centrifuged at 21,000 g at 4°C for 30 min, and then 10 µL of the supernatant was transferred to a sample vial for loading. The analytical liquid chromatography and mass spectrometry conditions were as follows: A Dr. Maisch GmbH ReproSil-Pur C18 AQ column (75 μm id × 20 cm, 3 μm) was used; the mobile phase was: A: water containing 0.1% formic acid, B: acetonitrile containing 0.1% formic acid; the flow rate was 300 nL / min; gradient elution was used: 4–11% B, 4 min; 11–21% B, 28 min; 21–30% B, 29 min; 30–42% B, 27 min; 42–95% B, 5 min; 95% B, 10 min. Two compensation voltages (-45 V and -65 V) were used for FAIMS separation. In data-dependent acquisition (DDA) mode, MS1 scan range: 350-1600 m / z, normalized AGC: 300%; maximum injection time: 80 ms; MS1 resolution: 60,000 (m / z 200); isolation window width: 1.6 m / z; HCD collision energy: 28%. MS / MS resolution: 60,000 (m / z 200); normalized AGC: 150%; maximum ion injection time: 118 ms; dynamic exclusion time: 30 s.

[0088] 3.2 Protein Discoverer 2.4 software was used to retrieve mass spectrometry data.

[0089] Mass spectrometry data were retrieved to obtain quantitative information on the whole peptide genome. Protein Discoverer 2.4 software was used, and the key parameters for data retrieval were set as follows: the database selected was the human database downloaded from Uniprot in September 2019; protease was set to No enzyme; the maximum errors for precursor and daughter ions were 10 ppm and 0.02 Da, respectively; variable modifications included oxidation of methionine and proline, cysteine ​​modification, and conversion of glutamine to pyroglutamic acid; the free dose ratio (FDR) for peptide level was set to 1%; label-free quantification was performed based on chromatographic area.

[0090] 4. Analysis of whole-peptide genome detection results

[0091] 4.1 Differential analysis and ROC curve analysis of peptides.

[0092] The p-value for comparison between the two groups was calculated using the t-test statistical method. The screening criteria for differentially regulated peptides were: fold change > 2 and p < 0.05 for upregulated peptides, and fold change < 0.5 and p < 0.05 for downregulated peptides. ROC curve analysis was performed on the differentially regulated proteins.

[0093] 4.2 Discovering stable amino acid sites for differentially expressed peptides with diagnostic potential.

[0094] In the three comparison groups, peptides with AUC > 0.9 were considered to have diagnostic potential; a total of 43 such differentially expressed peptides were identified. Figure 1 It contains 76 protein amino acid sites, of which 7 are stable amino acid sites.

[0095] 4.3 Identification of important amino acid sites related to the disease based on the grouped aa-score algorithm

[0096] Data analysis was performed using the R package aascore (1.0.0). First, grouped aa-scores were calculated for amino acid sites of the differentially expressed peptides. Then, the average aa-score for each site across three pairs of differential comparisons was calculated. The average aa-score was normalized to the full protein length to obtain the normalized mean aa-score. Transition points were extracted from the variable sites contained in the differentially expressed peptides, totaling 186. These sites were considered important amino acid sites associated with the disease. Figure 2 ).

[0097] Taking FGA protein as an example, this paper specifically analyzes the changing points and transition points of FGA protein. There are a total of 159 changing points in FGA protein, of which 3 sites (20, 101, and 122) are recorded in UniProt; there are a total of 14 transition points in FGA protein that are present in all subgroups compared to healthy controls, of which 3 sites (20, 101, and 122) are recorded in UniProt. Figure 3This result indicates that meaningful sites recorded in UniProt account for a higher proportion of transition points compared to changing points, suggesting that transition points are more meaningful for exploration and are more crucial for guiding targeted research.

[0098] 4.4 Screening for amino acid sites and polypeptide clusters with diagnostic potential

[0099] Comparative analysis of stable amino acid sites and disease-related important amino acid sites contained in differentially expressed peptides with diagnostic potential revealed that 6 sites were common amino acid sites. Figure 4 These shared amino acid sites, along with the first few disease-related important amino acid sites that have both high mean of aa-score and normalized mean of aa-score, are considered to have the greatest diagnostic potential. The polypeptide clusters of the amino acid sites with the greatest diagnostic potential include all differentially expressed polypeptides at that site.

[0100] 4.5 Discovering novel disease characteristics using grouped AA-score

[0101] When performing differential peptide analysis, the fold change of differentially expressed peptides in each subgroup compared to healthy controls was sorted from largest to smallest, and the top 20 peptides with the largest fold changes in upregulation and downregulation were displayed. Figure 5 A). Among the top 20 differentially expressed peptides, the degree of downregulation of AHSG [312-339] in the disease was found to be significantly different, and this peptide showed statistically significant downregulation differences in all subgroups. Figure 5 B), suggesting that enzyme activity may be downregulated at sites 312 and 339 of the AHSG protein. However, after analyzing and visualizing the AHSG protein using grouped aa-scores, it was found that the overall enzyme activity at site 339 was enhanced. Figure 5 C). Analysis of published plasma proteomics data revealed no significant change in the abundance of AHSG protein (which, theoretically, represents the combination of degraded and non-degraded protein forms) in patients with β-thalassemia. Figure 5 D), using Western blot analysis of the abundance of intact AHSG protein in patient plasma, it was found that AHSG protein did indeed show a trend of enhanced degradation in thalassemia patients. Figure 5E). Existing literature reports (reference: Stirnberg, M. et al. Cell surface serine protease matriptase-2 suppresses fetuin-A / AHSG-mediated induction of hepcidin. Biol Chem 396, 81-93, doi:10.1515 / hsz-2014-0120 (2015).) that AHSG protein can be specifically cleaved into A chain, B chain, and a linker chain in vivo. The cleavage site for the B chain and the linker chain is site 340. Site 340 was not found in the peptides identified by AHSG, but a large number of peptides with a C-terminus at 339 and an N-terminus at 341 were found, indicating that site 340 is cleaved into a single amino acid in vivo. Therefore, the enhanced AHSG degradation results in this example suggest that the cleavage activity at site 339 should be enhanced. Thus, amino acid site analysis based on grouped aa-scores is more reasonable than single peptide analysis. Furthermore, in vivo experiments have shown that the intact form of AHSG protein can induce HAMP expression. Based on this viewpoint, this embodiment illustrates that AHSG protein-induced HAMP expression is suppressed in thalassemia patients. Figure 5 F). HAMP protein (hepcidin) is known to be one of the potential therapeutic targets for thalassemia; therefore, dysregulation of AHSG degradation activity may be a new therapeutic target for the disease.

[0102] Example 2: Using a reference sample-based individual aa-score strategy, amino acid sites with potential for screening and diagnosing β-thalassemia were validated, and their representative polypeptide clusters were identified.

[0103] In this embodiment, the recruitment, enrollment, and isolation and enrichment of plasma peptide samples from individuals with β-thalassemia are the same as in Example 1. PRM targeted quantitative peptidomics technology is used to target and screen peptides contained in amino acid sites with β-thalassemia screening and diagnostic potential. ROC curve analysis is used to determine the diagnostic performance of representative peptide cluster markers at protein amino acid sites.

[0104] 1. Preparation of peptide samples from targeted analysis cohort plasma and mixed plasma

[0105] This cohort included 49 participants: 16 healthy controls, 17 carriers of mild β-thalassemia, and 16 patients with severe β-thalassemia. Specifically, plasma samples from the cohort were mixed in equal volumes, and peptides were extracted using the SPD method for use in both the preliminary and formal PRM experiments. The mixed peptide samples used in the formal PRM experiment were referred to as reference samples and were used to calculate the individual aa-score.

[0106] 2. PRM preliminary screening of amino acid sites and representative polypeptide clusters for targeted analysis.

[0107] 2.1 Establishment of PRM Pre-experiment Methods

[0108] The method was established using SpectroDive v12.1 (Biognosys) software. When developing the PRM detection method, the Pulsar engine was used to retrieve raw data from the entire peptide genome, generating a spectral library. Key search parameters were set as follows: non-enzymatic digestion; peptide length 7-40; modification: variable modification: Gln->pyro-Glu, Oxidation (P, M). Amino acid sites with diagnostic potential and their contained peptides were included in the PRM preliminary analysis list as much as possible, with the following parameter settings: precursor ion mass-to-charge ratio range, 350-1500; precursor ion charge number, 2-6; peptide length, 8-40; daughter ion mass-to-charge ratio range, 300-1800; maximum daughter ion charge number, 3; ion type, b, y ions; allowed neutral loss types, H2O, NH3, and no loss; top 6 daughter ions were selected. Unscheduled PRM was selected to establish the preliminary experimental method.

[0109] 2.2 PRM Pre-Experiment Data Acquisition

[0110] The LC-MS / MS instrument used in the parallel reaction monitoring (PRM) targeted quantitative analysis stage was an Orbitrap Eclipse Tribrid mass spectrometer (Thermo Fisher Scientific) equipped with an EASY-nLC1200 HPLC system. The chromatographic column and flow were the same as in step 2.2. The chromatographic elution gradient was set as follows: 5-10%, 3 min; 10-20% B, 22 min; 20-30% B, 22 min; 30-40% B, 13 min; 40-99% B, 4 min; 95% B, 9 min. The mass spectrometry parameters for PRM data acquisition were set as follows: full MS mass-to-charge ratio scan range 350-1200 m / z, resolution 60,000 (m / z 200), AGC 6 × 10⁻⁶. 5The maximum ion implantation time was 50 ms. The MS / MS resolution was 30,000 (m / z 200), the isolation window was 1.0 Da, the collision energy was 30%, and the AGC was 2 × 10⁻⁶. 5 The maximum ion implantation time is 80 ms.

[0111] 2.3 Analysis of PRM preliminary mass spectrometry data to screen representative polypeptide clusters at amino acid sites

[0112] Data were analyzed using SpectroDive v12.1 (Biognosys) software. Based on the preliminary PRM experimental results, the target peptides were further screened. The screening criteria were as follows: a. Targetable by PRM; b. Good peak shape with no interference; c. High abundance. Finally, 9 amino acid sites and their representative peptide clusters were selected for further analysis: PLXDC2

[90] _C, C3

[1320] _N, CDH1

[152] _C, AHSG

[339] _C, SRGN

[130] _N, SRGN

[72] _N, IGF2

[126] _C, APOC3

[21] _N, and ITIH4

[668] _C. Multiple peptides with diagnostic potential at the corresponding amino acid sites were preferentially selected as representative peptide clusters; when no peptides with diagnostic potential were found, representative peptide clusters were selected based on abundance.

[0113] 3. Establish a targeted analysis method to calculate the individual aa-score.

[0114] 3.1 Establish a targeted analysis method for the formal experiment, targeting the cohort sample and reference sample.

[0115] A targeted analysis method was established using a mixed sample. 10×iRT standard peptide (Biognosys) was added to the mixed sample for unscheduled PRM analysis, thereby establishing a scheduled PRM targeted quantitative analysis method that includes the m / z of the endogenous peptide precursor ion and the retention time of the precursor ion for targeted analysis. The retention time window was set to ±2.5 minutes.

[0116] Based on the aforementioned targeting method, scheduled PRM data collection for queue samples and references was formally conducted. When editing the sample list, reference samples were collected for the first and last shots, and an additional reference sample was interspersed between each group in the queue.

[0117] 3.2 PRM data analysis: Calculation of individual aa-score values ​​for amino acid sites based on reference samples.

[0118] Data was processed in SpectroDive software to exclude daughter ions with significant interfering signals. For each parent ion, at least three daughter ion pairs were selected for quantitative analysis. The data was then exported and further processed in R language.

[0119] Based on the average abundance of endogenous peptides in the reference samples, the relative content of the corresponding peptides in a single sample was calculated, and the individual aa-score of the representative peptide clusters at the amino acid sites was further calculated. The representative peptide clusters involved in the individual aa-score calculation of the three amino acid sites PLXDC2

[90] _C, CDH1

[152] _C, and IGF2

[126] _C are shown in Table 1.

[0120] Preferably, the representative polypeptide clusters at the two amino acid sites PLXDC2

[90] _C and IGF2

[126] _C are single polypeptides with diagnostic potential.

[0121] The differences in individual aa-scores of amino acid sites across different groups were calculated using the two independent samples Wilcoxon test. Figure 5 A). The three amino acid sites PLXDC2

[90] _C, CDH1

[152] _C, and IGF2

[126] _C showed statistically significant differences between the β-thalassemia group and the healthy control group, with p-values ​​of 5.97e-13, 1.77e-12, and 5.97e-13, respectively. The diagnostic potential of the individual aa-score of the amino acid site was calculated using the ROC curve analysis method in the pROC package of R language. Sequence B), PLXDC2

[90] _C had an AUC of 100% (95% CI: 1-1, p<0.05), with both specificity and sensitivity of 100%, and the optimal cutoff value for relative content was 0.4466; CDH1

[152] _C had an AUC of 99.8% (95% CI: 0.988-1, p<0.05), with specificity of 100.0% and sensitivity of 96.9%, and the optimal cutoff value for relative content was 0.0010; IGF2

[126] _C had an AUC of 100% (95% CI: 1-1, p<0.05), with both specificity and sensitivity of 100.0%, and the optimal cutoff value for relative content was 18.3369.

[0122] Table 1. Peptide clusters involved in individual aa-score calculations for amino acid sites of PLXDC2

[90] , CDH1

[152] _C, and IGF2

[126] when analyzed based on reference samples.

[0123] Gene name gene position aa position Whether the single polypeptide has diagnostic potential DTNRASVGQDSPEPR PLXDC2 PLXDC2[76-90] PLXDC2

[90] _C Yes VDTNRASVGQDSPEPR PLXDC2 PLXDC2[75-90] PLXDC2

[90] _C No AVDTNRASVGQDSPEPR PLXDC2 PLXDC2[74-90] PLXDC2

[90] _C No FLKAVDTNRASVGQDSPEPR PLXDC2 PLXDC2[71-90] PLXDC2

[90] _C Yes SGIQAELLTFPNSSPGLRRQ CDH1 CDH1[133-152] CDH1

[152] _C Yes SVSGIQAELLTFPNSSPGLRRQ CDH1 CDH1[131-152] CDH1

[152] _C Yes DTWKQSTQRL IGF2 IGF2[117-126] IGF2

[126] _C Yes Q[Gln->pyro-Glu]YDTWKQSTQRL IGF2 IGF2[115-126] IGF2

[126] _C No YDTWKQSTQRL IGF2 IGF2[116-126] IGF2

[126] _C Yes FFQYDTWKQSTQRL IGF2 IGF2[113-126] IGF2

[126] _C No QYDTWKQSTQRL IGF2 IGF2[115-126] IGF2

[126] _C No Figure 6

[0124] Example 3: The individual aa-score strategy based on relabeled peptides was used to verify amino acid sites with potential for screening and diagnosis of β-thalassemia, and representative peptide clusters were identified.

[0125] The recruitment, enrollment, and isolation and enrichment of plasma peptide samples for β-thalassemia patients in this embodiment are the same as in the previous embodiment.

[0126] 1. Targeted quantitative peptidomics (PRM) technology was used to target and screen peptides contained in amino acid sites with potential for screening and diagnosing β-thalassemia. Relabeled peptides were synthesized from representative peptide clusters at these amino acid sites. ROC curve analysis was used to determine the diagnostic performance of representative peptide cluster markers at protein amino acid sites.

[0127] 1. Preparation of peptide samples from targeted analysis cohort plasma and mixed plasma

[0128] This cohort included 49 participants, comprising 16 healthy controls, 17 carriers of mild β-thalassemia, and 16 patients with severe β-thalassemia. Specifically, plasma samples from the cohort were mixed in equal volumes, and peptides were extracted using the SPD method for use in preliminary PRM experiments and standard curve plotting experiments.

[0129] 2. PRM preliminary screening of amino acid sites and representative polypeptide clusters for targeted analysis.

[0130] 2.1 Establishment of PRM Pre-experiment Methods

[0131] The method was established using SpectroDive v12.1 (Biognosys) software. When developing the PRM detection method, the Pulsar engine was used to retrieve raw data from the entire peptide genome, generating a spectral library. Key search parameters were set as follows: non-enzymatic digestion; peptide length 7-40; modification: variable modification: Gln->pyro-Glu, Oxidation (P, M). Amino acid sites with diagnostic potential and their contained peptides were included in the PRM preliminary analysis list as much as possible, with the following parameter settings: precursor ion mass-to-charge ratio range, 350-1500; precursor ion charge number, 2-6; peptide length, 8-40; daughter ion mass-to-charge ratio range, 300-1800; maximum daughter ion charge number, 3; ion type, b, y ions; allowed neutral loss types, H2O, NH3, and no loss; top 6 daughter ions were selected. Unscheduled PRM was selected to establish the preliminary experimental method.

[0132] 2.2 PRM Pre-Experiment Data Acquisition

[0133] The LC-MS / MS instrument used in the parallel reaction monitoring (PRM) targeted quantitative analysis stage was an Orbitrap Eclipse Tribrid mass spectrometer (Thermo Fisher Scientific) equipped with an EASY-nLC1200 HPLC system. The chromatographic column and flow were the same as in step 2.2. The chromatographic elution gradient was set as follows: 5-10%, 3 min; 10-20% B, 22 min; 20-30% B, 22 min; 30-40% B, 13 min; 40-99% B, 4 min; 95% B, 9 min. The mass spectrometry parameters for PRM data acquisition were set as follows: full MS mass-to-charge ratio scan range 350-1200 m / z, resolution 60,000 (m / z 200), AGC 6 × 10⁻⁶. 5 The maximum ion implantation time was 50 ms. The MS / MS resolution was 30,000 (m / z 200), the isolation window was 1.0 Da, the collision energy was 30%, and the AGC was 2 × 10⁻⁶. 5 The maximum ion implantation time is 80 ms.

[0134] 2.3 Analysis of PRM preliminary experimental mass spectrometry data to screen representative polypeptide clusters at amino acid sites and synthesize relabeled polypeptides

[0135] Data were analyzed using SpectroDive v12.1 (Biognosys) software. Based on the preliminary PRM experimental results, the target peptides were further screened. The screening criteria were as follows: a. Targetable by PRM; b. Good peak shape with no interference; c. High abundance. Finally, 9 amino acid sites and their representative peptide clusters were selected for further analysis: PLXDC2

[90] _C, C3

[1320] _N, CDH1

[152] _C, AHSG

[339] _C, SRGN

[130] _N, SRGN

[72] _N, IGF2

[126] _C, APOC3

[21] _N, and ITIH4

[668] _C. Multiple peptides with diagnostic potential at the corresponding amino acid sites were preferentially selected as representative peptide clusters; when no peptides with diagnostic potential were found, representative peptide clusters were selected based on abundance. For the peptides in the representative peptide clusters, corresponding relabeled peptides were synthesized.

[0136] 3. Establish a targeted analysis method to calculate the individual aa-score.

[0137] 3.1 Establish a targeted analysis method for formal experiments, and conduct targeted analysis on the cohort samples.

[0138] Based on the abundance ratio of the target endogenous peptide in the preliminary experimental samples, a mixed labeled synthetic peptide was used as an internal standard. A targeted analysis method was established using the mixed sample. The internal standard and 10×iRT standard peptide (Biognosys) were added to the mixed sample for unscheduled PRM analysis. This established a scheduled PRM targeted quantitative analysis method that includes the m / z of the labeled peptide precursor ion, the m / z of the endogenous peptide precursor ion, and the retention time of the precursor ion for targeted analysis. The retention time window was set to ±2.5 minutes.

[0139] Based on the above method, the scheduled PRM data collection for the queue samples was formally carried out.

[0140] 3.2 Establish a standard curve for the relabeled peptides for absolute quantitative analysis.

[0141] Based on the expected abundance of the target peptide, serially diluted recalibrated peptide standards were prepared. The serially diluted recalibrated peptide standards and iRT peptides were added to the mixed sample for scheduled PRM analysis to establish a standard curve.

[0142] 3.3 PRM data analysis: Calculation of individual aa-score values ​​for amino acid sites using different methods

[0143] Data was processed in SpectroDive software to exclude daughter ions with significant interfering signals. For each parent ion, at least three daughter ion pairs were selected for quantitative analysis. The data was then exported and further processed in R language.

[0144] The representative polypeptide clusters involved in the individual aa-score calculations for the two amino acid sites PLXDC2

[90] _C and CDH1

[152] _C are shown in Table 2.

[0145] The calculation of individual aa-score based on relabeled peptides includes the following two methods: a. Calculate the absolute content of the corresponding peptide in a single sample based on the standard curve established by the relabeled peptide, thereby calculating the individual aa-score of the representative peptide cluster at the amino acid site; b. Calculate the relative content of the corresponding peptide in a single sample based on the relabeled peptide mixture standard, i.e., the internal control standard, thereby calculating the individual aa-score of the representative peptide cluster at the amino acid site.

[0146] Based on the calculation method of method a, the differences in individual aa-score of amino acid sites in different groups were calculated using the two independent samples Wilcoxon test. Figure 6A). The two amino acid sites PLXDC2

[90] _C and CDH1

[152] _C showed statistically significant differences between the β-thalassemia group and the healthy control group, with p-values ​​of 5.97e-13 and 6.21e-12, respectively. The diagnostic potential of the individual aa-score of the amino acid site was calculated using the ROC curve analysis method in the pROC package of R language. Figure 7 B), PLXDC2

[90] _C had an AUC of 100% (95% CI: 1-1, p<0.05), with both specificity and sensitivity at 100%, and an optimal cutoff value for absolute content of 0.5346 fmol / µL; CDH1

[152] _C had an AUC of 99.4% (95% CI: 0.977-1, p<0.05), with specificity at 100.0%, sensitivity at 96.9%, and an optimal cutoff value for absolute content of 6.2011 fmol / µL. (Optimal cutoff value obtained by maximizing the Youden index)

[0147] Based on the calculation method of method b, the differences in individual aa-scores of amino acid sites in different groups were calculated using the two independent samples Wilcoxon test. Figure 7 A). The two amino acid sites PLXDC2

[90] _C and CDH1

[152] _C showed statistically significant differences between the β-thalassemia group and the healthy control group, with p-values ​​of 5.97e-13 and 8.87e-13, respectively. The diagnostic potential of the individual aa-score of the amino acid site was calculated using the ROC curve analysis method in the pROC package of R language. Sequence B), PLXDC2

[90] _C had an AUC of 100% (95% CI: 1-1, p<0.05), with both specificity and sensitivity at 100%, and the optimal cutoff value for relative content was 0.0191; CDH1

[152] _C had an AUC of 100% (95% CI: 1-1, p<0.05), with both specificity and sensitivity at 100%, and the optimal cutoff value for relative content was 0.0745.

[0148] Table 2. Peptide clusters involved in the individual aa-score calculation of PLXDC2

[90] _C and CDH1

[152] _C amino acid sites when analyzed based on relabeled peptides.

[0149] Gene name gene position aa position DTNRASVGQDSPEPR PLXDC2 PLXDC2[76-90] PLXDC2

[90] _C FLKAVDTNRASVGQDSPEPR PLXDC2 PLXDC2[71-90] PLXDC2

[90] _C SGIQAELLTFPNSSPGLRRQ CDH1 CDH1[133-152] CDH1

[152] _C SVSGIQAELLTFPNSSPGLRRQ CDH1 CDH1[131-152] CDH1

[152] _C Figure 8

[0150] Example 4: Discovery of amino acid sites with diagnostic potential for colorectal cancer

[0151] The colorectal cancer patients and healthy participants involved in this invention were recruited by the First Affiliated Hospital of Henan University. The sample collection and clinical testing involved in this embodiment have all been approved by the Biomedical Science and Research Ethics Committee of Henan University, and all participants provided written informed consent.

[0152] This embodiment is based on the plasma whole peptidomome. The relative abundance of identified peptides is calculated using reference samples, and site clustering is performed on the peptides to calculate the individual aa-score of amino acid sites. Conventional omics data analysis and statistical methods, such as volcano plots, PCA analysis, and ROC analysis, can be used to analyze and mine the obtained amino acid site data, thereby discovering disease-related protein amino acid sites with diagnostic potential and their contained peptide clusters. This screening process includes the following aspects:

[0153] 1. Whole-peptide genome study population cohort

[0154] A total of 57 participants were included, comprising 29 healthy controls (Ctr) and 28 colorectal cancer patients (CRC). Specifically, plasma samples from the cohort were mixed in equal volumes, with 50 µL / tube prepared as five aliquots. After peptide enrichment and extraction, these aliquots served as reference samples for individual aa-score calculation.

[0155] 2. Isolation and enrichment of peptides in plasma samples

[0156] 2.1 Add DTT to 50 µL of plasma to a final concentration of 20 mM, and then incubate at 4°C for 1 h.

[0157] 2.2 Continue by adding IAM to the plasma to a final concentration of 40 mM, and then incubating at 4°C in the dark for 30 min.

[0158] 2.3 After incubation, add 250 µL of methyl tert-butyl ether (MTBE), 50 µL of deionized water and 150 µL of methanol to the sample in sequence, mix well and let stand at 4°C for 30 min.

[0159] 2.4 Centrifuge the above sample at 21,000 g, 4℃ for 30 min, and remove the supernatant after centrifugation.

[0160] 2.5 Add 500 µL MTBE and 100 µL deionized water to the supernatant above, mix well, and centrifuge at 1,000 g and 4 °C for 10 min.

[0161] 2.6 The above samples were divided into two phases. The upper layer, rich in hydrophobic interfering substances such as lipids, was removed, and the lower clear liquid was retained.

[0162] 2.7 Add 400µL of the upper phase of MTBE / MeOH / H2O (5:1:1, v / v) solvent to the hydrophilic phase containing the peptide, mix, and then perform a second defatting. Keep the lower clear liquid and dry it at 4℃.

[0163] 2.8 The obtained samples are processed by a desalting column and then ready for testing.

[0164] 3. Detection of the whole peptide genome using liquid chromatography-mass spectrometry (LC-MS / MS)

[0165] 3.1 Data acquisition was performed using an Orbitrap Exploris 480 mass spectrometer equipped with FAIMS Pro.

[0166] The LC-MS / MS instrument used was an OrbitrapExploris 480 mass spectrometer (Thermo Fisher Scientific) equipped with an EASY-nLC 1200 HPLC system and FAIMS Pro. The peptide was dissolved in 15 µL of 0.1% FA aqueous solution, centrifuged at 21,000 g at 4°C for 30 min, and then 10 µL of the supernatant was transferred to a sample vial for loading. The analytical liquid chromatography and mass spectrometry conditions were as follows: A Dr. Maisch GmbH ReproSil-Pur C18 AQ column (75 μm id × 20 cm, 3 μm) was used; the mobile phase was: A: water containing 0.1% formic acid, B: acetonitrile containing 0.1% formic acid; the flow rate was 300 nL / min; gradient elution was used: 4–11% B, 4 min; 11–21% B, 28 min; 21–30% B, 29 min; 30–42% B, 27 min; 42–95% B, 5 min; 95% B, 10 min. Two compensation voltages (-45 V and -65 V) were used for FAIMS separation. In data-dependent acquisition (DDA) mode, MS1 scan range: 350-1600 m / z, normalized AGC: 300%; maximum injection time: 80 ms; MS1 resolution: 60,000 (m / z 200); isolation window width: 1.6 m / z; HCD collision energy: 28%. MS / MS resolution: 60,000 (m / z 200); normalized AGC: 150%; maximum ion injection time: 118 ms; dynamic exclusion time: 30 s.

[0167] 3.2 Protein Discoverer 2.4 software was used to retrieve mass spectrometry data.

[0168] Mass spectrometry data were retrieved to obtain quantitative information on the whole peptide genome. Protein Discoverer 2.4 software was used, and the key parameters for data retrieval were set as follows: the database selected was the human database downloaded from Uniprot in September 2019; protease was set to No enzyme; the maximum errors for precursor and daughter ions were 10 ppm and 0.02 Da, respectively; variable modifications included oxidation of methionine and proline, cysteine ​​modification, and conversion of glutamine to pyroglutamic acid; the free dose ratio (FDR) for peptide level was set to 1%; label-free quantification was performed based on chromatographic area.

[0169] 4. Analysis of detection results and discovery of amino acid sites and amino acid clusters with diagnostic potential.

[0170] 4.1 Calculate the individual aa-score value of amino acid sites using reference samples.

[0171] The quantitative values ​​of peptides identified from five reference samples were averaged, and the relative abundance of peptides in each cohort was calculated based on this. The amino acid site classification and individualaa-score calculation of the peptides were performed using the R package aascore, thereby obtaining the quantitative data of amino acid sites.

[0172] 4.2 Site difference analysis and routine omics statistical analysis revealed amino acid sites and their amino acid clusters with diagnostic potential.

[0173] Data analysis and mining were performed using standard omics statistical methods. First, the p-value for comparison between the two groups was calculated using the t-test. The criteria for differentially expressed loci were: fold change > 2 and p < 0.05 for upregulated loci, and fold change < 0.5 and p < 0.05 for downregulated loci. The UMAP analysis results are as follows: Figure 8 As shown in Figure A. With at least 80% quantitative data in each group, ROC curve analysis was performed on the differentially expressed sites. Figure 9 B) Using volcano plots to display amino acid sites with AUC>90%, eight amino acid sites with AUC=100 were selected as sites with diagnostic potential.

[0174] Based on at least three quantitative values ​​per group, amino acid site analysis based on individual Aa-score and conventional peptidomics analysis (i.e., single-peptide analysis) were compared. In this dataset, 1766 amino acid sites were involved in the differential analysis, representing 2327 peptides; while conventional peptidomics analysis involved 1776 single peptides. This demonstrates that the individual Aa-score-based analysis strategy significantly improved the utilization of peptidomics data, significantly reduced data redundancy, and minimized the impact of missing values. ​ ).

[0175] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A single-site polypeptide clustering method based on the aa-score method, characterized in that, Includes the following steps: 1) Collect sample data; 2) Perform traditional peptide differential analysis; 3) Perform single-site cluster analysis on differentially expressed peptides and peptides with diagnostic potential: Extract the protein sites corresponding to the N and C ends of all differentially expressed peptides or peptides with diagnostic potential. When the number of peptides under a certain site is ≥2, it is a stable amino acid site; the remaining sites, i.e., when the number of peptides under a site is 1, are dynamic amino acid sites. Stable amino acid sites and dynamic amino acid sites are collectively referred to as variable sites; 4) Calculate the aa-score value of amino acid sites based on the grouped aa-score method or the individual aa-score method, identify amino acid sites that are related to diseases or have diagnostic potential, and achieve visual analysis; The specific calculation method for calculating the aa-score value of an amino acid site based on the grouped aa-score method is to directly accumulate the fold change values ​​of all differentially expressed polypeptides contained in that site, where the fold change value can be positive or negative. The aa-score value of the amino acid site calculated based on the individual aa-score method is the direct sum of the relative or absolute quantitative values ​​of the polypeptide. 5) For amino acid sites with diagnostic potential, identify representative polypeptide clusters and use the individual aa-score method to verify the diagnostic potential of representative polypeptide clusters based on amino acid sites.

2. According to the method of claim 1, in step 4), the individual aa-score value of the amino acid site is calculated using the reference sample, the amino acid site is subjected to differential analysis, and then the differential amino acid site is analyzed by ROC curve analysis to directly discover amino acid sites with diagnostic potential.

3. The method of claim 1, wherein the positive and negative values ​​of the grouped aa-score represent the enzyme activity in an activated or inhibited state at the site.

4. According to the method of claim 1, when the ratio of the average or median of peptide quantification in the experimental group to the average or median of peptide quantification in the control group is >1, fold change = ratio; When the ratio of the mean or median of peptide quantification in the experimental group to the mean or median of peptide quantification in the control group is less than 1, the fold change is -1 / ratio.

5. According to the method of claim 1, the grouped aa-score value of protein amino acid sites is calculated, and based on the changing trend of the slope of the cumulative grouped aa-score value, inflection sites, i.e., important amino acid sites related to the disease, are extracted from the variable sites, thereby indirectly discovering amino acid sites with diagnostic potential; among the important amino acid sites, the higher the normalized grouped aa-score value or the higher the grouped aa-score value, the higher its diagnostic potential; at the same time, the intersection sites of important amino acid sites and the stable sites to which the polypeptides with diagnostic potential belong also have higher diagnostic potential.

6. According to the method of claim 1, the grouped aa-score value of the protein amino acid sites is calculated, and the turning points are extracted from the variable sites based on the changing trend of the cumulative grouped aa-score slope. These sites are associated with diseases, have greater potential for biological function, and can be used for screening bioactive peptides.

7. The application of the method according to any one of claims 1-6 in constructing a peptidomics data analysis system.

8. The application of the method according to any one of claims 1-6 in screening disease diagnostic biomarkers, disease therapeutic targets, exploring enzyme-substrate relationships, and predicting bioactive peptides.