A method for analyzing and visualizing the plasma proteomic differences of individuals
By using multi-sample data correction and KEGG enrichment analysis, combined with the Somascan platform, the challenge of differential plasma proteomics analysis in personalized medicine has been solved, enabling personalized health monitoring and intervention effect evaluation, and providing personalized health change mapping and medication guidance.
Patent Information
- Application Number
- CN202411872887.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-12-18
AI Technical Summary
Existing technologies are insufficient for effectively analyzing differences in plasma proteomes between individuals in personalized medicine. Intergroup difference analysis cannot represent changes in individual health, and conventional methods cannot assess the effectiveness of interventions on individuals.
Multi-sample data correction methods such as hybridization normalization, median signal normalization, and plate-level normalization were employed. Differential protein analysis was performed using the Somascan platform, and influence scores were calculated and visualized through KEGG enrichment analysis and human function mapping to generate individualized analysis reports.
It enables precise analysis of changes in plasma proteomics before and after individual intervention, intuitively demonstrating the impact of intervention on human health and providing a basis for individual disease prevention and medication guidance.
Smart Images

Figure CN119993277B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of proteomics detection, and relates to an analysis method for mapping individual plasma protein changes to human health based on Somascan proteomics. BACKGROUND
[0002] Proteomics technology can be used to reveal the mechanism of various diseases, including cardiovascular disease, kidney disease, aging, respiratory disease, and neurological disease. Plasma proteomics mainly studies the composition and function of proteins in blood, which is of great significance for understanding human health and disease. Proteins in blood can provide important information about physiological status, metabolic processes, and disease risk. By analyzing the composition of proteins in blood, potential biomarkers can be identified for early diagnosis, disease monitoring, and treatment effectiveness evaluation. Blood collection as a convenient and non-invasive means enables plasma proteomics research to be more widely applied in clinical and large-scale population studies, providing convenience and feasibility for health management and disease monitoring. In addition, plasma proteomics research helps to reveal the pathogenesis of diseases, providing a basis for new drug development and personalized medicine. Therefore, studying plasma proteomics is of great significance for promoting medical science and clinical practice. Currently, mass spectrometry, olink, Somascan, and other platforms can detect plasma proteomics. However, the Somascan platform can measure up to 7k (v4.1) or 11K (v5.0) proteins in a small volume of biological sample (55ul of serum or plasma), enabling ultra-high abundance detection of plasma proteins and providing a protein detection range of 10log (fM-uM) with good reproducibility, with a median CV of less than 5%.
[0003] As a standard and conventional analysis method of proteomics, inter-group difference evaluation generally requires that the number of samples in each group be no less than 3. However, for individualized medicine or individual health monitoring in clinical practice, the applicability of inter-group difference analysis methods is limited. At the same time, due to genetic background, lifestyle, disease status, and other factors, plasma proteomes exhibit large individual differences, and inter-group differences are easily smoothed out by individual differences. Moreover, inter-group differences cannot represent the changes in individual health.
[0004] Defining individual health at the molecular level is one of the challenges in precision medicine. Individual plasma proteomes have unique and stable plasma protein profiles for up to 2 years, and have a strong correlation with clinical chemistry detection. In addition, conventional inter-group difference protein analysis can only show the differences between two groups before and after intervention, and cannot represent whether the intervention method is effective for the individual. SUMMARY
[0005] To solve the problems of the prior art, the present application provides:
[0006] A method for analyzing and visualizing the plasma proteomic differences of an individual, the method comprising the following steps
[0007] 1) Multiple sample data correction, preferably the correction method comprises hybrid normalization, median signal normalization, plate-level normalization, inter-plate calibration, which can be realized by experimental settings and algorithms; preferably, it is completed on the somalogic platform; further preferably, in order to further reduce the comparison bias, the pre-intervention and post-intervention samples are ensured to be in the same batch; preferably, the intervention is medication, exercise or dietary adjustment; further preferably, the medication is selected from small molecule drugs, antibody drugs, adoptive cell therapy (such as CAR-T or CAR-NK) or stem cell input, etc.
[0008] 2) Differential protein analysis, based on the data correction of step 1), the data measured before and after the intervention of the individual are directly compared to obtain differential protein data; preferably, the differential proteins are selected according to a specific fold change (|log2FC| >= 1.5);
[0009] 3) KEGG enrichment analysis based on the differential proteins of step 2), preferably, the top 20 pathways with a p_adjust value less than 0.05 are obtained as significant pathways;
[0010] 4) Mapping the functions of the human body in the KEGG database of 361 pathways, and counting the number of pathways in each mapping, and calculating the proportion of each mapping;
[0011] 5) Classifying the significant pathways obtained in step 3) according to the mapping obtained in 4), and calculating the influence value (escore) of each mapping, the escore calculation formula is as follows:
[0012] escore = ck / cr
[0013] Wherein ck (class kegg) represents the number of kegg in a certain mapping category of the top 20 pathways, and cr (class rate) represents the proportion of each mapping;
[0014] 6) Visualizing the mapping classification results of 5), first calculating the corresponding escore according to the mapping of step 5), then sorting the mapping according to the escore value, and finally visualizing the mapping classification information of the pathway through the mulberry chart; preferably, the KEGG pathway is sorted according to p_adjust.
[0015] In a further embodiment of the present application, the method further comprises:
[0016] 7) For the selected differential proteins, first classify according to the seven key pathways of the Somascan platform, the seven key pathways of the Somascan platform including cardiovascular disease related, inflammation and immune response related, metabolic disease related, cancer related, nerve related, cytokine related, and respiratory related proteins;
[0017] 8) According to the classification results of step 7), statistically analyze the up-regulation and down-regulation data of the differential proteins belonging to each classification and protein interaction mapping.
[0018] In further embodiments of the application, the human function mapping results in step 4) include immune, cardiovascular disease, neural function, endocrine function, carbohydrate metabolism, lipid metabolism, skin, muscle, kidney function, gene replication and repair function, liver detoxification function, glucose metabolism, alcohol metabolism, aging, bone and joint function, gonadal function, caffeine metabolism, stem cell homing or antioxidant function; preferably, the mapping classification used in the application is shown in Table 1.
[0019] In further embodiments of the application, step 9) is further included, which is to interpret the mapping level of the intervention effect based on the mapping results of 6).
[0020] In further embodiments of the application, step 10) is further included, which is to generate an individual systematic analysis report by comprehensively analyzing the above information.
[0021] The application also provides a computer element which stores a computer program for implementing any of the above methods.
[0022] The application also provides the method or the element for non-diagnostic purposes of evaluating the effect of intervention on human health.
[0023] Beneficial technical effects
[0024] The application provides a method for analyzing the plasma proteomics of an individual and visualizing the results, which maps the functions of 361 pathways of the human body, then classifies the significant pathways of intervention, and scores and ranks based on the mapping information, more directly reflects the effect of intervention on human health, and provides ideas for disease prevention, judgment and future medication guidance of individuals. BRIEF DESCRIPTION OF DRAWINGS
[0025] Figure 1 16 Somascan data analysis before and after human input stem cells - top 20 Keggpathway pathways enriched by differential proteins
[0026] Figure 2 . Data clustering before correction (top) and after correction (bottom) of different categories
[0027] Figure 3The relevant cardiovascular disease related differential protein map before and after the infusion of stem cells, there are 43 kinds of cardiovascular disease related differential proteins, of which 38 kinds of differential proteins are expressed by more than 1.5 times, and 5 kinds of differential proteins are expressed by less than 0.5 times
[0028] Figure 4 The relevant cardiovascular disease related differential protein interaction map before and after the infusion of stem cells of the individual in the embodiment of the application
[0029] Figure 5 The mapping result map of 361 human pathways in the application
[0030] Figure 6 The visualization of the top 20 pathway results in the embodiment of the application
[0031] Figure 7 The report composition in the embodiment of the application
[0032] Figure 8 The complete analysis process of the application DETAILED DESCRIPTION
[0033] As a standard and conventional analysis method of proteomics, the inter-group difference evaluation generally requires that the number of samples in each group is not less than 3. However, for individualized medical treatment or individual health monitoring in the clinic, the applicability of the inter-group difference analysis method is small. At the same time, due to genetic background, living habits, disease state and the like, the plasma proteome of individuals shows great individual difference, and the inter-group difference is easily smoothed by the individual difference. At the same time, the inter-group difference cannot represent the health change of the individual. By analyzing the individual plasma proteomic changes of 16 healthy people before and after the infusion of stem cells, it is also verified that the individual proteomic changes are quite different from others (such as Figure 1
[0034] Therefore, in order to support the definition of individual-based health, monitor the proteomic changes of individuals before and after health intervention, and systematically display the intervention effect, the application designs a plasma proteomic analysis method for individuals, and maps the functional annotation of the differential proteins to the human health changes, which more directly reflects the influence direction of the proteome on the human body.
[0035] The specific analysis method of the application comprises the following steps:
[0036] 1) Multi-sample data correction Before individualized differential protein analysis, the data needs to be corrected in multiple ways. The preferred correction methods include hybrid standardization, median signal standardization, plate-level standardization, and inter-plate calibration. These correction methods can be realized through experimental settings and algorithms; preferably, it is completed on the somalogic platform, and the correction effect is as shown in Figure 2 It is also preferred that, in order to further reduce the comparison bias, the pre-intervention and post-intervention samples are guaranteed to be on the same batch machine;
[0037] 2) Differential protein analysis, based on the data correction of step 1), directly compare the data measured before and after the individual, obtain differential protein data; preferably, according to the specific fold difference (|log2FC| >= 1.5) to select differential proteins;
[0038] 3) Kegg enrichment analysis based on differential proteins, preferably, through R package clusterprofiler to complete;
[0039] 4) Mapping the functions of the human body to the 361 pathways in the KEGG database, preferably, the mapping is obtained through literature query and / or module information in the KEGG database; count the number of pathways in each mapping, and calculate the percentage, the number of pathways and the percentage are shown as Figure 5 ;
[0040] 5) For the top 20 pathways with p_adjust value less than 0.05 obtained in step 3), classify according to the mapping of step 4), and calculate the influence value (escore), the escore calculation formula is as follows:
[0041] escore = ck / cr
[0042] Where ck (class kegg) represents the number of kegg in a certain mapping category of the top 20 pathways, and cr (class rate) represents the proportion of each mapping;
[0043] 6) Visualize the results of step 5), first calculate the corresponding escore according to the mapping of step 5), then sort the mapping according to the escore value, and finally visualize the mapping information through Sankey diagram (Sankey diagram), the KEGG pathway is arranged according to p_adjust, the results are shown as Figure 6 .
[0044] The method further comprises the following steps:
[0045] 7) For the selected differential proteins, first according to the seven key pathways of the Somascan platform, the seven key pathways of the Somascan platform include cardiovascular disease related, inflammation and immune response related, metabolic disease related, cancer related, neural related, cytokine related, and respiratory related proteins. The seven categories come from the seven panels designed by the Somascan platform according to 1178 papers for proteins in different directions. The specific literature can be viewed at https: / / somalogic.com / publications / .
[0046] 8) According to the classification results of step 3), statistical analysis of the up-regulation and down-regulation data of the differential proteins belonging to each category and protein interaction mapping are performed. The results are shown in Figure 3 and Figure 4 ;
[0047] The method further comprises the following steps:
[0048] 9) Based on the mapping results of step 8), the intervention effects are interpreted at the mapping level;
[0049] 10) Generating a system individualized analysis report.
[0050] Technical terms
[0051] SomaScan protein sequencing platform
[0052] The SomaScan protein sequencing platform discovers new disease biomarkers through high-throughput protein analysis technology. Currently, more than 11000 proteins can be detected in a single ultra-micro sample. The detection technology covers sample types such as blood, urine, and body fluids, and the sample sources include humans and other model animals. Only 55 microliters of sample are required to achieve protein detection, with a dynamic detection range of 10log, and it can detect very low abundance proteins.
[0053] The seven categories described herein come from the seven panels designed by the SomaScan platform according to 1178 papers for proteins in different directions. The specific literature can be viewed at https: / / somalogic.com / publications / .
[0054] Protein interaction network diagram
[0055] Protein interaction network, namely Protein-protein interaction (PPI) network, is an analysis method for describing the interaction relationship between a group of proteins. Because the interaction relationship between proteins is complex, each protein is a node, and the interaction between multiple proteins is connected into a line to form a network structure, so it is called protein interaction network.
[0056] Pathway enrichment analysis
[0057] Pathway enrichment analysis: focus on checking whether genes / metabolites in a gene set are enriched in a specific pathway. Enrichment analysis is performed by information of known biological pathways provided by KEGG. The specific calculation method is, for example:
[0058] 1) provide the number of differentially expressed genes n;
[0059] 2) calculate the number of differentially expressed genes k from pathway A;
[0060] 3) read the number of genes M from the KEGG background file;
[0061] 4) read the number of genes N in pathway A;
[0062] 5) calculate the enrichment fold (Fold Enrichment) or enrichment score, the formula is: EnrichmentScore=(k / n) / (N / M);
[0063] 6) calculate the p value to evaluate the significance of enrichment, such as hypergeometric distribution test and Fisher's exact test.
[0064] In the present application, KEGG enrichment analysis is performed by clusterProfiler R package.
[0065] In the present application, significant pathways are usually selected according to the corrected p value (p.adjust) of the enrichment result, and generally p.adjust<0.05 can be considered as significant enrichment, and the enrichment results of different studies are different.
[0066] Human function mapping
[0067] In the present application, human function mapping refers to mapping 361 pathways in KEGG database according to immunity, cardiovascular disease, neural function, endocrine function, etc. (see Table 1 and Figure 5 ) by literature query and / or module information in KEGG database, finally producing 20 mapping categories, and counting the number and proportion of pathways involved in each mapping.
[0068] In the present application, the significant pathways are further classified according to the above mapping, and the influence value (escore) is calculated, the escore calculation formula is as follows:
[0069] escore=ck / cr
[0070] wherein ck (class kegg) represents the number of significant pathways in each mapping category after the significant pathways are classified according to the above mapping, and cr (class rate) represents the proportion of each mapping, for example, 20 kegg pathways are obtained by enrichment analysis after intervention of a certain individual, after the above mapping classification, 3 pathways belong to muscle function, then ck = 3, and the proportion cr of muscle function mapping is 0.036, then the influence value escore of muscle function mapping is 83.33. The influence value (escore) introduced by the present application reflects the influence degree of intervention on different physiological functions of the individual, and the value is adjusted by the proportion of mapping, so that the core function affected by the intervention can be more accurately determined.
[0071] Table 1 Classification results of human function mapping
[0072]
[0073]
[0074]
[0075]
[0076]
[0077]
[0078]
[0079] Embodiment
[0080] The seven categories of differential proteins of a certain individual before and after infusion of stem cells are analyzed according to the method of the present application, the statistical graph of cardiovascular related differential proteins is shown in Figure 3 , the PPI interaction diagram is shown in Figure 4 , and the visualization result of the KEGG PATHWAY top20 pathway result of the differential proteins is shown in Figure 6 .
[0081] The above experimental results of the present application can form a Somascan proteomics detection report, the directory of the report is shown in Figure 7 , and the overall analysis process of the present application is shown in Figure 8 .
Claims
1. A method for analyzing and visualizing plasma proteomic differences of an individual, the method comprising the following steps: 1) data correction before and after intervention of the individual; 2) differential protein analysis, based on the data correction of step 1), comparing the data measured before and after intervention of the individual to obtain differential protein data; 3) KEGG enrichment analysis based on the differential proteins of step 2) to obtain significant pathways, and the pathways with a p_adjust value less than 0.05 and ranking in the top 20 are taken as significant pathways; 4) mapping human functions to the 361 pathways in the KEGG database, and counting the number of pathways in each mapping to calculate the proportion of each mapping; the human function mapping results include immunity, cardiovascular disease, neural function, endocrine function, carbohydrate metabolism, lipid metabolism, skin, muscle, kidney function, gene replication and repair function, liver detoxification function, glucose metabolism, alcohol metabolism, aging, bone and joint function, gonadal function, caffeine metabolism, stem cell homing or antioxidant function; 5) classifying the significant pathways obtained in step 3) according to the mapping of step 4), and calculating the influence value escore of each mapping, the escore calculation formula is as follows: wherein ck represents the KEGG number of a certain mapping category in the top 20 pathways, and cr represents the proportion of each mapping; ; 6) visualizing the mapping classification results of step 5), first calculating the corresponding escore according to the mapping of step 5), then sorting the mapping according to the escore value, and finally visualizing the mapping classification information of the pathways by a Sankey diagram, and the KEGG pathway sorting is arranged according to p_adjust. 2.The method of claim 1, further comprising: 7) for the differential proteins selected in step 2), first classifying according to the seven key pathways of the Somascan platform, the seven key pathways of the Somascan platform including cardiovascular disease related, inflammation and immune response related, metabolic disease related, cancer related, neural related, cytokine related, and respiratory related proteins; 8) according to the classification results of step 7), statistically analyzing the up-regulation and down-regulation data of the differential proteins in each classification and protein interaction mapping. 3.The method of claim 2, further comprising step 9), interpreting the intervention impact based on the mapping results of step 6). 4.The method of claim 3, further comprising step 10), generating an individual systematic analysis report by integrating information. 5.The method of any one of claims 1-4, wherein the intervention is medication, exercise or diet adjustment. 6.The method of claim 5, wherein the medication is selected from small molecule drugs, antibody drugs, adoptive cell therapy or stem cell input. 7.The method of any one of claims 1-4, wherein the correction method comprises hybrid normalization, median signal normalization, plate-level normalization, and inter-plate calibration, and the correction method is completed on the Somalogic platform, and the samples before and after intervention are subjected to the same batch. 8. The method of any one of claims 1-4, wherein step 2) selects differential proteins based on a specific fold change |log2FC| >= 1.
5.
9. A computer device, said device storing a computer program implementing the method of any one of claims 1-4.
10. The method of any one of claims 1-4 or the device of claim 9 for non-diagnostic use in assessing the effect of an intervention on human health.
Citation Information
Patent Citations
Protein spectrum-based spontaneous premature delivery risk prediction model and construction method thereof
CN117649939A
Method for determining efficacy of cordyceps sinensis based on system biological response
CN118150722A