Individual plasma proteomics difference analysis and result visualization method

By correcting and analyzing individual plasma protein group data, combining KEGG pathway enrichment and functional mapping, the influence value is calculated and visualized, the shortcomings of individual plasma proteome differential analysis in the prior art are solved, and the accurate mapping and intuitive reflection of the changes in individual health are achieved.

CN119993277AActive Publication Date: 2025-05-13LOTUSLAKE BIOMEDICAL TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411872887.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-13
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively analyze the differences in individual plasma proteomes, cannot accurately represent changes in individual health, and the inter-group difference analysis method is of little applicability.

Method used

Through steps such as multi-sample data correction, differential protein analysis, KEGG enrichment analysis, human function mapping and influence value calculation, individual plasma proteomic differential analysis is achieved, and the results are visualized through Sanchima.

Benefits of technology

It can intuitively reflect the impact of intervention on individual health, providing ideas for individual disease prevention, judgment and future medication guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993277A_ABST
    Figure CN119993277A_ABST
Patent Text Reader

Abstract

According to the individual plasma proteomics difference analysis and result visualization method provided by the invention, function mapping is carried out on 361 paths of a human body, and mapping information is scored and sorted, so that the influence of intervention on human health is more intuitively reflected, and a thought is provided for individual disease prevention and judgment, future medication guidance and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of proteomics detection, and relates to an analytical method for mapping the corresponding human health to individual plasma protein changes based on Somascan proteomics monitoring. Background Art

[0002] Proteomics technology can be used to reveal the mechanisms of a variety of diseases, including cardiovascular disease, kidney disease, aging, respiratory diseases, neurological diseases, etc. Plasma proteomics mainly studies the composition and function of proteins in the blood, which is of great significance for understanding human health and disease. Proteins in the blood can provide important information such as physiological state, metabolic processes and disease risks. By analyzing the composition of proteins in the blood, potential biomarkers can be identified for early diagnosis, disease monitoring and treatment effect evaluation. Blood collection, as a convenient and non-invasive means, enables plasma proteomics research to be more widely used in clinical and large-scale population research, and provides convenience and feasibility for health management and disease monitoring. In addition, plasma proteomics research can also help reveal the pathogenesis of diseases and provide a basis for new drug development and personalized medicine. Therefore, studying plasma proteomics is of great significance for promoting medical science and clinical practice. At present, mass spectrometry, olink, Somascan and other platforms can all detect proteomics in plasma. However, the Somascan platform can measure proteins ranging from 7k (v4.1) or 11K (v5.0) in small volumes of biological samples (serum plasma sample volume 55ul), achieving ultra-high abundance detection of plasma proteins and providing a protein detection range (fM-uM) spanning 10log with good reproducibility and median CV within 5%.

[0003] As a standard routine analysis method in proteomics, inter-group difference evaluation generally requires that the number of samples in each group should be no less than 3. However, the inter-group difference analysis method has little applicability in clinical personalized medicine or individual health monitoring. At the same time, due to genetic background, living habits, disease status, etc., the plasma proteome shows large individual differences, and the inter-group differences are easily smoothed out by the inter-individual differences. At the same time, the inter-group differences cannot represent the health changes of individuals.

[0004] Defining individual health at the molecular level is one of the challenges in the field of precision medicine. Individual plasma proteomes have unique and stable plasma protein profiles for up to 2 years and have a strong correlation with clinical chemistry tests. In addition, conventional analysis of intergroup differential proteins can only explain the differences between the two groups before and after the intervention, and cannot represent whether the intervention is effective for the individual. Summary of the invention

[0005] In order to solve the problems of the prior art, the present invention provides:

[0006] A method for analyzing individual plasma proteomics differences and visualizing the results, the method comprising the following steps

[0007] 1) Multi-sample data correction, preferably the correction methods include hybridization standardization, median signal standardization, plate-level standardization, and inter-plate calibration, which can be achieved through experimental settings and algorithms; preferably, it is completed on the somalogic platform; preferably, in order to further reduce the comparison bias, the samples before and after the intervention are guaranteed to be in the same batch; preferably, the intervention is medication, exercise or diet adjustment; more preferably, the medication is selected from small drug, antibody drug, adoptive cell therapy (such as CAR-T or CAR-NK) or stem cell infusion, etc.;

[0008] 2) differential protein analysis: based on the data correction in step 1), directly compare the data measured before and after the individual intervention to obtain differential protein data; preferably, differential proteins are selected based on a specific fold difference (|log2FC|>=1.5);

[0009] 3) Perform KEGG enrichment analysis based on the differentially expressed proteins in step 2), preferably obtaining pathways with a p_adjust value less than 0.05 and ranked in the top 20 as significant pathways;

[0010] 4) Perform human functional mapping on 361 pathways in the KEGG database, count the number of pathways in each mapping, and calculate the proportion of each mapping;

[0011] 5) The significant pathways obtained in step 3) are classified according to the mappings obtained in step 4), and the influence value (escore) of each mapping is calculated. The score calculation formula is as follows:

[0012] escore=ck / cr

[0013] Among them, ck (class kegg) represents the number of keggs belonging to a certain mapping type in the first 20 paths, and cr (classrate) represents the proportion of each mapping;

[0014] 6) Visualizing the mapping classification results of 5), firstly calculating the corresponding score according to the mapping of step 5), then sorting the mappings according to the score values, and finally visualizing the mapping classification information of the pathway through a mulberry diagram; preferably, the KEGG pathway ranking is arranged according to p_adjust.

[0015] In a further embodiment of the present invention, the method further comprises:

[0016] 7) For the selected differentially expressed proteins, they were first classified according to the seven key pathways of the Somascan platform, which include proteins related to cardiovascular disease, inflammation and immune response, metabolic disease, cancer, nerves, cytokines, and respiration;

[0017] 8) According to the classification results of step 7), statistical analysis of up-regulated and down-regulated data of differentially expressed proteins in each classification and protein interaction mapping were performed.

[0018] In a further embodiment of the present invention, the human body function mapping results in step 4) include immunity, cardiovascular disease, nervous function, endocrine function, carbohydrate metabolism, lipid metabolism, skin, muscle, kidney function, gene replication and repair function, liver detoxification function, blood sugar metabolism, alcohol metabolism, aging, bone and joint function, gonad function, caffeine metabolism, stem cell homing or antioxidant function; preferably, the mapping classification used in the present invention is as shown in Table 1.

[0019] In a further embodiment of the present invention, step 9) is further included, in which the intervention impact is interpreted at the mapping level based on the mapping result of 6).

[0020] In a further embodiment of the present invention, step 10) is also included to generate an individual systematic analysis report by integrating the above information.

[0021] The present invention also provides a computer component storing a computer program for implementing any one of the above methods.

[0022] The present invention also provides for the use of the method or the element for non-diagnostic purposes in assessing the effects of intervention on human health.

[0023] Beneficial technical effects

[0024] The present invention provides an individual plasma proteomics difference analysis and result visualization method, which maps the functions of 361 pathways in the human body, then classifies the intervention-significant pathways, and scores and sorts them based on the mapping information, which more intuitively reflects the impact of intervention on human health and provides ideas for individual disease prevention, judgment and future medication guidance. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 Somascan data analysis of 16 people before and after stem cell infusion - the top 20 Kegg pathways enriched with differential proteins Figure 2 Clustering of different types of data before (top) and after (bottom) correction

[0026] Figure 3The differential protein map related to cardiovascular disease before and after stem cell infusion. There are 43 differential proteins related to cardiovascular disease, of which 38 differential proteins were expressed more than 1.5 times, and 5 differential proteins were expressed less than 0.5 times.

[0027] Figure 4 Differential protein interaction diagrams related to cardiovascular diseases before and after stem cell infusion in an individual according to an embodiment of the present invention

[0028] Figure 5 Schematic diagram of the results of mapping 361 human pathways in this invention

[0029] Figure 6 Visualization of the first 20 pathway results in the embodiment of the present invention

[0030] Figure 7 The embodiment of the present invention reports the composition

[0031] Figure 8 The complete analysis process of the present invention DETAILED DESCRIPTION

[0032] As a standard routine analysis method of proteomics, intergroup difference evaluation generally requires that the number of samples in each group should be no less than 3. However, the intergroup difference analysis method is not applicable to clinical personalized medicine or individual health monitoring. At the same time, due to genetic background, living habits, disease status, etc., the plasma proteome of individuals shows large individual differences, which are easily smoothed out by inter-individual differences. At the same time, inter-group differences cannot represent individual health changes. By analyzing the changes in individual plasma proteomics of 16 healthy people before and after stem cell infusion, it was also verified that the individual proteomic changes were quite different from others (such as Figure 1 shown).

[0033] Therefore, in order to support the individual-based definition of health, monitor the proteomic changes before and after individual health interventions, and systematically display the intervention effects, the present invention designs an individual plasma proteomic analysis method, and maps the functional annotations of differential proteins to the changes in human health, which more intuitively reflects the direction of the impact of the proteome on the human body.

[0034] The specific analysis method of the present invention comprises the following steps:

[0035] 1) Multi-sample data correction: Before individual differential protein analysis, multiple corrections need to be performed on the data. The preferred correction methods include hybridization normalization, median signal normalization, plate-level normalization, and inter-plate calibration. These correction methods can be achieved through experimental settings and algorithms; preferably, they are completed on the somalogic platform. The correction effect is as follows: Figure 2. It is also preferred that, in order to further reduce the comparison bias, the samples before and after the intervention are ensured to be loaded on the machine in the same batch;

[0036] 2) differential protein analysis: based on the data correction in step 1), directly compare the data measured before and after the individual to obtain differential protein data; preferably, differential proteins are selected based on a specific fold difference (|log2FC|>=1.5);

[0037] 3) Perform KEGG enrichment analysis based on differentially expressed proteins, preferably using the R package clusterprofiler;

[0038] 4) Perform human function mapping on 361 human pathways in the KEGG database. Preferably, the mapping is obtained through literature query and / or module information in the KEGG database; count the number of pathways in each mapping and calculate the percentage. The number of pathways and percentage results are as follows: Figure 5 As shown;

[0039] 5) For the pathways with p_adjust values ​​less than 0.05 obtained in step 3) and ranked in the top 20, they are classified according to the mapping in step 4) and the influence value (escore) is calculated. The score calculation formula is as follows:

[0040] escore=ck / cr

[0041] Among them, ck (class kegg) represents the number of keggs belonging to a certain mapping type in the first 20 paths, and cr (classrate) represents the proportion of each mapping;

[0042] 6) Visualize the results of step 5). First, calculate the corresponding score according to the mapping of step 5). Then sort the mappings according to the score value. Finally, visualize the mapping information through a Sankey diagram. The KEGG pathway ranking is arranged according to p_adjust. The results are as follows: Figure 6 shown.

[0043] The method further comprises the steps of:

[0044] 7) For the selected differentially expressed proteins, they were first classified according to the seven key pathways of the Somascan platform, which include proteins related to cardiovascular disease, inflammation and immune response, metabolic disease, cancer, nerves, cytokines, and respiration. The seven major categories come from 7 panels designed by the Somascan platform for proteins in different directions based on 1,178 literature collections. The specific literature can be viewed at https: / / somalogic.com / publications / .

[0045] 8) According to the classification results of step 3), statistical analysis of the up-regulated and down-regulated data of differentially expressed proteins in each classification and protein interaction mapping were performed. The results are shown in Figure 3 and Figure 4 ;

[0046] The method further comprises the steps of:

[0047] 9) Based on the mapping results of step 8), interpret the impact of the intervention at the mapping level;

[0048] 10) Generate system individualized analysis report.

[0049] Technical terms

[0050] SomaScan Protein Sequencing Platform

[0051] The SomaScan protein sequencing platform discovers new disease biomarkers through high-throughput protein analysis technology. Currently, it can detect more than 11,000 proteins in ultra-trace samples at one time. The detection technology covers sample types including blood, urine, body fluids, etc. The sample sources include humans and other model animals. Only 55 microliters of sample are needed to detect proteins, and the dynamic detection range is 10log, which can detect extremely low-abundance proteins.

[0052] The seven major categories described in this article come from seven panels designed by the SomaScan platform based on 1,178 literature collections targeting proteins in different directions. Specific literature can be viewed at https: / / somalogic.com / publications / .

[0053] Protein interaction network diagram

[0054] Protein-protein interaction (PPI) network is an analytical method that describes the interaction relationship between a group of proteins. Because the interaction relationship between proteins is complex, each protein is a node, and the interaction between multiple proteins is connected into lines to form a network structure, so it is called a protein interaction network.

[0055] Pathway enrichment analysis

[0056] Pathway enrichment analysis: Focus on checking whether the genes / metabolites in the gene set are enriched in a specific pathway. Enrichment analysis is performed based on the information of known biological pathways provided by KEGG. Specific calculation methods include:

[0057] 1) Provide differentially expressed genes n;

[0058] 2) Calculate the number k of differentially expressed genes from pathway A;

[0059] 3) Read the number of genes M in the KEGG background file;

[0060] 4) Read the number of genes N in pathway A;

[0061] 5) Calculate the enrichment fold (Fold Enrichment) or enrichment score, the calculation formula is: Enrichment Score = (k / n) / (N / M);

[0062] 6) Calculate p-values ​​to assess the significance of enrichment, such as the hypergeometric test and Fisher's exact test.

[0063] In the present invention, KEGG enrichment analysis was performed using the clusterProfiler R package.

[0064] In the present invention, significant pathways are usually selected based on the corrected p-value (p.adjust) of the enrichment results. Generally, p.adjust < 0.05 can be considered as significant enrichment, and the results of enrichment in different studies vary.

[0065] Human body function mapping

[0066] In the present invention, human function mapping refers to mapping the 361 human pathways in the Kegg database according to immunity, cardiovascular disease, neural function, endocrine function, etc. (see Table 1 and Figure 5 ) is used for mapping, and finally 20 mapping categories are generated, and the number and proportion of pathways involved in each mapping are counted.

[0067] The present invention further classifies the significant pathways according to the above mapping and calculates the influence value (escore). The score calculation formula is as follows:

[0068] escore=ck / cr

[0069] Among them, ck (class kegg) represents the number of significant pathways in each mapping category after the pre-significant pathways are classified according to the above mapping, and cr (class rate) represents the weight of each mapping. For example, after an individual intervention, enrichment analysis obtains 20 kegg pathways. After the above mapping classification, 3 pathways belong to muscle function, then ck = 3, and the weight cr of muscle function mapping is 0.036, then the influence value score of muscle function mapping is 83.33. The present invention introduces an influence value (escore) to reflect the degree of influence of the intervention on different physiological functions of the individual, and adjusts this value through the mapping weight, so that the core function affected by the intervention can be determined more accurately.

[0070] Table 1 Classification results of human body function mapping

[0071]

[0072]

[0073]

[0074]

[0075]

[0076]

[0077]

[0078] Example

[0079] The seven major categories of differential proteins before and after an individual was infused with stem cells were analyzed according to the method of the present invention. The statistical graph of cardiovascular-related differential proteins is shown in FIG. Figure 3 As shown, the PPI interaction diagram is as follows Figure 4 As shown; the visualization results of the KEGG PATHWAY top20 pathway results of differentially expressed proteins are shown in Figure 6 .

[0080] The above experimental results of the present invention can be formed in a Somascan proteomics detection report, and the contents of the report are as follows: Figure 7 The overall analysis process of the present invention is shown in Figure 8 .

Claims

1. A method for analyzing individual plasma proteomics differences and visualizing the results, the method comprising the following steps: 1) Correction of individual data before and after intervention; 2) Differential protein analysis: based on the data correction in step 1), the differential protein data are obtained by comparing the data measured before and after the individual intervention; 3) Based on the differentially expressed proteins in step 2), KEGG enrichment analysis is performed to obtain significant pathways. Preferably, pathways with a p_adjust value (corrected p value) less than 0.05 and ranked in the top 20 are selected as significant pathways; 4) Perform human functional mapping on 361 human pathways in the KEGG database, count the number of pathways in each mapping, and calculate the proportion of each mapping; 5) The significant pathways obtained in step 3) are classified according to the mapping in step 4), and the influence value (escore) of each mapping is calculated. The score calculation formula is as follows: escore=ck / cr Among them, ck (class kegg) represents the number of KEGG entries belonging to a certain mapping type in the top 20 pathways, and cr (class rate) represents the proportion of each mapping; 6) Visualizing the mapping classification results of step 5), firstly calculating the corresponding score according to the mapping described in step 5), then sorting the mapping according to the score value, and finally visualizing the mapping classification information of the pathway through a Sankey diagram; preferably, the KEGG pathway ranking is arranged according to p_adjust.

2. The method of claim 1, further comprising: 7) For the differentially expressed proteins selected in step 2), they were first classified according to the seven key pathways of the Somascan platform, which include proteins related to cardiovascular disease, inflammation and immune response, metabolic disease, cancer, nerves, cytokines, and respiration; 8) According to the classification results of step 7), statistical analysis of up-regulated and down-regulated data of differentially expressed proteins in each classification and protein interaction mapping were performed.

3. The method according to claim 1, wherein the human function mapping results in step 4) include immunity, cardiovascular disease, nervous function, endocrine function, carbohydrate metabolism, lipid metabolism, skin, muscle, kidney function, gene replication and repair function, liver detoxification function, blood sugar metabolism, alcohol metabolism, aging, bone and joint function, gonad function, caffeine metabolism, stem cell homing or antioxidant function; preferably, the human function mapping is as shown in Table 1.

4. The method according to claim 2 further comprises step 9), performing a mapping-level interpretation of the intervention impact based on the mapping results of step 6).

5. The method according to claim 4 further comprises step 10), wherein the comprehensive information is used to generate an individual systematic analysis report.

6. The method according to any one of claims 1 to 5, wherein the intervention is medication, exercise or diet adjustment; further preferably, the medication is selected from small-drug, antibody drug, adoptive cell therapy (such as CAR-T or CAR-NK) or stem cell infusion, etc.

7. The method according to any one of claims 1 to 6, wherein the correction method comprises hybridization standardization, median signal standardization, plate-level standardization, and inter-plate calibration, and these correction methods are implemented through experimental settings and algorithms; preferably, they are completed on the somalogic platform; preferably, in order to further reduce the comparison bias, the samples before and after the intervention are guaranteed to be in the same batch.

8. The method according to any one of claims 1 to 7, wherein in step 2) differential proteins are selected according to a specific fold difference (|log2FC|>=1.5).

9. A computer component storing a computer program for implementing the method according to any one of claims 1 to 8.

10. Use of the method of any one of claims 1 to 8 or the element of claim 9 for non-diagnostic purposes of assessing the effects of an intervention on human health.

Citation Information

Patent Citations

  • Processing method of cow urine iTRAQ test data

    CN103163257A

  • Protein spectrum-based spontaneous premature delivery risk prediction model and construction method thereof

    CN117649939A

  • Method for determining efficacy of cordyceps sinensis based on system biological response

    CN118150722A

  • A computational method for mapping peptides to proteins using sequencing data

    EP2674758A1

  • Virtual inference of protein activity by regulon enrichment analysis

    US20170076035A1