Protein-centered multi-omics data analysis method

Through the multiomics data analysis method centered on protein, transcriptional regulation, post-translational modification regulation and metabolic regulation are integrated, and the protein importance score is calculated, which solves the problem that traditional analysis methods are difficult to find key molecular changes, and the judgment of the importance and authenticity of proteins is achieved.

CN120472979AInactive Publication Date: 2025-08-12THE AFFILIATED SIR RUN RUN SHAW HOSPITAL OF SCHOOL OF MEDICINE ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510555079.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional Multi-omic studies lack diversity at the post-translational modification omics level of proteins, resulting in one-sided analysis results and difficulty in finding key molecular changes data.

Method used

Provide a protein-centered multiomics data analysis method. By obtaining differential analysis data and prior knowledge, integrating transcriptional regulation, post-translational modification regulation, metabolic regulation and interaction sub-tables, calculating the omics type score and protein importance score, and screening molecular targets.

Benefits of technology

The judgment of the importance and authenticity of proteins is achieved, the complexity of differential analysis caused by the diversification of omics levels and data types is solved, and the comprehensiveness and accuracy of the analysis results are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472979A_ABST
    Figure CN120472979A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of molecular protein, and provides a protein-centered multi-omics data analysis method which comprises the following steps: analyzing multi-omics data difference, acquiring a known intermolecular regulation relation table, integrating known intermolecular regulation relations of differential protein, scoring protein importance, and analyzing based on protein importance scores. According to the method, the importance and difference authenticity of the proteins are judged by acquiring omics results of multiple types and integrating different types of regulation and control relationships between the proteins and between the proteins and metabolites.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of molecular protein technology, and in particular to a protein-centered multi-omics data analysis method. Background Art

[0002] Proteins are the most important executors of life activities. The most direct factors affecting their functions are protein abundance and post-translational modifications. There are many types of post-translational modifications, and the structure and composition of the same modification are also different. Therefore, post-translational modifications give proteins diverse functions and form a complex and precise molecular regulatory network behind life phenomena.

[0003] Traditional multi-omic studies simply rely on more than one omics group, without strict requirements for omics hierarchy. Furthermore, they lack diversity in protein post-translational modification omics, resulting in a less comprehensive and insightful understanding of the essence of life phenomena. Due to the limited omics hierarchy and number, single-omic and multi-omic studies often yield partial results that fail to fully reflect the overall picture. Consequently, their analyses can be biased and even erroneous. Summary of the Invention

[0004] In order to overcome the shortcomings of the existing technology, the purpose of the present invention is to provide a protein-centered multi-omics data analysis method to address the problem that the differential analysis results are numerous and complex due to the increase in the level and number of omics and the diversification of data types, and traditional analysis methods are difficult to find key molecular change data from them.

[0005] To achieve the above object, the present invention provides the following solutions:

[0006] A protein-centric multi-omics data analysis method, including:

[0007] Obtain differential analysis data and prior knowledge;

[0008] According to the differential analysis data, the detection and differential entries of each group are obtained to obtain a valid protein table;

[0009] Integrating the effective protein table and the prior knowledge to obtain an effective prior knowledge table; the effective prior knowledge table includes: a transcriptional regulator table, a post-translational modification regulator table, a metabolic regulator table, and an interaction regulator table;

[0010] Summarizing all the effective prior knowledge tables to obtain an effective prior knowledge usage table; the content of the effective prior knowledge usage table includes: the names of two proteins or metabolites with regulatory or interactive relationships, the direction of regulation or interaction, the type of regulation or interaction, and the degree of regulation;

[0011] Calculating the omics type scores according to the effective prior knowledge table to obtain an omics type scoring table;

[0012] The regulatory type score and direction score are calculated according to the omics type score table to obtain a protein importance score table, and molecular target screening is performed according to the protein importance score table.

[0013] Preferably, the differential analysis data includes: transcriptome data, proteome data, post-translational modification group data and metabolome data; the prior knowledge includes: transcription factor and target gene transcriptional regulation, enzyme and substrate post-translational modification regulation, enzyme and metabolite metabolic regulation and protein interaction.

[0014] Preferably, the omics type scores include: transcription score, protein score, post-translational modification score and metabolic score.

[0015] Preferably, the regulation type score includes: transcriptional regulation score, post-translational modification regulation score, metabolic regulation score and interaction regulation score.

[0016] Preferably, the directional scores include: upstream score, downstream score, self score and undirected interaction score.

[0017] Preferably, the calculation formula for the transcription score is:

[0018] Score 转录 =BL 转录 ×BD 转录 ×BA 转录 ×EM 转录 ;

[0019] The calculation formula of the protein fraction is:

[0020] Score 蛋白 =BL 蛋白 ×BD 蛋白 ×BA 蛋白 ×EM 蛋白 ;

[0021] The calculation formula for the post-translational modification score is:

[0022] Score 翻译后修饰 =BL 翻译后修饰 ×BD 翻译后修饰 ×BA 翻译后修饰 ×EM 翻译后修饰 ;

[0023] The calculation formula of the metabolic score is:

[0024] Score 代谢 =BL 代谢 ×BD 代谢×BA 代谢 ×EM 代谢 ;

[0025] Among them, Score 转录 、Score 蛋白 、Score 翻译后修饰 、Score 代谢 are the transcription score, the protein score, the post-translational modification score, and the metabolic score respectively; BL 转录 BL 蛋白 BL 翻译后修饰 BL 代谢 are the balance coefficients of transcriptome, proteome, post-translational modification group, and metabolome respectively; BD 转录 , BD 蛋白 , BD 翻译后修饰 , BD 代谢 are the balance coefficients of the sequencing depth of transcriptome, proteome, post-translational modification group, and metabolome, respectively; BA 转录 , BA 蛋白 , BA 翻译后修饰 , BA 代谢 are the balance coefficients of transcriptome, proteome, post-translational modification group, and metabolome differential analysis methods; EM 转录 , EM 蛋白 , EM 翻译后修饰 , EM 代谢 They are the enrichment coefficients of transcriptome, proteome, post-translational modification group, and metabolome, respectively.

[0026] Preferably, the calculation formula for the transcriptional regulation score is:

[0027] Score 转录调控 =Score 转录1 +Score 蛋白1 ;

[0028] The calculation formula for the post-translational modification regulation score is:

[0029] Score 翻译后修饰调控 =Score 蛋白1 +Score 翻译后修饰1 ;

[0030] The calculation formula of the metabolic regulation score is:

[0031] Score 代谢调控 =Score 代谢2 ;

[0032] The calculation formula of the interaction regulation score is:

[0033] Score 互作调控 =Score 转录1+Score 蛋白1 +Score 翻译后修饰1 ;

[0034] Among them, Score 转录调控 、Score 翻译后修饰调控 、Score 代谢调控 、Score 互作调控 are the transcriptional regulation score, the metabolic regulation score, and the interaction regulation score respectively; Score 转录1 、Score 蛋白1 、Score 翻译后修饰1 Score is the transcription score, protein score and post-translational modification score of the target protein respectively; 代谢2 is the metabolic fraction of the target metabolite.

[0035] Preferably, the calculation formula of the upstream score is:

[0036] Score 上游 =ER 转录调控 ×Score 自身 +ER 翻译后修饰调控 ×Score 自身 ;

[0037] The downstream fraction is calculated as:

[0038] Score 下游 =ER 转录调控 ×Score 转录调控 +ER 翻译后修饰调控 ×Score 翻译后修饰调控 +ER 代谢调控 ×

[0039] Score 代谢调控 ;

[0040] The calculation formula of the self-score is:

[0041] Score 自身 =Score 转录 +Score 蛋白 +Score 翻译后修饰 ;

[0042] The calculation formula of the undirected interaction score is:

[0043] Score 无向互作 =ER 无向互作 ×Score 无向互作 ;

[0044] Among them, Score 上游 、Score 下游 、Score自身 、Score 无向互作 are the upstream score, the downstream score, the self score and the undirected interaction score respectively; ER 转录调控 , ER 翻译后修饰调控 , ER 代谢调控 , ER 无向互作 They are the regulatory coefficients of transcriptional regulation, post-translational modification regulation, metabolic regulation and interaction regulation, respectively.

[0045] The present invention discloses the following technical effects:

[0046] The present invention provides a protein-centered multi-omics data analysis method. By obtaining a large number of omics results and integrating different types of regulatory relationships between proteins and proteins, and proteins and metabolites, it solves the problem that the difference analysis results are numerous and complex due to the increase in the level and number of omics and the diversification of data types, and traditional analysis methods are difficult to find key molecular change data from them, and realizes the judgment of protein importance and the authenticity of differences. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0048] Figure 1 Schematic diagram of the multi-omics SuperOmic data analysis process provided by an embodiment of the present invention;

[0049] Figure 2 This is a flowchart of the multi-omics SuperOmic data analysis provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0051] The purpose of the present invention is to provide a protein-centered multi-omics data analysis method to address the problem that the differential analysis results are numerous and complex due to the increase in the level and number of omics and the diversification of data types, and traditional analysis methods are difficult to find key molecular change data from them.

[0052] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0053] Figure 1 Schematic diagram of the multi-omics Super Omic data analysis process provided by the embodiment of the present invention, Figure 2 The multi-omics Super Omic data analysis flow chart provided in the embodiment of the present invention is as follows: Figure 1 and Figure 2 As shown, the present invention provides a protein-centered multi-omics data analysis method, comprising:

[0054] Step 100: Obtaining differential analysis data and prior knowledge;

[0055] Step 200: Obtain detection and difference items of each omics according to the difference analysis data to obtain a valid protein table;

[0056] Step 300: Integrate the effective protein table and the prior knowledge to obtain an effective prior knowledge table; the effective prior knowledge table includes: a transcriptional regulator table, a post-translational modification regulator table, a metabolic regulator table, and an interaction subtable;

[0057] Step 400: Summarizing all the effective prior knowledge tables to obtain an effective prior knowledge usage table; the content of the effective prior knowledge usage table includes: the names of two proteins or metabolites with regulatory or interactive relationships, the direction of regulation or interaction, the type of regulation or interaction, and the degree of regulation;

[0058] Step 500: Calculate the omics type score according to the effective prior knowledge usage table to obtain an omics type scoring table;

[0059] Step 600: Calculate the regulatory type score and direction score according to the omics type score table to obtain a protein importance score table, and perform molecular target screening according to the protein importance score table.

[0060] Preferably, the differential analysis data includes: transcriptome data, proteome data, post-translational modification group data and metabolome data; the prior knowledge includes: transcription factor and target gene transcriptional regulation, enzyme and substrate post-translational modification regulation, enzyme and metabolite metabolic regulation and protein interaction.

[0061] Preferably, the omics type scores include: transcription score, protein score, post-translational modification score and metabolic score.

[0062] Preferably, the regulation type score includes: transcriptional regulation score, post-translational modification regulation score, metabolic regulation score and interaction regulation score.

[0063] Preferably, the directional scores include: upstream score, downstream score, self score and undirected interaction score.

[0064] Preferably, the calculation formula for the transcription score is:

[0065] Score 转录 =BL 转录 ×BD 转录 ×BA 转录 ×EM 转录 ;

[0066] The calculation formula of the protein fraction is:

[0067] Score 蛋白 =BL 蛋白 ×BD 蛋白 ×BA 蛋白 ×EM 蛋白 ;

[0068] The calculation formula for the post-translational modification score is:

[0069] Score 翻译后修饰 =BL 翻译后修饰 ×BD 翻译后修饰 ×BA 翻译后修饰 ×EM 翻译后修饰 ;

[0070] The calculation formula of the metabolic score is:

[0071] Score 代谢 =BL 代谢 ×BD 代谢 ×BA 代谢 ×EM 代谢 ;

[0072] Among them, Score 转录 、Score 蛋白 、Score 翻译后修饰 、Score 代谢 are the transcription score, the protein score, the post-translational modification score, and the metabolic score respectively; BL 转录 BL 蛋白 BL 翻译后修饰 BL 代谢 are the balance coefficients of transcriptome, proteome, post-translational modification group, and metabolome respectively; BD 转录 , BD 蛋白 , BD 翻译后修饰 , BD 代谢 are the balance coefficients of the sequencing depth of transcriptome, proteome, post-translational modification group, and metabolome, respectively; BA 转录 , BA 蛋白 , BA 翻译后修饰, BA 代谢 are the balance coefficients of transcriptome, proteome, post-translational modification group, and metabolome differential analysis methods; EM 转录 , EM 蛋白 , EM 翻译后修饰 , EM 代谢 They are the enrichment coefficients of transcriptome, proteome, post-translational modification group, and metabolome, respectively.

[0073] Preferably, the calculation formula for the transcriptional regulation score is:

[0074] Score 转录调控 =Score 转录1 +Score 蛋白1 ;

[0075] The calculation formula for the post-translational modification regulation score is:

[0076] Score 翻译后修饰调控 =Score 蛋白1 +Score 翻译后修饰1 ;

[0077] The calculation formula of the metabolic regulation score is:

[0078] Score 代谢调控 =Score 代谢2 ;

[0079] The calculation formula of the interaction regulation score is:

[0080] Score 互作调控 =Score 转录1 +Score 蛋白1 +Score 翻译后修饰1 ;

[0081] Among them, Score 转录调控 、Score 翻译后修饰调控 、Score 代谢调控 、Score 互作调控 are the transcriptional regulation score, the metabolic regulation score, and the interaction regulation score respectively; Score 转录1 、Score 蛋白1 、Score 翻译后修饰1 Score is the transcription score, protein score and post-translational modification score of the target protein respectively; 代谢2 is the metabolic fraction of the target metabolite.

[0082] Preferably, the calculation formula of the upstream score is:

[0083] Score 上游 =ER 转录调控 ×Score自身 +ER 翻译后修饰调控 ×Score 自身 ;

[0084] The downstream fraction is calculated as:

[0085] Score 下游 =ER 转录调控 ×Score 转录调控 +ER 翻译后修饰调控 ×Score 翻译后修饰调控 +ER 代谢调控 ×

[0086] Score 代谢调控 ;

[0087] The calculation formula of the self-score is:

[0088] Score 自身 =Score 转录 +Score 蛋白 +Score 翻译后修饰 ;

[0089] The calculation formula of the undirected interaction score is:

[0090] Score 无向互作 =ER 无向互作 ×Score 无向互作 ;

[0091] Among them, Score 上游 、Score 下游 、Score 自身 、Score 无向互作 are the upstream score, the downstream score, the self score and the undirected interaction score respectively; ER 转录调控 , ER 翻译后修饰调控 , ER 代谢调控 , ER 无向互作 They are the regulatory coefficients of transcriptional regulation, post-translational modification regulation, metabolic regulation and interaction regulation, respectively.

[0092] Preferably, the calculation of the PCAS score takes into account the omics results at different levels, such as RNA, WCP, PTM, and META (metabolites only); the currently known major intermolecular regulatory modes TF, ES, EM, and INTACT; and the directions of intermolecular regulatory modes up, down, and undirected Intact (interacting molecules after eliminating upstream and downstream).

[0093] refer to Figure 2, K9, R78, K143, Y206, S301, K412, and T511 are lysine at position 9, arginine at position 78, lysine at position 143, tyrosine at position 206, serine at position 301, lysine at position 412, and threonine at position 511, respectively.

[0094] Specifically, the main principles are:

[0095] 1) Balance principle:

[0096] Balance the ability of different levels of omics (transcription / protein / post-translational modification / metabolism) to reflect the final functional activity of proteins; balance the impact of different omics sequencing depths on the results, that is, seek a balance between different omics levels of specific proteins rather than a balance between different omics levels of all proteins as a whole, so as to avoid excessive impact on the results when the number of differential entries in a certain omics is too small; balance the differences in results caused by different differential analysis algorithms (the stricter the algorithm, the higher the score of the result should be).

[0097] 2) Enrichment principle:

[0098] Since a protein usually has multiple modification sites, the principle of enrichment analysis is used to calculate the proportion of differentially modified sites to the total detected modification sites (to reduce the impact of differences due to errors in proteins containing a large number of modification sites, and at the same time give different importance to differentially post-translationally modified proteins); when there are multiple related molecules in the same regulatory mode and direction of the differential protein, the principle of enrichment analysis is used to calculate the proportion of differentially related molecules to the total detected related molecules (to avoid the situation where transcription factors containing a large number of target genes obtain too high scores).

[0099] 3) Fuzzy principle:

[0100] The regulatory relationships between proteins discovered so far are not comprehensive. For example, if A is a phosphokinase of B, in addition to increasing the phosphorylation of a certain site of B, it may also cause changes in B's acetylation or reduce the phosphorylation of other sites. Secondly, the credibility of whether or not there is regulation is greater than the direction of regulation. Therefore, the scoring does not currently consider the specific modification sites and regulatory directions of regulation.

[0101] Furthermore, integrating the required prior knowledge

[0102] 1) Required Documents

[0103] Differential analysis results: transcriptome, proteome, post-translational modification group, metabolome.

[0104] Prior knowledge: transcription factors and target genes; enzymes and substrates; enzymes and metabolites; protein interactions.

[0105] 2) Obtain the detection and difference items of each omics, and summarize transcription, protein, and post-translational modification as "gene", transcription and protein as "expression", and protein and post-translational modification as "function"; the corresponding difference items are marked as "gene difference" / "expression difference" / "function difference", and the final valid protein table with 14 columns is obtained.

[0106] 3) Categorize, import, and organize prior knowledge from all database sources.

[0107] 4) Obtain valid prior knowledge tables based on the effective protein table and the prior knowledge table: transcriptional regulator table (the "gene" item of the effective protein is in the transcriptional regulation column of the prior knowledge table, and the "expression" item of the effective protein is in the target molecule column of the prior knowledge table); post-translational modification regulator table (the "gene" item of the effective protein is in the enzyme column of the prior knowledge table, and the "function" item of the effective protein is in the substrate column of the prior knowledge table); metabolic regulator table (the "gene" item of the effective protein is in the gene name of the corresponding enzyme obtained from the metabolic item containing the KEGG number in the prior knowledge table); interaction subtable (the "gene" item of the effective protein is also in the prior knowledge table); each subtable contains a "regulation" column (yes / no), that is, the corresponding item is calculated using "gene difference" / "expression difference" / "function difference". Only direct effects are considered, and positive and negative feedback effects are not considered.

[0108] 5. Summarize the effective prior knowledge tables into one table to form an effective prior knowledge usage table containing genes (upstream), target molecules (downstream), direction (yes / no), type (transcriptional regulation / post-translational modification regulation / metabolic regulation / interaction), and regulation (yes / no). Duplicate values must be removed from entries of each type (after negating the interaction, merge it with the original table to form a new table and then remove duplicate values. Subsequently, only genes need to be queried). Interactions must also exclude entries with the same relationship as post-translational modification regulation types (both positive and negative are required).

[0109] Specifically, the balance parameters:

[0110] Omics level BL: BL 转录 =1,BL 蛋白 =2, BL 翻译后修饰 =2, BL 代谢 =2

[0111] Sequencing depth BD: BD 转录 =1, BD 蛋白 =1, BD 翻译后修饰 =1, BD 代谢 =1

[0112] Analytical methods:

[0113] BA:BA 转录 =√(all 转录 / Regulation 转录 )

[0114] BA 蛋白 =√(all 蛋白 / Regulation 蛋白 )

[0115] BA 翻译后修饰 =√(all 翻译后修饰 / Regulation 翻译后修饰 )

[0116] BA 代谢 =√(all 代谢 / Regulation 代谢 )

[0117] Here, "all" refers to all detected items, and "adjusted" refers to items with differences. The square root is used because, on the one hand, the differences in proportion may be real and not all caused by the algorithm, and on the other hand, it avoids overcorrecting the algorithm's bias.

[0118] Furthermore, the enrichment parameters:

[0119] Post-translational modification regulation of protein sites, so that the sum is the number of proteins, that is, the average value is 1: Post-translational modification regulation 转录 =1, post-translational modification regulation 蛋白 =1, post-translational modification regulation 代谢 =1, post-translational modification regulation 翻译后修饰 =(number of differential sites of the protein / total number of detected sites of the protein) × (total number of differential proteins ÷∑ number of differential sites of the protein / total number of detected sites of the protein)

[0120] Regulation mode, ER = the number of different molecules in this regulation type and direction ÷ the number of detected molecules in this regulation type and direction

[0121] Specifically, PCASscore calculation is divided into three levels. The first level is the omics type score. 组学 (Score 转录 、Score 蛋白 、Score 翻译后修饰 、Score 代谢 ), the second level is the control type Score regulation (Score 转录调控 、Score 翻译后修饰调控 、Score 代谢调控 、Score 互作 ), the third level is direction Score 方向 (Score 上游 、Score 下游 、Score 自身 、Score 无向互作 )

[0122] 1) Calculate the score of differentially expressed proteins and metabolites 组学 =BL×BD×BA×post-translational modification regulation

[0123] Score 转录 =BL 转录 ×BD 转录 ×BA 转录 × Post-translational modification regulation 转录 =1×1×√(total transcriptome detection items / differential items)×1

[0124] Score 蛋白 =BL 蛋白 ×BD 蛋白 ×BA 蛋白 × Post-translational modification regulation 蛋白 =2×1×√(total protein group detection items / differential items)×1

[0125] Score 翻译后修饰 =BL 翻译后修饰 ×BD 翻译后修饰 ×BA 翻译后修饰 × Post-translational modification regulation 翻译后修饰 =2×1×√(total detected items in the post-translational modification group / differential items)×post-translational modification regulation 翻译后修饰

[0126] Score 代谢 =BL 代谢 ×BD 代谢 ×BA 代谢 × Post-translational modification regulation 代谢 =2×1×√(total metabolomics test items / differential items)×1

[0127] BA 翻译后修饰 The number of entries is the number of sites rather than the number of proteins, among which post-translational modification regulation 翻译后修饰 The omics type score table is obtained by dividing the ratio of differentially modified sites on the protein to all detected sites on the protein by the sum of the values of all proteins.

[0128] 2) Calculate the PCASscore and Score of all differentially expressed proteins 自身 、Score 上游 、Score 下游 、Score 无向互 do.

[0129] The calculation method is: PACSscore = Score 自身 +Score 上游 +Score 下游 +Score 无向互作 ;

[0130] Score 自身 =Score 转录 +Score 蛋白 +Score 翻译后修饰 ;

[0131] Score 上游 =ER 转录调控 ×Score 自身 +ER 翻译后修饰调控 ×Score 自身 ; Obtain the type of entries containing the gene in the target molecule column from the effective prior knowledge usage table, and obtain the upstream table after classification; Obtain the entries containing the gene in the target molecule column from the effective prior knowledge usage (regulation = yes), and calculate the score of each entry = 1 / number of upstream differential regulation × Score 自身 ;

[0132] Where Score downstream = ER 转录调控 ×Score 转录调控 +ER 翻译后修饰调控 ×Score 翻译后修饰调控 +ER 代谢调控 ×Score 代谢调控 ; Get the entry containing the gene in the gene column from the effective prior knowledge usage; Get the entry containing the gene in the gene column from the effective prior knowledge usage (regulation = yes); ER 转录调控 =1 / number of transcriptional regulators in the downstream table; ER 翻译后修饰调控 =1 / number of post-translational modifications in the downstream table; ER 代谢调控 =1 / number of metabolic regulation in the downstream table; Score 转录调控 =Score 转录3 +Score 蛋白3 ;Score 翻译后修饰调控 =Score 蛋白3 +Score 翻译后修饰3 ;Score 代谢调控 =Score 代谢3 ;Score 转录3 、Score 蛋白3 、Score 翻译后修饰3 、Score 代谢3 corresponding target molecules;

[0133] Score 无向互作 =ER 无向互作 ×Score 无向互作 ; Use the effective prior knowledge to obtain the type of entries containing the gene in the gene column, and obtain the downstream table after taking the table; Use the effective prior knowledge (regulation = yes) to obtain the entries containing the gene in the gene column; ER互作 =1 / number of regulatory interactions in the downstream table; Score 无向互作 =Score 自身3 ;Score 自身3 corresponding target molecules;

[0134] When the protein itself is used as its upstream and downstream, the coefficient r=0.5 is additionally multiplied; the protein importance score table is obtained by sorting from large to small according to PCAS score.

[0135] The beneficial effects of the present invention are as follows:

[0136] The present invention achieves the judgment of protein importance and the authenticity of differences by obtaining a large number of types of omics results and integrating different types of regulatory relationships between proteins and proteins, and proteins and metabolites.

[0137] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0138] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. A protein-centric multi-omics data analysis method, characterized in that: include: Obtain differential analysis data and prior knowledge; According to the differential analysis data, the detection and differential entries of each group are obtained to obtain a valid protein table; Integrating the effective protein table and the prior knowledge to obtain an effective prior knowledge table; The effective prior knowledge table includes: a transcriptional regulator table, a post-translational modification regulator table, a metabolic regulator table, and an interaction regulator table; Summarizing all the effective prior knowledge tables to obtain an effective prior knowledge usage table; the content of the effective prior knowledge usage table includes: the names of two proteins or metabolites with regulatory or interactive relationships, the direction of regulation or interaction, the type of regulation or interaction, and the degree of regulation; Calculating the omics type scores according to the effective prior knowledge table to obtain an omics type scoring table; The regulatory type score and direction score are calculated according to the omics type score table to obtain a protein importance score table, and molecular target screening is performed according to the protein importance score table.

2. A protein-centric multi-omics data analysis method according to claim 1, characterized in that: The differential analysis data includes: transcriptome data, proteome data, post-translational modification group data and metabolome data; the prior knowledge includes: transcription factor and target gene transcriptional regulation, enzyme and substrate post-translational modification regulation, enzyme and metabolite metabolic regulation and protein interaction.

3. The protein-centric multi-omics data analysis method according to claim 1, characterized in that: The omics type scores include: transcription score, protein score, post-translational modification score and metabolic score.

4. The protein-centric multi-omics data analysis method according to claim 1, characterized in that: The regulation type scores include: transcriptional regulation score, post-translational modification regulation score, metabolic regulation score and interaction regulation score.

5. The protein-centric multi-omics data analysis method according to claim 1, characterized in that: The directional scores include: upstream score, downstream score, self score and undirected interaction score.

6. The protein-centric multi-omics data analysis method according to claim 3, characterized in that: The calculation formula of the transcription score is: Score 转录 =BL 转录 ×BD 转录 ×BA 转录 ×EM 转录 ; The calculation formula of the protein fraction is: Score 蛋白 =BL 蛋白 ×BD 蛋白 ×BA 蛋白 ×EM 蛋白 ; The calculation formula for the post-translational modification score is: Score 翻译后修饰 =BL 翻译后修饰 ×BD 翻译后修饰 ×BA 翻译后修饰 ×EM 翻译后修饰 ; The calculation formula of the metabolic score is: Score 代谢 =BL 代谢 ×BD 代谢 ×BA 代谢 ×EM 代谢 ; Among them, Score 转录 、Score 蛋白 、Score 翻译后修饰 、Score 代谢 are the transcription score, the protein score, the post-translational modification score, and the metabolic score respectively; BL 转录 BL 蛋白 BL 翻译后修饰 BL 代谢 are the balance coefficients of transcriptome, proteome, post-translational modification group, and metabolome respectively; BD 转录 , BD 蛋白 , BD 翻译后修饰 , BD 代谢 are the balance coefficients of the sequencing depth of transcriptome, proteome, post-translational modification group, and metabolome, respectively; BA 转录 , BA 蛋白 , BA 翻译后修饰 , BA 代谢 are the balance coefficients of transcriptome, proteome, post-translational modification group, and metabolome differential analysis methods; EM 转录 , EM 蛋白 , EM 翻译后修饰 , EM 代谢 They are the enrichment coefficients of transcriptome, proteome, post-translational modification group, and metabolome, respectively.

7. The protein-centric multi-omics data analysis method according to claim 6, characterized in that: The calculation formula of the transcriptional regulation score is: Score 转录调控 =Score 转录1 +Score 蛋白1 ; The calculation formula for the post-translational modification regulation score is: Score 翻译后修饰调控 =Score 蛋白1 +Score 翻译后修饰1 ; The calculation formula of the metabolic regulation score is: Score 代谢调控 =Score 代谢2 ; The calculation formula of the interaction regulation score is: Score 互作调控 =Score 转录1 +Score 蛋白1 +Score 翻译后修饰1 ; Among them, Score 转录调控 、Score 翻译后修饰调控 、Score 代谢调控 、Score 互作调控 are the transcriptional regulation score, the metabolic regulation score, and the interaction regulation score respectively; Score 转录1 、Score 蛋白1 、Score 翻译后修饰1 Score is the transcription score, protein score and post-translational modification score of the target protein respectively; 代谢2 is the metabolic fraction of the target metabolite.

8. The protein-centric multi-omics data analysis method according to claim 7, characterized in that: The calculation formula of the upstream fraction is: Score 上游 =IS 转录调控 ×Score 自身 +IS 翻译后修饰调控 ×Score 自身 ; The downstream fraction is calculated as: Score 下游 =IS 转录调控 ×Score 转录调控 +IS 翻译后修饰调控 ×Score 翻译后修饰调控 +IS 代谢调控 × Score 代谢调控 ; The calculation formula of the self-score is: Score 自身 =Score 转录 +Score 蛋白 +Score 翻译后修饰 ; The calculation formula of the undirected interaction score is: Score 无向互作 =IS 无向互作 ×Score 无向互作 ; Among them, Score 上游 、Score 下游 、Score 自身 、Score 无向互作 are the upstream score, the downstream score, the self score and the undirected interaction score respectively; ER 转录调控 , ER 翻译后修饰调控 , ER 代谢调控 , ER 无向互作 They are the regulatory coefficients of transcriptional regulation, post-translational modification regulation, metabolic regulation and interaction regulation, respectively.