Method for establishing protein interaction network of disease markers

By systematically integrating the physical and chemical characteristics and dynamic interaction information of proteins, a multivariate optimization equation system is constructed, which solves the problem of reliability and dynamic changes in the construction of disease marker networks in the existing technology, and realizes the accurate construction of disease marker protein interaction networks, identifys key nodes, and provides a reliable basis for disease research.

CN120299514APending Publication Date: 2025-07-11QINGDAO RAISECARE BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510450209.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, when building a network of protein interactions with disease markers, experimental methods require a lot of work and are low in reliability. Database mining methods are difficult to reflect the dynamic changes in tissue specificity and disease state. Prediction methods based on sequence homology ignore the complex interactions between protein spatial structure and physiological environment.

Method used

By collecting human tissue samples, affinity chromatography is used to enrich the target disease marker protein by affinity chromatography and mass spectrometry combined with liquid chromatography and real-time fluorescence confocal microscopy, a network optimization equation set that comprehensively considers the protein steric hinder effect, expression abundance correlation, colocal coefficient and interaction intensity, and a minimum dominance set algorithm is used to identify key protein nodes.

Benefits of technology

The precise construction of the protein interaction network of disease marker proteins is achieved, which can reflect dynamic changes in the disease state, provide a reliable basis for disease mechanism research and drug target discovery, reduce false positives, and improve the biological rationality and reliability of the network structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299514A_ABST
    Figure CN120299514A_ABST
Patent Text Reader

Abstract

A method for establishing a protein interaction network of disease markers belongs to the field of disease markers, and comprises the following steps: collecting a sample, extracting total protein, and carrying out protein quantitative determination; carrying out affinity chromatography by using a magnetic bead coupled specific antibody to enrich target disease marker protein; component identification is carried out by adopting a liquid chromatography-mass spectrometry technology; measuring the expression abundance of the protein components; carrying out dynamic imaging by adopting a real-time fluorescence confocal microscope; calculating an interaction intensity matrix and constructing an initial network topology structure; constructing and solving a network optimization equation set which comprehensively considers a protein steric hindrance effect, expression abundance correlation, a co-localization coefficient and interaction intensity; identifying key protein nodes by adopting a minimum dominating set algorithm; constructing a tissue-specific disease marker protein interaction network; according to the invention, innovative research is carried out aiming at the core technical problem of how to construct a tissue-specific disease marker protein interaction network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of disease markers, and more particularly, relates to a method for establishing a protein interaction network of disease markers. Background Art

[0002] Currently, the construction of protein interaction networks for disease markers mainly relies on two methods: experimental determination and database mining. Traditional experimental methods such as the yeast two-hybrid system and co-immunoprecipitation technology can directly obtain protein interaction information. However, these methods often require a large amount of experimental work, and there are differences between the experimental conditions and the real physiological environment. Therefore, the obtained interaction data may not conform to the actual situation. Methods based on database mining rely too much on existing data and are difficult to discover new interaction relationships. At the same time, this method cannot reflect the tissue specificity and dynamic change characteristics under disease states. In addition, although the prediction method based on sequence homology is simple to operate, since it only considers the one-dimensional sequence information of proteins and ignores the complex interactions of proteins in spatial structure and physiological environment, the reliability of the prediction results is relatively low. Generally speaking, the existing methods for constructing protein interaction networks have the technical problems that experimental determination requires a large amount of work and the reliability is relatively low. Summary of the Invention

[0003] In view of this, the present invention provides a method for establishing a protein interaction network of disease markers, which can solve the technical problems in the prior art that experimental determination requires a large amount of work and the reliability is relatively low. The method includes the following operating steps: collecting human tissue samples and extracting total proteins, and performing protein quantitative determination; enriching target disease marker proteins by affinity chromatography using magnetic bead-coupled specific antibodies; performing component identification by liquid chromatography-mass spectrometry; measuring the expression abundance of protein components; performing dynamic imaging using a real-time fluorescence confocal microscope; calculating the interaction strength matrix and constructing an initial network topology structure; constructing and solving a network optimization equation set that comprehensively considers protein steric hindrance effects, expression abundance correlations, co-localization coefficients, and interaction strengths; identifying key protein nodes using the minimum dominating set algorithm; and constructing a tissue-specific disease marker protein interaction network.

[0004] Among them, the step of collecting human tissue samples and extracting total proteins includes: measuring the total protein concentration using the Bradford protein quantification method, adjusting the total protein concentration to 50 to 100 micrograms per milliliter, measuring the total protein steric hindrance coefficient and the protein molecular weight, and measuring the number of total protein interaction sites.

[0005] Among them, the step of using the specific antibody conjugated with magnetic beads for affinity chromatography to enrich the target disease biomarker protein includes: measuring the binding degree value of the target disease biomarker protein and the specific antibody to generate a protein binding degree matrix.

[0006] Among them, the step of using liquid chromatography-mass spectrometry technology for component identification includes: obtaining protein component data, constructing an initial interaction matrix based on the protein component data, and calculating the number of interactions.

[0007] Among them, the step of measuring the expression abundance of protein components includes: obtaining protein expression abundance values, and calculating a node importance vector according to the protein expression abundance values.

[0008] Among them, the step of performing dynamic imaging using a real-time fluorescence confocal microscope includes: obtaining fluorescence intensity time series data, and calculating a protein co-localization coefficient matrix based on the fluorescence intensity time series data.

[0009] Among them, the step of constructing the initial network topology structure includes: obtaining a network topology matrix, and calculating the total number of network edges and network local density of the initial network topology structure.

[0010] Among them, the network optimization equation set includes: network reconstruction equation, network sparsity constraint equation, network connectivity constraint equation, network weight normalization equation, network node distance equation, network edge weight correlation equation, node degree correlation equation.

[0011] Among them, the minimum dominating set algorithm includes the following steps: calculating the domination degree of each protein node based on the final network weight matrix; selecting the protein node with the largest domination degree as the candidate key node; adding the candidate key node to the domination set and updating the set of undominated nodes; repeating the above steps until all protein nodes are dominated; determining the protein nodes in the domination set as the key protein nodes.

[0012] Among them, the input and output of each equation of the network optimization equation set include: the input of the network reconstruction equation includes the co-localization coefficient matrix, the interaction intensity matrix, the network topology matrix, and the network reconstruction error threshold, and the output is the network connection weight matrix; the input of the network sparsity constraint equation includes the network connection weight matrix, the network local density, the node importance vector, the network sparsity penalty coefficient, and the total number of network edges, and the output is the sparse constraint weight matrix; the input of the network connectivity constraint equation includes the sparse constraint weight matrix, the network topology matrix, the network connectivity threshold, and the node importance vector, and the output is the connectivity constraint weight matrix.

[0013] The effects of the present invention are as follows:

[0014] The method for constructing a protein interaction network proposed by the present invention integrates the physicochemical properties, dynamic interaction information, and expression abundance data of proteins, and innovatively designs a multivariate optimization equation system to achieve the precise construction of the interaction network of disease markers. Compared with the prior art, the method of the present invention has made significant progress in the following aspects: 1) By introducing the consideration of steric hindrance effect, the feasibility of the three-dimensional structure of protein interactions in the network is ensured. At the same time, the correlation constraint of expression abundance is also introduced to make the network structure more in line with biological reality. 2) Based on the dynamic imaging data obtained by real-time fluorescence confocal microscopy, not only the spatial co-localization of proteins is verified, but also a quantitative basis for the interaction strength is provided, thus making the calculation of network edge weights more accurate. 3) By constructing a multivariate optimization equation system, the all-round optimization of the network structure is achieved, such as network reconstruction, sparse constraint, connectivity guarantee, etc. This multi-dimensional optimization design enables the constructed network to better reflect the actual biological characteristics. 4) This method can identify the protein nodes that play key roles in the interaction network of disease markers, and based on this, construct a tissue-specific network. This network can not only reflect the dynamic changes of protein interactions in the disease state, but also provide a reliable basis for the study of disease mechanisms and the discovery of drug targets. It solves the technical problems of a large amount of work and low reliability in experimental determination existing in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a flowchart of the method provided by the present invention;

[0016] Figure 2 is a relationship diagram of protein molecular characteristics in Example 2;

[0017] Figure 3 is a relationship diagram of binding degree and distance in Example 2;

[0018] Figure 4 is a distribution diagram of the number of protein interactions in Example 2;

[0019] Figure 5 is an analysis diagram of fluorescence intensity correlation in Example 2. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0021] As Figure 1 shown, the present invention includes the following operating steps:

[0022] S10. Collect human tissue samples and extract total proteins. Determine the total protein concentration using the Bradford protein quantification method, adjust the total protein concentration to 50 to 100 micrograms per milliliter, measure the steric hindrance coefficient of the total protein and the protein molecular weight, and measure the number of interaction sites of the total protein.

[0023] S20. Use specific antibodies conjugated to magnetic beads for affinity chromatography to enrich target disease biomarker proteins. Measure the binding degree value of the target disease biomarker protein and the specific antibody to generate a protein binding degree matrix.

[0024] S30. Use liquid chromatography-mass spectrometry technology to identify the components of the target disease biomarker protein, obtain protein component data, construct an initial interaction matrix based on the protein component data, and calculate the number of interactions.

[0025] S40. Measure the expression abundance of the protein components to obtain protein expression abundance values, and calculate the node importance vector based on the protein expression abundance values.

[0026] S50. Use a real-time fluorescence confocal microscope to perform dynamic imaging on the protein components to obtain fluorescence intensity time series data, and calculate a protein colocalization coefficient matrix based on the fluorescence intensity time series data.

[0027] S60. Calculate an interaction strength matrix between protein components based on the fluorescence intensity time series data, construct an initial network topology structure according to the interaction strength matrix to obtain a network topology matrix, and calculate the total number of edges and network local density of the initial network topology structure.

[0028] S70. Construct a network optimization equation system and solve it to generate a final network weight matrix.

[0029] S80. Use the minimum dominating set algorithm to analyze the final network weight matrix to identify key protein nodes.

[0030] S90. Construct a tissue-specific disease biomarker protein interaction network based on the key protein nodes.

[0031] The network optimization equation system includes a network reconstruction equation, a network sparsity constraint equation, a network connectivity constraint equation, a network weight normalization equation, a network node distance equation, a network edge weight correlation equation, and a node degree correlation equation.

[0032] The network reconstruction equation is used to calculate the network connection weights. The inputs include the colocalization coefficient matrix, the interaction strength matrix, the network topology matrix, and the network reconstruction error threshold, and the output is a network connection weight matrix.

[0033] The network sparse constraint equation is used to optimize the number of network connections. The inputs include the network connection weight matrix, the network local density, the node importance vector, the network sparse penalty coefficient, and the total number of network edges. The output is the sparse constraint weight matrix;

[0034] The network connectivity constraint equation is used to ensure network connectivity. The inputs include the sparse constraint weight matrix, the network topology matrix, the network connectivity threshold, and the node importance vector. The output is the connectivity constraint weight matrix;

[0035] The network weight normalization equation is used to standardize the network weights. The inputs include the connectivity constraint weight matrix, the node importance vector, the total number of network edges, and the weight normalization coefficient. The output is the final network weight matrix;

[0036] The network node distance equation is used to limit the interaction distance between adjacent protein nodes. The inputs include the protein molecular weight, the steric hindrance coefficient, and the number of interaction sites. The output is the minimum allowable distance between nodes;

[0037] The network edge weight correlation equation is used to limit the ratio range of adjacent edge weights. The inputs include the co-localization coefficient matrix, the interaction strength matrix, and the node importance vector. The output is the edge weight ratio constraint;

[0038] The node degree correlation equation is used to limit the degree difference between adjacent nodes. The inputs include the protein expression abundance value, the number of interactions, and the network local density. The output is the node degree difference constraint;

[0039] The minimum dominating set algorithm includes:

[0040] In the first step, calculate the domination degree of each protein node based on the final network weight matrix;

[0041] In the second step, select the protein node with the maximum domination degree as the candidate key node;

[0042] In the third step, add the candidate key node to the dominating set and update the set of undominated nodes;

[0043] In the fourth step, repeat the first step to the third step until all protein nodes are dominated;

[0044] In the fifth step, determine the protein nodes in the dominating set as the key protein nodes.

[0045] The following describes the specific implementation manners of the above steps in detail.

[0046] The specific implementation of step S10 is as follows: First, collect human tissue samples and extract total proteins using conventional biochemical methods. Then, determine the concentration of the total proteins using the Bradford protein quantification method and adjust it to a concentration range of 50 to 100 micrograms per milliliter. Next, measure and analyze the characteristics of the total proteins using physicochemical parameters such as molecular weight and steric hindrance coefficient, including determining the number of interaction sites for each protein. This information provides important basic data for subsequent network construction.

[0047] The specific implementation of step S20 is as follows: First, use specific antibodies against the target disease biomarker proteins and enrich and isolate them by magnetic bead-coupled affinity chromatography. Then, measure the binding degree value between the target proteins and the specific antibodies and generate a protein binding degree matrix accordingly. This binding degree matrix reflects the mutual binding strength between different proteins and provides key parameters for the subsequent construction of network topology.

[0048] The specific implementation of step S30 is as follows: Use liquid chromatography-tandem mass spectrometry (LC-MS / MS) technology to identify the components of the enriched target disease biomarker proteins and obtain qualitative and quantitative data on the protein components. Based on these protein component data, construct an initial interaction matrix and calculate the number of interactions between each protein and other proteins. This information lays the foundation for the initial construction of network nodes and edges.

[0049] The specific implementation of step S40 is as follows: Measure the expression abundance value of each protein component. Based on these expression abundance data, calculate the importance vector of each node (protein). Node importance is an indicator to measure the status of the protein in the network, and it is related to factors such as the expression abundance, degree, and number of interactions of the node. This node importance vector provides a key reference basis for subsequent network optimization.

[0050] The specific implementation of step S50 is as follows: Use a real-time fluorescence confocal microscope to perform dynamic imaging observations on the enriched protein components. By analyzing the obtained fluorescence intensity time series data, calculate the colocalization coefficient matrix between each protein component. This colocalization coefficient reflects the degree of colocalization of proteins in time and space and provides important parameters for calculating the weights of network edges.

[0051] The specific implementation of step S60 is as follows: First, based on the fluorescence intensity time series data obtained in step S50, calculate the interaction strength matrix between each protein component. Then, combine the network topology matrix obtained in the previous steps to construct an initial network topology structure. Next, calculate indicators such as the total number of edges and local density of this initial network topology. This information provides a benchmark reference for subsequent network optimization.

[0052] The specific implementation of step S70 is as follows: construct a system of multivariate equations to optimize the network. This system of equations includes a network reconstruction equation, a network sparsity constraint equation, a network connectivity constraint equation, a network weight normalization equation, etc. By solving these optimization equations, the final network weight matrix can be generated. This network weight matrix reflects the interaction strength between each edge after optimization and is the core of constructing the disease biomarker interaction network.

[0053] The network reconstruction equation is used to calculate the network connection weights. The inputs include the colocalization coefficient matrix, the interaction strength matrix, the network topology matrix, and a network reconstruction error threshold (the value range is 0.01 - 0.03), and the output is the optimized network connection weight matrix. This equation comprehensively considers various factors such as the spatial correlation of proteins, the interaction strength, and the network topology structure, so as to obtain a network connection weight closer to the actual situation.

[0054] The network sparsity constraint equation is used to optimize the number of connections in the network. The inputs include the network connection weight matrix, the network local density, the node importance vector, the network sparsity penalty coefficient (the value range is 0.1 - 1.0), and the total number of network edges, and the output is the weight matrix after sparse constraint. By introducing factors such as node importance and network sparsity, this equation effectively controls the connection density of the network and avoids the false positive problem caused by over-connection.

[0055] The network connectivity constraint equation is used to ensure the connectivity of the network. The inputs include the sparse constraint weight matrix, the network topology matrix, the network connectivity threshold, and the node importance vector, and the output is the weight matrix after connectivity constraint. On the basis of maintaining the sparsity of the network structure, this equation ensures that information can be smoothly transmitted between each node, thus ensuring the overall connectivity of the network.

[0056] The network weight normalization equation is used to standardize the network weights. The inputs include the connectivity constraint weight matrix, the node importance vector, the total number of network edges, and a weight normalization coefficient (the value range is 0.001 - 0.01), and the output is the final network weight matrix. By introducing factors such as node importance, this equation ensures the reasonable distribution of network weights and makes it better reflect the relative importance between each edge.

[0057] The specific implementation of step S80 is as follows: The optimized network weight matrix is analyzed using the minimum dominating set algorithm. Specifically, it includes the following sub-steps: First, calculate the domination degree of each protein node based on the network weight matrix; then select the node with the maximum domination degree as the candidate key node; next, add the candidate key node to the domination set and update the set of undominated nodes; repeat the above process until all nodes are dominated; finally, determine the nodes in the domination set as the key protein nodes. This method can effectively identify the key nodes that play important roles in the disease biomarker interaction network.

[0058] The specific implementation of step S90 is as follows: Based on the key protein nodes identified in the previous steps, a tissue-specific disease biomarker protein interaction network is constructed. This network can reflect the specific changes in protein interactions under disease states and provide a reliable basis for disease mechanism research and drug target discovery.

[0059] The following details the specific calculation processes involved in the present invention.

[0060] 1. Calculation of the protein binding degree matrix:

[0061] The protein binding degree matrix is specifically expressed as follows:

[0062]

[0063] In the formula, B ij is the binding degree between the i-th protein and the j-th protein; C i , C j are the concentrations of the i-th and j-th proteins respectively; d ij is the spatial distance between the two proteins; r ij is the van der Waals interaction radius between the two proteins; r0 is the standard interaction distance, with a value of 10 nanometers; k1, k2 are binding constants; ε1 is the experimental error term, with a range of 0.01 - 0.05.

[0064] 2. Calculation of the node importance vector:

[0065] The node importance vector is specifically expressed as follows:

[0066]

[0067] In the formula, N i is the importance of the i-th node; E i is the expression abundance of the i-th protein; D i is the degree of the i-th node; I iis the number of interactions of the \(i\) -th protein; \(\alpha_1,\alpha_2,\alpha_3\) are weight coefficients, and \(\alpha_1+\alpha_2+\alpha_3 = 1\); \(\varepsilon_2\) is the calculation error, with a range of \(0.001\sim0.01\); \(n\) is the total number of nodes.

[0068] 3. Calculation of the colocalization coefficient matrix:

[0069] The colocalization coefficient matrix is specifically expressed as follows:

[0070]

[0071] In the formula, \(M\) ij is the colocalization coefficient of the \(i\) -th and \(j\) -th proteins; \(F\) it is the fluorescence intensity of the \(i\) -th protein at time point \(t\); is the average fluorescence intensity of the \(i\) -th protein; \(T\) is the total number of time points; \(\varepsilon_3\) is the measurement error, with a range of \(0.005\sim0.02\).

[0072] 4. Calculation of the interaction intensity matrix:

[0073] The interaction intensity matrix is specifically expressed as follows:

[0074]

[0075] In the formula, \(S\) ij is the interaction intensity of the \(i\) -th and \(j\) -th proteins; \(M\) ij is the colocalization coefficient; \(B\) ij is the binding degree; \(E\) i , \(E\) j is the protein expression abundance; \(E_0\) is the normalization parameter; \(\beta_1,\beta_2,\beta_3\) are weight coefficients; \(\varepsilon_4\) is the comprehensive error, with a range of \(0.01\sim0.05\).

[0076] 5. Calculation of the network reconstruction equation:

[0077] The network reconstruction equation is specifically expressed as follows:

[0078]

[0079] In the formula, \(W\) ij is the optimized network connection weight; \(S\) ij is the interaction intensity; \(M\) ij is the colocalization coefficient; \(N\) i , \(N\) j is the node importance; \(d\) ij is the spatial distance; \(T\) ij is the topological matrix element; \(\gamma_1,\gamma_2,\gamma_3\) are weight coefficients; \(\varepsilon_5\) is the reconstruction error, with a range of \(0.01\sim0.03\).

[0080] 6. Calculation of network sparse constraint equation:

[0081] The specific expression of the network sparse constraint equation is as follows:

[0082]

[0083] In the formula, L ij is the weight after sparse constraint; W ij is the network connection weight; λ is the sparse penalty coefficient, and its value range is 0.1 to 1.0; D l is the local density of the network; N i , N j is the node importance; E is the total number of network edges; δ1 is the adjustment coefficient; ε6 is the constraint error, and its range is 0.005 to 0.02.

[0084] 7. Calculation of network connectivity constraint equation:

[0085] The specific expression of the network connectivity constraint equation is as follows:

[0086]

[0087] In the formula, C ij is the weight after connectivity constraint; L ij is the sparse constraint weight; T ij is the topological matrix element; d ij is the spatial distance; N i , N j is the node importance; η1, η2 are the connectivity adjustment coefficients; ε7 is the connectivity error, and its range is 0.02 to 0.05.

[0088] 8. Calculation of network weight normalization equation:

[0089] The specific expression of the network weight normalization equation is as follows:

[0090]

[0091] In the formula, F ij is the final network weight; C ij is the connectivity constraint weight; N i , N j is the node importance; E is the total number of network edges; μ is the normalization coefficient; N is the total number of nodes; ε8 is the normalization error, and its range is 0.001 to 0.01.

[0092] 9. Calculation of network node distance equation:

[0093] The specific expression of the network node distance equation is as follows:

[0094]

[0095] In the formula, R ij is the minimum allowable distance between nodes; m i ,m j is the molecular weight of the protein; s i ,s j is the steric hindrance coefficient; p ij is the number of interaction sites; θ1, θ2 are distance adjustment coefficients; ε9 is the distance error, with a range of 0.5 to 2.0 nanometers.

[0096] 10. Calculation of the network edge weight correlation equation:

[0097] The network edge weight correlation equation is specifically expressed as follows:

[0098]

[0099] In the formula, K ij is the edge weight ratio constraint; M ij is the colocalization coefficient; S ij is the interaction strength; Ω i ,Ω j respectively represent the sets of nodes adjacent to nodes i and j; N i ,N j is the node importance; ρ1, ρ2 are weight coefficients, with a value range of 0.3 to 0.7; ε 10 is the constraint error, with a range of 0.01 to 0.03.

[0100] Parameter acquisition method:

[0101] 1) The colocalization coefficient M ij and the interaction strength S ij are obtained by calculating through the aforementioned equations;

[0102] 2) The adjacent node sets Ω i ,Ω j are determined by analyzing the network topology;

[0103] 3) The node importance N i ,N j is obtained by calculating the node importance vector;

[0104] 4) The weight coefficients ρ1, ρ2 are obtained by fitting experimental data through the least squares method.

[0105] 11. Calculation of the node degree correlation equation:

[0106] The node degree correlation equation is specifically expressed as follows:

[0107]

[0108] In the formula, D ij is the node degree difference constraint; E i , E j is the protein expression abundance; E max is the maximum expression abundance; I i , I j is the number of interactions; D l is the local density of the network; d ij is the spatial distance; σ1, σ2, σ3 are adjustment coefficients, and satisfy σ1 + σ2 + σ3 = 1; ε 11 is the degree difference error, and the range is 0.02 - 0.06.

[0109] Parameter acquisition method:

[0110] 1) The protein expression abundance E i , E j is obtained through mass spectrometry quantitative analysis;

[0111] 2) The number of interactions I i , I j is determined through protein interaction experiments;

[0112] 3) The local density D of the network l is obtained by calculating the ratio of the local connection number to the possible maximum connection number;

[0113] 4) The spatial distance d ij is obtained through structural analysis or molecular dynamics simulation.

[0114] The derivation and establishment processes of the above calculation process are described in detail below.

[0115] 1. Protein binding degree matrix equation The derivation of the equation is based on the basic principle of intermolecular forces. First, the electrostatic interaction is considered. According to Coulomb's law, the intermolecular force between charged particles is proportional to the product of the electric charges and inversely proportional to the square of the distance. In the protein system, the concentration can characterize the local charge density, so the concentration term C i C j and the distance term are introduced. Secondly, the contribution of van der Waals force is considered. Van der Waals force is a short-range force, and its magnitude decays exponentially with the increase of distance. Therefore, the exponential function is used to describe it, where r0 is the characteristic interaction distance; finally, considering the uncertainty of experimental measurement, the error term τ1 is introduced; the relative contributions of electrostatic interaction and van der Waals force are balanced by the adjustment coefficients k1 and k2, and the complete binding degree matrix equation is finally obtained.

[0116] For the determination process of the electrostatic interaction adjustment coefficient k1, the binding thermodynamic parameters between protein molecules under different concentration conditions need to be measured by isothermal titration calorimetry (ITC): First, prepare a series of protein solutions with different concentrations, add them to the sample cell and syringe of the ITC instrument respectively, and conduct titration experiments under constant temperature conditions, recording the heat change after each injection; by fitting the heat curve, obtain the binding constant and enthalpy change at different concentrations; combined with the electrostatic interaction contribution calculated by molecular dynamics simulation, determine the k1 value that best matches the experimental value through the least squares method.

[0117] The acquisition of the van der Waals force adjustment coefficient k2 requires measuring the force between protein molecules through atomic force microscopy (AFM) experiments: First, modify the target protein on the surfaces of the AFM probe and mica substrate respectively, and measure the force-distance curve between the probe and the sample in buffer solutions with different ionic strengths; by analyzing the approach and retraction segments of the force curve, extract the contribution of the van der Waals force; combined with the van der Waals potential surface obtained by molecular dynamics simulation, use the non-linear fitting method to determine the optimal k2 value.

[0118] The determination of the characteristic interaction distance r0 is based on the average size and interaction range of protein molecules: Measure the hydrodynamic radius of the protein by dynamic light scattering (DLS), combine the radius of gyration obtained by small-angle X-ray scattering (SAXS), and the van der Waals interaction decay curve calculated by molecular dynamics simulation to comprehensively determine the value of r0; considering that the size of most protein molecules is in the range of 1 - 10 nm, and the effective action distance of the van der Waals force is usually 2 - 3 times the molecular size, so r0 is set to 10 nm.

[0119] The range of the error term ε1 is determined based on the analysis of experimental repeatability: By repeatedly measuring the protein binding degree under the same conditions for multiple times, calculate the standard deviation of the measured values; considering the systematic and random errors brought by experimental operations, instrument precision, environmental factors, etc., combined with statistical analysis methods, determine the value range of ε1 to be 0.01 - 0.05, and this range can reasonably reflect the uncertainty of experimental measurements and ensure the reliability of calculation results.

[0120] 2. Node importance vector equation The construction process is based on the idea of multi-dimensional feature fusion. First, considering that the importance of a protein in a biological network is jointly affected by its expression level, topological features, and functional features, so the expression abundance E i , node degree D i and the number of interactions I iAs characteristic parameters; considering that these parameters have different dimensions and numerical ranges, a total normalization method is used for standardization; by introducing weight coefficients α1, α2, α3 to adjust the importance contributions of different features; finally, an error term ε2 is added to account for the uncertainty in the calculation process, thus obtaining a complete node importance calculation equation.

[0121] 3. Co-localization coefficient matrix equation The construction process of is based on the principle of Pearson correlation coefficient. This equation quantifies the spatial co-localization degree of different proteins by analyzing the correlation of their fluorescence intensities changing with time: First, obtain the time-series fluorescence intensity data, and calculate the average fluorescence intensity of each protein as a reference; then calculate the deviation of the fluorescence intensity at each time point from the average value, and calculate the sum of the products of the deviation values of the two proteins; to eliminate the influence of the absolute value of fluorescence intensity, a normalization term is introduced, that is, the geometric mean of the sum of the squares of the deviations of the two proteins; finally, considering the uncertainty in the experimental measurement and image processing processes, an error term ε3 is added; this method can not only reflect the spatial co-distribution characteristics of proteins, but also effectively eliminate the influence of background fluorescence and photobleaching.

[0122] Fluorescence intensity time-series data F it and F jt The acquisition of involves a complex experimental process: First, it is necessary to construct an expression vector labeled with fluorescent proteins, and fuse the coding sequence of the target protein with the genes of different fluorescent proteins (such as GFP and RFP); introduce the recombinant plasmid into cells by transient transfection or stable transfection, and perform confocal microscopy observation after an appropriate expression time; set appropriate parameters such as laser power, gain, and scanning speed, and continuously scan and image the selected cell area. The typical acquisition time interval is 1 - 5 seconds, and the total observation time is 5 - 10 minutes; use image analysis software to extract the fluorescence intensity values at each time point, and after background correction and photobleaching correction, obtain the standardized fluorescence intensity time-series data.

[0123] Average fluorescence intensity and The calculation of needs to consider multiple factors: First, it is necessary to preprocess the original fluorescence intensity data, including removing outliers and baseline drift correction; then considering the influence of cell autofluorescence, it is necessary to set untransfected cells as negative controls and deduct the background fluorescence value; for data observed for a long time, it is also necessary to consider the influence of the photobleaching effect and correct it through an exponential decay model; finally, perform time averaging on the processed data to obtain the characteristic fluorescence intensity value of each protein.

[0124] The range determination of the measurement error ε3 is based on the repeatability analysis of the system: by repeatedly measuring the colocalization coefficient of the same sample under the same conditions multiple times, the distribution characteristics of the measured values are analyzed; considering the influences of factors such as laser intensity fluctuations, detector noise, sample drift, etc., as well as the errors introduced in the image processing process, the value range of ε3 is determined to be 0.005 - 0.02, and this range can reasonably reflect the uncertainty of experimental measurements.

[0125] 4. Construction of the interaction strength matrix equation incorporates experimental data at multiple levels: first, considering the spatial colocalization characteristics of proteins, the colocalization coefficient M ij term is introduced; second, considering the physical interactions between molecules, the binding degree B ij term is introduced; third, considering the influence of the matching degree of protein expression levels on the interaction, it is described by an exponential function , where E0 is a normalization parameter; the relative contributions of each term are adjusted by the weight coefficients β1, β2, β3; finally, an error term ε4 is introduced to consider the influence of comprehensive errors; this multi-dimensional characterization method can more comprehensively reflect the strength characteristics of protein interactions.

[0126] The determination of the normalization parameter E0 is based on the statistical analysis of protein expression data: first, a large amount of protein expression profile data is collected, and the distribution characteristics of the expression levels are analyzed; the distribution of expression differences of known interacting protein pairs is calculated to determine the relationship between expression differences and interaction strength; through nonlinear regression analysis, the value of E0 is optimized so that the exponential decay term can reasonably reflect the influence of expression differences on the interaction; typically, the value of E0 is the median of the expression level differences in the dataset.

[0127] The optimization process of the weight coefficients β1, β2, β3 uses machine learning methods: first, a training dataset is constructed, including protein pairs with known interaction strengths and their respective characteristic parameters; using the method of cross-validation, the weight coefficients are optimized through the gradient descent algorithm to best match the predicted values with the experimentally measured interaction strengths; during the optimization process, the generalization ability and overfitting problems of the model are considered simultaneously, and the value range of the weight coefficients is constrained by adding regularization terms.

[0128] The range determination of the comprehensive error term ε4 needs to consider multiple error sources: including the propagation of measurement errors of the colocalization coefficient, measurement errors of the binding degree, measurement errors of expression levels, etc.; the total error range is calculated through the error propagation formula, and the rationality of the error estimation is verified through Monte Carlo simulation; considering various uncertain factors in practical applications, the value range of ε4 is set to 0.01 - 0.05, and this range can reasonably reflect the cumulative errors in the process of multi-source data fusion.

[0129] 5. Network reconstruction equation The construction process is based on the integration of multiple network features: First, consider the functional association features of protein pairs, and introduce the interaction strength S ij and the co-localization coefficient M ij as product terms. This non-linear combination can highlight protein pairs with high functional correlation. Second, consider the importance of nodes and spatial constraints, and introduce the ratio of the product of the node importance N i ,N j to the spatial distance d ij . This form not only emphasizes the connections between important nodes but also considers the restrictive effect of spatial distance. Third, introduce the topological matrix element T ij to maintain the basic topological features of the network; balance the contributions of each term through the weight coefficients γ1, γ2, γ3; finally, add the reconstruction error term ε5 to consider the uncertainty in the network reconstruction process.

[0130] The optimization method of the weight coefficients γ1, γ2, γ3 is based on network performance evaluation: First, construct a protein sub-network containing known functional annotations as training data; design evaluation indicators based on network structure and functional enrichment analysis, including network connectivity, modularity, functional similarity, etc.; use the genetic algorithm for parameter optimization, and evaluate the performance of different weight combinations through cross-validation; finally, select the weight coefficient combination that can achieve the optimal balance on multiple evaluation indicators.

[0131] The construction of the topological matrix element T ij requires integrating interaction data from multiple experimental sources: Protein interaction data obtained through methods such as yeast two-hybrid, co-immunoprecipitation, affinity purification-mass spectrometry, etc.; perform credibility scoring on the interactions obtained by each method based on factors such as experimental repeatability and literature support; weight and integrate the interaction data from different sources to form an initial topological relationship matrix; remove potential false-positive interactions through a network filtering algorithm.

[0132] The range determination of the reconstruction error ε5 is based on the stability analysis of network reconstruction: By sampling the data with replacement, construct multiple reconstructed network versions; calculate the structural similarity and functional consistency between different versions of the network; analyze the variation degree of network characteristic parameters, including degree distribution, clustering coefficient, centrality, etc.; based on these analyses, determine the reasonable range of the reconstruction error to be 0.01 - 0.03.

[0133] 6. Network Sparsity Constraint Equation is designed based on the sparsity characteristics of biological networks: First, consider the exponential decay constraint of weights, and implement the penalty for large-weight connections through the term, where λ is the sparsity penalty coefficient; second, introduce the compensation term of node importance Ensure that the necessary connections between important nodes are not overly weakened; control the sparsity of the network by adjusting the parameters λ and δ1; finally, add a constraint error term ε6 to account for the uncertainties in the constraint process.

[0134] The determination process of the sparse penalty coefficient λ needs to consider the biological characteristics of the network: by analyzing the degree distribution characteristics of the known protein interaction network, determine a reasonable range of network connection density; set different values of λ for network reconstruction, and analyze the characteristics such as degree distribution and connectivity of the reconstructed network; select the value of λ that can make the reconstructed network closest to the characteristics of the real biological network, and the usual value range is 0.1 - 1.0.

[0135] The optimization of the compensation coefficient δ1 is based on the principle of maintaining functional modules: first, identify the functional modules in the network, based on the known protein function annotations and interaction data; analyze the impact of different values of δ1 on the integrity of the functional modules; select the optimal value of δ1 that can achieve network sparsification while maintaining key functional modules; evaluate the stability of the parameters through cross - validation.

[0136] The range setting of the constraint error ε6 is based on the uncertainty analysis of network pruning: through multiple repeated network sparsification experiments, analyze the distribution characteristics of the edge weight changes; consider the weight change rules of different types of node connections; combine functional annotation information to evaluate the rationality of the error range; finally, determine the value range of ε6 as 0.005 - 0.02.

[0137] 7. Network connectivity constraint equation is constructed to maintain the overall connectivity of the network: on the basis of sparse constraints, strengthen the connections required by the topological structure through the term; maintain the indirect connections between important nodes through the term; adopt different distance - dependent forms (d ij and ) to reflect spatial constraints at different scales; balance the maintenance of connectivity and spatial limitations by adjusting the parameters η1, η2; finally, add a connectivity error term ε7 to account for the uncertainties in the constraint process.

[0138] The optimization process of the connectivity adjustment coefficients η1, η2 is based on network connectivity analysis: first, set the target range of network connectivity, based on the known biological network characteristics; construct test networks at different scales, including local modules and global networks; optimize the parameter combination through the simulated annealing algorithm to make the network maintain appropriate connectivity at different scales; at the same time, consider the rationality of spatial constraints to avoid generating long - range connections that do not conform to biological significance; finally, determine the optimal parameter combination through cross - validation.

[0139] The range determination of the connectivity error ε7 is based on the fluctuation analysis of network connectivity: by randomly perturbing the initial network, analyzing the change range of the connectivity index; considering the impact of node deletion and edge weight change on network connectivity; evaluating the reasonable range of connectivity change in combination with biological knowledge; comprehensively considering these factors, the value range of ε7 is set to 0.02 - 0.05.

[0140] 8. Network weight normalization equation is constructed using an improved cosine similarity normalization method: first, through the denominator term the connectivity constraint weights are standardized, and this form takes into account the overall connection strength of nodes; second, a correction term based on node importance is added, and the correction strength is adjusted by the normalization coefficient μ; finally, a normalization error term ε8 is introduced to consider the uncertainty in the calculation process.

[0141] The determination process of the normalization coefficient μ needs to consider the balance of the network structure: by analyzing the connection distribution characteristics of nodes with different importance levels in the network; setting different μ values for normalization processing and evaluating the structural characteristics of the processed network; selecting the optimal μ value that can highlight the role of important nodes while maintaining the overall network structure; verifying the rationality of the parameter selection by comparing the functional enrichment analysis results of the reconstructed network under different parameter conditions.

[0142] The range determination of the normalization error ε8 is based on the accuracy analysis of numerical calculations: considering the numerical rounding error and cumulative error in the calculation process; analyzing the stability of the normalization results in networks of different scales; evaluating the impact of error propagation on the network structure; based on these analyses, the value range of ε8 is determined to be 0.001 - 0.01.

[0143] 9. Network node distance equation is constructed based on the physicochemical properties of protein molecules: first, considering the influence of molecular weight and steric hindrance, the relative size of two protein molecules is characterized by the term, where m i , m j are the molecular weights, and s i , s j are the steric hindrance coefficients; second, considering the contribution of interaction sites, the number of interaction sites is represented by p ij ; the weights of different factors are adjusted by the parameters θ1 and θ2; finally, a distance error term ε9 is introduced to consider the uncertainty in the spatial distance estimation.

[0144] The optimization of the distance adjustment coefficients θ1 and θ2 is based on protein structure analysis: by analyzing the structures of known protein complexes in the PDB database, geometric features of the interaction interface are extracted; a relationship model between molecular weight, steric hindrance coefficient, the number of interaction sites, and the actual contact distance is established; a multiple regression method is used to optimize the parameters to best match the predicted minimum allowable distance with the experimental observation results; the reliability of the parameters is verified through cross-validation at the structural level.

[0145] The range setting of the distance error ε9 is based on the accuracy evaluation of structural analysis: considering the dynamic characteristics and conformational changes of protein structures; analyzing the range of distance changes at the interaction interfaces of different types of proteins; evaluating the influence of environmental factors such as temperature and pH on the distance between proteins; considering these factors comprehensively, the value range of ε9 is determined to be 0.5 - 2.0 nanometers.

[0146] 10. Network edge weight correlation equation is constructed based on the principle of local network structure consistency: first, the weight relationship of adjacent edges is considered, and local normalization of the co-localization coefficient and interaction strength is achieved through the denominator term where Ω j , Ω j represents the set of nodes adjacent to nodes i and j; secondly, the ratio term of node importance is introduced. This form not only considers the overall importance of the node pair but also avoids excessive importance differences through maximum normalization; the relative contributions of the two effects are balanced by the weight coefficients ρ1 and ρ2; finally, a constraint error term ε 10 is added to consider the uncertainty of weight ratio estimation.

[0147] The optimization process of the weight coefficients ρ1 and ρ2 is based on the stability analysis of the local network structure: first, test network subgraphs of different scales are constructed, including known functional modules and random sub-networks; the local characteristics of the edge weight distribution are analyzed, including weight gradient, aggregation degree, etc.; the parameter combination is optimized through the simulated annealing algorithm to make the network not only maintain the local correlation of edge weights but also reflect the influence of node importance; finally, the optimal parameter range is determined through cross-validation. Usually, the value range of ρ1 and ρ2 is 0.3 - 0.7.

[0148] The constraint error ε 10 The range determination needs to consider multiple sources of uncertainty: including measurement errors of co-localization coefficients and interaction strength, uncertainties in node importance calculation, changes in local network structure, etc.; the combined influence of these factors is evaluated through error propagation analysis; combined with the observed weight ratio fluctuation range in actual applications; finally, the value range of ε 10 is determined to be 0.01 - 0.03.

[0149] 11. Node degree correlation equation is constructed by comprehensively considering multiple biological features: First, the exponential term is used to describe the impact of the difference in expression levels on the difference in node degrees, where E max is the maximum expression abundance in the dataset; Second, the ratio term is used to reflect the relative difference in the number of interactions; Third, the term is used to consider the combined impact of local network density and spatial distance; The relative contributions of each term are balanced by the adjustment coefficients σ1, σ2, σ3; Finally, the degree difference error term ε 11 is added to consider the uncertainty in the estimation process.

[0150] The optimization of the adjustment coefficients σ1, σ2, σ3 is based on the characteristic analysis of known protein interaction networks: First, collect protein interaction network data of multiple species, and analyze the correlation between node degrees and expression levels and the number of interactions; Construct a prediction model based on these characteristics, and determine the optimal parameter combination through the maximum likelihood estimation method; Through cross-validation of different types of networks, verify the universality of the parameters; Finally, determine the optimal parameter combination that satisfies the constraint of σ1 + σ2 + σ3 = 1.

[0151] The range of the degree difference error ε 11 is set based on the analysis of network dynamic characteristics: Consider the impact of the spatio-temporal changes in protein expression levels on the network structure; Analyze the range of changes in the degree difference observed under different experimental conditions; Evaluate the stability of the degree difference during network reconstruction; Considering these factors, the value range of ε 11 is determined to be 0.02 - 0.06.

[0152] The overall design of these equations reflects multiple key considerations in biological network construction: First, improve the reliability of network construction through multi-level integration of experimental data; Second, ensure the biological rationality of the network structure through multiple constraint equations; Third, ensure the stability and repeatability of network reconstruction through parameter optimization and error control; Finally, identify key nodes through the minimum dominating set algorithm, providing important clues for subsequent functional research. The innovation of the entire equation system lies in: 1) The non-linear integration method of multi-source data; 2) The spatial constraint based on physical and chemical principles; 3) The weight optimization considering the local network structure; 4) The introduction of a dynamic adjustment mechanism for node importance. These features enable this method to more accurately describe the structural and functional characteristics of protein interaction networks.

[0153] In the prior art, the construction of protein interaction networks for disease markers mainly relies on the direct determination of experimental data or the mining and analysis of databases. Although traditional experimental methods such as the yeast two-hybrid system and co-immunoprecipitation technology can obtain protein interaction information, these methods often require a large amount of experimental work, and there are differences between the experimental conditions and the real physiological environment, resulting in the possibility that the obtained interaction data may not conform to the actual situation. The method based on database mining, on the other hand, overly relies on existing data, making it difficult to discover new interaction relationships and unable to reflect the tissue specificity and dynamic change characteristics under disease states. In addition, although the prediction method based on sequence homology is simple to operate, since it only considers the one-dimensional sequence information of proteins and ignores the complex interactions of proteins in spatial structures and physiological environments, the reliability of the prediction results is relatively low.

[0154] The method for establishing a protein interaction network for disease markers proposed by the present invention has significant technological innovation. This method systematically integrates the physicochemical properties, dynamic interaction information, and expression abundance data of proteins for the first time, and realizes the precise construction of the interaction network through the design of an innovative network optimization equation set. In the design of the equation set, by introducing the consideration of steric hindrance effects, the feasibility of the predicted interactions in the three-dimensional structure is ensured; by the constraint of expression abundance correlation, the rationality of the interactions at the biological level is ensured; through the dynamic imaging data obtained by real-time fluorescence confocal microscopy, not only the spatial co-localization of proteins is verified, but also a quantitative basis for the interaction strength is provided, thus realizing the precise calculation of the network edge weights.

[0155] The core advantage of the method of the present invention lies in its innovative network optimization equation set, which realizes multiple optimizations in the network construction process through the synergistic action of multiple key equations. The network reconstruction equation ensures the consistency between the network structure and experimental data, the network sparsity constraint equation effectively controls the connection density of the network, avoiding the false positive problem caused by over-connection. The network connectivity constraint ensures the integrity of information transmission, while the node distance constraint makes the network structure more in line with the actual biological characteristics. The introduction of the network edge weight correlation equation and the node degree correlation equation further improves the rationality and reliability of the network structure. This multi-dimensional optimization design enables the constructed network to better reflect the characteristics of the actual biological system.

[0156] In terms of actual application effects, the method of the present invention shows significant advantages. By considering the steric hindrance effect and integrating dynamic imaging data, this method can more accurately predict the interactions between proteins, greatly reducing the occurrence of false positives. At the same time, due to the integration of tissue-specific expression data and dynamic interaction information, this method can reflect the specific changes in the protein interaction network under disease states, providing a more reliable basis for disease mechanism research and drug target screening.

[0157] A specific Example 1 of the present invention is provided below. The specific implementation manners of each step in this Example 1 are described in detail as follows: The specific implementation manner of step S10 is: First, collect human tissue samples from the test subjects, such as the liver, kidneys, etc. Then, extract the total protein using conventional biochemical methods. Next, measure the concentration of the total protein using the Bradford protein quantification method and adjust it to the range of 50 to 100 micrograms per milliliter. After that, use physical and chemical parameters such as the molecular weight (m i ) and steric hindrance coefficient (s i ) to measure and analyze the characteristics of the total protein. Specifically, it includes measuring the number of interaction sites (p ij ) of each protein. This information provides important basic data for subsequent network construction. The purpose of this step is to obtain the target protein sample and measure its basic physical and chemical characteristics, laying a foundation for subsequent analysis and network construction.

[0158] The specific implementation manner of step S20 is: First, prepare specific monoclonal antibodies against the target disease biomarker protein. Then, use these specific antibodies to enrich and separate the target protein from the total protein sample by the method of magnetic bead-coupled affinity chromatography. Next, measure the binding degree value (B ij ) between the target protein and the specific antibody, and generate a protein binding degree matrix accordingly. The calculation formula of this binding degree matrix is: In the formula, C i and C j are the concentrations of the i-th and j-th proteins respectively, d ij is the spatial distance between them, r ij is their van der Waals interaction radius, r0 is the standard interaction distance (taking 10 nanometers), k1 and k2 are binding constants, and ε1 is the experimental error (taking values in the range of 0.01 to 0.05). This binding degree matrix reflects the mutual binding strength between different proteins and provides key parameters for the subsequent construction of network topology.

[0159] The specific implementation of step S30 is as follows: First, use liquid chromatography-mass spectrometry (LC-MS / MS) technology to identify the components of the target disease biomarker proteins enriched in step S20, and obtain qualitative and quantitative data of the protein components. Based on these protein component data, an initial interaction matrix is constructed. The calculation formula is as follows: In the formula, I ij is the number of interactions between the i-th and j-th proteins, and are the concentrations of the i-th and j-th proteins in the k-th component respectively, d ij and r ij are the same as before, and ε2 is the calculation error (ranging from 0.01 to 0.05). These data lay the foundation for the initial construction of network nodes and edges.

[0160] The specific implementation of step S40 is as follows: First, measure the expression abundance value (E i ) of each protein component. Then, based on these expression abundance data, calculate the importance vector (N i ) of each node (protein). The calculation formula of the node importance vector is as follows: In the formula, D i is the degree of the i-th node, I i is the number of interactions of the i-th protein, α1, α2, and α3 are weight coefficients, and satisfy α1 + α2 + α3 = 1, ε2 is the calculation error (ranging from 0.001 to 0.01), and n is the total number of nodes. This node importance vector provides a key reference basis for subsequent network optimization.

[0161] The specific implementation of step S50 is as follows: First, use a real-time fluorescence confocal microscope to dynamically image and observe the protein components enriched in step S20. By analyzing the obtained fluorescence intensity time series data (F it ), calculate the colocalization coefficient matrix (M ij ) between each protein component. The calculation formula of the colocalization coefficient matrix is: In the formula, and are the average fluorescence intensities of the i-th and j-th proteins respectively, T is the total number of time points, and ε3 is the measurement error (ranging from 0.005 to 0.02). This colocalization coefficient reflects the degree of colocalization of proteins in time and space, and provides an important parameter for calculating the weights of network edges.

[0162] The specific implementation of step S60 is as follows: First, based on the fluorescence intensity time series data obtained in step S50, calculate the interaction strength matrix (S ij)。 The calculation formula for the interaction strength matrix is as follows: In the formula, M ij is the co-localization coefficient, B ij is the degree of binding, E i and E j are the protein expression abundances, E0 is the normalization parameter, β1, β2, and β3 are the weight coefficients, and ε4 is the comprehensive error (ranging from 0.01 to 0.05). Then, combining with the network topology matrix (T ij ) obtained in the previous steps, the initial network topology structure is constructed. Next, calculate the total number of edges (E) and local density (D l ) and other indicators of this initial network topology. This information provides a benchmark reference for subsequent network optimization.

[0163] The specific implementation of step S70 is: construct a multi-variable optimization equation set to optimize the network. This equation set includes the following key equations:

[0164] Network reconstruction equation: In the formula, W ij is the optimized network connection weight, γ1, γ2, and γ3 are the weight coefficients, d ij is the spatial distance between proteins i and j, and ε5 is the reconstruction error (ranging from 0.01 to 0.03). This equation comprehensively considers factors such as interaction strength, spatial correlation, and network topology structure, so as to obtain a network connection weight closer to the actual situation.

[0165] Network sparsity constraint equation: In the formula, L ij is the weight after sparsity constraint, λ is the sparsity penalty coefficient (ranging from 0.1 to 1.0), δ1 is the adjustment coefficient, and ε6 is the constraint error (ranging from 0.005 to 0.02). This equation effectively controls the connection density of the network and avoids the false positive problem caused by over-connection.

[0166] Network connectivity constraint equation: In the formula, C ij is the weight after connectivity constraint, η1 and η2 are the connectivity adjustment coefficients, and ε7 is the connectivity error (ranging from 0.02 to 0.05). This equation ensures that information can be smoothly transmitted between each node and guarantees the overall connectivity of the network.

[0167] By solving these optimization equations, the final network weight matrix (F ij ) can be generated. This network weight matrix reflects the interaction strength between each edge after optimization and is the core of constructing the interaction network of disease markers.

[0168] The specific implementation of step S80 is as follows: The minimum dominating set algorithm is used to analyze the optimized network weight matrix (F ij ). This algorithm includes the following steps: 1) Calculate the domination degree of each protein node based on the network weight matrix; 2) Select the node with the largest domination degree as the candidate key node; 3) Add the candidate key node to the domination set and update the set of undominated nodes; 4) Repeat steps 1-3 until all nodes are dominated; 5) Determine the nodes in the domination set as the key protein nodes. This algorithm can effectively identify the key nodes that play important roles in the disease biomarker interaction network.

[0169] The specific implementation of step S90 is as follows: Based on the key protein nodes identified in the foregoing steps, a tissue-specific disease biomarker protein interaction network is constructed. This network can reflect the specific changes in protein interactions under disease states and provides a reliable basis for disease mechanism research and drug target discovery.

[0170] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Table 1 below.

[0171] Table 1 Variable Explanation Table

[0172]

[0173] The following provides Example 2 of a specific application scenario of the present invention:

[0174] In this example, hepatocellular carcinoma (HCC) is used as the research object, aiming to construct a HCC disease biomarker protein interaction network.

[0175] Sample collection and protein extraction First, liver tissue samples are collected from clinical HCC patients and healthy controls respectively. Total proteins are extracted using conventional biochemical methods. The total protein concentration is measured by the Bradford protein quantification method and adjusted to the range of 50-100 micrograms per milliliter. Then, the molecular weight (m i ) and steric hindrance coefficient (s i ) of each protein, as well as the number of interaction sites (p ij ) are measured. The relevant data are shown in Table 2.

[0176] Table 2 Protein Physicochemical Property Data

[0177] Protein <![CDATA[Molecular weight m i (kDa)]]> <![CDATA[Steric hindrance coefficient s i > <![CDATA[Number p of interaction sites ij > P1 55.8 1.2 12 P2 42.3 1.5 8 P3 68.1 1.4 10 P4 37.2 1.3 7 P5 75.4 1.6 14 … … … …

[0178] As Figure 2The figure shows the relationship among three key parameters of the five proteins P1 - P5 in Table 2: molecular weight (bar graph), steric hindrance coefficient (line graph), and number of interaction sites (numerical annotation).

[0179] For the determination of target protein enrichment and binding degree, specific monoclonal antibodies against HCC marker proteins were used, and the target proteins were enriched and separated from the total protein sample by magnetic bead - coupled affinity chromatography. Then, the binding degree (B ij ) between each pair of target proteins and the specific antibody was measured, and a binding degree matrix was generated. The binding degree calculation formula is: where C i and C j are the concentrations of the i - th and j - th proteins respectively, d ij is the spatial distance between them, r ij is their van der Waals interaction radius, r0 is the standard interaction distance (10 nanometers), k1 and k2 are binding constants, and ε1 is the experimental error (in the range of 0.01 - 0.05). Some binding degree data are shown in Table 3.

[0180] Table 3 Protein Binding Degree Matrix

[0181] P1 P2 P3 P4 P5 P1 - 0.82 0.91 0.75 0.88 P2 0.82 - 0.78 0.61 0.84 P3 0.91 0.78 - 0.69 0.93 P4 0.75 0.61 0.69 - 0.72 P5 0.88 0.84 0.93 0.72 -

[0182] Based on the binding degree calculation formula in the examples, as Figure 3 shown, the relationship between the binding degree between proteins and the spatial distance is presented, and the complete theoretical formula is annotated.

[0183] For protein component identification and interaction analysis, liquid chromatography - mass spectrometry (LC - MS / MS) technology was used to identify and quantitatively analyze the components of the enriched HCC marker proteins. The expression abundance data (E i ) of each protein component were obtained. Meanwhile, the number of interactions (I ij ) between each pair of proteins was calculated, and the relevant data are shown in Table 4.

[0184] Table 4 Protein Component Data and Number of Interactions

[0185]

[0186] The protein expression abundance in Table 4 was correlated with the total number of interactions. As Figure 4 shown, the correlation between the two was presented, and the overall change trend was reflected by the trend line.

[0187] For dynamic interaction imaging and analysis, real - time fluorescence confocal microscopy was used to dynamically image and observe the enriched protein components. The fluorescence intensity data (Fit )。Based on these fluorescence intensity time series data, the colocalization coefficient matrix (M ij ) between each protein component is calculated. Among them, and are the average fluorescence intensities of the i-th and j-th proteins respectively, T is the total number of time points, and ε3 is the measurement error (taking values in the range of 0.005 - 0.02). Some colocalization coefficient data are shown in Table 5.

[0188] Table 5 Protein colocalization coefficient matrix

[0189] P1 P2 P3 P4 P5 P1 - 0.78 0.84 0.72 0.81 P2 0.78 - 0.75 0.68 0.77 P3 0.84 0.75 - 0.65 0.85 P4 0.72 0.68 0.65 - 0.69 P5 0.81 0.77 0.85 0.69 -

[0190] Based on the colocalization coefficient calculation formula in the embodiment, as Figure 5 shown, the correlation scatter plot of the fluorescence intensities of two proteins is presented, and the complete mathematical formula is marked.

[0191] Then, based on these colocalization coefficient data, combined with the protein expression abundance and binding degree information obtained above, the interaction intensity matrix (S ij ) between each protein component is calculated. The interaction intensity calculation formula is: Among them, β1, β2, and β3 are weight coefficients, E0 is the normalization parameter of expression abundance, and ε4 is the comprehensive error (taking values in the range of 0.01 - 0.05). Some interaction intensity data are shown in Table 6.

[0192] Table 6 Protein interaction intensity matrix

[0193] P1 P2 P3 P4 P5 P1 - 0.76 0.82 0.69 0.79 P2 0.76 - 0.72 0.63 0.75 P3 0.82 0.72 - 0.61 0.83 P4 0.69 0.63 0.61 - 0.65 P5 0.79 0.75 0.83 0.65 -

[0194] Based on the various data obtained above, a system of equations containing multiple optimization equations such as network reconstruction, sparse constraint, and connectivity guarantee is constructed for network optimization. By solving these equations, the final network weight matrix (F ij ) is obtained.

[0195] Network reconstruction equation: Among them, γ1, γ2, and γ3 are weight coefficients, d ij is the spatial distance between proteins i and j, and ε5 is the reconstruction error (taking values in the range of 0.01 - 0.03). This equation comprehensively considers factors such as interaction intensity, spatial correlation, and network topology.

[0196] Network sparse constraint equation: Among them, λ is the sparse penalty coefficient (ranging from 0.1 to 1.0), δ1 is the adjustment coefficient, and ε6 is the constraint error (ranging from 0.005 to 0.02). This equation effectively controls the connection density of the network.

[0197] Network connectivity constraint equation: Among them, η1 and η2 are the connectivity adjustment coefficients, and ε7 is the connectivity error (ranging from 0.02 to 0.05). This equation ensures the overall connectivity of the network.

[0198] By solving these optimization equations, the final network weight matrix (F ij ) was obtained. Part of the data is shown in Table 7. This weight matrix reflects the interaction strength between each edge after optimization and is the basis for constructing the HCC disease biomarker interaction network.

[0199] Table 7 Final network weight matrix

[0200] P1 P2 P3 P4 P5 P1 - 0.71 0.77 0.64 0.74 P2 0.71 - 0.67 0.59 0.70 P3 0.77 0.67 - 0.57 0.78 P4 0.64 0.59 0.57 - 0.61 P5 0.74 0.70 0.78 0.61 -

[0201] Key node identification and network construction: The minimum dominating set algorithm was used to analyze the optimized network weight matrix, and three key protein nodes, P1, P3, and P5, were identified.

[0202] Finally, based on these three key nodes, the HCC disease biomarker protein interaction network was constructed. This network can not only reflect the dynamic change characteristics of protein interactions under the HCC state but also provide a reliable basis for disease mechanism research and drug target discovery.

[0203] Technical principle: The key to the method of the present invention lies in fully integrating the multi-dimensional information of proteins and accurately constructing the interaction network through an innovative mathematical model. First, this method collects human tissue samples, extracts total proteins and determines their physicochemical properties, such as molecular weight, steric hindrance coefficient, and the number of interaction sites, etc. These parameters provide the basic data for subsequent network construction. Secondly, specific antibodies against disease biomarker proteins are used to enrich and isolate the target proteins by affinity chromatography. Then, the binding degree between the target protein and the antibody is measured to generate a protein binding degree matrix. This lays the foundation for the initial construction of the network topology. Next, mass spectrometry technology is used to identify and quantitatively analyze the enriched protein components to obtain information such as protein expression abundance. At the same time, a real-time fluorescence confocal microscope is used to image the protein dynamics and measure the time series data of its fluorescence intensity. These data provide the basis for calculating the spatial co-localization coefficient and interaction strength of proteins. Based on the above-mentioned multi-dimensional information obtained, the present invention constructs a system of equations containing multiple optimization equations such as network reconstruction, sparse constraint, and connectivity guarantee. By solving these equations, not only can the final weight matrix of the network be obtained, but also the rationality and reliability of the network structure can be ensured. In addition, the present invention also uses the minimum dominating set algorithm to analyze the optimized network and effectively identify the key protein nodes. Finally, a tissue-specific disease biomarker interaction network is constructed based on these key nodes. In summary, the innovation of the method of the present invention lies in systematically integrating the multi-dimensional information of proteins and achieving the optimized construction of the network by constructing a fine mathematical model. This not only ensures the rationality of the network structure but also reflects the dynamic characteristics of the network under disease conditions, providing a reliable analysis tool for research in related fields.

[0204] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.

Claims

1. A method for establishing a protein interaction network of a disease biomarker, characterized in that, Including the following steps: Collect human tissue samples and extract total proteins, and perform protein quantitative determination; Use magnetic bead-coupled specific antibodies for affinity chromatography to enrich target disease marker proteins; Adopt liquid chromatography-mass spectrometry technology for component identification; determine the expression abundance of protein components; use a real-time fluorescence confocal microscope for dynamic imaging; calculate the interaction strength matrix and construct an initial network topology; construct and solve a network optimization equation set that comprehensively considers protein steric hindrance effects, expression abundance correlation, colocalization coefficient, and interaction strength; use the minimum dominating set algorithm to identify key protein nodes; construct a tissue-specific disease marker protein interaction network.

2. The method according to claim 1, wherein The step of collecting human tissue samples and extracting total proteins includes: using the Bradford protein quantification method to measure the total protein concentration, adjusting the total protein concentration to 50 to 100 micrograms per milliliter, measuring the protein steric hindrance coefficient and the protein molecular weight, and measuring the number of protein interaction sites.

3. The method according to claim 1, wherein The step of using magnetic bead-coupled specific antibodies for affinity chromatography to enrich target disease marker proteins includes: measuring the binding degree value of the target disease marker protein and the specific antibody to generate a protein binding degree matrix.

4. The method according to claim 1, wherein The step of adopting liquid chromatography-mass spectrometry technology for component identification includes: obtaining protein component data, constructing an initial interaction matrix based on the protein component data, and calculating the number of interactions.

5. The method according to claim 1, characterized in that The step of determining the expression abundance of protein components includes: obtaining protein expression abundance values and calculating the node importance vector according to the protein expression abundance values.

6. The method according to claim 1, wherein The step of using a real-time fluorescence confocal microscope for dynamic imaging includes: obtaining fluorescence intensity time series data and calculating a protein colocalization coefficient matrix based on the fluorescence intensity time series data.

7. The method according to claim 1, characterized in that The step of constructing the initial network topology includes: obtaining a network topology matrix and calculating the total number of network edges and network local density of the initial network topology.

8. The method according to claim 1, characterized in that, The network optimization equation set includes: a network reconstruction equation, a network sparsity constraint equation, a network connectivity constraint equation, a network weight normalization equation, a network node distance equation, a network edge weight correlation equation, and a node degree correlation equation.

9. The method according to claim 1, characterized in that, The minimum dominating set algorithm includes the following steps: Calculate the domination degree of each protein node based on the final network weight matrix; Select the protein node with the largest domination degree as the candidate key node; Add the candidate key node to the domination set and update the set of undominated nodes; repeat the above steps until all protein nodes are dominated; Determine the protein nodes in the domination set as the key protein nodes.

10. The method according to claim 8, wherein The input and output of each equation of the network optimization equation set include: The input of the network reconstruction equation includes the co-location coefficient matrix, the interaction strength matrix, the network topology matrix, and the network reconstruction error threshold, and the output is the network connection weight matrix; the input of the network sparsity constraint equation includes the network connection weight matrix, the network local density, the node importance vector, the network sparsity penalty coefficient, and the total number of edges in the network, and the output is the sparsity constraint weight matrix; the input of the network connectivity constraint equation includes the sparsity constraint weight matrix, the network topology matrix, the network connectivity threshold, and the node importance vector, and the output is the connectivity constraint weight matrix.