Disease marker subtype classification determination method

Through the combination of bidirectional electrophoresis and graph neural network, a protein interaction network is constructed, and the embedded neural network model is used to classify the subtype of disease markers, which solves the problem of insufficient recognition in the existing technology, and achieves high-precision classification of disease subtypes and molecular pathway feature recognition.

CN120279987APending Publication Date: 2025-07-08QINGDAO RAISECARE BIOTECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510438320.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing classification methods for the subtype of disease marker are not highly recognized enough to meet the needs of precision medicine. Traditional methods are difficult to fully reflect the interaction between disease molecular mechanisms and neglect proteins.

Method used

Bidirectional electrophoresis technology is used to separate protein components, combine mass spectrometry analysis and graph neural network methods to build a protein interaction network, feature extraction and clustering is performed through embedded neural network models, and combine bioinformatics analysis and functional annotation to achieve high-precision classification of subtypes of disease markers.

Benefits of technology

It improves the recognition and accuracy of the classification of subtypes of disease markers, provides more reliable data basis and molecular mechanism analysis, and supports precise medical treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279987A_ABST
    Figure CN120279987A_ABST
Patent Text Reader

Abstract

The invention provides a disease marker subtype classification determination method, and belongs to the technical field of disease marker subtype classification, and the method comprises the following steps: firstly, extracting protein from a patient sample, separating through a dimensional electrophoresis technology, and obtaining a protein expression matrix through Coomassie brilliant blue dyeing and an image analysis system; differentially expressed proteins are determined, and sequence information is obtained through mass spectrum identification. And extracting network node features by using a graph neural network method, and carrying out function annotation analysis and path identification. Patient samples are grouped by using a multi-level clustering method, and feature extraction and optimization are realized by training an embedded neural network model. And inputting the extracted features into a pre-trained function enrichment analysis model, and determining the molecular pathway features of the disease subtype. And finally, the reliability of a classification result is confirmed through an immunoblotting technology and independent patient queue verification, subtype classification of the disease markers is completed, and the problem that in the prior art, the recognition degree of subtype classification of the disease markers is not high enough is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of disease biomarker subtype classification, and more particularly, relates to a method for determining disease biomarker subtype classification. Background Art

[0002] Disease biomarker subtype classification can not only help doctors diagnose diseases more accurately, but also provide more targeted treatment plans for patients. Currently, in clinical practice, proteomic methods are mainly used to analyze and classify disease biomarkers. Traditional disease biomarker subtype classification methods mainly rely on single-dimensional protein expression data analysis. Common technical means include two-dimensional electrophoresis, mass spectrometry analysis, and Western blotting, etc. These methods have the following limitations in practical applications: First, single protein expression data is difficult to comprehensively reflect the molecular mechanism of diseases and is prone to one-sidedness of classification results; Second, traditional data analysis methods often ignore the interaction relationships between proteins and cannot reveal the complex network regulation mechanism of disease occurrence and development; Third, existing classification methods lack in-depth analysis of protein functions and biological pathways and are difficult to accurately identify the molecular pathway characteristics of different subtypes.

[0003] That is to say, there is a problem in the prior art that the recognition degree of disease biomarker subtype classification is not high enough, which is difficult to meet the high requirements of precision medicine for disease subtype classification. Summary of the Invention

[0004] In view of this, the present invention provides a method for determining disease biomarker subtype classification, which can solve the problem that the recognition degree of disease biomarker subtype classification in the prior art is not high enough.

[0005] The present invention is implemented as follows:

[0006] The present invention provides a method for determining disease biomarker subtype classification, including: extracting proteins from patient samples in a training set, separating protein components using two-dimensional electrophoresis technology, recording and obtaining a first protein expression matrix; displaying protein spots in the protein components by Coomassie Brilliant Blue staining, performing quantitative analysis on the protein spots, and recording and obtaining a second protein expression matrix; performing dimensionality reduction analysis on the second protein expression matrix, and preliminarily determining the number of differentially expressed proteins through the main component distribution characteristics; selecting protein spots with significant expression differences for mass spectrometry identification to obtain differential protein sequence information; constructing a protein interaction network, using a graph neural network method to extract network node features, and obtaining protein feature vectors; performing functional annotation analysis and pathway identification on the differential protein sequence information, and transforming the protein feature vectors into low-dimensional features; grouping patient samples using a multi-level clustering method based on the low-dimensional features, and training an embedded neural network model; using the embedded neural network model to extract features from new patient samples to determine the molecular pathway features of disease subtypes; verifying the expression levels of key proteins to obtain the disease biomarker subtype classification result. Specifically, it includes the following steps:

[0007] S10. Extract proteins from patient samples in a training set, separate protein components using two-dimensional electrophoresis technology, and obtain a first protein expression matrix through standardization and normalization processing;

[0008] S20. Display protein spots in the protein components by Coomassie Brilliant Blue staining, perform quantitative analysis on the protein spots using an image analysis system, and obtain a second protein expression matrix;

[0009] S30. Perform dimensionality reduction analysis on the second protein expression matrix, and preliminarily determine the number of differentially expressed proteins through the main component distribution characteristics;

[0010] S40. Select protein spots with significant expression differences for mass spectrometry identification to obtain differential protein sequence information;

[0011] S50. Construct a protein interaction network, use the differential protein sequence information as network nodes and protein-protein interactions as network connections, use a graph neural network method to extract network node features, and obtain protein feature vectors;

[0012] S60. Use bioinformatics analysis methods to perform functional annotation analysis and pathway identification on the differential protein sequence information, and use a feature dimensionality reduction method to transform the protein feature vectors into low-dimensional features;

[0013] S70. Group the patient samples using a multi-level clustering method based on the low-dimensional features, construct a protein expression reconstruction loss function and a protein feature vector loss function, train an embedded neural network model, and optimize the embedded neural network model through a cross-validation method;

[0014] S80. Use the embedded neural network model to extract features from new patient samples, input the extracted low-dimensional features into a pre-trained functional enrichment analysis model, and determine the molecular pathway features of the disease subtype;

[0015] S90. Verify the key protein expression levels using immunoblotting technology, perform a correlation analysis in combination with the patient clinical phenotype data, and verify the stability of the molecular pathway features through an independent patient cohort to obtain the disease biomarker subtype classification results.

[0016] Among them, the training set patient samples include at least 100 patient samples, including at least serum samples, tissue samples, peripheral blood mononuclear cell samples, and whole blood samples of diagnosed patients.

[0017] Use two-dimensional electrophoresis technology to separate protein components. Specifically, in the first dimension, isoelectric focusing is used to separate according to the isoelectric point of proteins, and in the second dimension, sodium dodecyl sulfate polyacrylamide gel electrophoresis is used to separate according to the molecular weight;

[0018] Show the protein spots in the protein components by Coomassie brilliant blue staining. Specifically, place the gel in Coomassie brilliant blue staining solution for 240 minutes, and use decolorizing solution for decolorization until the protein spots are shown;

[0019] Use an image analysis system to perform quantitative analysis on the protein spots. Specifically, use a gel imaging device to obtain an image, measure the optical density value of the protein spots through the image analysis system, perform background correction and normalization analysis, and form the second protein expression matrix;

[0020] Perform dimensionality reduction analysis on the second protein expression matrix. Specifically, use the principal component analysis method to map the second protein expression matrix to a low-dimensional principal component space;

[0021] Preliminarily determine the number of differentially expressed proteins through the principal component distribution characteristics. Specifically, analyze the aggregation trend of the data distribution in the low-dimensional principal component space, and determine the number of clusters in combination with the principal component contribution rate;

[0022] Select the protein spots with significant expression differences for mass spectrometry identification. Specifically, cut the protein spots from the gel, perform enzymatic digestion, and use liquid chromatography-mass spectrometry to perform peptide analysis and protein sequence identification;

[0023] Construct a protein - protein interaction network. Specifically, based on a protein - protein interaction database, construct the differential protein sequence information as network nodes and the protein - protein interaction relationships as network connections.

[0024] Use a graph neural network method to achieve network node feature extraction. Specifically, use a graph convolutional network to propagate and aggregate the network node features and learn to obtain the protein feature vectors.

[0025] Use bioinformatics analysis methods to perform functional annotation analysis and pathway identification on the differential protein sequence information. Specifically, compare the differential protein sequence information with a biological function annotation database to obtain gene ontology annotation information and biological pathway information.

[0026] Use a feature dimensionality reduction method to transform the protein feature vectors into low - dimensional features. Specifically, use the stochastic neighbor embedding method to transform the protein feature vectors into three - dimensional feature representations.

[0027] Based on the low - dimensional features, use a multi - level clustering method to group patient samples. Specifically, construct a patient sample similarity matrix and perform iterative clustering through the hierarchical clustering method.

[0028] Construct a protein expression reconstruction loss function and a protein feature vector loss function. Specifically, use the reconstruction error of the second protein expression matrix and the spatial distance of the protein feature vectors as optimization objectives to form an overall loss function.

[0029] The functional enrichment analysis model specifically performs pathway enrichment calculation on the low - dimensional features, obtains the significance scores of gene function sets, and identifies significant biological pathways.

[0030] The functional enrichment analysis model adopts a deep attention neural network structure, which consists of multiple functional units. At the input layer, it receives the low - dimensional feature vectors after dimensionality reduction processing, and calculates the association strength between the features and biological pathway nodes through an attention mechanism. The attention layer contains a multi - head self - attention module, which can capture the complex dependence relationships between features. Under the guidance of the attention weights, the enrichment calculation layer scores each biological pathway and uses an improved hypergeometric distribution test method to evaluate the enrichment significance. Finally, at the output layer, it generates pathway enrichment results, including statistical indicators such as significance scores and enrichment multiples.

[0031] The process of constructing the training dataset for the model is as follows: First, collect protein expression profile data with known functional annotations from public databases, and use standardized two-dimensional electrophoresis technology to obtain proteome expression information. At the same time, obtain the correspondence between proteins and biological pathways from databases such as STRING and KEGG, and establish a protein-pathway association network. Conduct quality control on the collected expression profile data to remove batch effects and technical noise. Then integrate the processed protein expression data with pathway annotation information to construct training sample pairs. Considering the balance of data distribution, use stratified sampling to divide the dataset into a training set and a validation set at a ratio of 8:2. In addition, an independent test set is also constructed to evaluate the generalization performance of the model.

[0032] During the model training process, first randomly initialize the network parameters, and use the Xavier method to optimize the initial weight distribution. The training uses the mini-batch stochastic gradient descent method, and in each iteration cycle, input a batch of training samples for forward calculation. Evaluate the model performance by calculating the cross-entropy loss between the predicted pathway enrichment results and the true annotations. The loss function takes into account both the accuracy of pathway significance prediction and the estimation error of enrichment fold. Update the model parameters using the Adam optimizer based on the calculated gradient information. To prevent overfitting, introduce Dropout and L2 regularization strategies. Stop training when the performance on the validation set no longer improves or reaches the preset number of iterations. Finally, evaluate the prediction effect of the model on the independent test set, and verify the reliability of important pathways through biological experiments.

[0033] The structure of the described embedded neural network model includes:

[0034] A three-layer encoding network for compressing the second protein expression matrix into a feature representation;

[0035] A three-layer decoding network for reconstructing the feature representation into protein expression data;

[0036] A first mathematical model unit, including a feature mapping function, a weight optimization function, and a loss calculation function;

[0037] A second mathematical model unit, including a feature reconstruction function, an error correction function, and a feedback regulation function;

[0038] The specific content of the first mathematical model unit includes:

[0039] The feature mapping function is used to transform the second protein expression matrix into a network feature representation, with the second protein expression matrix as the input and the initial feature mapping result as the output;

[0040] The weight optimization function is used to adjust the network connection parameters, with the initial feature mapping result and the patient grouping identifier as the input and the optimized network parameters as the output;

[0041] The loss calculation function is used to evaluate the model performance. The input is the model prediction grouping and the actual grouping of patients, and the output is the loss calculation value;

[0042] The second mathematical model unit specifically includes:

[0043] The feature reconstruction function is used for data reconstruction. The input is the output data of the three-layer decoding network, and the output is the reconstructed protein expression data;

[0044] The error correction function is used to eliminate the reconstruction deviation. The input is the reconstructed protein expression data and the second protein expression matrix, and the output is the corrected protein expression data;

[0045] The feedback regulation function is used for parameter optimization. The input is the corrected protein expression data and the second protein expression matrix, and the output is the network regulation signal.

[0046] Compared with the prior art, the beneficial effects of a method for determining disease biomarker subtype classification provided by the present invention are as follows: By innovatively integrating multiple technical modules, the problems existing in the prior art are effectively solved. First, the present invention adopts two-dimensional electrophoresis technology combined with mass spectrometry analysis to achieve high-throughput and high-precision acquisition of proteomics data. Through standardization processing and normalization processing, the data quality and comparability are significantly improved.

[0047] At the data analysis level, the present invention innovatively constructs a feature extraction method based on graph neural network. By introducing protein interaction network information into the analysis framework, it can not only capture the complex interaction relationships between proteins, but also deeply explore the biological information contained in the network structure. This method significantly improves the effect of feature extraction and provides a more reliable data basis for subsequent classification analysis.

[0048] The embedded neural network model designed by the present invention has a unique dual optimization mechanism. By simultaneously optimizing the protein expression reconstruction loss and the feature vector loss, the model can better maintain the essential features of the data and improve the classification accuracy. Especially when dealing with high-dimensional data, the model shows excellent dimensionality reduction effect and feature extraction ability.

[0049] In terms of biological interpretability, the present invention establishes a complete molecular mechanism analysis framework by integrating functional annotation analysis and pathway identification. This not only improves the interpretability of the classification results, but also provides an important molecular mechanism basis for the formulation of clinical treatment plans. Through the deep attention neural network structure, the present invention can accurately identify the molecular pathway characteristics specific to different subtypes, providing reliable theoretical support for precision medicine. Therefore, the present invention solves the problem of insufficient recognition degree of disease biomarker subtype classification existing in the prior art. Brief Description of the Drawings

[0050] Figure 1 It is a flowchart of the method provided by the present invention;

[0051] Figure 2 It is the two-dimensional electrophoresis protein expression heat map in Example 2;

[0052] Figure 3 It is the correlation diagram between protein expression and clinical indexes in Example 2. Detailed Embodiments

[0053] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0054] As Figure 1 shown, it is a flowchart of a method for determining disease biomarker subtype classification provided by the present invention, and this method includes the following steps:

[0055] S10. Extract proteins from the training set of patient samples, separate the protein components using two-dimensional electrophoresis technology, and obtain the first protein expression matrix through standardization and normalization processing;

[0056] S20. Display the protein spots in the protein components by Coomassie brilliant blue staining, and perform quantitative analysis on the protein spots using an image analysis system to obtain the second protein expression matrix;

[0057] S30. Perform dimensionality reduction analysis on the second protein expression matrix, and preliminarily determine the number of differentially expressed proteins through the principal component distribution characteristics;

[0058] S40. Select the protein spots with significant expression differences for mass spectrometry identification to obtain the differential protein sequence information;

[0059] S50. Construct a protein interaction network, use the differential protein sequence information as network nodes and the protein-protein interactions as network connections, and adopt a graph neural network method to realize the extraction of network node features to obtain protein feature vectors;

[0060] S60. Use bioinformatics analysis methods to perform functional annotation analysis and pathway identification on the differential protein sequence information, and adopt a feature dimensionality reduction method to transform the protein feature vectors into low-dimensional features;

[0061] S70. Based on the low-dimensional features, use a multi-level clustering method to group the patient samples, construct a protein expression reconstruction loss function and a protein feature vector loss function, train an embedded neural network model, and optimize the embedded neural network model through a cross-validation method;

[0062] S80. Extract features from new patient samples using an embedded neural network model, and input the extracted low-dimensional features into a pre-trained functional enrichment analysis model to determine the molecular pathway features of disease subtypes;

[0063] S90. Verify the expression levels of key proteins using immunoblotting technology, conduct a correlation analysis in combination with patient clinical phenotype data, and verify the stability of molecular pathway features through an independent patient cohort to obtain the disease marker subtype classification results.

[0064] The following describes the specific implementation manners of the above steps in detail:

[0065] The specific implementation manner of step S10 is to extract proteins from the patient samples in the training set and separate the protein components using two-dimensional electrophoresis technology. First, the researchers collect at least 100 patient samples from the training set, including serum samples, tissue samples, peripheral blood mononuclear cell samples, and whole blood samples of diagnosed patients. Then, protein extraction methods are used to isolate proteins from these samples. To separate the protein components, two-dimensional electrophoresis technology is used. In the first dimension, isoelectric focusing is used to separate according to the isoelectric point of the proteins, and in the second dimension, sodium dodecyl sulfate polyacrylamide gel electrophoresis is used to separate according to the molecular weight. Next, the separated protein components are subjected to standardization and normalization processing to eliminate the dimensional differences in the expression levels of different proteins and make the data distribution more reasonable. The standardization uses the Z-score standardization method, that is, subtracting the average value of each protein spot from its expression value and then dividing by the standard deviation; the normalization uses the Min-Max normalization method, that is, linearly mapping the standardized expression value to the interval of [0,1]. After these processes, the first protein expression matrix is obtained. The purpose of this step is to isolate proteins from the original samples and preprocess them to lay a foundation for subsequent analysis.

[0066] The specific implementation manner of step S20 is to display the protein spots in the protein components by Coomassie brilliant blue staining and conduct quantitative analysis of the protein spots using an image analysis system. First, the gel is placed in Coomassie brilliant blue staining solution for 240 minutes of staining to make the protein spots appear. Then, a gel imaging device is used to obtain an image, and the optical density value of each protein spot is measured through the image analysis system, and background correction and standardization analysis are performed. This method can quantitatively reflect the relative expression level of each protein spot and obtain the second protein expression matrix. The purpose of this step is to visually display the protein components and use image analysis technology to quantitatively analyze the expression of each protein spot to provide basic data for subsequent statistical analysis.

[0067] The specific implementation of step S30 is to perform dimensionality reduction analysis on the second protein expression matrix and preliminarily determine the number of differentially expressed proteins based on the principal component distribution characteristics. First, the principal component analysis (PCA) method is used to map the high-dimensional second protein expression matrix to a low-dimensional principal component space. This process involves eigenvalue decomposition. By calculating the eigenvectors and eigenvalues of the covariance matrix, a principal component matrix is obtained, and then the original data is projected onto the principal component space. Next, the aggregation trend of data points in the low-dimensional principal component space is analyzed, and the number of clusters is determined in combination with the principal component contribution rate to preliminarily determine the number of differentially expressed proteins. The purpose of this step is to use an unsupervised dimensionality reduction method to identify samples with different protein expressions, providing guidance for subsequent selection of differential proteins and functional analysis.

[0068] The specific implementation of step S40 is to select protein spots with significant expression differences for mass spectrometry identification to obtain differential protein sequence information. First, according to the analysis results of the previous step, protein spots with significant expression differences are selected. Then, these protein spots are cut out from the gel, enzymatically digested, and the peptides of the proteins are analyzed and sequenced using a liquid chromatography-mass spectrometry instrument. Through this method, the sequence information of these differential proteins can be obtained. The purpose of this step is to determine the key proteins causing disease typing, providing basic data for subsequent functional analysis and network construction.

[0069] The specific implementation of step S50 is to construct a protein interaction network and use the graph neural network method to extract network node features to obtain protein feature vectors. First, based on the protein interaction information in the public database, the previously obtained differential protein sequence information is used as network nodes, and the interaction relationships between them are constructed as network connections. Next, the graph convolutional network (GCN) is used to propagate and aggregate the features of network nodes. Through the message passing mechanism, the feature vector representation of each node (protein) is learned. This method considers the topological relationship between nodes and can depict the complex interactions between proteins as a whole. The purpose of this step is to use the advantages of graph neural networks to extract biologically meaningful feature representations from the protein interaction network, providing a basis for subsequent functional analysis.

[0070] The specific implementation of step S60 is to perform functional annotation analysis and pathway identification on the differential protein sequence information using bioinformatics analysis methods, and to transform the protein feature vectors into low-dimensional features using feature dimensionality reduction methods. First, the differential protein sequence information is compared with a biological function annotation database to obtain gene ontology annotation information and biological pathway information. This bioinformatics analysis method can reveal the characteristics of these differential proteins in terms of biological functions and signal pathways. Next, the t-distributed Stochastic Neighbor Embedding (t-SNE) method is used to transform the protein feature vectors obtained above into low-dimensional feature representations. This method can better preserve the local structural relationships between data points in the high-dimensional space and enhance the separability of samples in the low-dimensional space. The purpose of this step is to extract the functional characteristics of differential proteins through bioinformatics analysis and use dimensionality reduction techniques to compress high-dimensional features into a low-dimensional space, laying a foundation for subsequent clustering analysis and model training.

[0071] The specific implementation of step S70 is to group patient samples using a multi-level clustering method based on low-dimensional features, and to construct a protein expression reconstruction loss function and a protein feature vector loss function to train an embedded neural network model. First, the hierarchical clustering algorithm is used to perform multi-level clustering on the low-dimensional features to obtain the grouping results of patient samples. This clustering method can reflect the hierarchical similarity between samples. Next, two loss functions are constructed: one is the protein expression reconstruction loss function, whose goal is to minimize the difference between the protein expression data reconstructed by the model and the original expression data; the other is the protein feature vector loss function, whose goal is to minimize the distance between sample feature vectors. These two loss functions are combined as the optimization objective for training the embedded neural network model. Through cross-validation, the parameters of this neural network model are continuously optimized so that it can effectively extract protein expression features and complete the sample grouping task. The purpose of this step is to establish an end-to-end model that can automatically learn protein expression features and perform disease typing.

[0072] The specific implementation of step S80 is to use the previously trained embedded neural network model to extract features from new patient samples, and input the extracted low-dimensional features into a pre-trained functional enrichment analysis model, so as to determine the molecular pathway features of disease subtypes. First, use the neural network structure of the encoding part to extract features from the protein expression data of new samples, and obtain a low-dimensional feature vector representation. Then, input these feature vectors into another pre-trained deep attention neural network model for functional enrichment analysis. This model contains multiple functional units, calculates the association strength between features and biological pathways through the attention mechanism, and evaluates the enrichment significance of each pathway based on an improved hypergeometric distribution test method. Finally, the model can output the enrichment results of each biological pathway, including statistical indicators such as significance scores and enrichment multiples, so as to identify the key molecular pathways related to disease classification. The purpose of this step is to use the trained model to extract features and perform functional enrichment analysis on new samples to determine the molecular mechanism features of disease subtypes.

[0073] The specific implementation of step S90 is to use immunoblotting technology to verify the expression levels of key proteins (including BDNF, IL-6, and PDGF), perform correlation analysis in combination with patient clinical phenotype data, and verify the stability of the molecular pathway features through an independent patient cohort, and finally obtain the disease biomarker subtype classification result. First, use immunoblotting experimental technology to verify the previously identified key differential proteins and measure their expression levels in different patient samples. Then, perform correlation analysis on these protein expression data and the clinical phenotypes of patients (such as symptoms, biochemical indicators, etc.) to explore the relationship between protein expression changes and clinical phenotypes. Finally, by collecting an independent patient sample cohort, repeat the above molecular pathway enrichment analysis to verify the stability and reliability of the previously identified key pathway features in the new samples. Combining the above analysis results, the final disease biomarker subtype classification result is obtained. The purpose of this step is to use biological experimental means to verify key differential proteins and molecular pathways to ensure that the established disease classification model has high reliability and stability.

[0074] The calculations involved in the present invention are described in detail below.

[0075] 1. Standardization and normalization in step S10:

[0076] The specific process of standardizing the first protein expression matrix is represented as follows:

[0077]

[0078] In the formula, Z ij is the standardized protein expression value; X ijis the original protein expression value, i represents the i-th sample, and j represents the j-th protein spot; μ j is the average expression value of the j-th protein spot; σ j is the standard deviation of the j-th protein spot.

[0079] The normalization process is specifically expressed as follows:

[0080]

[0081] In the formula, N ij is the normalized protein expression value; min(Z j ) is the minimum value of the standardized value of the j-th protein spot; max(Z j ) is the maximum value of the standardized value of the j-th protein spot.

[0082] 2. Quantitative analysis in step S20:

[0083] The optical density quantitative analysis of the protein spot is specifically expressed as follows:

[0084]

[0085] In the formula, D ij is the corrected optical density value of the j-th protein spot in the i-th sample; I ij is the original optical density value; B i is the background optical density value of the i-th sample; α is the correction coefficient, and its value range is 0.8 to 1.2.

[0086] 3. Principal component analysis in step S30:

[0087] The eigenvalue decomposition of the principal component analysis is expressed as follows:

[0088]

[0089] In the formula, C is the covariance matrix; X is the protein expression matrix after centering; n is the number of samples; V is the eigenvector matrix; Λ is the diagonal eigenvalue matrix.

[0090] The data after dimensionality reduction is expressed as follows:

[0091] Y = XV k ;

[0092] In the formula, Y is the data matrix after dimensionality reduction; V k is the matrix composed of the eigenvectors corresponding to the first k principal components, and the value of k is determined according to the cumulative contribution rate, usually requiring the cumulative contribution rate to be greater than 85%.

[0093] 4. Graph neural network feature extraction in step S50:

[0094] The node feature update formula is as follows:

[0095]

[0096] In the formula, is the feature representation of node i at the (l + 1)-th layer; N(i) is the set of neighbor nodes of node i; c ij is the relationship weight between nodes; W (l) is the weight matrix at the l-th layer; b (l) is the bias term; σ is the activation function.

[0097] 5. Feature dimensionality reduction in step S60:

[0098] The objective function of stochastic neighbor embedding is expressed as follows:

[0099]

[0100] In the formula, p ij is the similarity between point i and point j in the high-dimensional space; q ij is the similarity between point i and point j in the low-dimensional space.

[0101] The similarity calculation formula is as follows:

[0102]

[0103] In the formula, x i is the point in the high-dimensional space; y i is the point in the low-dimensional space; σ i is the bandwidth parameter of the Gaussian kernel.

[0104] 6. Construction of the loss function in step S70:

[0105] The protein expression reconstruction loss function is specifically expressed as follows:

[0106]

[0107] In the formula, X i is the original protein expression matrix; is the reconstructed protein expression matrix; n is the number of samples; Θ is the model parameter; λ1 is the regularization coefficient, and its value range is 0.001 to 0.1; ∥·∥ F represents the Frobenius norm; ∥·∥2 represents the L2 norm.

[0108] The protein feature vector loss function is specifically expressed as follows:

[0109]

[0110] In the formula, hi and h j are the feature vectors of sample i and sample j; S ij is the sample similarity matrix; α is the balance coefficient, and its value range is 0.1 to 1; m is the boundary parameter, and its value range is 1 to 5.

[0111] The overall loss function is expressed as follows:

[0112]

[0113] In the formula, β is the feature loss weight coefficient, and its value range is 0.1 to 1; γ is the network parameter regularization coefficient, and its value range is 0.0001 to 0.01; L is the number of network layers; W (l) is the weight matrix of the l-th layer.

[0114] 7. Functional enrichment analysis in step S80:

[0115] The formula for calculating the significance of pathway enrichment is as follows:

[0116]

[0117] In the formula, N is the total number of proteins; M is the number of proteins contained in a certain pathway; n is the number of differentially expressed proteins; k is the number of differentially expressed proteins belonging to this pathway.

[0118] The formula for calculating the enrichment multiple is as follows:

[0119]

[0120] In the formula, E is the enrichment multiple.

[0121] 8. Feature mapping function in the embedded neural network model:

[0122] h (1) = σ(W (1) x + b (1) );

[0123] h (2) = σ(W (2) h (1) + b (2) );

[0124] h (3) = σ(W (3) h (2) + b (3) );

[0125] In the formula, h (l) is the hidden layer feature of the l-th layer; W (l) is the weight matrix of the l-th layer; b (l)is the bias vector of the l-th layer; σ is the activation function, and the ReLU function is usually selected.

[0126] The feature reconstruction function is expressed as follows:

[0127]

[0128] In the formula, is the reconstructed protein expression data.

[0129] 9. The error correction function is expressed as follows:

[0130]

[0131] In the formula, is the corrected reconstructed data; δ is the correction coefficient, and its value range is 0 to 1.

[0132] 10. The feedback regulation function is expressed as follows:

[0133]

[0134] In the formula, ΔW (l) is the update amount of the l-th layer weight; η is the learning rate, and its value range is 0.0001 to 0.01; μ is the momentum coefficient, and its value range is 0.8 to 0.99.

[0135] 11. Attention mechanism calculation:

[0136] The attention weight calculation formula is as follows:

[0137]

[0138] In the formula, A ij is the attention weight matrix; e ij is the attention score; Q i is the query vector; K j is the key vector; W q and W k are trainable weight matrices; d k is the dimension of the key vector.

[0139] Multi-head attention output calculation:

[0140] MultiHead(Q, K, V) = Concat(head1,..., head h )W O ;

[0141] head i = Attention(QW i Q , KW iK , VW i V );

[0142] Where W i Q , W i K , W i V is the parameter matrix of the i-th head; W O is the output projection matrix; h is the number of attention heads.

[0143] 12. Network parameter initialization:

[0144] Xavier initialization method:

[0145]

[0146] Where W ij is an element of the weight matrix; n in is the input dimension; n out is the output dimension; U represents a uniform distribution.

[0147] The principles and meanings of these formulas are explained as follows:

[0148] 1. Standardization and normalization are performed using Z-score standardization and Min-Max normalization, aiming to eliminate the dimensional differences in the expression levels of different protein spots and make the data distribution more reasonable;

[0149] 2. Logarithmic transformation is used for optical density quantitative analysis to convert multiplicative noise into additive noise and introduce a correction coefficient to adjust measurement errors;

[0150] 3. Principal component analysis reduces the dimension through eigenvalue decomposition, retains the main variation information, and reduces data redundancy;

[0151] 4. Graph neural network feature extraction considers the topological relationship between nodes and learns node representations through the message passing mechanism;

[0152] 5. Stochastic neighbor embedding preserves the local structural relationship between data points, and the t-distribution is used to enhance the separability in the low-dimensional space;

[0153] 6. The loss function design includes three parts: reconstruction error, feature preservation, and regularization, which balances the reconstruction ability and feature extraction ability of the model;

[0154] 7. Pathway enrichment analysis is based on hypergeometric distribution testing to evaluate the enrichment degree of differential proteins in biological pathways;

[0155] 8. The deep neural network model adopts a multi-layer structure and extracts hierarchical features through non-linear transformation;

[0156] 9. The attention mechanism introduces a query-key-value mechanism, enhancing the model's attention to important features;

[0157] 10. The parameter initialization adopts the Xavier method, which helps the stable propagation of gradients in the initial stage of training. The Xavier method is a method for initializing network parameters that can maintain the smoothness of gradients and is beneficial to the training of the network.

[0158] Specifically, the principle of the present invention is as follows: From the perspective of data acquisition and preprocessing, the present invention uses two-dimensional electrophoresis technology for protein separation. This method can separate proteins simultaneously according to their isoelectric points and molecular weights, featuring high resolution and high repeatability. Through Coomassie brilliant blue staining and quantitative analysis by an image analysis system, protein expression information can be accurately obtained. The principle basis of this method is the migration characteristics of proteins in an electric field. By optimizing the electrophoresis conditions and staining protocols, the accuracy and reliability of data acquisition are ensured.

[0159] In terms of feature extraction, the present invention innovatively introduces the graph neural network method. Its core principle is to convert the topological structure information in the protein interaction network into a learnable feature representation. Through graph convolution operations, each protein node can aggregate the information of its neighbor nodes, thereby capturing the local structure features in the network. This feature extraction method based on the message passing mechanism can effectively integrate the interaction relationships between proteins and provide a richer feature representation.

[0160] The embedded neural network model designed by the present invention adopts an encoder-decoder structure. Its working principle is to map high-dimensional data to a low-dimensional latent space through a multi-layer neural network and then reconstruct it back to the original space. This structural design can effectively retain the essential features of the data while achieving the purpose of dimensionality reduction. In particular, by introducing a dual loss function, both the accuracy of data reconstruction and the distribution of the feature space are considered, thereby improving the performance of the model.

[0161] In terms of functional enrichment analysis, the present invention adopts a deep attention neural network structure. Its principle is to adaptively learn the association strength between features and biological pathways through the attention mechanism. The multi-head self-attention module can capture the complex dependence relationships between features. This design enables the model to more accurately identify key molecular pathway features. At the same time, an improved hypergeometric distribution test method provides reliable statistical support.

[0162] The technical solution of the present invention has a complete logical chain: First, the reliability of the basic data is ensured through high-quality data collection. Then, the graph neural network is used to extract the network structure features. Next, the embedded neural network is used to achieve feature optimization and dimensionality reduction. Finally, the molecular pathway features are determined based on functional enrichment analysis. This multi-level and multi-angle analysis strategy can comprehensively and accurately achieve the subtype classification of disease markers.

[0163] A specific embodiment 1 of the present invention is provided below. The specific implementation of each step in this embodiment 1 is described in detail as follows: The specific implementation of step S10 involves standardization processing and normalization processing, which can be described by the following formula:

[0164] Standardization processing:

[0165] Among them, Z ij is the standardized protein expression value, X ij is the original protein expression value. The subscript i represents the i-th sample, and the subscript j represents the j-th protein spot; μ j is the average expression value of the j-th protein spot, and σ is the standard deviation of the j-th protein spot. Standardization processing can eliminate the dimensional difference of different protein expression levels and make the data distribution more reasonable.

[0166] Normalization processing:

[0167] Among them, N ij is the normalized protein expression value; min(Z j ) is the minimum value of the standardized value of the j-th protein spot; max(Z j ) is the maximum value of the standardized value of the j-th protein spot. Normalization processing can map the data into the interval [0,1] and further eliminate the dimensional difference.

[0168] The quantitative analysis of the protein spot optical density involved in step S20 can be described by the following formula:

[0169]

[0170] Among them, D ij is the corrected optical density value of the j-th protein spot in the i-th sample; I ij is the original optical density value; B i is the background optical density value of the i-th sample; α is the correction coefficient, and the value range is 0.8 to 1.2. Logarithmic conversion can convert multiplicative noise into additive noise, and introducing the correction coefficient can adjust the measurement error.

[0171] The dimensionality reduction process of principal component analysis involved in step S30 can be described by the following formula:

[0172] Eigenvalue decomposition of the covariance matrix:

[0173] where C is the covariance matrix; X is the protein expression matrix after centering; n is the number of samples; V is the matrix of eigenvectors; and Λ is the diagonal eigenvalue matrix.

[0174] Data representation after dimensionality reduction: Y = XV k ;

[0175] where Y is the data matrix after dimensionality reduction; V k is the matrix composed of the eigenvectors corresponding to the first k principal components, and the value of k is determined according to the cumulative contribution rate, usually requiring the cumulative contribution rate to be greater than 85%. Principal component analysis maps high-dimensional data to a low-dimensional principal component space through eigenvalue decomposition, retaining the main variation information.

[0176] The feature extraction process involving the graph neural network in step S50 can be described by the following formula:

[0177] Node feature update formula:

[0178] where is the feature representation of node i at the (l + 1)-th layer; N(i) is the set of neighbor nodes of node i; c ij is the relationship weight between nodes; W (l) is the weight matrix at the l-th layer; b (l) is the bias term; and σ is the activation function. The graph neural network learns a feature representation rich in biological significance through the message passing mechanism, utilizing the topological relationship between nodes.

[0179] The dimensionality reduction process involving stochastic neighbor embedding in step S60 can be described by the following formula:

[0180] Objective function:

[0181] where p ij is the similarity between point i and point j in the high-dimensional space; q ij is the similarity between point i and point j in the low-dimensional space.

[0182] Similarity calculation:

[0183]

[0184] where x i is the point in the high-dimensional space; y i is the point in the low-dimensional space; σ iis the bandwidth parameter of the Gaussian kernel. Stochastic Neighbor Embedding enhances the separability of samples in the low-dimensional space by preserving the local structural relationships of data points in the high-dimensional space.

[0185] The construction of the loss function involved in step S70 can be described by the following formula:

[0186] Protein expression reconstruction loss function:

[0187] where, X i is the original protein expression matrix; is the reconstructed protein expression matrix; n is the number of samples; Θ is the model parameter; λ1 is the regularization coefficient, and its value range is 0.001 to 0.1; ∥·∥ F represents the Frobenius norm; ∥·∥2 represents the L2 norm.

[0188] Protein feature vector loss function:

[0189] where, h i and h j are the feature vectors of sample i and sample j; S ij is the sample similarity matrix; α is the balance coefficient, and its value range is 0.1 to 1; m is the boundary parameter, and its value range is 1 to 5.

[0190] Overall loss function:

[0191] where, β is the feature loss weight coefficient, and its value range is 0.1 to 1; γ is the network parameter regularization coefficient, and its value range is 0.0001 to 0.01; L is the number of network layers; W (l) is the weight matrix of the l-th layer.

[0192] This loss function design takes into account both the reconstruction ability and the feature extraction ability of the model, and prevents overfitting through the regularization term.

[0193] The functional enrichment analysis involved in step S80 can be described by the following formula:

[0194] Pathway enrichment significance calculation:

[0195] where, N is the total number of proteins; M is the number of proteins contained in a certain pathway; n is the number of differentially expressed proteins; k is the number of differentially expressed proteins belonging to this pathway. This formula is based on the hypergeometric distribution test to evaluate the enrichment degree of differentially expressed proteins in biological pathways.

[0196] Enrichment fold calculation:

[0197] Among them, E is the enrichment multiple. The enrichment multiple reflects the relative enrichment of differential proteins in a certain pathway.

[0198] The feature mapping function of the embedded neural network model is also involved in step S80:

[0199] h (1) = σ(W (1) x + b (1) );

[0200] h (2) = σ(W (2) h (1) + b (2) );

[0201] h (3) = σ(W (3) h (2) + b (3) );

[0202] Among them, h (l) is the hidden layer feature of the l-th layer; W (l) is the weight matrix of the l-th layer; b (l) is the bias vector of the l-th layer; σ is the activation function, and usually the ReLU function is selected. This feature mapping function extracts feature representations rich in biological significance from the original protein expression data through multi-layer non-linear transformations.

[0203] Feature reconstruction function:

[0204] Among them, is the reconstructed protein expression data. This function reconstructs the features using the neural network structure of the decoding part, enabling the model to effectively model the complex distribution of protein expression.

[0205] The error correction function and the feedback regulation function are also involved in step S80:

[0206] Error correction function:

[0207] Among them, is the reconstructed data after correction; δ is the correction coefficient, and its value range is 0 to 1. This function eliminates the bias introduced in the reconstruction process by introducing the correction coefficient.

[0208] Feedback regulation function:

[0209] Among them, ΔW (l)is the update amount of the weights in the l-th layer; η is the learning rate, with a value range of 0.0001 to 0.01; μ is the momentum coefficient, with a value range of 0.8 to 0.99. This function updates the model parameters using the gradient descent algorithm, and the momentum term can accelerate the convergence speed and improve stability.

[0210] The calculation formula of the attention mechanism is also involved in step S80:

[0211] Attention weight calculation:

[0212]

[0213] where, A ij is the attention weight matrix; e ij is the attention score; Q i is the query vector; K j is the key vector; W q and W k are trainable weight matrices; d k is the dimension of the key vector. The attention mechanism enhances the model's attention to important features by calculating the association strength between features and path nodes.

[0214] Multi-head attention output calculation: MultiHead(Q, K, V) = Concat(head1,..., head h )W O ;

[0215] head i = Attention(QW i Q , KW i K , VW i V );

[0216] where, W i Q , Wx K , W i V are the parameter matrices of the i-th head; W O is the output projection matrix; h is the number of attention heads. The multi-head attention structure can capture complex dependencies between features.

[0217] The initialization of network parameters is also involved in step S70:

[0218] Xavier initialization method:

[0219] where, W ij is the element of the weight matrix; nin is the input dimension; n out is the output dimension; U represents a uniform distribution. This initialization method helps the stable propagation of gradients in the initial stage of training.

[0220] To better understand and implement the present invention, Example 2 of a specific application scenario of the present invention is provided below: A research team plans to use the disease biomarker subtype classification determination method proposed by the present invention to conduct disease typing research on stroke patients. First, they collected serum samples, peripheral blood mononuclear cell samples, and whole blood samples from 100 confirmed stroke patients in a certain tertiary hospital as the training set. At the same time, corresponding samples from 50 healthy controls were also collected as the comparison group.

[0221] Sample pretreatment and protein extraction

[0222] The researchers first preprocessed the collected patient and control samples. The serum samples were directly frozen and stored; the peripheral blood mononuclear cell samples were separated by Ficoll-Paque density gradient centrifugation, and the centrifuged cell layer was frozen and stored; the whole blood samples were mixed well with lysis buffer and then centrifuged, and the erythrocyte membrane layer was obtained and frozen and stored. Then, using a conventional protein extraction method, total proteins were isolated from the above three biological samples. Protein concentration was determined by the BCA method, and the sample concentration was adjusted to 2 mg / mL.

[0223] Two-dimensional electrophoresis and protein spot quantification

[0224] The extracted total protein samples were separated by two-dimensional electrophoresis. In the one-dimensional isoelectric focusing (IEF) direction, an 18 cm IPG strip (pH 3-10) was used, and the voltage was set to 200 V for 1 h, 500 V for 1 h, 1000 V for 1 h, and 8000 V for 12 h to complete the isoelectric point separation of proteins. In the two-dimensional direction, a 12.5% SDS-PAGE gel was used, and the voltage was set to 80 V for 30 min and 120 V for 4-5 h to complete the molecular weight separation. After electrophoresis, the gel was soaked in Coomassie Brilliant Blue staining solution and decolorized with decolorizing solution after 240 min until protein spots were shown.

[0225] Then, a gel imaging system was used to obtain the electrophoresis gel image, and ImageMaster 2D software was used to quantitatively analyze the optical density value of each protein spot. The specific steps are as follows: First, the image was background-corrected to calculate the original optical density value (I) of each protein spot; then, according to the formula the optical density value was corrected to obtain the corrected optical density value (D), where B represents the background optical density value of the sample, and α is the correction coefficient with a value of 0.9. Finally, the protein expression data of all samples were organized into a two-dimensional matrix as the input for subsequent analysis. As Figure 2As shown, this figure shows the expression pattern of two-dimensional electrophoresis of proteins. The horizontal axis represents the isoelectric point of proteins, ranging from 3 to 10; the vertical axis represents the molecular weight of proteins, ranging from 10 to 200. The color of the heat map from blue to red represents the change in protein expression level, where blue indicates low expression and red indicates high expression.

[0226] Analysis of protein expression differences

[0227] To preliminarily identify samples with different protein expressions, the researchers performed principal component analysis (PCA) on the standardized protein expression matrix. First, calculate the covariance matrix where X is the centered expression matrix and n is the number of samples. Then, perform eigenvalue decomposition on C: C = VΛV T , to obtain the eigenvector matrix V and the eigenvalue diagonal matrix Λ. Select the first k principal components, where the value of k is determined according to the cumulative contribution rate, usually requiring more than 85%. Finally, project the original expression matrix X into the principal component space to obtain the reduced-dimensional data representation Y = XV k .

[0228] By observing the PCA results, the researchers found that there were obvious clustering distributions of stroke patient samples and healthy controls in the principal component space. Further analyzing the contribution rate of the principal components, it was found that the first 3 principal components could explain approximately 90% of the total variation. Therefore, it was preliminarily determined that there were approximately 30 - 50 significantly differentially expressed proteins in the stroke patient samples.

[0229] Differential protein identification and network construction

[0230] Based on the above PCA analysis results, the researchers selected 30 protein spots with the most significant differential expressions, cut them out from the gel, and performed enzymatic digestion and liquid chromatography - tandem mass spectrometry (LC - MS / MS) analysis. Through database comparison, the amino acid sequence information of these 30 proteins was successfully identified.

[0231] Next, the researchers constructed a protein - protein interaction network using this differential protein sequence information. First, obtain the known protein - protein interaction relationships from the STRING database, use the identified differential proteins as network nodes, and their interaction relationships as network edges. Then, use a graph neural network (GNN) to extract features from this network. Specifically, use a graph convolutional network (GCN) to aggregate and propagate the features of each node to learn a 32 - dimensional protein feature vector representation h i .

[0232] Through this method, the researchers not only obtained the characteristic representations of each differential protein, but also fully considered the interaction relationships between them. These characteristics will provide an important basis for subsequent disease typing.

[0233] Biological function analysis

[0234] To further explore the biological significance behind these differential proteins, the researchers performed functional annotation and pathway analysis on them. First, the sequence information of the identified differential proteins was aligned with the Gene Ontology (GO) database to obtain their functional category information, which involves biological processes, cellular components, and molecular functions, etc.

[0235] Next, the signaling pathways involved by these differential proteins were searched through the KEGG database. The results showed that these differential proteins were mainly concentrated in biological pathways related to nervous system development, inflammatory response, coagulation function, etc. For example, some of these proteins were involved in the neurotrophic factor signaling pathway, interleukin signaling pathway, and platelet activation pathway, etc. The abnormal activities of these pathways may be the key molecular mechanisms leading to the occurrence and development of stroke.

[0236] To further verify the importance of these biological functions, the researchers adopted a feature dimensionality reduction method based on t-Distributed Stochastic Neighbor Embedding (t-SNE) to compress the 32-dimensional protein feature vectors into a 3-dimensional space. In the low-dimensional feature space, the stroke patient samples and healthy control samples could be well separated, indicating that the extracted protein features indeed contained key information related to the disease.

[0237] As Figure 3 shown, this figure shows the correlation between the expression levels of three key proteins (BDNF, IL-6, and PDGF) and the clinical NIHSS score. The scatter plot contains actual data points and the fitted regression line, and the correlation coefficient and significance level are marked in the figure.

[0238] Machine learning model training

[0239] Based on the above-extracted protein expression features, the researchers designed an embedded neural network model for automatically learning disease typing features and completing classification tasks. This model consists of two parts: an encoding network and a decoding network:

[0240] Encoding network:

[0241] h (1) =σ(W (1) x + b (1) );

[0242] h (2) =σ(W (2) h(1) +b (2) );

[0243] h (3) = σ(W (3) h (2) +b (3) );

[0244] where, n (l) is the hidden layer feature of the l-th layer, W (l) and b (l) are the weight matrix and bias vector of this layer respectively, and σ is the ReLU activation function.

[0245] Decoder network:

[0246]

[0247] This part reconstructs the protein expression data using a multi-layer perceptron structure.

[0248] During the training process, a joint optimization strategy is adopted to simultaneously minimize the reconstruction loss L rec and the feature loss L fea :

[0249]

[0250] where, X i and are the original and reconstructed protein expression data respectively; h i and h j are the feature vectors of samples i and j; S ij is the sample similarity; λ1, α, β, γ are hyperparameters. In this way, the model can not only learn effective feature representations but also reconstruct the original protein expression profile.

[0251] To further improve the performance of the model, the researchers also introduced an attention mechanism into the functional enrichment analysis module. Specifically, this module consists of multiple functional units, including attention calculation, enrichment scoring, and output generation. The attention calculation formula is as follows:

[0252]

[0253] where, A ij is the attention weight, e ij is the attention score, Q i is the query vector, K j is the key vector, W q and W k are trainable parameters. This attention mechanism can adaptively assign weights to different biological pathways, enhancing the model's attention to key features.

[0254] The research team trained and optimized the embedded neural network model using cross-validation. On the training set, the classification accuracy of the model reached over 85%; on the independent test set, the classification accuracy remained at about 80%, indicating that the model has good generalization ability.

[0255] Biological verification and clinical correlation analysis

[0256] To further verify the correlation between the extracted protein features and the pathogenesis of stroke, the researchers selected 3 proteins with the most significant differential expression, namely BDNF, IL-6, and PDGF, and verified their expression levels in patient and control samples using immunoblotting technology. The results showed that BDNF was significantly downregulated in stroke patients, while IL-6 and PDGF were significantly upregulated, which was consistent with the previous proteomic analysis results.

[0257] At the same time, the researchers also collected the clinical index data of these patients, including age of onset, stroke severity score (NIHSS score), prognostic function score (mRS score), etc. Through correlation analysis, it was found that the expression level of BDNF was significantly positively correlated with the prognostic function (r = 0.68, p < 0.01), while the expression levels of IL-6 and PDGF were significantly positively correlated with the stroke severity (r = 0.72, p < 0.01; r = 0.61, p < 0.05). This further confirmed that these differential proteins may be key molecular markers reflecting the pathological mechanism of stroke.

[0258] To verify the stability on independent samples, the research team collected another 50 stroke patients and 30 healthy controls' biological samples and repeated the above proteomic analysis and machine learning classification. The results showed that on the new validation set, the classification accuracy of the embedded neural network model still maintained at about 78%, which was very close to the results of the training set, indicating that this typing method has good generalization ability. At the same time, the immunoblot experiment also re-verified the abnormal expression characteristics of BDNF, IL-6, and PDGF in stroke patients.

[0259] Based on the above results, the disease biomarker subtype classification method constructed by this research team can automatically extract the key molecular features related to the onset of stroke from proteomic data and achieve disease typing based on these features. The identified differential proteins are not only related to the pathogenesis of stroke in terms of biological function but also show good correlation in clinical indicators. This method provides new technical support for the precise diagnosis and individualized treatment of stroke. The research team summarized the application of this method and made the following several tables:

[0260] Table 1 Sample Information of the Training Set and Validation Set

[0261] Sample category Training set Validation set Stroke patients 100 cases 50 cases Healthy control 50 cases 30 cases Total 150 cases 80 cases

[0262] Table 2 Functional Annotations and Pathway Information of Differentially Expressed Proteins

[0263] Differential protein Functional category Involved pathway BDNF Neurotrophic factor activity Nerve growth factor signaling pathway IL-6 Inflammatory response; cytokine activity Interleukin signaling pathway PDGF Platelet-derived growth factor activity Platelet activation and coagulation function

[0264] Table 3 Correlation Analysis of Key Protein Expression and Clinical Indicators

[0265] Protein Related indicators Correlation coefficient p value BDNF Prognostic function (mRS score) 0.68 <0.01 IL-6 Stroke severity (NIHSS score) 0.72 <0.01 PDGF Stroke severity (NIHSS score) 0.61 <0.05

[0266] Table 4 Performance Evaluation of the Disease Subtyping Model

[0267] Indicator Training set Validation set Classification accuracy 85.3% 78.1% Sensitivity 83.7% 75.4% Specificity 87.2% 81.3%

[0268] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Table 5 below.

[0269] Table 5 Variable Explanation Table

[0270]

[0271]

[0272] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.

Claims

1. A method for determining the classification of disease biomarker subtypes, characterized in that, Including: Extract proteins from the patient samples in the training set and separate protein components, and record the first protein expression matrix; Stain to show protein spots in the protein components, perform quantitative analysis on the protein spots, and record the second protein expression matrix; perform dimensionality reduction analysis on the second protein expression matrix to determine the number of differentially expressed proteins; Select protein spots with significant expression differences for mass spectrometry identification to obtain differential protein sequence information; Construct a protein interaction network, use the graph neural network method to realize the extraction of network node features, and obtain protein feature vectors; perform functional annotation analysis and pathway identification on the differential protein sequence information, and transform the protein feature vectors into low-dimensional features; Group the patient samples using the multi-level clustering method based on the low-dimensional features, and train an embedded neural network model; use the embedded neural network model to extract features from new patient samples to determine the molecular pathway features of the disease subtypes; Verify the expression levels of key proteins to obtain the classification results of disease biomarker subtypes.

2. The method for determining a disease biomarker subtype classification according to claim 1, wherein Perform quantitative analysis on the protein spots. Specifically, use a gel imaging device to obtain images, measure the optical density values of the protein spots through an image analysis system, perform background correction and normalization analysis, and form the second protein expression matrix.

3. The method for determining a disease biomarker subtype classification according to claim 2, wherein Perform dimensionality reduction analysis on the second protein expression matrix. Specifically, use the principal component analysis method to map the second protein expression matrix to a low-dimensional principal component space.

4. The method for determining a disease biomarker subtype classification according to claim 3, wherein Determine the number of differentially expressed proteins through the principal component distribution characteristics. Specifically, analyze the aggregation trend of the data distribution in the low-dimensional principal component space, and determine the number of clusters in combination with the principal component contribution rate.

5. The method for determining a disease marker subtype classification according to claim 4, wherein Select protein spots with significant expression differences for mass spectrometry identification. Specifically, cut the protein spots from the gel, perform enzymatic digestion, and use a liquid chromatography-mass spectrometry instrument for peptide segment analysis and protein sequence identification.

6. The method for determining a disease marker subtype classification according to claim 5, wherein Construct a protein interaction network. Specifically, based on the protein interaction database, construct the differential protein sequence information as network nodes and the protein interaction relationships as network connections.

7. A method for determining a disease biomarker subtype classification according to claim 6, characterized in that, Use the graph neural network method to realize the extraction of network node features. Specifically, use a graph convolutional network to propagate and aggregate the network node features and learn to obtain protein feature vectors.

8. A method for determining a disease marker subtype classification according to claim 7, characterized in that Perform functional annotation analysis and pathway identification on the differential protein sequence information. Specifically, compare the differential protein sequence information with the biological function annotation database to obtain gene ontology annotation information and biological pathway information.

9. The method for determining a disease biomarker subtype classification according to claim 8, wherein Transform the protein feature vectors into low-dimensional features. Specifically, use the stochastic neighbor embedding method to transform the protein feature vectors into three-dimensional feature representations.

10. The method for determining a disease marker subtype classification according to claim 9, wherein Group the patient samples using the multi-level clustering method based on the low-dimensional features. Specifically, construct a similarity matrix of patient samples and perform iterative clustering through the hierarchical clustering method.

Citation Information

Cited By

  • Brain disease subtype identification method and system based on comparative generation learning

    CN120580516A