Proteome marker identification method and system, electronic equipment and storage medium

By using a variational autoencoder for data augmentation and marker recognition when processing exosome proteome data, the problems of poor accuracy and low reliability in multi-group sample comparisons are solved, and higher data robustness and marker recognition accuracy are achieved.

CN120126577AActive Publication Date: 2025-06-10SHENZHEN HUIXIN LIFE TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202411025586.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2025-06-10
Estimated Expiration
2044-07-29

AI Technical Summary

Technical Problem

The prior art has problems of poor accuracy and low reliability when processing exosome proteome data, especially comparisons of multiple groups of samples.

Method used

Using a data augmentation model based on a variational autoencoder, supervised training is performed on mass spectrometry quantitative data, new sample data is generated, and group sets are divided through principal component analysis and similarity analysis to identify potential markers.

Benefits of technology

It significantly improves the robustness of data and the accuracy of marker recognition, overcomes the limitations of the existing technology, and provides a more effective solution for the field of bioinformatics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126577A_ABST
    Figure CN120126577A_ABST
Patent Text Reader

Abstract

The invention discloses a proteome marker identification method and system, electronic equipment and a storage medium, and the method comprises the steps: adding a one-hot coded grouping tag to protein abundance data in mass spectrum quantitative data after data preprocessing, and coding group information corresponding to the protein abundance data into a binary vector; performing supervised training on the data enhancement model, so that a logarithmic variance output by an encoder of the data enhancement model is cut to be within a range of (-10, 10), and generating new sample data by using a decoder of the trained data enhancement model; according to the result of principal component analysis, dividing the new sample data with the similarity reaching a set threshold value into a group set, identifying potential markers through the group set by utilizing a differential expression analysis or random forest method, sorting the potential markers according to the difference multiple or importance parameters output by a specified data filling model, and obtaining the potential markers. And multi-group exosome proteome marker identification results are obtained. According to the invention, the accuracy and reliability of marker identification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for identifying proteomic markers, and belongs to the technical field of bioinformatics processing. Background Art

[0002] Currently, in the field of proteomic analysis of biological information, differential protein analysis (DPA) is a common method for identifying protein markers with significantly different abundances in multiple groups of samples (such as multiple disease sample groups).

[0003] In related technologies, the main steps of differential protein analysis generally include: A1) Data collection and preprocessing: Obtain protein abundance data through proteomic sequencing, and perform quality control, missing value filling, and normalization processing. A2) Differential protein analysis: Use statistical methods (such as t-test, Wilcoxon test, etc.) and machine learning methods (such as random forest, support vector machine, etc.) to analyze the preprocessed data to identify proteins with significant differences. A3) Marker identification: Screen potential protein markers based on statistical significance and fold change.

[0004] The differential protein analysis method has been widely used in proteomic research, but there are obvious limitations in processing exosome samples and multi-group sample comparisons, mainly manifested in the following aspects:

[0005] B1) Consistency problem of exosome samples:

[0006] High complexity: The preparation and detection of exosome samples are more complex than those of cell samples, resulting in lower consistency between repeated samples in proteomic sequencing results.

[0007] High noise: Due to the high noise background of the data, it is difficult to accurately capture reasonable and robust differences, resulting in insufficient accuracy in differential protein identification.

[0008] B2) Small sample size (i.e., insufficient sample diversity):

[0009] Cost and scarcity: The preparation cost of exosome samples is high, and the samples are rare. Usually, the number of omics sequencing samples for each disease is 3-5, and the sample size is small.

[0010] Limitations in statistical analysis: The small sample size limits the effectiveness of statistical analysis and the generalization ability of the model, resulting in reduced reliability in marker identification.

[0011] B3) Complexity of multi-group comparison:

[0012] The problems of noise and sample size are exacerbated: In the case of multi-group comparison, the above problems B1) and B2) (data noise and small sample size) have a greater impact on the analysis results because at this time, comparisons need to be made among multiple groups, increasing the data complexity.

[0013] Limitations of traditional methods: When directly using traditional statistical test methods (such as Wilcox test) or machine learning methods (such as random forest) to determine potential exosomal protein markers of diseases, it is easy to cause insufficient marker identification or misjudgment due to noise and sample size problems.

[0014] In summary, the existing technologies have significant limitations in processing exosomal proteome data, especially in the comparison of multi-group samples. There is an urgent need for a new method to improve the accuracy and reliability of marker identification. Summary of the Invention

[0015] To this end, the present invention provides a proteome marker identification method and system to solve the problems of poor accuracy and low reliability in the existing technology when processing exosomal proteome data, especially in the comparison of multi-group samples.

[0016] To achieve the above object, the present invention provides the following technical solutions: In the first aspect, a proteome marker identification method is provided, including:

[0017] Adding one-hot encoded grouping labels to the protein abundance data in the mass spectrometry quantitative data, and encoding the group information corresponding to the protein abundance data into a binary vector;

[0018] Constructing a data augmentation model based on a variational autoencoder, using the protein abundance data and the one-hot encoded grouping labels to perform supervised training on the data augmentation model, and using the decoder of the trained data augmentation model to generate new sample data;

[0019] Performing principal component analysis and similarity analysis on the new sample data; According to the results of the principal component analysis, dividing the new sample data with similarity reaching a set threshold into a group set;

[0020] Identifying potential markers using a specified method, and sorting the potential markers according to the importance parameters of the potential markers to obtain the identification results of multi-group exosomal proteome markers.

[0021] As a preferred solution of the proteome marker identification method, it further includes performing data preprocessing on the mass spectrometry quantitative data, and the data preprocessing includes data quality control, data normalization, and missing value filling.

[0022] As a preferred solution for the method of identifying proteomic markers, during the data quality control process, protein abundance data with missing values in more than a specified number of samples is removed;

[0023] During the data normalization process, the protein abundance data is normalized using a specified normalization method;

[0024] During the missing value filling process, a specified data filling model is used to fill in the missing values in the protein abundance data after data normalization.

[0025] As a preferred solution for the method of identifying proteomic markers, the loss function of the data augmentation model includes a reconstruction error function and a KL divergence function; an optimizer is used to minimize the loss function of the data augmentation model to optimize the parameters of the data augmentation model. For example, the Adam (Adaptive Moment Estimation) optimizer is an efficient optimization algorithm used in deep learning, which has been proven effective in practice and has become one of the default optimizers for many deep learning models.

[0026] As a preferred solution for the method of identifying proteomic markers, the expression of the reconstruction error function in the loss function of the data augmentation model is:

[0027]

[0028] In the formula, recon_batch represents the reconstructed protein abundance data; data represents the original input protein abundance data; recon_batch i represents the i-th item of the reconstructed protein abundance data; data i represents the i-th item of the original input protein abundance data.

[0029] As a preferred solution for the method of identifying proteomic markers, the expression of the KL divergence function in the loss function of the data augmentation model is:

[0030]

[0031] In the formula, μ represents the mean of the latent space; D represents the dimension of the latent space; μ i represents the mean of the i-th dimension of the latent space; σ 2 i represents the variance of the i-th dimension of the latent space.

[0032] As a preferred solution for the method of identifying proteomic markers, during the supervised training of the data augmentation model, the logarithmic variance output by the encoder of the data augmentation model is clipped to the range (-10, 10).

[0033] As a preferred solution of the proteomic biomarker identification method, during the principal component analysis of the new sample data:

[0034] Calculate the covariance matrix of the protein abundance data, solve the eigenvalues and eigenvectors of the covariance matrix, select the eigenvectors corresponding to the covariance matrix reaching the set eigenvalues as the principal components, and project the protein abundance numbers onto the determined principal components;

[0035] During the similarity analysis of the new sample data after principal component analysis: Evaluate the similarity between the samples of the specified group according to the distribution of the data points after dimensionality reduction by principal component analysis.

[0036] As a preferred solution of the proteomic biomarker identification method, use the differential expression analysis or random forest method to identify potential biomarkers through the group set, rank the potential biomarkers according to the fold change or the importance parameters filled in by the specified data in the model output, and obtain the identification results of the multi-group exosomal proteomic biomarkers.

[0037] In a second aspect, the present invention also provides a proteomic biomarker identification system, including:

[0038] A data preparation module, used to add one-hot encoded grouping labels to the protein abundance data in the mass spectrometry quantitative data, and encode the group information corresponding to the protein abundance data into a binary vector;

[0039] A data augmentation model processing module, used to construct a data augmentation model based on a variational autoencoder, use the protein abundance data and the one-hot encoded grouping labels to perform supervised training on the data augmentation model, and use the decoder of the trained data augmentation model to generate new sample data;

[0040] A group set generation module, used to perform principal component analysis and similarity analysis on the new sample data; According to the results of the principal component analysis, divide the new sample data with similarity reaching the set threshold into a group set;

[0041] A biomarker identification module, used to identify potential biomarkers using a specified method, rank the potential biomarkers according to the importance parameters of the potential biomarkers, and obtain the identification results of the multi-group exosomal proteomic biomarkers.

[0042] As a preferred solution of the proteomic biomarker identification system, it further includes a data preprocessing module, used to perform data preprocessing on the mass spectrometry quantitative data, and the data preprocessing includes data quality control, data normalization, and missing value filling;

[0043] During the data quality control process, protein abundance data with missing abundance values in more than a specified number of samples is removed;

[0044] During the data normalization process, the protein abundance data is normalized using a specified normalization method;

[0045] During the missing value filling process, the protein abundance data after data normalization is filled with missing values using a specified data filling model.

[0046] As a preferred solution of the proteomic biomarker identification system, in the data augmentation model processing module:

[0047] The loss function of the data augmentation model includes a reconstruction error function and a KL divergence function; the Adam optimizer is used to minimize the loss function of the data augmentation model to optimize the parameters of the data augmentation model;

[0048] The expression of the reconstruction error function in the loss function of the data augmentation model is:

[0049]

[0050] In the formula, recon_batch represents the reconstructed protein abundance data; data represents the original input protein abundance data; recon_batch i represents the i-th item of reconstructed protein abundance data; data i represents the i-th item of original input protein abundance data;

[0051] The expression of the KL divergence function in the loss function of the data augmentation model is:

[0052]

[0053] In the formula, μ represents the mean of the latent space; D represents the dimension of the latent space; μ i represents the mean of the i-th dimension of the latent space; σ 2 i represents the variance of the i-th dimension of the latent space;

[0054] In the data augmentation model processing module, during the supervised training of the data augmentation model, the log variance output by the encoder of the data augmentation model is clipped to the range (-10, 10).

[0055] As a preferred solution of the proteomic biomarker identification system, in the group set generation module:

[0056] Calculate the covariance matrix of the protein abundance data, solve the eigenvalues and eigenvectors of the covariance matrix, select the eigenvectors corresponding to the covariance matrix that reaches the set eigenvalue as the principal components, and project the protein abundance data onto the determined principal components;

[0057] During the similarity analysis of the new sample data after principal component analysis: evaluate the similarity between the samples of the specified group according to the distribution of the data points after dimensionality reduction by principal component analysis.

[0058] As a preferred solution of the proteomic biomarker recognition system, in the biomarker recognition module:

[0059] Identify potential biomarkers through the group set using differential expression analysis or random forest method, rank the potential biomarkers according to the fold change or the importance parameter filled in the model output of the specified data, and obtain the recognition result of the multi-group exosomal proteomic biomarkers.

[0060] In a third aspect, the present invention provides an electronic device, including: one or more processors;

[0061] A storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors implement the proteomic biomarker recognition method described in any implementation manner of the first aspect above.

[0062] In a fourth aspect, the present invention provides a computer-readable medium storing a computer program thereon, wherein when the program is executed by a processor, the proteomic biomarker recognition method described in any implementation manner of the first aspect above is implemented.

[0063] The present invention has the following advantages: adding one-hot encoded grouping labels to the protein abundance data in the mass spectrometry quantitative data, and encoding the group information corresponding to the protein abundance data into a binary vector; constructing a data augmentation model based on a variational autoencoder, using the protein abundance data and the one-hot encoded grouping labels to perform supervised training on the data augmentation model, and using the decoder of the trained data augmentation model to generate new sample data; performing principal component analysis and similarity analysis on the new sample data; according to the results of the principal component analysis, dividing the new sample data with similarity reaching the set threshold into a group set; identifying potential biomarkers using a specified method, and ranking the potential biomarkers according to the importance parameters of the potential biomarkers to obtain the recognition result of the multi-group exosomal proteomic biomarkers.

[0064] The present invention is more suitable for data augmentation in exosomal proteomics, avoiding gradient explosion during model training; it can retain more information, enabling the data augmentation model to capture the complex features of the data more comprehensively and improve the overall performance; a large number of new samples that conform to the potential features (similarity) of the original data and have a certain degree of diversity (difference) can be obtained, increasing the sample size and sample diversity, greatly reducing the difficulty of downstream analysis; when processing exosomal proteomics data and multi-group sample comparison, it significantly improves the robustness of the data and the accuracy of biomarker identification, overcomes the limitations of the prior art, and provides a more effective solution for the field of bioinformatics. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are merely exemplary, and for those of ordinary skill in the art, other implementation drawings can be obtained by extension based on the provided drawings without creative efforts.

[0066] Figure 1 Schematic flowchart of the proteomic biomarker identification method provided in the embodiment of the present invention;

[0067] Figure 2 Schematic diagram of similarity analysis in the proteomic biomarker identification method provided in the embodiment of the present invention, where (a) is the dimensionality reduction result of the real sample; (b) is the dimensionality reduction result of the simulated sample generated during the 10th iteration of the data augmentation model training; (c) is the dimensionality reduction result of the simulated sample generated during the 90th iteration of the data augmentation model training; (d) is the dimensionality reduction result of the simulated sample generated during the 180th iteration of the data augmentation model training;

[0068] Figure 3 Schematic diagram of the results of tear exosomal protein biomarkers for Parkinson's disease (compared with Alzheimer's disease) identified by the proteomic biomarker identification method provided in the embodiment of the present invention;

[0069] Figure 4 Receiver operating characteristic curve of the results of tear exosomal protein biomarkers for Parkinson's disease (compared with Alzheimer's disease) identified by the proteomic biomarker identification method provided in the embodiment of the present invention;

[0070] Figure 5 Schematic diagram of the architecture of the proteomic biomarker identification system provided in the embodiment of the present invention;

[0071] Figure 6 Schematic diagram of the architecture of the electronic device provided in the embodiment of the present invention. Detailed implementation manners

[0072] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0073] Embodiment 1

[0074] Refer to Figure 1 , Embodiment 1 of the present invention provides a method for identifying proteomic markers, including the following steps:

[0075] S1. Obtain exosomal proteomic mass spectrometry quantitative data, where the mass spectrometry quantitative data includes protein abundance data of a preset group; perform data preprocessing on the mass spectrometry quantitative data, and the data preprocessing includes data quality control, data normalization, and missing value filling;

[0076] S2. Add one-hot encoded grouping labels to the protein abundance data in the mass spectrometry quantitative data after the data preprocessing, and encode the group information corresponding to the protein abundance data into a binary vector;

[0077] S3. Construct a data augmentation model based on a variational autoencoder, where the loss function of the data augmentation model includes a reconstruction error function and a KL divergence function; use an Adam optimizer to minimize the loss function of the data augmentation model to optimize the parameters of the data augmentation model;

[0078] S4. Use the protein abundance data after the data preprocessing and the one-hot encoded grouping labels to perform supervised training on the data augmentation model, so that the log variance output by the encoder of the data augmentation model is clipped to the range of (-10, 10), and use the decoder of the trained data augmentation model to generate new sample data;

[0079] S5. Perform principal component analysis on the new sample data generated by the decoder of the data augmentation model, and perform similarity analysis on the new sample data after the principal component analysis;

[0080] S6. According to the results of the principal component analysis, divide the new sample data with similarity reaching a set threshold into a group set, and use the group set to identify potential markers by differential expression analysis or random forest method, and sort the potential markers according to the fold change or the importance parameters filled by the specified data output by the model to obtain the identification result of multi-group exosomal proteomic markers.

[0081] In this embodiment, in step S1, exosomal proteomic mass spectrometry quantitative data is obtained, including protein abundance data of multiple groups (such as different disease sample groups and healthy groups). Among them, during the data quality control of the mass spectrometry quantitative data, protein abundance data with missing values in more than 2 / 3 of the samples is removed to ensure the reliability and stability of the data. During the data normalization of the mass spectrometry quantitative data, the specified normalization method is used to normalize the protein abundance data to eliminate systematic biases between samples. During the missing value filling of the mass spectrometry quantitative data, the specified data filling model is used to fill the missing values in the protein abundance data after data normalization. Random forest can infer and fill in the missing values based on data characteristics, improving the integrity of the data.

[0082] Among them, at least one of Cycloess, log2, median, mean, vsn, quantile, rlr, and gi can be used as the specified normalization method. The Cycloess method is preferred.

[0083] Specifically, the Cycloess (Cyclic Locally Weighted Scatterplot Smoothing) normalization method is a technique for normalizing gene expression data, used to correct systematic biases in the data. This method smooths the data through local regression (Loess) to remove technical variations and maintain biological variations. The basic principle is based on the MA plot (M-A plot), where M represents the log ratio and A represents the average log intensity, and the dependence between M and A is eliminated by iteratively applying local weighted regression (loess).

[0084] The following are the detailed steps and formulas of the Cycloess normalization method:

[0085] 1) For each pair of samples i and j, calculate the M and A values:

[0086] M = log 2 (X i ) - log 2 (X j )

[0087] A = (log 2 (X i ) + log 2 (X j )) / 2

[0088] Among them, X i and X j represent the expression values of samples i and j respectively;

[0089] 2) Apply loess regression to the M - A plot to obtain the fitted curve f(A); Loess regression formula: Loess regression uses locally weighted polynomial regression. Loess regression fits a local polynomial for each point (x 0 , y 0 ) in the dataset. During the fitting process, points closer to (x 0 , y 0 ) receive larger weights, while points farther away receive smaller weights. The fitting formula is:

[0090] y = β 0 + β 1 (x - x 0 ) + β 2 (x - x 0 ) 2 +... + β p (x - x 0 ) p

[0091] where p is usually taken as 1 or 2; β 0 , β 1 , β 2 ... β p are regression coefficients, and these coefficients are estimated by weighted least squares. Usually, the tricube function is used as the weight function:

[0092] w(x) = (1 - |x| 3 ) 3 , when |x| < 1;

[0093] w(x) = 0, when |x| ≥ 1;

[0094] Here, x is the standardized distance, usually obtained by dividing the actual distance by a certain bandwidth parameter.

[0095] 3) Calculate the adjusted expression value:

[0096] X i_adjusted = X i * 2 (-f(A) / 2)

[0097] X j_adjusted = X j * 2 (f(A) / 2)

[0098] 4) Iterate in a loop, repeating steps 1) - 3) until the trend in the M - A plot disappears or the preset number of iterations is reached. The preset number of iterations is set to 3 in this application.

[0099] Thus, in this way, the loess regression can capture the non-linear relationships in the data, enabling the Cycloess method to effectively handle the non-linear deviations in the protein mass spectrometry data and enhancing the data consistency.

[0100] In this embodiment, in step S2, one-hot encoded grouping labels are added to the end of the protein abundance data after normalization and missing value imputation, encoding the group information as binary vectors for subsequent data augmentation model training.

[0101] Due to the problems of low consistency between repeated samples and high data noise in exosomal proteomics data, the data itself is too complex to provide sufficient effective information. Therefore, conventional unsupervised training methods cannot be directly used to train the model, and additional manual guidance for the model is necessary. By adding one-hot encoded grouping labels to the end of the protein abundance data after normalization and missing value imputation, the learning effect of the data augmentation model can be significantly improved.

[0102] In this embodiment, in step S3, the architecture of the data augmentation model is as follows:

[0103] Encoder: Receives the preprocessed protein abundance data and one-hot encoded grouping labels as input. The encoder maps the high-dimensional input data to the latent space to represent the latent variable distribution of the data.

[0104] Latent space: Generation of latent variables, using the mean and variance output by the encoder, generating latent variables through the reparameterization trick.

[0105] Decoder: Similar to the inverse process of the encoder, restores the data set based on the vectors in the latent space.

[0106] Since the purpose is data augmentation rather than directly performing a discrimination task, there is no need to test the discrimination ability of the data augmentation model on new samples. Therefore, it is not necessary to use an unsupervised training method. Essentially, the present invention maximizes its data augmentation ability at the expense of the discrimination ability of the model.

[0107] Specifically, the data augmentation model can learn the true features of exosomal proteomics data and reduce them to the latent variable space, where the samples in this space are represented as lower-dimensional vectors, minimizing the interference of noise. Based on the latent variable space, the data augmentation model can generate a large number of new samples with high consistency among samples within the group, thereby improving the accuracy and robustness of differential protein identification.

[0108] Among them, the loss function of the data augmentation model includes a reconstruction error function and a Kullback-Leibler divergence function. The reconstruction error measures the difference between the data generated by the decoder and the original input data, using the Mean Squared Error (MSE). The KL divergence measures the difference between the latent variable distribution and the standard normal distribution.

[0109] Specifically, the expression of the reconstruction error function in the loss function of the data augmentation model is:

[0110]

[0111] In the formula, recon_batch represents the reconstructed protein abundance data; data represents the original input protein abundance data; recon_batch i represents the i-th item of reconstructed protein abundance data; data i represents the i-th item of original input protein abundance data.

[0112] Specifically, the expression of the KL divergence function in the loss function of the data augmentation model is:

[0113]

[0114] In the formula, μ represents the mean of the latent space; D represents the dimension of the latent space; μ i represents the mean of the i-th dimension of the latent space; σ 2 i represents the variance of the i-th dimension of the latent space.

[0115] Among them, adding the above reconstruction error and KL divergence loss gives the complete ELBO (Evidence Lower Bound) loss:

[0116]

[0117] By using the Adam optimizer to minimize the loss function, the parameters of the data augmentation model are optimized so that the data augmentation model can effectively generate samples similar to the input data distribution from the latent space.

[0118] In this embodiment, in step S4, the protein abundance data that has been quality controlled, missing value imputed, and standardized is horizontally concatenated with the one-hot encoded group labels for supervised training of the data augmentation model, and the log variance output by the encoder is clipped to the range (-10, 10).

[0119] Specifically, due to the complex data characteristics of exosome proteomics, there are a large number of noises or outliers. These values may generate unstable gradients during the training process. Especially when the magnitude of the noise or outlier is much larger than that of the normal data, they may cause a substantial increase in the gradient during backpropagation, ultimately leading to gradient explosion. Gradient explosion refers to the situation where, during backpropagation, the gradient gradually becomes larger and even approaches infinity. This will cause the weight update to be too large, making the model unstable during the training process. Gradient explosion may cause the model weights to diverge rapidly after several updates, preventing the model from learning effective feature representations. By taking the nearest boundary value for any log variance value outside the range of (-10, 10) and leaving the values within the range unchanged, the magnitude of the gradient can be effectively controlled, which helps to stabilize the training process of the data augmentation model and prevent it from being too large to cause gradient explosion.

[0120] In many bioinformatics studies, the encoder of the data augmentation model is directly used to reduce the dimension of the data and infer the features (proteins) with greater contribution. This usually requires the latent space dimension output by the encoder to be consistent with the number of groups. In the present invention, the data augmentation model is designed not to directly use the encoder but the decoder, allowing the encoder to output a higher latent space dimension without being restricted by the number of groups. This design can retain more information, enabling the data augmentation model to capture the complex features of the data more comprehensively and improve the overall performance. At the same time, when the learning effect of the data augmentation model is not good, the dimension of the latent space can be dynamically adjusted, which provides greater flexibility for downstream applications.

[0121] Specifically, using the decoder of the trained data augmentation model to generate a large number of new samples makes it possible for downstream analysis methods that require a large number of samples. Since the data augmentation model has learned the features of the data and mapped them into the latent variable space, the encoder can generate new samples that conform to the latent features of the original data based on the latent variable space. During the generation process, by randomly sampling latent variables from a prior distribution that follows a standard normal distribution, diversity is introduced into the newly generated samples. Therefore, by using the decoder of the trained data augmentation model to generate a large number of new samples, a large number of new samples that conform to the latent features of the original data (similarity) and have a certain degree of diversity (difference) can be obtained. The increased sample size and diversity greatly reduce the difficulty of downstream analysis.

[0122] In this embodiment, in step S5, during the principal component analysis of the new sample data generated by the decoder of the data augmentation model: calculating the covariance matrix of the protein abundance data, solving the eigenvalues and eigenvectors of the covariance matrix, selecting the eigenvectors corresponding to the covariance matrix reaching the set eigenvalue as the principal components, and projecting the protein abundance numbers onto the determined principal components; during the similarity analysis of the new sample data after principal component analysis: evaluating the similarity between samples of a specified group according to the distribution of data points after dimensionality reduction by principal component analysis.

[0123] Specifically, observe the distribution of data points after dimensionality reduction by principal component analysis and evaluate the similarity between samples of different groups. The closer the data points are on the plane, the higher the similarity. See Figure 2 , which is an application example of the data augmentation model in tear exosome proteomics samples. Use the principal component analysis method to reduce the dimensionality of omics data. Figure 2 Part (a) in it is the dimensionality reduction result of real samples. Samples of each group are mixed together and cannot be distinguished, indicating that the data pattern is complex and the noise is large; Figure 2 Part (b) in it is the dimensionality reduction result of the simulated samples generated by the data augmentation model when it is trained to the 10th iteration. At this time, the model still cannot well distinguish samples of each group; Figure 2 Part (c) in it is the dimensionality reduction result of the simulated samples generated by the data augmentation model when it is trained to the 90th iteration. At this time, the model gradually has a tendency to distinguish samples of each group; Figure 2 Part (d) in it is the dimensionality reduction result of the simulated samples generated by the data augmentation model when it is trained to the 180th iteration. At this time, the relative distance relationship between samples of each group is consistent with experience, indicating that the data augmentation model can learn the latent representation of samples and output new samples that are enhanced and have a certain degree of diversity.

[0124] In this embodiment, in step S6, according to the results of principal component analysis, several groups with higher similarity can be defined as a group set. First, identify the markers of several group sets with lower similarity, and then identify the markers of each group within the set. Perform differential expression analysis on the generated new samples or use a specified data filling model to judge important proteins to obtain proteins with large differences between groups or proteins that contribute greatly to phenotype classification. And sort the potential markers according to the fold change of proteins or the importance output by the specified data filling model to obtain a list of potential markers for further consideration and verification in the future.

[0125] Among them, the specified data filling model can be at least one of a random forest model, fixed-value filling, mean / median / mode filling, K-nearest neighbor method (KNN), regression-based filling (such as linear regression, random forest regression), genotype filling (specifically for genotype data), interpolation method (such as linear interpolation, polynomial interpolation), and multiple imputation.

[0126] See Figure 3 , for the tear exosome protein markers of Parkinson's disease (compared with Alzheimer's disease) identified by this method, additional samples were detected using enzyme-linked immunosorbent assay (hereinafter referred to as ELISA). Figure 3 The p-value in represents the confidence level of the T one-tailed test, and HSGP2 is the gene name of this protein marker. The results show that there are significant differences in HSGP2 of tear exosomes between Parkinson's disease and Alzheimer's disease samples.

[0127] See Figure 4 , used Figure 3 The detection results of were used to draw a receiver operating characteristic curve (ROC) to illustrate the performance of the marker in distinguishing the two diseases. AUC represents the area under the ROC curve, and the closer it is to 1, the better the distinguishing performance.

[0128] By applying the method of the embodiment of the present invention to the tear exosome proteome data of five nerve-related diseases and healthy human samples, the effectiveness of the method was verified. By comparing the markers identified from the new samples generated based on the enhanced simulation dataset with the markers directly identified from the real samples, the performance of the new method was evaluated. The results show that the data enhancement model can well learn the latent representation of the mass spectrometry data and reflect the similarities between diseases through the generated new samples. The theoretical performance of the markers generated based on the data enhancement simulation dataset is better than that of the markers directly identified from the real samples, proving the superiority of the data enhancement module.

[0129] In summary, the present invention obtains mass spectrometry quantitative data of exosome proteomes, where the mass spectrometry quantitative data includes protein abundance data of a preset group; performs data preprocessing on the mass spectrometry quantitative data, and the data preprocessing includes data quality control, data normalization, and missing value filling; adds one-hot encoded grouping labels to the protein abundance data in the mass spectrometry quantitative data after the data preprocessing, and encodes the group information corresponding to the protein abundance data into a binary vector; constructs a data augmentation model based on a variational autoencoder, and the loss function of the data augmentation model includes a reconstruction error function and a KL divergence function; uses an Adam optimizer to minimize the loss function of the data augmentation model to optimize the parameters of the data augmentation model; uses the protein abundance data after the data preprocessing and the one-hot encoded grouping labels to perform supervised training on the data augmentation model, so that the logarithmic variance output by the encoder of the data augmentation model is clipped to the range of (-10, 10), and uses the decoder of the trained data augmentation model to generate new sample data; performs principal component analysis on the new sample data generated by the decoder of the data augmentation model, and performs similarity analysis on the new sample data after the principal component analysis; according to the results of the principal component analysis, divides the new sample data with similarity reaching a set threshold into a group set, and uses differential expression analysis or a random forest method to identify potential markers through the group set, and ranks the potential markers according to the fold change or the importance parameters filled by the specified data in the model output to obtain the identification result of multi-group exosome proteome markers. The present invention is more suitable for data augmentation of exosome proteomics and avoids gradient explosion during model training; can retain more information, enabling the data augmentation model to capture the complex features of the data more comprehensively and improve the overall performance; can obtain a large number of new samples that conform to the potential features of the original data (similarity) and have a certain diversity (difference), increasing the sample size and sample diversity, greatly reducing the difficulty of downstream analysis; when processing exosome proteome data and multi-group sample comparison, significantly improves the robustness of the data and the accuracy of marker identification, overcomes the limitations of the prior art, and provides a more effective solution for the field of bioinformatics.

[0130] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.

[0131] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0132] Embodiment 2

[0133] Referring to Figure 5 , Embodiment 2 of the present invention also provides a proteomic biomarker identification system, including:

[0134] A data preparation module 100, configured to add one-hot encoded group tags to the protein abundance data in the mass spectrometry quantitative data, and encode the group information corresponding to the protein abundance data into a binary vector;

[0135] A data augmentation model processing module 200, configured to construct a data augmentation model based on a variational autoencoder, use the protein abundance data and the one-hot encoded group tags to perform supervised training on the data augmentation model, and generate new sample data using the decoder of the trained data augmentation model;

[0136] A group set generation module 300, configured to perform principal component analysis and similarity analysis on the new sample data; according to the results of the principal component analysis, divide the new sample data with similarity reaching a set threshold into a group set;

[0137] A biomarker identification module 400, configured to identify potential biomarkers using a specified method, sort the potential biomarkers according to the importance parameters of the potential biomarkers, and obtain the identification results of the multi-group exosomal proteomic biomarkers.

[0138] In this embodiment, a data preprocessing module 500 is further included, configured to perform data preprocessing on the mass spectrometry quantitative data, and the data preprocessing includes data quality control, data normalization, and missing value filling;

[0139] During the data quality control process, remove the protein abundance data with missing abundance values in more than a specified number of samples;

[0140] During the data normalization process, perform normalization processing on the protein abundance data using a specified normalization method;

[0141] During the missing value filling process, perform missing value filling on the protein abundance data after data normalization using a specified data filling model.

[0142] In this embodiment, in the data augmentation model processing module 200:

[0143] The loss function of the data augmentation model includes a reconstruction error function and a KL divergence function; the Adam optimizer is used to minimize the loss function of the data augmentation model to optimize the parameters of the data augmentation model;

[0144] The expression of the reconstruction error function in the loss function of the data augmentation model is:

[0145]

[0146] In the formula, recon_batch represents the reconstructed protein abundance data; data represents the original input protein abundance data; recon_batch i represents the i-th item of reconstructed protein abundance data; data i represents the i-th item of original input protein abundance data.

[0147] The expression of the KL divergence function in the loss function of the data augmentation model is:

[0148]

[0149] In the formula, μ represents the mean of the latent space; D represents the dimension of the latent space; μ i represents the mean of the i-th dimension of the latent space; σ 2 i represents the variance of the i-th dimension of the latent space.

[0150] In the data augmentation model processing module 200, during the supervised training process of the data augmentation model, the log variance output by the encoder of the data augmentation model is clipped to the range of (-10, 10).

[0151] In this embodiment, in the group set generation module 300:

[0152] Calculate the covariance matrix of the protein abundance data, solve the eigenvalues and eigenvectors of the covariance matrix, select the eigenvectors corresponding to the covariance matrix reaching the set eigenvalues as the principal components, and project the protein abundance numbers onto the determined principal components;

[0153] During the similarity analysis of the new sample data after principal component analysis: According to the distribution of the data points after dimensionality reduction by principal component analysis, evaluate the similarity between the specified group samples.

[0154] In this embodiment, in the biomarker identification module 400:

[0155] Identifying potential markers through the group set using differential expression analysis or random forest method, ranking the potential markers according to the fold change or the importance parameter filled by the specified data in the model output, and obtaining the identification results of the multi-group exosome proteome markers.

[0156] It should be noted that the information interaction, execution process, etc. between the above-mentioned system modules, due to being based on the same concept as the method embodiment in Embodiment 1 of the present application, bring the same technical effects as the method embodiment of the present application. The specific content can be seen in the description of the method embodiment shown above in the present application, and will not be repeated here.

[0157] Embodiment 3

[0158] Reference Figure 6 , which shows a schematic structural diagram of an electronic device (such as a computing device) 600 suitable for implementing some embodiments of the present disclosure. The electronic devices in some embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage ranges of the embodiments of the present disclosure.

[0159] As Figure 6 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0160] Generally, the following devices may be connected to the I / O interface 605: an input device 603 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6An electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or have all the devices shown. Instead, more or fewer devices may be implemented or had. Figure 6 Each block shown in Figure 6 may represent a device or, as needed, multiple devices.

[0161] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program code for performing the methods shown in the flowcharts. In such some embodiments, the computer program can be downloaded and installed from a network through a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by a processing device 601, the above functions defined in the methods of some embodiments of the present disclosure are performed.

[0162] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0163] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0164] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect through the Internet).

[0165] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0166] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes: a first acquisition unit, a determination unit, a training unit, a second acquisition unit, a fine-tuning unit, and an input unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the input unit can also be described as "a unit that inputs the synthetic information to be detected into the above-mentioned target intelligent synthetic information recognition model to obtain a synthetic information detection result".

[0167] The functions described above herein can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.

[0168] Although the present invention has been described in detail above with general descriptions and specific embodiments, on the basis of the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of the present invention claimed.

Claims

1. A method for identifying proteomic markers, characterized in that: include: Add one-hot encoded grouping labels to the protein abundance data in the mass spectrometry quantitative data, and encode the group information corresponding to the protein abundance data into a binary vector; Constructing a data enhancement model based on a variational autoencoder, using the protein abundance data and the one-hot encoded grouping labels to perform supervised training on the data enhancement model, and using the decoder of the trained data enhancement model to generate new sample data; Performing principal component analysis and similarity analysis on the new sample data; dividing the new sample data whose similarity reaches a set threshold into a group set according to the result of the principal component analysis; Potential markers are identified using a specified method, and the potential markers are ranked according to importance parameters of the potential markers to obtain multiple groups of exosome proteome marker identification results.

2. The method for identifying proteomic markers according to claim 1, characterized in that: The method also includes performing data preprocessing on the mass spectrometry quantitative data, wherein the data preprocessing includes data quality control, data normalization and missing value filling.

3. The method for identifying proteomic markers according to claim 2, characterized in that: In the data quality control process, protein abundance data with abundance values ​​exceeding a specified number of samples being missing values ​​are removed; During the data normalization process, a specified normalization method is used to normalize the protein abundance data; During the missing value filling process, a specified data filling model is used to fill missing values ​​in the protein abundance data after data normalization.

4. The method for identifying proteomic markers according to claim 1, characterized in that: The loss function of the data enhancement model includes a reconstruction error function and a KL divergence function; using an optimizer to minimize the loss function of the data enhancement model to optimize the parameters of the data enhancement model; The expression of the reconstruction error function in the loss function of the data enhancement model is: In the formula, recon_batch represents the reconstructed protein abundance data; data represents the original input protein abundance data; recon_batch i Represents the reconstructed protein abundance data of item i; data i Represents the i-th protein abundance data of the original input; The expression of the KL divergence function in the loss function of the data enhancement model is: In the formula, μ represents the mean of the latent space; D represents the dimension of the latent space; μ i represents the mean of the i-th dimension of the latent space; σ 2 i represents the variance of the i-th dimension of the latent space.

5. The method for identifying proteomic markers according to claim 1, characterized in that: During the supervised training of the data augmentation model, the logarithmic variance of the encoder output of the data augmentation model is clipped to the range of (-10, 10).

6. The method for identifying proteomic markers according to claim 1, characterized in that: During the principal component analysis of the new sample data: Calculate the covariance matrix of protein abundance data, solve the eigenvalue and eigenvector of the covariance matrix, select the eigenvector corresponding to the covariance matrix that reaches the set eigenvalue as the principal component, and project the protein abundance number onto the determined principal component; In the process of similarity analysis of new sample data after principal component analysis: based on the distribution of data points after dimensionality reduction by principal component analysis, the similarity between samples in the specified group is evaluated.

7. The method for identifying proteomic markers according to claim 1, characterized in that: Potential markers are identified by using differential expression analysis or random forest method through the group set, and the potential markers are ranked according to the difference multiple or the importance parameter output by the specified data filling model to obtain multiple groups of exosome proteome marker identification results.

8. A proteomic marker identification system, using the proteomic marker identification method according to any one of claims 1 to 7, characterized in that: include: A data preparation module is used to add one-hot encoded grouping labels to the protein abundance data in the mass spectrometry quantitative data, and encode the group information corresponding to the protein abundance data into a binary vector; A data enhancement model processing module is used to construct a data enhancement model based on a variational autoencoder, use the protein abundance data and the one-hot encoded grouping labels to perform supervised training on the data enhancement model, and use the decoder of the trained data enhancement model to generate new sample data; A group set generation module is used to perform principal component analysis and similarity analysis on the new sample data; according to the result of the principal component analysis, the new sample data whose similarity reaches a set threshold is divided into a group set; The marker identification module is used to identify potential markers using a specified method, sort the potential markers according to the importance parameters of the potential markers, and obtain multiple groups of exosome proteome marker identification results.

9. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the proteomic marker identification method according to any one of claims 1 to 7.

10. A computer readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the proteomic marker identification method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method, system and medium for integrally screening proteome clinical biomarkers

    CN114550832A

  • Bladder cancer metabolism marker screening method and system based on deep learning

    CN114997303A

  • Method for analysis of omics data

    US20240087677A1

  • Method for analysis of omics data

    WO2022161824A1

  • Sample preparation for glycoproteomic analysis that includes diagnosis of disease

    WO2023164672A2