A protein post-translational modification data enhancement method based on generative adversarial networks

Through the data enhancement method based on the generative adversarial network, the ESM-2 model and RP-CGAN are used to generate pseudo-samples, which solves the problem of model selection and data imbalance in the prediction of protein S-sulfination site, improves prediction performance and stability, and has cross-domain application potential.

CN120412756BActive Publication Date: 2025-08-29JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510920403.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-08-29
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient model selection, scarce data samples and category imbalance in the prediction of protein S-sulfination sites, resulting in degradation of prediction performance and insufficient generalization capabilities, and lack of effective data enhancement methods.

Method used

Using the protein post-translation modified data enhancement method based on the generative adversarial network, features are extracted through the ESM-2 pre-trained model, and pseudo-samples are generated by constructing the RP-CGAN model. Combined with Softplus relative adversarial loss and gradient regular term optimization, high-quality pseudo-samples are screened for classifier training.

Benefits of technology

It significantly improved the recognition and generalization ability of positive samples, improved the problem of category imbalance, improved the accuracy of classifiers, AUC, F1 score and G-mean, and enhanced the stability of the model and the potential for cross-domain expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412756B_ABST
    Figure CN120412756B_ABST
Patent Text Reader

Abstract

This invention is applicable to the field of proteomics and provides a method for enhancing protein post-translational modification data based on generative adversarial networks (GANs). This method can effectively alleviate class imbalance and enhance the ability to identify positive samples. It generates minority class pseudo samples through an improved cluster-enhanced conditional GAN ​​(RP-CGAN), and extracts features using the ESM-2 pre-trained protein language model to improve both positive sample identification and generalization capabilities. The method integrates ESM-2 feature extraction with RP-CGAN data enhancement technology to enhance overall prediction performance and classification stability, significantly improving key metrics for multiple classifiers. Multi-dimensional indicator screening and adaptive regulation ensure that the enhanced data approximates the true distribution. Leveraging a lightweight architecture and convergence strategy, the method optimizes training efficiency and model stability. The method possesses strong stability and cross-domain potential, enabling migration to other protein modification prediction tasks and possessing potential for application in cross-domain imbalanced classification scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of proteomics, and in particular relates to a protein post-translational modification data enhancement method based on a generative adversarial network. Background Art

[0002] Sulfenylation is a post-translational modification (PTM) of proteins, defining covalent and chemical modifications of protein residues. It plays a crucial role in regulating various biological functions, including cardiovascular homeostasis. Identifying S-sulfenylation sites is crucial for further understanding their regulatory functions. Over the past few decades, proteomics techniques such as chromatin immunoprecipitation, liquid chromatography, and mass spectrometry have proven effective in detecting S-sulfenylation sites. However, these laboratory techniques have significant drawbacks, such as high time and experimental costs. Therefore, the development of effective computational methods is urgent to improve the efficiency of identifying sulfenylation sites in proteins, reduce labor and time costs, and accelerate the study of their potential roles. Deep learning algorithms still have considerable room for development in predicting S-sulfenylation sites.

[0003] Currently, some methods for predicting protein S-sulfenylation sites have incorporated deep learning and natural language processing (NLP) techniques to mine hidden contextual information within sequences. Compared to traditional machine learning models based on handcrafted features, these methods offer improved modeling capabilities and predictive effectiveness. However, existing deep learning methods still suffer from the following significant drawbacks:

[0004] First, in terms of model selection, existing methods often use lightweight embedding methods or simplified language models rather than high-performance protein pre-trained models like ESM-2. As a large-scale protein language model, ESM-2 can extract richer contextual semantics and structural features from amino acid sequences, and its feature expression capabilities far exceed those of traditional natural language processing (NLP) or shallow network embedding methods. However, ESM-2 has not yet been effectively applied in the field of S-sulfenylation site prediction.

[0005] Secondly, at the data level, although previous studies have used the publicly available iSulf-Cys dataset, one of the most comprehensive S-sulfenylation site datasets available, which screens modification sites in human proteins from the NCBI database, this dataset suffers from sample scarcity and a severe imbalance between positive and negative classes. Among 8,169 sequence data, there are only 1,045 S-sulfenylation site samples. This data distribution can easily lead to overfitting of the model to negative samples during training, resulting in poor prediction performance.

[0006] Third, in terms of technical application, existing technologies have yet to incorporate generative models such as generative adversarial networks (GANs) to effectively augment biological sequence data. Given the limited number of positive samples (sequences containing S-sulfenylation sites), the lack of appropriate data augmentation methods will limit the model's learning capabilities, especially when faced with unknown samples in real-world scenarios, further reducing the model's generalization ability.

[0007] To solve the above problems, the present invention proposes a protein post-translational modification data enhancement method based on generative adversarial networks. Summary of the Invention

[0008] The purpose of the present invention is to provide a protein post-translational modification data enhancement method based on generative adversarial networks, aiming to solve the problems raised in the above background technology.

[0009] The purpose of the present invention is achieved through the following technical solutions:

[0010] A protein post-translational modification data enhancement method based on a generative adversarial network comprises the following steps:

[0011] Step S1: data preprocessing and feature extraction;

[0012] Read the protein sequence file, divide it into training set and test set, use the ESM-2 pre-trained protein language model for feature encoding, and generate standardized feature vectors;

[0013] Step S2: Sample enhancement of clustering-enhanced conditional generative adversarial network;

[0014] Construct an RP-CGAN model consisting of a generator and a discriminator; the generator input contains a random noise vector, a target category label, and a category center vector generated by K-means clustering, and outputs a target category pseudo sample; the discriminator adopts a dual-output structure to judge the sample authenticity and category label respectively, and combines Softplus Relative adversarial loss, binary cross entropy loss, and gradient regularization term optimize training stability;

[0015] Step S3: Pseudo-sample screening;

[0016] The trained RP-CGAN model is used to generate pseudo samples exceeding the target number. Through a multi-metric constrained screening mechanism, pseudo samples that are closest to the distribution of real samples are selected.

[0017] Step S4: classifier training and evaluation;

[0018] Merge the real training data with the filtered pseudo samples and input them into the classifier for training to optimize the classification performance;

[0019] Step S5: predict output;

[0020] Use the trained classifier to predict the new protein sequence and output the modification site probability score.

[0021] Furthermore, the data preprocessing and feature extraction steps include:

[0022] With the target residue as the center, 10 amino acid residues upstream and downstream were extracted to form a 21-residue window sequence;

[0023] Load the ESM-2 pre-trained protein language model and alphabet;

[0024] Use the get_batch_converter() method to convert the protein sequence into an input format that the model can recognize;

[0025] Extract the token-level embedding representation of the 33rd layer of the model as the feature of each residue;

[0026] The original dimension embedding is retained for each sequence to form a structure-aware feature matrix;

[0027] The feature matrix of each sequence is averaged to obtain the final feature vector of the sequence.

[0028] Furthermore, the generator design of the RP-CGAN model includes:

[0029] The input condition is a random noise vector , category labels And the category center vector obtained by K-means clustering , generates samples close to the real distribution through the joint input splicing mechanism ;

[0030] The step of generating the category center vector includes:

[0031] K-means clustering is performed on the minority and majority class samples respectively, and the optimal number of clusters is determined by the DBI indicator; the optimal cluster centers of the minority and majority classes are extracted as the prior condition input of the generator.

[0032] Furthermore, the loss function of the RP-CGAN model includes:

[0033] Softplus Relative adversarial loss of variants:

[0034] ;

[0035] in, To counter the loss function; Represents the distribution of real samples The following sample Seek expectations; Represents the distribution of generated samples The following sample Seek expectations; is the true sample feature vector; Pseudo samples generated by the generator; Softplus The function is defined as ;

[0036] is the authenticity output of the discriminator;

[0037] Binary cross entropy loss:

[0038] ;

[0039] in, is the classification loss function; Represents the joint distribution of sample features and labels Seek expectations; is the sample feature vector; is the true category label; For the discriminator to sample The category prediction output value of

[0040] Gradient regularization term:

[0041] R1 regularization term:

[0042] ;

[0043] R2 regularization term:

[0044] ;

[0045] in, is the gradient regularization term for real samples; is the gradient regularization term for the generated samples; Represents the discriminator output with respect to the real sample input gradient; Represents the discriminator output with respect to the generated sample input gradient; Represents the square of the L2 norm of the vector;

[0046] Total loss function:

[0047] Discriminator loss function:

[0048] ;

[0049] Generator loss function:

[0050] ;

[0051] in, is the total loss function of the discriminator; is the total loss function of the generator; is the adversarial loss term of the discriminator; is the adversarial loss term of the generator; is the classification loss term of the discriminator; is the classification loss term of the generator; is the weight of classification loss; and They are Regularization term and The weight of the regularization term.

[0052] Furthermore, the pseudo-sample screening step includes:

[0053] Calculate sensitivity and specificity through confusion matrix and dynamically determine the number of pseudo sample enhancements ; Generate using generator pseudo samples; use AutoEncoder to construct feature space, calculate the Pearson correlation coefficient, Euclidean distance and minimum mean square error of three types of distance indicators between pseudo samples and real samples, normalize and weight the three types of distance indicators to obtain a comprehensive score; select the top sample with the lowest comprehensive score pseudo samples, forming the final enhanced sample set .

[0054] Furthermore, the classifier training and evaluation steps include:

[0055] Random forest, support vector machine, XGBoost, and LightGBM were used for training; classification performance was evaluated using accuracy, F1 score, AUC, MCC, and G-mean value.

[0056] Furthermore, the prediction output step includes:

[0057] Features are extracted from the new protein sequence and input into the trained classifier; the probability score of each residue being a modification site is output, and the modification site is determined based on the threshold.

[0058] Compared with the prior art, the present invention has the following beneficial effects:

[0059] 1. Effectively alleviate category imbalance and enhance the ability to identify positive samples: The improved cluster-enhanced conditional generative adversarial network (RP-CGAN) is used to generate minority (positive) pseudo samples. The high-dimensional semantic features extracted by the ESM-2 pre-trained protein language model are combined to capture protein sequence context and structural information to supplement scarce positive data. The generator introduces category labels and real sample cluster centers as conditional inputs to fit the target category distribution. The discriminator's dual-output structure guides the generation of samples close to the true distribution. The data from the embodiment shows that the enhanced classifier's sensitivity (Sen), F1 score, MCC, G-mean and other balance indicators are significantly improved, effectively alleviating the model bias caused by category imbalance and improving the ability to identify positive samples and generalize.

[0060] 2. Significantly improve the overall prediction performance and classification stability: Integrate ESM-2 feature extraction and RP-CGAN data enhancement technology to provide rich and balanced training data for the classifier, and generate samples through Softplus The relative adversarial loss and gradient regularization term (R1 / R2) optimize stability and avoid the vanishing gradient problem of traditional GAN. On mainstream classifiers such as RF, SVM, XGBoost, and LightGBM, the enhanced accuracy, AUC, F1 score, G-mean and other key indicators are significantly improved, verifying the model's predictive stability and accuracy under imbalanced data.

[0061] 3. Ensure that the augmented data is close to the true distribution: A comprehensive evaluation of generated samples is performed through a multi-dimensional similarity metric screening mechanism (Pearson correlation coefficient, Euclidean distance, and autoencoder latent space error). Adaptive sample quantity control (dynamically determining the augmentation scale based on training set performance) is combined to retain pseudo-samples that are highly similar to true samples to avoid distribution drift. This mechanism enhances the credibility of the augmented data and ensures that the generated samples are integrated into the true distribution without introducing noise. This provides reliable input for subsequent classifier training and indirectly supports performance improvement.

[0062] 4. Optimizing training efficiency and model stability: Through lightweight architecture design (such as the generator joint input splicing mechanism) and convergence optimization strategies (Early Stopping, GradScaler mixed-precision training, and gradient regularization), training stability is improved and convergence is accelerated. The model exhibits a more robust parameter update trend during training, and the augmented samples can quickly connect to the classifier training process, demonstrating efficiency advantages in engineering practice.

[0063] 5. Strong stability and cross-domain expansion potential: RP-CGAN has passed SoftplusRelative adversarial loss and gradient regularization alleviate the training instability of traditional GAN. The classifier performance in the embodiment is steadily improved without crashing, which proves the reliability of the generative model. The method is based on a general imbalanced data augmentation framework and can be migrated to other protein modification prediction tasks. It has application potential in cross-domain imbalanced classification scenarios such as financial fraud detection and medical image analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 Flow chart of the method of the present invention.

[0065] Figure 2 This is the structural diagram of RP-CGAN. DETAILED DESCRIPTION

[0066] In order to have a clearer understanding of the technical features, objectives and beneficial effects of the present invention, the technical solution of the present invention is now described in detail below, but it should not be understood as limiting the scope of implementation of the present invention.

[0067] The present invention provides a protein post-translational modification data enhancement method based on generative adversarial networks, the flow chart of which is as follows: Figure 1 As shown, the method includes the following steps:

[0068] Step S1: data preprocessing and feature extraction;

[0069] Input Normalization: FASTA-formatted protein sequence files were read, containing positive samples (labeled 1, indicating modification) and negative samples (labeled 0, indicating unmodified). 21-residue window sequences were constructed, centered around the target residue, with 10 amino acid residues upstream and downstream. Sequences were capitalized, redundancy removed, and low-quality sequences removed. Each sequence was assigned a corresponding label (1 for modification, 0 for unmodified) to provide standardized input for model training. The dataset partitioning is shown in Table 1.

[0070] Table 1 Dataset division

[0071] Number of positive samples Number of negative samples total training set 900 6856 7756 Test set 145 268 413 total 1045 7124 8169

[0072] Sequence Encoding: We use Facebook's open-source ESM-2 pre-trained protein language model (650M parameters) to embed the window sequence. This model, based on the Transformer architecture and pre-trained on millions of protein sequences, is capable of capturing complex contextual relationships and implicit structural information between amino acid residues.

[0073] The specific steps are as follows:

[0074] Load the ESM-2 pre-trained protein language model and alphabet;

[0075] Use the get_batch_converter() method to convert the protein sequence into an input format that the model can recognize;

[0076] Extract the token-level embedding representation of the 33rd layer of the model as the feature of each residue;

[0077] The original dimension embedding (1280 dimensions) is retained for each sequence to form a structure-aware feature matrix;

[0078] The feature matrix of each sequence is averaged to obtain the final feature vector of the sequence.

[0079] Step S2: Sample enhancement of clustering enhanced conditional generative adversarial network (RP-CGAN);

[0080] First, when the number of S-sulfenylation samples is limited and their distribution is highly unbalanced, traditional supervised models find it difficult to construct effective classification boundaries in high-dimensional, nonlinear tasks. Data augmentation is an important means to improve model performance. Data augmentation methods based on generative adversarial networks (GANs) can be used to generate data samples in minority classes.

[0081] Goodfellow et al. proposed the Generative Adversarial Network (GANs) in 2014, which consists of two parts: the generative model G and the discriminative model D. GANs uses the idea of ​​zero-sum game to train the generative model through the adversarial process between the generative model G and the discriminative model D. The probability distribution of the real data x is , the prior distribution of the input noise z is The generative model G is a parameter The multi-layer perceptron can be expressed as a mapping function , which maps the input noise z to the generated data The discriminant model D is also a parameter The multi-layer perceptron, its discrimination result is expressed as The objective function of GANs It can be defined as follows:

[0082] ;

[0083] in, Represents the distribution of real data; represents the distribution of generated data; Indicates that x comes from real data rather than generated data; Indicates expectation. During the training process, the discriminant model D is trained to enhance the discriminant ability of the discriminator to maximize the recognition of whether the label comes from the training sample or the G distribution, that is, to maximize , and train the generative model G so that Minimization can be described as the objective function Even though the generative adversarial network has powerful generation capabilities and high generation quality, it still has well-known shortcomings such as network training difficulties, gradient disappearance, and model instability.

[0084] Conditional Generative Adversarial Networks (CGANs) are an extension of classic generative adversarial networks (GANs) and were first proposed by Mirza and Osindero in 2014. Their core idea is to add prior conditional information c (such as category labels, text descriptions, and image features) to the inputs of the generator G and the discriminator D to control the attributes or categories of generated samples.

[0085] ;

[0086] Compared with the original GANs that can only generate "unconditional" samples, CGAN has the ability to "generate on demand", so it is widely used in tasks such as image generation, image-to-image translation, sequence modeling, and molecular design.

[0087] In 2024, Huang et al. proposed an enhanced relative loss (RpGAN) with a zero-centered gradient penalty objective to improve stability. They mathematically showed that RpGAN with gradient penalty has the same local convergence guarantee as regularized classic GANs, and the regularization method ensures adversarial resistance.

[0088] 2. Based on the traditional CGAN structure, this paper designs a generative adversarial network with classification output and regularization constraints to enhance the number and diversity of positive samples. It is named Cluster Enhanced Conditional Generative Adversarial Network (RP-CGAN). Its structure is as follows: Figure 2 As shown in Figure 2, RP-CGAN consists of a generator G and a discriminator D.

[0089] Generator G:

[0090] Input: random noise vector , the category label you want to construct (1 is positive), the category center vector obtained by clustering category samples ( d is the sample feature dimension).

[0091] To obtain the category center vector , the feature vector set of the real samples of the minority class and the majority class and Perform K-means clustering with the objective function:

[0092] ;

[0093] in, K is the number of clusters; For the clusters; belongs to the kth cluster The sample feature vector of ; is the center of the cluster.

[0094] For a given sample ,in is the sample feature vector, is the true label, and the optimal category center extraction algorithm is shown in Algorithm 1. Suppose a category is clustered into k clusters by KMeans, and its corresponding category center set is , where each center vector , which is consistent with the feature dimension of the original sample.

[0095]

[0096] In the optimal category center extraction algorithm, The number of clusters that minimizes the DBI of the minority class sample clustering results (i.e., the optimal number of clusters for the minority class); The number of clusters that minimizes the DBI of the majority class sample clustering result (that is, the optimal number of clusters for the majority class); best_dbi_min is the current optimal DBI value for the minority class clustering; best_dbi_maj is the current optimal DBI value for the majority class clustering; kmeans_min_final is the final KMeans clustering model that performs best for the minority class samples among all the attempted cluster numbers; kmeans_maj_final is the final KMeans clustering model that performs best for the majority class samples among all the attempted cluster numbers; kmeans_min_final.cluster_centers_ is the set of cluster centers corresponding to the optimal clustering model for the minority class; kmeans_maj_final.cluster_centers_ is the set of cluster centers corresponding to the optimal clustering model for the majority class.

[0097] Eventually, the generator will and Concatenate into a joint input vector and generate samples:

[0098] ;

[0099] The concatenation of the joint input vectors guides the generation of samples that better align with the true distribution, effectively enhancing the generator's ability to model the target category and improving the targetedness and accuracy of sample generation. This design not only alleviates the single-sample distribution issue common in traditional generative adversarial networks, but also accelerates network convergence by introducing prior structural information, improving training stability and generation quality.

[0100] Discriminator D: has dual output and , respectively judging "real / fake" and "predicted category" to assist the generator in generating target category samples in a targeted manner.

[0101] 3. Loss function and training strategy;

[0102] Introduced based on the original generative adversarial network loss:

[0103] Softplus Relative adversarial loss of variants:

[0104] ;

[0105] in, To counter the loss function; is the true sample feature vector; Pseudo samples generated by the generator; From the real sample distribution; The distribution of pseudo samples generated by the generator; Represents the distribution of real samples The following sample Seek expectations; Represents the distribution of generated samples The following sample Seek expectations; Softplus The function is defined as ;

[0106] is the authenticity output of the discriminator.

[0107] Adversarial Loss Adoption Softplus Form replaces traditional and , which can avoid the vanishing gradient problem of the log function and enhance training stability. This loss term simultaneously penalizes both real samples being misclassified as fake and fake samples being misclassified as real, thereby guiding the discriminator to learn a more accurate real-fake boundary.

[0108] Binary Cross Entropy Loss (BCE):

[0109] ;

[0110] in, is the classification loss function; Represents the joint distribution of sample features and labels Seek expectations; is the sample feature vector; is the true category label; For the discriminator to sample The category prediction output value of

[0111] In addition to outputting an adversarial score, the discriminator also outputs a class prediction probability. To this end, this paper introduces a binary cross-entropy loss function to measure the deviation between the discriminator's predicted labels and the true labels. This loss term effectively enhances the ability to utilize conditional information, guiding the discriminator to distinguish between positive and negative sample categories, and also prompting the generator to synthesize samples that are more consistent with the target class distribution.

[0112] Gradient regularization term (R1 / R2):

[0113] R1 regular term (for real samples):

[0114] ;

[0115] R2 regular term (used to generate samples):

[0116] ;

[0117] in, is the gradient regularization term for real samples; is the gradient regularization term for the generated samples; Represents the discriminator output with respect to the real sample input gradient; Represents the discriminator output with respect to the generated sample input gradient; Represents the square of the L2 norm of a vector.

[0118] The present invention introduces a gradient regularization mechanism in GAN training, in which the R1 regularization term imposes a penalty on the gradient of real samples to prevent the discriminator from overfitting the training data; the R2 regularization term is used to limit the gradient fluctuation of generated samples, effectively alleviating training instability and gradient explosion problems.

[0119] Total loss function:

[0120] Discriminator loss function:

[0121] ;

[0122] Generator loss function:

[0123] ;

[0124] in, is the total loss function of the discriminator; is the total loss function of the generator; is the adversarial loss term of the discriminator; is the adversarial loss term of the generator; is the classification loss term of the discriminator; is the classification loss term of the generator; is the weight of classification loss; and are the weights of the R1 regularization term and the R2 regularization term, which take the same value in this task.

[0125] Training control mechanism: Set Early Stopping to prevent overfitting; use mixed precision training through GradScaler to accelerate convergence.

[0126] After RP-CGAN training is completed, the weight parameters of the generator G are saved for subsequent expansion of the training set.

[0127] Step S3: Pseudo-sample screening;

[0128] The number of augmented pseudo samples is determined based on the difficulty of training. The trained RP-CGAN model is then used to generate pseudo samples exceeding the target number. A discriminant-guided and multi-metric-constrained pseudo sample screening algorithm (as shown in Algorithm 2) is then used to select high-quality pseudo samples to enhance the rationality of training data and reduce the risk of distribution drift. The specific process includes: determining the number of pseudo sample augmentations based on classifier performance; constructing a feature space using AutoEncoder; and comprehensively evaluating pseudo sample quality using three metrics: the Pearson correlation coefficient (PCC), Euclidean distance (ED), and minimum mean square error (MSE). Finally, pseudo samples that are most similar to real samples are selected as augmented data to provide reliable input for subsequent classifier training.

[0129]

[0130]

[0131] In the pseudo sample screening algorithm with discriminant guidance and multi-index constraint, TP is the true positive example, that is, the number of positive examples correctly predicted by the model as positive; TN is the true negative example, that is, the number of negative examples correctly predicted by the model as negative; FP is the false positive example, that is, the number of negative examples incorrectly predicted by the model as positive; FN is the false negative example, that is, the number of positive examples incorrectly predicted by the model as negative. is an indicator function, which takes a value of 1 when the condition in the brackets is true, and 0 otherwise; N is the total number of negative samples; P is the total number of positive samples; is the basic enhancement coefficient, is a tiny constant that prevents the denominator from being zero; is the set of Pearson's maximum correlation coefficients of all pseudo samples; For the The maximum value of the Pearson correlation coefficient between the pseudo samples and all real samples; is the set of minimum Euclidean distances between all pseudo samples and real samples; For the The minimum Euclidean distance between a pseudo sample and all real samples; is the minimum error set of feature space for all pseudo samples; For the The minimum mean square error between pseudo samples and all real samples in the latent space representation; For the The standardized score of the Pearson correlation coefficient of the pseudo samples; For the The normalized score of the Euclidean distance of the pseudo samples; For the The standardized score of the MSE of the pseudo sample latent space; For the The comprehensive score of the pseudo samples.

[0132] Step S4: classifier training and evaluation;

[0133] After sample augmentation, the real training data was combined with the filtered pseudo samples and fed into several classic classifiers for supervised learning, including random forest (RF), support vector machine (SVM), XGBoost, and LightGBM. Classifier performance was evaluated using various metrics, including accuracy, F1-score, area under the curve (AUC), Matthews correlation coefficient (MCC), and geometric mean (G-mean). A grid search was performed to automatically optimize the classification threshold within a range of 0.1-0.9 to assess the classifier's ability to identify S-sulfenylation sites.

[0134] Step S5: predict output;

[0135] Use the trained classifier to predict the new protein sequence and output the probability score of each residue being an S-sulfenylation site. The modification site is determined based on the probability score to achieve automatic identification.

[0136] The specific implementation of the present invention is described in detail below with reference to specific embodiments.

[0137] Example 1: To verify the effectiveness of the proposed RP-CGAN and its pseudo-sample screening mechanism in the S-sulfenylation site prediction task, four mainstream classifiers (RF, SVM, XGBoost, and LightGBM) were selected for comparative testing. Performance indicators included accuracy (Acc), sensitivity (Sen), specificity (Spe), precision (Pre), Matthews correlation coefficient (MCC), F1 score, area under the curve (AUC), and G-mean value. Table 2 shows the comparison of classification performance before and after enhancement:

[0138] Table 2 Comparison of classification performance between unenhanced and enhanced

[0139]

[0140] As can be seen from the data in the table, among all classifiers, the performance indicators after enhancement have improved to varying degrees, especially in terms of balance indicators such as Sen, F1 score, MCC and Gmean value, which shows that the generated samples have effectively improved the problem of class imbalance. In addition, RF&SVM has the most obvious improvement. RF's F1 has increased from 0.5312 to 0.5674, and G-mean has increased from 0.5984 to 0.6408, proving that the enhanced data has improved the ability to identify positive samples; XGBoost maintains high Spe while compensating for low Sen: Although the Sen value is still lower than other models, both F1 and AUC have increased, indicating that pseudo samples have an effect in improving recall rate; LightGBM is the most balanced, maintaining a high overall performance before and after enhancement, and is currently the classifier option with the best comprehensive indicators.

[0141] The above are only preferred embodiments of the present invention. It should be pointed out that for those skilled in the art, several variations and improvements can be made without departing from the concept of the present invention. These should also be regarded as the scope of protection of the present invention. These will not affect the effect of the implementation of the present invention and the practicality of the patent.

Claims

1. A protein post-translational modification data enhancement method based on generative adversarial networks, characterized in that: The following steps are involved: Step S1: data preprocessing and feature extraction; Read the protein sequence file, divide it into training set and test set, use the ESM-2 pre-trained protein language model for feature encoding, and generate standardized feature vectors; Step S2: Sample enhancement of clustering-enhanced conditional generative adversarial network; A RP-CGAN model consisting of a generator and a discriminator was constructed. The generator input consists of a random noise vector, a target category label, and a category center vector generated by K-means clustering, and outputs a pseudo sample of the target category. The discriminator adopts a dual-output structure to judge the authenticity of the sample and the category label respectively, and combines Softplus relative adversarial loss, binary cross entropy loss, and gradient regularization to optimize training stability. Step S3: Pseudo-sample screening; The trained RP-CGAN model is used to generate pseudo samples exceeding the target number. Through a multi-metric constrained screening mechanism, pseudo samples that are closest to the distribution of real samples are selected. Step S4: classifier training and evaluation; Merge the real training data with the filtered pseudo samples and input them into the classifier for training to optimize the classification performance; Step S5: predict output; Use the trained classifier to predict the new protein sequence and output the modification site probability score; The generator design of the RP-CGAN model includes: The input conditions are random noise vector z~N(0,1) and category label y c ∈{0,1+ and the category center vector obtained by K-means clustering Generate samples close to the real distribution through joint input splicing mechanism The step of generating the category center vector includes: Perform K-means clustering on the minority and majority class samples respectively, and determine the optimal number of clusters using the DBI indicator; extract the optimal cluster centers of the minority and majority classes as the prior condition input of the generator; The pseudo sample screening step comprises: Calculate sensitivity and specificity through confusion matrix and dynamically determine the number of pseudo sample enhancements n * ; Use the generator to generate m×n * pseudo samples; use AutoEncoder to construct feature space, calculate the Pearson correlation coefficient, Euclidean distance and minimum mean square error of three types of distance indicators between pseudo samples and real samples, normalize and weightedly fuse the three types of distance indicators to obtain a comprehensive score; select the top n samples with the lowest comprehensive scores * pseudo samples, forming the final enhanced sample set 2. The protein post-translational modification data enhancement method based on generative adversarial network according to claim 1, characterized in that The data preprocessing and feature extraction steps include: With the target residue as the center, 10 amino acid residues upstream and downstream were extracted to form a 21-residue window sequence; Load the ESM-2 pre-trained protein language model and alphabet; Use the get_batch_converter() method to convert the protein sequence into an input format that the model can recognize; Extract the token-level embedding representation of the 33rd layer of the model as the feature of each residue; The original dimension embedding is retained for each sequence to form a structure-aware feature matrix; The feature matrix of each sequence is averaged to obtain the final feature vector of the sequence.

3. The protein post-translational modification data enhancement method based on generative adversarial network according to claim 1, characterized in that The loss function of the RP-CGAN model includes: Relative adversarial loss of Softplus variants: in, To counter the loss function; Represents the distribution of real samples p real Find the expectation of the sample x under; Represents the generated sample distribution p G The following sample Find the expectation; x is the true sample feature vector; is a pseudo sample generated by the generator; the Softplus function is defined as softplus(x) = log(1+e x );D real (·) is the authenticity output of the discriminator; Binary cross entropy loss: in, is the classification loss function; E (x,y) It means to find the expectation of the joint distribution (x,y) of sample features and labels; x is the sample feature vector; y∈{0,1+ is the true category label; D cls (x) is the class prediction output value of the discriminator for sample x; Gradient regularization term: R1 regularization term: R2 regularization term: Among them, R1 is the gradient regularization term for real samples; R2 is the gradient regularization term for generated samples; Represents the gradient of the discriminator output with respect to the true sample input x; Represents the discriminator output with respect to the generated sample input The gradient of ||·|| 2 Represents the square of the L2 norm of the vector; Total loss function: Discriminator loss function: Generator loss function: in, is the total loss function of the discriminator; is the total loss function of the generator; is the adversarial loss term of the discriminator; is the adversarial loss term of the generator; is the classification loss term of the discriminator; is the classification loss term of the generator; cls is the weight of classification loss; R1 and λ R2 are the weights of the R1 regularization term and the R2 regularization term, respectively.

4. The protein post-translational modification data enhancement method based on generative adversarial network according to claim 1, characterized in that The classifier training and evaluation steps include: Random forest, support vector machine, XGBoost, and LightGBM were used for training; classification performance was evaluated using accuracy, F1 score, AUC, MCC, and G-mean value.

5. The protein post-translational modification data enhancement method based on generative adversarial network according to claim 1, characterized in that The prediction output step includes: Features are extracted from the new protein sequence and input into the trained classifier; the probability score of each residue being a modification site is output, and the modification site is determined based on the threshold.

Citation Information

Patent Citations

  • Cutter wear condition detection method based on improved conditional generative adversarial network

    CN114346761A

  • Power system load prediction method based on conditional guidance diffusion process

    CN119338283A