Metabonomics data batch effect correction method based on deep learning
Through a two-stage correction method, combined with SERRF and generative adversarial network, the problem of batch effect in metabolomics is solved, data quality and reproducibility are improved, and biological information is retained.
Patent Information
- Application Number
- CN202510428984.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-04
AI Technical Summary
In metabolomics research, it is difficult to effectively eliminate batch effects, especially batch effects, in mass spectrometry data acquisition process, which leads to data consistency and accuracy problems, and existing methods have problems such as overfitting, signal loss or overcorrection.
A two-stage correction method is adopted: firstly, the in-batch effect is corrected by the SERRF method, and then the interbatch effect is corrected by a joint framework of the generative adversarial network and the autoencoder, combining the centroid determination and mutual nearest neighbor alignment loss to ensure data consistency and avoid loss of biological information.
It significantly improves the quality and reproducibility of metabolomics data, effectively eliminates batch effects, retains biological differences, and provides reliable data support.
Smart Images

Figure CN120260672A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of metabolomics data processing. Batch effect correction can significantly improve the quality and reproducibility of metabolomics data by eliminating systematic biases. Aiming at the complexity of data drift between batches in large-scale metabolomics data, the present invention determines the pairwise correction order of multi-batch data through centroids, and constructs a joint correction framework by combining random forest and deep adversarial alignment network, eliminating data errors while retaining true metabolic differences. Background Art
[0002] Metabolomics is the science of studying small molecule metabolites in living organisms, covering qualitative and quantitative analysis of various metabolites such as amino acids, fatty acids, and sugars. Through high-throughput analysis techniques, metabolomics can systematically reveal the metabolic characteristics of organisms under different physiological or pathological states, providing important basis for understanding disease mechanisms, identifying biomarkers, promoting early disease diagnosis and personalized treatment, etc. In addition, metabolomics can also promote the development of drug discovery and precision medicine by exploring the interactions between intracellular metabolites.
[0003] In metabolomics research, nuclear magnetic resonance and chromatography-mass spectrometry are two main analysis platforms. Compared with nuclear magnetic resonance technology, mass spectrometry combined with different ionization techniques can significantly improve the separation and identification ability of metabolites in complex biological samples, thus improving the detection throughput and sensitivity. However, in large-scale clinical cohort studies, systematic errors caused by factors such as equipment operation, experimental operation, and environmental conditions during the mass spectrometry data acquisition process are still an important challenge in metabolomics research.
[0004] Batch effect refers to the systematic bias in the measurement signals of metabolite ion peaks due to different experimental batches, usually manifested as signal drift or intensity change, affecting the consistency and accuracy of data. Batch effects can be divided into within-batch effects and between-batch effects. The within-batch effects are mainly measurement errors caused by fluctuations in the instrument operating state and accumulation of pollutants during the measurement of a batch of samples. The between-batch effects are usually caused by factors such as column replacement, ion source cleaning, differences in the operation of different experimental personnel, or changes in environmental conditions (temperature, humidity, pressure). Batch effects may lead to data bias, especially in studies with uneven distribution of case / control samples, which may mask true biological differences. Therefore, effective correction of batch effects is crucial for improving data quality and ensuring the reliability of omics research.
[0005] Currently, batch effect correction methods are mainly divided into three categories: internal standard-based methods, quality control (QC) sample-based methods, and data-driven methods. The internal standard-based method adds isotope-labeled internal standards before mass spectrometry experiments, and captures the changes in sample analysis through the internal standards, so as to eliminate the systematic bias caused by non-biological factors. However, the internal standard-based method has limitations. The addition of internal standards requires additional experimental steps and costs. Among all detected metabolites, only a part of metabolites with similar chemical properties to the internal standards can be well corrected.
[0006] The quality control (QC) sample-based method inserts multiple QC samples into the analysis sequence of research samples, and corrects the batch effect according to the QC intensity of different batches. The QC sample-based method can usually correct complex batch effects. However, due to the limited number of QC samples, the QC sample-based method may have overfitting problems when training the model, thus reducing the generalization ability of the model in actual samples.
[0007] The data-driven method directly uses research samples for correction without relying on internal standards or QC samples. Since the data-driven method does not require any additional experimental elements, such as internal standards or quality control samples, they help to reduce the overall research cost and complexity. However, it has problems such as strong dependence on the quality of the dataset, sensitivity to noisy data, easy to cause signal loss or overcorrection, and difficulty in dealing with complex batch effects.
[0008] Deep learning has also been applied to batch effect correction in recent years. The deep learning-based batch effect correction method has powerful nonlinear modeling capabilities, can automatically extract features and adapt to different datasets, and has remarkable effects in processing high-dimensional data. But at the same time, there may be a risk of overcorrection, that is, while over-eliminating the differences between batches, it may also cover up the true biological information. In addition, the complexity of deep learning models and the sensitivity of hyperparameters also increase the difficulty of model training and adjustment. Summary of the Invention
[0009] The present invention proposes a batch effect correction method based on deep learning, which is applicable to metabolomics datasets containing quality control samples (QC). The metabolomics data batch effect correction method of the present invention includes two stages: within-batch correction and between-batch correction. In the within-batch correction stage, the SERRF method is used to model the injection order of QC samples, identify and correct systematic drift effects, and eliminate data biases caused by factors such as injection order within the same batch. In the between-batch correction stage, first, the centroids of each batch are calculated, and the pairwise correction order is determined based on the distances between batches; based on the joint framework of the generative adversarial network and the autoencoder, adversarial training is used to reduce systematic biases between batches, and at the same time, the mutual nearest neighbor (MNN) alignment loss is introduced to ensure the distribution consistency of samples from different batches in the latent space and avoid biological information loss caused by overcorrection.
[0010] The technical solution adopted by the present invention is as follows:
[0011] A metabolomics data batch effect correction method based on deep learning, comprising the following steps:
[0012] Step 1: Data cleaning and missing value processing
[0013] First, clean the metabolomics data. When the missing ratio of a certain feature in all batch samples exceeds 20%, then delete this feature; when the missing ratio is lower than 20%, then fill it with half of the minimum non-zero value of this feature in all batch samples.
[0014] Step 2: Correct the within-batch effect using the SERRF method
[0015] Assume that the dataset contains N batches, B = {B1, B2,..., B N}, batch B t (1 ≤ t ≤ N) contains n t + n qt samples, and the sample data is where is the QC sample of batch B t ; the present invention first corrects the within-batch data drift for the data of each batch B t .
[0016] 1) Standardize all samples of batch B t :
[0017]
[0018] where, represents the intensity of sample i at metabolite j, represents the mean of metabolite j in batch B t , represents batch Bt Standard deviation of metabolite j
[0019] 2) For each metabolite j (j = 1, 2, …, m), using the QC samples within batch B t where the metabolite is the response variable, and using the injection order l t and the values of other metabolites in the QC samples except this metabolite as predictor variables to train a random forest model Calculate the predicted value of the systematic error of metabolite j within batch B t :
[0020]
[0021] where is the predicted value of the systematic error of the j-th metabolite in batch B t , and is the random forest model for fitting the systematic error
[0022] 3) Use the systematic error to correct the sample data of batch B t :
[0023]
[0024] where represents the intensity of the j-th metabolite of the i-th sample in batch B t after removing the within-batch effect, and represents the mean value of metabolite j in all samples in batch B t
[0025] Step Three: Calculate the distance between batches and determine the pair of batches with the closest distribution
[0026] If the number of batches N in the current dataset > 1, that is, there are multiple batches in the dataset, then select the two batches with the closest distribution from the current N batches for between-batch correction in the following steps; if N = 1, then the correction ends
[0027] 1) Calculate the centroid of each batch in the current N batches of data. The centroid is defined as:
[0028]
[0029] where represents the metabolite intensity vector of sample i after correction in Step Two
[0030] 2) Based on the batch centroids, calculate the Euclidean distance between any two batches to determine the two batches with the closest distance (B a , B b ), 1 ≤ a, b ≤ N, a ≠ b.
[0031] Step Four: Calculate batch B a , B b The mutual nearest neighbor pairs of samples between
[0032] When performing between-batch correction, to ensure the distribution consistency of samples from different batches in the latent space, the present invention adopts the mutual nearest neighbor (MNN) matching method.
[0033] 1) For any and the Euclidean distance is used to measure the similarity between them. For the samples in batch B a determine the K samples in batch B that are closest to it in Euclidean distance, and form a set b For each sample in batch B b determine the K samples in batch B that are closest to it in Euclidean distance, and form a set a
[0034] 2) Determine the set of mutual nearest neighbor sample pairs between batches B a , B b . If and are located in each other's K-nearest neighbor sample sets, then the sample pair is a mutual nearest neighbor pair. The set of mutual nearest neighbor samples between batches B a , B b is defined as:
[0035] and
[0036] Step Five: Build an autoencoder to eliminate the batch effect between B a , B b
[0037] Use a framework based on an autoencoder combined with an adversarial training and mutual nearest neighbor batch effect correction model. Extract the embedded representation of the samples through the encoder, and optimize the model using multiple loss functions to perform between-batch correction on the selected batch pair (B a , B b ). The loss functions include reconstruction loss, mutual nearest neighbor alignment loss, and discriminator loss.
[0038] 1) Train the autoencoder and use it to extract the embedded representation of the sample data. The autoencoder consists of an encoder E(·) and a decoder D(·). The encoder compresses the sample data into the latent space, while the decoder reconstructs the sample data from the latent representation. The goal of training the autoencoder is to minimize the reconstruction loss, that is, to minimize the difference between the input data and the decoded data. Batch B a and B b both contain normal samples and QC samples, and the total number of samples is respectively: N a = n a + n qa , N b = n b + n qb ; The reconstruction loss is defined as:
[0039]
[0040] 2) Calculate the mutual nearest neighbor alignment loss. Use the mutual nearest neighbor pairs in step four to calculate the nearest neighbor alignment loss and optimize the encoder. This loss is used to force the samples in different batches to align in the latent space and reduce the distribution difference between batches.
[0041]
[0042] 3) Train the batch discriminator. In the present invention, the batch discriminator is trained to measure the batch information in the sample embedded representation through the cross-entropy loss, so as to help the encoder learn to remove the batch effect and make the data distribution more consistent between different batches. The discriminator loss is:
[0043]
[0044] where, and respectively represent the predicted probabilities of the discriminator for the samples in batch B a and B b . Use the cross-entropy loss to optimize the discriminator so that it can predict the batch labels of the samples as accurately as possible.
[0045] 4) Adversarial training. The final training goal is achieved by optimizing the game between the encoder and the discriminator through adversarial training. The goals of the encoder E and the decoder D are to minimize the reconstruction loss and maximize the misclassification of the discriminator at the same time. The total optimization function is:
[0046]
[0047] where, λ b and λ MNN are hyperparameters. After the training is completed, the best model will be saved for inter-batch correction.
[0048] Step Six:
[0049] 1) Use the trained decoder to generate corrected sample data from the optimized latent representation, and merge the corrected B a ,B b data into the same batch, denoted as batch B a ; Delete batch B from the batch set b , and update the number of batches N = N - 1.
[0050] 2) Repeat the process from step three to step five above until all data is normalized to the same batch and the correction process terminates.
[0051] The batch correction method of the present invention based on the centroid-determined paired correction order of multi-batch data. At the same time, combining the advantages of random forest and deep adversarial alignment network, a two-stage joint correction framework is constructed. While retaining the true metabolic differences of biological samples, the errors in the data are eliminated, providing reliable data support for metabolomics research. Description of the Drawings
[0052] Figure 1 It is a diagram of a metabolomics batch effect correction model based on deep learning. Detailed Embodiment
[0053] The following uses 7 datasets (see Table 1) to illustrate the specific implementation of the technical solution of the present invention and conduct a comparative analysis with existing batch effect removal methods.
[0054] Table 1: Statistical information of 7 datasets
[0055]
[0056] The implementation environment of the present invention is the Pytorch platform, and the model training includes the following steps:
[0057] 1) Data preprocessing: When the proportion of missing values of a certain feature exceeds 20%, delete this feature; otherwise, fill the missing values with half of the smallest non-zero value of this feature in the samples.
[0058] 2) Intra-batch effect removal: Use the QC samples of each batch to establish a random forest model to estimate the systematic error within the batch, and use the predicted systematic error to achieve intra-batch effect removal.
[0059] 3) Removal of inter-batch variation: The training stage is divided into three stages: pre-training of the autoencoder, pre-training of the discriminator, and adversarial training. The number of epochs for the three stages is set to 1000, 10, and 700 respectively. When the sum of the reconstruction error and the weighted embedding distance of the quality control (QC) samples is minimized, the current model is recorded as the best state and updated; when the weighted comprehensive score does not decrease for 5 consecutive epochs, the optimization is stopped and the model state at the best weighted comprehensive score is rolled back.
[0060] 4) Parameter settings: The dimension of the input layer of the autoencoder is set to 1000, the dimension of the bottleneck layer is set to 500, the dimension of the output layer is set to 1000, the input dimension of the discriminator is set to 500, the learning rate of the autoencoder is set to 0.0002, the learning rate of the discriminator is set to 0.005, and the Adam optimization algorithm is used to optimize the model.
[0061] 5) Model evaluation: To evaluate the data quality after removing batch effects, calculate the proportion of features with a relative standard deviation (RSD) less than 0.3 among all features on the QC samples. To evaluate the retention of biological variation in the data after removing batch effects, we constructed a support vector machine based on the reconstructed data, performed five-fold cross-validation, and compared it with other baseline methods. The performance of the classifier is measured by the average classification accuracy, which reflects the overall prediction ability of the model.
[0062] Table 2 shows the comparison results of the classification accuracy between the present invention and 5 baseline algorithms on 7 datasets, and Table 3 shows the comparison results of the proportion of features with RSD < 0.3 on the corrected QC samples of the present invention on 7 datasets. Among them, the comparison methods include Combat, Harmony, QC-RFSC, SERRF, and NormAE.
[0063] The experimental results show that the present invention has better performance than other baseline algorithms on multiple datasets, demonstrating the effectiveness of this method in removing batch effects and standardizing metabolomics data.
[0064] Table 2: Performance comparison with baseline methods on 7 datasets (accuracy)
[0065] Dataset number The present invention NormAE Combat Harmony SERRF QC-RFSC 1 0.9669 0.9185 0.9593 0.9565 0.9607 0.8993 2 0.835 0.7728 0.7675 0.7507 0.7324 0.7689 3 0.8868 0.8696 0.8727 0.8738 0.8768 0.8704 4 0.8867 0.7901 0.8947 0.8864 0.8859 0.8792 5 0.6826 0.6015 0.8242 0.9049 0.8067 0.8342 6 0.875 0.8037 0.8604 0.8732 0.8766 0.8694 7 1 1 1 1 1 1 W / T / L 0 / 1 / 6 2 / 1 / 4 1 / 1 / 5 2 / 1 / 4 1 / 1 / 5
[0066] From the experimental results in Table 2 and Table 3, it can be concluded that the present invention has excellent performance compared with the 5 baseline algorithms in comparison. Specifically, the present invention achieved the highest proportion of features with RSD < 0.3 on 6 of the datasets, indicating that the present invention can better remove batch effects; and achieved the highest accuracy on 4 of the datasets, indicating that the present invention can retain biological variation while removing batch effects.
[0067] Table 3: Proportion of features with RSD < 0.3 after calibration
[0068] Dataset The present invention NormAE Combat Harmony SERRF QC-RFSC 1 89.29% 81.75% 14.28% 9.13% 78.17% 11.90% 2 100% 50.94% 28.30% 15.09% 94.34% 81.13% 3 100% 98.31% 83.09% 81.39% 100% 100% 4 91.59% 91.36% 34.64% 33.06% 80.21% 85.08% 5 88.1% 99.16% 88.95% 87.65% 95.43% 93.37% 6 99.94% 95.25% 39.99% 40.58% 91.79% 93.34% 7 95.92% 83.67% 75.51% 75.51% 85.71% 83.67% W / T / L 1 / 0 / 6 1 / 0 / 6 0 / 0 / 7 1 / 1 / 5 1 / 1 / 5
Claims
1. A batch effect correction method for metabolomics data based on deep learning, Its features include the following steps: Step 1: Data cleaning and missing value handling First, clean the metabolomics data. When the missing ratio of a certain feature in all batch samples exceeds 20%, delete the feature; When the missing ratio is lower than 20%, fill it with half of the minimum non-zero value of the feature in all batch samples; Step 2: Use the SERRF method to correct the within-batch effect Assume that the dataset contains N batches, B = {B1, B2, …, B N}, and batch B t (1 ≤ t ≤ N) contains n t + n qt samples, and the sample data is wherein is the QC sample of batch B t ; the present invention first corrects the data drift within each batch B t for the data of the batch 1) Standardize all samples of batch B t : Among them, represents the intensity of sample i at metabolite j, represents batch B t the mean of metabolite j in it, represents batch B t the standard deviation of metabolite j in it; 2) For each metabolite j (j = 1, 2, …, m), using the QC samples within batch B t as the response variable for that metabolite, and using the injection order l t and the values of other metabolites in the QC samples except that metabolite as the predictor variables to train a random forest model Calculate the predicted value of the systematic error of metabolite j within batch B t : Among them, is the predicted value of the systematic error of the j-th metabolite in batch B, t and is a random forest model for fitting the systematic error. 3) Use systematic error to calibrate the sample data of batch B t : Among them, represents the intensity of the j-th metabolite of the i-th sample in batch B after removing the within-batch effect; t represents the mean value of metabolite j of all samples in batch B; t Step 3: Calculate the distance between batches and determine the pair of batches with the closest distribution If the number of batches N in the current dataset > 1, that is, the dataset has multiple batches, select the 2 batches with the closest distribution from the current N batches for between-batch correction in the following steps; if N = 1, the correction ends; 1) Calculate the centroid of each batch in the data of the current N batches. The centroid is defined as: Among them, represents the metabolite intensity vector of sample i after calibration in step two; 2) Calculate the Euclidean distance between any two batches based on the batch centroid, and determine the two batches with the closest distance (B a , B b ), where 1 ≤ a, b ≤ N and a ≠ b; Step 4: Calculate batch B a , B b mutual nearest neighbor pairs between samples When performing between-batch correction, to ensure the consistency of the distribution of samples in the latent space, the mutual nearest neighbor (MNN) matching method is adopted; 1) For any and the Euclidean distance is used to measure the similarity between them; for batch B a each sample in determine the K samples with the closest Euclidean distance to it in batch B b to form a set For each sample in batch B b each sample in determine the K samples with the closest Euclidean distance to it in batch B a to form a set 2) Determine batch B a , B b among the nearest neighbor sample pairs to each other; if and are located in each other's K-nearest neighbor sample sets, then the sample pair is a mutual nearest neighbor pair; the mutual nearest neighbor sample set between batches B a , B b is defined as: Step 5: Establish an autoencoder to eliminate the batch effect between B a ,B b ; Use an autoencoder-based framework combined with an adversarial training and mutual nearest neighbor batch effect correction model to extract the embedded representation of samples through the encoder, and optimize the model using multiple loss functions to perform inter-batch correction on the selected batch pairs (B a , B b ). The loss functions include reconstruction loss, mutual nearest neighbor alignment loss, and discriminator loss; 1) Train an autoencoder to extract the embedded representation of the sample data using the autoencoder. The autoencoder consists of an encoder E(·) and a decoder D(·). The encoder compresses the sample data into the latent space, while the decoder reconstructs the sample data from the latent representation. The objective of training the autoencoder is to minimize the reconstruction loss, that is, to minimize the difference between the input data and the decoded data, batch B a and B b both contain normal samples and QC samples, and their total number of samples are respectively: N a = n a + n qa , N b = n b + n qb ; The reconstruction loss is defined as: 2) Calculate the mutual nearest neighbor alignment loss. Use the mutual nearest neighbor pairs in Step 4 to calculate the nearest neighbor alignment loss and optimize the encoder. This loss is used to force the samples of different batches to align in the latent space and reduce the distribution difference between batches; 3) Train the batch discriminator. In the present invention, the batch discriminator is trained. The batch information in the sample embedding representation is measured by the cross-entropy loss, so as to help the encoder learn to remove the batch effect and make the distribution of data more consistent between different batches. The discriminator loss is: Among them, and respectively represent the predicted probabilities of the discriminator for batches B a and B b samples. The cross-entropy loss is used to optimize the discriminator so that it can predict the batch labels of the samples as accurately as possible; 4) Adversarial training. The final training goal is achieved by optimizing the game between the encoder and the discriminator through adversarial training. The goal of the encoder E and the decoder D is to minimize the reconstruction loss while maximizing the misclassification of the discriminator. The total optimization function is: Among them, λ b and λ MNN are hyperparameters. After training is completed, the best model will be saved for between-batch correction; Step 6: 1) Use the trained decoder to generate corrected sample data from the optimized latent representation, and merge the corrected data of B a ,B b into the same batch, denoted as batch B a ; Delete batch B from the batch set b , and update the number of batches N = N - 1; 2) Repeat the process of the aforementioned Steps 3 to 5 until all data are normalized to the same batch, and the correction process terminates.
Citation Information
Cited By
Metabonomics data batch effect correction method based on generative adversarial network
CN121281627A
A method for batch effect correction of metabolomics data based on generative adversarial network
CN121281627B