Method for single-cell metabolomics analysis based on mass spectrometry
By using the SCSR-Unet super-resolution model and the hybrid loss function CL, the contradiction between resolution and sensitivity in high-speed scanning mode of linear ion trap mass spectrometers was resolved, achieving mass spectrometry resolution improvement without hardware upgrades. This significantly improved the accuracy of single-cell metabolomics analysis and the revelation of heterogeneous characteristics of bladder cancer cells.
Patent Information
- Application Number
- CN202510864885.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-06-26
AI Technical Summary
Existing linear ion trap mass spectrometers suffer from a contradiction between resolution and sensitivity in high-speed scanning mode, resulting in the inability to separate isotope peaks and adjacent mass-to-charge ratio ions. This affects the accuracy of single-cell metabolite identification and cell subtype classification. Furthermore, existing deep learning algorithms have not been optimized for the low abundance and high noise characteristics of single-cell metabolic data and lack systematic verification of the heterogeneity of tumor cells such as bladder cancer.
The SCSR-Unet super-resolution model, combined with the hybrid loss function CL, is used to reconstruct high-resolution data from mass spectrometry data in low-resolution mode, thereby improving the accuracy and resolution of signal reconstruction. Deep learning modeling can overcome the resolution bottleneck without hardware upgrades. Single-cell metabolomics analysis is performed by combining data preprocessing and machine learning methods.
It significantly improved mass spectrometry resolution and single-cell typing accuracy, fully demonstrated the heterogeneity of single cells in bladder cancer, and enhanced the effective features and classification accuracy of single-cell metabolomics analysis.
Smart Images

Figure CN120374399B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to mass spectrometry detection technology, and particularly to a single-cell metabolomics analysis method based on mass spectrometry technology. Background Technology
[0002] Single-cell metabolomics provides a crucial tool for revealing tumor heterogeneity and drug response mechanisms by analyzing the metabolic characteristics of individual cells. Mass spectrometry, with its high sensitivity and wide dynamic range, has become a core tool for single-cell metabolic analysis. However, existing linear ion trap mass spectrometers (such as LTQ XL) face a trade-off between resolution and sensitivity in high-speed scanning modes (such as Turbo mode). Increasing the scanning speed reduces resolution, leading to the inability to separate isotope peaks and adjacent mass-to-charge ratio ions, severely impacting the accuracy of metabolite identification and cell subtype classification.
[0003] Traditional solutions often rely on hardware iterations (such as orbital trap mass spectrometry) or external condition optimization, but these suffer from high equipment costs and low throughput. While deep learning-based super-resolution algorithms exist to improve mass spectrometry resolution, they are not optimized for the low abundance and high noise characteristics of single-cell metabolic data and lack systematic validation for the heterogeneity of tumor cells such as bladder cancer. Furthermore, existing single-cell metabolic data analysis workflows are fragmented, traditional dimensionality reduction methods (such as PCA) are prone to losing key biological information, and machine learning models lack standardized procedures. Summary of the Invention
[0004] To address the shortcomings of the existing technical solutions, this invention provides a single-cell metabolomics analysis method based on mass spectrometry.
[0005] The objective of this invention is achieved through the following technical solution:
[0006] A single-cell metabolomics analysis method based on mass spectrometry includes the following steps:
[0007] (A1) Collect population cell data, including multiple average mass spectra of various population cells in low-resolution and high-resolution modes;
[0008] (A2) Construct the SCSR-Unet super-resolution model using the population cell data, and train, validate and test the model;
[0009] In network construction, the loss function CL used is:
[0010] ;
[0011] output i It is the prediction result of the sample, label iHere, α is the label corresponding to the sample, i represents the i-th sample, α is the hyperparameter controlling the loss weight, and N is the number of samples;
[0012] (A3) Acquire single-cell data, including average mass spectrometry data of multiple single cells in low-resolution and high-resolution modes;
[0013] (A4) The average mass spectrometry data in low resolution mode is passed through the SCSR-Unet model to obtain high-resolution mass spectrometry data with predicted absolute intensity. The average mass spectrometry data in high resolution mode is preprocessed simultaneously and used for single-cell metabolomics analysis.
[0014] Compared with the prior art, the present invention has the following beneficial effects.
[0015] 1. The hybrid loss function used in modeling improves the accuracy of signal reconstruction;
[0016] A hybrid loss function was designed to optimize the numerical error of the signal while constraining the overall directional consistency of the spectrum, thus solving the problem of mutual limitation between mass spectrometer resolution and sensitivity. The optimization of parameter α significantly improved the ability to distinguish overlapping peaks.
[0017] 2. Improved mass spectrometry resolution can be achieved without upgrading hardware technology;
[0018] The SCSR-Unet algorithm uses deep learning modeling to reconstruct high-resolution mass spectrometry data from high-speed scanning in low-resolution mode into predicted high-resolution data, which is equivalent to the hardware level of low-speed scanning in high-resolution mode. It can overcome the resolution bottleneck of ion trap mass spectrometry without replacing high-cost hardware such as orbital traps.
[0019] 3. The accuracy of single-cell typing has been significantly improved;
[0020] The improved resolution of mass spectrometry data increases the number of effective features in single-cell metabolomics, which in turn significantly improves the accuracy of subtyping and fully reveals the heterogeneity of single bladder cancer cells. Attached Figure Description
[0021] The disclosure of this invention will become more readily understood with reference to the accompanying drawings. It will be readily understood by those skilled in the art that these drawings are merely illustrative of the technical solutions of this invention and are not intended to limit the scope of protection of this invention. In the drawings:
[0022] Figure 1 This is a flowchart illustrating the analytical method of the present invention;
[0023] Figure 2 The Turbo absolute intensity mass spectra of UMUC3 bladder cancer cell population before and after cubic spline interpolation in the range of 100Da-101Da;
[0024] Figure 3 These are the loss curves and accuracy curves for the training and validation sets.
[0025] Figure 4 These are absolute intensity mass spectra of UMUC3 bladder cancer cell population data at 132, 136, and 148 Da under Turbo, Normal, and SR conditions;
[0026] Figure 5 This is a relative intensity mass spectrum of UMUC3 bladder cancer cell population data at 132, 136, and 148 Da under normal and SR conditions;
[0027] Figure 6 These are the t-SNE and UMAP clustering results for Normal, SR, and Turbo scenarios;
[0028] Figure 7 These are the RF confusion matrix and ROC curves for Normal, SR, and Turbo scenarios;
[0029] Figure 8 These are heatmaps and primary and secondary arginine comparison spectra under Normal, SR, and Turbo conditions. Detailed Implementation
[0030] Figures 1-8 The following description illustrates optional embodiments of the invention to teach those skilled in the art how to implement and reproduce the invention. Some conventional aspects have been simplified or omitted to teach the technical solutions of the invention. Those skilled in the art should understand that variations or substitutions derived from these embodiments will be within the scope of the invention. Those skilled in the art should understand that the following features can be combined in various ways to form multiple variations of the invention. Therefore, the invention is not limited to the optional embodiments described below, but is defined only by the claims and their equivalents. Example 1
[0031] This embodiment uses a single-cell metabolomics analysis method based on mass spectrometry technology, such as... Figure 1 As shown, it includes the following steps:
[0032] (A1) Collect population cell data, including multiple average mass spectra of various population cells in low-resolution and high-resolution modes;
[0033] (A2) Construct the SCSR-Unet super-resolution model using the population cell data, and train, validate and test the model;
[0034] In network construction, the loss function CL used is:
[0035] ;
[0036] output i It is the prediction result of the sample, label i Here, α is the label corresponding to the sample, i represents the i-th sample, α is the hyperparameter controlling the loss weight, and N is the number of samples;
[0037] (A3) Acquire single-cell data, including average mass spectrometry data of multiple single cells in low-resolution and high-resolution modes;
[0038] (A4) The average mass spectrometry data in low resolution mode is passed through the SCSR-Unet model to obtain high-resolution mass spectrometry data with predicted absolute intensity. The average mass spectrometry data in high resolution mode is preprocessed simultaneously and used for single-cell metabolomics analysis.
[0039] To improve the accuracy of predictions, further, in step (A2), the model is constructed by including the following steps:
[0040] (A21) Data alignment: The average mass spectrometry data in the low-resolution mode is interpolated to obtain the same equidistant mass-to-charge ratio data as the average mass spectrometry data in the high-resolution mode.
[0041] (A22) Data encoding: Multiple segments of equal length are extracted from the average mass spectrometry data in low-resolution and high-resolution modes. The segments are arranged in order of mass-to-charge ratio into a two-dimensional matrix. After normalization, training pairs are generated. The low-resolution segments are used as input data, and the high-resolution segments are used as target labels.
[0042] (A23) Network Construction. A 2D U-Net network structure is used, where both input and output are 32×32 two-dimensional matrices, representing low-resolution data and predicted high-resolution data, respectively. The 2D U-Net structure consists of three downsampling layers and three upsampling layers. Each downsampling and upsampling layer contains two two-dimensional convolutional layers to extract features from the input data. Furthermore, skip connections are added between the downsampling and upsampling layers to directly pass the low-level features extracted during downsampling to the corresponding upsampling layer, effectively preserving detail and improving reconstruction results. In SR-Unet, mean squared error (MSE) is used as the loss function. However, considering that MSE is a loss function based on the difference of the original numerical values, primarily focusing on the squared difference between each data point, it often overemphasizes absolute error during training, easily neglecting the relative direction between data points. Cosine similarity (CS), on the other hand, focuses on the angular difference between two vectors, is unaffected by the absolute size of the data, and is more suitable for measuring the relative relationship between samples. For mass spectrometry data, comparisons of different mass-to-charge ratios (m / z) or intensities are often involved. Cosine similarity can better capture the directionality of the data, helping the network learn the global structure and patterns of the data. Therefore, to obtain better training results, the loss function is set as a combination of MSE and CS loss functions, with the loss function CL being: .
[0043] output i It is the prediction result of the sample, label i α is the label corresponding to the sample, i represents the i-th sample, α is the hyperparameter controlling the loss weight, and N is the number of samples.
[0044] The calculation uses cosine similarity loss. , α is a hyperparameter that controls the two loss weights, and its value ranges from 0 to 1. By adjusting α, the influence between the directionality (measured by CS) and amplitude difference (measured by MSE) between the prediction vectors can be balanced. This setting of the loss function can optimize the amplitude accuracy of the mass spectrometry signal while ensuring the consistency of the spectral shape, thereby improving the final super-resolution effect.
[0045] (A24) Input the average mass spectrometry data of the low-resolution mode to be tested into the trained SCSR-Unet super-resolution model to obtain the predicted high-resolution mass spectrometry data.
[0046] (A25) Recover the predicted high-resolution mass spectrometry data according to the normalized scale in (A22), and stitch the recovered data segments in mass-to-charge ratio order to generate the predicted high-resolution mass spectrometry data of absolute intensity.
[0047] To effectively preserve detailed information and improve reconstruction results, a skip connection is added between the upsampling and downsampling layers in the U-Net network structure to directly pass the low-level features extracted during downsampling to the corresponding upsampling layer.
[0048] To improve analytical accuracy, further, in step (A4), the high-resolution mass spectrometry data of the predicted absolute intensity and the average mass spectrometry data in the high-resolution mode are used as input mass spectrometry data, and the input mass spectrometry data are subjected to data preprocessing, visualization analysis, cell typing, and differential metabolite analysis.
[0049] To extract effective single-cell metabolomics features, the data preprocessing is further performed as follows:
[0050] Extract the identified mass spectrometry peaks and their corresponding ion intensities from the input mass spectrometry data to generate a list of metabolic peaks;
[0051] Noise reduction processing;
[0052] Remove background signals from the sampling solvent;
[0053] Missing values were processed using the "80% rule" and the K-value nearest neighbor algorithm to obtain single-cell feature data for analysis.
[0054] To further demonstrate the differences in single-cell metabolic profiles, the visualization analysis is as follows:
[0055] The single-cell feature data to be analyzed is dimensionality reduced, and high-dimensional data in a low-dimensional space is obtained through multivariate analysis. The dimensionality reduction methods used are t-distributed random neighborhood embedding, uniform manifold approximation, and projection, all of which are unsupervised nonlinear dimensionality reduction methods.
[0056] To differentiate single-cell mass spectrometry data of different subtypes, a random forest method is further used in cell typing to distinguish the high-dimensional data of the low-dimensional space of different subtypes, and the performance of the random forest model is evaluated using ROC curves.
[0057] To reveal the differential metabolites screened, a heatmap was further used to display the relative intensities of potential metabolite ions in each single cell during the discovery of differential metabolites.
[0058] Metabolites in the heatmap are sorted according to hierarchical clustering, and horizontal sample clustering shows the differences in metabolite abundance among different subtypes of single cells. Example 2
[0059] Example 1 of the present invention describes the application of a single-cell metabolomics analysis method based on mass spectrometry in the single-cell analysis of bladder cancer.
[0060] This invention uses a single-cell metabolite analysis mass spectrometer (brand: Huayi Ningchuang). The mass spectrometer has a microscopic imaging module, a microscopic manipulation module, and a mass spectrometry detection module. The mass spectrometry detection module can be used in conjunction with a pulsed-DC-ESI source and an LTQ-XL mass spectrometer to obtain mass spectrometry data.
[0061] like Figure 1 As shown, the analysis method in this embodiment includes the following steps:
[0062] (A1) Population Cell Data Acquisition. 5637, 5637-epi, and UMUC3 bladder cancer cells were routinely cultured in RPMI 1640 medium (37°C, 5% CO2) supplemented with 10% FBS. After cell adhesion, the cells were washed with ammonium formate solution to remove excess salt from the cell surface. The cells were then resuspended in 0.9% ammonium formate to obtain a cell suspension. The resuspended cell suspension was centrifuged with an extraction solvent to obtain the supernatant, which was then sonicated in a water bath for 10 min, followed by centrifugation at 12000 rpm for 10 min. The supernatant was collected as the population cell extract. The population cells were then passed through a cell filter to obtain a clean population cell metabolite extract. Then, the average mass spectra of 43 LTQ XL cells from three bladder cancer populations (5637, 5637-epi, and UMUC3) were acquired using a single-cell mass spectrometry system in Turbo (low resolution) and Normal (high resolution) modes of the LTQ-XL mass spectrometer. These average mass spectra were obtained from 14 5637 bladder cancer cells, 13 5637-epi bladder cancer cells, and 16 UMUC3 bladder cancer cells, respectively.
[0063] (A2) Super-resolution model construction. The collected data was used to construct the SCSR-Unet super-resolution network. Data from 5637 and 5637-epi cell populations were used for model training and validation, while data from the UMUC3 cell population was used for testing.
[0064] (A21) Data Alignment. Considering the discrete sampling characteristics of low-resolution mass spectrometry (as shown in Table 1, where the M / Z values are unevenly spaced in Turbo mode), cubic spline interpolation is used to resample the mass axis into continuous data points with 0.1 Da intervals. For example, if the original data is discretely distributed in the range of 100-101 Da (100.0, 100.3, 100.7, 101.0 Da), interpolation can generate a complete spectrum covering all mass points, such as... Figure 2 As shown, it significantly improves the smoothness of spectral lines.
[0065] Table 1 shows the Turbo intensity values before and after interpolation of the UMUC3 bladder cancer population.
[0066] M / Z Turbo_Intensity Turbo_Interpolation_Intensity 100 3191.592 3191.592 100.1 - 2655.242 100.2 - 2286.79 100.3 2051.584 2051.584 100.4 - 1914.974 100.5 - 1842.308 100.6 - 1798.935 100.7 1750.205 1750.205 100.8 - 1671.159 100.9 - 1575.606 101 1487.05 1487.05
[0067] (A22) Data Encoding. To enhance the model's ability to capture local features, a 2D U-Net structure was adopted. After data alignment, to reduce the impact of differences between different peak positions on algorithm performance and improve training speed, 500 102.4 Da fragments were randomly selected from each spectrum in the quality range of 100–1000 Da for training, validation, and testing in Turbo and Normal modes, resulting in a total of 21,500 data fragments. The 2D U-Net structure requires both input and output to be two-dimensional data; therefore, the 102.4 Da fragments were arranged into a 32*32 two-dimensional matrix in mass-to-charge ratio for the model. Next, maximum intensity normalization was performed on each selected fragment to generate training pairs, where low-resolution fragments were used as input data and high-resolution fragments were used as target labels. Finally, the resulting 5637 and 5637-epi population data training pairs were split into training and validation data in an 8:2 ratio; the former was used for network training, and the latter for model validation. All UMUC3 population data training pairs were used for testing.
[0068] (A23) Network Construction and Evaluation. Using the Adam optimizer, the learning rate was set to a fixed 1 × 10⁻³. L2 regularization was applied to the weights to improve generalization ability and avoid overfitting. Training was stopped after 50 iterations. To obtain the optimal model, training and validation were performed by setting the loss function α to 0, 0.3, 0.5, 0.7, and 1. The loss was lowest when α=0, as shown below. Figure 3 As shown in (a) and (d), due to the different properties and dimensions of MSE Loss and CS, the value of the loss function may be affected to varying degrees when their proportions in the loss function change. This leads to the loss function values not being directly comparable under different proportions, and it may be necessary to pay more attention to the effects of MSE and CS. When α=0.5, its MSE is the lowest in the training set and validation set, at 0.009 and 0.019 respectively, as shown in (a) and (d). Figure 3 As shown in (b) and (e), its CS is the highest in the training set, at 0.831, as... Figure 3 As shown in (c); the highest CS in the validation set is 0.877 when α=0.3, followed by 0.871 when α=0.5, as shown in (c). Figure 3As shown in (f). Table 2 shows the performance on the test set when α is set to 0, 0.3, 0.5, 0.7, and 1. When α=0.5, the MSE is the lowest at 0.012, while the CS is the highest at 0.868. Therefore, α=0.5 is selected as the loss function for the model. During the training process, the optimal parameters are obtained through multiple experiments. After training and validation, the model has converged, and the accuracy on the test set is 92.8%. The accuracy here is set to... .
[0069] Table 2 shows the MSE, CS, and accuracy values for the test set.
[0070] Test MSE Loss Test Cosine Similarity Test Accuracy α=0 0.014 0.803 0.986 α=0.3 0.013 0.862 0.950 α=0.5 0.012 0.868 0.928 α=0.7 0.016 0.839 0.883 α=1 0.019 0.843 0.843
[0071] (A24) High-resolution spectral prediction. Based on the detection results of UMUC3 bladder cancer cell metabolites by the LTQ XL mass spectrometer in Turbo and Normal modes, this study improved the spectral quality using the SCSR-Unet algorithm.
[0072] First, the cell populations of three different subtypes of bladder cancer were analyzed, and 12 potential differentially expressed amino acid metabolites were identified, as shown in Table 3. Figure 4 This is a mass spectrum of the absolute intensities of three metabolite ions from the UMUC3 bladder cancer cell population: leucine (m / z -> 132), adenine (m / z -> 136), and glutamic acid (m / z -> 148). Among these three characteristic ions, [the following is a partial translation of the original text, which is incomplete and requires further context]. Figure 4 (a), (b), and (c) are absolute intensity mass spectra of UMUC3 bladder cancer cell population data at 132, 136, and 148 Da in Turbo mode. The Turbo mode causes peak overlap due to insufficient resolution. Figure 4 (d), (e), and (f) show the absolute intensity mass spectra of the UMUC3 bladder cancer cell population at 132, 136, and 148 Da in both Normal and SR modes. Both the Normal mode and SR spectra effectively separated the target peak. After maximum intensity normalization, the signal intensity of the SR spectra at the three characteristic peaks reached 1.33, 1.48, and 2.15 times that of the Normal mode, respectively, with an average fold increase of 1.65, as shown in Table 4.
[0073] Table 3 lists 12 potential metabolites that can be detected in bladder cancer cell populations.
[0074]
[0075] Table 4 shows the parameters of the UMUC3 bladder cancer population cell data at 132, 136, and 148 Da.
[0076] M / Z 132 136 148 average value The intensity of SR / Normal before normalization 1.33 1.48 2.15 1.65 The FWHM (Da) of SR after normalization 0.30 0.30 0.29 0.30 The FWHM (Da) of Normal after normalization 0.22 0.20 0.43 0.28 The mass shift (%) of SR / Normal 10.00 0 20.00 10 The relative intensity deviation (%) of SR / Normal 1.3 0.17 3.04 4.51
[0077] To quantitatively evaluate algorithm performance, Figure 5 (a), (b), and (c) show the normalized peak values of the Normal and SR modes at 132, 136, and 148 Da for the UMUC3 bladder cancer cell population data. The full width at half maximum (FWHM) of the three peaks in the SR mode are 0.30, 0.30, and 0.29 Da (mean 0.30 Da), which are basically equivalent to those in the Normal mode (0.22, 0.20, and 0.43 Da (mean 0.28 Da)). The algorithm predicts an average mass shift of 10% in the spectrum, with the average relative intensity deviation controlled within 4.51%.
[0078] (A3) Single-cell data acquisition. Single cells were obtained by dropping the resuspension from step 1 onto oxygenated PDMS. Then, 296 average mass spectra of three types of bladder single cells (5637, 5637-epi, and UMUC3) were acquired using a single-cell mass spectrometry analysis system in Turbo and Normal modes of the LTQ-XL mass spectrometer. These average mass spectra were obtained from 100 5637 bladder cancer single cells, 93 5637-epi bladder cancer single cells, and 103 UMUC3 bladder cancer single cells, respectively.
[0079] (A4) Analysis of single-cell metabolomics. The collected low-resolution data were processed through the SCSR-Unet network to obtain high-resolution data after super-resolution prediction. This data was then used for single-cell metabolomics analysis.
[0080] (A41) Data Preprocessing. To extract key metabolomics information from single cells and reduce noise interference and the influence of exogenous pollutants during the detection process, a series of preprocessing steps were performed on the high-resolution and low-resolution mass spectrometry data obtained from the super-resolution prediction. First, all identified mass spectrometry peaks (i.e., m / z values) and their corresponding ion intensities were extracted from the mass spectra to generate a list of metabolic peaks. Second, noise reduction was performed using filtering methods. Specifically, the signal was decomposed using wavelet transform, noise was removed using adaptive thresholding, and the signal was smoothed using a Savitzky-Golay filter. Subsequently, the background signal of the sampling solvent was removed, and the total ion intensity of all detected ion intensities was normalized to form a matrix data containing the detected ions and their relative intensities. Finally, missing values were processed using the "80% rule" and the K-nearest neighbor algorithm to obtain the final single-cell feature data to be analyzed. Through the above preprocessing steps, the dimensionality of the dataset was significantly reduced while preserving the effective metabolomics information of single cells.
[0081] (A42) Visualization Analysis. t-SNE and UMAP were used to visualize the metabolic profiles of single-cell bladder cancer cells, and the differences between cells were displayed in a two-dimensional space. For example... Figure 6 As shown in (a) and (c), the t-SNE clustering results of the three bladder cancer cell subtypes under different conditions present the following situations: Figure 6 (a) High-resolution mass spectrometry data acquired in Normal mode. Figure 6 (b) Super-resolution mass spectrometry data obtained through the SCSR-Unet network. Figure 6 (c) Low-resolution mass spectrometry data acquired in Turbo mode. Under all three conditions, the t-SNE parameters were set to a learning rate of 50 and a perplexity of 10. As shown in the figure, after super-resolution processing, the differentiation of the three bladder cancer cell subtypes was significantly improved, with the distance between different groups being significantly greater than the distance within groups, approaching the cell subtype differentiation effect in Normal mode. The original data in Turbo mode, however, showed weaker differentiation ability. The Calinski-Harabasz Index (CH index), Davies-Bouldin Index (DB index), and Silhouette Coefficient (SC) were used to evaluate the clustering effect, as shown in Table 5. The values of all three indicators after super-resolution were higher than those in Turbo mode. To further verify this hypothesis and to visualize the clustering effect in another way, UMAP was used for dimensionality reduction analysis. Figure 6(d), (e), and (f) represent the UMAP clustering results of high-resolution mass spectrometry data acquired in Normal mode, super-resolution mass spectrometry data obtained through the SCSR-Unet network, and low-resolution mass spectrometry data acquired in Turbo mode, respectively, with parameters set to n_neighbors=8 and min_dist=0.05. The results show that, consistent with the visualization of t-SNE, the super-resolution data performs better in distinguishing different bladder cancer cell subtypes. Similarly, the values of the three indicators after super-resolution are all higher than those in Turbo mode.
[0082] Table 5 shows the dimensionality reduction results based on t-SNE and UMAP.
[0083]
[0084] (A43) Cell typing. To evaluate the classification performance of different subtypes of bladder cancer cells under three conditions, a machine learning method with mature applications in metabolomics research—Random Forest (RF)—was used to classify the phenotypic distribution of preprocessed single-cell datasets and summarize the results. In the RF algorithm, the construction of decision trees is the core of model performance, and the number of decision trees (n_estimators) and the depth (max_depth) are key parameters affecting the efficiency and classification performance of RF. Here, to ensure a fair comparison of the three datasets, the same model architecture and hyperparameters were used. The parameters n_estimators=100 and max_depth=10 were set. In addition, all three models used an 8:2 ratio to divide the training and test sets.
[0085] Figure 7 The classification results of the test set are shown under three conditions. Figure 7 (a), (b), and (c) represent the confusion matrices for Normal mode, after super-resolution, and Turbo mode, respectively, showing the degree of matching between the model's predicted classes and the actual classes. The diagonal elements of the confusion matrix represent the number of correctly classified samples, and the off-diagonal elements represent the distribution of misclassified samples. Furthermore, Figure 7Figures (d), (e), and (f) show the ROC curves for the three modes, used to evaluate the model's classification performance at different thresholds. The closer the ROC curve is to the upper left corner, the stronger the model's discriminative ability. As can be seen from the figures, the super-resolution mode exhibits the most outstanding classification performance, approaching that of the Normal mode, while the Turbo mode's classification performance is slightly inferior. To better evaluate the model's classification accuracy, Table 6 presents the RF-based classification results for the three scenarios. The accuracy, precision, recall, and F1 score (all 0.99) after super-resolution are all greater than the results of the Turbo mode (>0.87), and similar to the results of the Normal mode (>0.98). This demonstrates that the super-resolution model can significantly improve single-cell classification performance.
[0086] Table 6 shows the classification results based on radio frequency (RF).
[0087] Normal SCSR Turbo Accuracy 0.98 0.99 0.87 Precision 0.99 0.99 0.88 Recall 0.98 0.99 0.87 F1 score 0.98 0.99 0.87
[0088] (A44) Metabolite heterogeneity analysis of bladder cancer cell subtypes. In metabolomics research, screening for potential differential metabolites is a crucial step in data analysis. To verify whether super-resolution models can facilitate the discovery of differential metabolites, 12 potential amino acid differential metabolites detected in bladder cancer cell populations were used to create heatmaps of single-cell data to visually demonstrate the relative intensity of each metabolite ion in a single cell. Figure 8 Images (a), (b), and (c) respectively present heatmap results of single-cell data in Normal mode, after super-resolution, and in Turbo mode. The heatmap results after super-resolution are particularly crucial, clearly showing a high degree of consistency in the single-cell metabolic profiles of the same cell subtype, while significant differences are observed between different subtypes. Furthermore, its distribution characteristics are quite similar to those in Normal mode, indicating that the super-resolution model performs well in preserving cellular metabolic features. Taking arginine (m / z->175) as an example, in the heatmap after super-resolution, its intensity is highest in 5637-epi cells, followed by UMUC3 cells, and lowest in 5637 cells. This clear differential distribution is significant in metabolomics research, helping researchers quickly locate the expression differences of key metabolites in different cell subtypes. However, in Turbo mode, this difference is not significant, further highlighting the advantage of the super-resolution model in enhancing differential metabolite expression.
[0089] Figure 8 (d) shows the secondary mass spectrum of arginine (m / z->175) in single-cell detection, from which daughter ions 116 and 158 can be detected, providing more in-depth mass spectrometry information for the metabolic characteristics analysis of arginine. Figure 8(e) and (f) show the first-order mass spectra in the range of 174 to 177 under Turbo and SR conditions. It can be clearly seen from the figure that the metabolite ion arginine can be clearly separated after super-resolution (m / z->175). This result strongly proves that the super-resolution model can effectively improve the mass spectrometry resolution.
Claims
1. A single-cell metabolomics analysis method based on mass spectrometry, characterized in that, Includes the following steps: (A1) Collect population cell data, including multiple average mass spectra of various population cells in low-resolution and high-resolution modes; (A2) Construct the SCSR-Unet super-resolution model using the population cell data, and train, validate and test the model; In network construction, the loss function CL used is: ; output i It is the prediction result of the sample, label i Here, α is the label corresponding to the sample, i represents the i-th sample, α is the hyperparameter controlling the loss weight, and N is the number of samples; The method for constructing the SCSR-Unet super-resolution model includes the following steps: (A21) Data alignment: interpolate the average mass spectrometry data in the low-resolution mode to obtain the same equidistant mass-to-charge ratio data as the average mass spectrometry data in the high-resolution mode; (A22) Data encoding: Multiple segments of equal length are extracted from the average mass spectrometry data in low-resolution and high-resolution modes. The segments are arranged in order of mass-to-charge ratio into a two-dimensional matrix. After normalization, training pairs are generated. The low-resolution segments are used as input data, and the high-resolution segments are used as target labels. (A23) Network construction: Using a two-dimensional U-Net network structure, the SCSR-Unet super-resolution model is constructed using the loss function described above; (A24) Input the average mass spectrometry data of the low-resolution mode to be tested into the trained SCSR-Unet super-resolution model to obtain the predicted high-resolution mass spectrometry data. (A25) Recover the predicted high-resolution mass spectrometry data according to the normalized scale in (A22), and stitch the recovered data segments in mass-to-charge ratio order to generate the predicted high-resolution mass spectrometry data of absolute intensity. (A3) Acquire single-cell data, including average mass spectrometry data of multiple single cells in low-resolution and high-resolution modes; (A4) The average mass spectrometry data in low resolution mode is passed through the SCSR-Unet model to obtain high-resolution mass spectrometry data with predicted absolute intensity. The average mass spectrometry data in high resolution mode is preprocessed simultaneously and used for single-cell metabolomics analysis. In step (A4), the high-resolution mass spectrometry data of the predicted absolute intensity and the average mass spectrometry data in the high-resolution mode are used as input mass spectrometry data, and the input mass spectrometry data are subjected to data preprocessing, visualization analysis, cell typing and differential metabolite analysis. The visualization analysis involves reducing the dimensionality of the single-cell feature data to be analyzed and obtaining high-dimensional data in a low-dimensional space through multivariate analysis. The dimensionality reduction method is: t-distributed random neighborhood embedding, uniform manifold approximation, and projection.
2. The single-cell metabolomics analysis method according to claim 1, characterized in that, In the U-Net network structure, a skip connection is added between the upsampling layer and the downsampling layer to directly pass the low-level features extracted during the downsampling process to the corresponding upsampling layer.
3. The single-cell metabolomics analysis method according to claim 1, characterized in that, The data preprocessing is as follows: Extract the identified mass spectrometry peaks and their corresponding ion intensities from the input mass spectrometry data to generate a list of metabolic peaks; Noise reduction processing; Remove background signals from the sampling solvent; Missing values were processed using the "80% rule" and the K-value nearest neighbor algorithm to obtain single-cell feature data for analysis.
4. The single-cell metabolomics analysis method according to claim 1, characterized in that, The cell typing method uses a random forest approach to distinguish the high-dimensional data in the low-dimensional space of different subtypes, and the performance of the random forest model is evaluated using ROC curves.
5. The single-cell metabolomics analysis method according to claim 1, characterized in that, In the discovery of these differential metabolites, a heatmap was used to show the relative intensities of potential metabolite ions in each single cell; The metabolites in the heatmap are sorted according to hierarchical clustering, and the horizontal sample clustering shows the differences in metabolite abundance among different subtypes of single cells.
6. The single-cell metabolomics analysis method according to claim 1, characterized in that, The cell population includes 5637, 5637-epi, and UMUC3 bladder cancer single cells.
Citation Information
Patent Citations
Zero sample learning-based mass spectrum image super-resolution reconstruction method
CN118014843A
MALDI matrix and MALDI method
EP2060919A1