Single-cell metabonomics analysis method based on mass spectrum technology

Through the SCSR-Unet super-segment model and the mixed loss function CL, the contradiction between resolution and sensitivity of the linear ion trap mass spectrometer in high-speed scanning mode is solved, and the reconstruction of high-resolution mass spectrometry data is achieved, the accuracy and typing ability of single-cell metabolomics analysis are improved, and the heterogeneity of single-cell bladder cancer is revealed.

CN120374399AActive Publication Date: 2025-07-25CHINA INNOVATION INSTR CO LTD

Patent Information

Application Number
CN202510864885.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

The existing linear ion trap mass spectrometer has a contradiction in resolution and sensitivity in high-speed scanning mode, resulting in the inability to separate isotope peaks and adjacent mass-to-charge ratio ions, affecting the accuracy of single-cell metabolites identification and cell subtype classification. The existing deep learning algorithms have not been optimized for the low abundance and high noise characteristics of single-cell metabolic data, and lack of systematic verification of the heterogeneity of tumor cells such as bladder cancer.

Method used

The SCSR-Unet super-segment model is adopted, combined with the mixed loss function CL, and the low-resolution mass spectrometry data is reconstructed into high-resolution data through deep learning, improving signal reconstruction accuracy and resolution, using the 2D U-Net structure to retain detailed information, and optimizing the direction consistency of the mass spectrometry data through the loss function combined with cosine similarity and mean square error.

Benefits of technology

The mass spectrometry resolution can be significantly improved without hardware upgrades, the accuracy of single-cell typing, reveal the heterogeneity characteristics of single-cell bladder cancer single cells, enhance the effective characteristics of single-cell metabolomic analysis, and improve the accuracy of typing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374399A_ABST
    Figure CN120374399A_ABST
Patent Text Reader

Abstract

The invention belongs to a mass spectrometry technology, and particularly provides a single cell metabonomics analysis method based on the mass spectrometry technology, which comprises the following steps: (A1) collecting population cell data including a plurality of average mass spectrograms of a plurality of population cells in a low-resolution mode and a high-resolution mode; (A2) constructing an SCSR-Unet super-division model by using the population cell data, and training, verifying and testing the model; a loss function CL is used in network construction; (A3) collecting single cell data, wherein the single cell data comprises an average mass spectrum of a plurality of single cells in a low-resolution mode and a high-resolution mode; and (A4) enabling the average mass spectrum data in the low-resolution mode to pass through an SCSR-Unet model to obtain high-resolution mass spectrum data of predicted absolute intensity, and synchronously preprocessing the high-resolution mass spectrum data and the mass spectrum data in the high-resolution mode for single-cell metabonomics analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to mass spectrometry detection technology, and particularly to a single-cell metabolomics analysis method based on mass spectrometry technology. Background Art

[0002] Single-cell metabolomics provides an important means for revealing tumor heterogeneity, drug response mechanisms, etc. by analyzing the metabolic characteristics of individual cells. Mass spectrometry technology has become the core tool for single-cell metabolism analysis due to its high sensitivity and wide dynamic range. However, existing linear ion trap mass spectrometers (such as LTQ XL) face the contradiction between resolution and sensitivity in high-speed scanning modes (such as Turbo mode), that is, increasing the scanning speed will reduce its resolution, resulting in the inability to separate isotope peaks and ions with adjacent mass-to-charge ratios, seriously affecting the accuracy of metabolite identification and cell subtype classification.

[0003] Traditional solutions mostly rely on hardware iteration (such as orbitrap mass spectrometry) or external condition optimization, but there are problems such as high equipment cost and low throughput. Although there are already super-resolution algorithms based on deep learning for improving mass spectrometry resolution, they are not optimized for the low-abundance and high-noise characteristics of single-cell metabolic data, and lack systematic verification of tumor cell heterogeneity such as bladder cancer. In addition, the existing single-cell metabolomics analysis process is fragmented, traditional dimensionality reduction methods (such as PCA) are prone to losing key biological information, and machine learning models also lack a standardized process. Summary of the Invention

[0004] To solve the deficiencies in the above-mentioned prior art solutions, the present invention provides a single-cell metabolomics analysis method based on mass spectrometry technology.

[0005] The object of the present invention is achieved through the following technical solutions: A single-cell metabolomics analysis method based on mass spectrometry technology, comprising the following steps: (A1) Collect population cell data, including multiple average mass spectra of multiple population cells in low-resolution mode and high-resolution mode; (A2) Use the population cell data to construct an SCSR-Unet super-resolution model, and perform model training, verification, and testing; In network construction, the loss function CL used is: ; output i is the prediction result of the sample, label i is the label corresponding to the sample, i represents the i-th sample, α is a hyperparameter controlling the loss weight, and N is the number of samples; (A3) Collect single-cell data, including average mass spectrometry data of multiple single cells in low-resolution mode and high-resolution mode; (A4) The average mass spectrometry data in the low-resolution mode is processed through the SCSR-Unet model to obtain the predicted high-resolution mass spectrometry data of the absolute intensity, which is synchronously preprocessed with the average mass spectrometry data in the high-resolution mode for single-cell metabolomics analysis.

[0006] Compared with the prior art, the beneficial effects of the present invention are as follows.

[0007] 1. The hybrid loss function used in the modeling improves the signal reconstruction accuracy; A hybrid loss function is designed to optimize the numerical error of the signal while constraining the overall direction consistency of the spectral shape, solving the problem of mutual limitation between the resolution and sensitivity of the mass spectrometer. The optimization of the parameter α significantly improves the discrimination ability of overlapping peaks.

[0008] 2. The improvement of the mass spectrometry resolution can be achieved without enhancing the hardware technology; The SCSR-Unet algorithm reconstructs the mass spectrometry data scanned at high speed in the low-resolution mode into the predicted high-resolution data through deep learning modeling, which is equivalent to the hardware level of slow scanning in the high-resolution mode. Without replacing high-cost hardware such as the orbitrap, the resolution bottleneck of the ion trap mass spectrometer can be broken through.

[0009] 3. The accuracy of single-cell typing is significantly improved; The improvement of the resolution of the mass spectrometry data increases the effective features of single-cell metabolomics, thereby significantly improving the typing accuracy and fully demonstrating the heterogeneous characteristics of bladder cancer single cells. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Referring to the accompanying drawings, the disclosure of the present invention will become more understandable. It is easy for those skilled in the art to understand that these drawings are only used to illustrate the technical solutions of the present invention and are not intended to limit the protection scope of the present invention. In the figures: Figure 1 is a schematic flowchart of the analysis method of the present invention; Figure 2 is the Turbo absolute intensity mass spectrometry diagram of UMUC3 bladder cancer population cells before and after cubic spline interpolation in the range of 100 Da - 101 Da; Figure 3 is the loss curve and accuracy curve diagram of the training set and the validation set; Figure 4 is the absolute intensity mass spectrometry diagram of UMUC3 bladder cancer population cell data at 132, 136, and 148 Da under Turbo, Normal, and SR conditions; Figure 5It is the relative intensity mass spectrometry diagram at 132, 136, and 148 Da for the UMUC3 bladder cancer population cell data under the Normal and SR conditions; Figure 6 It is the t-SNE and UMAP clustering results under the Normal, SR, and Turbo conditions; Figure 7 It is the RF confusion matrix and ROC curve under the Normal, SR, and Turbo conditions; Figure 8 It is the heat map, primary comparison spectrum of arginine, and secondary spectrum under the Normal, SR, and Turbo conditions. Detailed implementation manners

[0011] Figure 1 - Figure 8 The following description and the following illustrate alternative specific embodiments of the present invention to teach those skilled in the art how to implement and reproduce the present invention. Some conventional aspects have been simplified or omitted for the purpose of teaching the technical solutions of the present invention. Those skilled in the art should understand that variations or substitutions derived from these specific embodiments will fall within the scope of the present invention. Those skilled in the art should understand that the following features can be combined in various ways to form multiple variations of the present invention. Thus, the present invention is not limited to the following alternative specific embodiments, but is only defined by the claims and their equivalents. Example 1

[0012] This example is a single-cell metabolomics analysis method based on mass spectrometry technology, as Figure 1 shown, and includes the following steps: (A1) Collect population cell data, including multiple average mass spectrometry diagrams of various population cells in low-resolution mode and high-resolution mode; (A2) Use the population cell data to construct an SCSR-Unet super-resolution model, and perform model training, validation, and testing; In network construction, the loss function CL used is: ; output i is the prediction result of the sample, label i is the label corresponding to the sample, i represents the i-th sample, α is a hyperparameter controlling the loss weight, and N is the number of samples; (A3) Collect single-cell data, including average mass spectrometry data of various single cells in low-resolution mode and high-resolution mode; (A4) Pass the average mass spectrometry data in low-resolution mode through the SCSR-Unet model to obtain the predicted high-resolution mass spectrometry data of absolute intensity, and synchronously preprocess it with the average mass spectrometry data in high-resolution mode for single-cell metabolomics analysis.

[0013] In order to improve the accuracy of prediction, further, in step (A2), the method of constructing the model includes the steps of: (A21) data alignment, performing subtraction on the average mass spectrum data in the low-resolution mode to obtain equally spaced mass-to-charge ratio data that is the same as the average mass spectrum data in the high-resolution mode; (A22) Data encoding: extract multiple fragments of equal length from the average mass spectrum data in low-resolution and high-resolution modes, arrange the fragments in a two-dimensional matrix in order of mass-to-charge ratio, and generate training pairs after normalization, with the low-resolution fragments as input data and the high-resolution fragments as target labels; (A23) Network construction. 2D U-Net is used as the network structure, in which the input and output are both two-dimensional matrices of size 32×32, representing low-resolution data and predicted high-resolution data, respectively. The 2D U-Net structure consists of 3 downsampling layers and 3 upsampling layers, each of which contains 2 two-dimensional convolutional layers to extract the features of the input data. In addition, a jump connection is added between the downsampling layer and the upsampling layer to directly transfer the underlying features extracted during the downsampling process to the corresponding upsampling layer, thereby effectively retaining the detail information and improving the reconstruction effect. In SR-Unet, the mean square error (MSE) is used as the loss function. However, considering that MSE is a loss function based on the difference in raw numerical values, it mainly focuses on the square difference between each data point. Therefore, it often pays too much attention to the absolute error during training and tends to ignore the relative direction between data. Cosine similarity (CS) focuses on the angular difference between two vectors, is not affected by the absolute size of the data, and is more suitable for measuring the relative relationship between samples. For mass spectrometry data, it usually involves comparisons of different mass-to-charge ratios (m / z) or intensities, and cosine similarity can better capture the directionality of the data and help the network learn the global structure and regularity of the data. Therefore, in order to obtain better training results, the loss function is set to a loss function that combines MSE and CS. The loss function CL is: .

[0014] output i is the prediction result of the sample, label i is the label corresponding to the sample, i represents the i-th sample, α is the hyperparameter that controls the loss weight, and N is the number of samples.

[0015] The loss of cosine similarity is calculated. , 。α is a hyperparameter that controls the weights of two losses, and its value range is (0-1). By adjusting α, the influence between the directionality (measured by CS) and the amplitude difference (measured by MSE) of the prediction vectors can be balanced. Setting the loss function in this way can not only optimize the amplitude accuracy of the mass spectrometry signal but also ensure the consistency of the spectral shape, thereby improving the final super-resolution effect.

[0016] (A24) Input the average mass spectrometry data in the low-resolution mode to be tested into the trained SCSR-Unet super-resolution model to obtain the predicted high-resolution mass spectrometry data. (A25) Restore the predicted high-resolution mass spectrometry data according to the normalization scale in (A22), and splice the restored data segments in the order of mass-to-charge ratio to generate the predicted high-resolution mass spectrometry data of the absolute intensity.

[0017] To effectively retain detailed information and improve the reconstruction effect, further, in the U-Net network structure, skip connections are added between the upsampling layer and the downsampling layer to directly transfer the underlying features extracted during the downsampling process to the corresponding upsampling layer.

[0018] To improve the analysis accuracy, further, in step (A4), the predicted high-resolution mass spectrometry data of the absolute intensity and the average mass spectrometry data in the high-resolution mode are used as the input-end mass spectrometry data, and data preprocessing, visualization analysis, cell typing, and differential metabolites are performed on the input-end mass spectrometry data.

[0019] To extract effective single-cell metabolomics features, further, the data preprocessing is as follows: Extract the identified mass spectrometry peaks and their corresponding ion intensities from the input-end mass spectrometry data to generate a metabolite peak list. Denoising processing. Remove the background signal of the sampling solvent. Use the "80% rule" and the K-nearest neighbor algorithm to process the missing values to obtain the single-cell feature data to be analyzed.

[0020] To display the differences in single-cell metabolic profiles, further, the visualization analysis is as follows: Perform dimensionality reduction on the single-cell feature data to be analyzed to obtain high-dimensional data in a low-dimensional space through multivariate analysis. The dimensionality reduction methods used are: t-distributed stochastic neighbor embedding and uniform manifold approximation and projection, both of which are unsupervised non-linear dimensionality reduction methods.

[0021] To distinguish single-cell mass spectrometry data of different subtypes, further, in cell typing, the random forest method is used to distinguish the high-dimensional data in the low-dimensional space of different subtypes, and the ROC curve is used to evaluate the performance of the random forest model.

[0022] To reveal the screened differential metabolites, further, in the discovery of differential metabolites, a heatmap was used to display the relative intensities of potential metabolite ions in each single cell; The metabolites in the heatmap were sorted according to hierarchical clustering, and the horizontal sample clustering showed the differences in metabolite abundances among different subtypes of single cells. Example 2

[0023] An application example of the single-cell metabolomics analysis method based on mass spectrometry technology according to Example 1 of the present invention in the single-cell analysis of bladder cancer.

[0024] The present invention uses a single-cell metabolite analysis mass spectrometer (brand: Huayi Ningchuang). This mass spectrometer has a microscopic imaging module, a micromanipulation module, and a mass spectrometry detection module. Among them, the mass spectrometry detection module can adopt a combined scheme of a pulsed direct current electrospray ion source (Pulsed-dc-ESI) and an LTQ-XL mass spectrometer to obtain mass spectrometry data.

[0025] As Figure 1 shown, the analysis method of this example includes the following steps: (A1) Collection of population cell data. The 5637, 5637-epi, and UMUC3 bladder cancer cells were routinely cultured in RPMI 1640 medium supplemented with 10% FBS (37 °C, 5% CO2). After the cells adhered to the wall, the cells were washed with ammonium formate solution to remove the excess salts on the cell surface, and the cells were resuspended with 0.9% ammonium formate to obtain a cell suspension. The supernatant was obtained by centrifuging the resuspended solution with an extractant, placed in a water bath and ultrasonically broken for 10 min, and then centrifuged at 12,000 rpm for 10 min. The supernatant was taken as the population cell extract, and the population cells were passed through a cell filter to obtain a clean population cell metabolite extract. Then, 43 LTQ XL average mass spectrometry maps of the three bladder cancer population cells of 5637, 5637-epi, and UMUC3 were respectively collected by a single-cell mass spectrometry analysis system in the Turbo (low resolution) and Normal (high resolution) modes of the LTQ-XL mass spectrometer, which were the average mass spectrometry maps of 14 5637 bladder cancer population cells, 13 5637-epi bladder cancer population cells, and 16 UMUC3 bladder cancer population cells, respectively.

[0026] (A2) Construction of a super-resolution model. The collected data was used for the construction of the SCSR-Unet super-resolution network. The data of 5637 and 5637-epi population cells were used for the training and validation of the model, and the data of UMUC3 population cells were used for testing.

[0027] (A21)Data alignment. Considering the discrete sampling characteristics of low-resolution mass spectrometry, as shown in Table 1, the M / Z value intervals are uneven in Turbo mode. The cubic spline interpolation method is used to resample the mass number axis into continuous data points at 0.1 Da intervals. For example, the original data is discretely distributed in the range of 100 - 101 Da (100.0, 100.3, 100.7, 101.0 Da). After interpolation, a complete spectrum covering all mass points can be generated, as Figure 2 shown, significantly improving the spectral line smoothness.

[0028] Table 1 shows the Turbo intensity values of UMUC3 bladder cancer population cells before and after interpolation.

[0029] M / Z Turbo_Intensity Turbo_Interpolation_Intensity 100 3191.592 3191.592 100.1 - 2655.242 100.2 - 2286.79 100.3 2051.584 2051.584 100.4 - 1914.974 100.5 - 1842.308 100.6 - 1798.935 100.7 1750.205 1750.205 100.8 - 1671.159 100.9 - 1575.606 101 1487.05 1487.05 (A22)Data encoding. To enhance the model's ability to capture local features, a 2D U-Net structure is adopted. After data alignment, to reduce the impact of differences between different peak positions on the algorithm performance and improve the training rate, in Turbo and Normal modes, 500 fragments of 102.4 Da are randomly selected from each spectrum in the mass range of 100 - 1000 Da for training, validation, and testing, with a total of 21,500 data fragments. The 2D U-Net structure requires both the input and output to be two-dimensional data. Therefore, the 102.4 Da fragments are arranged in a 32*32 two-dimensional matrix according to the mass-to-charge ratio order for the model. Next, maximum intensity normalization is performed on each intercepted fragment to generate training pairs, where the low-resolution fragment is used as the input data and the high-resolution fragment is used as the target label. Finally, the obtained 5637 and 5637-epi population data training pairs are split into training data and validation data in an 8:2 ratio. The former is used for network training, and the latter is used for model validation. The UMUC3 population data training pairs are all used for testing.

[0030] (A23)Network construction and evaluation. Using the Adam optimizer, the learning rate is set to a fixed 1 × 10−3, and L2 regularization is applied to the weights to improve the generalization ability and avoid overfitting. Training stops after 50 epochs. To obtain the optimal model, here the training and validation are performed by setting α of the loss function to 0, 0.3, 0.5, 0.7, and 1. When α = 0, the loss is the lowest, as Figure 3 (a), (d) shown. However, due to the different properties and dimensions of MSE Loss and CS, when their proportions in the loss function change, the value of the loss function may be affected to varying degrees. This results in the fact that at different proportions, the values of the loss function may not be directly comparable, and more attention may need to be paid to the effects of MSE and CS. When α = 0.5, the MSE in the training set and validation set is the lowest, 0.009 and 0.019 respectively, asFigure 3 (b) and (e) as shown; its CS in the training set is the highest, which is 0.831, as Figure 3 shown in (c); the highest CS in the validation set is 0.877 when α = 0.3, and the cosine similarity at α = 0.5 is the second highest, which is 0.871, as Figure 3 shown in (f). As shown in Table 2, when α is set to 0, 0.3, 0.5, 0.7, and 1, its performance on the test set. When α = 0.5, its MSE is the lowest, which is 0.012, and at the same time, its CS is the highest, which is 0.868. Therefore, α = 0.5 is selected for the loss function of the model. During the training process, the optimal parameters are obtained through multiple experiments. After training and validation, the model has converged, and the accuracy rate on the test set is 92.8%. The accuracy rate here is set as .

[0031] Table 2 shows the MSE values, CS values, and accuracy rate values of the test set.

[0032] Test MSE Loss Test Cosine Similarity Test Accuracy α=0 0.014 0.803 0.986 α=0.3 0.013 0.862 0.950 α=0.5 0.012 0.868 0.928 α=0.7 0.016 0.839 0.883 α=1 0.019 0.843 0.843 (A24) High-resolution spectral map prediction. Aiming at the detection results of metabolites of UMUC3 bladder cancer cells by LTQ XL mass spectrometer in Turbo and Normal modes, this study achieved the improvement of spectral quality through the SCSR-Unet algorithm.

[0033] First, the cells of three different subtypes of bladder cancer populations were detected and analyzed, and 12 potential amino acid differential metabolites were identified, as shown in Table 3, Figure 4 which are the mass spectrometry maps of the absolute intensities of three metabolite ions of UMUC3 bladder cancer population cells. These three metabolite ions are leucine (leucine, m / z->132), adenine (adenine, m / z->136), and glutamic (glutamic acid, m / z->148). In the detection of these three characteristic ions, Figure 4 (a), (b), and (c) are the absolute intensity mass spectrometry maps of UMUC3 bladder cancer population cell data at 132, 136, and 148 Da in Turbo mode. The Turbo mode results in peak overlap due to insufficient resolution. Figure 4 (d), (e), and (f) are the absolute intensity mass spectrometry maps of UMUC3 bladder cancer population cell data at 132, 136, and 148 Da in Normal mode and algorithm reconstruction (SR) cases. Both the Normal mode and the algorithm reconstruction (SR) spectral maps can effectively separate the target peaks. After the maximum intensity normalization processing, the signal intensities of the SR spectral map at the three characteristic peaks reach 1.33, 1.48, and 2.15 times that of the Normal mode respectively, and the average multiple is 1.65 times, as shown in Table 4.

[0034] Table 3 shows 12 potential metabolites detectable in the bladder cancer cell population.

[0035]

[0036] Table 4 shows the parameters of the UMUC3 bladder cancer cell population data at 132, 136, and 148 Da.

[0037] M / Z 132 136 148 Average value The intensity of SR / Normal before normalization 1.33 1.48 2.15 1.65 The FWHM (Da) of SR after normalization 0.30 0.30 0.29 0.30 The FWHM (Da) of Normal after normalization 0.22 0.20 0.43 0.28 The mass shift (%) of SR / Normal 10.00 0 20.00 10 The relative intensity deviation (%) of SR / Normal 1.3 0.17 3.04 4.51 To quantitatively evaluate the algorithm performance, Figure 5 (a), (b), and (c) are the normalized comparison of the characteristic peaks of the Normal and SR modes of the UMUC3 bladder cancer cell population data at 132, 136, and 148 Da. The full width at half maximum (FWHM) of the three peaks in the SR mode are 0.30, 0.30, and 0.29 Da (mean 0.30 Da), which are basically equivalent to 0.22, 0.20, and 0.43 Da (mean 0.28 Da) in the Normal mode. The average mass shift of the algorithm-predicted spectrum is 10%, and the average relative intensity deviation is controlled within 4.51%.

[0038] (A3) Single-cell data acquisition. The resuspension obtained in step 1 was dropped onto the oxygenated PDMS to obtain single cells. Then, 296 LTQ XL average mass spectra of three bladder single cells, namely 5637, 5637-epi, and UMUC3, were collected by the single-cell mass spectrometry analysis system in the Turbo and Normal modes of the LTQ-XL mass spectrometer, which are the average mass spectra of 100 5637 bladder cancer single cells, 93 5637-epi bladder cancer single cells, and 103 UMUC3 bladder cancer single cells, respectively.

[0039] (A4) Analysis of single-cell metabolomics. The collected low-resolution data was passed through the SCSR-Unet network to obtain the super-resolved and predicted high-resolution data. Subsequently, it was used for the analysis of single-cell metabolomics.

[0040] (A41) Data preprocessing. To extract key metabolomics information from single cells and reduce the influence of noise interference and exogenous pollutants during detection, a series of preprocessing steps were performed on the high-resolution mass spectrometry data and low-resolution mass spectrometry data predicted by super-resolution. First, all identified mass spectrometry peaks (i.e., m / z values) and their corresponding ion intensities were extracted from the mass spectrometry spectra to generate a metabolite peak list. Second, a filtering method was applied for noise reduction. The specific processing steps were to first decompose the signal through wavelet transform, then apply an adaptive threshold to remove noise, and finally use the Savitzky-Golay (SG) filter to smooth the signal. Subsequently, the background signal of the sampling solvent was removed, and the total ion intensity normalization was performed on all detected ion intensities to form matrix data containing detected ions and their relative intensities. Finally, the "80% rule" and the K-nearest neighbor algorithm were used to process missing values to obtain the final single-cell feature data to be analyzed. Through the above preprocessing steps, the dimension of the dataset was significantly reduced, and at the same time, the effective metabolomics information of single cells was completely retained.

[0041] (A42) Visualization analysis. t-SNE and UMAP were used to perform visualization analysis on the single-cell metabolomic profiles of bladder cancer, and the differences between cells were shown through a two-dimensional space. As Figure 6 shown in (a) and (c), the t-SNE clustering results of three bladder cancer cell subtypes under different conditions respectively showed the following situations: Figure 6 (a) was the high-resolution mass spectrometry data collected in the Normal mode, Figure 6 (b) was the super-resolution mass spectrometry data obtained through the SCSR-Unet network, Figure 6 (c) was the low-resolution mass spectrometry data collected in the Turbo mode. Under the three conditions, the parameters of t-SNE were all set to a learning rate of 50 and a perplexity of 10. It can be seen from the figure that after super-resolution processing, the discrimination effect of the three bladder cancer cell subtypes was significantly improved, the distance between different groups was significantly greater than the distance within the group, which was close to the discrimination effect of cell subtypes in the Normal mode, while the original data in the Turbo mode showed a weak discrimination ability. Here, the Calinski-Harabasz Index (CH index), Davies-Bouldin Index (DB index), and Silhouette Coefficient (SC) were used to evaluate the clustering effect, as shown in Table 5. Among them, the values of the three indexes after super-resolution were all higher than those in the Turbo mode. To further verify this speculation, at the same time, the clustering effect was visualized in another way, and UMAP was used for dimensionality reduction analysis. Figure 6(d), (e), and (f) respectively correspond to the UMAP clustering results of the high-resolution mass spectrometry data collected in the Normal mode, the super-resolution mass spectrometry data obtained through the SCSR-Unet network, and the low-resolution mass spectrometry data collected in the Turbo mode. The parameter settings are n_neighbors = 8 and min_dist = 0.05. The results show that, consistent with the visualization of t-SNE, the data after super-resolution have better performance in distinguishing different bladder cancer cell subtypes. Similarly, the values of the three indicators after super-resolution are all higher than those in the Turbo mode.

[0042] Table 5 shows the dimensionality reduction results based on t-SNE and UMAP.

[0043]

[0044] (A43) Cell typing. To evaluate the classification effect of different subtypes of bladder cancer cells under three conditions, a relatively mature machine learning method in metabolomics research - random forest (RF) - was used to perform phenotypic classification on the preprocessed single-cell data sets of each group and summarize the results. In the RF algorithm, the construction of the decision tree is the core of the model performance, and the number of decision trees (n_estimators) and depth (max_depth) are the key parameters affecting the running efficiency and classification effect of RF. Here, to make a fair comparison of the three data sets, the same model architecture and hyperparameters were used. The parameters n_estimators = 100 and max_depth = 10 were set. In addition, all three models used an 8:2 ratio to divide the training set and the test set.

[0045] Figure 7 The classification results of the test sets under three conditions are shown. Figure 7 (a), (b), and (c) respectively correspond to the confusion matrices in the Normal mode, after super-resolution, and in the Turbo mode, showing the matching degree between the predicted categories and the actual categories of the model. The diagonal elements of the confusion matrix represent the number of correctly classified samples, and the non-diagonal elements represent the distribution of misclassified samples. In addition, Figure 7(d), (e), and (f) respectively show the ROC curves under three modes to evaluate the classification performance of the model at different thresholds. The closer the ROC curve is to the upper left corner, the stronger the discrimination ability of the model. As can be seen from the figure, the classification effect of the super-resolution mode is the most prominent, close to the classification performance of the Normal mode, while the classification effect of the Turbo mode is slightly inferior to that of the super-resolution mode. To better evaluate the classification accuracy of the model, as shown in Table 6, the classification results based on RF are presented. The accuracy, precision, recall, and F1 score values (all 0.99) after super-resolution are greater than the corresponding index results in the Turbo case (>0.87) and are similar to those in the Normal case (>0.98). Thus, it can be seen that the super-resolution model can greatly improve the single-cell classification effect.

[0046] Table 6 shows the classification results based on RF.

[0047] Normal SCSR Turbo Accuracy 0.98 0.99 0.87 Precision 0.99 0.99 0.88 Recall 0.98 0.99 0.87 F1-score 0.98 0.99 0.87 (A44)Analysis of metabolite heterogeneity of bladder cancer cell subtypes. In metabolomics research, screening for potential differential metabolites is a crucial step in data analysis. To verify whether the super-resolution model is helpful for the discovery of differential metabolites, 12 potential amino acid differential metabolites detected in the bladder cancer population cells were used to draw a heat map of single-cell data to visually display the relative intensity of each single-cell metabolite ion. Figure 8 (a), (b), and (c) successively present the heat map results of single-cell data in the Normal mode, after super-resolution, and in the Turbo mode. The heat map result after super-resolution is particularly crucial. It clearly shows that the single-cell metabolic profiles of the same cell subtype exhibit high consistency, while significant differences are shown between different subtypes. Moreover, its distribution characteristics are relatively close to those in the Normal mode, indicating that the super-resolution model has good performance in retaining cell metabolic characteristics. Taking arginine (m / z->175) as an example, in the heat map after super-resolution, its intensity is the highest in 5637-epi cells, followed by UMUC3 cells, and the lowest in 5637 cells. This clear differential distribution is of great significance in metabolomics research and can help researchers quickly locate the expression differences of key metabolites in different cell subtypes. However, in the Turbo mode, this difference is not significant, further highlighting the advantage of the super-resolution model in enhancing the differential expression of metabolites.

[0048] Figure 8 (d)shows the secondary mass spectrum of arginine (m / z->175) in single-cell detection. Daughter ions 116 and 158 can be detected from it, providing more in-depth mass spectrometry information for the metabolic feature analysis of arginine. Figure 8(e) and (f) show the first-order mass spectra in the range of 174 to 177 in the Turbo and SR cases. It can be clearly seen from the figure that after super-resolution, the metabolite ion arginine (m / z -> 175) can be clearly separated, which strongly proves that the super-resolution model can effectively improve the mass spectrometry resolution.

Claims

1. A method for single-cell metabolomics analysis based on mass spectrometry technology, characterized in that, It includes the following steps: (A1)Collect population cell data, including multiple average mass spectra of various population cells in low-resolution mode and high-resolution mode; (A2)Construct an SCSR-Unet super-resolution model using the population cell data, and perform model training, validation, and testing; In network construction, the loss function CL used is: ; output i is the prediction result of the sample, and label i is the label corresponding to the sample, i represents the i-th sample, α is the hyperparameter controlling the loss weight, and N is the number of samples; (A3)Collect single-cell data, including average mass spectrometry data of various single cells in low-resolution mode and high-resolution mode; (A4)The average mass spectrometry data in low-resolution mode is passed through the SCSR-Unet model to obtain predicted high-resolution mass spectrometry data of absolute intensity, which is preprocessed synchronously with the average mass spectrometry data in high-resolution mode for single-cell metabolomics analysis.

2. The single-cell metabolomics analysis method according to claim 1, characterized in that, In step (A2), the method of constructing the SCSR-Unet super-resolution model includes the steps of: (A21)Data alignment, perform difference on the average mass spectrometry data in low-resolution mode to obtain equidistant mass-to-charge ratio data identical to that of the average mass spectrometry data in high-resolution mode; (A22)Data encoding, intercept multiple segments of equal length in the average mass spectrometry data in low-resolution and high-resolution modes, arrange the segments in order of mass-to-charge ratio to form a two-dimensional matrix, and generate training pairs after normalization. The low-resolution segments are used as input data, and the high-resolution segments are used as target labels; (A23)Network construction, use a two-dimensional U-Net network structure, and construct an SCSR-Unet super-resolution model using the loss function; (A24)Input the average mass spectrometry data in low-resolution mode to be tested into the trained SCSR-Unet super-resolution model to obtain predicted high-resolution mass spectrometry data; (A25)Restore the predicted high-resolution mass spectrometry data according to the normalization scale in (A22), and splice the restored data segments in order of mass-to-charge ratio to generate the predicted high-resolution mass spectrometry data of absolute intensity.

3. The single-cell metabolomics analysis method according to claim 2, wherein In the U-Net network structure, a skip connection is added between the upsampling layer and the downsampling layer to directly transfer the underlying features extracted during the downsampling process to the corresponding upsampling layer.

4. The single-cell metabolomics analysis method according to claim 1, wherein In step (A4), the predicted high-resolution mass spectrometry data of absolute intensity and the average mass spectrometry data in high-resolution mode are used as input-end mass spectrometry data, and data preprocessing, visualization analysis, cell typing, and differential metabolites steps are taken for the input-end mass spectrometry data.

5. The single-cell metabolomics analysis method according to claim 4, characterized in that The data preprocessing is as follows: Extract the identified mass spectrometry peaks and their corresponding ion intensities from the input-end mass spectrometry data to generate a metabolite peak list; Noise reduction processing; Remove the background signal of the sampling solvent; Use the "80% rule" and the K-nearest neighbor algorithm to process missing values to obtain single-cell feature data to be analyzed.

6. The single-cell metabolomics analysis method according to claim 5, characterized in that The visualization analysis is as follows: Perform dimensionality reduction on the single-cell feature data to be analyzed, and obtain high-dimensional data in a low-dimensional space through multivariate analysis.

7. The single-cell metabolomics analysis method according to claim 6, wherein The method of dimensionality reduction is: t-distributed stochastic neighbor embedding and uniform manifold approximation and projection.

8. The single-cell metabolomics analysis method according to claim 6, wherein The cell typing uses the random forest method to distinguish the high-dimensional data in the low-dimensional space of different subtypes, and the ROC curve is used to evaluate the performance of the random forest model.

9. The single-cell metabolomics analysis method according to claim 4, wherein In the discovery of the differential metabolites, a heat map is used to display the relative intensities of potential metabolite ions in each single cell; The metabolites in the heat map are sorted according to hierarchical clustering, and the horizontal sample clustering shows the differences in metabolite abundances among single cells of different subtypes.

10. The single-cell metabolomics analysis method according to claim 1, characterized in that, The population cells include 5637, 5637-epi, and UMUC3 bladder cancer single cells.

Citation Information

Patent Citations

  • Single cell on-line extended pyrolysis mass spectrometry flow analysis method

    CN116429667A

  • Image super-division method and image data processing method

    CN116883236A

  • Heterogeneity-based cell metabolism network modeling method and application thereof

    CN117476092A

  • Zero sample learning-based mass spectrum image super-resolution reconstruction method

    CN118014843A

  • Mass spectrum imaging space super-resolution reconstruction method and system based on label propagation network

    CN119323519A

Cited By

  • A system and method for deep analysis of same-piece bipolar single-cell metabolomics

    CN122551882A