Mass spectrometry-based quality control method and system
By employing quality control methods for mass spectrometry flow cytometry data, the problems of batch effects and signal drift in CyTOF data were solved, enabling systematic evaluation and optimization of data quality. This improved the accuracy and reliability of the analysis, especially in high-dimensional single-cell analysis, ensuring the accuracy of staining effects and the credibility of experimental results.
Patent Information
- Application Number
- CN202510515691.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-04-23
AI Technical Summary
CyTOF data are subject to batch effects and signal drift due to differences in experimental conditions, instrument detection, and sample processing, which affects the accuracy of data analysis.
This paper presents a quality control method based on mass spectrometry flow cytometry data, including data reading, signal expression distribution analysis, dispersion assessment, batch difference assessment, cluster analysis, and staining effect assessment. By calculating indicators such as MMD, Wasserstein Distance, and KS test statistic, the method identifies and quantifies signal differences between batches, and uses the flowsom clustering algorithm to optimize data quality.
It improves data quality and consistency, enhances the reliability and repeatability of analysis, ensures comparability and consistency between different batches of data, improves the accuracy of staining effect assessment, and reduces errors caused by batch effects.
Smart Images

Figure CN120432012B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical technology, specifically to a quality control method and system based on mass spectrometry flow cytometry data. Background Technology
[0002] CyTOF (Cytometry by Time-of-Flight) is a high-dimensional single-cell analysis technique based on mass spectrometry, widely used in fields such as immunology, cancer research, and systems biology. Compared to traditional flow cytometry (FACS), CyTOF technology can use metal isotope-labeled antibodies, significantly reducing fluorescence overlap and allowing the number of detection channels to be expanded to 40-50 or even more, thus enabling higher-dimensional single-cell phenotypic analysis.
[0003] However, due to variations in experimental conditions, instrument detection, and sample processing, CyTOF data may suffer from batch effects or signal drift, affecting the accuracy of data analysis. Therefore, an efficient and systematic quality control method is needed to evaluate and optimize CyTOF data. Summary of the Invention
[0004] This invention addresses the shortcomings of existing technologies by providing a quality control method and system based on mass spectrometry flow cytometry data.
[0005] In a first aspect, the present invention provides a quality control method based on mass spectrometry flow cytometry data, comprising the following steps:
[0006] Data Reading: Read each FCS file in the specified folder;
[0007] Signal representation distribution analysis: For each FCS file, histograms and cumulative probability density plots are plotted for each channel to evaluate the consistency of signal representation distribution;
[0008] Discreteness assessment: Calculate the absolute median deviation of each channel in each data point to determine the degree of dispersion of the channel signal values;
[0009] Batch difference assessment: Calculate the maximum mean difference of each channel, EMD distance, and KS test statistic between two batches to determine the signal difference between batches;
[0010] Cluster analysis: Clustering is performed using the Flowsom clustering algorithm. Initially, several subclusters are formed, and then the consistency clustering method is used to merge them into a set number of subclusters.
[0011] Staining effect evaluation: The average overlap frequency value of the channels is calculated in each data set, and the minimum value of all subgroups is taken as the final average overlap frequency value to determine whether the staining effect is good.
[0012] Secondly, the present invention provides a quality control system based on mass spectrometry flow cytometry data, the system comprising:
[0013] The data reading module is used to read each FCS file in a specified folder;
[0014] The signal representation distribution analysis module is used to plot histograms and cumulative probability density maps for each channel for each FCS file to evaluate the consistency of the signal representation distribution.
[0015] The dispersion assessment module is used to calculate the absolute median deviation of each channel in each data point to determine the dispersion of the channel signal values.
[0016] The batch difference assessment module is used to calculate the maximum mean difference of each channel between two batches, the EMD distance, and the KS test statistic to determine the signal differences between batches.
[0017] The clustering analysis module uses the Flowsom clustering algorithm to perform clustering. Initially, several subclusters are formed, and then the consistency clustering method is used to merge them into a set number of subclusters.
[0018] The staining effect evaluation module is used to calculate the average overlap frequency value of the channels in each data set, and take the minimum value of all subgroups as the final average overlap frequency value to determine whether the staining effect is good.
[0019] Thirdly, the present invention provides a computer-readable storage medium storing multiple instructions adapted for loading and execution by a processor, such as a quality control method based on mass spectrometry streaming data.
[0020] Fourthly, the present invention provides an electronic device including a processor, a memory, and a transceiver. The memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to enable the electronic device to perform a quality control method based on mass spectrometry streaming data.
[0021] The beneficial effects of this invention are:
[0022] Improving data quality and consistency: Through a systematic quality control process, including signal expression distribution analysis, dispersion assessment, batch difference assessment, cluster analysis, and staining effect evaluation, the accuracy and consistency of mass spectrometry flow cytometry data can be comprehensively evaluated and optimized. In particular, calculating indicators such as MMD, Wasserstein Distance, and KSStat can effectively identify and quantify signal differences between batches, thereby providing a high-quality data foundation for subsequent data analysis and reducing errors caused by batch effects.
[0023] Enhancing the reliability and reproducibility of the analysis: This method ensures comparability and consistency between different batches of data through standardized data preprocessing (such as arcsinh transformation and zscale transformation) and systematic quality control steps. This not only improves the reliability of individual experiments but also enhances data reproducibility across experiments and laboratories, making the research results more credible and facilitating comparison and validation between different studies.
[0024] Improving the accuracy of staining effect assessment: This method can accurately assess staining effect by calculating the AOF (Average Overlap Frequency) value. The smaller the AOF value, the higher the separation between negative and positive peaks, and the better the staining effect. This indicator provides researchers with a tool to quantitatively assess staining quality, which helps to optimize experimental conditions, improve the success rate of experiments and the availability of data. Especially in high-dimensional single-cell analysis, good staining effect is the key to obtaining reliable data. Attached Figure Description
[0025] Figure 1 This is a flowchart illustrating a quality control method based on mass spectrometry flow cytometry data in an embodiment of the present invention.
[0026] Figure 2 Histograms and cumulative probability density plots of the signal representation distribution on CD3 for two batches of data;
[0027] Figure 3 The MAD value for each channel (displayed by marker name) in each data set;
[0028] Figure 4 The difference in the proportion of each subgroup across different batches. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] like Figure 1 As shown in the embodiments of this application, a quality control method based on mass spectrometry flow cytometry data is provided, the method comprising:
[0031] Step 1: Read each fcs file in the specified folder.
[0032] Step 2: Based on Step 1, draw the histogram and cumulative probability density map (CDF) for each channel in each fcs file.
[0033] Step 3: Based on Step 1, calculate the MAD value (median absolute deviation) for each channel in each data set. i -median(x)|)) determines the degree of dispersion of the signal values in the channel.
[0034] Step 4: Based on Step 1, calculate the MMD (Maximum Mean Difference), Wasserstein Distance (EMD), and KSStat (KS test statistic) for each channel between the two batches to determine the signal differences between the batches.
[0035] Step 5: Based on Step 1, use the Flowsom clustering algorithm to perform clustering, initially creating 100 subclusters, and then use the consistency clustering method to merge them into 30 subclusters.
[0036] Step 6: Based on step 5, calculate the AOF (average overlap frequency) value of each channel in each data set. All subpopulations of each data set should be calculated once, and then the minimum value should be taken as the final AOF to determine whether the staining effect is good.
[0037] In some embodiments, step 1 specifically involves reading the fcs file, reading the fcs file in the given folder according to the given folder and metadata, and reading the batch information in the metadata, wherein the batch information is used in subsequent steps to calculate MMD, WassersteinDistance, KSStat indicators and calculate the difference in the proportion of subgroups between batches, and reading the given panel file to record the channels to be analyzed and the corresponding markers.
[0038] Furthermore, if a metadata file is specified, it will be read according to the file name specified in the metadata. The metadata must contain two columns: file and batch. File is the name of the fcs file, and batch is the batch information corresponding to each file. If no metadata is specified, each piece of data is treated as a batch.
[0039] Furthermore, the panel file must contain two columns: channel and marker.
[0040] Furthermore, the metadata file and panel file can be in CSV or Excel format.
[0041] In some embodiments, step 2 specifically involves plotting histograms and cumulative probability density maps (CDFs) of the signal representation for each channel pairwise across all batches. If there are multiple input data sets, then all data sets are plotted pairwise (or between two batches if batch information for the samples is specified) to determine the consistency of the distribution of the two data sets in this channel.
[0042] Furthermore, before plotting, the signal representation values were subjected to an Arcsinh transformation, and the cofactor was set to 5. The histogram and cumulative probability density plot show the distribution of signal intensity in the channel. The overlap between two batches of data in the plot indicates whether the signals between batches in this channel are relatively consistent. (See...) Figure 2 .
[0043] In some embodiments, step 3 specifically involves: calculating the MAD value of each channel in each data set (fcs file) to determine the dispersion of the channel's signal values. If there are multiple data sets, a bar chart is also plotted, see [link to relevant documentation]. Figure 3 .
[0044] Furthermore, the signal representation value is transformed using Arcsinh before calculating MAD, and the cofactor is set to 5. The MAD value is used to measure the dispersion of the channel signal; the larger the MAD, the more dispersed the channel signal, and vice versa.
[0045] In some embodiments, step 4 specifically involves calculating the MMD, Wasserstein Distance, and KSStat for each channel between two batches to determine the signal difference between the batches. If only one data point is available, this step is skipped.
[0046] Furthermore, the signal representation values were converted using Arcsinh before calculation, and the cofactor was set to 5. These metrics are used to measure the distance between channel signal values in batches. The larger their values, the greater the distance, indicating less consistency in signal distribution between batches and a more severe batch effect, which needs to be removed in subsequent analysis.
[0047] In some embodiments, step 5 specifically involves: performing clustering using flowsom, and before clustering, performing an arcsinh transformation with cofactor set to 5.
[0048] Furthermore, for each batch of data, a z-scale transformation is performed on each channel to be clustered to reduce batch effects. If each batch contains multiple samples, a rank-sum test is performed on the batch based on the proportion of each subgroup in each data set, and a box plot is generated. (See [link to relevant documentation]). Figure 4Box plots show the proportion of each subpopulation of cells in each sample within each batch, and statistical significance is demonstrated based on the rank-sum test results. The more asterisks (*) marking the significance, the more statistically significant the difference in subpopulation proportions between batches.
[0049] In some embodiments, step 6 specifically involves calculating the AOF (Average Overlap Frequency) value for each channel in each data set based on the clustering results of step 5. A smaller AOF value indicates a more separated negative and positive peak in the channel's signal value, with less overlap between the two peaks. This value can be used to determine the effectiveness of the staining; a larger AOF value indicates a poor staining effect.
[0050] This embodiment utilizes a variety of statistical methods to efficiently and accurately assess the quality of different batches of CyTOF data, identify potential systematic errors or batch effects, and provide optimization suggestions, thereby improving the reliability of single-cell proteomics research.
[0051] Based on the same concept as the above method, this application also provides a quality control system based on mass spectrometry flow cytometry data, the system comprising:
[0052] The data reading module is used to read each FCS file in a specified folder;
[0053] The signal representation distribution analysis module is used to plot histograms and cumulative probability density maps for each channel for each FCS file to evaluate the consistency of the signal representation distribution.
[0054] The dispersion assessment module is used to calculate the absolute median deviation of each channel in each data point to determine the dispersion of the channel signal values.
[0055] The batch difference assessment module is used to calculate the maximum mean difference of each channel between two batches, the EMD distance, and the KS test statistic to determine the signal differences between batches.
[0056] The clustering analysis module uses the Flowsom clustering algorithm to perform clustering. Initially, several subclusters are formed, and then the consistency clustering method is used to merge them into a set number of subclusters.
[0057] The staining effect evaluation module is used to calculate the average overlap frequency value of the channels in each data set, and take the minimum value of all subgroups as the final average overlap frequency value to determine whether the staining effect is good.
[0058] Based on the same concept as the above method, this application also provides an electronic device, including a processor, a memory, and a transceiver. The memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to enable the electronic device to perform a quality control method based on mass spectrometry streaming data.
[0059] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory, random access memory, magnetic disks, or optical disks.
[0060] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.
[0061] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A quality control method based on mass spectrometry flow cytometry data, characterized in that, Includes the following steps: Data Reading: Read each FCS file in the specified folder; Signal representation distribution analysis: For each FCS file, histograms and cumulative probability density plots are plotted for each channel to evaluate the consistency of signal representation distribution; Discreteness assessment: Calculate the absolute median deviation of each channel in each data point to determine the degree of dispersion of the channel signal values; Batch difference assessment: Calculate the maximum mean difference of each channel, EMD distance, and KS test statistic between two batches to determine the signal difference between batches; Cluster analysis: Clustering is performed using the Flowsom clustering algorithm. Initially, several subclusters are formed, and then the consistency clustering method is used to merge them into a set number of subclusters. Staining effect evaluation: The average overlap frequency value of the channels is calculated in each data set, and the minimum value of all subgroups is taken as the final average overlap frequency value to determine whether the staining effect is good.
2. The quality control method based on mass spectrometry flow cytometry data according to claim 1, characterized in that, In the data reading step: If a metadata file is specified, it is read according to the file name specified in the metadata file, which contains the file name and batch information; Based on the input panel file, record the channels and corresponding markers used for subsequent analysis. The panel file contains two columns: channels and markers. If no metadata file is specified, each piece of data is processed as a separate batch.
3. The quality control method based on mass spectrometry flow cytometry data according to claim 1 or 2, characterized in that, In the signal expression distribution analysis step: Before plotting the histogram and cumulative probability density plot, the signal expression values are transformed using arcsinh, and the cofactor is set to 5. For multiple datasets, all datasets are plotted pairwise. If batch information for the samples is specified, plotting is performed between two batches.
4. The quality control method based on mass spectrometry flow cytometry data according to claim 1, characterized in that, In the dispersion assessment step: Before calculating the absolute median deviation, the signal representation values are subjected to an arcsinh transformation. Set cofactor to 5; If there are multiple data points, plot a bar chart of MAD values to visually demonstrate the dispersion of signals in each channel.
5. The quality control method based on mass spectrometry flow cytometry data according to claim 1 or 2, characterized in that, In the batch difference assessment step: Before calculating the maximum mean difference, EMD distance, and KS test statistic, the signal expression values were transformed using arcsinh, and the cofactor was set to 5.
6. The quality control method based on mass spectrometry flow cytometry data according to claim 1, characterized in that, In the cluster analysis step: Before clustering, each batch of data is transformed using the arcsinh transformation, with cofactor set to 5; For each batch of data, a zscale transformation is performed on each channel to be clustered to reduce batch effects.
7. The quality control method based on mass spectrometry flow cytometry data according to claim 6, characterized in that, If there are multiple samples in each batch, a rank-sum test is performed on the batch based on the proportion of each subgroup in each data point, and a box plot is drawn to show the differences in the proportion of subgroups between batches.
8. A quality control system based on mass spectrometry flow cytometry data, characterized in that, The system includes: The data reading module is used to read each FCS file in a specified folder; The signal representation distribution analysis module is used to plot histograms and cumulative probability density maps for each channel for each FCS file to evaluate the consistency of the signal representation distribution. The dispersion assessment module is used to calculate the absolute median deviation of each channel in each data point to determine the dispersion of the channel signal values. The batch difference assessment module is used to calculate the maximum mean difference of each channel between two batches, the EMD distance, and the KS test statistic to determine the signal differences between batches. The clustering analysis module uses the Flowsom clustering algorithm to perform clustering. Initially, several subclusters are formed, and then the consistency clustering method is used to merge them into a set number of subclusters. The staining effect evaluation module is used to calculate the average overlap frequency value of the channels in each data set, and take the minimum value of all subgroups as the final average overlap frequency value to determine whether the staining effect is good.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions that are adapted to be loaded by a processor and executed as described in any one of claims 1 to 7.
10. An electronic device, characterized in that, It includes a processor, a memory, and a transceiver. The memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as claimed in any one of claims 1 to 7.
Citation Information
Patent Citations
Big-data analyzing method and mass spectrometric system using the same method
US20170358434A1
Automatic ion population control for charge detection mass spectrometry
WO2023076583A1