Quality control method and system based on mass spectrum flow data
Through the quality control method of mass spectrometry flow data, the problems of batch effect and signal drift in CyTOF data are solved. Through evaluation and cluster analysis of multiple statistical indicators, the quality of CyTOF data is optimized, and the accuracy and reliability of data analysis are improved.
Patent Information
- Application Number
- CN202510515691.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-04-23
AI Technical Summary
Due to differences in experimental conditions, instrument detection and sample processing, there are batch effects and signal drifts in CyTOF data, which affects the accuracy of data analysis.
The quality control method based on mass spectrometry flow data is adopted, including signal expression distribution analysis, degree of discreteness evaluation, batch difference evaluation, cluster analysis and staining effect evaluation. By calculating indicators such as MMD, Wasserstein Distance, KS test statistics and AOF values, the signal difference between batches is identified and quantified, and the data quality is optimized.
It improves the accuracy and consistency of CyTOF data, reduces errors caused by batch effects, enhances the reliability and repeatability of data analysis, and improves the success rate and credibility of single-cell proteomics research.
Smart Images

Figure CN120432012A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical technology, and in particular to a quality control method and system based on mass spectrometry flow data. Background Art
[0002] Cytometry by Time-of-Flight (CyTOF) is a high-dimensional single-cell analysis technology based on mass spectrometry, widely used in fields such as immunology, cancer research, and systems biology. Compared to traditional flow cytometry (FACS), CyTOF technology can use metal isotope-labeled antibodies, significantly reducing fluorescence overlap and expanding the number of detection channels to 40-50 or even more, thereby enabling higher-dimensional single-cell phenotyping.
[0003] However, due to differences in experimental conditions, instrument detection, and sample processing, CyTOF data may exhibit batch effects or signal drift, affecting the accuracy of data analysis. Therefore, an efficient and systematic quality control method is needed to evaluate and optimize CyTOF data. Summary of the Invention
[0004] In view of the deficiencies in the prior art, the present invention provides a quality control method and system based on mass spectrometry flow data.
[0005] In a first aspect, the present invention provides a quality control method based on mass spectrometry flow data, comprising the following steps:
[0006] Data reading: read each FCS file in the specified folder;
[0007] Signal expression distribution analysis: For each FCS file, a histogram and cumulative probability density plot of each channel were drawn to evaluate the consistency of signal expression distribution;
[0008] Discreteness evaluation: Calculate the absolute median deviation of each channel in each data to determine the discreteness of the channel signal value;
[0009] Batch difference evaluation: Calculate the maximum mean difference, EMD ground motion distance and KS test statistic of each channel between two batches to determine the signal difference between batches;
[0010] Cluster analysis: Flowsom clustering algorithm was used for clustering. Several subgroups were initially clustered and then merged into a set number of subgroups using the consensus clustering method.
[0011] Staining effect evaluation: Calculate the average overlapping frequency value of each channel in each data, and take the minimum value of all subpopulations as the final average overlapping frequency value to determine whether the staining effect is good.
[0012] In a second aspect, the present invention provides a quality control system based on mass spectrometry flow data, the system comprising:
[0013] The data reading module is used to read each FCS file in the folder according to the specified folder;
[0014] Signal expression distribution analysis module, used to draw the histogram and cumulative probability density diagram of each channel for each FCS file to evaluate the consistency of signal expression distribution;
[0015] The discrete degree evaluation module is used to calculate the absolute median deviation of each channel in each data to determine the discrete degree of the channel signal value;
[0016] Batch difference evaluation module, which is used to calculate the maximum mean difference, EMD ground motion distance and KS test statistic of each channel between two batches to determine the signal difference between batches;
[0017] Cluster analysis module, using flowsom clustering algorithm for clustering, initially clustering several subgroups, and then using the consensus clustering method to merge into a set number of subgroups;
[0018] The staining effect evaluation module is used to calculate the average overlapping frequency value of the channel in each data, and take the minimum value of all subpopulations as the final average overlapping frequency value to determine whether the staining effect is good.
[0019] In a third aspect, the present invention provides a computer-readable storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded and executed by a processor, such as a quality control method based on mass spectrometry flow data.
[0020] In a fourth aspect, the present invention provides an electronic device comprising a processor, a memory and a transceiver, wherein the memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs a quality control method based on mass spectrometry flow data.
[0021] Beneficial effects of the present invention:
[0022] Improve data quality and consistency: Through a systematic quality control process, including signal expression distribution analysis, dispersion assessment, batch variation assessment, cluster analysis, and staining effect evaluation, the accuracy and consistency of mass spectrometry flow cytometry data can be comprehensively evaluated and optimized. In particular, the calculation of metrics such as MMD, Wasserstein Distance, and KSStat can effectively identify and quantify signal differences between batches, providing a high-quality data foundation for subsequent data analysis and reducing errors caused by batch effects.
[0023] Enhanced reliability and reproducibility of analysis: This method ensures comparability and consistency between different batches of data through standardized data preprocessing (such as ArcSinh transformation and ZScale transformation) and systematic quality control steps. This not only improves the reliability of a single experiment, but also enhances the reproducibility of data across experiments and laboratories, making research results more credible and facilitating comparison and verification between different studies.
[0024] Improving the accuracy of staining effect assessment: This method accurately evaluates staining quality by calculating the average overlap frequency (AOF). A smaller AOF value indicates greater separation between negative and positive peaks, indicating better staining quality. This metric provides a tool for quantitatively assessing staining quality, helping to optimize experimental conditions, improve experimental success rates, and enhance data usability. This is particularly true in high-dimensional single-cell analysis, where good staining is crucial for obtaining reliable data. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 Schematic diagram of a quality control method based on mass spectrometry data in an embodiment of the present invention;
[0026] Figure 2 The histogram and cumulative probability density plot of the signal expression distribution on CD3 for two batches of data;
[0027] Figure 3 The MAD value of each channel (displayed by marker name) in each data;
[0028] Figure 4 The difference in the proportion of each subpopulation between different batches. DETAILED DESCRIPTION
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0030] like Figure 1 As shown, the embodiment of the present application provides a quality control method based on mass spectrometry flow data, the method comprising:
[0031] Step 1: According to the specified folder, read each fcs file in the folder.
[0032] Step 2: Based on step 1, draw the histogram (histgram) and cumulative probability density plot (CDF) of each channel for each fcs file.
[0033] Step 3: Based on step 1, calculate the MAD value (absolute median deviation median(|x i -median(x)|)) determines the discreteness of the signal value of the channel.
[0034] Step 4: Based on step 1, calculate the MMD (maximum mean difference), WassersteinDistance (EMD ground motion distance), and KSStat (KS test statistic) of each channel between the two batches to determine the signal difference between the batches.
[0035] Step 5: Based on step 1, the flowsom clustering algorithm was used for clustering. 100 subgroups were initially clustered and then merged into 30 subgroups using the consensus clustering method.
[0036] Step 6: Based on step 5, calculate the AOF (average overlap frequency) value of the channel in each data. All subpopulations of each data should be calculated once, and then take the minimum value as the final AOF to determine whether the staining effect is good.
[0037] In some embodiments, step 1 specifically reads the fcs file, reads the fcs file in the folder according to the given folder and meta information, and reads the batch information in the meta information, wherein the batch information is used in subsequent steps to calculate MMD, WassersteinDistance, KSStat indicators and calculate the differences in subpopulation proportions between batches, reads the given panel file, and records the channels to be analyzed and the corresponding markers.
[0038] Furthermore, if the meta information file is specified, the data is read according to the file name specified in the meta information. The meta information must contain two columns: file and batch, where file is the name of the fcs file and batch is the batch information corresponding to each file. If no meta information is specified, each data is treated as a batch.
[0039] Furthermore, the panel file must contain two columns: channel and marker.
[0040] Furthermore, the meta information file and the panel file can be in CSV format or Excel format.
[0041] In certain embodiments, step 2 specifically involves plotting histograms and cumulative probability density plots (CDFs) of the signal expression for each channel between all pairs of batches. If there are multiple input data sets, all data sets are plotted pairwise (or between two batches if sample batch information is specified) to identify the consistency of the expression distribution of the two data sets on this channel.
[0042] Furthermore, the signal expression values were arcsinh transformed before plotting, and the cofactor was set to 5. The histogram and cumulative probability density plot show the distribution of signal intensity in the channel. The overlap of the two batches of data in the figure can be used to determine whether the signals between batches in this channel are relatively consistent, see Figure 2 .
[0043] In some embodiments, the step 3 is specifically: calculating the MAD value of each channel in each data (fcs file) to determine the discrete degree of the signal value of the channel. If there are multiple data, a histogram is also drawn, see Figure 3 .
[0044] Furthermore, the signal expression values were arcsinh transformed before calculating the MAD, and the cofactor was set to 5. The MAD value is used to measure the discreteness of the channel signal. The larger the MAD, the more discrete the channel signal, and vice versa.
[0045] In some embodiments, step 4 specifically involves calculating the MMD, WassersteinDistance, and KSStat indicators for each channel between the two batches to determine the signal difference between the batches. If there is only one data point, this step is skipped.
[0046] Furthermore, the signal expression values were arcsinh transformed before calculation, and the cofactor was set to 5. These indicators are used to measure the distance between channel signal values between batches. The larger their values, the greater the distance, which also indicates that the consistency of signal distribution between batches is less, and the batch effect is more serious. Subsequent analysis needs to remove the batch effect.
[0047] In some embodiments, the step 5 specifically includes: using flowsom to perform clustering, and before clustering, first performing arcsinh transformation, and setting the cofactor to 5.
[0048] Furthermore, for each batch of data, perform a zscale transformation on each channel to be clustered to reduce the batch effect. If there are multiple samples in each batch, perform a rank sum test on the batch based on the proportion of each subgroup in each data, and draw a box plot, see Figure 4The box plot shows the proportion of each subpopulation cell number in each sample in each batch to the total cell number of the sample, and shows the statistical significance based on the results of the rank sum test. The more asterisks (*) marked with significance, the more statistically significant the differences in subpopulation proportions between batches are.
[0049] In certain embodiments, step 6 specifically comprises calculating the AOF (average overlap frequency) value for each channel in each data set based on the clustering results of step 5. A smaller AOF value indicates that the negative and positive peaks of the signal value in that channel are more separated, and the overlap between the two peaks is smaller. This value can be used to determine whether the staining effect is good; a larger AOF value indicates that the staining effect is poor.
[0050] This example uses a combination of statistical methods to efficiently and accurately evaluate the quality of different batches of CyTOF data, identify potential systematic errors or batch effects, and provide optimization suggestions, thereby improving the reliability of single-cell proteomics research.
[0051] Based on the same concept as the above method, an embodiment of the present application further provides a quality control system based on mass spectrometry flow data, the system comprising:
[0052] The data reading module is used to read each FCS file in the folder according to the specified folder;
[0053] Signal expression distribution analysis module, used to draw the histogram and cumulative probability density diagram of each channel for each FCS file to evaluate the consistency of signal expression distribution;
[0054] The discrete degree evaluation module is used to calculate the absolute median deviation of each channel in each data to determine the discrete degree of the channel signal value;
[0055] Batch difference evaluation module, which is used to calculate the maximum mean difference, EMD ground motion distance and KS test statistic of each channel between two batches to determine the signal difference between batches;
[0056] Cluster analysis module, using flowsom clustering algorithm for clustering, initially clustering several subgroups, and then using the consensus clustering method to merge into a set number of subgroups;
[0057] The staining effect evaluation module is used to calculate the average overlapping frequency value of the channel in each data, and take the minimum value of all subpopulations as the final average overlapping frequency value to determine whether the staining effect is good.
[0058] Based on the same concept as the above method, an embodiment of the present application also provides an electronic device, including a processor, a memory and a transceiver, the memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs a quality control method based on mass spectrometry flow data.
[0059] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiments; and the aforementioned storage medium includes: mobile storage devices, read-only memories, random access memories, magnetic disks or optical disks, and other media that can store program codes.
[0060] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0061] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A quality control method based on mass spectrometry data, characterized in that: The following steps are involved: Data reading: read each FCS file in the specified folder; Signal expression distribution analysis: For each FCS file, a histogram and cumulative probability density plot of each channel were drawn to evaluate the consistency of signal expression distribution; Discreteness evaluation: Calculate the absolute median deviation of each channel in each data to determine the discreteness of the channel signal value; Batch difference evaluation: Calculate the maximum mean difference, EMD ground motion distance and KS test statistic of each channel between two batches to determine the signal difference between batches; Cluster analysis: Flowsom clustering algorithm was used for clustering. Several subgroups were initially clustered and then merged into a set number of subgroups using the consensus clustering method. Staining effect evaluation: Calculate the average overlapping frequency value of each channel in each data, and take the minimum value of all subpopulations as the final average overlapping frequency value to determine whether the staining effect is good.
2. The quality control method based on mass spectrometry data according to claim 1, characterized in that: In the data reading step: If a meta information file is specified, the file is read according to the file name specified in the meta information, and the meta information file includes the file name and batch information; Record the channels and corresponding markers for subsequent analysis based on the incoming panel file. The panel file contains two columns: channel and marker. If no metadata file is specified, each data is processed as a separate batch.
3. The quality control method based on mass spectrometry data according to claim 1 or 2, characterized in that: In the signal expression distribution analysis step: Before drawing the histogram and cumulative probability density plot, the signal expression values were arcsinh transformed, and the cofactor was set to 5; For multiple data, all data are plotted between each other. If the batch information of the sample is specified, it is plotted between two batches.
4. The quality control method based on mass spectrometry data according to claim 1, characterized in that: In the discrete degree evaluation step: Before calculating the absolute median deviation value, the signal expression value was arcsinh transformed. cofactor is set to 5; If there are multiple data, draw a histogram of the MAD value to intuitively show the degree of discreteness of each channel signal.
5. The quality control method based on mass spectrometry data according to claim 1 or 2, characterized in that: In the batch difference assessment step: Before calculating the maximum mean difference, EMD ground motion distance, and KS test statistic, the signal expression values were arcsinh transformed, and the cofactor was set to 5.
6. The quality control method based on mass spectrometry data according to claim 1, characterized in that: In the cluster analysis step: Before clustering, each batch of data was arcsinh transformed, and the cofactor was set to 5; For each batch of data, zscale conversion is performed on each channel to be clustered to reduce the batch effect.
7. The quality control method based on mass spectrometry data according to claim 6, characterized in that: If there are multiple samples in each batch, a rank sum test is performed on the batches based on the proportion of each subpopulation in each data, and a box plot is drawn to show the differences in subpopulation proportions between batches.
8. A quality control system based on mass spectrometry data, characterized in that: The system comprises: The data reading module is used to read each FCS file in the folder according to the specified folder; Signal expression distribution analysis module, used to draw the histogram and cumulative probability density diagram of each channel for each FCS file to evaluate the consistency of signal expression distribution; The discrete degree evaluation module is used to calculate the absolute median deviation of each channel in each data to determine the discrete degree of the channel signal value; Batch difference evaluation module, which is used to calculate the maximum mean difference, EMD ground motion distance and KS test statistic of each channel between two batches to determine the signal difference between batches; Cluster analysis module, using flowsom clustering algorithm for clustering, initially clustering several subgroups, and then using the consensus clustering method to merge into a set number of subgroups; The staining effect evaluation module is used to calculate the average overlapping frequency value of the channel in each data, and take the minimum value of all subpopulations as the final average overlapping frequency value to determine whether the staining effect is good.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: The electronic device comprises a processor, a memory and a transceiver, wherein the memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Large-batch unicellular ATAC-seq data quality controlling and analyzing method
CN107368701A
High-dimensional data visualization analysis method and system
CN116364193A
Mass spectrum streaming data analysis report generation method and system
CN118262800A
Big-data analyzing method and mass spectrometric system using the same method
US20170358434A1
Automatic ion population control for charge detection mass spectrometry
WO2023076583A1