Detection kit for peripheral blood whole immune cell lineage and door analysis model thereof

By constructing a gate analysis model based on metal isotope-labeled antibodies and a grid search algorithm, the problems of low efficiency, large sample consumption, and high cost of high-dimensional data analysis in clinical applications of mass spectrometry flow cytometry are solved. This enables rapid and automated analysis of high-dimensional data from multiple samples, reduces detection costs, and is suitable for large-sample clinical research and point-of-care diagnosis.

CN121805115APending Publication Date: 2026-04-07ARMY MEDICAL UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-12
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Mass cytometry faces challenges in clinical applications, including insufficient efficiency and accuracy in high-dimensional data analysis, high consumption of valuable clinical samples, and high testing costs, making it difficult to meet the needs of large-scale research and timely diagnosis.

Method used

A mixture of three different metal isotope-labeled anti-human CD45 antibodies and 28 metal-labeled antibodies were used, combined with lymphocyte separation medium and sucrose concentration gradient centrifugation, to separate PBMC components. The mixed samples were then detected by mass spectrometry flow cytometry. A grid search algorithm was used to fit the global optimal threshold and construct a gate analysis model to achieve rapid and automated analysis of high-dimensional data from multiple samples.

Benefits of technology

It enables rapid analysis of high-dimensional mass spectrometry flow cytometry data from multiple samples, reduces the amount of precious samples used, improves analytical efficiency and accuracy, and lowers detection costs, making it suitable for large-sample clinical research and point-of-care diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121805115A_ABST
    Figure CN121805115A_ABST
Patent Text Reader

Abstract

The invention relates to a detection kit for a peripheral blood complete immune cell lineage and a door analysis model thereof. The detection kit comprises a monoclonal antibody which is combined with a peripheral mononuclear cell antigen and is marked by 33 rare metal elements. The detection method comprises the following steps: (1) a PBMC (peripheral blood mononuclear cell) acquisition unit: carrying out PBMC separation on a human peripheral blood sample; (2) a living cell bar code marking unit: performing dyeing marking on a plurality of PBMC samples by using a CD45 antibody combination marked by three different metal elements; (3) a sample dyeing unit: mixing a plurality of samples marked by bar codes, and performing dyeing incubation by using prepared 28 metal marked antibodies mix; and (4) a data acquisition unit: performing detection by using a mass spectrometer to obtain original data of the mixed samples. According to the method, rapid and accurate analysis of multi-sample high-dimensional data can be realized, the consumption of precious samples is remarkably reduced, and the cost benefit is considered.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of human peripheral blood immune cell lineage detection and analysis, and particularly relates to a peripheral blood whole immune cell lineage detection kit and a gating analysis model thereof. BACKGROUND

[0002] Precise typing and quantitative analysis of human peripheral blood mononuclear cell subpopulations (including classical monocytes, non-classical monocytes, dendritic cells, NK cells, etc.) are the core basis for assisting the evaluation of immune status of the body and the determination of pathological mechanisms of diseases. For example, in autoimmune diseases such as rheumatoid arthritis, the imbalance of monocyte subpopulation ratio will directly affect the release of inflammatory factors; in tumor immunity, the maturity of dendritic cells and the activity of NK cells are key indicators for predicting response; and in the recovery period of COVID-19 infection, the dynamic changes of peripheral blood mononuclear cell subpopulations can reflect the recovery process of immune function. Therefore, high-dimensional and precise analysis of such cell subpopulations has irreplaceable value for assisting clinical and basic research.

[0003] Mass cytometry, as the only high-dimensional analysis technology that can simultaneously detect 30-50 markers on the cell surface and intracellularly, has the advantages of low signal interference and high-parameter parallel detection, and has broken through the limitation of traditional flow cytometry which can only detect 10-15 markers, becoming a core tool for analyzing the heterogeneity of peripheral blood mononuclear cell subpopulations. This technology combines metal isotope-labeled antibodies with cells, uses inductively coupled plasma mass spectrometry to detect isotope signals, and can simultaneously capture the expression information of cell surface differentiation antigens (such as CD14, CD16, CD11c), intracellular cytokines (such as TNF-α, IL-6) and signaling pathway proteins (such as p-STAT3) at the single-cell level, providing technical support for the fine typing and functional analysis of mononuclear cell subpopulations, and has been widely applied in clinical research in the fields of autoimmune diseases, tumors, infectious diseases, etc.

[0004] However, although mass cytometry has the advantage of high-dimensional analysis, it still faces three major technical bottlenecks in its clinical scale application, which seriously restricts its value.

[0005] First, the analysis efficiency and accuracy of high-dimensional data are insufficient. Current mass spectrometry flow cytometry data analysis mainly relies on manual gate setting or semi-automated analysis tools: manual gate setting requires operators to set gates step by step based on experience, sequentially selecting white blood cells, monocytes, and then further subdividing subpopulations from the total cells. Completing a full set of analyses (including data verification) for a single sample often takes 3-5 hours. Moreover, differences in the judgment of "boundary cells" among different operators can easily lead to deviations in the accuracy and repeatability of the results. In multi-center studies, the data consistency pass rate is less than 60%. Although semi-automated tools can simplify some steps, they still rely on manually preset parameters, such as distance thresholds for clustering algorithms and combinations of subpopulation markers. For complex samples, such as abnormal cell subpopulations in patients with immune disorders, their ability to identify is limited, and they are prone to missed or misclassification, making it difficult to meet the clinical demand for "precise subtyping".

[0006] Secondly, the consumption of valuable clinical samples is significant. Clinical peripheral blood samples are mostly collected via minimally invasive venipuncture, with a typical yield of only 5-10 mL per patient per test. After centrifugation to separate mononuclear cells, the usable sample volume is further reduced. Current analytical methods require processing samples from 1-2 patients per test, and technical replication (usually 2-3 times) is necessary to ensure data reliability. This means a single patient sample can only support one valid analysis. For large-scale clinical studies (such as disease cohort analyses of over 200 patients), not only do patients need to be sampled multiple times to supplement the sample, increasing their physical burden and cooperation difficulty, but some patients (such as the elderly or critically ill patients) may also be unable to provide sufficient samples, compromising the integrity of the research cohort and affecting the representativeness of the research conclusions.

[0007] Third, the testing cost is mismatched with clinical application needs. Current methods require individual sample labeling, washing, and loading: the cost of consumables such as metal-labeled antibodies, washing buffers, and centrifuge tubes per sample, plus the time required for mass spectrometry flow cytometry for each batch of samples, results in a high total cost per sample analysis after adding instrument depreciation and labor costs. Furthermore, the inability to process multiple samples simultaneously means that in emergency clinical scenarios (such as rapid immune assessment of severely infected patients), the time from sample collection to obtaining analysis results can be 8-12 hours, failing to meet the demand for "point-of-care diagnosis." In large-scale disease screening (such as community-level immunization screening), the contradiction between high cost and low efficiency is even more pronounced, limiting the promotion of mass spectrometry flow cytometry technology to primary healthcare institutions or large-scale public health projects.

[0008] In summary, while mass cytometry provides a high-dimensional tool for peripheral blood mononuclear cell subset analysis, its large-scale application in clinical, large-sample cohort studies, and public health fields is severely limited by technical bottlenecks such as low analysis efficiency, high sample consumption, and high cost. Summary of the Invention

[0009] To address the current technical challenge of lacking rapid and accurate analysis of high-throughput, high-dimensional flow cytometry data, this invention provides a detection kit for peripheral blood whole immune cell lineages and its gate analysis model. This kit enables rapid analysis and statistical processing of high-dimensional mass spectrometry flow cytometry data from multiple samples, offering advantages of high accuracy and sensitivity. All samples used are derived from the peripheral blood of patients, allowing for the collection of peripheral immune cell information from 10 patients in a single detection reaction. This effectively reduces the amount of valuable samples required, demonstrating good operability, stability, and repeatability, while also significantly reducing detection costs.

[0010] The technical problem solved by this invention is achieved by the following technical solution:

[0011] The present invention aims to provide a detection kit for peripheral blood complete immune cell lineages, characterized in that,

[0012] It includes a mixture of anti-human CD45 antibodies labeled with three different metal isotopes and 28 metal-labeled antibodies.

[0013] Using human lymphocyte separation medium, PBMC components in fresh human blood samples were separated by sucrose concentration gradient centrifugation. Live cell barcoding was performed using a mixture of anti-human CD45 antibodies labeled with three different metal isotopes. After mixing 5-10 samples, they were stained with multiple metal-labeled antibodies.

[0014] Furthermore, the metal isotopes are selected from three of 110Cd, 111Cd, 195Pt, 196Pt, and 198Pt.

[0015] Three combinations of anti-human CD45 antibodies with different metal labels were selected, and the antibody information is shown in the table below:

[0016] .

[0017] Furthermore, information on 28 metal-labeled antibodies is shown in the table below, including antibody combinations targeting T-cell markers, B-cell markers, monocyte markers, NK-cell markers, and dendritic cell markers.

[0018] .

[0019] Furthermore, the volume of each antibody is 50-100uL.

[0020] A gate analysis model for peripheral blood whole immune cell lineages, including

[0021] (1) Data processing unit: Using the above kit, the raw data of mixed samples are obtained by mass spectrometry flow cytometry, and the raw data are subjected to quality control cleaning, inverse hyperbolic sine transformation and single sample data decoding.

[0022] (2) Core model training unit: Based on the manually labeled training dataset, the global optimal threshold is fitted for all preset immune cell subset gates using a grid search algorithm to form a global gate parameter set;

[0023] (3) Hierarchical filtering unit: enforces hierarchical logic during gate parameter training and cell recognition;

[0024] (4) Automatic identification output unit: The new sample is automatically classified using the global gate parameter set, and the cell-feature matrix and the absolute number and relative proportion of each immune cell subpopulation are output.

[0025] (5) Visualization output unit: Generates a visual cell population hierarchy diagram based on the optimal gate parameters and hierarchical structure.

[0026] Furthermore, the formula for the inverse hyperbolic sine transform is as follows: ,in: : Original signal value; Cofactor; : The transformed value.

[0027] Furthermore, the immune cell types in the manually labeled training dataset are shown in the table below:

[0028]

[0029]

[0030] .

[0031] Furthermore, the training process of the core model training unit includes:

[0032] a. Constructing the cell-feature matrix DATA Matrix-A Its dimension is N×(M+L), where N is the total number of cell events, M is the number of detected markers, and L is the number of preset gates;

[0033] b. Define an objective function based on the F1 score to optimize the overall classification error;

[0034] c. The system generates candidate threshold pairs within a defined search interval using a grid search algorithm;

[0035] d. Evaluate the total error value of each candidate threshold across all training samples;

[0036] e. Select the candidate threshold pair with the smallest total error value as the optimal threshold.

[0037] Furthermore, the overall error of a single sample under a specific gate is quantized as (1 - F1 score), and the F1 score is calculated using the following formula: in, =Number of true positive cells / (Number of true positive cells + Number of false positive cells); =Number of true positive cells / (Number of true positive cells + Number of false negative cells)

[0038] Furthermore, the specific steps of the grid search algorithm include:

[0039] Step 1: For each training sample, initially fit a candidate threshold within the subset of cells passed through by the parent gate;

[0040] Step 2: Collect candidate thresholds from all training samples to form a candidate threshold set;

[0041] Step 3: Determine the global search interval and generate a candidate threshold sequence with a preset fixed step size;

[0042] Step 4: Combine candidate threshold sequences of different dimensions in pairs to generate all possible candidate threshold pairs;

[0043] Step 5: Evaluate the total error of each candidate threshold pair across all training samples.

[0044] Furthermore, evaluating the total error of each candidate threshold pair across all training samples includes:

[0045] 1) Use thresholds to automatically gate all training samples and divide them into four subgroups;

[0046] 2) For the target subgroup, compare the automated gatening results with the manually labeled tags, and calculate...

[0047] Calculate the F1 score and overall error (1 - F1) for each sample;

[0048] 3) Sum the combined errors of all samples for the target subgroup to obtain the corresponding candidate threshold.

[0049] The total error value.

[0050] Furthermore, the hierarchical filtering unit ensures that during the training phase, parameter learning for subpopulations occurs only within the subset of cells that their parent gates have passed through; during the cell recognition phase, cells must pass through the parent gates level by level before they can be judged by the sub-gates.

[0051] This invention discloses a detection gating analysis model for peripheral blood whole immune cell lineages, the details of which include:

[0052] ① PBMC Acquisition Unit: Monocyte population is collected from peripheral blood of the sample to be tested. Monocytes are isolated from fresh peripheral blood using centrifugation with lymphocyte separation medium, washed with DPBS, and then viable cell counting is performed, with cell viability >85%.

[0053] ② PBMC pre-processing unit: single-sample live cell staining "barcode", which is a mixture of anti-human CD45 antibodies labeled with three different metal isotopes (including but not limited to 110Cd / 111Cd / 195Pt / 196Pt / 198Pt-anti-HuCD45). After labeling, the single sample cells are washed, and then 5-10 sample cells are mixed and stained with 28 metal-labeled antibody mix.

[0054] ③ Sample mass spectrometry flow cytometry data acquisition unit: The test results are output in FCS format as the raw data of the mixed sample, denoted as DATAraw, through mass spectrometry flow cytometry.

[0055] ④ Raw data cleaning unit: This unit performs quality control cleaning on the raw data DATAraw, including removing interference signals such as double cells, cell debris, dead cells, and multimers, to improve data quality and obtain a filtered high-quality single-cell dataset, denoted as Datafilter.

[0056] ⑤ The original data transformation unit then performs an inverse hyperbolic sine transform, and the output transformed dataset is denoted as DataArcsinh; the formula for performing the inverse hyperbolic sine transform on the data using Datafilter is as follows: ,in: : Original signal value; Cofactor (usually 5); The transformed value. The output transformed dataset is denoted as DataArcsinh and is used by downstream units.

[0057] ⑥ Single-sample data acquisition unit: When using multiplex detection technology, this unit is responsible for decoding the mixed data DataArcsinh back to each independent original sample. Using the CATALYST package in the R language environment, DataArcsinh is decoded, and according to the CD45-barcode encoding scheme used in the experiment, each cell event is assigned back to its original sample, generating mass spectrometry flow cytometry data (DATAunits) based on the detected samples.

[0058] ⑦ The core model training unit uses a training dataset of 67 immune cell subtypes annotated with artificial gating. It fits the globally optimal threshold for gating all preset immune cell subpopulations using a grid search algorithm, forming a global gating parameter set. .

[0059] The first step, based on current literature reports and professional consensus, is for immunologists to use the gate method to label...

[0060] A training dataset of immune cell typing is used as the initial standard for model training, with approximately 25-40 cases in the training dataset.

[0061] In the second step, each sample DATAunit is organized into a structured cell-feature matrix.

[0062] DATA Matrix-A Its dimensions are This represents the total number of cellular events in the sample. This indicates the number of cell surface markers detected. This indicates the number of preset gates in the system. The matrix before this... Listed as cell marker expression levels, later List the Boolean values ​​for the subgroup affiliation of manually labeled individuals (1 represents belonging, 0 represents not belonging).

[0063] The third step is to define a single-channel objective function. Taking the learning of gate parameters for CD8+ T cell subsets (further subdivided into CD8+TNaive, Tcm, Tem, and Temra) as an example, the objective function optimized in this unit is defined as the sum of the comprehensive classification errors of all training samples at that gate. The comprehensive error of a single sample at a specific gate is quantified as (1 - F1 score). The F1 score is calculated as follows: .in, =Number of true positive cells / (Number of true positive cells + Number of false positive cells); =Number of true positive cells / (Number of true positive cells + Number of false negative cells)

[0064] Step 4: Determine the search scope

[0065] a. In-sample preliminary fitting: For each sample in the training set, first obtain the results through CD8+ T fine-tuning.

[0066] The cell subset of the parent phylum. Within this subset, based on the expression levels of surface markers (such as CD45RA and CCR7), four cell subpopulation regions are initially defined by fitting two straight lines perpendicular to the coordinate axes. This process generates a pair of candidate segmentation thresholds for each sample (one for the CD45RA dimension and one for the CCR7 dimension).

[0067] b. Threshold collection: Collect candidate thresholds fitted to all N training samples in the CD45RA dimension and CCR7 dimension, forming candidate threshold set X and candidate threshold set Y respectively.

[0068] c. Determine the global search interval: Take the minimum value of set X respectively. and maximum value This constitutes the search range of the CD45RA dimension. Take the minimum value of set Y. and maximum value This constitutes the search range of CCR7 dimensions. .

[0069] Step 5: Generate candidate threshold pairs

[0070] Within a defined two-dimensional search interval, candidate threshold pairs are systematically generated. (In the CD45RA dimension interval) Within this range, a candidate threshold sequence is generated with a fixed step size of 0.01. Within the CCR7 dimensional range Within this range, a candidate threshold sequence is generated with a fixed step size of 0.01. .Will and Perform pairwise combinations to generate all possible candidate threshold pairs. .

[0071] Step 6, Single threshold pair evaluation:

[0072] For each candidate threshold pair Perform the following operations:

[0073] a. Use this threshold to... For all Automatic implementation of a subset of CD8+ T cells from each training sample

[0074] The phylum is divided into four subgroups.

[0075] b. For the target subgroup (e.g., Temra), compare the automated gatening results with manually labeled tags.

[0076] Compare and calculate the F1 score and overall error (1 - F1) for each sample.

[0077] c. Accumulate all The combined error of each sample for the Temra subgroup yields the candidate threshold pair corresponding to...

[0078] The total error value.

[0079] Step 7: Determining the optimal threshold:

[0080] After traversing all candidate threshold pairs, the candidate threshold pair with the smallest total error value is selected and determined as the splitting threshold pair.

[0081] The optimal two-dimensional threshold for CD8+ T cell subsets ( , ).

[0082] Step 8: Repeat steps 3 through 7 above to obtain the optimal gating parameters for all preset immune cell subsets (such as CD4+ T cells, NK cells, B cells, and their respective sub-subsets), forming a global gating parameter set. .

[0083] ⑧ The hierarchical filtering unit enforces hierarchical logic during the fitting (training) of gate parameters and cell recognition (prediction). Subgroup parameter learning takes place entirely within the subset of cells that pass through the parent gate.

[0084] Hierarchical logic is enforced during the fitting (training) of gate parameters and cell recognition (prediction).

[0085] During the training phase, parameter learning for subpopulations occurs only within the subset of cells that their parent gate has passed through. During the model prediction phase: for new samples, cells must pass through each parent gate sequentially before being considered for a sub-gate. For example, CD4+ T and CD8+ are analyzed only at the CD3+γδTCR-T cell level.

[0086] ⑨ Automatic Immune Cell Recognition and Output Unit. Utilizing a global gating parameter set. Automated cell typing of new samples was performed, and a cell-feature matrix (DATAmatrix) integrating cell expression level information with the classification results was constructed. Based on the DATAmatrix, the absolute number and relative proportion of immune cell subsets for each sample were statistically analyzed.

[0087] This unit utilizes a pre-trained global gate parameter set. For test set samples Perform automated cell typing and output structured analysis results.

[0088] Hierarchical gate execution: The system applies the corresponding optimal threshold layer by layer to each sample according to the preset gate hierarchy, performs Boolean logic judgments on each cell event in the sample, and finally completes the automated identification of all preset immune cell subpopulations. The system generates a cell-feature matrix DATA for each sample that integrates cell expression information and its classification results. Matrix-B The dimension of this matrix is... The first part of the matrix. Column storage for each cell The standardized expression levels on each marker, followed by The columns store the Boolean (0 or 1) result of each cell passing through each gate, where 1 represents passing and 0 represents failing. Based on the above data...Matrix A matrix that statistically analyzes the absolute number and relative proportion of each immune cell subset in each sample.

[0089] ⑩ The visualization output unit generates a visualized cell population hierarchy map based on the optimal gate parameters and gate hierarchy structure. For each gate in the system, a corresponding two-dimensional scatter plot is generated.

[0090] Compared with the prior art, the beneficial technical effects of the present invention are as follows:

[0091] 1. Advantages in Analytical Efficiency and Accuracy: Existing technologies rely on manual gatening or semi-automated tools to analyze high-throughput, high-dimensional flow cytometry data. Single-sample analysis can take several hours and is highly subjective, easily leading to insufficient accuracy and repeatability due to differences in human experience. This invention, by constructing a dedicated rapid gatening analysis model, enables rapid synchronous analysis and statistical analysis of multi-sample, high-dimensional mass spectrometry flow cytometry data, significantly reducing analysis time while avoiding human error, and significantly improving the accuracy and sensitivity of analysis results, meeting the clinical demand for efficient and accurate data interpretation.

[0092] 2. Advantage in utilizing precious samples: Clinical peripheral blood samples are mostly collected via minimally invasive procedures, and the total amount is limited and precious. Existing technologies can only process 1-2 patient samples per test, resulting in high sample consumption in large-scale studies and easy interruption of research due to insufficient samples. This invention can collect peripheral immune cell information from 10 patients in one test reaction, significantly reducing the amount of sample used per unit, significantly improving the utilization efficiency of precious clinical samples, reducing the sampling burden on patients, and providing sample guarantee for large-scale clinical studies (such as disease cohort analysis).

[0093] 3. Advantages in operation and result stability: Existing streaming data analysis methods have cumbersome operation procedures, require high operator skills, and make it difficult to ensure data consistency in multi-center studies. The analytical model of this invention has a high degree of standardization, the operation steps are simple and easy to understand, and no complex professional skills are required to get started. At the same time, through unified gate logic and statistical algorithms, the stability and repeatability of test results from different batches and different laboratories are ensured, which facilitates the promotion and application of clinical multi-center studies.

[0094] 4. Advantages in controlling detection costs: In existing technologies, samples need to be processed individually, resulting in high consumption of reagents and consumables, long instrument occupancy time, and high unit sample detection costs, making it difficult to scale up applications. This invention, through a multi-sample simultaneous detection design, reduces the cost of reagents and consumables required per sample, as well as the instrument time cost, effectively lowering the unit sample detection cost. At the same time, it reduces the waste of instrument resources and is more suitable for large-scale scenarios such as large-sample clinical immune assessment and disease screening, enhancing the clinical translation and promotion value of the technology. Attached Figure Description

[0095] Figure 1This is a simplified flowchart illustrating the construction process of a gate analysis model for peripheral blood whole immune cell lineages according to the present invention.

[0096] Figure 2 This is a diagram of the machine gate strategy for each immune cell group in PBMC in this invention.

[0097] Figure 3 This is a machine gate strategy diagram of the functional subsets of innate and adaptive immune cells in PBMCs in this invention.

[0098] Figure 4 This is a machine gate strategy diagram of the functional and differentiation subtypes of CD4+ T cells and CD8+ T cells in PBMCs in this invention.

[0099] Figure 5 This is a graph showing the machine gate model prediction results of CyTOF data from 10 PBMC samples in this invention.

[0100] Figure 6 This invention compares the results of manual and machine-based gate analysis of CyTOF data from 10 samples under the same gate analysis strategy. The results are presented as statistical analysis of the proportion of each cell subtype in the 10 samples (the manual gate analysis was independently performed by three experienced immunologists).

[0101] Figure 7 This is an analysis of the dispersion of the differences in the results of manual and machine gated models for each cell subtype in this invention, with coefficients of variation (CV) all below 30%. Detailed Implementation

[0102] The technical solution of the present invention will be further described in detail below with reference to specific embodiments. It should be understood that the following embodiments are merely illustrative and explanatory of the present invention and should not be construed as limiting the scope of protection of the present invention. All technologies implemented based on the above content of the present invention are covered within the scope of protection intended by the present invention.

[0103] In addition, unless otherwise specified, all raw materials, reagents, instruments and equipment used in this invention can be obtained by purchasing them from the market or prepared by existing methods.

[0104] Table 1: Reagent Sources

[0105] .

[0106] A gate analysis model for detecting peripheral blood whole immune cell lineages includes:

[0107] (1) Freshly drawn blood (1-3 mL, anticoagulated) from 40 cases was centrifuged using a density gradient centrifugation method with lymphocyte separation medium to separate the white membrane of cells containing PBMCs. After washing with DPBS, the cell count showed a viability of at least >85% and a number >0.3×10⁻⁶. 6 It can be used for subsequent experiments.

[0108] (2) CyTOF detection experimental procedure, the specific steps are as follows:

[0109] ① For every 10 human PBMC samples, 0.3 × 10⁻⁶ samples were collected from each sample. 6 Cells were brought to a final volume of 49 μL with FACS Buffer (DPBS solution containing 2% BSA), and 1 μL of Human TruStain FcX was added. Cells were then blocked at room temperature.

[0110] ② Add 50 μL of premixed CD45-Barcode incubation solution to each sample. The formula is shown in Table 2. After mixing, incubate at room temperature for 0.5 h.

[0111] ③ Add 500 μL FACS Buffer to each sample, resuspend the cells, centrifuge at 300g for 5 min at room temperature, and discard the supernatant; repeat once.

[0112] ④ After mixing 10 samples, resuspend the cells, centrifuge at 300g for 5 min at room temperature, discard the supernatant, add 100 μL of antibody incubation solution (as shown in Table 3), and incubate at room temperature for 30 min.

[0113] ⑤ Add 1000 μL FACS Buffer, resuspend the cells, centrifuge at 300g for 5 min at room temperature, discard the supernatant; repeat once.

[0114] ⑥ Add 500 μL of 1.6% paraformaldehyde solution, fix at room temperature for 10 min, then centrifuge at 500g at room temperature for 5 min and discard the supernatant;

[0115] ⑦ Prepare a final concentration of 125 nM Ir staining solution using Fix and Perm Buffer. Resuspend the cells in 500 μL and incubate at room temperature for 1 hour or overnight at 4°C to stain and fix the DNA. Centrifuge at 500g for 5 minutes at room temperature and discard the supernatant.

[0116] ⑧ Add 1000 μL FACS Buffer to resuspend the cells, centrifuge at 500g for 5 min at room temperature, and discard the supernatant; repeat once.

[0117] ⑨ Prepare a 10% EQ beads solution using Cell Acquisition Buffer, add 1 mL to resuspend the cells, and transfer the solution to a 5 mL flow cytometry tube with a filter using a pipette;

[0118] ⑩ The filtered cell suspension was analyzed by mass cytometry to obtain the raw mass cytometry data of the sample.

[0119] Table 2: CD45-Barcode Combinations for Live Cell Staining

[0120]

[0121] Table 3: Incubation solutions for metal-labeled antibody combinations used

[0122] .

[0123] (3) Data cleaning is performed on DATAraw, including removing interference signals such as double cells, cell debris, dead cells, and multimers to improve data quality and obtain a filtered, high-quality single-cell dataset, denoted as Datafilter. Datafilter performs an inverse hyperbolic sine transform. The transform formula is as follows: ,in: : Original signal value; Cofactor (usually 5); The transformed value. The output transformed dataset is denoted as DataArcsinh.

[0124] (4) Decode the mixed sample data DataArcsinh back to each independent original sample. First, in the R language environment, call the CATALYST package, and according to the CD45-barcode encoding scheme used in the experiment, assign each cell event back to its original sample, generating mass spectrometry flow cytometry data (DATAunit1-10) in units of samples.

[0125] (5) The core model is based on a manually labeled training dataset, and uses a grid search algorithm to fit the global single-channel optimal threshold for all preset immune cell subpopulations.

[0126] The first step involves immunologists labeling a training dataset of immune cell typing using the gate mapping method. This dataset serves as the initial standard for model training and contains approximately 25 cases.

[0127] The second step involves organizing each sample DATAunit into a structured cell-feature matrix DATA. Matrix-A Its dimensions are This represents the total number of cellular events in the sample. This indicates the number of cell surface markers detected. This indicates the number of preset gates in the system. The matrix before this... Listed as cell marker expression levels, later List the Boolean values ​​for the subgroup affiliation of manually labeled individuals (1 represents belonging, 0 represents not belonging).

[0128] The third step is to define a single-channel objective function. Taking the learning of gate parameters for CD8+ T cell subsets (further subdivided into CD8+TNaive, Tcm, Tem, and Temra) as an example, the optimized objective function is defined as the sum of the comprehensive classification errors of all training samples at that gate. The comprehensive error of a single sample at a specific gate is quantified as (1 - F1 score). The F1 score is calculated as follows: .in, = Number of true positive cells / (Number of true positive cells + Number of false positive cells); = Number of true positive cells / (Number of true positive cells + Number of false negative cells).

[0129] Step 4, Determine the search range. a. Preliminary in-sample fitting: For each sample in the training set, first obtain the results through CD8+ T fine-tuning.

[0130] The cell subset of the parent phylum. Within this subset, based on the expression levels of surface markers (such as CD45RA and CCR7), four cell subpopulation regions are initially defined by fitting two straight lines perpendicular to the coordinate axes. This process generates a pair of candidate segmentation thresholds for each sample (one for the CD45RA dimension and one for the CCR7 dimension).

[0131] b. Threshold collection: Collect candidate thresholds fitted to all N training samples in the CD45RA dimension and CCR7 dimension, forming candidate threshold set X and candidate threshold set Y respectively.

[0132] c. Determine the global search interval: Take the minimum value of set X respectively. and maximum value This constitutes the search range of the CD45RA dimension. Take the minimum value of set Y. and maximum value This constitutes the search range of CCR7 dimensions. .

[0133] Step 5: Generate candidate threshold pairs

[0134] Within a defined two-dimensional search interval, candidate threshold pairs are systematically generated. (In the CD45RA dimension interval) Within this range, a candidate threshold sequence is generated with a fixed step size of 0.01. Within the CCR7 dimensional range Within this range, a candidate threshold sequence is generated with a fixed step size of 0.01. .Will and Perform pairwise combinations to generate all possible candidate threshold pairs. .

[0135] Step 6, Single threshold pair evaluation:

[0136] For each candidate threshold pair Perform the following operations:

[0137] a. Use this threshold to... For all An automated gating process was performed on a subset of CD8+ T cells from the training samples, dividing them into four subpopulations.

[0138] b. For the target subgroup (e.g., Temra), compare the automated gatening results with the manually labeled tags, and calculate the F1 score and overall error (1 - F1) for each sample.

[0139] c. Accumulate all The combined error of each sample for the Temra subgroup is used to obtain the total error value corresponding to the candidate threshold pair.

[0140] Step 7: Determining the optimal threshold:

[0141] After traversing all candidate threshold pairs, the candidate threshold pair with the smallest total error value is selected and determined as the optimal two-dimensional threshold for classifying CD8+ T cell subsets. , ).

[0142] Step 8: Repeat steps 3 through 7 above to obtain the optimal gating parameters for all preset immune cell subsets (such as CD4+ T cells, NK cells, B cells, and their respective sub-subsets), forming a global gating parameter set. .

[0143] During the fitting (training) and cell identification (prediction) processes of gate parameters, hierarchical logic is enforced. In the training phase, parameter learning for subpopulations occurs only within the subset of cells that their parent gates have passed through. In the model prediction phase: for new samples, cells must pass through parent gates level by level before entering a sub-gate for evaluation. For example, CD4+ T and CD8+ are analyzed only at the CD3+γδTCR-T cell level.

[0144] (6) Automatic recognition and output of immune cells

[0145] Using a pre-trained global gate parameter set For test set samples The system performs automated cell typing and outputs structured analysis results. Hierarchical gatening execution: Following a preset gate hierarchy, the system applies optimal thresholds layer by layer to each sample, performs Boolean logic judgments on each cellular event in the sample, and ultimately completes the automated identification of all preset immune cell subsets. The system generates a cell-feature matrix (DATA) for each sample that integrates cell expression level information and its classification results. Matrix-B The dimension of this matrix is... The first part of the matrix Column storage for each cell The standardized expression levels on each marker, followed by The columns store the Boolean (0 or 1) result of each cell passing through each gate, where 1 represents passing and 0 represents failing. Based on the above data... Matrix A matrix that statistically analyzes the absolute number and relative proportion of each immune cell subset in each sample.

[0146] (7) Visualization output: Based on the optimal parameters and hierarchical structure of the gates, a visualization diagram is generated. For each gate in the system, a corresponding two-dimensional scatter plot is generated.

[0147] Figures 2-4 This invention utilizes 28 CyTOF antibodies to detect 67 immune cell populations in PBMCs, generating gated logic and visualization output results. These results are represented as a two-dimensional scatter plot of each cell subpopulation and its percentage within the next higher-level cell class. Figures 5-7 This invention illustrates the comparison and standard deviation of all cell types in the same sample between the analytical model and manual gating.

[0148] This invention is a one-click rapid gating analysis model for human peripheral blood mononuclear cell (PBMC) subsets based on mass cytometry. Peripheral blood samples are treated with lymphocyte separation fluid to isolate PBMCs, and live cell barcodes are applied to the PBMCs. Multiple samples are mixed in a single reaction for staining with metal-labeled antibodies, followed by mass cytometry analysis to obtain a flow cytometry dataset. A training dataset of 67 immune cell subtypes annotated using artificial gating, combined with negative thresholds for cell surface markers, serves as the core modeling parameter. A grid search algorithm fits globally optimal thresholds for all preset immune cell subset gatings, forming a global gating parameter set. Simultaneously, hierarchical logic is enforced to train cell subsets to learn parameters at parent gates. Test set samples undergo automated cell subtyping within the model, integrating cell expression information with the cell-feature matrix of the classification results. The final output is the absolute number and relative proportion of immune cell subsets. This invention provides a high-precision, rapid, and automated PBMC immune cell subset gating analysis model, offering auxiliary assessment and treatment options for clinical practice. It can enable rapid and accurate analysis of multi-sample, high-dimensional data, significantly reduce the consumption of precious samples, and take cost-effectiveness into account, so as to break through the current technical barriers and promote the transformation of mass spectrometry flow cytometry technology from a "scientific research tool" to a "routine clinical testing method".

Claims

1. A detection kit for peripheral blood complete immune cell lineages, characterized in that, Including many different Metal isotope-labeled anti-human CD45 antibody and 28 metal-labeled antibodies.

2. The detection kit for peripheral blood complete immune cell lineages as described in claim 1, characterized in that, The metal isotopes used to label anti-human CD45 antibodies were 110Cd, 111Cd, 195Pt, 196Pt, and 198Pt.

3. The detection kit for peripheral blood complete immune cell lineages as described in claim 1, characterized in that, The 28 metal-labeled antibodies include:

4. A gate analysis model for peripheral blood whole immune cell lineages, characterized in that, include (1) Data processing unit: Using the kit described in claims 1-3, the raw data of mixed samples is obtained by mass spectrometry flow cytometry, and the raw data is subjected to quality control cleaning, inverse hyperbolic sine transformation and single sample data decoding; (2) Core model training unit: Based on the manually labeled training dataset, the global optimal threshold is fitted for all preset immune cell subset gates using a grid search algorithm to form a global gate parameter set; (3) Hierarchical filtering unit: enforces hierarchical logic during gate parameter training and cell recognition; (4) Automatic identification output unit: The new sample is automatically classified using the global gate parameter set, and the cell-feature matrix and the absolute number and relative proportion of each immune cell subpopulation are output. (5) Visualization output unit: Generates a visual cell population hierarchy diagram based on the optimal gate parameters and hierarchical structure.

5. The gate analysis model for peripheral blood whole immune cell lineages as described in claim 4, characterized in that, The formula for the inverse hyperbolic sine transform is as follows: ,in: : Original signal value; Cofactor; : The transformed value.

6. The gate analysis model for peripheral blood whole immune cell lineages as described in claim 4, characterized in that, The manually labeled training dataset includes 66 immune cell subtypes:

7. The gate analysis model for peripheral blood whole immune cell lineages as described in claim 4, characterized in that, The training process of the core model training unit includes: a. Constructing the cell-feature matrix DATA Matrix-A Its dimension is N×(M+L), where N is the total number of cell events, M is the number of detected markers, and L is the number of preset gates; b. Define an objective function based on the F1 score to optimize the overall classification error; c. The system generates candidate threshold pairs within a defined search interval using a grid search algorithm; d. Evaluate the total error value of each candidate threshold across all training samples; e. Select the candidate threshold pair with the smallest total error value as the optimal threshold.

8. The gate analysis model for peripheral blood whole immune cell lineages as described in claim 7, characterized in that, The overall error of a single sample under a specific gate is quantized as (1 - F1 score), and the F1 score is calculated as follows: in, =Number of true positive cells / (Number of true positive cells + Number of false positive cells); =Number of true positive cells / (Number of true positive cells + Number of false negative cells) 9. The gate analysis model for peripheral blood whole immune cell lineages as described in claim 7, characterized in that, The specific steps of the grid search algorithm include: Step 1: For each training sample, initially fit a candidate threshold within the subset of cells passed by the parent gate; Step 2: Collect candidate thresholds from all training samples to form a candidate threshold set; Step 3: Determine the global search interval and generate a candidate threshold sequence with a preset fixed step size; Step 4: Combine candidate threshold sequences of different dimensions in pairs to generate all possible candidate threshold pairs; Step 5: Evaluate the total error of each candidate threshold pair across all training samples.

10. The gate analysis model for peripheral blood whole immune cell lineages as described in claim 9, characterized in that... It lies in, For each candidate threshold pair, the total error across all training samples is evaluated as follows: 1) Use thresholds to automatically gate all training samples and divide them into four subgroups; 2) For the target subgroup, compare the automated gatening results with the manually labeled tags, and calculate... Calculate the F1 score and overall error (1 - F1) for each sample; 3) Sum the combined errors of all samples for the target subgroup to obtain the corresponding candidate threshold. The total error value.

Citation Information

Patent Citations

  • Drawing method of acute hepatic failure mouse liver full-immune cell characteristic spectrum

    CN110333357A

  • Mass spectrum flow platform trace cell sample detection method using platinum chimeric carrier cells

    CN115561146A

  • Application of quantum dot labeled antibody reagent in cell detection

    CN115792219A

  • Non-small cell lung cancer early typing system based on peripheral blood immune cell map

    CN116337727A

  • Mass spectrum streaming data analysis report generation method and system

    CN118262800A