An early screening system for atherosclerotic coronary heart disease based on peripheral blood immune cell atlas
By using a peripheral blood immune cell atlas screening system, mass spectrometry flow cytometry and random forest classifier, the problem of inaccurate early screening of atherosclerotic coronary heart disease in existing technologies has been solved, achieving efficient and non-invasive early screening and diagnosis of atherosclerotic coronary heart disease.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2022-09-20
- Publication Date
- 2026-04-17
AI Technical Summary
Current technologies lack effective disease prediction models for different stages of coronary atherosclerosis pathological progression, especially failing to systematically incorporate peripheral blood immune system characteristics into cardiovascular disease risk factors, resulting in inaccurate early screening for atherosclerotic coronary heart disease.
A screening system based on peripheral blood immune cell atlas is adopted, which achieves early screening and diagnosis of atherosclerotic coronary heart disease by acquiring peripheral blood single cells, mass spectrometry flow cytometry detection, data conversion unit, immune cell clustering and feature vector acquisition, combined with random forest binary classifier.
It achieves highly accurate and sensitive screening for atherosclerotic coronary heart disease, provides a rapid, non-invasive, and stable analytical method, and can distinguish between healthy individuals and patients with atherosclerosis, as well as low-risk and high-risk patients, to help develop effective treatment plans.
Smart Images

Figure CN115547485B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of early screening technology for atherosclerotic coronary heart disease, and more particularly to an early screening system for atherosclerotic coronary heart disease based on peripheral blood immune cell atlases. Background Technology
[0002] Cardiovascular disease (CVD) is the leading cause of death worldwide, accounting for more than one-third of all deaths. Atherosclerosis (AS), a major underlying pathogenesis of CVD, is a chronic inflammatory disorder characterized by endothelial dysfunction, immune cell activation, and the formation of lipid-rich atherosclerotic plaques in medium and large arteries. This leads to a series of cardiovascular and cerebrovascular complications, seriously endangering human health. Early diagnosis of atherosclerosis and objectively and accurately assessing its severity are of paramount clinical importance.
[0003] Over the past few decades, various predictive models and imaging techniques for cardiovascular diseases, such as intravascular ultrasound (IVUS), targeted acoustic contrast imaging, magnetic resonance imaging (MRI), acoustic density quantification (AD), and optical coherence tomography (OCT), have been applied to clinical diagnosis. While these models have shown good detection effectiveness in distinguishing between patients with and without atherosclerosis, there is currently a lack of models that can effectively predict the pathological progression of coronary atherosclerosis at different stages. Although novel immune biomarkers related to cardiovascular disease are being discovered and applied to early clinical diagnostic examinations, current cardiovascular disease prediction models do not systematically incorporate the immune system, particularly peripheral blood immune system characteristics, as a risk factor for cardiovascular disease into the modeling process. Therefore, developing an early screening system for atherosclerotic coronary artery disease that targets characteristic changes in the peripheral immune system has significant clinical translational application value. Summary of the Invention
[0004] To address the current technical problem of lacking effective models for predicting the pathological progression of coronary atherosclerosis at different stages, this invention provides an early screening system for atherosclerotic coronary heart disease based on peripheral blood immune cell atlases. This system enables early screening and diagnosis of atherosclerotic coronary heart disease, boasting high accuracy and sensitivity. Furthermore, the analysis samples are derived from patient peripheral blood, exhibiting good operability, non-invasiveness, stability, and repeatability.
[0005] The specific technical solution of this invention is as follows:
[0006] An early screening system for atherosclerotic coronary heart disease based on peripheral blood immune cell atlas, comprising:
[0007] The peripheral blood single-cell acquisition unit is used to isolate single cells from the peripheral blood of patients to be diagnosed, obtaining peripheral blood single-cell samples. The mass cytometry detection unit is used to split the peripheral blood single-cell samples in two, and then perform mass cytometry cell detection experiments on the peripheral blood single-cell samples using metal isotope-coupled protein-specific biomarkers (T-panel marker and M-panel marker) for T cells and myeloid cells respectively, obtaining raw cell mass cytometry data of the samples. The raw mass cytometry data of T cells and myeloid cells are respectively denoted as Data. rawT and Data rawM ;
[0008] The raw data conversion unit is used to process the raw data from the cell mass cytometry flow cytometry of the sample. rawT and Data rawM The data were converted to obtain the converted single-cell mass cytometry data. transT and Data transM The conversion formulas are as follows: Data transT =sinh -1 (Data rawT / 5), Data transM =sinh -1 (Data rawM / 5);
[0009] Peripheral blood immune cell cluster acquisition unit, used to separately process Data transT and Data transM Cluster analysis was performed to obtain the proportions of T cells, NK cells, B cells, and myeloid cells in peripheral blood samples, denoted as Data. clusterTcell Data clusterNKcell Data clusterBcell and Data clusterMcell ;
[0010] Subgroup merging unit is used to collect subgroup information obtained from cluster analysis of two sets of single-cell mass cytometry data of peripheral blood samples. clusterTcell Data clusterNKcell Data clusterBcell and Data clusterMcell By merging the data, we obtain the peripheral blood immune cell subset information, i.e., the peripheral blood immune cell atlas, denoted as Data. immune-cell ;
[0011] The feature vector acquisition unit is used to obtain features based on peripheral blood immune cell atlas data.immune-cell Calculate the feature vector Q composed of the proportion of each immune cell subpopulation in the sample, where the proportion of a certain immune cell subpopulation = the total number of cells in the sample belonging to a certain cell subpopulation / the total number of cells in the sample.
[0012] The classification acquisition unit is used to substitute the feature vector Q into a pre-trained Random Forest binary classifier to calculate the sample classification and obtain the prediction results of the early screening system for atherosclerotic coronary heart disease.
[0013] Preferably, the peripheral blood single-cell acquisition unit uses a gradient centrifugation method with Ficoll separation solution to separate mononuclear cells from the peripheral blood of the patient to be diagnosed, thereby obtaining a peripheral blood sample single cell.
[0014] As a preferred embodiment, the specific biomarkers for metal isotope-coupled proteins of T cells and myeloid cells in the mass spectrometry flow cytometry detection unit are shown in Tables 1 and 2, respectively:
[0015] Table 1. Specific biomarkers for metal isotope-coupled proteins in T cells.
[0016]
[0017] Table 2. Specific biomarkers for metal isotope-coupled proteins in myeloid cells.
[0018]
[0019] Preferably, in the peripheral blood immune cell cluster acquisition unit, the Data... transT The clustering analysis method includes the following steps: Clustering is performed using the phenotyping by accelerated refined community-partitioning (PARC) algorithm; subpopulations are merged based on the expression of metal isotope-coupled protein-specific biomarkers; T cell and NK cell-related subpopulations are selected and subjected to secondary clustering using the phenotyping by accelerated refined community-partitioning algorithm; the corresponding peripheral blood sample subpopulation proportions are denoted as Data. clusterTcell and Data clusterNKcell .
[0020] Preferably, in the peripheral blood immune cell cluster acquisition unit, the Data... transMThe clustering analysis method includes the following steps: clustering is performed using a phenotypic analysis algorithm based on accelerated fine-grained community partitioning; subpopulations are merged based on the expression of metal isotope-coupled protein-specific biomarkers; B cell and myeloid cell-related subpopulations are selected and subjected to secondary clustering using a phenotypic analysis algorithm based on accelerated fine-grained community partitioning; the corresponding peripheral blood sample subpopulation proportions are denoted as Data. clusterBcell and Data clusterMcell .
[0021] Furthermore, during the clustering process using a phenotypic analysis algorithm based on accelerated fine community segmentation, subgroups with a subgroup percentage of less than 0.1% are filtered out.
[0022] Preferably, in the subgroup merging unit, the subgroup information Data obtained from the cluster analysis of two groups of single-cell mass cytometry data of peripheral blood samples is used. clusterTcell Data clusterNKcell Data clusterBcell and Data clusterMcell The merging method used is hierarchical clustering.
[0023] Preferably, in the classification acquisition unit, feature selection is performed before substituting the feature vector Q into the pre-trained random forest binary classifier; during the acquisition of the pre-trained random forest binary classifier, the selected features are used to train the random forest binary classifier; the feature selection method includes the following steps: randomly selecting multiple samples, performing normalization processing, then using a random forest model to construct a classifier, and simultaneously using 10-fold cross-validation to obtain the average feature importance of the random forest model, selecting features with an average feature importance greater than 0.01 as important features, iterating in this way for 800 to 1200 times, counting the number of times each important feature appears, and if the number is greater than 300, then the important feature is selected.
[0024] Preferably, the early screening system for atherosclerotic coronary heart disease is a disease prediction (DP) system used to distinguish between healthy individuals and patients with atherosclerosis. The method for obtaining the pre-trained random forest binary classifier in the classification acquisition unit includes the following steps: randomly selecting several low-risk atherosclerotic coronary heart disease patients, several high-risk atherosclerotic coronary heart disease patients, and several healthy individuals; obtaining the feature vector Q of each sample through a peripheral blood single-cell acquisition unit, a mass spectrometry flow cytometry detection unit, a raw data conversion unit, a peripheral blood immune cell clustering acquisition unit, a subpopulation merging unit, and a feature vector acquisition unit, as the training dataset for training the random forest binary classifier; randomly selecting several low-risk atherosclerotic coronary heart disease patients, several high-risk atherosclerotic coronary heart disease patients, and several healthy individuals; obtaining the feature vector Q of each sample through the same method, as the test dataset for testing the trained random forest binary classifier; and training the random forest binary classifier using the feature vector Q of all samples.
[0025] Preferably, the early screening system for atherosclerotic coronary heart disease is a disease progression prediction (DPP) system used to distinguish between low-risk and high-risk atherosclerotic coronary heart disease patients. The method for obtaining the pre-trained random forest binary classifier in the classification acquisition unit includes the following steps: randomly selecting several low-risk and high-risk atherosclerotic coronary heart disease patients, and obtaining the feature vector Q of each sample through a peripheral blood single-cell acquisition unit, a mass spectrometry flow cytometry detection unit, a raw data conversion unit, a peripheral blood immune cell cluster acquisition unit, a subgroup merging unit, and a feature vector acquisition unit, as the training dataset for training the random forest binary classifier; randomly selecting several low-risk, high-risk, and healthy individuals, and obtaining the feature vector Q of each sample using the same method, as the prediction model for testing the trained random forest binary classifier; and training the random forest binary classifier using the feature vector Q of all samples.
[0026] Compared with the prior art, the present invention has the following advantages:
[0027] (1) The system of the present invention has high accuracy and sensitivity for early screening and prediction of atherosclerotic coronary heart disease. It can be used to distinguish between healthy individuals and patients with atherosclerosis, as well as between patients with low-risk and high-risk atherosclerotic coronary heart disease. It is fast, efficient, and cost-effective. It can provide richer judgment indicators for early screening and diagnosis of atherosclerotic coronary heart disease and help clinicians to formulate effective treatment plans for patients in a timely manner.
[0028] (2) The analytical samples used in the system of the present invention are derived from the peripheral blood of patients, and have good operability, non-invasiveness, stability and repeatability. Attached Figure Description
[0029] Figure 1 This is a simplified flowchart of an early screening system for atherosclerotic coronary heart disease according to the present invention.
[0030] Figure 2 This is a receiver operating characteristic (ROC) curve graph of the sample based on immune-clinical feature classification compared with the results of immune feature classification and clinical feature classification; where "clinical features" refers to clinical features, "immune features" refers to immune features, and "combined features" refers to immune-clinical features.
[0031] Figure 3 This invention compares the prediction results of the immune-clinical feature classification based on the present invention with those of the immune feature classification and clinical feature classification in terms of sensitivity and specificity; where "clinicalfeatures" refers to clinical features, "immune features" refers to immune features, and "combined features" refers to immune-clinical features. Detailed Implementation
[0032] The present invention will be further described below with reference to embodiments.
[0033] General Implementation Examples
[0034] An early screening system for atherosclerotic coronary heart disease based on peripheral blood immune cell atlas, comprising:
[0035] The peripheral blood single-cell acquisition unit is used to isolate single cells from the peripheral blood of patients to be diagnosed, obtaining peripheral blood single-cell samples. The mass cytometry detection unit is used to split the peripheral blood single-cell samples in two, and then perform mass cytometry cell detection experiments on the peripheral blood single-cell samples using metal isotope-coupled protein-specific biomarkers (T-panel marker and M-panel marker) for T cells and myeloid cells respectively, obtaining raw cell mass cytometry data of the samples. The raw mass cytometry data of T cells and myeloid cells are respectively denoted as Data. rawT and Data rawM ;
[0036] The raw data conversion unit is used to process the raw data from the cell mass cytometry flow cytometry of the sample. rawT and Data rawM The data were converted to obtain the converted single-cell mass cytometry data. transT and Data transM The conversion formulas are as follows: Data transT =sinh -1 (Data rawT / 5), Data transM =sinh -1 (Data rawM / 5);
[0037] Peripheral blood immune cell cluster acquisition unit, used to separately process Data transT and Data transM Cluster analysis was performed to obtain the proportions of T cells, NK cells, B cells, and myeloid cells in peripheral blood samples, denoted as Data. clusterTcell Data clusterNKcell Data clusterBcell and Data clusterMcell ;
[0038] Subgroup merging unit is used to collect subgroup information obtained from cluster analysis of two sets of single-cell mass cytometry data of peripheral blood samples. clusterTcell Data clusterNKcell Data clusterBcell and Data clusterMcell By merging the data, we obtain the peripheral blood immune cell subset information, i.e., the peripheral blood immune cell atlas, denoted as Data. immune-cell ;
[0039] The feature vector acquisition unit is used to obtain features based on peripheral blood immune cell atlas data. immune-cellCalculate the feature vector Q composed of the proportion of each immune cell subpopulation in the sample, where the proportion of a certain immune cell subpopulation = the total number of cells in the sample belonging to a certain cell subpopulation / the total number of cells in the sample.
[0040] The classification acquisition unit is used to substitute the feature vector Q into a pre-trained Random Forest binary classifier to calculate the sample classification and obtain the prediction results of the early screening system for atherosclerotic coronary heart disease.
[0041] Optionally, the peripheral blood single-cell acquisition unit uses a gradient centrifugation method with Ficoll separation solution to separate mononuclear cells from the peripheral blood of the patient to be diagnosed, thereby obtaining a peripheral blood sample single cell.
[0042] Optionally, the specific biomarkers for metal isotope-coupled proteins of T cells and myeloid cells in the mass spectrometry flow cytometry detection unit are shown in Table 1 and Table 2, respectively:
[0043] Table 1. Specific biomarkers for metal isotope-coupled proteins in T cells.
[0044]
[0045] Table 2. Specific biomarkers for metal isotope-coupled proteins in myeloid cells.
[0046]
[0047] Optionally, in the peripheral blood immune cell cluster acquisition unit, the Data... transT The clustering analysis method includes the following steps: Clustering is performed using the phenotyping by accelerated refined community-partitioning (PARC) algorithm; subpopulations with a proportion less than 0.1% are filtered out; subpopulations are merged based on the expression of metal isotope-coupled protein-specific biomarkers; T cell and NK cell-related subpopulations are selected and subjected to secondary clustering using the phenotyping by accelerated refined community-partitioning algorithm; the corresponding peripheral blood sample subpopulation proportions are denoted as Data. clusterTcell and Data clusterNKcell .
[0048] Optionally, in the peripheral blood immune cell cluster acquisition unit, the Data... transMThe clustering analysis method includes the following steps: clustering is performed using a phenotypic analysis algorithm based on accelerated fine-grained community partitioning; subpopulations with a proportion less than 0.1% are filtered out; subpopulations are merged based on the expression of metal isotope-coupled protein-specific biomarkers; B-cell and myeloid cell-related subpopulations are selected and subjected to secondary clustering using a phenotypic analysis algorithm based on accelerated fine-grained community partitioning; the corresponding peripheral blood sample subpopulation proportion information is denoted as Data. clusterBcell and Data clusterMcell .
[0049] Optionally, in the subgroup merging unit, the subgroup information Data obtained from the cluster analysis of two groups of single-cell mass cytometry data of peripheral blood samples is used. clusterTcell Data clusterNKcell Data clusterBcell and Data clusterMcell The merging method used is hierarchical clustering.
[0050] Optionally, in the classification acquisition unit, feature selection is performed before substituting the feature vector Q into the pre-trained random forest binary classifier; during the acquisition of the pre-trained random forest binary classifier, the selected features are used to train the random forest binary classifier; the feature selection method includes the following steps: randomly selecting multiple samples, performing normalization processing, then using a random forest model to construct a classifier, and simultaneously using 10-fold cross-validation to obtain the average feature importance of the random forest model, selecting features with an average feature importance greater than 0.01 as important features, iterating in this way for 800 to 1200 times, counting the number of times each important feature appears, and if the number is greater than 300, then the important feature is selected.
[0051] As one specific implementation, the early screening system for atherosclerotic coronary heart disease is a disease prediction (DP) system used to distinguish between healthy individuals and patients with atherosclerosis. The method for obtaining the pre-trained random forest binary classifier in the classification acquisition unit includes the following steps: randomly selecting several low-risk atherosclerotic coronary heart disease patients, several high-risk atherosclerotic coronary heart disease patients, and several healthy individuals; obtaining the feature vector Q of each sample through a peripheral blood single-cell acquisition unit, a mass spectrometry flow cytometry detection unit, a raw data conversion unit, a peripheral blood immune cell clustering acquisition unit, a subpopulation merging unit, and a feature vector acquisition unit, as the training dataset for training the random forest binary classifier; randomly selecting several low-risk atherosclerotic coronary heart disease patients, several high-risk atherosclerotic coronary heart disease patients, and several healthy individuals; obtaining the feature vector Q of each sample through the same method, as the test dataset for testing the trained random forest binary classifier; and training the random forest binary classifier using the feature vector Q of all samples.
[0052] As another specific implementation, the early screening system for atherosclerotic coronary heart disease is a disease progression prediction (DPP) system used to distinguish between low-risk and high-risk atherosclerotic coronary heart disease patients. The method for obtaining the pre-trained random forest binary classifier in the classification acquisition unit includes the following steps: randomly selecting several low-risk and high-risk atherosclerotic coronary heart disease patients, and obtaining the feature vector Q of each sample through a peripheral blood single-cell acquisition unit, a mass spectrometry flow cytometry detection unit, a raw data conversion unit, a peripheral blood immune cell cluster acquisition unit, a subgroup merging unit, and a feature vector acquisition unit, as a training dataset for training the random forest binary classifier; randomly selecting several low-risk, high-risk, and healthy individuals, and obtaining the feature vector Q of each sample using the same method, as a prediction model for testing the trained random forest binary classifier; and training the random forest binary classifier using the feature vector Q of all samples.
[0053] Example 1
[0054] An early screening system for atherosclerotic coronary heart disease based on peripheral blood immune cell atlas, comprising:
[0055] The peripheral blood single-cell acquisition unit is used to isolate single cells from the peripheral blood of patients to be diagnosed, obtaining peripheral blood single-cell samples. The mass cytometry detection unit is used to split the peripheral blood single-cell sample in two, and then perform mass cytometry detection experiments on the peripheral blood single-cell samples using specific biomarkers for T cells and myeloid cells, respectively, to obtain raw mass cytometry data of the samples. The raw mass cytometry data of T cells and myeloid cells are denoted as Data. rawT and Data rawM ;
[0056] The raw data conversion unit is used to process the raw data from the cell mass cytometry flow cytometry of the sample. rawT and Data rawM The data were converted to obtain the converted single-cell mass cytometry data. transT and Data transM The conversion formulas are as follows: Data transT =sinh -1 (Data rawT / 5), Data transM =sinh -1 (Data rawM / 5);
[0057] Peripheral blood immune cell cluster acquisition unit, used to separately process Data transT and Data transM Cluster analysis was performed to obtain the proportions of T cells, NK cells, B cells, and myeloid cells in peripheral blood samples, denoted as Data. clusterTcell Data clusterNKcell Data clusterBcell and Data clusterMcell ;
[0058] Subgroup merging unit is used to collect subgroup information obtained from cluster analysis of two sets of single-cell mass cytometry data of peripheral blood samples. clusterTcell Data clusterNKcell Data clusterBcell and Data clusterMcell By merging the data, we obtain the peripheral blood immune cell subset information, i.e., the peripheral blood immune cell atlas, denoted as Data. immune-cell ;
[0059] The feature vector acquisition unit is used to obtain features based on peripheral blood immune cell atlas data. immune-cell Calculate the feature vector Q composed of the proportion of each immune cell subpopulation in the sample, where the proportion of a certain immune cell subpopulation = the total number of cells in the sample belonging to a certain cell subpopulation / the total number of cells in the sample.
[0060] The classification acquisition unit is used to substitute the feature vector Q into a pre-trained random forest binary classifier to calculate the sample classification and obtain the prediction results of the early screening system for atherosclerotic coronary heart disease.
[0061] Using the above system, 13 healthy individuals, 38 patients with low-risk atherosclerotic coronary heart disease, and 32 patients with high-risk atherosclerotic coronary heart disease underwent early screening for atherosclerotic coronary heart disease, and the results were verified. The specific steps are as follows: (1) Obtain peripheral blood single cell samples:
[0062] Peripheral blood samples were obtained from 13 healthy individuals, 38 patients with low-risk atherosclerotic coronary artery disease, and 32 patients with high-risk atherosclerotic coronary artery disease. The samples were processed using Ficoll separation solution gradient centrifugation to obtain cell pellets, i.e., single cells.
[0063] (2) Mass cytometry detection:
[0064] Using pre-designed metal isotope-coupled protein-specific biomarkers for T cells and myeloid cells (Tables 1 and 2), peripheral blood cell pellets were analyzed by mass cytometry (CyTOF) to obtain raw single-cell mass cytometry data. rawT and DatarawM :
[0065]
[0066] Among them, Data rawT For raw single-cell mass cytometry data of sample T cells, Data rawM This is raw mass cytometry data for single-cell myeloid cells in the sample, where m is the number of metal isotope-coupled protein-specific biomarkers, n is the number of cells in a single sample, and a ij The expression level of the metal isotope-coupled protein-specific biomarker in the j-th cell of a single sample.
[0067] The CyTOF detection experimental procedure is as follows: Figure 1 As shown, the specific steps are as follows:
[0068] a) Count the cells in the cell pellet obtained in step (1), and take 2 × 10⁻⁶ cells. 6 Cell count;
[0069] b) Prepare a 0.25 μM 194Pt (1 mM) live / dead staining solution using phosphate-buffered saline (PBS) (pH 7.4). Resuspend the cells in 100 μL of the 194Pt live / dead staining solution and stain on ice for 5 min. This solution will be used to distinguish between live and dead cells in subsequent data analysis.
[0070] c) Add 500 μL FACS Buffer to each sample, resuspend the cells, centrifuge at 400 g / 5 min at 4°C, and discard the supernatant;
[0071] d) Add 50 μL of blocking solution to each sample, resuspend the cells, and block on ice for 20 min;
[0072] e) Add the metal isotope-coupled protein specific biomarker (Table 1 or Table 2) mixture (2 μL of each metal antibody), stain on ice for 30 min, and perform extracellular staining on the sample;
[0073] f) Add 1 mL of FACS Buffer to each sample, resuspend the cells, centrifuge at 400 g for 5 min at 4°C, discard the supernatant, and repeat 2-3 times;
[0074] g) Prepare a final concentration of 300 nM Ir staining solution using Fix and Perm Buffer. Take 500 μL of each sample to resuspend the cells, incubate at room temperature for 1 hour, stain and fix the DNA;
[0075] h) Add 5 mL of FACS Buffer to each sample, resuspend the cells, centrifuge at 800 g / 5 min at 4 °C, and discard the supernatant;
[0076] i) Add 0.5 mL of ddH2O to each sample to resuspend the cells and transfer them to a 5 mL flow cytometer tube (12×75 mm) with a filter, and filter twice;
[0077] j) The filtered cell suspension was analyzed using mass flow cytometry.
[0078] Table 1. Specific biomarkers for metal isotope-coupled proteins in T cells.
[0079]
[0080] Table 2. Specific biomarkers for metal isotope-coupled proteins in myeloid cells.
[0081]
[0082] (3) For Data respectively rawT and Data rawM Perform raw data transformation:
[0083] Use the following formulas to apply the data respectively. rawT and Data rawM The data was converted to obtain the converted single-cell mass cytometry data. transT and Data transM :
[0084] Data transT =sinh -1 (Data rawT / 5), Data transM =sinh -1 (Data rawM / 5).
[0085] (4) For Data respectively transT and Data transM Cluster analysis was performed to obtain peripheral blood immune cell populations: Data transT and Data transM Each sample in the study contained 100,000 cells, for a total of 8,300,000 cells.
[0086] Data transT Using the PARC algorithm for clustering, subpopulations with a percentage below 0.1% were filtered out, resulting in 46 subpopulations. Subpopulations were then merged based on the expression of metal isotope-coupled protein-specific biomarkers, yielding 11 T cell subpopulations and 4 NK cell subpopulations. Further PARC clustering was applied to the T cell and NK cell subpopulations, filtering out subpopulations with a percentage below 0.1%, resulting in 26 and 13 subpopulations respectively. The corresponding peripheral blood sample subpopulation percentages are denoted as Data.clusterTcell and Data clusterNKcell .
[0087] Data transM Using the PARC algorithm for clustering, subpopulations with a percentage below 0.1% were filtered out, resulting in 54 subpopulations. Subpopulations were then merged based on the expression of metal isotope-coupled protein-specific biomarkers, yielding 2 B-cell subpopulations and 5 myeloid cell subpopulations. Further PARC clustering was applied to the B-cell and myeloid cell subpopulations, filtering out subpopulations with a percentage below 0.1%, resulting in 12 and 15 subpopulations respectively. The corresponding peripheral blood sample subpopulation percentages are denoted as Data. clusterBcell and Data clusterMcell .
[0088] (5) Merging subgroups:
[0089] Hierarchical clustering method is used for Data clusterTcell Data clusterNKcell Data clusterBcell and Data clusterMcell Merging and clustering were performed to obtain peripheral blood sample immune cell subset information, namely peripheral blood immune cell atlas data, which includes 47 immune cell subsets: 26 T cell subsets, 2 B cell subsets, 4 NK cell subsets, and 15 monocyte subsets. immune-cell :
[0090]
[0091] Where, β ij β represents the expression value of the metal isotope-coupled protein-specific biomarker in the j-th cell of a single sample. cj This refers to the cell subpopulation to which the j-th cell in a single sample belongs.
[0092] (6) Obtain the feature vector:
[0093] Based on peripheral blood immune cell atlas data immune-cell The feature vector Q is calculated based on the proportion of each immune cell subset in the sample. Specifically, for a single sample, the feature vector Q is calculated based on its respective Data. immune-cell Calculate the percentage of each immune cell subset q k The eigenvectors Q(q1,q2,q3,…,q) are composed of (k=1,2,3,…,30). i ,…,q 30 ), q k = Total number of cells belonging to immune cell subset k in a single sample / Total number of cells in a single sample.
[0094] (7) Features were screened from 47 immune cell subsets for training the Random Forest binary classifier:
[0095] A certain number of samples were randomly selected (disease prediction model: 50 cases; disease progression prediction: 25 cases), normalized, and then a Random Forest model was used to build a classifier. At the same time, 10-fold cross-validation was used to obtain the average feature importance of the Random Forest model. Features with an average feature importance greater than 0.01 were selected as important features. This process was repeated 1000 times, and the frequency of each important feature was counted. If the frequency was greater than 300, the important feature was selected.
[0096] (8) Obtain the pre-trained Random Forest binary classifier:
[0097] Two disease risk prediction models were established: a disease prediction (DP) model to distinguish between healthy individuals and patients with atherosclerosis; and a disease progression prediction (DPP) model to distinguish between low-risk and high-risk patients with atherosclerotic coronary artery disease.
[0098] For the disease prediction model, 25 low-risk atherosclerotic coronary heart disease patients, 25 high-risk atherosclerotic coronary heart disease patients, and 50 healthy individuals were randomly selected (sample data were obtained from the bootstrap resampling strategy of 13 healthy individuals). The feature vector Q of each sample was obtained using the method in steps (1) to (6), which served as the training dataset for training the Random Forest binary classifier. The feature vector Q of the remaining samples (13 low-risk atherosclerotic coronary heart disease patients, 7 high-risk atherosclerotic coronary heart disease patients, and 10 healthy individuals) was obtained using the method in steps (1) to (6), which served as the test dataset for testing the trained Random Forest binary classifier. The feature vector Q of these samples was used to train the Random Forest binary classifier, where the algorithm parameters n_estimators = 200, max_depth = 7, and min_samples_leaf = 5.
[0099] For the disease progression prediction model, this patent randomly selected 25 low-risk atherosclerotic coronary heart disease patients and 25 high-risk atherosclerotic coronary heart disease patients. The feature vector Q of each sample was obtained using the methods in steps (1) to (6), which served as the training dataset for training the Random Forest binary classifier. The feature vector Q of each of the remaining samples (13 low-risk atherosclerotic coronary heart disease patients and 7 high-risk atherosclerotic coronary heart disease patients) was obtained using the methods in steps (1) to (6), which served as the test dataset for testing the trained Random Forest binary classifier. The feature vector Q of these samples was used to train the Random Forest binary classifier, where the algorithm parameters n_estimators = 200, max_depth = 7, and min_samples_leaf = 5.
[0100] (9) Obtain sample classification:
[0101] Substitute the feature vector Q obtained in step (6) into the pre-trained Random Forest binary classifier to calculate the sample classification and obtain the prediction results of the early screening system for atherosclerotic coronary heart disease.
[0102] The immune feature prediction results of the 83 samples obtained through the above steps, based on the Random Forest binary classifier, were combined with the prediction results of the patient's clinical manifestations and the combined prediction results of immune and clinical manifestations to plot receiver operating characteristic (ROC) curves. These ROC curves were then used to compare the performance of the prediction methods. The results are as follows: Figure 2 As shown, in both disease prediction (DP) and disease progression prediction (DPP), the ROC curves for the immune-clinical feature prediction results (AUC = 0.99 in DP, AUC = 0.88 in DPP) are above the ROC curves for the boundary classification using immune feature prediction and clinical feature prediction, respectively (AUC = 0.99 and AUC = 0.96 in DP, AUC = 0.80 and AUC = 0.75 in DPP), and have larger AUC values, indicating that the classification method has higher operational performance.
[0103] Sensitivity and accuracy analyses were performed on the immune feature prediction results based on the Random Forest binary classifier for 83 samples obtained through the above steps, as well as the prediction results of the combined immune-clinical feature prediction results. The results are as follows: Figure 3As shown, the sensitivity and accuracy of the early screening system for atherosclerotic coronary heart disease based on the combined prediction of immune and clinical manifestation characteristics (DP model: sensitivity = 98.6%, accuracy = 100%; DPP model: sensitivity = 93.8%, accuracy = 89.5%) are both superior to the prediction results using immune characteristics (DP model: sensitivity = 98.6%, accuracy = 92.3%; DPP model: sensitivity = 87.5%, accuracy = 86.8%), which are also superior to the prediction results using clinical manifestation characteristics (DP model: sensitivity = 94.2%, accuracy = 92.3%; DPP model: sensitivity = 87.5%, accuracy = 86.8%). This indicates that the early screening system for atherosclerotic coronary heart disease of the present invention has high accuracy and sensitivity, and combining it with the prediction results of clinical manifestation characteristics can further improve accuracy and sensitivity.
[0104] Unless otherwise specified, the raw materials and equipment used in this invention are all commonly used in the field; unless otherwise specified, the methods used in this invention are all conventional methods in the field.
[0105] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Any simple modifications, alterations, and equivalent transformations made to the above embodiments based on the technical essence of the present invention shall still fall within the protection scope of the present invention.
Claims
1. An early screening system for atherosclerotic coronary heart disease based on peripheral blood immune cell atlas, characterized in that, include: The peripheral blood single-cell acquisition unit is used to separate single cells from the peripheral blood of patients to be diagnosed and obtain peripheral blood single-cell samples. The mass cytometry unit is used to split peripheral blood single-cell samples into two halves. The two halves are then subjected to mass cytometry analysis using metal isotope-coupled protein-specific biomarkers targeting T cells and myeloid cells, respectively. The raw mass cytometry data for T cells and myeloid cells are denoted as follows: Data rawT and Data rawM ; The following are specific biomarkers for metal isotope-coupled proteins (MCPs) in myeloid cells and T cells, with "||" indicating myeloid cells on the left and T cells on the right: 1# Label-89Y||89Y Marker-CD45||CD45 2# Label-115In||115In Marker-CD3||CD3 3# Label-139La||139La Marker-CD47||CD47 4# Label-141Pr||141Pr Marker-CD56||CD56 5# Label-142Nd||142Nd Marker-CD19||CD19 / TCRγδ 6# Label-143Nd||143Nd Marker-CD184||CD196 7# Label-144Nd||145Nd Marker-CD38||CD95 8# Label-145Nd||146Nd Marker-CD115||CD123 9# Label-146Nd||147Sm Marker-CD54||CD66b 10# Label-147Sm||148Nd Marker-CD15||CD33 11# Label-148Nd||149Sm Marker-CD33||CD25 12# Label-149Sm||150Nd Marker-CD169||CD14 13# Label-150Nd||151Eu Marker-CD14||CD38 14# Label-151Eu||152Sm Marker-CD36L1||CD39 15# Label-152Sm||153Eu Marker-FceRIa||CD274 16# Label-153Eu||154Sm Marker-CD274||Ki67 17# Label-154Sm||155Gd Marker-CD163||CD45RA 18# Label-155Gd||156Gd Marker-CD206||CD11c 19# Label-156Gd||157Gd Marker-CD24||CD68 20# Label-157Gd||158Gd Marker-CD172a||CD197 21# Label-158Gd||159Tb Marker-CD204||CD357 22# Label-159Tb||160Gd Marker-CD11c||CD28 23# Label-160Gd||161Dy Marker-CD319||CD152 24# Label-161Dy||162Dy Marker-CD66b||FoxP3 25# Label-162Dy||163Dy Marker-CD32||CD183 26# Label-163Dy||164Dy Marker-CD68||RORγ 27# Label-164Dy||165Ho Marker-CD192||CD161 28# Label-165Ho||166Er Marker-ProMBP-1||CD27 29# Label-166Er||167Er Marker-CX3CR1||CD278 30# Label-167Er||168Er Marker-CD36||T-bet 31# Label-168Er||169Tm Marker-CD95||CD15 32# Label-169Tm||170Er Marker-CD40||CD127 33# Label-170Er||171Yb Marker-CD86||GATA-3 34# Label-171Yb||172Yb Marker-CD64||CD272 35# Label-172Yb||173Yb Marker-CD117||Granzyme B 36# Label-173Yb||174Yb Marker-Siglec-8||CD279 37# Label-174Yb||175Lu Marker-CD279||CD16 38# Label-175Lu||176Yb Marker-CD16||HLA-DR 39# Label-176Yb||197Au Marker-HLA-DR||CD4 40# Label-198Pt||198Pt Marker-Ki67||CD8a 41# Label-209Bi||209Bi Marker-CD11b||CD11b The raw data conversion unit is used to process the raw cell mass cytometry data of the samples. Data rawT and Data rawM The data were converted to obtain the converted single-cell mass cytometry data. Data transT and Data transM ; The conversion formulas are as follows: Data transT = sinh -1 ( Data rawT / 5), Data transM = sinh -1 ( Data rawM / 5); Peripheral blood immune cell fragmentation unit, used for separate processing of... Data transT and Data transM Cluster analysis was performed to obtain the proportions of T cells, NK cells, B cells, and myeloid cells in peripheral blood samples, denoted as follows: Data clusterTcell , Data clusterNKcell , Data clusterBcell and Data clusterMcell ; Subgroup merging unit, used for... Data clusterTcell , Data clusterNKcell , Data clusterBcell and Data clusterMcell By merging the data, we obtain the peripheral blood immune cell subset information, i.e., the peripheral blood immune cell atlas, denoted as... Data immune-cell ; Feature vector acquisition unit, used to obtain features based on peripheral blood immune cell atlas Data immune-cell Calculate the feature vector Q composed of the proportion of each immune cell subpopulation in the sample, where the proportion of a certain immune cell subpopulation = the total number of cells in the sample belonging to a certain cell subpopulation / the total number of cells in the sample. The classification acquisition unit is used to substitute the feature vector Q into a pre-trained random forest binary classifier to calculate the sample classification and obtain the prediction results of the early screening system for atherosclerotic coronary heart disease.
2. The early screening system for atherosclerotic coronary heart disease as described in claim 1, characterized in that, The peripheral blood single-cell acquisition unit uses a gradient centrifugation method with sucrose separation solution to separate mononuclear cells from the peripheral blood of the patient to be diagnosed, thus obtaining a peripheral blood sample single cell.
3. The early screening system for atherosclerotic coronary heart disease as described in claim 1, characterized in that, In the peripheral blood immune cell population acquisition unit, for Data transT The clustering analysis method includes the following steps: clustering is performed using a phenotypic analysis algorithm based on accelerated fine-grained community partitioning; subpopulations are merged based on the expression of metal isotope-coupled protein-specific biomarkers; T cell and NK cell-related subpopulations are selected and subjected to secondary clustering using the phenotypic analysis algorithm based on accelerated fine-grained community partitioning; the corresponding peripheral blood sample subpopulation proportions are denoted as follows. Data clusterTcell and Data clusterNKcell .
4. The early screening system for atherosclerotic coronary heart disease as described in claim 1, characterized in that, In the peripheral blood immune cell population acquisition unit, for Data transM The clustering analysis method includes the following steps: clustering is performed using a phenotypic analysis algorithm based on accelerated fine-grained community partitioning; subpopulations are merged based on the expression of metal isotope-coupled protein-specific biomarkers; B-cell and myeloid cell-related subpopulations are selected and subjected to secondary clustering using the phenotypic analysis algorithm based on accelerated fine-grained community partitioning; the corresponding peripheral blood sample subpopulation proportions are denoted as follows. Data clusterBcell and Data clusterMcell .
5. The early screening system for atherosclerotic coronary heart disease as described in claim 3 or 4, characterized in that, During the clustering process using a phenotypic analysis algorithm based on accelerated fine community segmentation, subgroups with a subgroup percentage of less than 0.1% are filtered out.
6. The early screening system for atherosclerotic coronary heart disease as described in claim 1, characterized in that, In the subgroup merging unit, for Data clusterTcell , Data clusterNKcell , Data clusterBcell and Data clusterMcell The merging method used is hierarchical clustering.
7. The early screening system for atherosclerotic coronary heart disease as described in claim 1, characterized in that, In the classification acquisition unit, feature selection is performed before substituting the feature vector Q into the pre-trained random forest binary classifier. During the acquisition of the pre-trained random forest binary classifier, the selected features are used to train the random forest binary classifier. The feature selection method includes the following steps: randomly selecting multiple samples, performing normalization processing, and then using a random forest model to construct a classifier. At the same time, 10-fold cross-validation is used to obtain the average feature importance of the random forest model. Features with an average feature importance greater than 0.01 are regarded as important features. This process is iterated 800 to 1200 times, and the frequency of each important feature is counted. If the frequency is greater than 300, the important feature is selected.
8. The early screening system for atherosclerotic coronary heart disease as described in claim 1, characterized in that, The aforementioned early screening system for atherosclerotic coronary heart disease is a disease prediction system used to distinguish between healthy individuals and patients with atherosclerosis. The method for acquiring the pre-trained random forest binary classifier in the classification acquisition unit includes the following steps: randomly selecting several low-risk atherosclerotic coronary heart disease patients, several high-risk atherosclerotic coronary heart disease patients, and several healthy individuals; obtaining the feature vector Q of each sample through a peripheral blood single-cell acquisition unit, a mass spectrometry flow cytometry detection unit, a raw data conversion unit, a peripheral blood immune cell clustering acquisition unit, a subpopulation merging unit, and a feature vector acquisition unit, as the training dataset for training the random forest binary classifier; randomly selecting several low-risk atherosclerotic coronary heart disease patients, several high-risk atherosclerotic coronary heart disease patients, and several healthy individuals; obtaining the feature vector Q of each sample through the same method, as the test dataset for testing the trained random forest binary classifier; and training the random forest binary classifier using the feature vector Q of all samples.
9. The early screening system for atherosclerotic coronary heart disease as described in claim 1, characterized in that, The aforementioned early screening system for atherosclerotic coronary heart disease is a disease progression prediction system used to distinguish between low-risk and high-risk patients with atherosclerotic coronary heart disease. The method for acquiring the pre-trained random forest binary classifier in the classification acquisition unit includes the following steps: randomly selecting several low-risk and high-risk patients with atherosclerotic coronary heart disease, and obtaining the feature vector Q of each sample through a peripheral blood single-cell acquisition unit, a mass spectrometry flow cytometry detection unit, a raw data conversion unit, a peripheral blood immune cell cluster acquisition unit, a subgroup merging unit, and a feature vector acquisition unit, as the training dataset for training the random forest binary classifier; randomly selecting several low-risk and high-risk patients with atherosclerotic coronary heart disease, and obtaining the feature vector Q of each sample using the same method, as the prediction model for testing the trained random forest binary classifier; and training the random forest binary classifier using the feature vector Q of all samples.
Citation Information
Patent Citations
Hepatocellular carcinoma early screening system based on peripheral blood immune cell map
CN114171189A
Method for establishing liver cancer screening model based on mass spectrum flow technology
CN114937493A