Method for constructing differential diagnosis model of lupus nephritis and membranous nephropathy
The differential diagnosis model for lupus nephritis and membranous nephropathy constructed through flow cytometry and feature screening solves the technical difficulties of early differentiation, achieves high-accuracy and low-invasive diagnosis, and is suitable for immediate diagnosis scenarios.
Patent Information
- Application Number
- CN202510691904.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-09
AI Technical Summary
Existing technologies make it difficult to accurately distinguish between lupus nephritis and membranous nephropathy in the early stages, especially when SLE symptoms are not obvious. In addition, renal puncture and pathological biopsy are very traumatic to patients, making it difficult to distinguish between the two based on clinical characteristics.
Through flow cytometry data, features were screened and a differential diagnosis model for lupus nephritis and membranous nephropathy was constructed. TabPFNClassifier was used for classification, and the RFE method was used to select features. Combined with pDC, CD4T, effector CD4T, Th2, CD8T, CD38+HLA-DR+CD8T, and CD38+PD-1+CD8T cells, high-accuracy identification was achieved.
It achieves high-accuracy differentiation between lupus nephritis and membranous nephropathy in the early stages, reduces trauma to patients, is suitable for immediate diagnosis scenarios, and has a short response time.
Smart Images

Figure CN120613100A_ABST
Abstract
Description
A method for constructing a differential diagnosis model for lupus nephritis and membranous nephropathy Technical Field
[0001] The present invention belongs to the field of intelligent medical technology, and specifically relates to a method for constructing a differential diagnosis model for lupus nephritis and membranous nephropathy. Background Art
[0002] Systemic lupus erythematosus (SLE) is a chronic autoimmune disease that affects multiple organ systems, particularly women of childbearing age. Early identification of kidney damage in SLE patients is crucial for improving prognosis.
[0003] The current gold standard for diagnosing lupus nephritis (LN) and membranous nephropathy (MN) is renal puncture followed by pathological biopsy, which is highly invasive to the patient. In the early stages of the disease, both are characterized by heavy proteinuria and low serum albumin. Clinically, distinguishing lupus nephritis from membranous nephropathy is a crucial research topic. This is particularly true in early-stage lupus nephritis, where SLE symptoms are less pronounced, making it difficult to distinguish from membranous nephropathy. Furthermore, many MN patients are idiopathic MN patients who are negative for PLA2R-specific antibodies, further complicating the clinical distinction between LN and MN.
[0004] Therefore, it is necessary to rebuild a model that can accurately distinguish lupus nephritis (LN) and membranous nephropathy (MN) to address the shortcomings of existing diagnostic models. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method for constructing a differential diagnosis model for lupus nephritis and membranous nephropathy. The present invention performs flow cytometry on lupus nephritis patients and healthy people, uses the detection data to train the model, screens the characteristics that distinguish lupus nephritis (LN) and membranous nephropathy (MN), and evaluates and constructs a differential diagnosis model with high accuracy.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for constructing a differential diagnosis model for lupus nephritis and membranous nephropathy, comprising the following steps:
[0008] S1. Collect flow cytometry data from lupus nephritis patients and healthy controls;
[0009] S2. Perform data cleaning and transformation on the flow cytometry data, converting non-numerical features into numbers; imputation of missing values using the k-nearest neighbor algorithm from the impute package;
[0010] S3. Randomly sample from the cleaned and transformed dataset and split the samples into training and validation sets in a ratio of 7:3.
[0011] S4. Split the training set into a training subset and a test set, and iteratively select a specified number of features using the RFE method;
[0012] S5. Use TabPFNClassifier for classification tasks, train on features selected from the training dataset, and then evaluate the model performance;
[0013] The cell types finally included in the constructed model in S4 are pDC, CD4T, effector CD4T, Th2, CD8T, CD38+HLA-DR+CD8T, and CD38+PD-1+CD8T.
[0014] Furthermore, the flow cytometry data in S1 is constructed as follows:
[0015] (1) Take 5 mL of EDTA-anticoagulated whole blood from patient 2 and dilute it to 5 mL with PBS. Add 5 mL of ficoll and 5 mL of diluted blood sample to a 15 mL centrifuge tube and centrifuge.
[0016] (2) After the centrifugation in step (1), aspirate the PBMCs and add PBS to the centrifuge volume of 15 mL;
[0017] (3) After the centrifugation in step (2), remove the supernatant, add 5 mL of red blood cell lysis buffer and gently blow the cells apart. After 2 minutes, add PBS to 15 mL and centrifuge;
[0018] (4) After the centrifugation in step (3), remove the supernatant and resuspend the cells in 100 μL of flow cytometry staining buffer. Then, filter the cell suspension using a 70-mesh sieve and transfer it to a 1.5 mL eppendorf tube. Add 4 μL of HumanTruStainFcX to each tube. TM , incubate at room temperature for 5 min;
[0019] Then, three incubation antibody combinations were added: 1. FITC CD4, APC-Cy7 CD3, PE CD25, PerCP-Cy5.5 CD27, APC CXCR5, and PE-Cy7 CD196; 2. APC CD8, PE HLA-DR, PE-Cy7 CD138, and FITC PD-1; 3. APC CD14, PerCP-Cy5.5 CD16, APC-Cy7 CD11C, and FITC CD123, and the cells were mixed by flicking and incubated at 4°C for 25 min.
[0020] (5) After step (4), 1 mL of PBS was added to each tube to wash the unbound antibody, the supernatant was discarded by centrifugation, 300 μL of PBS was added to each tube, and the tubes were transferred to flow cytometry tubes. The fluorescence compensation adjustment was completed on the flow cytometer, and the test results were analyzed;
[0021] (6) The flow cytometric analysis method includes data acquisition, analysis and recording; wherein multiple data channels record the light scattering and fluorescence emitted by each particle when it passes through the sensing area from each detector, and the analysis system classifies and counts the particles, wherein each particle is presented as a series of digitized parameter values.
[0022] Further, the flow cytometer is configured to be triggered at selected parameters to distinguish target particles from background and noise;
[0023] The trigger is used to detect a preset threshold value of a parameter, and is used as a means of detecting a particle passing through the laser beam, and an event exceeding the preset threshold value of the selected parameter is detected;
[0024] For particles or other components in the assay medium whose response is below a threshold, no data is collected. The trigger parameter may be detecting forward scattered light caused by the particle passing through the light beam, and then detecting and collecting light scattering and fluorescence data of the particle.
[0025] Furthermore, the non-numerical feature conversion in S2 uses LabelEncoder of sklearn.preprocessing to encode the non-numerical target label into a numerical value from 0 to n_classes-1.
[0026] Furthermore, the k-nearest neighbor algorithm in S2 finds the nearest k samples for each missing value and fills the missing value with the mean or mode of these neighbors.
[0027] Furthermore, a random seed set.seed(123) is set during random sampling in S3 to ensure the repeatability of the results.
[0028] Furthermore, the test set size in S4 is 33%.
[0029] Furthermore, the evaluation indicators in S5 include: accuracy, confusion matrix, precision, recall rate, F1-Score and AUC.
[0030] Contains at least the following beneficial technical effects:
[0031] By performing flow cytometry on lupus nephritis patients and healthy people, we finally included pDC, CD4T, effector CD4T, Th2, CD8T, CD38+HLA-DR+CD8T, and CD38+PD-1+CD8T through screening features. We used the test data for model training and evaluation to construct a highly accurate identification model for lupus nephritis (LN) and membranous nephropathy (MN). BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a flow chart of the flow gate;
[0033] Figure 2 shows the model construction and performance verification; Figure 3 shows the ROC curves for individual diagnosis of 7 cell types. DETAILED DESCRIPTION
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0035] Example
[0036] S1. Data Collection
[0037] A total of 181 patients with lupus nephritis and 100 patients with membranous nephropathy were selected to obtain flow cytometry data; specifically:
[0038] (1) Take about 2.5 mL of EDTA anticoagulated whole blood from the patient and dilute it with PBS to 5 mL. Add 5 mL of ficoll to a 15 mL centrifuge tube, then slowly layer 5 mL of the diluted blood sample, and then centrifuge at 1440 RPM for 20 min.
[0039] (2) After centrifugation, four layers can be observed in the centrifuge tube: from top to bottom, the plasma and platelet layer, the PBMC layer, the ficol layer, the polymorphonuclear leukocyte layer, and the red blood cell layer. The PBMC (buffy coat layer) is aspirated, of which 70-90% are lymphocytes (70-90%), of which T cells account for the vast majority of lymphocytes (45-70%). PBS is then added to 15 mL, and the tube is centrifuged at 500 x g for 5 min.
[0040] (3) After centrifugation, remove the supernatant and add about 5 mL of red blood cell lysis buffer to gently blow the cells apart. After 2 minutes, add PBS to 15 mL and then centrifuge at 500 × g for 5 minutes.
[0041] (4) After centrifugation, remove the supernatant and resuspend the cells in about 100 μL of flow cytometry staining buffer. Then filter the cell suspension through a 70-mesh sieve and transfer it to a 1.5 mL ep tube and add Human TruStain FcX (Fc Receptor Blocking Solution), adding 4 μL to each tube and incubating at room temperature for 5 minutes; then add three incubation antibody combinations: 1. FITC CD4, APC-Cy7 CD3, PE CD25, PerCP-Cy5.5 CD27, APC CXCR5 and PE-Cy7 CD196; 2. APC CD8, PE HLA-DR, PE-Cy7 CD138 and FITC PD-1; 3. APC CD14, PerCP-Cy5.5 CD16, APC-Cy7CD11C and FITC CD123; gently flick to mix and incubate at 4°C for 25 minutes.
[0042] (5) After incubation, add 1 mL of PBS to each tube to wash away unbound antibodies and centrifuge at 0.5 x g for 5 min. Discard the supernatant, add 300 μL of PBS to each tube, transfer to a flow cytometer, and perform flow cytometric analysis. Complete fluorescence compensation adjustment and analyze the test results. All test indicators are strictly set up with blank, isotype, FMO, negative, and positive control groups.
[0043] (6) The flow cytometric analysis method includes data acquisition, analysis and recording; wherein multiple data channels record the light scattering and fluorescence emitted by each particle when it passes through the sensing area from each detector, and the analysis system classifies and counts the particles, wherein each particle is presented as a series of digitized parameter values.
[0044] The flow cytometer is configured to trigger at selected parameters to distinguish target particles from background and noise;
[0045] The trigger is used to detect a preset threshold value of a parameter, and is used as a means of detecting a particle passing through the laser beam, and an event exceeding the preset threshold value of the selected parameter is detected;
[0046] For particles or other components in the assay medium whose response is below a threshold, no data is collected. The trigger parameter may be detecting forward scattered light caused by the particle passing through the light beam, and then detecting and collecting light scattering and fluorescence data of the particle.
[0047] Figure 1 (A) shows the flow cytometry gating process for CD4 T cells, effector CD4 T cells, and TH2 cells;
[0048] Figure 1(B) shows the flow cytometry gating process for CD8 T cells, CD38+HLA-DR+CD8+ T cells, and CD38+PD-1+CD8+ T cells;
[0049] Figure 1(C) shows the flow cytometry gating process for plasmacytoid dendritic cells (PDCs).
[0050] S2. Data cleaning and transformation
[0051] Perform data cleaning and transformation on the flow cytometry data, use LabelEncoder from sklearn.preprocessing to encode non-numeric target labels into values from 0 to n_classes-1; imputation is performed on missing values using the k-nearest neighbor algorithm from the impute package;
[0052] S3. Divide the training set and validation set
[0053] Random sampling was performed from the cleaned and converted dataset, and the samples were divided into training and validation sets in a ratio of 7:3; the number of lupus nephritis patients and 64 membranous nephropathy patients included in the training set was 105; the number of lupus nephritis patients and 36 membranous nephropathy patients included in the validation set was 76.
[0054] S4. Feature Selection
[0055] The training set is divided into training subsets and test sets, and the RFE method is used to iteratively select a specified number of features.
[0056] The training features include:
[0057] CD14pCD16pcell (CD14 + CD16 + )、CD14nCD16pcell(CD14 - CD16 + )、CD14nCD16ncell(CD14 - CD16 - )、CD14pCD16ncell(CD14 + CD16 - ), NK(HLA - DR - CD56 + )、Pdc(CD 14 - CD 16 - CD11c int CD 123 + ) and cDCs (CD14 - CD16 - CD11c + CD123 - HLA - DR + );CD4T(CD19- CD3 + CD8 - CD4 + )、CD8T(CD19 - CD3 + CD4 - CD8 + )、 CD8T(CD3 + CD4 - CD8 + CD27 + CD45RA + )、E M R A C D 8T(C D 3 + CD 4 - C D 8 + C D 27 - C D 45R A + )、C MC D 8T(C D 3 + C D 4 - C D 8 + C D 27 + C D 45R A - )、E MCD 8T(C D 3 + C D 4 - C D 8 + C D 27 - C D 45R A - )、 C D 4T(C D 3 + C D 8 - C D 4 + C D27 + C D 45R A + )、E M R AC D 4T(C D 3 + C D 4 + C D 8 - C D 27 - C D 45R A + )、C MC D 4T(C D3 + C D 4 + C D 8 - C D 27 + C D 45R A - )、E MCD4T(CD3 + CD4 + CD8 - CD27 - CD45RA -)、CXCR5+CD4+T (CXCR5 + CD4 + )、Tfh (CXCR5 + CD4 + PD - 1 + )、Eff CD4 T (CD4 + CD25low)、Treg (CD4 + CD25 + )、Th1 (CD4 + CD25 low CXCR3 + CD196 - )、Th2 ((CD4 + CD25 low CXCR3 - CD196 - )、Th17 (CD4 + CD25 [[ID=3... (注:原文中存在一些不完整或表述不太清晰的地方,翻译尽量忠实于原文呈现。) (后面还有部分内容未完整翻译,你可以补充完整原文以便我继续为你准确翻译。) low CXCR3 - CD196 + )、CD3+T (CD3 + )、DNT (CD3)、DNT (CD3 + CD4 - CD8 - )、DPT (CD3 + CD4 + CD8 + )、Act CD8 T (CD8 + CD38 + HLA - DR + )、Ex CD8 T (CD8 + CD38 + HLA - DR + );CD19+B (CD19<0000... (翻译继续) + )、ASC (CD19 + CD38 + CD27 + )、non - ASC (CD19 + 除去CD38 + CD27 + [[ID=... (翻译继续) )、Trans B (CD19 + CD38 + CD27 - CD24 + )、SWM B (CD19 + IgD - CD27 +)、UNSWM B(CD19 + IgD + CD27 + )、DNB(CD19 + IgD - CD27 - ), B(CD19 + IgD + CD27 - ), DN1(CD19 + IgD - CD27 - CD21 + CXCR5 + )、DN 2(CD 19 + Ig D - CD 27 - CD 21 - CXCR 5 in t ) and ABC(CD 19 + Ig D - CD 27 - CD 21 - CXC R5 in t CD 11c + ).
[0058] Recursive feature elimination (RFE) was used to screen features for inclusion in a differential diagnosis model for lupus nephritis and membranous nephropathy. RFE selects features by recursively considering smaller and smaller feature sets. RFE aims to select the most important features by recursively reducing the size of the feature set examined. It uses model accuracy to identify which features contribute most to the prediction outcome.
[0059] The working process of RFE is as follows:
[0060] 1. Train a model and use the model's coef_ or feature_importances_ attribute to get the importance of each feature;
[0061] 2. Based on the obtained feature importance, remove the least important feature from the current feature set.
[0062] 3. Repeat steps 1 and 2 until the desired number of features is reached.
[0063] 4. At each step, RFE trains a new model and evaluates the importance of the remaining features. This process continues until the preset number of features is reached.
[0064] This approach can effectively reduce the number of features while retaining the features that contribute most to the model.
[0065] The features incorporated by the above methods include pDC, CD4T, effector CD4T, Th2, CD8T, CD38+HLA-DR+CD8T, and CD38+PD-1+CD8T.
[0066] S5. Training model and evaluation
[0067] TabPFNClassifier is used for classification tasks, trained on features selected from the training dataset, and then the performance of the model is evaluated.
[0068] 1. Input feature vector construction
[0069] Data source: Flow cytometry test results of patients, including immune cell subset proportion data.
[0070] Feature dimension: Input vector (x∈R d ), where (d) is the number of immune cell characteristics.
[0071] 2. Forward propagation calculation
[0072] Step 1: Input the preprocessed feature vector into the Transformer encoder of TabPFNClassifier and capture the nonlinear interaction relationship h between immune cells through the multi-head attention mechanism enc :
[0073] [h enc =TransformerEncoder(x norm )]
[0074] Step 2: Generate predicted probability distribution through the probabilistic inference layer:
[0075] [p(y|x)=softmax(W·h enc +b)]
[0076] Where (W) and (b) are the classification layer parameters, and (y∈{0,1]) represents lupus nephritis negative / positive.
[0077] 3. Real-time inference efficiency
[0078] The model only requires a single forward propagation in a single inference, with a response time of <10ms (based on GPU acceleration), making it suitable for clinical instant diagnosis scenarios.
[0079] 4. Hierarchical decision logic:
[0080] 1. If the model output is (P(y=1|x)>0.7), lupus nephritis is directly diagnosed;
[0081] 2. If the model output is (0.4 ≤ p(y = 1 | x) ≤ 0.7), additional renal biopsy verification is recommended;
[0082] 3. If the model output is (p(y=1|x)<0.4), exclude the diagnosis based on clinical signs.
[0083] Figure 2 (A) shows the ROC curves of the training and test data sets of the model for the combined differential diagnosis of lupus nephritis and membranous nephropathy using PDC, CD4T, Eff CD4T, Th2, CD8T, Act CD8T, and Ex CD8T cells;
[0084] Figure 2(B) shows the confusion matrix of the training set and test set data;
[0085] Figure 2(C) shows the ROC curve of the validation set data of the model for the combined differential diagnosis of lupus nephritis and membranous nephropathy using PDC, CD4T, Eff CD4T, Th2, CD8T, Act CD8T, and Ex CD8T cells;
[0086] Figure 2(D) is the confusion matrix of the validation set data.
[0087] Figure 3 shows the ROC curves for the diagnosis of seven cell types individually, where A, B, C, D, E, F, and G are the ROC curves for pDC cells, CD4T cells, Eff CD4T, Th2, CD8T, Act CD8T, and Ex CD8T, respectively. The diagnostic results show that the diagnostic effect of the seven cell types alone is not as good as the combined diagnostic effect of PDC, CD4T, Eff CD4T, Th2, CD8T, Act CD8T, and Ex CD8T cells in the present invention.
[0088] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A method for modeling a differential diagnosis model for lupus nephritis and membranous nephropathy, characterized in that: The following steps are involved: S1. Collect flow cytometry data from lupus nephritis patients and healthy controls; S2. Perform data cleaning and transformation on the flow cytometry data, converting non-numerical features into numbers; imputation of missing values using the k-nearest neighbor algorithm from the impute package; S3. Randomly sample from the cleaned and transformed dataset and split the samples into training and validation sets in a ratio of 7:
3. S4. Split the training set into a training subset and a test set, and iteratively select a specified number of features using the RFE method; S5. Use TabPFNClassifier for classification tasks, train on features selected from the training dataset, and then evaluate the model performance; The cell types finally included in the constructed model in S4 are pDC, CD4T, effector CD4T, Th2, CD8T, CD38+HLA-DR+CD8T, and CD38+PD-1+CD8T.
2. The modeling method according to claim 1, characterized in that The flow cytometry data in S1 is constructed as follows: (1) Take 5 mL of EDTA-anticoagulated whole blood from patient 2 and dilute it to 5 mL with PBS. Add 5 mL of ficoll and 5 mL of diluted blood sample to a 15 mL centrifuge tube and centrifuge. (2) After the centrifugation in step (1), aspirate the PBMCs and add PBS to the centrifuge volume of 15 mL; (3) After the centrifugation in step (2), remove the supernatant, add 5 mL of red blood cell lysis buffer and gently blow the cells apart. After 2 minutes, add PBS to 15 mL and centrifuge; (4) After the centrifugation in step (3), remove the supernatant and resuspend the cells in 100 μL of flow cytometry staining buffer. Then, filter the cell suspension using a 70-mesh sieve and transfer it to a 1.5 mL eppendorf tube. Add 4 μL of HumanTruStainFcX to each tube. TM , incubate at room temperature for 5 min; Then, three incubation antibody combinations were added:
1. FITC CD4, APC-Cy7 CD3, PE CD25, PerCP-Cy5.5CD27, APC CXCR5, and PE-Cy7 CD196; 2. APC CD8, PE HLA-DR, PE-Cy7 CD138, and FITC PD-1; 3. APC CD14, PerCP-Cy5.5 CD16, APC-Cy7 CD11C, and FITC CD123, and the cells were mixed by flicking and incubated at 4°C for 25 min. (5) After step (4), 1 mL of PBS was added to each tube to wash the unbound antibody, the supernatant was discarded by centrifugation, 300 μL of PBS was added to each tube, and the tubes were transferred to flow cytometry tubes. The fluorescence compensation adjustment was completed on the flow cytometer, and the test results were analyzed; (6) The flow cytometric analysis method includes data acquisition, analysis and recording; wherein multiple data channels record the light scattering and fluorescence emitted by each particle when it passes through the sensing area from each detector, and the analysis system classifies and counts the particles, wherein each particle is presented as a series of digitized parameter values.
3. The modeling method according to claim 2, characterized in that The flow cytometer is configured to trigger at selected parameters to distinguish target particles from background and noise; The trigger is used to detect a preset threshold value of a parameter, and is used as a means of detecting a particle passing through the laser beam, and an event exceeding the preset threshold value of the selected parameter is detected; For particles or other components in the assay medium whose response is below a threshold, no data is collected. The trigger parameter may be detecting forward scattered light caused by the particle passing through the light beam, and then detecting and collecting light scattering and fluorescence data of the particle.
4. The modeling method according to claim 1, characterized in that The non-numeric feature conversion in S2 uses LabelEncoder of sklearn.preprocessing to encode the non-numeric target label into a numerical value from 0 to n_classes-1.
5. The modeling method according to claim 1, characterized in that The k-nearest neighbor algorithm in S2 finds the nearest k samples for each missing value and fills the missing value with the mean or mode of these neighbors.
6. The modeling method according to claim 1, characterized in that The random seed set.seed(123) is set during random sampling in S3 to ensure the repeatability of the results.
7. The modeling method according to claim 1, characterized in that The test set size in S4 is 33%.
8. The modeling method according to claim 1, characterized in that: The evaluation indicators in S5 include: accuracy, confusion matrix, precision, recall, F1-Score and AUC.
Citation Information
Cited By
Biomarker for membranous nephropathy diagnosis, membranous nephropathy diagnosis model and construction method of membranous nephropathy diagnosis model
CN121114419A
Membranous nephropathy diagnosis biomarker, membranous nephropathy diagnosis model and construction method thereof
CN121114419B