Membranous nephropathy diagnosis biomarker, membranous nephropathy diagnosis model and construction method thereof
By combining biomarkers and using a random forest model, the challenge of non-invasive diagnosis of membranous nephropathy has been solved, especially in PLA2R antibody-negative patients, achieving high sensitivity and specificity in diagnosis and providing a non-invasive and accurate diagnostic tool.
Patent Information
- Application Number
- CN202511352062.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2045-09-22
AI Technical Summary
The diagnosis of membranous nephropathy in the current technology relies on invasive renal biopsy or non-invasive tests with insufficient sensitivity, which makes it difficult to achieve high sensitivity and high specificity in non-invasive diagnosis, especially for patients with negative PLA2R antibodies, resulting in a diagnostic blind spot.
A combination of biomarkers, including 12 peripheral blood immune cell subsets such as non-classical monocytes, CD4+ T cells, and effector CD4+ T cells, was used to construct a non-invasive diagnostic model. Characteristic biomarkers were screened through flow cytometry and machine learning algorithms to achieve efficient diagnosis.
It achieves high sensitivity and specificity in the diagnosis of membranous nephropathy, especially demonstrating strong diagnostic efficacy in PLA2R antibody-negative patients, with an AUC of 0.9374, significantly improving the diagnostic accuracy and coverage of existing technologies.
Smart Images

Figure CN121114419B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical technology and relates to a combination of peripheral blood cell biomarkers for the diagnosis of membranous nephropathy (MN), as well as a machine learning diagnostic model and interactive application system based on the combination. Background Technology
[0002] Membranous nephropathy is one of the main pathological types leading to nephrotic syndrome in adults, and its diagnosis has long relied on invasive renal biopsy as the "gold standard." However, renal biopsy is an invasive procedure with certain risks, such as bleeding and infection, and it is not suitable for or acceptable to all patients. Furthermore, it is difficult to use for frequent disease monitoring.
[0003] In recent years, the discovery of anti-phospholipase A2 receptor (PLA2R) antibodies has been a major breakthrough in the non-invasive diagnosis of membranous nephropathy. As a serological marker, according to multiple meta-analyses, anti-PLA2R antibodies have a specificity of approximately 90%, making them a powerful tool for diagnosing membranous nephropathy (MN) and leading to their widespread use in clinical practice.
[0004] However, the main drawback of this biomarker is its insufficient sensitivity. Literature reports a diagnostic sensitivity of approximately 65-78%, meaning that a significant proportion (about 22-35%) of patients with membranous nephropathy will test negative for serum PLA2R antibodies. For this antibody-negative patient group, non-invasive diagnosis remains a significant clinical challenge, and physicians often still need to return to renal biopsy for a definitive diagnosis. Furthermore, a recent study by Zand L et al. in *Kidney International* revealed that even in cases where PLA2R is negative using standard immunofluorescence (IF) methods, a portion can still be detected using the more advanced and sensitive laser microdissection / mass spectrometry (LMD / MS) technique.
[0005] In conclusion, whether due to the invasiveness and high risk of renal biopsy or the sensitivity limitations of existing non-invasive testing methods, there is an urgent clinical need to develop a non-invasive biomarker and diagnostic system that does not rely on traditional PLA2R antibody detection, and possesses high specificity and sensitivity. This system could assist in or even partially replace renal biopsy, enabling early and accurate diagnosis of all types of membranous nephropathy patients (especially those who are negative for PLA2R antibodies using conventional methods). This is of great significance for guiding clinical decision-making and improving patient prognosis. Summary of the Invention
[0006] To address the aforementioned technical problems in the prior art, this invention provides biomarkers for the diagnosis of membranous nephropathy, a diagnostic model for membranous nephropathy, and a method for constructing the same. These biomarkers, diagnostic models, and methods for constructing the same are intended to solve the technical problems of low sensitivity in the diagnosis of membranous nephropathy in the prior art, and the need for further invasive detection.
[0007] This invention provides a biomarker combination for the in vitro diagnosis of membranous nephropathy, the combination comprising the following twelve peripheral blood immune cell subsets: non-classical monocytes (CD14nCD16p cells), CD4+ T cells (CD4T), effector CD4+ T cells (EffCD4T), type 2 helper T cells (Th2), double-negative T cells (DNT), central memory CD4+ T cells (CMCD4T), plasmacytoid dendritic cells (pDC), antibody-secreting cells (ASC), total T cells (CD3T), unconverted memory B cells (USM B), classical monocytes (CD14pCD16n cells), and natural killer cells (NK).
[0008] The present invention also provides the application of the above-described combination of biomarkers in the preparation of products for the auxiliary diagnosis of membranous nephropathy.
[0009] Furthermore, the product includes flow cytometry antibody assemblies, kits, or software or online application systems or chips containing machine learning models trained on the biomarkers.
[0010] Furthermore, the machine learning model is a random forest model.
[0011] This invention also provides a method for constructing a diagnostic model for membranous nephropathy, comprising the following steps:
[0012] 1. A step for screening characteristic biomarkers,
[0013] 1.1) Peripheral venous blood samples were obtained from patients with membranous nephropathy. Flow cytometry was used to detect the immune cell subsets in the mononuclear cells of each sample, as well as the expression level of each immune cell subset in the mononuclear cells. The detected immune cell subsets were defined as candidate biomarkers.
[0014] 1.2) Construct a random forest model as the initial selection model, input the expression level of each candidate biomarker in the mononuclear cells of each peripheral venous blood sample into the initial selection model, and use the initial selection model to calculate the average Gini impurity reduction value of each candidate biomarker.
[0015] 1.3) Construct a random forest model as a temporary model, and let n=1;
[0016] 1.4) From the candidate biomarkers, the top n candidate biomarkers ranked from highest to lowest in terms of average Gini impurity decrease value were selected, and the expression levels of the selected n candidate biomarkers in mononuclear cells of each peripheral venous blood sample were incorporated into the temporary model. The AUC value of the temporary model was detected by 10-fold cross-validation.
[0017] A temporary model that incorporates the expression levels of n candidate biomarkers in mononuclear cells of each peripheral venous blood sample is defined as temporary model n.
[0018] If n>1, and the AUC value of temporary model n-1 is greater than 0.9, and the AUC value of temporary model n-1 is greater than the AUC value of temporary model n, then the candidate biomarker included in temporary model n-1 is selected as a biomarker for membranous nephropathy, and then proceed to step 1.6).
[0019] 1.5) If n is less than the total number of candidate biomarkers, let n = n + 1 and go back to step 1.4); otherwise, select the candidate biomarkers included in the nth provisional model as biomarkers for membranous nephropathy and go back to step 1.6).
[0020] 1.6) Screening complete;
[0021] 2) The steps for constructing and evaluating the diagnostic model are as follows: using the selected biomarkers, 10-fold cross-validation is used to train and evaluate the final random forest model; the optimal diagnostic threshold of 0.576 is determined by maximizing the Youden exponent.
[0022] Furthermore, one step in screening characteristic biomarkers includes:
[0023] 1) Feature importance quantification: The random forest algorithm is used to calculate the "mean Gini impurity reduction" score for each cell marker using R language;
[0024] 2) Iterative feature selection: Based on the "mean Gini impurity reduction" score of each cell marker, the importance is ranked, and features are added to the candidate model one by one from high to low. 10-fold cross-validation is used to evaluate the change in model performance after each addition of a new feature.
[0025] 3) Final model construction: After determining the optimal feature combination, use this combination to train and evaluate the final diagnostic model.
[0026] Furthermore, in step 1), the model is built by calling the randomForest() function in the randomForest package of R language. During this process, the parameter importance = TRUE is set, instructing the function to calculate and store the “mean Gini impurity decrease” score for each cell marker while generating 500 decision trees.
[0027] Furthermore, in step 2), using R language and a for loop, the features are sorted according to the "Average Gini Impurity Decrease" score in importance_df, starting with the most important features, and then incorporated into a temporary sub-model one by one. In each loop, a complete 10-fold cross-validation is performed using the train() function of the caret package in R language, and the AUC value that the current feature combination can achieve is recorded; finally, the optimal combination of 12 biomarkers is determined.
[0028] Furthermore, a combination of 12 biomarkers was ultimately identified; these 12 biomarkers are: non-classical monocytes (CD14nCD16p cells), CD4+ T cells (CD4T), effector CD4+ T cells (EffCD4T), type 2 helper T cells (Th2), double-negative T cells (DNT), central memory CD4+ T cells (CMCD4T), plasmacytoid dendritic cells (pDC), antibody-secreting cells (ASC), total T cells (CD3T), unconverted memory B cells (USM B), classical monocytes (CD14pCD16n cells), and natural killer cells (NK).
[0029] The present invention also provides a diagnostic model for membranous nephropathy, which is obtained by the above-described construction method.
[0030] This invention also provides a diagnostic aid system for membranous nephropathy, comprising:
[0031] 1) Data acquisition module, used to detect the expression levels of various membranous nephropathy markers in peripheral venous blood samples in mononuclear cells, wherein the membranous nephropathy markers are obtained by screening using the above-mentioned screening method for membranous nephropathy markers;
[0032] 2) The auxiliary diagnostic module has a built-in pre-trained membranous nephropathy diagnostic model constructed above. The membranous nephropathy diagnostic model outputs a predictive conclusion on whether the peripheral venous blood sample provider belongs to the membranous nephropathy patient based on the expression level of various membranous nephropathy markers detected by the data acquisition module in mononuclear cells.
[0033] Furthermore, the biomarkers for membranous nephropathy are a combination of the aforementioned 12 biomarkers, specifically: non-classical monocytes (CD14nCD16p cells), CD4+ T cells (CD4T), effector CD4+ T cells (EffCD4T), type 2 helper T cells (Th2), double-negative T cells (DNT), central memory CD4+ T cells (CMCD4T), plasmacytoid dendritic cells (pDC), antibody-secreting cells (ASC), total T cells (CD3T), unconverted memory B cells (USM B), classical monocytes (CD14pCD16n cells), and natural killer cells (NK).
[0034] This invention also provides a method for assisting in the diagnosis of membranous nephropathy, comprising the following steps:
[0035] (1) Obtain the expression level data of a set of cell markers in peripheral blood samples of the subjects;
[0036] (2) Input the expression level data into a pre-trained diagnostic model for membranous nephropathy;
[0037] (3) Based on the output of the membranous nephropathy diagnostic model, determine the risk or probability of the subject having membranous nephropathy.
[0038] Furthermore, in step (3), 0.576 is used as the diagnostic threshold to determine the risk or probability of a subject having membranous nephropathy.
[0039] Furthermore, step (1) includes the following sub-steps:
[0040] (1) Obtain peripheral blood samples from the subjects and dilute them with PBS;
[0041] (2) Peripheral blood mononuclear cells were separated by density gradient centrifugation;
[0042] (3) Perform erythrocyte lysis on isolated peripheral blood mononuclear cells;
[0043] (4) Resuspend the cells in flow cytometry staining buffer, add Fc receptor blocker, and then add fluorescently labeled antibodies against multiple markers for incubation;
[0044] (5) After incubation, the cells were washed to remove unbound antibodies, and then flow cytometry was used for detection and analysis.
[0045] Furthermore, the excitation and emission wavelength range of the fluorescent dye of the fluorescently labeled antibody is 300 nm to 810 nm.
[0046] Furthermore, the fluorescent dyes include: APC, percp-cy5.5, PE, FITC, PE-CY7, APC-CY7, BV421, and BV510.
[0047] This invention also provides an interactive diagnostic system, which includes a data input module, a diagnostic model invocation module constructed using the above method, and a result display module. The data input module, invocation module, and result display module are connected via communication signals. This achieves full automation from data input to diagnostic conclusion output, greatly improving the model's usability and clinical translation potential.
[0048] This invention scientifically and unbiasedly evaluates the potential efficacy of the master model in solving specific clinical challenges by constructing an independent validation cohort consisting of a target subgroup (anti-PLA2R-negative MN patients) and a control group, and performing novel cross-validation within this cohort. This method ensures the objectivity of the evaluation results.
[0049] This invention has at least the following beneficial technical effects:
[0050] Superior diagnostic performance and efficacy: The model of this invention has an AUC of up to 0.9374, demonstrating a superior level of diagnostic accuracy, which is better than traditional models based on a single serum biomarker or clinical indicator, and significantly improved compared to the published cell combination models.
[0051] Successfully Overcoming a Blind Spot in Clinical Diagnosis: It is particularly noteworthy that the model of this invention still demonstrates strong diagnostic efficacy in a subgroup of membranous nephropathy patients who are negative for existing non-invasive anti-PLA2R antibody markers. Independent 10-fold cross-validation proves that the model's AUC value remains at an extremely high level in this key subgroup. This demonstrates that the 12-cell marker combination of this invention may capture independent pathophysiological features different from the anti-PLA2R antibody pathway, successfully solving the major challenge of existing technologies failing to diagnose approximately 22-35% of serologically negative patients, and possessing extremely high clinical application value and complementarity.
[0052] Advanced machine learning methodology: This invention preferably employs a random forest model, which, compared to the traditional logistic regression model, can more effectively capture the complex nonlinear relationships and interactions between immune cells. Furthermore, this invention uses rigorous 10-fold cross-validation for model evaluation, resulting in more robust performance and a better reflection of the model's true generalization ability compared to the single 7:3 data split of previous models.
[0053] Scientific Feature Selection and Model Optimization: This invention employs a strategy of random forest importance ranking combined with iterative verification to objectively select the 12 most contributing key indicators from over 40 candidate indicators. Simultaneously, by maximizing the Youden index, the optimal diagnostic threshold of 0.576 was determined, achieving an ideal balance between sensitivity (89.0%) and specificity (86.0%).
[0054] Innovative Applications and Rigorous External Validation: This invention not only constructs a model but also develops it into an interactive web application system. Crucially, the model underwent independent external validation, achieving an AUC of 0.9125 in a cohort of 40 new patients and 9 new healthy individuals. This demonstrates the model's robustness, generalization ability, and reliability in real-world scenarios.
[0055] Compared with existing technologies, the technical effects of this invention are positive and significant. This invention utilizes the Random Forest Recursive Feature Elimination (RF-RFE) algorithm to screen the optimal combination of biomarkers from high-dimensional flow cytometry data and constructs a random forest diagnostic model. This model, after 10-fold cross-validation, demonstrates superior performance in distinguishing membranous nephropathy patients from healthy controls, achieving an area under the curve (AUC) of 0.9374. It also exhibits high accuracy in independent validations, including in the challenging subgroup of PLA2R antibody negativity. This invention provides a novel, high-precision, non-invasive, machine learning-based auxiliary diagnostic strategy for membranous nephropathy, particularly for patients with serum PLA2R antibody negativity. Attached Figure Description
[0056] Figure 1 illustrates a representative multi-step sequential gating strategy for identifying various key immune cell subsets in peripheral blood. The analytical workflow begins with the initial separation of lymphocytes based on their physical characteristics (forward and lateral scattering, FSC / SSC), followed by rigorous screening of individual cells using pulse signal widths (e.g., FSC-A vs FSC-H). Subsequently, downstream cell subsets are identified step-by-step using a series of two-dimensional scatter plots based on different combinations of fluorescently labeled antibodies. For example, CD4+ and CD8+ subsets are distinguished from CD3+ T cells, and further functional subsets such as Th1 / Th2 and memory T cells are defined based on markers such as CXCR3, CD196, CCR7, and CD45RO; or CD19+ B cells, various monocytes, and dendritic cells are directly delineated from lymphocytes. This figure clearly illustrates the phenotypic definitions and analytical pathways for all target cell populations in this invention, providing a rigorous methodological basis for subsequent quantification and modeling.
[0057] Figure 2This section presents a recursive feature evaluation graph based on a random forest model. The horizontal axis represents cell subpopulation features ranked by importance, the left vertical axis represents feature importance scores, and the right vertical axis represents the model's average AUC value. The line graph clearly depicts the trend of model diagnostic performance (AUC) as the number of features increases, showing that performance peaks at the first 12 features.
[0058] Figure 3 The results show the receiver operating characteristic (ROC) curves of the final 12-biomarker model across the entire cohort, with an AUC value of 0.9374 clearly indicated.
[0059] Figure 4 The ROC curve of the model of this invention in the anti-PLA2R antibody-negative membranous nephropathy subgroup is presented. The curve was generated using an independent 10-fold cross-validation method, and its excellent AUC value demonstrates the model's strong ability to address blind spots in clinical diagnosis.
[0060] Figure 5: Screenshot (A) of the interactive web application system demonstrating the final result of this invention. The left side shows the numerical input area for 12 biomarkers, and the right side shows the diagnostic results display area, which can display the predicted disease probability and the final diagnostic conclusion in real time. It also demonstrates successful diagnostic verification on two independent external samples (predicting a 91.2% disease probability for a patient diagnosed with MN via renal biopsy (C), and a 18.8% disease probability for a healthy person (B)), proving the accuracy and practicality of the system.
[0061] Figure 6 The ROC curve of the model is shown in an independent external validation cohort consisting of 40 new MN patients and 9 new healthy controls. Its excellent AUC value (0.9125) demonstrates the model's generalization ability. Detailed Implementation
[0062] Unless otherwise specified, all raw materials or instruments used in the following embodiments of the present invention are commercially available.
[0063] Example 1
[0064] I. Research Subjects
[0065] The membranous nephropathy patients described in this invention were 100 patients diagnosed by renal biopsy, and the control group consisted of 57 healthy controls. All participants signed informed consent forms, and the study was approved by the ethics review committee, complying with national ethics review regulations.
[0066] II. Sample Collection and Processing
[0067] 1. Sample collection: Collect 2-3 mL of peripheral venous blood from the subject and use EDTA anticoagulation tubes for anticoagulation.
[0068] 2. PBMC separation: Dilute EDTA-anticoagulated whole blood with an equal volume of PBS and slowly add it onto an equal volume of Ficoll-PaquePLUS separation solution. Centrifuge horizontally (1800-2000 rpm, 20-25 minutes, room temperature, turn off the centrifuge and brake).
[0069] 3. PBMC collection and washing: Carefully aspirate the mononuclear cell (PBMC) layer into a new centrifuge tube, add PBS to resuspend and centrifuge to wash the cells (1500 rpm, 5-10 minutes).
[0070] 4. Red blood cell lysis: Discard the supernatant, add an appropriate amount of red blood cell lysis buffer (1X RBC Lysis Buffer) to lyse the red blood cells, incubate for several minutes, then add PBS to stop the lysis, and centrifuge to wash the cells.
[0071] 5. Cell resuspending: Discard the supernatant, resuspend the cells in flow cytometry staining buffer (e.g., PBS containing BSA and NaN3), and adjust the cell concentration to approximately 1-5 × 10^6 cells / mL.
[0072] III. Flow Cytometry Detection
[0073] 1. Fc receptor blockade: Take the cell suspension with the adjusted concentration, add an appropriate amount of Fc receptor blocker (such as HumanTruStain FcX™) to each tube, and incubate at room temperature for 5-10 minutes.
[0074] 2. Antibody Staining and Gating Strategy: Based on the experimental design, we configured multiple multicolor antibody panels to comprehensively identify and characterize various target immune cell subsets in peripheral blood. Fluorescently labeled antibody combinations were added to the cell suspension, with all antibodies added at the recommended concentration. After gentle mixing, the mixture was incubated at 4°C in the dark for 25-30 minutes. We analyzed the target cells using a rigorous multi-step sequential gating strategy, the complete and representative analytical process of which is detailed in Figure 1. To focus the core findings with the greatest diagnostic value, only 12 key cell subsets screened using the machine learning method described in this invention are listed below, along with the core antibody combinations necessary for their identification. Detailed typing strategies for other cell types (such as Treg, Th1, memory T cell subtypes, B cell subtypes, etc.) can be found in Figure 1.
[0075] Figure 1A Antibody-secreting cells (ASC): anti-CD19, anti-CD27, anti-CD38.
[0076] Figure 1BUnconverted memory B cells (USM B): Anti-CD19, anti-CD27, anti-IgD.
[0077] Figure 1C Central memory CD4+ T cells (CM CD4T): anti-CD3, anti-CD4, anti-CCR7, anti-CD45RO.
[0078] Figure 1D CD4+ T cells (CD4T): anti-CD3, anti-CD4; effector CD4+ T cells (EffCD4T): anti-CD3, anti-CD4, anti-CD25, anti-CD127; type 2 helper T cells (Th2): anti-CD3, anti-CD4, anti-CD25, anti-CD127, anti-CXCR3, anti-CD196.
[0079] Figure 1E Double-negative T cells (DNT): anti-CD3, anti-CD4, anti-CD8a; Total T cells (CD3T): anti-CD3; Natural killer cells (NK): anti-CD3, anti-CD56.
[0080] Figure 1F Non-classical monocytes (CD14nCD16p cells): anti-HLA-DR, anti-CD14, anti-CD16; Classical monocytes (CD14pCD16n cells): anti-HLA-DR, anti-CD14, anti-CD16; Plasma-like dendritic cells (pDCs): anti-CD14, anti-CD16, anti-CD11c, anti-CD123.
[0081] 3. Washing and resuspending: After incubation, wash the cells with flow cytometry staining buffer, centrifuge and discard the supernatant. Resuspend the cells with an appropriate amount of flow cytometry staining buffer.
[0082] 4. Analytical Process: Transfer the cell suspension to flow cytometry tubes and perform analysis using a flow cytometer (BD FACSCanto II). Set appropriate voltage and compensation, and collect a sufficient number of cell events for analysis.
[0083] IV. Data Analysis and Results
[0084] 1. Screening of key biomarkers (corresponding to...) Figure 2 )
[0085] 1.1 Method Description
[0086] To identify the most valuable combination of core biomarkers for the diagnosis of membranous nephropathy (MN) from high-dimensional flow cytometry data and to build a high-performance predictive model, we employed a rigorous multi-step machine learning approach.
[0087] The core idea of this method is as follows: First, it utilizes the inherent mechanism of the Random Forest model to quantify and rank the importance of all candidate biomarkers (features); then, through systematic iterative modeling and cross-validation, it finds the optimal combination of features that maximizes model prediction performance (measured by AUC) while minimizing the number of features. The entire process mainly includes three key stages:
[0088] ① Feature importance quantification: Using the random forest algorithm, the "Mean Decrease Gini" score of each cell marker is calculated.
[0089] ② Iterative feature selection: Based on the importance ranking obtained in the first step, features are added to the candidate model one by one from high to low, and 10-fold cross-validation is used to evaluate the change in model performance after each addition of a new feature.
[0090] ③ Final model construction: After determining the optimal feature combination, use this combination to train and evaluate the final diagnostic model on the entire dataset.
[0091] 1.2 Detailed Explanation of Mean Decrease Gini
[0092] 1.2.1 Definition and function.
[0093] Mean Decrease Gini is a feature importance metric derived from the Random Forest algorithm. Its core function is to quantify the average contribution of each feature (e.g., a cellular marker) to improving the "purity" or "determinism" of the model's decisions. In our diagnostic task, a higher Mean Decrease Gini score for a feature indicates that the feature plays a more crucial and decisive role in distinguishing patients with membranous nephropathy (MN) from healthy controls (HC). Therefore, we can use this score to rank all candidate markers, thereby identifying the features with the greatest diagnostic potential.
[0094] 1.2.2 Gini Impurity and Calculation Formula
[0095] "Gini Impurity" is a metric used in decision trees to assess the "mixed" nature of a data node. In a classification task, if all samples within a node belong to the same category (e.g., all are MN patients), then the node is "completely pure," and its Gini Impurity is 0. Conversely, the more uniformly the samples from each category are mixed, the higher the node's Gini Impurity. For a data node containing C categories, its Gini Impurity G is calculated using the following formula:
[0096]
[0097] Where pi is the proportion of samples belonging to category i in the data node.
[0098] In our binary classification problem (MN vs. HC), the formula simplifies to:
[0099]
[0100] : Represents the proportion of patients with membranous nephropathy (MN) in a node. : Represents the proportion of the Healthy Control (HC) sample in a node.
[0101] When a decision tree uses a feature (such as pDC) to split at a node, it produces two child nodes. The "Gini impurity reduction" resulting from this split is equal to the Gini impurity of the parent node before the split, minus the weighted average of the Gini impurities of the two child nodes. The larger this reduction value, the stronger the "purification" effect of the feature in this split.
[0102] 1.2.3 The meaning of "average"
[0103] Random forests consist of hundreds or thousands of decision trees. A feature may be used for splitting on multiple trees and at multiple nodes in the forest. The "average Gini impurity reduction" score of that feature is the average of the "Gini impurity reduction values" produced by that feature on all trees in the forest and at all the split points it participates in. This final average score comprehensively reflects the overall importance of that feature in the entire model system.
[0104] 1.3. Summary of R code implementation
[0105] Our implementation in R strictly follows the above methodology: Calculating Importance Scores: We first build the model by calling the `randomForest()` function from the `randomForest` package. During this process, we specifically set the parameter `importance = TRUE`, instructing the function to calculate and store the Mean Decrease Gini score for each feature while generating 500 decision trees. Iterative Selection and Evaluation: Next, we use a `for` loop, strictly following the order of Mean Decrease Gini scores in `importance_df`, starting with the most important features and incorporating them one by one into a temporary sub-model. In each iteration, we perform a complete 10-fold cross-validation using the `train()` function from the `caret` package and record the AUC value achievable by the current feature combination.
[0106] Through this series of rigorous calculations and verification steps, we were finally able to determine the optimal combination of 12 biomarkers, laying a solid foundation for subsequent efficient and accurate diagnostic models.
[0107] Specifically, in a step of screening characteristic biomarkers,
[0108] 1.1) Peripheral venous blood samples were obtained from patients with membranous nephropathy. Flow cytometry was used to detect the immune cell subsets in the mononuclear cells of each sample, as well as the expression level of each immune cell subset in the mononuclear cells. The detected immune cell subsets were defined as candidate biomarkers.
[0109] 1.2) Construct a random forest model as the initial selection model, input the expression level of each candidate biomarker in the mononuclear cells of each peripheral venous blood sample into the initial selection model, and use the initial selection model to calculate the average Gini impurity reduction value of each candidate biomarker.
[0110] 1.3) Construct a random forest model as a temporary model, and let n=1;
[0111] 1.4) From the candidate biomarkers, the top n candidate biomarkers ranked from highest to lowest in terms of average Gini impurity decrease value were selected, and the expression levels of the selected n candidate biomarkers in mononuclear cells of each peripheral venous blood sample were incorporated into the temporary model. The AUC value of the temporary model was detected by 10-fold cross-validation.
[0112] A temporary model that incorporates the expression levels of n candidate biomarkers in mononuclear cells of each peripheral venous blood sample is defined as temporary model n.
[0113] If n>1, and the AUC value of temporary model n-1 is greater than 0.9, and the AUC value of temporary model n-1 is greater than the AUC value of temporary model n, then the candidate biomarker included in temporary model n-1 is selected as a biomarker for membranous nephropathy, and then proceed to step 1.6).
[0114] 1.5) If n is less than the total number of candidate biomarkers, let n = n + 1 and go back to step 1.4); otherwise, select the candidate biomarkers included in the nth provisional model as biomarkers for membranous nephropathy and go back to step 1.6).
[0115] 1.6) Screening complete;
[0116] 2) The steps for constructing and evaluating the diagnostic model are as follows: using the selected biomarkers, 10-fold cross-validation is used to train and evaluate the final random forest model; the optimal diagnostic threshold of 0.576 is determined by maximizing the Youden exponent.
[0117] Specifically, a combination of 12 biomarkers was finally selected. These 12 biomarkers are: non-classical monocytes, CD4+ T cells, effector CD4+ T cells, type 2 helper T cells, double-negative T cells, central memory CD4+ T cells, plasmacytoid dendritic cells, antibody-secreting cells, total T cells, unconverted memory B cells, classical monocytes, and natural killer cells.
[0118] 2. Construction and performance evaluation of the master diagnostic model (corresponding to...) Figure 3 We used 12 selected biomarkers and performed 10-fold cross-validation on the entire dataset (100 MN patients and 57 healthy controls) to train and evaluate the final random forest model. The optimal diagnostic threshold was determined to be 0.576 by maximizing the Youden exponent. Under this condition, the model demonstrated excellent and balanced performance: AUC=0.9374, sensitivity 89.0%, specificity 86.0%, positive predictive value (PPV) 91.8%, negative predictive value (NPV) 81.7%, and F1 score 90.4%. 3.
[0120] 4. In-depth validation in the anti-PLA2R negative subgroup (corresponding to) Figure 4 To verify the efficacy of this invention in addressing clinical diagnostic challenges, we conducted a rigorous and in-depth validation study targeting the critical subgroup of anti-PLA2R antibody-negative individuals.
[0121] We constructed a dedicated, independent validation cohort from the overall sample. This cohort consisted of two groups: a patient group (all 19 patients with anti-PLA2R antibody-negative membranous nephropathy) and a control group (all 57 healthy controls). This cohort, specifically for in-depth validation, contained 76 samples. Subsequently, we performed a novel, independent 10-fold cross-validation within this independent cohort of 76 individuals. The core objective of this was to evaluate the intrinsic diagnostic capability of our selected combination of 12 biomarkers in distinguishing between "anti-PLA2R negative patients" and "healthy individuals" in this specific and challenging clinical scenario. The results showed that the model built based on these 12 biomarkers still exhibited extremely high diagnostic performance in this key subgroup, with an AUC as high as 0.9012. This result fully demonstrates the strong clinical application potential of the biomarker combination proposed in this invention, especially in addressing diagnostic blind spots that existing serological methods cannot cover.
[0122] 4. Development and Validation of the Diagnostic System (corresponding to Figure 5): To facilitate the application of the model, we used the Shiny framework in R to encapsulate the trained random forest master model into an interactive web application system. This system allows users to input the detection values of 12 key cellular markers. The system backend then calls the pre-stored model for calculation and returns in real time the predicted probability of the patient having membranous nephropathy and a diagnostic suggestion (high risk / low risk) based on a 0.576 threshold. Successful diagnostic validation on independent external samples demonstrates the accuracy and practicality of the system.
[0123] 5. External validation of the model (corresponding to...) Figure 6 To rigorously evaluate the model's generalization ability and clinical applicability, we validated the model using an independent external cohort consisting of 40 new patients with membranous nephropathy (MN) and 9 new healthy controls (HC).
[0124] The model demonstrated excellent performance in distinguishing between MN patients and healthy controls in an external validation cohort. Key performance indicators showed that the area under the curve (AUC) was as high as 0.9125 (95% confidence interval: 0.8266–0.9984). This excellent external validation result fully demonstrates that the proposed combination of 12 biomarkers and diagnostic model possesses high robustness and excellent generalization ability, providing strong evidence for its clinical application.
[0125] In summary, the present invention provides a complete technical solution from biomarker discovery to clinical application verification, offering a powerful tool for the non-invasive diagnosis of membranous nephropathy, and achieving significant technical breakthroughs, especially in solving the diagnostic challenges of anti-PLA2R antibody-negative patients.
Claims
1. A biomarker combination for the in vitro diagnosis of membranous nephropathy, characterized in that, The combination comprises the following twelve peripheral blood immune cell subpopulations: non-classical monocytes, CD4+ T cells, effector CD4+ T cells, type 2 helper T cells, double negative T cells, central memory CD4+ T cells, plasmacytoid dendritic cells, antibody secreting cells, total T cells, unswitched memory B cells, classical monocytes, and natural killer cells.
2. Use of the biomarker combination of claim 1 in the preparation of a product for aiding the diagnosis of membranous nephropathy.
3. Use according to claim 2, characterized in that, The product comprises a combination of flow cytometry antibodies, a kit, or a chip based on a machine learning model trained on the biomarkers.
4. Use according to claim 3, characterized in that, The machine learning model is a random forest model.
5. A method for constructing a diagnostic model for membranous nephropathy, characterized by, The method comprises the following steps: 1) a step of screening characteristic biomarkers, which specifically comprises the following steps: 1.1) obtaining peripheral venous blood samples from patients with membranous nephropathy, detecting immune cell subpopulations in mononuclear cells of each sample using a flow cytometer, and determining the expression levels of each immune cell subpopulation in mononuclear cells, and defining each detected immune cell subpopulation as a candidate biomarker; 1.2) constructing a random forest model as a preliminary model, inputting the expression levels of each candidate marker in mononuclear cells of each peripheral venous blood sample into the preliminary model, and calculating the average Gini impurity reduction value of each candidate marker using the preliminary model; 1.3) constructing a random forest model as a temporary model, and setting n = 1; 1.4) selecting the top n candidate markers from each candidate marker in terms of descending average Gini impurity reduction value, and inputting the expression levels of the selected n candidate markers in mononuclear cells of each peripheral venous blood sample into the temporary model, and detecting the AUC value of the temporary model using a 10-fold cross-validation method; the temporary model comprising the expression levels of the n candidate markers in mononuclear cells of each peripheral venous blood sample is defined as the n th temporary model; if n > 1, and the AUC value of the n-1 th temporary model is greater than 0.9, and the AUC value of the n-1 th temporary model is greater than the AUC value of the n th temporary model, then the candidate markers included in the n-1 th temporary model are selected as the biomarkers for membranous nephropathy, and the method proceeds to step 1.6); 1.5) if n is less than the total number of candidate markers, then n = n + 1, and the method proceeds to step 1.4), otherwise the candidate markers included in the n th temporary model are selected as the biomarkers for membranous nephropathy, and the method proceeds to step 1.6); 1.6) the screening is completed, and the obtained biomarker combination comprises twelve peripheral blood immune cell subpopulations: non-classical monocytes, CD4+ T cells, effector CD4+ T cells, type 2 helper T cells, double negative T cells, central memory CD4+ T cells, plasmacytoid dendritic cells, antibody secreting cells, total T cells, unswitched memory B cells, classical monocytes, and natural killer cells; 2) The step of building and evaluating the diagnostic model, using the selected biomarkers, to train and evaluate the final random forest model by 10-fold cross-validation; the optimal diagnostic threshold is determined by maximizing the Youden index, which is 0.
576.
6. The method according to claim 5, wherein the method is characterized by, A step of screening feature biomarkers includes: 1) Feature importance quantification: using the random forest algorithm, the "average Gini impurity reduction" score of each cell marker is calculated using R language; 2) Iterative feature screening: according to the "average Gini impurity reduction" score of each cell marker, the importance is ranked, and the features are added to the candidate model one by one from high to low, and the change in model performance after adding new features each time is evaluated using 10-fold cross-validation; 3) Final model building: after determining the optimal feature combination, the final diagnostic model is trained and evaluated using the combination.
7. The method according to claim 6, wherein the method is characterized by, In step 1), the randomForest() function in the randomForest package in R language is called to build the model, and in this process, the parameter importance = TRUE is set, instructing the function to calculate and store the "average Gini impurity reduction" score of each cell marker while generating 500 decision trees.
8. The method according to claim 6, wherein, In step 2), using R language, a for loop is used to sort the "average Gini impurity reduction" score in importance_df, starting from the most important feature, and adding it to a temporary sub-model one by one. In each loop, a complete 10-fold cross-validation is performed using the train() function in the caret package in R language, and the AUC value that can be achieved by the current feature combination is recorded. Finally, the optimal combination of 12 biomarkers is determined.
9. An interactive diagnostic system characterized by, The system comprises a data input module, a membrane nephropathy diagnostic model calling module constructed using the method of claim 5, and a result display module, wherein the data input module, the calling module, and the result display module are connected through communication signals.
Citation Information
Patent Citations
Method for constructing differential diagnosis model of lupus nephritis and membranous nephropathy
CN120613100A
Method for diagnosing microcirculation inflammation or antibody-mediated rejection in kidney transplant recipients using the frequency of HLA-DR+ T cells in peripheral blood
KR1020170089334A