Membrane nephropathy intelligent staging and diagnosis and treatment method based on knowledge graph reasoning

By constructing a causal graph model specific to drug intervention for membranous nephropathy, integrating multi-source data and eliminating spurious associations, and combining protein interaction strength quantification algorithms and hierarchical analysis, the problem of insufficient consistency and reliability in diagnosis and treatment decisions in existing technologies has been solved, achieving precision and adaptability of personalized drug intervention.

CN121862385APending Publication Date: 2026-04-14GUANGZHOU INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-30
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies lack systematic tools to integrate multi-dimensional information in the diagnosis and treatment of membranous nephropathy, making it difficult to accurately predict disease progression and drug efficacy. In particular, the consistency and reliability of diagnosis and treatment decisions are insufficient in complex comorbidity scenarios, and there is a lack of clear biological basis and dynamic iterative optimization capabilities, which cannot meet the needs of precision medicine.

Method used

We constructed a causal graph model specifically for drug intervention in membranous nephropathy, integrated multi-source data through knowledge graph reasoning, eliminated false association edges, quantified the net correlation strength between drug intervention and efficacy, coupled the target protein network with the causal reasoning link, and combined the pathological characteristics differences of different stages of membranous nephropathy to construct a stage-appropriate drug screening model. We then iteratively optimized the model parameters using real-world cohort data.

Benefits of technology

It establishes a clear and genuine causal relationship between drug intervention and efficacy, provides more precise support for diagnostic and treatment decision-making, meets the requirements for personalized drug prioritization and adaptability to actual clinical changes, and ensures the long-term adaptability of the model and the practicality of the treatment plan.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121862385A_ABST
    Figure CN121862385A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of knowledge graph and medical fusion, in particular to a membranous nephropathy intelligent staging and diagnosis and treatment method based on knowledge graph reasoning, which comprises the following steps of: constructing a membranous nephropathy exclusive causal graph model, integrating multi-source data and defining a core node; hierarchical analysis, tendency score matching and a backdoor adjustment formula are combined to process hybrid variables, interference is effectively stripped, real causal association of drugs and curative effects is defined, the limitation that in the prior art, dependency correlation analysis is prone to generating false association is broken through, and accurate logic support is provided for diagnosis and treatment decisions. On the basis of an STRING database adaptation algorithm, a target protein network and a causal link are coupled, a drug onset molecular mechanism is mined, dual verification of virtual and in-vitro cell experiments is combined, a biological basis is enhanced, and the problems that a traditional method is fuzzy in mechanism and single in verification are solved. And an adaptive drug screening model is constructed in combination with disease staging, so that personalized sorting is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of knowledge graph and medical integration technology, and more specifically, to a method for intelligent staging and diagnosis of membranous nephropathy based on knowledge graph reasoning. Background Technology

[0002] Membranous nephropathy is a common glomerular disease in clinical practice. Its diagnosis and treatment require the integration of multi-dimensional information, including clinical indicators, pathological features, drug responses, and individual patient conditions. Furthermore, the pathological mechanisms and treatment needs differ significantly across different stages of the disease, making it highly dependent on personalized and precise treatment. Current clinical diagnosis and treatment largely rely on physician experience and lack systematic tools to integrate multi-source information, making it difficult to accurately predict disease progression and drug efficacy. This is especially true in complex comorbidity scenarios, where the consistency and reliability of treatment decisions are insufficient.

[0003] Existing AI-based diagnostic and treatment technologies for membranous nephropathy generally suffer from a lack of causal analysis and insufficient handling of confounding variables. Most methods rely solely on single-modality data for correlation analysis, failing to effectively distinguish the correlation and causality between drug intervention and efficacy. They are also susceptible to confounding factors such as concomitant treatment and patient baseline differences, leading to the risk of spurious associations in the selected intervention programs. Consequently, these interventions are difficult to directly implement in clinical practice and cannot provide reliable logical support for diagnostic and treatment decisions.

[0004] Meanwhile, existing technologies often fail to deeply integrate research findings at the molecular mechanism level, making it difficult to establish a clear link between drug intervention and pathological changes. This results in treatment plans lacking sufficient biological basis, and validation systems are often limited to a single dimension, exhibiting limited stability and generalization ability. Furthermore, most methods lack adaptation mechanisms designed for the staging characteristics of membranous nephropathy and lack dynamic iterative optimization capabilities, failing to adapt to individual differences and data update needs in actual clinical diagnosis and treatment, thus failing to meet the requirements of precision medicine development. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to provide an intelligent staging and treatment method for membranous nephropathy based on knowledge graph reasoning.

[0006] To achieve the above objectives, the present invention provides the following technical solution: A knowledge graph-based intelligent staging and diagnosis method for membranous nephropathy, comprising the following steps: Step 1: Construct a causal graph model specifically for drug intervention in membranous nephropathy. The core nodes are determined based on all elements of clinical diagnosis and treatment. Through literature mining and training with real-world cohort data, the causal relationship paths between each core node are defined and false association edges are removed to form a structured causal relationship framework. Step 2: Based on the structured causal framework, quantify and eliminate confounding variables, clarify the net association strength between drug intervention measures and efficacy evaluation indicators, and distinguish between their correlation and causality; Step 3: Using a protein interaction strength quantification algorithm adapted to the STRING database, couple key nodes of the target protein network with the causal inference link to determine the direction and magnitude of regulation, locate key nodes that show significant changes after the implementation of the drug intervention, and construct the molecular mechanism causal link of the drug intervention, key nodes of the target protein network, and pathological change indicators of membranous nephropathy. Step 4: Verify the authenticity and stability of the causal link of the molecular mechanism through a multi-dimensional verification scheme to strengthen the biological basis of the molecular mechanism of drug efficacy; Step 5: Combining the differences in pathological characteristics of different stages of membranous nephropathy, construct a stage-appropriate drug screening model based on the aforementioned molecular mechanism causal link to achieve priority ranking of personalized drug intervention measures; Step 6: Apply the personalized drug intervention plan to a real-world cohort, and iteratively optimize the model parameters by tracking data to form a complete diagnosis and treatment closed loop.

[0007] Furthermore, the core nodes include drug intervention measures, patient individual baseline characteristics, accompanying treatment plans, target protein network nodes, pathological change indicators, and efficacy evaluation indicators. The dedicated causal graph model for membranous nephropathy drug intervention achieves precise mapping between the core nodes and knowledge graph entities through its built-in structured causal relationship framework and a standardized JSON data interface.

[0008] Furthermore, the confounding variables mentioned in step 2 include the type of concomitant treatment, patient age, disease course, and underlying diseases. A combination of stratified analysis and propensity score matching is used to remove the interference of the confounding variables on the association between the drug intervention and the efficacy evaluation indicators.

[0009] Furthermore, the direction and magnitude of regulation are determined using a protein-protein interaction strength quantification algorithm adapted to the STRING database. The specific technical solution is as follows: First, target proteins related to membranous nephropathy are extracted from the STRING database. Low-confidence indirect interaction data are removed, and direct physical interaction relationships are retained to form an initial protein-protein interaction dataset. Then, an analytic hierarchy process (AHP) is used for each candidate protein... Assign weights Determine the interaction probability Based on the core formula The interaction strength between the target protein and candidate proteins is quantified one by one, with the summation range covering all pre-processed candidate proteins. After calculation Normalization is performed, and intensity threshold screening and control amplitude judgment are based on S.

[0010] Furthermore, by combining differences in expression levels and the strength of interactions, the regulatory direction of key nodes on target proteins can be determined: if the expression level of a candidate protein is upregulated and S ≥ the threshold, it is determined to be activation-type regulation; if the expression level of a candidate protein is downregulated and S ≥ the threshold, it is determined to be inhibition-type regulation, providing a clear regulatory basis for constructing a causal link between drug-protein-pathological molecular mechanisms.

[0011] Furthermore, the four types of confounding variables were first categorized and standardized preprocessed to fit the subsequent analysis model: concomitant treatment type, patient age, disease duration, and underlying disease. The preprocessed data should conform to a normal distribution or be free from severe skewness. Using concomitant treatment type as the core stratification variable, sub-strata were first divided according to concomitant treatment type. Within each sub-stratum, patients were further grouped according to age and disease duration, with underlying disease serving as an in-stratum correction factor. A two-level stratification structure of main stratum and subgroup was constructed. By stratifying, patients with different concomitant treatment backgrounds, age ranges, and disease durations were isolated, initially eliminating systematic differences in confounding variables between groups. A backdoor adjustment formula was then used for quantitative stripping.

[0012] Furthermore, the types of accompanying treatments are categorized according to ATC drug classification codes, excluding symptomatic treatments without clear pharmacological effects; Patient age: According to the guidelines for the diagnosis and treatment of membranous nephropathy, patients were grouped into three groups: ≤40 years old, 41-60 years old, and >60 years old, with each group accounting for ≥20% of the sample size. Disease duration: Based on the time from onset to intervention, patients were divided into three groups: ≤1 year, 1-3 years, and >3 years, to avoid non-linear interference from the duration of the disease on the efficacy assessment. Underlying diseases: Focusing on three common complications: hypertension, diabetes and hyperlipidemia, coded in binary categories of presence or absence, and those with overlapping complications are marked separately.

[0013] Furthermore, based on the stratification, PSM was implemented for the remaining confounding variables within each stratum to further remove micro-level interference, specifically as follows: ① Matching variable setting: All four types of confounding variables were included in the propensity score calculation model, with the dependent variable being whether or not the target drug intervention was received, and the independent variables being the four types of pre-processed confounding variables; ② Score calculation: The propensity score for each patient was calculated using a logistic regression model, with a model fit R² ≥ 0.6, ensuring the explanatory power of confounding variables on intervention allocation; ③ Matching strategy: A 1:1 nearest neighbor matching method was used, with a caliper range set to 0.02. After matching, a balance test was performed on the confounding variables between the intervention group and the control group within each stratum to ensure that the SMD < 0.05 among all variable groups, indicating no statistically significant difference; ④ Extreme value handling: Samples with propensity scores in the top 5% and bottom 5% were removed to avoid matching bias caused by extreme values.

[0014] Furthermore, the backdoor adjustment formula is used for quantitative stripping, specifically as follows: ① Weight assignment: The influence weight of each type of confounding variable is calculated using the backdoor adjustment formula; ② Net association strength calculation: Based on the adjusted weights, the association strength between drug intervention and efficacy is corrected, using the following formula: ,in To correct the net correlation strength, For uncorrected correlation strength, For the weights of the j-th type of confounding variables, ③ Error control: By iteratively optimizing the weight parameters, ensure that the error of removing confounding variables is ≤5%, and finally obtain the true causal association strength between drug intervention and efficacy evaluation indicators, providing a precise data foundation for subsequent molecular mechanism linkage coupling.

[0015] Furthermore, the model parameters are optimized by iteratively backtracking data. These model parameters include node association weights of the structured causal framework, backdoor adjustment formula parameters, and protein-protein interaction strength thresholds.

[0016] Compared with the prior art, the present invention has the following beneficial effects: The method of this invention constructs a causal graph model specific to membranous nephropathy, integrates multi-source data and accurately defines core nodes, and combines hierarchical analysis, propensity score matching and backdoor adjustment formula to process confounding variables, effectively removes interference from irrelevant factors, and clarifies the true causal relationship between drug intervention and efficacy. This breaks through the limitation of existing technologies that rely solely on correlation analysis and are prone to producing false associations, and provides more accurate logical support for diagnosis and treatment decisions. Based on a protein-protein interaction strength quantification algorithm adapted to the STRING database, this invention successfully couples the target protein network with the causal inference link to uncover the molecular mechanism of drug efficacy. Simultaneously, through multi-dimensional verification via virtual intervention experiments and in vitro cell experiments, it strengthens the biological basis of the treatment plan, addressing the problems of traditional methods lacking clear mechanistic support and having a single verification system. Combining the pathological characteristics differences of different stages of membranous nephropathy, this invention constructs a stage-adapted drug screening model to achieve personalized drug priority ranking. Furthermore, through dynamic iterative optimization using real-world cohort data, the model parameters are continuously corrected to ensure that the plan can adapt to actual clinical changes and individual differences. This not only meets the personalized needs of precision medicine but also ensures the long-term adaptability of the model, providing a more targeted and practical solution for clinical diagnosis and treatment. Attached Figure Description

[0017] Figure 1 A flowchart of an intelligent staging and treatment method for membranous nephropathy based on knowledge graph reasoning; Figure 2 A flowchart for removing the interference of confounding variables on the association between drug intervention and efficacy evaluation indicators. Detailed Implementation

[0018] Reference Figures 1 to 2 A method for intelligent staging and diagnosis of membranous nephropathy based on knowledge graph reasoning, characterized by the following steps: Step 1: Construct a causal graph model specifically for drug intervention in membranous nephropathy. Core nodes are determined based on all elements of clinical diagnosis and treatment. Through literature mining and training with real-world cohort data, causal pathways between core nodes are defined, and spurious edges are eliminated to form a structured causal framework. Core nodes include drug intervention measures, patient baseline characteristics, concomitant treatment plans, target protein network nodes, pathological change indicators, and efficacy evaluation indicators. The causal graph model specifically for drug intervention in membranous nephropathy achieves precise mapping between the core nodes and knowledge graph entities through its built-in structured causal framework and a standardized JSON data interface.

[0019] Step 2: Based on the structured causal framework, quantify and eliminate confounding variables, clarify the net association strength between drug intervention measures and efficacy evaluation indicators, and distinguish between their correlation and causality.

[0020] Step 3: Couple key nodes of the target protein network with the causal reasoning link, locate key nodes that show significant changes after the implementation of the drug intervention, and construct the molecular mechanism causal link of the drug intervention - key nodes of the target protein network - pathological change indicators of membranous nephropathy.

[0021] Step 4: Verify the authenticity and stability of the causal link of the molecular mechanism through a multi-dimensional verification scheme to strengthen the biological basis of the molecular mechanism of drug efficacy.

[0022] Step 5: Combining the differences in pathological characteristics of different stages of membranous nephropathy, construct a stage-appropriate drug screening model based on the aforementioned molecular mechanism causal link to achieve priority ranking of personalized drug intervention measures.

[0023] Step 6: Apply the personalized drug intervention plan to a real-world cohort, and iteratively optimize the model parameters by tracking data to form a complete diagnosis and treatment closed loop.

[0024] The confounding variables mentioned in step 2 include the type of concomitant treatment, patient age, disease duration, and underlying diseases. A combination of stratified analysis and propensity score matching is used to isolate the confounding variables from interfering with the association between the drug intervention and the efficacy evaluation indicators. The specific technical solution is as follows: (1) Classification and preprocessing of confounding variables: First, the four types of confounding variables were classified and standardized for preprocessing to adapt to the subsequent analysis model: ① Type of accompanying treatment (categorical variable): Classified according to ATC drug classification codes (such as ARBs, anticoagulants, etc.), and symptomatic treatments without clear pharmacological effects (such as fluid replacement) were removed; ② Patient age (continuous variable): Grouped into three groups according to the guidelines for the diagnosis and treatment of membranous nephropathy: ≤40 years, 41-60 years, and >60 years, with each group accounting for ≥20% of the sample size; ③ Disease duration (continuous variable): Divided into three groups according to the time from onset to intervention: ≤1 year, 1-3 years, and >3 years, to avoid nonlinear interference of disease duration on efficacy assessment; ④ Underlying diseases (multi-category variable): Focusing on the three common complications of hypertension, diabetes, and hyperlipidemia, coded by "presence" and "absence", and those with overlapping complications were marked separately (the proportion was controlled within 30% of the total sample). The preprocessed data must conform to a normal distribution (continuous variable) or have no severe skewness (categorical variable). Outliers (such as age > 80 years, disease duration > 10 years) are removed according to the 3σ criterion to ensure sample validity. (2) Stratification analysis implementation: The "companion therapy type" is the core stratification variable (because this variable has the strongest interference with drug efficacy). First, sub-layers are divided according to the companion therapy type (such as no companion therapy layer, ARB companion therapy layer, anticoagulant companion therapy layer, etc.). Within each sub-layer, patients are grouped according to age and disease duration. The underlying disease is used as the correction factor within the layer to construct a two-level stratification structure of "main layer - subgroup". The stratification principle is that the sample size of each layer is ≥ 50 cases (derived from the total sample size of ≥ 500 cases in the instructions). The distribution of confounding variables within the layer is balanced (standardized mean difference between groups SMD < 0.1). By stratification, patients with different companion therapy backgrounds, age and disease duration ranges are isolated, and the systematic differences of confounding variables between groups are initially eliminated, reducing the macroscopic interference on the "drug intervention-efficacy" association. (3) Propensity Score Matching (PSM) Refinement and Balancing: Based on stratification, PSM is implemented for the remaining confounding variables within each stratum (such as differences in age and underlying diseases within the same concomitant treatment stratum) to further remove micro-level interference: ① Matching Variable Setting: All four types of confounding variables are included in the propensity score calculation model. The dependent variable is "whether the target drug intervention is received" (such as whether immunosuppressants are used), and the independent variables are the four types of pre-processed confounding variables; ② Score Calculation: The propensity score of each patient is calculated using a logistic regression model (range 0-1). The model fit R² ≥ 0.6 to ensure the explanatory power of confounding variables on intervention allocation; ③ Matching Strategy: The 1:1 nearest neighbor matching method is adopted, and the caliper range is set to 0.02 (optimized based on the characteristics of membranous nephropathy samples). After matching, the confounding variables of the intervention group and the control group within each stratum are tested for balance to ensure that the SMD < 0.05 among all variable groups, with no statistical difference; ④ Extreme Value Handling: Samples with propensity scores in the top 5% and bottom 5% are removed to avoid matching bias caused by extreme values.(4) Quantitative stripping using backdoor adjustment formula: ① Weight assignment: Calculate the influence weight of each type of confounding variable using the backdoor adjustment formula, with a value range of 0.1-0.8 and an adjustment step size of 0.05-0.1. Among them, the weight of the accompanying treatment type is the highest (0.3-0.6), and the weights of age, disease course, and underlying disease are each allocated 0.1-0.2; ② Net association strength calculation: Based on the adjusted weights, the association strength of "drug intervention-efficacy" is corrected, and the formula is: ,in To correct the net correlation strength, For uncorrected correlation strength, For the weights of the j-th type of confounding variables, The interference value of the variable on the association; ③ Error control: By iteratively optimizing the weight parameters, ensure that the error of the confounding variable stripping is ≤5% (the optimal effect is achieved when the weight is 0.3-0.6 and the step size is 0.05), and finally obtain the true causal association strength between the drug intervention measures and the efficacy evaluation indicators, so as to provide a precise data foundation for subsequent molecular mechanism link coupling. (5) Validation and adaptation: After the processing is completed, 200 independent samples are used for validation, and the stripping effect of the combined treatment of stratification + PSM + backdoor adjustment and the single method (only stratification or only PSM) is compared to ensure that the false association rate after the combined treatment is ≤3.2%, which is consistent with the validation standard of the causal graph model of drug intervention for membranous nephropathy; at the same time, adapt to the staging characteristics of membranous nephropathy, and perform the above process for patients in stages I-IV respectively, and the weight of the underlying disease can be appropriately increased for late-stage patients (stages III-IV).

[0025] The influence weights of the confounding variables are quantified using a backdoor adjustment formula, which includes a theoretical prototype formula and an adapted simplified formula: the theoretical prototype formula is... The simplified application formula adapted to this solution is as follows: ,in Let X be the average causal effect of drug intervention on efficacy, and let X be the drug intervention measure ( In order to accept intervention, (for comparison) As an indicator for evaluating therapeutic efficacy, For a set of mixed variables, For the probabilities of values ​​taken by the confounding variable; The weights for confounding variables are set at 0.1-0.8, with an adjustment step size of 0.05-0.1. Based on data from 200 patients, the confounding variable stripping error is minimized (≤5%) when the weights are 0.3-0.6 and the step size is 0.05.

[0026] A protein-protein interaction strength quantification algorithm adapted to the STRING database (core formula is...) Where S is the interaction strength, The weight of protein i (values ​​range from 0.1 to 0.9). To determine the direction and magnitude of the interaction probability between protein i and the target protein, the interaction strength threshold of the algorithm is set to 0.6-0.8. The located key nodes must satisfy the expression level difference (Fold Change) > 2 and P < 0.05. The specific technical solution is as follows: (1) Data acquisition and preprocessing from the STRING database: First, the interaction data of target proteins related to membranous nephropathy (such as nephrin, podocin, glomerular basement membrane proteins, etc., corresponding to the protein entities encoded by UniProt ID in the core node) were extracted from the STRING database. The screening criteria were set as follows: the species was limited to humans (Taxonomy ID: 9606), and the interaction types included experimental verification, literature mining and homology sequence prediction. Indirect interaction data with low confidence (database built-in confidence score < 0.4) were removed, and direct physical interaction relationships were retained to form the initial protein interaction dataset.

[0027] (2) Protein weight Assignment: The Analytic Hierarchy Process (AHP) was used to assign values ​​to each candidate protein. Assign weights (Values ​​range from 0.1 to 0.9), and the weighting is based on: ① the protein's central role in the pathological mechanism of membranous nephropathy (e.g., podocyte proteins have a weight ≥0.6, and inflammatory factor-related proteins have a weight 0.3-0.5); ② literature support (the frequency of mentions of the protein in core journals related to membranous nephropathy in the past 10 years, with a weight increase of 0.1-0.2 for frequencies ≥50); ③ experimental verifiability (proteins that are easily detected in in vitro cell experiments have a higher weight than proteins that are difficult to detect). After weighting, a consistency test (CR < 0.1) is required to ensure that the weight allocation is reasonable and a standardized weight matrix is ​​finally formed.

[0028] (3) Interaction probability Sure: This is derived from the interaction confidence score conversion built into the STRING database. The database confidence score ranges from 0 to 1, which is directly mapped to the interaction probability (i.e., confidence score = interaction probability), representing the protein. The possibility of direct interaction with the target protein. Multiple candidate proteins targeting the same target protein. Requires removal Low-probability interactions (<0.3) are avoided to prevent interference with the strength calculation results.

[0029] (4) Interaction strength Calculation: Based on core formula The interaction strength between the target protein and candidate proteins is quantified one by one, with the summation range covering all pre-processed candidate proteins. After calculation, for Normalization is performed (normalization range 0-1) to eliminate intensity scale differences between different target proteins, which facilitates subsequent threshold screening. The higher the value, the stronger the interaction between the candidate protein and the target protein, and the more significant the regulatory response after drug intervention.

[0030] (5) Intensity threshold screening and regulation amplitude judgment: The interaction intensity threshold was set to 0.6-0.8 (based on the validation results of 50 protein pairs in the terminology definition, the link construction accuracy reached 89.2% when the threshold was 0.7, which was taken as the preferred threshold), and the following were screened out. Candidate proteins: ① When When the correlation is weak, the control amplitude is set to "mild"; ② When At that time, it was determined to be a significant control-regulatory association, and the regulatory amplitude was set to "significant". Simultaneously, the values ​​of each candidate protein were recorded. The specific value serves as the core basis for the correlation strength weight in subsequent molecular mechanism links.

[0031] (6) Key node localization and expression level verification: For the selected candidate proteins with high interaction strength, expression level difference detection and statistical verification were performed: ① Expression level detection adopted dual verification using RNA-seq (transcriptome level) and Western Blot (protein level) to obtain protein expression level data of the drug intervention group and the control group respectively; ② The expression level difference FoldChange value (expression level of intervention group / expression level of control group) was calculated, requiring FoldChange>2 (i.e., protein expression level is upregulated or downregulated by more than 2 times); ③ The statistical P value was calculated by independent samples t test (comparison between two groups) or one-way ANOVA (comparison between multiple groups), requiring P<0.05 (the difference is statistically significant), and proteins with no significant difference in expression level or statistical insignificance were removed, finally locating the key nodes with significant changes after drug intervention. P value: The core indicator in statistical testing to determine whether the experimental results are caused by random factors. In essence, it is the probability of observing the current result (or a more extreme result) when the null hypothesis of "no real difference between two / multiple groups of data" is true. In this protocol, the difference in protein expression levels between the drug intervention group and the control group is used to verify the significance of the difference. P < 0.05 indicates that the probability of the difference being caused by random factors is < 5%, and the difference can be judged to be statistically significant (i.e., the change in protein expression is the true effect of drug intervention); P ≥ 0.05 indicates that the difference may be due to random fluctuations, and the corresponding candidate proteins need to be removed. Calculation logic: Based on the quantitative data of protein expression from in vitro cell experiments (Western Blot gray values, ELISA concentration values), the independent samples t-test is used for comparison between two groups (sample size n=3 / group, degrees of freedom df=4, the t-value is calculated by formula and then mapped to the p-value). One-way ANOVA is used for comparison of multiple groups. The F-statistic is calculated first and then the p-value is calculated. SPSS, R language and other tools are commonly used for automatic calculation, provided that the data are normally distributed and homogeneous in variance (outliers are removed according to the 3σ criterion).

[0032] (7) Determining the direction of regulation: Combining the differences in expression levels and the strength of interaction, the direction of regulation of the target protein by key nodes is determined: ① If the expression level of the candidate protein is upregulated and S≥ the threshold, it is determined to be "activation regulation"; ② If the expression level of the candidate protein is downregulated and S≥ the threshold, it is determined to be "inhibition regulation", providing a clear basis for the regulatory relationship in constructing the causal link of the molecular mechanism of "drug-protein-pathology".

[0033] The multi-dimensional verification scheme combines virtual intervention experiments with in vitro cell experiments. The virtual intervention experiments are implemented using the COMSOL Multiphysics 6.0 simulation platform (time step 0.1s, iterations ≤1000, convergence accuracy 1e-6). The in vitro cell experiments use human glomerular mesangial cells. The results of the in vitro cell experiments have a higher weight than those of the virtual intervention experiments. The specific technical solution is as follows: (1) Pre-experimental preparation: First, based on the "drug-protein-pathology" molecular mechanism causal link constructed in step 3, the core verification targets of the two experiments were identified. The expression changes of target proteins (such as nephrin and podocin) and downstream pathological indicators (urine protein quantification corresponding to cell secretion factor levels) were used as the core verification indicators to ensure that the verification dimensions of the two experiments were consistent and the results could be cross-compared. At the same time, the experimental materials and platform were calibrated: the COMSOL platform was loaded with the nephrology-related physical field module (fluid dynamics + biochemical reaction module). Human glomerular mesangial cells (purchased from ATCC cell bank, model HBZY-1) were tested for mycoplasma (negative) and their activity was verified (survival rate ≥95%) before use. Drug reagents were used with clinical-grade purity (≥98%) to avoid impurities from interfering with the experimental results.

[0034] (2) Implementation of virtual intervention experiment: ① Simulation model construction: Based on real human kidney physiological parameters (glomerular diameter, blood flow rate, tissue permeability, etc.), a three-dimensional simplified glomerular mesangial area model is constructed. The protein interaction strength (S value) and drug regulation direction quantified in step 3 are used as model input parameters to build a dynamic simulation link of "drug intervention → protein expression regulation → pathological index response". The model grid division accuracy is set to 0.01mm to ensure accurate calculation of the reaction diffusion process.

[0035] ② Boundary and initial conditions setting: Boundary conditions are configured according to the human physiological environment (blood flow velocity 0.02m / s, temperature 37℃, pH 7.4), and initial conditions correspond to the baseline state before drug intervention (protein expression level and pathological indicators are taken as the mean of the cohort sample); drug intervention parameters are converted to model concentrations according to the clinical routine dose (e.g., the immunosuppressant tacrolimus is set to 5-20nmol / L) to simulate the dynamic response process within 72 hours after a single administration.

[0036] ③ Simulation calculation and result extraction: The transient solver is used to start the calculation with a time step of 0.1s (balancing computational efficiency and accuracy, and avoiding curve distortion due to excessive step size). The maximum number of iterations is set to 1000 (it can automatically terminate if the convergence accuracy of 1e-6 is reached in advance). After the calculation is completed, the core results are extracted, including the curve of target protein expression over time, the peak and steady-state values ​​of pathological indicators (such as simulated urinary protein secretion), and a virtual intervention experiment report is generated.

[0037] (3) In vitro cell experiments: ① Cell culture and grouping: Human glomerular mesangial cells were seeded in 6-well plates (cell density 5×10⁶ cells per well). 5(1) were placed in a 37℃, 5% CO2 constant temperature incubator (95% humidity) and cultured for 24h until they adhered to the wall; they were divided into "drug intervention group + blank control group" and each group was set with 3 replicates. The intervention group was added with the corresponding concentration of drug (consistent with the concentration of the virtual experiment, which can be finely adjusted to adapt to the cell experiment scenario), and the control group was added with an equal volume of physiological saline. They were cultured simultaneously for 72h.

[0038] ② Sample processing and detection: After culture, a dual detection method was used to obtain indicator data: First, Western blotting was used to detect the expression level of target proteins in cells (total protein was extracted, subjected to SDS-PAGE electrophoresis, transferred to a membrane, incubated with primary / secondary antibodies, and then the gray value was quantitatively analyzed by the Bio-Rad ChemiDoc XRS+ gel imaging system); Second, ELISA was used to detect the secretion level of pathologically related factors (such as albumin and inflammatory factor TNF-α) in the cell supernatant, corresponding to the pathological indicator response in the virtual experiment.

[0039] ③ Experimental repeatability assurance: The relative standard deviation (RSD) of the detection results between replicates in the same batch of experiments is ≤8%. If it exceeds the range, the experiment should be repeated. At the same time, a positive control (a group with known effective drug intervention) should be set up to ensure the effectiveness of the experimental system. The difference in the expression level of the target protein in the positive control group should meet the requirement of Fold Change > 2, which is consistent with the core validation standard.

[0040] (4) Result weighting and fusion: The analytic hierarchy process (AHP) was used to set the result weights. The weights for in vitro cell experiments were assigned as 0.6-0.7, and the weights for virtual intervention experiments were assigned as 0.3-0.4. The core basis for this was that in vitro cell experiments are closer to the physiological microenvironment and can reflect the drug regulation effect at the real cell level, while virtual experiments have model simplification errors and therefore have lower weights than in vitro experiments. The fusion method was as follows: the core indicators of the two experiments (the rate of change in protein expression and the magnitude of change in pathological indicators) were weighted and summed separately to obtain a comprehensive verification value. If the deviation between the comprehensive verification value and the causal link prediction value in step 3 was ≤15%, the link verification was considered successful.

[0041] (5) Deviation handling and experimental iteration: If the deviation between the two experimental results is ≤30%, the virtual experimental model parameters are corrected based on the in vitro cell experiment results (such as adjusting the protein interaction rate constant); if the deviation exceeds 30%, the in vitro cell experiment is repeated 2-3 times, each time using the same drug concentration gradient and incubation conditions. If the deviation is still not eliminated after repetition, return to step 3 to reconstruct the molecular mechanism causal link to ensure the rigor of the verification system and the authenticity of the link.

[0042] The reverse iterative optimization cycle is set to 3-6 months. When the amount of new real-world data exceeds 15%-20% of the existing sample size, an emergency iteration process is initiated. During iteration, the node association weights of the structured causal relationship framework are corrected first, and the gradient descent method is used for optimization with a learning rate of 0.001-0.01.

[0043] When the deviation between the virtual intervention experiment and the in vitro cell experiment exceeds 30%, the in vitro cell experiment is repeated 2-3 times, with the same concentration gradient of the drug used in each experiment, and the incubation conditions are 37°C and a 5% CO2 constant temperature incubator. If the deviation is still not eliminated, the molecular mechanism causal link is reconstructed.

[0044] When reversing the association weights of core nodes, the correction range is limited to ±10% to ±15%. If the range is exceeded, L2 regularization is used to constrain the parameter values ​​(regularization coefficient). Calculate the objective function to avoid overfitting the model; In step 1, when removing the false association edges, a dual verification mechanism of "clinical verification + biological experiment" is adopted. Association edges without experimental support are directly removed, and controversial association edges are reviewed by more than three nephrology specialists to determine whether to retain them. Based on the differences in pathological characteristics across different stages of membranous nephropathy, a stage-appropriate drug screening model was constructed using the aforementioned molecular mechanism causal pathway. A weighted summation algorithm was employed to calculate the comprehensive drug score, as shown in the formula below. Where α is the efficacy weight (0.6), β is the safety weight (0.4), E is the drug efficacy assessment value (range 0-1, converted based on the net association strength in step 2), and A is the adverse reaction weight value (assigned values ​​of 0.2, 0.5, and 0.8 according to mild, moderate, and severe, respectively). The personalized drug intervention measures are prioritized by sorting the scores from high to low. The personalized drug intervention program was applied to a real-world cohort, and the model parameters were optimized iteratively by tracking data. The model parameters included the node association weights of the structured causal framework, backdoor adjustment formula parameters, and protein interaction strength thresholds. The iteration cycle was set to 3-6 months. When the amount of new data exceeded 15%-20% of the existing sample size, an emergency iteration process was initiated to form a complete diagnosis and treatment closed loop.

[0045] The data sources for literature mining in step 1 are limited to SCI core journals (JCR Q1-Q2) and Chinese core journals (Peking University core journals) from the past ten years. The extraction method follows the PRISMA criterion. The extracted core node association data is reviewed and confirmed by more than two medical experts before being used for training the causal graph model.

[0046] The specific construction process of constructing a causal graph model for drug intervention in membranous nephropathy is as follows: (1) Data collection and standardized preprocessing: First, collect dual-source data. One is real-world data from the idiopathic membranous nephropathy cohort registration platform (sample size ≥ 500 cases, covering patients in stages I-IV, including those diagnosed with complete clinical data, excluding patients with secondary membranous nephropathy and severe liver and kidney failure complications), extracting individual patient baseline characteristics, drug intervention records, accompanying treatment plans, pathological test results, and efficacy follow-up data; the other is literature from SCI core journals (JCR Q1-Q2) and Peking University core journals in the past ten years (focusing on drug intervention and molecular mechanism research in idiopathic membranous nephropathy, excluding review literature), extracting data such as target protein network association and drug-pathological response patterns. Standardize the collected data: Clinical indicators are uniformly selected according to the "Guidelines for the Diagnosis and Treatment of Kidney Diseases", drug intervention measures adopt ATC drug classification coding, target protein nodes adopt UniProt ID coding, and the data format is unified and entity mapping is achieved by connecting with the membranous nephropathy knowledge graph through a standardized JSON data interface (UTF-8 format, response time ≤500ms). (2) Refined definition and coding of core nodes: Taking the full elements of clinical diagnosis and treatment as the core, six core nodes are determined and dedicated coding is completed to ensure that the nodes are adapted to the drug intervention scenario of membranous nephropathy: ① Drug intervention measures (coded D001-D999): including drug type (glucocorticoids, immunosuppressants, etc.), dosage, administration frequency, course of treatment, and associated drug target information; ② Patient individual baseline characteristics (coded P001-P999): including age, renal function grade, course of disease, comorbidities, and genetic background, highlighting the susceptibility factors of membranous nephropathy; ③ Accompanying treatment plan (coded C001-C999): including antihypertensive drugs, anticoagulants and other non-core kidney disease drugs, and labeled ④ Target protein network nodes (coded M001-M999): focusing on podocyte proteins (nephrin, podocin), glomerular basement membrane proteins, and inflammatory factor-related proteins, and linking them to interaction information in the STRING database; ⑤ Pathological change indicators (coded B001-B999): including quantitative urine protein, serum creatinine, eGFR, and renal tissue pathological score, with reference thresholds set according to the stage of membranous nephropathy; ⑥ Efficacy evaluation indicators (coded E001-E999): including urine protein remission rate, relapse rate, and eGFR change range, corresponding to evaluation criteria for different drug intervention cycles.(3) Causal Link Mining and Construction: Based on standardized data, causal links were constructed through a dual approach of "literature mining + real-world data inference": ① Literature mining: Natural language processing (NLP) tools were used to extract the relationships between core nodes, and nearly 100 core journals were selected to identify "drug-protein", "protein-pathology", and "drug-efficacy" associations (such as glucocorticoids → nephrin upregulation → decreased urinary protein) to form an initial link set; ② Real-world data inference: Propensity score matching and backdoor adjustment formulas were used to eliminate confounding interference, quantify the causal association strength between nodes, retain links with an association strength ≥0.6 (based on the STRING database confidence standard), and supplement individual difference links not covered by the literature (such as drug response differences among patients with different baseline renal function). Finally, a core link of "drug intervention → target protein network node → pathological change index → ​​efficacy evaluation index" was formed, while the moderating effect of patient baseline characteristics and accompanying treatment regimens on the core link (such as the influence of hypertension on the efficacy of immunosuppressants) was also considered. (4) False association removal and link optimization: A triple verification mechanism of "clinical verification + biological experiment + expert review" is adopted to remove false associations: ① Clinical verification: Compare data from 200 independent patients and delete links that conflict with clinical diagnosis and treatment rules (such as drug-protein associations without pharmacological basis); ② Biological experiment verification: Verify the authenticity of core protein links through in vitro cell experiments (human glomerular mesangial cell model) and remove links that have not been verified by experiments; ③ Expert review: Three or more nephrology specialists review controversial links (such as low-strength association links) to determine whether to retain or remove them. (5) Model structure adaptation and initialization: A causal graph model is built using a Bayesian network structure, with six core nodes as network nodes and causal links as directed edges, and each edge is assigned a corresponding association strength weight (based on the backdoor adjustment formula to quantify the results). In view of the differences in pathological characteristics of membranous nephropathy stages I-IV, a dynamic adaptation interface for staging is reserved, and the node weights can be adjusted according to the patient's stage (such as strengthening the weights of renal function-related indicators in late-stage patients). After initializing the model, it was calibrated using 1000 sets of core node mapping samples to ensure that the model output was ≥85% consistent with actual clinical diagnosis and treatment results, forming the final causal graph model specific to membranous nephropathy drug intervention. This provides a basic framework for subsequent confounding processing and molecular link coupling. The model constructed through the above steps not only ensures specific adaptability to the diagnosis and treatment of membranous nephropathy, but also ensures the accuracy of causal inference through multi-source data fusion and multiple verifications, avoiding the link bias problem caused by a single data source in existing technologies.

[0047] The above formulas are all dimensionless calculations, and the preset parameters in the formulas should be set by those skilled in the art according to the actual situation.

[0048] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for intelligent staging and diagnosis of membranous nephropathy based on knowledge graph reasoning, characterized in that, The steps are as follows: Step 1: Construct a causal graph model specific to drug intervention for membranous nephropathy. The core nodes are determined based on all elements of clinical diagnosis and treatment. Through literature mining and training with real-world cohort data, the causal relationship paths between each core node are defined and false association edges are removed to form a structured causal relationship framework. Step 2: Based on the structured causal framework, quantify and eliminate confounding variables, clarify the net association strength between drug intervention measures and efficacy evaluation indicators, and distinguish between their correlation and causality; Step 3: Using a protein interaction strength quantification algorithm adapted to the STRING database, couple key nodes of the target protein network with the causal inference link to determine the direction and magnitude of regulation, locate key nodes that show significant changes after the implementation of the drug intervention, and construct the molecular mechanism causal link of the drug intervention, key nodes of the target protein network, and pathological change indicators of membranous nephropathy. Step 4: Verify the authenticity and stability of the causal link of the molecular mechanism through a multi-dimensional verification scheme to strengthen the biological basis of the molecular mechanism of drug efficacy; Step 5: Combining the differences in pathological characteristics of different stages of membranous nephropathy, construct a stage-appropriate drug screening model based on the aforementioned molecular mechanism causal link to achieve priority ranking of personalized drug intervention measures; Step 6: Apply the personalized drug intervention plan to a real-world cohort, and iteratively optimize the model parameters by tracking data to form a complete diagnosis and treatment closed loop.

2. The intelligent staging and diagnosis method for membranous nephropathy based on knowledge graph reasoning according to claim 1, characterized in that, The core nodes include drug intervention measures, patient baseline characteristics, accompanying treatment plans, target protein network nodes, pathological change indicators, and efficacy evaluation indicators. The dedicated causal graph model for membranous nephropathy drug intervention achieves precise mapping between the core nodes and knowledge graph entities through its built-in structured causal relationship framework and a standardized JSON data interface.

3. The intelligent staging and diagnosis method for membranous nephropathy based on knowledge graph reasoning according to claim 1, characterized in that, The confounding variables mentioned in step 2 include the type of concomitant treatment, patient age, disease course, and underlying diseases. A combination of stratified analysis and propensity score matching is used to isolate the confounding variables from interfering with the association between the drug intervention and the efficacy evaluation indicators.

4. The intelligent staging and diagnosis method for membranous nephropathy based on knowledge graph reasoning according to claim 1, characterized in that, The direction and magnitude of regulation were determined using a protein-protein interaction strength quantification algorithm adapted to the STRING database. The specific technical solution is as follows: First, target proteins related to membranous nephropathy were extracted from the STRING database. Low-confidence indirect interaction data were removed, and direct physical interactions were retained to form an initial protein-protein interaction dataset. Then, an analytic hierarchy process (AHP) was used for each candidate protein... Assign weights Determine the interaction probability Based on the core formula The interaction strength between the target protein and candidate proteins is quantified one by one, with the summation range covering all pre-processed candidate proteins. After calculation Normalization is performed, and intensity threshold screening and control amplitude judgment are based on S.

5. The intelligent staging and diagnosis method for membranous nephropathy based on knowledge graph reasoning according to claim 4, characterized in that, By combining differences in expression levels and the strength of interactions, the regulatory direction of key nodes on target proteins can be determined: if the expression level of a candidate protein is upregulated and S ≥ the threshold, it is determined to be activation-type regulation; if the expression level of a candidate protein is downregulated and S ≥ the threshold, it is determined to be inhibition-type regulation, providing a clear regulatory basis for constructing a causal link between drug-protein-pathological molecular mechanisms.

6. The intelligent staging and diagnosis method for membranous nephropathy based on knowledge graph reasoning according to claim 2, characterized in that, First, four types of confounding variables were categorized and standardized preprocessed to fit the subsequent analysis model: type of concomitant treatment, patient age, disease duration, and underlying disease. The preprocessed data should conform to a normal distribution or be free from severe skewness. Using type of concomitant treatment as the core stratification variable, substrata were first divided according to type of concomitant treatment. Within each substratum, patients were further grouped according to age and disease duration, with underlying disease as an in-stratum correction factor, constructing a two-level stratification structure of main stratum and subgroup. By stratifying, patients with different concomitant treatment backgrounds, age ranges, and disease durations were isolated, initially eliminating systematic differences in confounding variables between groups. A backdoor adjustment formula was then used for quantitative stripping.

7. The intelligent staging and diagnosis method for membranous nephropathy based on knowledge graph reasoning according to claim 6, characterized in that, Accompanying treatment type: Classified according to ATC drug classification codes, excluding symptomatic treatments without clear pharmacological effects; Patient age: According to the guidelines for the diagnosis and treatment of membranous nephropathy, patients were grouped into three groups: ≤40 years old, 41-60 years old, and >60 years old, with each group accounting for ≥20% of the sample size. Disease duration: Based on the time from onset to intervention, patients were divided into three groups: ≤1 year, 1-3 years, and >3 years, to avoid non-linear interference from the duration of the disease on the efficacy assessment. Underlying diseases: Focusing on three common complications: hypertension, diabetes and hyperlipidemia, coded in binary categories of presence or absence, and those with overlapping complications are marked separately.

8. The intelligent staging and diagnosis method for membranous nephropathy based on knowledge graph reasoning according to claim 6, characterized in that, Based on the stratification, PSM was implemented for the remaining confounding variables within each stratum to further remove micro-level interference, as follows: ① Matching variable setting: All four types of confounding variables were included in the propensity score calculation model, with the dependent variable being whether or not the target drug intervention was received, and the independent variables being the four types of pre-treated confounding variables; ② Score calculation: The propensity score for each patient was calculated using a logistic regression model, with a model fit R² ≥ 0.6, ensuring the explanatory power of confounding variables on intervention allocation; ③ Matching strategy: The 1:1 nearest neighbor matching method was adopted, and the caliper range was set to 0.

02. After matching, the balance test of confounding variables between the intervention group and the control group in each layer was performed to ensure that the SMD between all variable groups was <0.05 and there was no statistical difference. ④ Extreme value handling: Samples with propensity score in the top 5% and bottom 5% were removed to avoid matching bias caused by extreme values.

9. The intelligent staging and diagnosis method for membranous nephropathy based on knowledge graph reasoning according to claim 6, characterized in that, The backdoor adjustment formula is used for quantitative stripping, specifically as follows: ① Weight assignment: The influence weight of each type of confounding variable is calculated using the backdoor adjustment formula; ② Net association strength calculation: Based on the adjusted weights, the association strength between drug intervention and efficacy is corrected using the following formula: ,in The corrected net correlation strength, For uncorrected correlation strength, For the weights of the j-th type of confounding variables, ③ Error control: By iteratively optimizing the weight parameters, ensure that the error of removing confounding variables is ≤5%, and finally obtain the true causal association strength between drug intervention and efficacy evaluation indicators, providing a precise data foundation for subsequent molecular mechanism linkage coupling.

10. The intelligent staging and diagnosis method for membranous nephropathy based on knowledge graph reasoning according to claim 1, characterized in that, The model parameters are optimized by iteratively backtracking data. These parameters include node association weights in the structured causal framework, backdoor adjustment formula parameters, and protein-protein interaction strength thresholds.