Method and system for evaluating injury state of near-end renal tubule
By preprocessing urine samples and inspection information and feature dimensionality reduction, combined with multiple machine learning algorithms to train evaluation models, the accuracy and stability of the evaluation of proximal tubular injury status are solved, and a fast and accurate non-invasive diagnosis is achieved.
Patent Information
- Application Number
- CN202510471431.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art has problems of insufficient accuracy and insufficient stability in the evaluation of proximal tubules injury status, especially the difficulty in standardizing the results of urinary amino acid detection. Traditional statistical analysis methods have limitations in processing complex biological data, resulting in high risk of misdiagnosis and missed diagnosis.
By collecting patient urine samples and related inspection information, inputting an evaluation model for feature dimensionality reduction optimization after preprocessing, combining multiple machine learning algorithms to train models, generate state evaluation results and visualize them, and push them to the user.
The rapid, accurate and non-invasive assessment of proximal tubular injury and Fanconi syndrome was achieved, reducing the complexity of the diagnosis process and improving the application value of urinary amino acid analysis in clinical practice.
Smart Images

Figure CN120496795A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medical technology, and in particular to a method and a system for evaluating the state of proximal renal tubular damage. Background Art
[0002] The proximal renal tubules are crucial for renal function, responsible for the reabsorption of most water, electrolytes, and small molecules in the primary urine. However, damage to these tubules can lead to Renal Fanconi Syndrome (RFS), causing severe metabolic disorders such as hypokalemia, hypophosphatemia, hypouricemia, panaminoaciduria, and metabolic acidosis. Because early symptoms are often atypical, current diagnosis relies primarily on multiple biochemical tests and tubular physiology tests, such as 24-hour urinalysis, phosphate clearance tests, and renal glucose threshold tests, to assess tubular dysfunction. However, these testing procedures are cumbersome, often requiring repeated sampling and comprehensive interpretation based on multiple laboratory parameters, increasing the risk of misdiagnosis and missed diagnosis. Furthermore, the limited clinical application of some testing methods, particularly those involving tubular physiology tests, which can only be performed in a few specialized hospitals, makes it difficult for most patients to obtain a clear diagnosis in a timely manner.
[0003] Urinary amino acid analysis is considered an important biomarker for assessing proximal renal tubular function. Studies have shown that proximal renal tubular damage can lead to abnormal excretion of multiple amino acids in urine. However, due to the wide variety of urinary amino acids, the diagnostic value of different amino acids under different etiologies is not completely consistent. Existing analysis methods often rely on empirical judgment and lack a standardized feature screening mechanism, making it difficult to directly use urine amino acid test results for clinical evaluation. Traditional statistical analysis methods have limitations when processing complex biological data and cannot fully explore the changing patterns of urine amino acids under different pathological conditions. Therefore, how to use data-driven methods to extract the most diagnostically valuable features from urine amino acid test results while reducing computational complexity and improving the stability and generalizability of the evaluation is an urgent problem that needs to be solved. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a method and system for assessing the status of proximal renal tubular damage, so as to at least solve the problems of insufficient accuracy and insufficient stability in existing proximal renal tubular damage status assessment schemes.
[0005] In order to achieve the above-mentioned purpose, the first aspect of the present invention provides a method for evaluating the status of proximal renal tubular damage, which includes: collecting detection information of a target patient and performing preprocessing on the detection information; wherein, the detection information includes basic information of a urine sample of the target patient and test result information of the target patient; calling a corresponding evaluation model based on the preprocessed detection information to output a corresponding status evaluation result based on the evaluation model; wherein, the evaluation model is trained based on historical detection information after feature dimensionality reduction; performing visualization processing on the status evaluation result and pushing it to the user end.
[0006] Optionally, the patient's basic information includes: any one or more of the patient's age, patient's gender, patient's cause of disease and patient's medical history information; the test result information includes the urine amino acid concentration information of the target patient's urine sample; the test result information also includes any one or more of urine electrolyte and glucose concentration information, blood electrolyte concentration information and blood gas result information.
[0007] Optionally, preprocessing is performed on the detection information, including: performing a missing value check on the detection information and filling in missing values based on a multiple interpolation method; and performing standardization processing on the detection information after the missing value filling is completed to obtain preprocessed detection information.
[0008] Optionally, the training rules of the evaluation model include: collecting historical detection information and performing preprocessing on the historical detection information; constructing a training data set based on the preprocessed historical detection information, and performing feature dimensionality reduction processing on the training data set to obtain training samples; performing model training based on a preselected machine learning algorithm and the training samples to obtain an evaluation model.
[0009] Optionally, performing feature dimensionality reduction processing on the training data set includes: based on the amino acid-transporter network, screening urine amino acid features related to proximal renal tubular multiple transport abnormalities in the training data set to obtain a screening feature set; in the screening feature set, calculating the PageRank importance score of each urine amino acid in the amino acid-transporter network, and performing ranking of each urine amino acid based on the scoring result; performing filtering of urine amino acids with lower scores based on the ranking result and user needs to obtain a training data set after feature dimensionality reduction.
[0010] Optionally, the filtering of urine amino acids with lower scores is performed based on the ranking results and user needs to obtain a training data set after feature dimensionality reduction, including: using XGBoost as the base model to calculate the importance score of each feature in the ranking results; using a recursive feature elimination method to eliminate low-contribution features with low importance scores in the ranking results; based on SHAP value analysis, calculating the contribution of each urine amino acid in the model decision after eliminating the low-contribution features respectively, filtering urine amino acids with contributions lower than a preset contribution threshold, and obtaining a training data set after feature dimensionality reduction.
[0011] Optionally, the preselected machine learning algorithm includes: a combination of any one or more algorithms among random forest algorithm, ensemble learning algorithm, support vector machine algorithm, gradient boosting decision tree algorithm and neural network algorithm; the model training is performed based on the preselected machine learning algorithm and the training sample to obtain an evaluation model, including: splitting the training sample into a training set and a test set, performing model training based on the training set and the preselected machine learning algorithm to obtain an initial model; performing verification on the initial model based on the test set, and using the model that passes the verification as the evaluation model.
[0012] Optionally, visualization processing is performed on the status assessment results and pushed to the user end, including: calculating the contribution of each urine amino acid and electrolyte feature in the status assessment results based on SHAP analysis, and using corresponding graphical methods to visualize the data; wherein, the graphical method is a bar chart, a heat map or a scatter plot; through the front-end visualization interface, the prediction results, feature importance rankings and individual urine amino acid abnormalities are displayed, and interactive analysis and trend tracking are opened; in response to interactive analysis requests and / or trend tracking requests, an assessment report for the corresponding patient is generated and pushed to the patient end through the cloud.
[0013] The second aspect of the present invention provides a proximal renal tubular injury status assessment system, which includes: an acquisition unit for collecting detection information of a target patient and performing preprocessing on the detection information; wherein the detection information includes basic information of a urine sample of the target patient and test result information of the target patient; an evaluation unit for calling a corresponding evaluation model based on the preprocessed detection information to output a corresponding status assessment result based on the evaluation model; wherein the evaluation model is trained based on historical detection information after feature dimensionality reduction; and a push unit for performing visualization processing on the status assessment result and pushing it to a user end.
[0014] On the other hand, the present invention provides a computer-readable storage medium having instructions stored thereon, which, when executed on a computer, enable the computer to execute the above-mentioned method for assessing the proximal renal tubular injury status.
[0015] Through the above technical solution, the solution of the present invention collects urine samples and related test information from target patients and preprocesses the data to ensure the integrity and standardization of the data. The preprocessed test information is input into a trained evaluation model, which is trained based on feature dimensionality reduction optimization of historical data. It can effectively extract key features related to proximal renal tubular damage and improve diagnostic accuracy. After calculation, the model generates the patient's status assessment results, and further performs visualization processing to make the diagnostic results intuitively presented, and pushes them to the user end to facilitate doctors or patients to view and make decisions. This solution reduces the complexity of the existing diagnostic process, improves the application value of urine amino acid analysis in clinical practice, and realizes rapid, accurate and non-invasive evaluation of proximal renal tubular damage and Fanconi syndrome.
[0016] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following detailed description, they are used to explain the embodiments of the present invention, but do not constitute a limitation of the embodiments of the present invention. In the accompanying drawings:
[0018] Figure 1 is a flowchart of the steps of a method for assessing proximal renal tubular injury status provided by one embodiment of the present invention;
[0019] Figure 2 This is a system structure diagram of a proximal renal tubular injury status assessment system provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0020] The following describes the specific embodiments of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention.
[0021] Figure 1 FIG. 1 is a flow chart of a method for assessing proximal renal tubular injury status provided by one embodiment of the present invention. Figure 1 As shown, an embodiment of the present invention provides a method for assessing the state of proximal renal tubular damage, the method comprising:
[0022] Step S10: collecting detection information of the target patient and performing preprocessing on the detection information.
[0023] Specifically, the test information includes basic information of the target patient's urine sample and test result information of the target patient. The basic information of the patient includes any one or more of the patient's age, gender, cause of disease, and medical history; the test result information includes the concentration information of urine amino acids (a total of 21 amino acids, including: lysine, threonine, arginine, ornithine, aspartic acid, glutamic acid, alanine, leucine, tryptophan, histidine, proline, methionine, valine, glycine, tyrosine, isoleucine, phenylalanine, serine, aspartic acid, glutamine, and citrulline) in the target patient's urine sample; the test result information also includes any one or more of urine electrolyte and glucose concentration information, blood electrolyte concentration information, and blood gas result information.
[0024] Furthermore, preprocessing is performed on the detection information, including: performing a missing value check on the detection information and filling in the missing values based on a multiple interpolation method; and performing standardization processing on the detection information after the missing value filling is completed to obtain preprocessed detection information.
[0025] In an embodiment of the present invention, the present solution optimizes the integrity and stability of urine amino acid test data through a systematic data collection and preprocessing approach, thereby enhancing the accuracy and scalability of the proximal renal tubular injury diagnostic model. During the data collection phase, urine samples and relevant test results are collected from target patients, and basic patient information, including age, gender, etiology, and medical history, is simultaneously recorded. This basic information has a significant impact on urine amino acid metabolism. For example, urine amino acid metabolism patterns vary between individuals of different age groups, while specific genetic diseases or long-term medication use may also affect renal tubular transport function. Therefore, incorporating these factors into data collection helps improve the model's personalized diagnostic capabilities. In addition, urine sample test information includes urine amino acid concentration, electrolyte concentration, glucose concentration, and blood-related indicators such as serum sodium, potassium, phosphorus, and uric acid. This data is used to assess whether proximal renal tubular function is impaired. Combined with blood gas analysis results, it can further determine whether the patient has metabolic acid-base imbalance, thereby assisting in the accurate diagnosis of renal tubular injury.
[0026] Furthermore, after data collection is complete, the test information needs to be preprocessed to ensure data quality for input into the model. Because urine and blood test data can be affected by factors such as patient physiological status, sampling method, and laboratory measurement error, some test parameters may contain missing values. Directly using these values for model training could compromise prediction reliability. Therefore, the data is first checked for missing values and filled using multiple imputation methods. For numerical data such as urine amino acid concentrations and electrolyte levels, K-nearest neighbor imputation or multivariate regression imputation methods are used to fill missing values based on the data's inherent correlations to minimize sample information loss. For categorical variables such as etiology and gender, mode imputation or Bayesian imputation is used to maintain data consistency and statistical plausibility. Furthermore, after filling missing values, the data needs to be normalized to prevent differences in numerical scales across different test parameters from impacting model calculations. Specifically, to prevent extreme outliers from interfering with model training, for urine amino acid measurements outside the physiological reference range, the high-line value is used as the cutoff point, the fold change is calculated, and the corresponding normalized value is assigned to minimize the impact of outliers on the model. In addition, for detection values within the normal range, a value of 0 is assigned to reduce noise interference on non-pathological samples, allowing the model to focus more on identifying pathological features.
[0027] Based on the solution of the present invention, the integrity, stability and comparability of the data are significantly improved, providing high-quality input data for the construction and training of subsequent diagnostic models. By optimizing the data acquisition and processing process, the present invention can effectively reduce the diagnostic bias caused by missing data or scale imbalance, and improve the application value of urine amino acid detection in the assessment of proximal renal tubular damage. In particular, by combining multi-dimensional data such as urine amino acid concentration, urine electrolyte concentration, and blood biochemical indicators, and adopting standardization and outlier processing strategies, the model is made more applicable in different patient groups, and the accuracy and stability of diagnosis are improved. Ultimately, the intelligent diagnostic system constructed by the present invention can more accurately identify proximal renal tubular damage, and push the evaluation results to the user end, realizing rapid, non-invasive, and automated proximal renal tubular function assessment, and providing strong technical support for clinical diagnosis and personalized treatment.
[0028] Step S20: calling a corresponding evaluation model based on the preprocessed detection information to output a corresponding state evaluation result based on the evaluation model.
[0029] Specifically, the training rules of the evaluation model include: collecting historical detection information and performing preprocessing on the historical detection information; constructing a training data set based on the preprocessed historical detection information, and performing feature dimensionality reduction processing on the training data set to obtain training samples; performing model training based on a preselected machine learning algorithm and the training samples to obtain an evaluation model.
[0030] Specifically, performing feature dimensionality reduction processing on the training data set includes: based on the amino acid-transporter network, screening urine amino acid features related to proximal renal tubular multiple transport abnormalities in the training data set to obtain a screening feature set; in the screening feature set, calculating the PageRank importance score of each urine amino acid in the amino acid-transporter network, and performing ranking of each urine amino acid based on the scoring result; based on the ranking result and user needs, filtering urine amino acids with lower scores to obtain a training data set after feature dimensionality reduction.
[0031] Furthermore, the filtering of urine amino acids with lower scores is performed based on the ranking results and user needs to obtain a training data set after feature dimensionality reduction, including: using XGBoost as the base model to calculate the importance score of each feature in the ranking results; using a recursive feature elimination method to eliminate low-contribution features with low importance scores in the ranking results; based on SHAP value analysis, after eliminating low-contribution features respectively, calculating the contribution of each urine amino acid in the model decision, filtering urine amino acids with contributions lower than a preset contribution threshold, and obtaining a training data set after feature dimensionality reduction.
[0032] Furthermore, the preselected machine learning algorithm includes: a combination of any one or more algorithms among random forest algorithm, ensemble learning algorithm, support vector machine algorithm, gradient boosting decision tree algorithm and neural network algorithm; the model training is performed based on the preselected machine learning algorithm and the training sample to obtain an evaluation model, including: splitting the training sample into a training set and a test set, performing model training based on the training set and the preselected machine learning algorithm to obtain an initial model; performing verification on the initial model based on the test set, and using the model that passes the verification as the evaluation model.
[0033] In an embodiment of the present invention, the core goal of the evaluation model is to accurately evaluate the patient's proximal tubular function status based on urine amino acid detection data, and ultimately determine whether there is damage or Fanconi syndrome. In order to achieve efficient and accurate model construction, it is necessary to collect a large amount of historical detection information and preprocess it to ensure the integrity, stability and consistency of the data. Subsequently, feature dimensionality reduction is performed based on these preprocessed data to extract the most diagnostically valuable biomarkers, and then the final evaluation model is constructed using a machine learning algorithm. This technical solution improves the application value of urine amino acids in the diagnosis of proximal tubular damage by optimizing feature screening, dimensionality reduction and model training methods, and enhances the interpretability, stability and generalization ability of the evaluation model.
[0034] Furthermore, before model training, a large amount of historical test data needs to be collected, including the patient's urine sample test results, blood test indicators, basic information, etc. Since these data may have problems such as missing values, noise or inconsistent scales, they need to be subjected to standardized preprocessing, including missing value filling, normalization and outlier correction, to obtain a high-quality training data set. Based on the processed data set, samples for training are constructed, and feature dimensionality reduction is further performed to screen urinary amino acids and other related indicators that are most influential in the evaluation of renal tubular function. After data preprocessing and feature screening are completed, a pre-selected machine learning algorithm is used for model training, and a stable and reliable evaluation model is finally obtained.
[0035] Furthermore, the feature dimensionality reduction of the training dataset includes the following steps:
[0036] 1) Amino Acid Transporter Network (AATN) Screening: The proximal renal tubules rely on multiple transporters for amino acid reabsorption, such as SLC3A1, SLC6A19, SLC7A9, and SLC1A1. Damage to these transporters can lead to abnormal excretion of specific urinary amino acids. Therefore, an amino acid transporter network (AATN) was constructed based on the urinary amino acid metabolic pathways to identify urinary amino acid signatures closely associated with proximal renal tubular dysfunction in the training dataset.
[0037] In the AATN, each urinary amino acid is treated as a network node and associated with its corresponding transporter and function. First, the importance score of each urinary amino acid in the network is calculated using the PageRank algorithm. The urinary amino acids are then ranked according to the scores to obtain a set of key urinary amino acids that have a significant impact on renal tubular function. This method prioritizes the retention of amino acids that are sensitive to proximal tubular damage, such as citrulline, phenylalanine, aspartic acid, and glutamine, while eliminating urinary amino acids that may be unrelated to pathological changes or have a minor impact.
[0038] 2) Recursive elimination of low-contribution urine amino acid features: After the PageRank ranking is completed, there may still be some urine amino acids with low contributions, which affects the model calculation efficiency. In order to further optimize the feature dimensionality reduction process, the present invention adopts XGBoost as the base model to calculate the importance of each urine amino acid, and combines the recursive feature elimination (RFE, Recursive Feature Elimination) method to gradually eliminate features with low contribution. The specific execution process of RFE includes: calculating the initial feature importance based on XGBoost, sorting the urine amino acid features, and deleting the features with low ranking. RFE is used to recursively remove low-contribution features, and each iteration eliminates urine amino acids that have a smaller impact on model performance, and recalculate the influence of the remaining features until the model performance (such as AUC value) reaches the optimal state. The model interpretability is analyzed based on the SHAP (Shapley Additive Explanations) value, the contribution of the urine amino acid feature in the final model decision is calculated, and a contribution threshold is set to further eliminate features with low contribution.
[0039] This feature dimensionality reduction method can significantly reduce unnecessary urine amino acid data, improve model calculation efficiency, and enhance the model's ability to identify disease states.
[0040] Furthermore, the present invention uses multiple machine learning algorithms to train the evaluation model to ensure the robustness and efficiency of the model. The pre-selected machine learning algorithms include:
[0041] 1) Random Forest (RF): Suitable for processing multi-feature data, it can avoid overfitting and provide good feature selection capabilities.
[0042] 2) Ensemble learning algorithm (XGBoost): It can optimize feature selection and improve model prediction accuracy, and is suitable for large-scale data training.
[0043] 3) Support Vector Machine (SVM): Suitable for high-dimensional data classification, can effectively handle nonlinear relationships and improve classification performance.
[0044] 4) Gradient Boosting Decision Tree (GBDT): Suitable for prediction tasks and has strong feature expression capabilities.
[0045] 5) Neural Network (ANN): Suitable for complex pattern recognition tasks, it can improve the generalization ability of the model through multi-layer feature abstraction.
[0046] To ensure the robustness and generalization of the evaluation model, the present invention uses a stratified sampling method, which divides the training data set into a training set and a test set based on the class ratio. 80% of the data is used for training the model to ensure that it can fully learn the relationship between urine amino acid characteristics and proximal renal tubular damage; 20% of the data is used for testing to evaluate the model's predictive ability on unseen data to avoid overfitting or underfitting.
[0047] In a possible embodiment, in the model training phase, training is first performed based on the training set, and a grid search (Grid Search) combined with a cross-validation (Cross-Validation) method is used to optimize hyperparameters, such as learning rate (learning rate), decision tree depth (tree depth), minimum sample split number (min_samples_split) and other key parameters to ensure that the model has the best performance in a complex data environment. After training, the model is verified on the test set, and indicators such as AUC (area under the curve), sensitivity (Sensitivity), specificity (Specificity), and F1 score are calculated to comprehensively evaluate the classification effect of the model and ensure its applicability and stability in different patient groups. Finally, based on the performance comparison of multiple machine learning algorithms, the model with the highest AUC value and the most stable performance is selected as the final evaluation model, and SHAP (Shapley Additive Explanations) is used for interpretability analysis to quantify the contribution of each urine amino acid feature in the model decision, optimize the feature weight distribution, and ensure that the model has higher credibility and interpretability in clinical applications.
[0048] In one possible embodiment, in an intelligent diagnostic system for the assessment of proximal renal tubular damage using urine amino acid detection, different data features have different statistical distributions and medical significance. Therefore, it is crucial to select an appropriate data-driven model. In order to optimize the performance of the model, the present invention proposes a pre-selection scheme for a machine learning algorithm for model training based on different data features, ensuring that the final evaluation model can not only accurately identify pathological changes in urine amino acid characteristics, but also maintain good stability and generalization capabilities in a complex medical data environment.
[0049] Specifically, based on the data type and feature distribution, a preliminary analysis of urine amino acid concentrations, electrolyte indicators and related blood biochemical data was performed. Urine amino acid data are usually continuous numerical variables, and their distribution may be skewed (such as abnormal increases in certain amino acids), so they are suitable for tree models (such as XGBoost, random forests) or deep learning methods (such as neural networks). At the same time, the patient's basic information (such as gender, cause, and past medical history) is categorical data and can be processed by support vector machines (SVM) or logistic regression to establish pattern distinctions for specific categories of patients. In addition, blood and urine electrolyte data not only have a wide numerical distribution, but may also have nonlinear correlations with urine amino acid variables, so ensemble learning (such as GBDT) can further optimize model performance and extract potential interactions.
[0050] Based on the above data characteristics, the pre-selected machine learning algorithm scheme is proposed as follows:
[0051] 1) Random Forest (RF): Applicable to high-dimensional data, can automatically select key features, and is suitable for preliminary feature screening.
[0052] 2) XGBoost (Extreme Gradient Boosting): It excels in feature selection and classification tasks and can enhance the ability to identify the nonlinear relationship between urinary amino acids and renal tubular damage.
[0053] 3) Support vector machine (SVM): It is suitable for situations where the data distribution is complex but the number of features is small, and can be used to analyze the correlation between urine amino acids and the patient's medical history.
[0054] 4) Gradient Boosting Decision Tree (GBDT): Suitable for nonlinear data, can integrate urine and blood electrolyte information, and improve the diagnostic stability of the model.
[0055] 5) Neural Network (ANN): Suitable for large-scale data learning, able to simulate the characteristics of complex metabolic networks, improve the generalization ability of the model, and especially has advantages in cross-population applicability.
[0056] Finally, during the model training phase, different algorithms were evaluated using indicators such as AUC (area under the curve), sensitivity, specificity, and F1 score, and the model with the highest AUC and the most stable performance was selected as the final evaluation model to ensure the efficiency and clinical feasibility of urine amino acid detection in the assessment of proximal renal tubular damage.
[0057] Step S30: Perform visualization processing on the status assessment result and push it to the user end.
[0058] Specifically, the status assessment results are visualized and pushed to the user end, including: calculating the contribution of each urine amino acid and electrolyte feature in the status assessment results based on SHAP analysis, and using corresponding graphical methods to visualize the data; wherein, the graphical method is a bar chart, a heat map or a scatter plot; through the front-end visualization interface, the prediction results, feature importance rankings and individual urine amino acid abnormalities are displayed, and interactive analysis and trend tracking are opened; in response to interactive analysis requests and / or trend tracking requests, an assessment report for the corresponding patient is generated and pushed to the patient end through the cloud.
[0059] In this embodiment of the present invention, to improve the interpretability and clinical usability of urine amino acid testing in the diagnosis of proximal renal tubular injury, the state assessment results generated by the model are visualized and pushed to the user. This process not only intuitively demonstrates the association between urine amino acids and renal tubular injury but also provides personalized diagnostic analysis and trend tracking, helping doctors and patients more clearly understand the assessment results and providing a basis for subsequent intervention and management.
[0060] Specifically, in order to quantify the impact of urinary amino acids and electrolytes in diagnostic evaluation, the present invention uses SHAP (Shapley Additive Explanations) to analyze and calculate the contribution of each feature in the model decision. The SHAP value is based on the principles of game theory and can explain the impact of each input variable on the final prediction result of the model, thereby clarifying which urinary amino acid features play a decisive role in the diagnosis of proximal renal tubular damage. For example, certain urinary amino acids (such as citrulline, phenylalanine, and aspartic acid) have a high contribution, indicating that they play an important role in the diagnosis of Fanconi syndrome, while amino acids with lower contributions may have a weaker correlation with the disease state. In order to make the contribution information intuitive, this system uses a variety of visualization methods such as bar graphs, heat maps, and scatter plots:
[0061] 1) Histogram: Used to display the SHAP value ranking of each urinary amino acid and electrolyte index, helping doctors understand the most important biomarkers.
[0062] 2) Heat map: used to display abnormal patterns of urine amino acids in different patients. The color gradient indicates the degree of abnormality of different urine amino acids, facilitating rapid screening of high-risk patients.
[0063] 3) Scatter plot: used to analyze the individual distribution of urine amino acid concentrations and, combined with the statistical distribution curve, to show the degree of deviation between the patient's indicators and those of the normal population.
[0064] Furthermore, based on the visual data processing results, the system displays the evaluation model's prediction results, feature importance rankings, and individual patient urine amino acid abnormalities on the front-end interface. Doctors or patients can query specific urine amino acid indicators through the interactive interface and further conduct interactive analysis and trend tracking, mainly including:
[0065] 1) Dynamic monitoring of individual urine amino acids: Analyze the patient's historical test data and generate a urine amino acid change trend chart to help doctors assess the progression of the disease.
[0066] 2) Multi-patient data comparison: allows doctors to compare urine amino acid profiles of different patients, identify common or specific features, and improve diagnostic reliability.
[0067] 3) Customized threshold warning: The system supports doctors to set personalized abnormal thresholds and provides automatic reminders when urine amino acid or electrolyte indicators exceed the warning range to assist clinical decision-making.
[0068] In one possible implementation, the corresponding system architecture includes:
[0069] 1) Front-end: Build the user interface based on HTML5 and Python;
[0070] 2) Backend: Developed in Python and integrated with machine learning methods;
[0071] 3) Database: Use MySQL to store patient sample data and analysis results.
[0072] Furthermore, based on user interaction, the system can respond to interactive analysis requests and / or trend tracking requests to generate a detailed assessment report for a specific patient. The assessment report includes:
[0073] 1) Basic information of the patient (age, gender, cause of disease, medical history, etc.).
[0074] 2) Urine amino acid test results and abnormal item annotation (based on visual analysis to display the patient's urine amino acid levels).
[0075] 3) Personalized analysis and clinical advice (combining model analysis results to provide personalized health advice).
[0076] 4) Prediction of disease progression trends (based on historical data, providing possible development trends).
[0077] Generated assessment reports can be pushed to doctors or patients via the cloud, allowing doctors to review the patient's condition online or allowing patients to access test results independently, reducing doctors' need for repeated explanations and improving diagnosis and treatment efficiency. Furthermore, reports can be accessed from multiple devices (PC and mobile), and integrated with electronic medical records (EMRs) to enable remote medical data sharing and improve the utilization of medical resources.
[0078] Figure 2 FIG. 1 is a system structure diagram of a proximal renal tubular injury status assessment system provided by one embodiment of the present invention. Figure 2 As shown, an embodiment of the present invention provides a proximal tubular injury status assessment system, the system comprising: an acquisition unit, for acquiring detection information of a target patient and performing preprocessing on the detection information; wherein the detection information comprises basic information of a urine sample of the target patient and test result information of the target patient; an evaluation unit, for calling a corresponding evaluation model based on the preprocessed detection information, so as to output a corresponding status assessment result based on the evaluation model; wherein the evaluation model is obtained by training based on historical detection information after feature dimensionality reduction; a push unit, for performing visualization processing on the status assessment result and pushing it to a user end.
[0079] An embodiment of the present invention further provides a computer-readable storage medium having instructions stored thereon, which, when executed on a computer, enables the computer to execute the above-mentioned method for assessing the state of proximal renal tubule damage.
[0080] Those skilled in the art will appreciate that all or part of the steps in the methods of the aforementioned embodiments can be accomplished by instructing the relevant hardware through a program, which is stored in a storage medium and includes a number of instructions for causing a single-chip microcomputer, chip, or processor to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0081] The above describes in detail the optional embodiments of the present invention in conjunction with the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the technical concept of the embodiments of the present invention, a variety of simple modifications can be made to the technical solutions of the embodiments of the present invention, and these simple modifications all fall within the scope of protection of the embodiments of the present invention. It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner unless there is any contradiction. In order to avoid unnecessary repetition, the embodiments of the present invention will no longer describe the various possible combinations separately.
[0082] In addition, the various embodiments of the present invention may be arbitrarily combined, and as long as they do not violate the concept of the embodiments of the present invention, they should also be regarded as the contents disclosed in the embodiments of the present invention.
Claims
1. A method for assessing proximal renal tubular injury status, characterized in that: The method comprises: Collecting test information of the target patient and performing preprocessing on the test information; wherein, The test information includes basic information of the target patient's urine sample and test result information of the target patient; Based on the pre-processed detection information, a corresponding evaluation model is called to output a corresponding state evaluation result based on the evaluation model; wherein, The evaluation model is obtained by training based on historical detection information after feature dimensionality reduction; Visualization processing is performed on the status assessment result and pushed to the user end.
2. The method according to claim 1, characterized in that The patient's basic information includes: Any one or more of the patient's age, gender, etiology, and medical history; The test result information includes urine amino acid concentration information of the target patient's urine sample; The test result information also includes any one or more of urine electrolyte and glucose concentration information, blood electrolyte concentration information, and blood gas result information.
3. The method according to claim 1, characterized in that Performing preprocessing on the detection information, including: Performing a missing value check on the detection information and performing missing value filling based on a multiple imputation method; The detection information after missing value filling is standardized to obtain preprocessed detection information.
4. The method according to claim 1, wherein The training rules of the evaluation model include: Collecting historical detection information and performing preprocessing on the historical detection information; Constructing a training data set based on the preprocessed historical detection information, and performing feature dimensionality reduction processing on the training data set to obtain training samples; Model training is performed based on a preselected machine learning algorithm and the training samples to obtain an evaluation model.
5. The method according to claim 4, characterized in that Performing feature dimensionality reduction processing on the training data set includes: Based on the amino acid-transporter network, urinary amino acid features associated with multiple abnormalities in proximal renal tubular transport are screened in the training data set to obtain a screening feature set; In the screening feature set, calculating the PageRank importance score of each urinary amino acid in the amino acid-transporter network, and performing ranking of each urinary amino acid based on the score result; Based on the ranking results and user needs, urine amino acids with lower scores are filtered to obtain a training dataset after feature dimensionality reduction.
6. The method according to claim 5, characterized in that The filtering of urine amino acids with lower scores is performed based on the ranking results and user needs to obtain a training data set after feature dimensionality reduction, including: Using XGBoost as the base model, calculate the importance score of each feature in the ranking results; A recursive feature elimination method is used to remove low-contribution features with low importance scores from the sorting results; Based on SHAP value analysis, the contribution of each urinary amino acid in the model decision was calculated after removing low-contribution features. The urinary amino acids with contributions lower than the preset contribution threshold were filtered out to obtain the training data set after feature dimensionality reduction.
7. The method according to claim 4, characterized in that Preselected machine learning algorithms include: A combination of any one or more of the following algorithms: random forest algorithm, ensemble learning algorithm, support vector machine algorithm, gradient boosting decision tree algorithm, and neural network algorithm; The performing model training based on the preselected machine learning algorithm and the training sample to obtain an evaluation model includes: Splitting the training samples into a training set and a test set, performing model training based on the training set and a preselected machine learning algorithm to obtain an initial model; The initial model is verified based on the test set, and the model that passes the verification is used as the evaluation model.
8. The method according to claim 1, characterized in that Performing visualization processing on the status assessment results and pushing them to the user end includes: The contribution of each urinary amino acid and electrolyte characteristic in the status assessment results was calculated based on SHAP analysis, and the data were visualized using corresponding graphical methods; The graphical representation is a bar graph, a heat map or a scatter plot; The prediction results, feature importance rankings, and individual urine amino acid abnormalities are displayed through a front-end visualization interface, enabling interactive analysis and trend tracking. In response to interactive analysis requests and / or trend tracking requests, an assessment report for the corresponding patient is generated and pushed to the patient via the cloud.
9. A proximal renal tubular injury status assessment system, characterized in that: The system comprises: The acquisition unit is used to acquire the detection information of the target patient and perform preprocessing on the detection information; wherein, The test information includes basic information of the target patient's urine sample and test result information of the target patient; An evaluation unit is used to call a corresponding evaluation model based on the preprocessed detection information to output a corresponding state evaluation result based on the evaluation model; wherein, The evaluation model is obtained by training based on historical detection information after feature dimensionality reduction; The push unit is used to perform visualization processing on the status assessment result and push it to the user end.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the method for assessing the proximal renal tubular injury state according to any one of claims 1 to 8.