Method for extracting kidney disease and its staging probability distribution vector based on real-time endogenous fluorescence imaging data of urine protein gel electrophoresis
By using real-time endogenous fluorescence imaging via urinary protein gel electrophoresis and deep learning technology, the problems of high subjectivity, low automation, and lack of probabilistic output in kidney disease staging have been solved. This has enabled efficient and accurate extraction of the probability distribution vector for kidney disease staging, supporting clinical decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI JIAOTONG UNIV
- Filing Date
- 2026-02-11
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies for kidney disease staging suffer from high subjectivity, low automation, and a lack of probabilistic output, making it impossible to automatically and efficiently extract the probability distribution vector of kidney disease and its staging from urine protein test data.
The study employed real-time intrinsic fluorescence imaging of urine protein gel electrophoresis (SDS-GE-IFI) combined with deep learning (DL) technology. By acquiring real-time intrinsic fluorescence imaging data of urine protein gel electrophoresis, a pre-trained convolutional neural network model was used to extract high-dimensional feature vectors, and a classifier was used to generate probability distribution vectors for kidney disease staging.
It has achieved fully automated processing of kidney disease staging, improved detection efficiency, reduced consumable costs, enhanced the reliability and accuracy of results, provided quantitative staging possibilities, and supported clinical decision-making.
Smart Images

Figure CN121709222B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioanalytical technology, and in particular to a method for extracting the probability distribution vector of kidney disease and its stages based on real-time endogenous fluorescence imaging data from urinary protein gel electrophoresis. Background Technology
[0002] The location and accurate staging of kidney damage in chronic kidney disease (CKD) are core aspects of clinical treatment decisions, directly impacting patient prognosis and the development of personalized treatment plans. Currently, the location of kidney damage relies on the analysis of different molecular weight urinary protein components and their levels to predict tubular and / or glomerular lesions; while kidney disease staging mainly relies on the integrated analysis of multiple clinical indicators, such as urinary albumin testing, urinary creatinine levels, urinary protein-creatinine ratio (UACR), and clinical symptoms and signs. However, this comprehensive assessment method has significant limitations: firstly, it relies on subjective experience and lacks standardized and automated processes, making staging results susceptible to human influence; secondly, existing methods typically output a single indicator or qualitative judgment, failing to provide a probabilistic staging vector, making it difficult to quantify staging uncertainty or express the possibility of multiple stages, thus limiting the accuracy, precision, and clinical applicability of staging.
[0003] Regarding urine protein detection technology, while existing methods can provide some data support, none can directly generate probability distribution vectors for kidney disease and its stages. For example:
[0004] Immunoturbidimetry has the advantage of rapid quantification, but it requires individual detection of each protein biomarker (such as IgG, TRF, HSA and α1-MG), and cannot simultaneously obtain full-spectrum urinary protein fingerprint information. Moreover, the detection cost increases exponentially with the number of biomarkers, making it difficult to support multi-biomarker combined staging analysis.
[0005] Traditional protein gel electrophoresis (SDS-GE) can achieve full-spectrum separation of urinary proteins, providing molecular weight and relative content data, which can be used to differentiate glomerular / tubular damage. However, this method relies on Coomassie brilliant blue staining, which has problems such as low sensitivity, cumbersome operation (such as manual staining, destaining, and scanning), and long time consumption (usually several hours). Moreover, the output results are mostly qualitative images or simple chromatograms, which require manual interpretation and cannot automatically extract high-dimensional feature vectors to support probabilistic staging.
[0006] Although capillary electrophoresis (CE) is widely used in clinical serum protein and hemoglobin analysis, due to the low content of urinary protein and its own insufficient sensitivity, CE is difficult to use directly for urine samples without urinary protein pre-enrichment pretreatment, and it is even more impossible to achieve efficient extraction of probability vectors.
[0007] Furthermore, while existing electrophoresis techniques (such as CN111413390A) attempt to optimize the electrophoresis process, they remain limited to single-protein detection or gel modification, failing to address the issue of automatic extraction of staging vectors. In the prediction of kidney disease and its staging, urinary protein fingerprint data (such as the molecular weight distribution, concentration, and proportion of various urinary protein markers) contains rich information about kidney disease and its staging. However, existing methods lack advanced algorithms such as deep learning for feature abstraction and probabilistic modeling, resulting in low staging efficiency and limited accuracy.
[0008] Therefore, there is an urgent need in this field for a method that can automatically and efficiently extract the probability distribution vector of kidney disease and its stage from urine protein test data, in order to solve the following core problems existing in the prior art:
[0009] Highly subjective: Relies on comprehensive human judgment and lacks objective, standardized vector output;
[0010] Low level of automation: The detection process is cumbersome, involves a lot of manual operation, and is greatly affected by human factors, making it impossible to achieve integrated processing from data acquisition to probability vector generation;
[0011] Lack of probabilistic output: Existing methods cannot quantify the uncertainty of kidney disease staging, making it difficult to support clinical decision-making regarding the possibility of multiple stages. Summary of the Invention
[0012] To address the aforementioned problems, the purpose of this invention is to propose a method for extracting the probability distribution vector of nephropathy and its stages based on real-time intrinsic fluorescence imaging (IFI) data from urinary protein gel electrophoresis.
[0013] The objective of this invention can be achieved through the following technical solutions:
[0014] This invention provides a method for extracting the probability distribution vector of nephropathy and its stage based on real-time endogenous fluorescence imaging data from urinary protein gel electrophoresis, comprising the following steps:
[0015] S1. Obtain the molecular weight bands and their gray values of urinary proteins extracted by real-time intrinsic fluorescence imaging technology of urinary protein gel electrophoresis.
[0016] S2. Based on the pre-established standard curve of concentration-grayscale value of urinary protein biomarker standard, the grayscale value data is converted into the absolute quantitative concentration data of the corresponding urinary protein component;
[0017] S3. Integrate the quantitative concentration data of all urinary protein components to generate a digital map characterizing the overall distribution pattern of urinary protein components of different molecular weights;
[0018] S4. Input the digitized map into a pre-trained convolutional neural network model, and calculate and output a high-dimensional feature vector representing the composition characteristics of urinary protein through the forward propagation of the model.
[0019] S5. Input the high-dimensional feature vector into the classifier to calculate the probability distribution vector associated with different kidney disease stages. The probability distribution vector is used as kidney disease stage prediction information.
[0020] Furthermore, in S1, the specific process of obtaining the grayscale data of each molecular weight band of urinary protein extracted by real-time intrinsic fluorescence imaging technology of urinary protein gel electrophoresis includes:
[0021] Using a CMOS camera equipped with a 275 nm ultraviolet LED excitation source and a 360 nm filter, real-time fluorescence images were acquired during polyacrylamide gel electrophoresis. The molecular weight bands of each urinary protein were automatically identified by the accompanying image analysis software, and their corresponding pixel gray values were calculated.
[0022] Furthermore, in S2, the concentration-grayscale value standard curve of the urinary protein marker standard is established by detecting, but not limited to, known immunoglobulin IgG, human transferrin TRF, human serum albumin HSA and α1-microglobulin α1-MG standard solutions. That is, the calibration equation is obtained by linear or nonlinear fitting with concentration as the abscissa and the corresponding band grayscale value as the ordinate.
[0023] In the process of converting the grayscale data into the absolute quantitative concentration data of the corresponding urinary protein components, the grayscale value of the sample to be tested is substituted into the calibration equation to calculate the absolute concentration of each urinary protein component.
[0024] Furthermore, in S3, the specific process of integrating the quantitative concentration data of all components to generate a digital map characterizing the overall distribution pattern of urinary protein components of different molecular weights includes:
[0025] Using the molecular weight of urinary protein as the x-axis and the relative or absolute concentration as the y-axis, the concentration data of each component are reconstructed into a structured data matrix containing the molecular weight identifier and concentration value of urinary protein.
[0026] Furthermore, in S4, the convolutional neural network model is trained using the following process:
[0027] A training set was constructed by collecting urinary protein digitization maps of clinically diagnosed samples containing various kidney diseases and their stages;
[0028] The training set is divided into a training subset, a validation subset, and a test subset;
[0029] The VGG19 network architecture was used as the convolutional neural network model, with the digitized map as input and the corresponding kidney disease stage label as the supervision signal.
[0030] The model parameters are dynamically adjusted by optimizing the loss function, and the hyperparameters are tuned using a validation subset to determine the network structure.
[0031] The model's generalization ability and accuracy are validated on a subset of tests until the model converges, resulting in a pre-trained model capable of extracting high-dimensional feature vectors.
[0032] Further, the high-dimensional feature vector is input into the classifier, S5, to calculate the probability distribution vector associated with different stages of kidney disease. The specific process includes:
[0033] After receiving the high-dimensional feature vector, the classifier calculates the probability value of the sample belonging to each stage of kidney disease through multi-class logistic regression.
[0034] Arrange all probability values in order of stage category to form a probability distribution vector, where the sum of all elements in the vector is 1, and the stage of kidney disease corresponding to the highest probability value is the main stage tendency predicted by the model.
[0035] Furthermore, the classifier is a Softmax function.
[0036] Furthermore, the high-dimensional feature vector specifically includes the following feature parameters:
[0037] a) Band distribution characteristics of high molecular weight region, medium molecular weight region and low molecular weight region according to molecular weight;
[0038] b) Concentration ratio characteristics including, but not limited to, IgG, transferrin, albumin and α1-microglobulin;
[0039] c) Spatial weight features of albumin strips generated based on gradient-weighted class activation mapping.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] 1) This invention achieves fully automated extraction of the probability distribution vectors of kidney disease and its stages by integrating urine protein gel electrophoresis-real-time intrinsic fluorescence imaging (SDS-GE-IFI) with deep learning technology, which has significant beneficial effects. First, this method eliminates the cumbersome manual operations such as staining, destaining and scanning in traditional urine protein gel electrophoresis detection. The entire process, from grayscale data acquisition to concentration conversion (S2), is based on standardized curves, making the operation simple and fast. It only requires a small amount of urine sample, which greatly improves detection efficiency and reduces consumable costs.
[0042] 2) Secondly, by generating a digital map (S3) and inputting it into a pre-trained convolutional neural network model, high-dimensional feature vectors can be automatically extracted, avoiding subjective interpretation errors, ensuring the objectivity and repeatability of data processing, and providing an efficient and low-cost kidney disease screening tool for clinical use.
[0043] 3) Furthermore, this invention demonstrates superior technical performance: On one hand, the high sensitivity of real-time imaging based on IFI, combined with deep learning (DL) feature extraction, significantly improves the accuracy of kidney disease and its staging. The probability distribution vector (S5) can quantify the likelihood of different stages, outputting richer predictive information to support clinical decision-making. On the other hand, the entire method achieves end-to-end automation from data acquisition to vector generation, which not only improves the analysis speed (e.g., model prepropagation only requires milliseconds) but also enhances the reliability of the results, providing a reliable data foundation for accurate staging and personalized treatment of kidney disease, and has broad clinical application prospects. Attached Figure Description
[0044] Figure 1 A schematic diagram of the process for using the DL-SDS-GE-IFI method to assist in kidney disease screening.
[0045] Figure 2 This is a schematic diagram of deep learning using the SDS-GE-IFI method and the VGG19 architecture.
[0046] Figure 3 The results of SDS-GE-IFI for four common urinary protein standards (IgG, TRF, HAS, α1-MG) in urine samples.
[0047] Figure 4 The relationship between the gray values of four urinary protein standards and urinary protein concentrations based on SDS-GE-IFI electrophoresis separation and detection.
[0048] Figure 5A SDS-GE-IFI electrophoresis image of a urine sample from a healthy person.
[0049] Figure 5B SDS-GE-IFI electrophoresis images of urine samples from diabetic patients and patients with kidney disease at different stages.
[0050] Figure 6 From Figure 5B Chromatogram of the conversion of standard protein molecular weight electrophoresis pattern in lane 1 of urine protein SDS-GE-IFI electrophoresis imaging.
[0051] Figure 7 From Figure 5BChromatograms of urinary protein electrophoresis patterns transformed from those of patients with chronic kidney disease in lane 7 of SDS-GE-IFI electrophoresis imaging, along with molecular weight, peak area, and relative content.
[0052] Figure 8 This study compares different stages of nephropathy based on SDS-GE-IFI urinary protein analysis.
[0053] Figure 9 Performance analysis of training for 7 CNN models.
[0054] Figure 10 This provides the classification performance data for VGG19 on the test set.
[0055] Figure 11 This provides a visual interpretation of how a DL model processes SDS-GE-IFI electrophoresis patterns using gradient-weighted class activation mapping (Grad-CAM). Detailed Implementation
[0056] Based on its fundamental concept, this invention provides a novel method for detecting and staging nephropathy using deep learning (DL)-urine protein SDS-GE-IFI. Urinary proteins are directly separated by SDS-GE electrophoresis. Next, endogenous ultraviolet fluorescence (IFI) at approximately 360 nm is generated by excitation with 275 nm ultraviolet light and the aromatic amino acids (tryptophan, tyrosine, and phenylalanine) of the protein, enabling qualitative and quantitative analysis of urinary protein components. Further DL training of the urinary protein fingerprint and comparative analysis with clinical results yields feature vectors for nephropathy staging. These feature vectors characterize the patterns and trends of urinary protein changes across different nephropathy stages.
[0057] The probability distribution vector of kidney disease and its stages described in this invention is the core predictive information ultimately output by this method. Specifically, it represents a discrete probability sequence that has undergone mathematical normalization. The dimension of this vector strictly corresponds to the total number of clinical stages of kidney disease. Each element in the vector is a probability value between 0 and 1, representing the likelihood of the urine sample being classified as stage 1, stage 2, up to stage N, and the sum of the probabilities of all elements is always 1. Its generation principle lies in the fact that the high-dimensional feature vector extracted in previous steps has encoded complex pattern information in the urinary protein digitization map that is deeply related to the nature and severity of kidney damage. When this feature vector is input into the Softmax classifier, the classifier uses its inherent multi-class logistic regression calculation to comprehensively interpret and weight these deep features, thereby mapping and transforming the abstract, high-dimensional feature representation into an intuitive, interpretable, and probabilistically axiom-compliant confidence distribution for each clinical stage. The value of this probability distribution vector lies in the fact that it not only provides a clear staging prediction through the category indicated by its maximum probability value, but more importantly, it objectively and quantitatively demonstrates the uncertainty inherent in the prediction result itself and other staging tendencies that the sample may have through the distribution pattern of the entire vector. This provides clinical decision-making with much richer reference information than a single judgment conclusion, and directly solves the key defects of the existing technology in staging judgment that are subjective, rigid and unable to quantify the possibility.
[0058] At the conceptual level, this invention provides a method for automatically extracting probability vectors for kidney disease staging from urine protein test data. Its core mechanism lies in constructing a standardized processing and analysis chain from raw biological samples to quantitative decision-making information. This method, by integrating highly sensitive label-free detection technology, a standardized data reconstruction process, and an intelligent feature learning model, achieves objective, accurate, and information-rich probabilistic predictions of kidney disease and its staging.
[0059] The entire process begins with the acquisition of high-quality data from urine protein samples. Using gel electrophoresis, the protein components in the urine are precisely separated according to their molecular weight under an electric field, forming bands arranged in molecular weight order. These bands are then directly detected using real-time intrinsic fluorescence imaging. The principle is that the aromatic amino acids endogenously present in proteins fluoresce under ultraviolet light of a specific wavelength, and the fluorescence intensity is proportional to the protein content. By configuring an imaging system with a specific excitation source and filters, these fluorescence signals can be captured non-destructively and in real-time during electrophoresis, and the pixel grayscale values corresponding to each molecular weight band can be automatically analyzed.
[0060] After acquiring grayscale data, it needs to be converted into quantitative information with clear physiological significance. To this end, the method pre-establishes a calibration curve, or standard curve, from grayscale values to absolute concentrations by detecting a series of urinary protein biomarker standards with known concentrations. By substituting the grayscale values of the sample into the mathematical formula corresponding to this curve, the precise absolute quantitative concentration of each urinary protein component can be calculated. Thus, the compositional information of urinary proteins is transformed from analog image signals into a series of structured digital concentration values.
[0061] However, isolated concentration values are insufficient to reflect the overall distribution pattern of urinary proteins. Therefore, the next crucial step is to systematically integrate the quantitative concentration data of all components. Specifically, by using the molecular weight of urinary protein molecules as the x-axis and their corresponding absolute concentrations as the y-axis, the molecular weight concentration pairs of all components are reconstructed to form a continuous or discrete digital distribution map. This map comprehensively reflects the complete composition and abundance relationship of proteins of different molecular weights in urine, and is a key step in transforming complex biological samples into a standardized digital format that can be recognized and learned by deep learning models.
[0062] Subsequently, the digitized atlas was fed into a pre-trained convolutional neural network model for analysis. This neural network model had been trained on a large amount of urinary protein atlas data annotated with clinical staging results. Its multi-level convolutional and pooling structures can automatically learn and extract deep, abstract feature patterns from the atlas. The model performs forward propagation calculations on the input atlas, ultimately outputting a high-dimensional feature vector at a specific network layer. This vector is a highly condensed representation of the original digitized atlas, with a dimension far exceeding the original input, containing complex feature information highly correlated with the location and severity of kidney damage. These features often exceed the scope that can be discovered through manual interpretation.
[0063] Finally, the aforementioned high-dimensional feature vectors are input into a classifier for decision inference. The classifier typically employs a multi-class logistic regression model, such as the Softmax function, to interpret the information carried by the feature vectors and calculate the probability that the current sample belongs to each preset kidney disease stage. The classifier's output is a strict probability distribution vector, where each element represents the predicted probability of the sample belonging to the corresponding stage, and the sum of all elements is 1. This probability distribution vector is the final output of this method. It not only indicates the most likely stage prediction but, more importantly, quantitatively presents the probability distribution of all possible stages, thus clearly revealing the confidence level and uncertainty of the model's predictions. This provides clinicians with much richer auxiliary decision-making information than a single judgment.
[0064] In summary, this method obtains objective data from label-free real-time fluorescence imaging, achieves precise quantification through standard curves, and integrates the data into a standardized digital atlas to unify the data format. Then, leveraging the powerful pattern recognition capabilities of convolutional neural networks, it automatically extracts deep biomarker features from the atlas, and finally, a classifier transforms these features into intuitive and information-rich probabilistic staging predictions. This entire process achieves end-to-end automated analysis from urine samples to staging probability vectors, effectively solving the core problems of traditional methods such as high subjectivity, low automation, and lack of quantitative uncertainty output.
[0065] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. Component models, material names, connection structures, circuit structures, control methods, algorithms, and other features not explicitly described in this technical solution are considered common technical features disclosed in the prior art.
[0066] In this invention, the training effect is represented by a loss function (eq. S1). The meaning of loss is the sample... The predictability of the class ( ) and real labels ( Differences between ) j The performance was evaluated using accuracy (eq. S2), precision (eq. S3), recall (eq. S4), and F1 score (eq. S5). Four performance metrics were used: true negative (TN), true positive (TP), false negative (FN), and false positive (FP). Higher predictive accuracy, sensitivity, and specificity values indicate better model performance.
[0067] (S1)
[0068] (S2)
[0069] (S3)
[0070] (S4)
[0071] (S5)
[0072] Example 1
[0073] The purpose of this invention is to provide a new method for the separation and detection of urine protein based on SDS-GE-IFI in order to solve the above-mentioned problems, so as to achieve simple, rapid, sensitive, accurate and efficient separation and detection of urine protein and protein fingerprinting. It is a further technological development and specific application of the previous electrophoresis technology (CN202211345703.9 and CN202010596037.0), which will not be described in detail here.
[0074] Based on previous technologies, this embodiment first establishes a simple, rapid, sensitive, accurate, efficient, and low-cost qualitative and quantitative analysis technique for urinary protein to detect urinary protein markers including but not limited to IgG, TRF, HSA, and α1-MG, facilitating the analysis of renal tubular and glomerular diseases. Then, based on efficient protein fingerprinting (DL) analysis, the severity of kidney disease can be rapidly analyzed. Clinical comparative data show that this invention can not only be easily used for histological localization of kidney diseases, but also accurately assist in the analysis of the severity and stage of kidney disease.
[0075] See Figure 1 , Figure 2 , Figure 3 and Figure 4 The present invention is achieved through the following scheme:
[0076] S1. Using major urinary protein biomarkers for nephropathy, including but not limited to human IgG (150 kDa), transferrin (TRF, 78 kDa), HSA (66.5 kDa), and human α1-microglobulin (α1-MG, 33 kDa), a standard curve for protein biomarkers was constructed: solutions of standard urinary protein at concentrations including but not limited to IgG, TRF, HSA, and α1-MG were prepared at concentrations of 1, 2, 5, 10, 20, 50, 100, and 200 mg / mL. The mixed solutions were separated and detected using SDS-GE-IFI technology. The IFI imaging images and peak diagrams were automatically obtained and analyzed by the analysis software, and a standard curve was plotted with urinary protein concentration as the x-axis and urinary protein band gray value as the y-axis.
[0077] The aforementioned supporting software is existing commercially mature software, and its specific operation process is as follows: First, it receives raw images of endogenous fluorescence gel electrophoresis from a CMOS camera in real time or offline, and preprocesses the images through background subtraction and noise filtering algorithms; then, based on the preset molecular weight standard (Marker) lanes, it automatically identifies and locates the center positions of protein bands with different migration distances in each sample lane on the gel, and defines the boundaries of each band through band detection and segmentation algorithms (such as peak recognition based on vertical density projection); next, the software calculates the average pixel gray value in each band region, and automatically converts the gray value into the absolute quantitative concentration of the corresponding component using a preset concentration-gray value calibration equation established by a known concentration standard;
[0078] The software integrates the molecular weight (calculated from a migration distance-logarithmic standard curve) and absolute concentration data of all identified bands, reconstructs and outputs a structured data matrix or vector with molecular weight as the x-axis and concentration as the y-axis. This is the digital spectrum used as input to the convolutional neural network. The core logic of this software lies in achieving a fully automated and standardized conversion from raw fluorescence images to structured quantitative data, making it a key module connecting the front-end electrophoresis experiment with the back-end deep learning model.
[0079] S2. Urine samples and urine protein standards were directly treated 1:1 with loading buffer, and the urine protein samples to be tested were separated and detected using SDS-GE-IFI technology to obtain the corresponding gray values of urine protein bands.
[0080] S3. Based on the standard curve and gray value obtained in steps S1 and S2, the gray value is substituted into the standard curve to obtain the concentration of urinary protein components in the test solution, thus realizing the qualitative and quantitative analysis of the molecular weight of different components of the urinary protein sample.
[0081] S4. Based on the qualitative and quantitative analysis of the molecular weight of different components of urinary protein samples, the urinary protein component map is further input into the artificial intelligence DL model and the DL model is trained; then the trained DL model is used to test the clinical urinary protein component map to obtain clinical detection auxiliary information (feature information) for different severity stages of kidney disease with high accuracy.
[0082] The detection principle of this invention is as follows: Urinary protein components exhibit significant differences in molecular weight and abundance. SDS-GE technology is used to separate urinary proteins at different molecular weight positions. Based on 275 nm ultraviolet excitation, aromatic amino acids (tryptophan, tyrosine, and phenylalanine) produce intrinsic fluorescence. Under specific conditions, a CCD camera is used to image the focusing results, and the accompanying software performs chromatographic analysis to determine the corresponding band gray values and proportions. The band gray values are positively correlated with the substance concentration. Quantitative detection of urinary proteins is achieved by detecting the gray values of the urinary proteins.
[0083] In specific implementation, the urine protein quantification method based on SDS-GE-IFI mentioned above uses urine as the test sample, and the required sample volume is 5~50 μL.
[0084] In specific implementation, the aforementioned SDS-GE-IFI-based method for quantifying urinary protein includes an SDS-GE-IFI detection system comprising, but not limited to, an LED light source, a CMOS camera, a 360 nm filter, and a semi-open gel electrophoresis (GEA) apparatus. The UV LED light source is equipped with four LED panels for irradiating the gel in the gel mold. In each panel, 42 LEDs are arranged in a 3 × 14 array, producing uniform deep UV light at 275 nm. The CMOS camera has a quantum efficiency of 50% or higher for capturing UV fluorescence at approximately 360 nm. The GEA apparatus consists of a semi-open electrophoresis tank, a gel plate, an anolyte, and a catholyte.
[0085] In specific implementation, the urine protein detection method based on SDS-GE-IFI includes a sample loading solution comprising 5-50 μL of extracted urine, 5 mg / mL standard, 1-2 × sample buffer, 1 × electrophoresis buffer, 8-18% separating gel, and 5% stacking gel.
[0086] In practice, based on the obtained urine protein electrophoresis pattern and related quantitative data, the following steps are performed to generate a probability distribution vector:
[0087] Electrophoresis imaging results of all patient samples were analyzed using specialized software. Based on established standard curves, the grayscale values of each protein band identified in each lane were converted into corresponding absolute quantitative concentrations. Subsequently, a complete urinary protein concentration distribution curve was plotted for each sample, with the molecular weight (logarithm, log MW) of the urinary protein components on the x-axis and their absolute concentration on the y-axis. This curve was discretized to generate a fixed-length one-dimensional feature vector, or directly saved as a standard-sized two-dimensional grayscale image (where pixel position encodes molecular weight and pixel intensity encodes concentration). This step integrates the quantitative concentration data of all urinary protein components, generating a digital map characterizing the overall distribution pattern of urinary protein components with different molecular weights.
[0088] Training and high-dimensional feature vector extraction of convolutional neural network models:
[0089] a) Dataset preparation: The generated digital atlases are paired with the clinically diagnosed kidney disease stage labels to form the model training set. The sets are then randomly divided into training, validation, and test sets according to a predetermined ratio.
[0090] b) Model Architecture and Training: The VGG19 network architecture is used as the base model. A digitized map (e.g., in image format) is taken as input. The original classification head is removed, and the convolutional bases are retained for feature extraction. After this, a global average pooling layer is added, followed by a fully connected layer containing multiple neurons. The output of this layer is defined as a "high-dimensional feature vector representing the compositional characteristics of urinary proteins."
[0091] c) Model Training: For the classification task, a new fully connected layer with the same number of neurons as the number of classification categories is added after the high-dimensional feature vectors as the classification head, using the Softmax activation function. The model is trained using the training set, with cross-entropy as the loss function and Adam as the optimizer. During training, performance is monitored using the validation set to prevent overfitting.
[0092] Generation of probability distribution vectors and model evaluation:
[0093] a) Probability Vector Generation: The digitized map of the test set samples is input into the trained model. The model performs forward propagation calculations: First, a 512-dimensional high-dimensional feature vector is extracted through convolutional base layers and global average pooling layers. Then, this feature vector is input into the final classification head (fully connected layer + Softmax). The output of the Softmax layer is a probability distribution vector, in the form P = [p1, p2, p3, p4, p5], where p i This represents the probability that the sample is predicted to be in the i-th period, and satisfies Σp i = 1.
[0094] b) Model performance evaluation: Based on the probability distribution vector output by the model, the stage corresponding to the highest probability value is taken as the final predicted stage, which is used to calculate the confusion matrix, accuracy, and AUC value. Simultaneously, the probability distribution vector itself is also provided to clinicians as reference information to quantify staging uncertainty and assist in decision-making.
[0095] Verification Example 1
[0096] Step 1. Urine collection and pretreatment
[0097] Take 20-50 μL of fresh urine, add it to a 1.5 mL centrifuge tube, add 10-50 μL of loading buffer, shake well, and place in a constant temperature metal bath at 95℃ for 5 min. It should be used for SDS-GE-IFI separation and detection experiments on the same day.
[0098] Step 2. Gel preparation and sample loading
[0099] (1) Prepare a separating gel according to the standard formula, including but not limited to 5-19% gel: 2.3~0.2 mL ultrapure water, 1.3~3.06 mL 30% polyacrylamide, 1.3 mL 1.5 M pH 8.8 Tris-HCl, 0.05 mL 10% SDS, 0.05 mL 10% ammonium persulfate, and 2 μL TEMED. Add the above solutions sequentially to a glass beaker, mix well, and then pour into a prepared protein gel glass plate. Finally, add anhydrous ethanol and press flat. After the separating gel solidifies, discard the upper layer of ethanol. Prepare a stacking gel according to the following formula, including but not limited to 5% gel: 1.4 mL ultrapure water, 0.33 mL 30% polyacrylamide, 0.25 mL 1.0 M pH 6.8 Tris-Gly, 0.02 mL 10% SDS, 0.02 mL 10% ammonium persulfate, and 2 μL TEMED. After pouring, add a comb and allow to solidify at room temperature.
[0100] (2) Add the mixed sample solution evenly into the gel sample wells, and then add the electrophoresis buffer. Place the GEA into the GE-IFI electrophoresis apparatus.
[0101] Step 3. Electrophoresis process
[0102] (1) After placing the GEA into the electrophoresis apparatus, set the operating parameters: 150 V for 10 min, 200 V for 30 min;
[0103] (2) During electrophoresis, protein migration can be monitored in real time and photographed at any time (usually automatically every 10-30 minutes).
[0104] in Figure 3 Image A shows the electrophoresis diagram of a molecular weight standard (M) and a mixed sample of urine tetraproteins (Mix). Figure 3 Figure B shows the results of a series of diluted standard urinary protein samples to test the lowest detection sensitivity.
[0105] Imaging parameters include 1 s exposure time and gain 2 (see...) Figure 3 (A). Serial dilution of standard urinary protein samples yielded the lowest detection sensitivity (see [reference]). Figure 3 (B) In this invention, gain 2 refers to the amplification factor of the camera sensor signal being 2.
[0106] Step 4. Calculation of urine protein concentration
[0107] Based on the SDS-GE-IFI chromatogram obtained from the analysis software, the gray values of the urinary protein bands in the test samples can be calculated. A standard curve is then plotted with the gray values of the standard urinary protein bands as the ordinate and the urinary protein concentration as the abscissa (refer to...). Figure 4 ).
[0108] Substituting the gray values of the sample bands into the standard curve, the concentration of the urinary protein in the injection solution can be determined. Figure 5A This is an SDS-GE-IFI electrophoresis image of a urine sample from a healthy person. Figure 5B SDS-GE-IFI electrophoresis images of urine samples from diabetic patients and patients with kidney disease at different stages.
[0109] Step 5. Comparison of urinary protein prediction and clinical staging in patients with different degrees of nephropathy based on SDS-GE-IFI.
[0110] SDS-GE-IFI was used to detect proteins in urine samples from patients with different kidney diseases, and electrophoretic patterns were obtained. Figure 5A and Figure 5B ) and related chromatograms ( Figure 6 and Figure 7 Therefore, utilizing Figure 5A and Figure 5B SDS-GE-IFI electrophoresis patterns, combined with Figure 4 A standard curve can be used to determine the molecular weight, content, and relative content of the main markers of urinary protein.
[0111] Figure 8 In the middle A, the SDS-GE-IFI electrophoresis patterns of urine samples from patients with different stages of nephropathy are compared. Figure 8 B represents a comparison of the chromatographic peaks converted from the above electrophoretic chromatograms; Figure 8 The comparison between the predicted results of SDS-GE-IFI analysis of urine samples from 100 patients with different stages of nephropathy and the actual clinical results is shown in the middle C.
[0112] If SDS-GE-IFI is used alone to analyze the prediction of urinary protein (especially HSA content) in patients with different stages of nephropathy, the difference from the actual clinical results is large, reaching up to 9% (see [link to relevant documentation]). Figure 8 (C). Therefore, the error is relatively large when using SDS-GE-IFI data alone to analyze the classification and staging of kidney disease.
[0113] Step 6. Artificial intelligence DL is introduced to predict urinary protein in patients with different stages of nephropathy using SDS-GE-IFI.
[0114] Based on the urine protein electrophoresis pattern and related quantitative data obtained in step 5, the following sub-steps are performed to generate a probability distribution vector:
[0115] Step 6.1. Digital map generation:
[0116] Will Figure 5BElectrophoresis imaging results of all patient samples were analyzed using the accompanying software. Based on the standard curve established in step 2 (or S2), the grayscale values of each protein band identified in each lane were converted into corresponding absolute quantitative concentrations. Subsequently, a complete urinary protein concentration distribution curve was plotted for each sample, with the molecular weight (logarithm, log MW) of the urinary protein components as the x-axis and their absolute concentration as the y-axis. This curve was discretized to generate a fixed-length one-dimensional feature vector, or it was directly saved as a standard-sized two-dimensional grayscale image (where pixel position encodes molecular weight and pixel intensity encodes concentration). This step integrates the quantitative concentration data of all urinary protein components, generating a digital map characterizing the overall distribution pattern of urinary protein components with different molecular weights.
[0117] Step 6.2. Training the Convolutional Neural Network Model and Extracting High-Dimensional Feature Vectors:
[0118] a) Dataset preparation: The generated digital atlases are paired with the clinically diagnosed kidney disease stage labels to form the model training set. This set is then randomly divided into training, validation, and test sets in a 7:2:1 ratio.
[0119] b) Model Architecture and Training: The VGG19 network architecture is used as the base model. A digitized map (e.g., in image format) is taken as input. The original classification head is removed, and the convolutional bases are retained for feature extraction. After this, a global average pooling layer is added, followed by a fully connected layer with 512 neurons. The output of this layer is defined as a "high-dimensional feature vector representing the compositional characteristics of urinary proteins".
[0120] c) Model Training: For the classification task, a new fully connected layer with the same number of neurons as the number of classification categories is added after the high-dimensional feature vectors as the classification head, using the Softmax activation function. The model is trained using the training set, with cross-entropy as the loss function and Adam as the optimizer. During training, performance is monitored using the validation set to prevent overfitting.
[0121] Step 6.3. Generation of probability distribution vectors and model evaluation:
[0122] a) Probability Vector Generation: The digitized map of the test set samples is input into the trained model. The model performs forward propagation calculations: First, a 512-dimensional high-dimensional feature vector is extracted through convolutional base layers and global average pooling layers. Then, this feature vector is input into the final classification head (fully connected layer + Softmax). The output of the Softmax layer is a probability distribution vector, in the form P = [p1, p2, p3, p4, p5], where p i This represents the probability that the sample is predicted to be in the i-th period, and satisfies Σp i= 1. For example, if the model output for a certain sample is [0.02, 0.08, 0.85, 0.05, 0.00], it means that the probability of it belonging to period 3 is the highest (85%).
[0123] b) Model performance evaluation: Based on the probability distribution vector output by the model, the period corresponding to the maximum probability value is taken as the final predicted period, which is used to calculate... Figure 10 The confusion matrix, accuracy (95.95%), and AUC (0.96) are shown. Simultaneously, the probability distribution vector itself is also provided to clinicians as reference information for quantifying staging uncertainty and assisting decision-making.
[0124] This validation example systematically evaluates the feature learning capabilities of seven convolutional neural networks in a multi-classification task of urine pathology images. The experimental design employs the classic three-point method to divide the dataset: optimizing model parameters on an independent training set, adjusting hyperparameters on a validation set, and determining the network structure. Finally, the model's generalization ability is validated on an independent test set. During model training, the loss function (eq S1) and accuracy (eq S2) are dynamically monitored to monitor model convergence, ensuring stable generalization performance on the independent test set. Figure 9 As shown, VGG19, VGG16, and ResNet50 exhibit significant advantages in learning efficiency, with their loss functions converging quickly and accuracy improving rapidly. Notably, the performance differences between the models gradually narrow when the training epochs reach 50. Finally, all models reach a convergence state where the loss value approaches 0 and the training accuracy approaches 1, indicating the effectiveness of the training process.
[0125] Figure 10 Figure A shows a comparison of the area under the receiver operating characteristic (AUC) of this convolutional neural network model with that of traditional detection methods; Figure 10 In the diagram, B represents the confusion matrix of the model's predictions on the test set. The scale bar indicates the number of SDS-GE-IFI classification patterns, ranging from 0 to 49.
[0126] In the model performance evaluation phase, three core metrics—precision (eq S3), recall (eq S4), and F1 score (eq S5)—were calculated using the validation set. As shown in Table 1, the seven models exhibited significant differences across these three metrics. Specifically, the Vovnet39 model showed significantly lower classification efficiency than the other models (average values for all metrics < 0.79), while VGG19 demonstrated the best performance, with all three metrics exceeding 0.97. Based on the stable performance and clinical consistency of the VGG19 model in urine sample feature extraction, it is more suitable for urine sample identification and classification applications (see Table 1). Figure 10 ).
[0127] Specifically, to further validate the clinical applicability of the selected model, the area under the receiver operating characteristic curve (AUC-ROC) index was used to evaluate the VGG19 test set (see [link to relevant documentation]). Figure 10 (See Figure A). Experimental results show that the model's AUC value reaches over 0.96, approaching the performance of a perfect classifier (AUC = 1), and significantly outperforming the random guessing baseline (AUC = 0.5), confirming its excellent sample discrimination ability. Furthermore, VGG19 is further evaluated using the confusion matrix (see Figure A). Figure 10 (B) The confusion matrix shows the classification ability of VGG19 for three different grades of urine samples. The results show that VGG19 can accurately predict most urine samples in the test set, with a classification accuracy of 95.95%. In summary, a robust disease screening workflow based on GE-IFI electrophoresis and the VGG19 model was established. This process optimizes traditional screening methods and improves the accuracy of disease screening.
[0128] To address the inherent black-box nature of deep learning models, a Gradient Weighted Class Activation Mapping (Grad-CAM) technique is employed to construct an interpretability analysis framework. This technique extracts gradient information of the target class from the final convolutional layer, generating high-resolution feature activation heatmaps to quantitatively reveal the spatial weight distribution of key biomarkers in the neural network decision-making process.
[0129] Figure 11 China A and Figure 11 In the image, B represents the SDS-GE-IFI electrophoresis images of two different urine samples (urine sample A and urine sample B) and their corresponding feature activation heatmap distributions generated by the model. For example... Figure 11 As shown, the activation heatmap visually presents the differences in feature importance through color coding: the red highlighted areas correspond to the core recognition regions of the electrophoretic bands, highly consistent with the typical high-molecular-weight urinary protein bands in clinical analysis; the blue areas are distributed in the electrophoretic background noise region, indicating that the model effectively suppresses non-specific interference. Notably, the model focuses on medium-molecular-weight protein bands (the region where albumin is located), exhibiting a precise local feature capture pattern, rather than a global feature-dependent pattern with lower clinical value. This feature selection specificity is highly consistent with the clinical nephrology analysis pathway, indicating that the model's predictions conform to basic logic and further verifying the model's reliability.
[0130] The AI-based SDS-GE-IFI assisted system developed in this technical solution has four distinct features (Table 2). First, the method is extremely simple. Urine samples can be directly loaded onto the electrophoresis apparatus without complex pretreatment processes such as desalting and protein condensation. After electrophoresis, the staining step is eliminated, saving detection time and complex manual staining and destaining operations. Furthermore, DL simplifies the processes of image recognition, urine protein quantification, and threshold determination, making disease analysis more convenient.
[0131] Secondly, this method is highly efficient. It is a deep learning-based SDS-GE-IFI method capable of classifying 742 images in a dataset in approximately 10 ms, while traditional SDS-GE-IFI analysis software takes 30 minutes or more to several hours to process. Furthermore, this method significantly improves the efficiency of GE-IFI analysis.
[0132] Third, this method has high accuracy in staging kidney disease. The accuracy of DL-based GE-IFI detection has increased from 91% for GE-IFI alone to 95.95%, which is comparable to the accuracy of clinicians combining patient medical history information with multiple detection modalities. Therefore, this method is highly accurate.
[0133] The above description of the embodiments is provided to enable those skilled in the art to understand and use the invention. It will be apparent to those skilled in the art that various modifications can be made to these embodiments, and the general principles described herein can be applied to other embodiments without inventive effort. Therefore, the present invention is not limited to the above embodiments, and any improvements and modifications made by those skilled in the art based on the disclosure of the present invention without departing from the scope of the invention should be within the protection scope of the present invention.
[0134] Table 1. Performance metrics (precision, recall, and F1 score) of seven models across three CNN architectures on the validation dataset.
[0135]
[0136] Table 2. Comprehensive comparison of GE-IFI with existing methods.
[0137]
[0138] a) N / A – Data missing. b) SPR: Surface Plasmon Resonance. c) GE: Polyacrylamide gel electrophoresis.
[0139] The references mentioned in Table 2 are as follows:
[0140] [1] Y. Jia, G. Liu, G. Xu, X. Li, Z. Shi, C. Cheng, D. Xu, Y. Lu, Q.Liu, Battery-free and wireless tag for in situ sensing of urinary albumin / creatinine ratio (ACR) for the assessment of albuminuria, Sensors andActuators B: Chemical 367 (2022).
[0141] [2] L. Zhang, H. Wang, H. Zhang, N. Zhang, X. Zheng, W. Li, X. Qiu,D. Yu, Development of a portable multiplexed instrument for multi-proteinsdetection in human urine using surface plasmon resonance, Sensors andActuators B: Chemical 369 (2022).
[0142] [3] M. Salinas, M. López-Garrigós, E. Flores, J. Lugo, C. Leiva-Salinas, Urinary albumin strip assay as a screening test to replacequantitative technology in certain conditions, Clinical Chemistry andLaboratory Medicine (CCLM) 57(2) (2018) 204-209.
[0143] [4] Z. Yu, Y. Cao, Y. Tian, W. Ji, K.-E. Chen, Z. Wang, J. Ren, H. Xiao, L. Zhang, W. Liu, L. Fan, Q. Zhang, C. Cao, Real-time and quantitative protein detection via polyacrylamide gel electrophoresis and online intrinsicfluorescence imaging, Analytica Chimica Acta 1291 (2024).
[0144] The above description of the embodiments is provided to enable those skilled in the art to understand and use the invention. It will be apparent to those skilled in the art that various modifications can be made to these embodiments, and the general principles described herein can be applied to other embodiments without inventive effort. Therefore, the present invention is not limited to the above embodiments, and any improvements and modifications made by those skilled in the art based on the disclosure of the present invention without departing from the scope of the invention should be within the protection scope of the present invention.
Claims
1. A method for extracting the probability distribution vector of nephropathy and its stage based on real-time endogenous fluorescence imaging data from urinary protein gel electrophoresis, characterized in that, Includes the following steps: S1. Obtain the molecular weight bands and their gray values of urinary proteins extracted by real-time intrinsic fluorescence imaging technology of urinary protein gel electrophoresis. S2. Based on the pre-established standard curve of concentration-grayscale value of urinary protein biomarker standard, the grayscale value data is converted into the absolute quantitative concentration data of the corresponding urinary protein component; S3. Integrate the quantitative concentration data of all urinary protein components to generate a digital map characterizing the overall distribution pattern of urinary protein components of different molecular weights; S4. Input the digitized map into a pre-trained convolutional neural network model, and calculate and output a high-dimensional feature vector representing the composition characteristics of urinary protein through the forward propagation of the model. S5. Input the high-dimensional feature vector into the classifier to calculate the probability distribution vector associated with different kidney disease stages. The probability distribution vector is used as kidney disease stage prediction information.
2. The method for extracting the probability distribution vector of nephropathy and its stage based on real-time endogenous fluorescence imaging data of urinary protein gel electrophoresis according to claim 1, characterized in that, In S1, the specific process of obtaining the grayscale data of each molecular weight band of urinary protein extracted by gel electrophoresis intrinsic fluorescence imaging technology includes: Using a CMOS camera equipped with a 275 nm ultraviolet LED excitation source and a 360 nm filter, real-time intrinsic fluorescence images were acquired during gel electrophoresis. The molecular weight bands of each urinary protein were automatically identified by the accompanying image analysis software, and their corresponding pixel gray values were calculated.
3. The method for extracting the probability distribution vector of nephropathy and its stage based on real-time endogenous fluorescence imaging data of urinary protein gel electrophoresis according to claim 1, characterized in that, In S2, the standard curve of concentration-gray value of the urine protein marker standard is established by detecting standard solutions including immunoglobulin IgG, human transferrin TRF, human serum albumin HSA and α1-microglobulin α1-MG. That is, the calibration equation is obtained by linear or nonlinear fitting with concentration as the abscissa and the corresponding band gray value as the ordinate. In the process of converting the grayscale data into the absolute quantitative concentration data of the corresponding urinary protein components, the grayscale value of the sample to be tested is substituted into the calibration equation to calculate the absolute concentration of each urinary protein component.
4. The method for extracting the probability distribution vector of nephropathy and its stage based on real-time endogenous fluorescence imaging data of urinary protein gel electrophoresis according to claim 1, characterized in that, In S3, the specific process of integrating the quantitative concentration data of all components to generate a digital map characterizing the overall distribution pattern of urinary protein components of different molecular weights includes: Using the molecular weight of urinary proteins as the x-axis and the relative or absolute concentration as the y-axis, the concentration data of each component are reconstructed into a structured data matrix of urinary protein molecular weight identifiers and concentration values.
5. The method for extracting the probability distribution vector of nephropathy and its stage based on real-time endogenous fluorescence imaging data of urinary protein gel electrophoresis according to claim 1, characterized in that, In S4, the convolutional neural network model is trained by the following process: A training set was constructed by collecting urinary protein digitization maps of clinically diagnosed samples containing various kidney diseases and their stages; The training set is divided into a training subset, a validation subset, and a test subset; The VGG19 network architecture was used as the convolutional neural network model, with the digitized map as input and the corresponding kidney disease stage label as the supervision signal. The model parameters are dynamically adjusted by optimizing the loss function, and the hyperparameters are tuned using a validation subset to determine the network structure. The model's generalization ability and accuracy are validated on a subset of tests until the model converges, resulting in a pre-trained model capable of extracting high-dimensional feature vectors.
6. The method for extracting the probability distribution vector of nephropathy and its stage based on real-time endogenous fluorescence imaging data of urinary protein gel electrophoresis according to claim 1, characterized in that, In S5, the probability distribution vectors associated with different stages of kidney disease are calculated. The specific process includes: After receiving the high-dimensional feature vector, the classifier calculates the probability value of the sample belonging to each stage of kidney disease through multi-class logistic regression. Arrange all probability values in order of stage category to form a probability distribution vector, where the sum of all elements in the vector is 1, and the stage of kidney disease corresponding to the highest probability value is the main stage tendency predicted by the model.
7. The method for extracting the probability distribution vector of nephropathy and its stage based on real-time endogenous fluorescence imaging data of urinary protein gel electrophoresis according to claim 1, characterized in that, In S5, the classifier is the Softmax function.
8. The method for extracting the probability distribution vector of nephropathy and its stage based on real-time endogenous fluorescence imaging data of urinary protein gel electrophoresis according to claim 1, characterized in that, In S4, the high-dimensional feature vector specifically includes the following feature parameters: a) Band distribution characteristics of high molecular weight region, medium molecular weight region and low molecular weight region according to molecular weight; b) Concentration ratio characteristics of IgG, transferrin, albumin and α1-microglobulin; c) Spatial weight features of albumin strips generated based on gradient-weighted class activation mapping.