Mama-based membranous nephropathy multi-mode pathological image quantitative analysis system and Mama-based membranous nephropathy multi-mode pathological image quantitative analysis method
By using a Mamba-based multimodal pathological image quantitative analysis system, combined with weakly supervised learning and iterative optimization strategies, the problems of strong subjectivity and insufficient quantification of lesion features in the diagnosis of membranous nephropathy have been solved. This system achieves multimodal fusion analysis and efficient quantification, improving diagnostic accuracy and automation efficiency, and providing precise quantitative indicators and interpretable reports.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TAIYUAN UNIVERSITY OF TECHNOLOGY
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-15
AI Technical Summary
Existing diagnostic systems for membranous nephropathy suffer from problems such as strong subjectivity in the diagnostic process, limited single-modal analysis, insufficient quantification of lesion characteristics, and complex model calculations. These limitations make it difficult to achieve multimodal fusion and efficient quantitative analysis, resulting in limited diagnostic accuracy and clinical applicability.
A multimodal pathological image quantitative analysis system based on Mamba is adopted, including image preprocessing, macroscopic lesion analysis, microscopic structure analysis, thickness quantification, and feature fusion and prediction modules. Combining weakly supervised multi-instance learning and iterative optimization strategies, the system achieves accurate segmentation and thickness quantification of the glomerular basement membrane through a state-space model, and uses gradient boosting decision trees to predict disease staging, providing interpretable analysis reports.
It enables multimodal fusion analysis, improves the comprehensiveness and accuracy of diagnosis, enhances segmentation precision, addresses the problem of scarce clinical data, automates the entire process, and enhances clinical trust through interpretable reports, providing precise quantitative evidence.
Smart Images

Figure CN122048837A_ABST
Abstract
Description
Technical Field
[0001] This invention provides a quantitative analysis system and method for multimodal pathological images of membranous nephropathy based on Mamba, belonging to the field of computer vision and image segmentation technology. Background Technology
[0002] Membranous nephropathy (MN) is one of the main types of chronic kidney disease. Its diagnosis mainly relies on the observation of spike-like protrusions under light microscopy (LM) and the measurement of glomerular basement membrane (GBM) thickness under electron microscopy (EM). Current diagnostic methods have the following prominent shortcomings: First, they rely on qualitative assessment by pathologists, which leads to large individual differences in judgment and poor reproducibility. Second, the spike-like protrusions in LM images are diffusely distributed and have fine structures, making manual annotation time-consuming, laborious, and difficult to achieve quantitative analysis. Third, the GBM boundaries in electron microscopy images are blurred and curved, making it easy to miss slight thickenings in the 100–400 nm range, resulting in insufficient accuracy in thickness measurement. Fourth, existing automated analysis methods are mostly limited to single-modal studies, which are out of sync with the dual-modal diagnostic process of "LM initial screening + TEM confirmation" in clinical practice. Fifth, mainstream CNN models are difficult to capture long-range contextual information, while Transformer models are difficult to apply directly to high-resolution pathological slide images due to their high secondary computational complexity. Existing technologies such as DeepLab-v3 and RADS-Net have not yet effectively solved the core challenges of multimodal fusion and efficient quantitative analysis, thus their diagnostic accuracy and clinical applicability remain significantly limited. Summary of the Invention
[0003] To address the technical problems of existing membranous nephropathy diagnostic systems, such as strong subjectivity in the diagnostic process, limited single-modal analysis, insufficient quantification of lesion features, and complex model calculations, this invention proposes a Mamba-based multimodal pathological image quantitative analysis system and method for membranous nephropathy.
[0004] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: a quantitative analysis system for multimodal pathological images of membranous nephropathy based on Mamba, comprising: an image preprocessing module for receiving and standardizing the pathological images of patients, the pathological images including a set of light microscopic full-field slice images containing glomeruli and electron microscopic images containing glomeruli;
[0005] The macroscopic lesion analysis module integrates the first analysis model, which is used to process the preprocessed light microscopy full-field slice image set to realize weakly supervised detection and quantification of nail-like lesions in the glomerulus, and output nail-like quantitative indicators and spatial location information of nail-like lesions.
[0006] The microstructure analysis module integrates a second analysis model to process preprocessed electron microscopy images, achieve precise segmentation of the glomerular basement membrane, and output a segmentation mask for the glomerular basement membrane.
[0007] The thickness quantification module is used to extract the skeleton and perform geometric measurements on the segmented glomerular basement membrane region, and output the glomerular basement membrane thickness quantification index by combining the scale bar information in the electron microscopy image.
[0008] The feature fusion and prediction module receives the pin-like quantification index from the macroscopic lesion analysis module and the glomerular basement membrane thickness quantification index from the thickness quantification module. It then fuses the pin-like quantification index and the glomerular basement membrane thickness quantification index into a multidimensional feature vector. Based on the multidimensional feature vector, it generates disease staging prediction results and interpretability analysis reports.
[0009] The results visualization module is used to graphically present the spatial location information of nail-like lesions, the segmentation mask of the glomerular basement membrane, the quantitative indicators of nail-like lesions, the quantitative indicators of glomerular basement membrane thickness, the prediction results of disease staging, and the interpretability analysis report.
[0010] Furthermore, the first analysis model is a weakly supervised multi-instance learning model constructed based on a state-space model. In the first analysis model, a single glomerular image in the light microscopy full-field slice image set is regarded as a package, and the image blocks divided into glomerular images are regarded as instances. By aggregating the prediction results of the image blocks, the quantification index of the glomerular nail-like lesion and the spatial location information of the nail-like lesion are obtained.
[0011] Furthermore, the first analysis model employs an iterative optimization strategy during training: In the initial stage of training, the first analysis model assigns the initial label of all instances to the label of their respective bags to perform weakly supervised training on the image patches. The initial label of all instances is assigned to the label of their respective bags, i.e., patient-level labels. During training, high-confidence positive sample image patches and high-confidence negative sample image patches are periodically selected based on the confidence level predicted by the first analysis model. The selected high-confidence positive sample image patches and the high-confidence negative sample image patches are then used to update the training set to reduce the noise brought by the patient-level labels. This iterative process is repeated until the performance of the first analysis model converges.
[0012] Furthermore, the second analysis model adopts an encoder-decoder structure, in which both the encoder and decoder integrate a feature extraction module based on a state-space model, which has the ability to perform local scanning and cross-window interaction.
[0013] Each stage of the encoder and decoder includes an improved state space module; the improved state space module sequentially performs local window scanning and cross-window scanning after window shifting, and fuses the features obtained from local window scanning and cross-window scanning after window shifting through residual connection to output a segmentation mask of the glomerular basement membrane.
[0014] Furthermore, the thickness quantization module includes:
[0015] The image post-processing unit is used to smooth and fill holes in the segmentation mask of the glomerular basement membrane.
[0016] The skeleton extraction unit is used to extract a single-pixel-width skeleton line from the segmentation mask after smoothing and hole filling.
[0017] The thickness quantification index aggregation unit is used to convert the pixel-level glomerular basement membrane thickness quantification index value into actual physical thickness units based on scale information, and output the glomerular basement membrane thickness quantification index.
[0018] Furthermore, the feature fusion and prediction module includes:
[0019] The feature fusion unit is used to perform factor analysis on the input pin quantification index and glomerular basement membrane thickness quantification index, extract uncorrelated key feature factors, and construct a multidimensional feature vector characterizing the severity of the patient's condition.
[0020] The classification and prediction unit uses a gradient boosting decision tree model to predict the stage of membranous nephropathy based on the multidimensional feature vectors, and outputs accurate stage prediction results.
[0021] The interpretability analysis unit, based on the SHAP value analysis method, quantifies the contribution of each key feature factor to the disease staging prediction results, clarifies the influence weight of each indicator on the disease staging, and generates an interpretability analysis report.
[0022] A quantitative analysis method for multimodal pathological images of membranous nephropathy based on Mamba, applied to the aforementioned system, includes the following steps:
[0023] Step S1: Acquire and standardize the pathological images of the patient to be analyzed. The pathological images include a set of light microscopic full-field section images containing glomeruli and electron microscopic images containing glomeruli, and identify scale information from the electron microscopic images.
[0024] Step S2: Process the standardized light microscopy full-field slice image set through the first analysis model to realize weakly supervised detection and quantification of nail-like lesions in the glomerulus, and output nail-like quantitative indicators and spatial location information of nail-like lesions;
[0025] Step S3: Process the standardized electron microscopy image through the second analysis model and output the segmentation mask of the glomerular basement membrane to segment out the glomerular basement membrane region;
[0026] Step S4: Based on the segmentation mask of the glomerular basement membrane, perform skeleton extraction and geometric measurement on the segmented glomerular basement membrane region, and calculate the quantitative index of glomerular basement membrane thickness by combining the scale bar information in the electron microscopy image.
[0027] Step S5: Integrate the quantitative index of the nail process with the quantitative index of the glomerular basement membrane thickness, and perform factor analysis on the quantitative index of the nail process and the index of the basement membrane thickness to extract key feature factors;
[0028] Step S6: Generate disease staging prediction results and interpretability analysis report based on key feature factors.
[0029] Furthermore, in step S2, the first analysis model takes image blocks divided from glomerular images in the light microscopy full-field slice image set as input, outputs the confidence level that each image block contains a nail-like lesion, and locates the spatial position of the nail-like lesion based on this confidence level and calculates the nail-like quantification index.
[0030] Furthermore, in step S3, the second analysis model uses a feature extraction module integrated in its encoder-decoder structure, which has the ability to perform local scanning and cross-window interaction, to extract and reconstruct features from the electron microscope image, thereby achieving accurate segmentation of the irregular and blurred-boundary basement membrane.
[0031] Furthermore, step S4 involves calculating the quantitative index of glomerular basement membrane thickness, including the following process:
[0032] The actual length in the scale information of the electron microscope image is identified by optical character recognition technology, and the conversion ratio between pixels and actual length is calculated. The geometric measurement value of basement membrane thickness calculated based on pixel units is converted into the actual physical thickness in nanometers according to the conversion ratio, and the actual thickness of glomerular basement membrane is output. The mean, standardization and maximum value extraction of the actual thickness of multiple glomerular basement membranes of the same patient are performed to obtain the quantitative index of glomerular basement membrane thickness.
[0033] The advantages of this invention over the prior art are as follows:
[0034] 1. Multimodal fusion analysis: Combining light microscopy full-field section images of glomeruli with electron microscopy images, the condition of membranous nephropathy is characterized from two dimensions: macroscopic nail lesions and microscopic basement membrane structure, thereby improving the comprehensiveness and accuracy of diagnosis;
[0035] 2. High segmentation accuracy: The second analysis model achieves local scanning and cross-window interaction through a weakly supervised multi-instance learning model based on the state space model, which can accurately segment the irregular and blurred boundary glomerular basement membrane, laying the foundation for thickness quantification.
[0036] 3. Weakly supervised training adapted to clinical data: The first analysis model adopts a weakly supervised multi-instance learning and iterative optimization strategy, which does not require precise pixel-level labels and is adapted to the problem of scarce labeled data in clinical scenarios.
[0037] 4. Full-process automation and interpretability: It automates the entire process from image preprocessing to disease prediction, greatly improving analysis efficiency; at the same time, it provides interpretable reports through SHAP value analysis, enhancing clinical trust.
[0038] 5. Precise quantitative indicators: The thickness quantification module, combined with a scale bar, enables precise conversion from pixel units to physical units, providing accurate quantitative basis for clinical diagnosis and disease assessment. Attached Figure Description
[0039] The present invention will be further described below with reference to the accompanying drawings:
[0040] Figure 1 This is a schematic diagram of the system of the present invention;
[0041] Figure 2 This is a schematic diagram of the process of the present invention. Detailed Implementation
[0042] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicate relative orientations or positional relationships and are used only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this invention, unless otherwise stated, "a plurality of" means two or more.
[0043] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0044] like Figures 1 to 2As shown, this invention provides a Mamba-based multimodal pathological image quantitative analysis system for membranous nephropathy, comprising: an electron microscopy image acquisition device and a processor. The electron microscopy image acquisition device includes a light microscopy scanner and an electron microscope. The light microscopy scanner is a KF-PRO-005-EX digital slide scanner, equipped with a 40× objective lens and a 10× eyepiece (pixel resolution: 0.25μm / pixel), used to scan PASM-stained kidney tissue sections and generate a light microscopy full-field-of-view slide image set containing glomeruli. In this embodiment, the electron microscope is a transmission electron microscopy (TEM) imaging system, and the transmission electron microscope uses a JEM-1400FLASH system (120kV accelerating voltage) to capture ultrastructural images of the glomerular basement membrane (GBM), with a magnification range of 2500×-15000×, outputting electron microscopy images containing glomeruli.
[0045] The processor integrates an image preprocessing module, which is communicatively connected to a macroscopic lesion analysis module, a microscopic structure analysis module, and a thickness quantization module. Both the macroscopic lesion analysis module and the thickness quantization module are communicatively connected to a feature fusion and prediction module. All of these modules are communicatively connected to a result visualization module.
[0046] The image preprocessing module is used to receive and standardize the patient's pathological images, which include a set of light microscopic full-field section images containing glomeruli acquired by a light microscopy scanner and electron microscopic images containing glomeruli acquired by a transmission electron microscopy imaging system.
[0047] Specifically, the image preprocessing module performs the following operations:
[0048] Image standardization: The full-field-of-view slice image set of the optical microscope is uniformly scaled and color normalized, and grayscale correction is performed to eliminate staining differences. In this embodiment, it is uniformly scaled to 1024×1024 resolution; the electron microscope image is standardized to 512×512 resolution, and Gaussian noise filtering (σ=1.0) is used to remove electronic noise interference.
[0049] Glomerular segmentation: Multi-tissue segmentation algorithms such as MSMTSeg (Multi-staining Multi-tissue Segmentation Framework) are used to automatically detect and crop individual glomerular images from the full-field-of-view light microscope slice images.
[0050] Scale bar annotation: Mark the scale bar (e.g., "5μm" or "10μm") in a fixed corner area of the electron microscope image to provide a length reference for subsequent conversion of actual glomerular basement membrane thickness;
[0051] Data Augmentation and Splitting: Data augmentation operations such as random flipping, rotation, and brightness / contrast adjustment are performed on the light microscopy full-field section image set and electron microscopy images of the glomerular basement membrane used for training. The images are then divided into training, validation, and test sets according to a preset ratio. In this embodiment, data augmentation operations such as horizontal flipping, ±15° rotation, and ±10% brightness / contrast adjustment are performed on the light microscopy full-field section image set and electron microscopy images of the glomerular basement membrane used for training to expand the diversity of the image dataset. The training, validation, and test sets are randomly divided in a 7:1:2 ratio to ensure consistent data distribution.
[0052] The macroscopic lesion analysis module integrates a first analysis model to process the preprocessed light microscopy full-field slice image set, realize weakly supervised detection and quantification of nail-like lesions in the glomerulus, and output nail-like quantitative indicators and spatial location information of nail-like lesions.
[0053] Specifically, the first analysis model is a weakly supervised multi-instance learning model built on the state-space model (Mamba). Its core design logic is as follows: a single glomerular image in the light microscopy full-field slice image set is regarded as a "package", and image blocks of glomerular images divided according to a preset size are regarded as "instances". By aggregating the prediction results of image blocks, the quantitative index of glomerular nail-like lesions and the spatial location information of nail-like lesions are obtained.
[0054] The first analysis model employs an iterative optimization strategy during training: In the early stages of training, the first analysis model assigns the initial label of all instances to the label of their respective packages to perform weakly supervised training on the image patches. The initial label of all instances is the label of their respective packages, i.e., the patient-level label. During training, high-confidence positive sample image patches (determined to contain nail lesions) and high-confidence negative sample image patches (determined to not contain nail lesions) are periodically selected based on the confidence level predicted by the first analysis model. The selected high-confidence positive sample image patches and the high-confidence negative sample image patches are then used to update the training set to reduce the noise brought by the patient-level labels, until the performance of the first analysis model converges.
[0055] In this embodiment, during the training phase:
[0056] Each glomerular image is treated as a "bag" and cropped into 64×64 image tiles using a sliding window (step size 32 pixels) as "instances". Each "bag" contains 256 "instances".
[0057] A 64×64 image patch is processed and a 2048-dimensional high-dimensional feature vector is extracted using a MedMamba backbone network (containing 12 MedMamba blocks, each block containing grouped convolution (group=4), channel shuffling operation, and LayerNorm normalization). The 2048-dimensional high-dimensional feature vector is then mapped to a 2D output through a fully connected layer (1024 hidden layer dimensions, GELU activation function). The Softmax function is used to output the instance confidence scores for "containing nail lesions (positive) / not containing nail lesions (negative)".
[0058] A Top-k iterative optimization strategy is adopted. In the early stage of training, the initial labels of all instances are assigned to the labels of their respective packages (patient-level labels). Every certain number of training epochs (e.g., 10 epochs), the top-5 instances with the highest confidence are selected as high-confidence positive instances, and the bottom-5 instances are selected as high-confidence negative instances. These are used to update the temporary training set image data to reduce label noise. This iteration is repeated until the validation set is reached. The score showed no improvement for 200 consecutive epochs. The label of a "package" is determined by the maximum confidence score of the instances it contains. When the confidence score of any instance is ≥0.7, the "package" is considered positive (containing a nail lesion), and otherwise negative (not containing a nail lesion), thus achieving glomerular classification.
[0059] In the first analytical model inference stage, the segmented image patches are used as input, and the output is the confidence score of each image patch containing a nail-like lesion. Based on this confidence score, the nail-like lesion region is located, and the nail-like quantification index is further calculated. At the same time, the spatial location information of the nail-like lesion is output.
[0060] Pin-like lesion extraction: Based on the confidence scores of the above examples, a heatmap of the pin-like lesion region of the glomerular image is generated. The region is then binarized with a confidence score ≥ 0.7 as the threshold to obtain the spatial location information of the pin-like lesion (only the pin-like lesion regions with confidence scores that meet the threshold are retained), thus realizing the localization of the pin-like lesion.
[0061] Area calculation: The contourArea function in OpenCV is used to calculate the pixel area of the spike region. and the total area of the corresponding glomerular region. Through the formula R= / Calculate the area R of the pisiform process of a single glomerulus;
[0062] Patient-level index aggregation: To reduce sampling error for a single glomerulus, the percentage R of the nail-like area in at least three glomeruli from the same patient was selected, and the mean percentage of the nail-like area was calculated. and the maximum percentage of the area of the spike The average proportion of the nail protrusion area and the average and maximum percentage of the area of the nail protrusion As a core quantitative indicator for macroscopic lesion analysis.
[0063] The microstructure analysis module integrates a second analysis model to process preprocessed electron microscopy images, achieve precise segmentation of the glomerular basement membrane, and output a segmentation mask of the glomerular basement membrane (GBM) (GBM regions are marked as 1 and background regions are marked as 0 in the mask).
[0064] Specifically, the second analysis model adopts an encoder-decoder symmetric structure. Its core improvement is that both the encoder and the decoder integrate a feature extraction module based on a state space model (specifically LocalMamba, i.e., a visual state space model with window selective scanning) that has the ability to perform local scanning and cross-window interaction. Furthermore, each stage of the encoder and decoder includes an improved state space module (hereinafter referred to as the "SMamba module").
[0065] The second analysis model's overall architecture details are as follows: The encoder comprises four stages, each integrating 1, 14, and 1 SMamba modules respectively. The encoder's input is a preprocessed electron microscope image. After initial convolutional layers extract basic features, the first-layer feature map is obtained. Subsequent feature maps are obtained by downsampling the output of the previous SMamba module through 3×3 convolutional layers (stride set to 2). This downsampling process simultaneously reduces the feature map size and enhances semantic features. The decoder also comprises four stages. Each stage performs upsampling through transposed convolutions (3×3 kernel size, stride 2, padding 1) to restore the spatial resolution of the feature map. The upsampled feature map is then fused with the corresponding feature map from the encoder stage via skip connections to supplement detailed features. The output layer uses 1×1 convolutions to convert the 64-dimensional feature map into a 1-channel glomerular basement membrane segmentation mask.
[0066] SMamba module core design: Each SMamba module integrates three core mechanisms—L-SS2D (local 2D scanning), Swing window shifting operation, and SWL-SS2D (shifted window local 2D scanning). It sequentially performs local window scanning and cross-window scanning after window shifting, and fuses the features obtained from the two types of scans through residual connections. The specific process is as follows:
[0067] Input preprocessing: Receive the feature map output from the previous stage. (The encoder displays the feature map after downsampling from the previous stage, while the decoder displays the feature map after transposed convolution output or skip connection fusion from the previous stage.) First, normalization is performed using LayerNorm, then a 1×1 convolution is used to adjust the number of channels to C. mid =2C, enhancing feature representation capabilities;
[0068] L-SS2D Local Scanning: The preprocessed feature map is divided into non-overlapping windows at three scales: 2×2, 7×7, and global. Pixels within each window are scanned sequentially along a Z-shaped path, realizing the transformation from two-dimensional spatial features to a one-dimensional sequence. Learnable weights α, β, and γ (satisfying α+β+γ=1) are introduced to weightedly fuse the sequence features at the three scales, obtaining the local basic features. ;
[0069] Swin window shifting operation: Differentiated shifting operations are performed for windows of different sizes—a 1-pixel horizontal / vertical shift is applied to a 2×2 window, a 3-pixel horizontal / vertical shift is applied to a 7×7 window (symmetric padding is applied to the feature map before shifting to ensure its size is an integer multiple of the window size), and a (H / 2, W / 2) cyclic shift is applied to the global window; after shifting, the window is scanned again along a Z-shaped path to obtain the cross-window interactive feature sequence. The spatial structure of the feature map is restored through inverse shifting operations, enabling cross-window feature interaction;
[0070] SWL-SS2D Shifted Window Scan: Repeat the L-SS2D local scanning process on the shifted feature map, focusing on capturing detailed features at the window boundary (adapting to the structural characteristics of the glomerular basement membrane curvature and blurred boundaries), and output optimized features. ;
[0071] Feature fusion and output: and Element-wise summation, followed by 1×1 convolution to restore the number of channels to C, then applying the GELU activation function and performing residual connection with the module input to output the final feature map. =F+GELU(Conv 1×1 ( + )).
[0072] Through the above design, the second analysis model can accurately extract the features of the glomerular basement membrane in electron micrographs, and achieve precise segmentation of the irregular and blurred basement membrane.
[0073] Second analysis model configuration: AdamW optimizer (learning rate 10) is used. -3 Weight decay rate 10 -4 The loss function adopts a combination of Dice loss and cross-entropy loss (weight ratio 1:1); the weights are initialized based on the pre-trained LocalVMamba-S encoder, and training stops when the Dice similarity coefficient on the validation set is stable (fluctuation is less than 0.005 within 200 consecutive epochs).
[0074] The thickness quantification module is used to extract the skeleton and perform geometric measurements on the segmented glomerular basement membrane region, and calculate the quantification index of glomerular basement membrane thickness by combining the scale bar information in the electron microscopy image. Specifically, it includes the following units:
[0075] Image post-processing unit: Smooths and fills holes in the segmentation mask of the glomerular basement membrane to eliminate noise and holes in the segmentation mask and improve the accuracy of subsequent measurements. In this embodiment, Gaussian smoothing (σ=1.5) is performed on the segmentation mask of the glomerular basement membrane to eliminate isolated noise points, fill tiny holes, and improve boundary continuity;
[0076] Skeleton extraction unit: Extracts single-pixel-width skeleton lines from the smoothed and hole-filled segmentation mask to represent the central trajectory of the glomerular basement membrane. In this embodiment, the Zhang-Suen thinning algorithm is used to extract the skeleton lines of the glomerular basement membrane, which are the central axis of the glomerular basement membrane, ensuring the accuracy of the glomerular basement membrane thickness measurement;
[0077] Thickness calculation unit: Calculate the maximum inscribed circle diameter of the point within the segmentation mask along the skeleton line point by point, and use it as the glomerular basement membrane thickness (in pixels) corresponding to that point.
[0078] Thickness Quantification Index Aggregation Unit: Based on scale information, this unit converts the pixel-level glomerular basement membrane thickness quantification index value into actual physical thickness units and outputs the glomerular basement membrane thickness quantification index.
[0079] Specifically, the thickness quantification index aggregation unit is used to perform the following tasks:
[0080] Scale identification and conversion ratio calculation: PaddleOCR technology is used to identify the scale information (e.g., "5μm", i.e., the actual length corresponding to the scale) in electron microscope images; simultaneously, the pixel length of this scale in the image is measured. The formula k = actual length / Calculate the conversion ratio (unit: nm / pixel; note that the actual length of the scale bar should be converted from μm to nm, 1μm = 1000nm).
[0081] Actual physical thickness conversion: Multiply all the inscribed circle diameters (in pixels) obtained from the thickness calculation unit by the conversion ratio k one by one to obtain the actual glomerular basement membrane thickness (in nm) corresponding to each point, and output the actual thickness of each glomerular basement membrane.
[0082] Index Aggregation and Processing: To reduce measurement errors for single glomerular basement membrane regions, actual thickness data from multiple glomerular basement membrane regions of the same patient were selected, and the maximum value was extracted. Simultaneously, these actual thickness data were averaged and standardized to obtain a quantitative index of glomerular basement membrane thickness. In this embodiment, actual glomerular basement membrane thickness data from at least five glomerular basement membrane regions of the same patient were selected, and the mean actual glomerular basement membrane thickness was calculated. Standard deviation of actual glomerular basement membrane thickness and the maximum actual glomerular basement membrane thickness ,Will , , The thickness of the glomerular basement membrane is a core quantitative indicator for the analysis of microscopic lesions.
[0083] The feature fusion and prediction module receives the pin-like projection quantitative index from the macroscopic lesion analysis module and the glomerular basement membrane thickness quantitative index from the thickness quantitative module. It then fuses these two indices into a multidimensional feature vector, and generates a disease staging prediction result and an interpretable analysis report based on this multidimensional feature vector. Specifically, it includes the following units:
[0084] The feature fusion unit is used to perform factor analysis on the input pin quantification index and glomerular basement membrane thickness index, extract uncorrelated key feature factors, and construct a multidimensional feature vector characterizing the severity of the patient's condition.
[0085] In this embodiment, the input pin quantification index and glomerular basement membrane thickness quantification index are the core pin quantification index and the core glomerular basement membrane thickness quantification index, respectively. The process of constructing a multidimensional feature vector characterizing the severity of the patient's condition is as follows:
[0086] Suitability test: Passed Bartlett's test of sphericity (p < 10). -4 The correlation between the quantitative indicators was verified, and the applicability of factor analysis was verified by the KMO test (KMO=65.1%) to ensure that the indicators are suitable for factor decomposition.
[0087] Common factor extraction: Principal component analysis was used to extract common factors with eigenvalues ≥1, resulting in two key eigenfactors, namely Factor 1 (the dominant factor for spikes), denoted as: =0.82 +0.79 Factor 2 (the dominant factor for glomerular basement membrane thickness), denoted as: =0.85 +0.73 0.77 This constitutes the key feature vector for disease staging prediction. , ].
[0088] The classification and prediction unit uses a gradient boosting decision tree model (such as the XGBoost model) to predict the stage of membranous nephropathy (MN) based on the multidimensional feature vectors and outputs accurate disease stage results.
[0089] In this embodiment, the classification prediction unit is used to perform the following tasks:
[0090] Dataset Construction: Each sample corresponds to one patient, and the sample information includes key feature vectors. , [and clinical diagnostic labels (labels are divided into three categories: no membranous nephropathy, moderate membranous nephropathy, and severe membranous nephropathy, corresponding to no membranous nephropathy, moderate membranous nephropathy, and severe membranous nephropathy, respectively);]
[0091] Model configuration and training: Five-fold cross-validation was used to optimize the XGBoost model parameters. The final parameters of the XGBoost model were determined as follows: the maximum tree depth was set to 5, the learning rate was set to 0.1, the number of base learners was set to 200, and both the L1 regularization coefficient and the L2 regularization coefficient were set to 0.1. In multi-class classification tasks, the XGBoost model adopts the form of an objective function that outputs the predicted probability of each class.
[0092] Model evaluation: Calculate Recall, Precision, and F1 score (weighted average by category) on the independent test set to comprehensively evaluate the model's classification performance.
[0093] The interpretability analysis unit, based on the SHAP value analysis method, quantifies the contribution of each key feature factor to the prediction results, clarifies the influence weight of each indicator on the disease stage, and generates an interpretability analysis report.
[0094] In this embodiment, the key feature factors are quantified using the SHAP (SHapley Additive exPlanations) value analysis method. , The marginal contribution of each factor to the prediction results is analyzed; a feature contribution heatmap is generated to visually demonstrate the influence of each factor on the prediction of different stages, and a single-sample explanation report is output to clarify the "increased proportion of spikes" (corresponding to...). Increased thickness), "GBM thickness increased" (corresponding to Increase the specific impact weight of disease stage (progression from moderate MN to severe MN) to enhance the clinical credibility of the prediction results.
[0095] The results visualization module is used to present the spatial location information of nail-like lesions output by the macroscopic lesion analysis module, the glomerular basement membrane segmentation mask output by the microstructure analysis module, the nail-like lesion quantitative index output by the macroscopic lesion analysis module, the glomerular basement membrane thickness quantitative index output by the thickness quantification module, the disease staging prediction results and interpretability analysis report output by the feature fusion and prediction module in a graphical manner (including heat map, labeled map, statistical table, etc.), so as to facilitate medical staff to view intuitively and make clinical diagnoses.
[0096] In this embodiment, the specific functions of the result visualization module include:
[0097] Lesion visualization: Generate two types of core lesion annotation maps - glomerular classification result map (with hot icons marking peg areas, where red indicates high-confidence peg areas) and glomerular basement membrane segmentation boundary map (with white lines precisely marking the outline of the glomerular basement membrane), clearly showing the location of the lesion;
[0098] Indicator Visualization: Quantitative indicators are displayed using a dual visualization approach—precisely presenting core quantitative indicator values (including the average percentage of spike area) in tabular form. Maximum percentage of nail protrusion area Mean actual glomerular basement membrane thickness Standard deviation of actual glomerular basement membrane thickness The actual maximum thickness of the glomerular basement membrane (etc.), comparing the patient's various indicators with normal reference values in the form of bar charts to intuitively reflect the degree of abnormality of the indicators;
[0099] Prediction results display: Displays the predicted probability distribution of patients from "no MN (non-membranous nephropathy) - moderate MN (moderate membranous nephropathy) - severe MN (severe membranous nephropathy)", highlights the final diagnosis results, and includes a SHAP interpretation chart to clarify the impact of key feature factors on the prediction results;
[0100] Data synchronization: All visualization results can be exported to DICOM format and synchronized to the clinical PACS system. Doctors can also interact to zoom in and view detailed areas to meet clinical diagnostic needs.
[0101] Real-time display: The visualization module is connected to the processor and equipped with a display screen, which can display all the above-mentioned visualization content in real time and supports touch interaction.
[0102] A quantitative analysis method for multimodal pathological images of membranous nephropathy based on Mamba, applied to the aforementioned system, includes the following steps:
[0103] Step S1: Obtain pathological images of the patient to be analyzed. The pathological images include a set of light microscopic full-field section images containing glomeruli and electron microscopic images containing glomeruli. The light microscopic full-field section images and the electron microscopic images are standardized by the image preprocessing module, and the scale information is identified from the electron microscopic images.
[0104] In this embodiment, step S1 includes the following steps:
[0105] Data acquisition: A full-field light microscopy image set of kidney biopsy specimens was acquired using a light microscope scanner, and transmission electron microscopy images were acquired using a transmission electron microscope.
[0106] Preprocessing execution: Following the operating standards of the image preprocessing module, standardization processing, glomerular segmentation, scale bar annotation, and data augmentation operations are completed for the light microscope full-field slice image set and electron microscope images; then, the processed dataset is divided into training set, validation set, and test set for subsequent model training and validation.
[0107] Step S2: The standardized light microscopy full-field section image set is processed by the first analysis model to achieve weakly supervised detection and quantification of nail-like lesions within the glomerulus, outputting nail-like lesion quantification indicators and spatial location information of nail-like lesions. Specifically, the first analysis model takes image blocks divided from glomerular images in the light microscopy full-field section image set as input, outputs a confidence level representing that each image block contains nail-like lesions, and locates the nail-like region and calculates the nail-like lesion quantification indicators based on this confidence level.
[0108] In this embodiment, step S2 includes the following steps:
[0109] Model Construction: The first analysis model (weakly supervised multi-instance learning model) was built using MedMamba as the backbone network. The input was a 64×64 glomerular slice. Basic features were extracted through three 3×3 convolutional layers (with a stride of 1, containing batch normalization (BN) layers and ReLU activation function). MedMamba was the optimized state space model (Mamba).
[0110] Training configuration: Patient-level labels (labeled as "with pedicle lesions" or "without pedicle lesions") are used as "bag" labels, and the temporary dataset is updated using a Top-k iterative optimization strategy; the Adam optimizer is selected as the optimizer (learning rate set to 10). -5 The loss function is the cross-entropy loss function; after training until the F1 score on the validation set is stable, the optimal model weights are saved (the file is named checkpoint_best.pth).
[0111] Step S3: Process the standardized electron microscope image using the second analysis model to output a segmentation mask for the glomerular basement membrane, thereby segmenting the glomerular basement membrane region. Specifically, the second analysis model uses a feature extraction module integrated in its encoder-decoder structure, which has the ability to perform local scanning and cross-window interaction, to extract and reconstruct features from the electron microscope image, achieving accurate segmentation of the irregular and blurred-boundary basement membrane.
[0112] In this embodiment, step S3 includes the following steps:
[0113] Model Setup: A U-Net model with an encoder-decoder structure is built. Both the encoder and decoder integrate an improved SMamba module (containing three core mechanisms: L-SS2D local two-dimensional scanning, Swing window shifting, and SWL-SS2D shifted window local two-dimensional scanning).
[0114] Training configuration: The AdamW optimizer is selected (learning rate set to 10). -3 The weight decay rate is set to 10. -4 The loss function adopts a combination of Dice loss and cross-entropy loss (with a weight ratio of 1:1); the weights are initialized based on the pre-trained LocalVMamba-S encoder, and after training until the Dice similarity coefficient on the validation set is stable, the optimal model weights are saved (the file is named segment_best.pth).
[0115] Step S4: The glomerular basement membrane segmentation mask is extracted and geometrically measured using the thickness quantification module. Combined with the scale bar information identified in Step S1, the glomerular basement membrane thickness quantification index is calculated. The specific process includes: identifying the actual length of the scale bar and calculating the conversion ratio; converting pixel-level thickness to nanometer-level actual physical thickness; and performing mean, standardization, and maximum value extraction on the actual thicknesses of multiple glomerular basement membranes from the same patient to finally obtain the thickness quantification index. Specifically, the calculation of the glomerular basement membrane thickness quantification index includes the following processes:
[0116] The actual length in the scale information of the electron microscope image is identified by optical character recognition technology, and the conversion ratio between pixels and actual length is calculated. The geometric measurement value of basement membrane thickness calculated based on pixel units is converted into the actual physical thickness in nanometers according to the conversion ratio, and the actual thickness of glomerular basement membrane is output. The mean, standardization and maximum value extraction of the actual thickness of multiple glomerular basement membranes of the same patient are performed to obtain the quantitative index of glomerular basement membrane thickness.
[0117] In this embodiment, step S4 includes the following steps:
[0118] Modal quantization of the full-field-of-view light microscopy slice image set: The first analysis model (multi-instance learning model) trained in step S2 is applied to the test set of glomerular images to output the confidence score of the nail lesion in each slice and generate a heat map of the nail lesion region; the proportion of nail lesion area in a single glomerulus is calculated based on the heat map, and further aggregated to obtain the mean proportion of nail lesion area in the same patient. and the maximum percentage of the area of the spike ;
[0119] Electron microscopy image modality quantization: The U-Net model trained in step S3 is applied to the test set of glomerular basement membrane images to obtain a glomerular basement membrane segmentation mask; Gaussian smoothing, skeleton extraction, and inscribed circle diameter calculation are performed on the segmentation mask sequentially; unit conversion is performed using the scale information marked in step S1 to obtain the actual GBM thickness; finally, the average actual glomerular basement membrane thickness of the same patient is obtained by aggregation. Standard deviation of actual glomerular basement membrane thickness and the maximum actual glomerular basement membrane thickness .
[0120] Step S5: The feature fusion unit of the feature fusion and prediction module fuses the pin process quantitative index and the glomerular basement membrane thickness quantitative index, and performs factor analysis on the pin process quantitative index and the basement membrane thickness quantitative index to extract uncorrelated key feature factors and construct a multidimensional feature vector to characterize the severity of the patient's condition.
[0121] In this embodiment, step S5 includes the following steps:
[0122] Feature screening: Factor analysis was performed on all quantitative indicators extracted in step S4 to screen out key feature factors (including the average percentage of spike area). Mean actual glomerular basement membrane thickness wait);
[0123] Model training: The XGBoost algorithm was used to construct a three-class prediction model for "no MN (non-membranous nephropathy) - moderate MN (moderate membranous nephropathy) - severe MN (severe membranous nephropathy)", and the XGBoost model parameters were optimized through five-fold cross-validation.
[0124] Model Evaluation and Interpretation: The model performance was validated on a test set containing paired samples of multiple light microscopy full-view slice images / electron microscopy images. Recall, precision, and F1 score were calculated. The contribution of each key feature factor to the prediction results was quantified using the SHAP (SHapley Additive exPlanations) value analysis method, and an interpretability analysis report was generated.
[0125] Step S6: The classification prediction unit of the feature fusion and prediction module generates disease staging prediction results based on key feature factors, and the interpretability analysis unit generates an interpretable analysis report.
[0126] Regarding the specific structure of this invention, it should be noted that the connection relationships between the various component modules used in this invention are definite and achievable. Except as specifically described in the embodiments, their specific connection relationships can bring about corresponding technical effects and solve the technical problems proposed by this invention without relying on the execution of corresponding software programs. The models of the components, modules, and specific components appearing in this invention, the connection methods between them, and the conventional usage methods and expected technical effects brought about by the above technical features, unless specifically described, are all publicly disclosed content in patents, journal articles, technical manuals, technical dictionaries, and textbooks that can be obtained by those skilled in the art before the application date, or belong to conventional technology, common knowledge, and other existing technologies in this field. There is no need to elaborate, which makes the technical solution provided in this case clear, complete, and achievable, and can reproduce or obtain corresponding physical products based on this technical means.
[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A quantitative analysis system for multimodal pathological images of membranous nephropathy based on Mamba, characterized in that, include: The image preprocessing module is used to receive and standardize the patient's pathological images, which include a set of light microscopic full-field section images containing glomeruli and electron microscopic images containing glomeruli. The macroscopic lesion analysis module integrates the first analysis model, which is used to process the preprocessed light microscopy full-field slice image set to realize weakly supervised detection and quantification of nail-like lesions in the glomerulus, and output nail-like quantitative indicators and spatial location information of nail-like lesions. The microstructure analysis module integrates a second analysis model to process preprocessed electron microscopy images, achieve precise segmentation of the glomerular basement membrane, and output a segmentation mask for the glomerular basement membrane. The thickness quantification module is used to extract the skeleton and perform geometric measurements on the segmented glomerular basement membrane region, and output the glomerular basement membrane thickness quantification index by combining the scale bar information in the electron microscopy image. The feature fusion and prediction module receives the pin-like quantification index from the macroscopic lesion analysis module and the glomerular basement membrane thickness quantification index from the thickness quantification module. It then fuses the pin-like quantification index and the glomerular basement membrane thickness quantification index into a multidimensional feature vector. Based on the multidimensional feature vector, it generates disease staging prediction results and interpretability analysis reports. The results visualization module is used to graphically present the spatial location information of nail-like lesions, the segmentation mask of the glomerular basement membrane, the quantitative indicators of nail-like lesions, the quantitative indicators of glomerular basement membrane thickness, the prediction results of disease staging, and the interpretability analysis report.
2. The Mamba-based multimodal pathological image quantitative analysis system for membranous nephropathy according to claim 1, characterized in that, The first analysis model is a weakly supervised multi-instance learning model based on a state-space model. In the first analysis model, a single glomerular image in the light microscopy full-field slice image set is regarded as a package, and the image blocks divided into glomerular images are regarded as instances. By aggregating the prediction results of the image blocks, the quantitative index of the glomerular nail-like lesion and the spatial location information of the nail-like lesion are obtained.
3. The Mamba-based multimodal pathological image quantitative analysis system for membranous nephropathy according to claim 2, characterized in that, The first analysis model adopts an iterative optimization strategy during training: In the early stage of training, the first analysis model assigns the initial label of all instances to the label of their respective bags to perform weakly supervised training on the image patch. The initial label of all instances is assigned to the label of their respective bags, which is also the patient-level label. During training, positive and negative sample image patches with high confidence are periodically selected based on the confidence predicted by the first analysis model. The selected positive and negative sample image patches with high confidence are then used to update the training set to reduce noise from patient-level labels. This iterative process is repeated until the performance of the first analysis model converges.
4. The Mamba-based multimodal pathological image quantitative analysis system for membranous nephropathy according to claim 1, characterized in that, The second analysis model adopts an encoder-decoder structure, in which both the encoder and decoder integrate a feature extraction module based on a state-space model, which has the ability to perform local scanning and cross-window interaction. Each stage of the encoder and decoder includes an improved state space module; the improved state space module sequentially performs local window scanning and cross-window scanning after window shifting, and fuses the features obtained from local window scanning and cross-window scanning after window shifting through residual connection to output a segmentation mask of the glomerular basement membrane.
5. The Mamba-based multimodal pathological image quantitative analysis system for membranous nephropathy according to claim 1, characterized in that, The thickness quantization module includes: The image post-processing unit is used to smooth and fill holes in the segmentation mask of the glomerular basement membrane. The skeleton extraction unit is used to extract a single-pixel-width skeleton line from the segmentation mask after smoothing and hole filling. The thickness quantification index aggregation unit is used to convert the pixel-level glomerular basement membrane thickness quantification index value into actual physical thickness units based on scale information, and output the glomerular basement membrane thickness quantification index.
6. The Mamba-based multimodal pathological image quantitative analysis system for membranous nephropathy according to claim 1, characterized in that, The feature fusion and prediction module includes: The feature fusion unit is used to perform factor analysis on the input pin quantification index and glomerular basement membrane thickness quantification index, extract uncorrelated key feature factors, and construct a multidimensional feature vector characterizing the severity of the patient's condition. The classification and prediction unit uses a gradient boosting decision tree model to predict the stage of membranous nephropathy based on the multidimensional feature vectors, and outputs accurate stage prediction results. The interpretability analysis unit, based on the SHAP value analysis method, quantifies the contribution of each key feature factor to the disease staging prediction results, clarifies the influence weight of each indicator on the disease staging, and generates an interpretability analysis report.
7. A quantitative analysis method for multimodal pathological images of membranous nephropathy based on Mamba, applied to the system as described in any one of claims 1-6, characterized in that, Includes the following steps: Step S1: Acquire and standardize the pathological images of the patient to be analyzed. The pathological images include a set of light microscopic full-field section images containing glomeruli and electron microscopic images containing glomeruli, and identify scale information from the electron microscopic images. Step S2: Process the standardized light microscopy full-field slice image set through the first analysis model to realize weakly supervised detection and quantification of nail-like lesions in the glomerulus, and output nail-like quantitative indicators and spatial location information of nail-like lesions; Step S3: Process the standardized electron microscopy image through the second analysis model and output the segmentation mask of the glomerular basement membrane to segment out the glomerular basement membrane region; Step S4: Based on the segmentation mask of the glomerular basement membrane, perform skeleton extraction and geometric measurement on the segmented glomerular basement membrane region, and calculate the quantitative index of glomerular basement membrane thickness by combining the scale bar information in the electron microscopy image. Step S5: Integrate the quantitative index of the nail process with the quantitative index of the glomerular basement membrane thickness, and perform factor analysis on the quantitative index of the nail process and the index of the basement membrane thickness to extract key feature factors; Step S6: Generate disease staging prediction results and interpretability analysis report based on key feature factors.
8. The method for quantitative analysis of multimodal pathological images of membranous nephropathy based on Mamba according to claim 7, characterized in that, In step S2, the first analysis model takes image blocks divided from glomerular images in the light microscopy full-field slice image set as input, outputs the confidence level that each image block contains a nail-like lesion, and locates the spatial position of the nail-like lesion based on this confidence level and calculates the nail-like quantification index.
9. The method for quantitative analysis of multimodal pathological images of membranous nephropathy based on Mamba according to claim 7, characterized in that, In step S3, the second analysis model uses the feature extraction module integrated in its encoder-decoder structure, which has the ability to perform local scanning and cross-window interaction, to extract and reconstruct features from the electron microscope image, thereby achieving accurate segmentation of the irregular and blurred-boundary basement membrane.
10. The method for quantitative analysis of multimodal pathological images of membranous nephropathy based on Mamba according to claim 7, characterized in that, Step S4 involves calculating the quantitative index of glomerular basement membrane thickness, including the following procedures: The actual length in the scale information marked in the electron microscope image is identified by optical character recognition technology, and the conversion ratio between pixels and actual length is calculated. The geometric measurement value of the basement membrane thickness calculated based on pixel units is converted into the actual physical thickness in nanometers according to the conversion ratio, and the actual thickness of the glomerular basement membrane is output. The actual thickness of the glomerular basement membrane of multiple glomeruli from the same patient was averaged, standardized, and the maximum value was extracted to obtain a quantitative index of glomerular basement membrane thickness.