Multi-mode kidney pathology picture processing method and device, electronic equipment and medium
By employing a multi-network model processing method, the type of kidney disease is automatically identified and Oxford classification labels are added, solving the problem of difficulty in identifying Oxford classification of IgA nephropathy in existing technologies, and achieving efficient and accurate multimodal kidney pathology image processing.
Patent Information
- Application Number
- CN202410507490.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-25
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies cannot effectively identify the Oxford classification of IgA nephropathy, and the processing of multimodal renal pathology images relies on manual labor, resulting in a large workload and low accuracy.
A multi-network model processing method is adopted, including a pre-trained first network model to classify kidney disease types in immunofluorescence images, and when the images are identified as a preset type, a second network model is used to identify pathological structures in PAS-stained pathological images and add Oxford classification labels.
It automatically identifies kidney disease types and adds Oxford classification labels, significantly reducing manual workload and offering fast processing speed and high accuracy.
Smart Images

Figure CN120852265A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method, apparatus, electronic device, and medium for processing multimodal kidney pathology images. Background Technology
[0002] With the continuous development of medical imaging technology, medical image processing has significant application value in computer-aided medical classification. However, there is currently no artificial intelligence system that can directly identify primary IgA nephropathy and assess the Oxford classification. Most existing research on kidney image processing focuses on building AI models to identify kidney pathological structures. However, these models can only identify a limited number of lesion types and are mostly based on transplanted kidney pathological images, falling far short of the requirements for identifying the pathological structures of kidney diseases. This greatly limits the application of AI methods in IgA nephropathy.
[0003] In existing technologies, Zeng CH et al. used a segmentation-classification method to segment and identify four types of glomeruli in PAS-stained pathological images of IgA nephropathy: global sclerosis glomeruli, segmental sclerosis glomeruli, crescentic glomeruli, and other glomeruli. However, they could not identify the specific classifications required for assessing the Oxford classification of IgA nephropathy, such as mesangial proliferative glomeruli, endothelial proliferative glomeruli, tubular atrophy / interstitial fibrosis, and crescentic glomeruli. Therefore, it is difficult to assess the Oxford classification of IgA nephropathy. Another difficulty in identifying the pathological structures related to the Oxford classification of IgA nephropathy is that multiple lesions may coexist in the same glomerulus, such as mesangial proliferation, endothelial proliferation, and crescentic structures. This not only makes annotation difficult for doctors but also increases the difficulty of the model in the structural identification process. Although Jaugey et al. directly classified the MEST-C score of the Oxford classification based on Masson staining images, they could not identify specific pathological structures, and the consistency between the model and pathologists still needs to be improved. In summary, the processing of multimodal renal pathology images still relies on manual processing by professional nephrologists, resulting in a large workload and an inability to guarantee accuracy. Summary of the Invention
[0004] This invention provides a method, apparatus, electronic device, and medium for processing multimodal kidney pathology images, which addresses the shortcomings of existing technologies where multimodal kidney pathology image processing is labor-intensive and relies on manual labor. It utilizes multiple networks to automatically identify the types of kidney diseases in multimodal kidney pathology images and add Oxford classification tags.
[0005] According to a first aspect of the present invention, a method for processing multimodal kidney pathology images is provided, the method comprising:
[0006] Obtain multimodal kidney pathology images to be processed, wherein the multimodal kidney pathology images to be processed include immunofluorescence images and PAS-stained pathology images corresponding to the immunofluorescence images;
[0007] The immunofluorescence image is input into a pre-trained first network model, wherein the first network model is used to classify the kidney disease type from the input immunofluorescence image to obtain the kidney disease type;
[0008] Determine whether the kidney disease type output by the pre-trained first network model belongs to the preset kidney disease type;
[0009] If the disease belongs to a preset type of kidney disease, the PAS-stained pathological image corresponding to the immunofluorescence image is input into a pre-trained second network model. The second network model is used to identify the pathological structure of the input PAS-stained pathological image and output labels for the identified pathological structure.
[0010] Oxford typing labels are added to the multimodal kidney pathology images to be processed based on the labels and pathological structures output by the pre-trained second network model.
[0011] In some possible implementations, the first network model includes a label-aware attention subnetwork, a feature extraction subnetwork, and a two-branch contrastive learning subnetwork connected in sequence;
[0012] The labeled attention subnetwork is used to mark key areas of interest and the corresponding order of the marked key areas of interest on immunofluorescence images;
[0013] The feature extraction subnetwork is used to extract features according to the key interest regions and the order corresponding to the marked key interest regions;
[0014] The dual-branch contrastive learning subnetwork includes a coarse-grained branch and a fine-grained branch. The coarse-grained branch is used to select one of seven preset kidney disease labels from the features extracted by the feature extraction subnetwork to obtain a first label for the multimodal kidney pathology image to be processed. The seven preset kidney disease labels include IgA nephropathy, membranous nephropathy, diabetic nephropathy, lupus nephritis, ANCA-associated nephropathy, membranoproliferative glomerulonephritis, and amyloidosis nephritis. The coarse-grained branch is used to select one of two preset etiology labels from the features extracted by the feature extraction subnetwork to obtain a second label for the multimodal kidney pathology image to be processed, based on the first label of IgA nephropathy. The two preset etiology labels include primary and secondary etiology.
[0015] In some possible implementations, determining whether the kidney disease type output by the pre-trained first network model belongs to a preset kidney disease type includes:
[0016] Determine whether the first label output by the pre-trained first network model is IgA nephropathy, and determine whether the second label output by the pre-trained first network model is primary nephropathy;
[0017] If the first label is IgA nephropathy and the second label is primary, then the pre-trained first network model is determined to classify the multimodal kidney pathology images to be processed into a preset kidney disease type.
[0018] If the first label is not IgA nephropathy and / or the second label is not primary, then the pre-trained first network model is determined to classify the multimodal renal pathology images to be processed as not belonging to the preset renal disease type.
[0019] In some possible implementations, the second network model includes a segmentation subnet, a glomerular classification subnet, and a MESC evaluation subnet connected in sequence;
[0020] The segmentation subnet takes PAS-stained pathological images as input and is used to segment the PAS-stained pathological images according to nine preset semantic categories to obtain regions corresponding to each preset semantic category. The nine preset semantic categories include background, cortex, medulla, glomerulus, tubular atrophy / interstitial fibrosis, interstitial inflammatory cell aggregation, medium and small arteries, arterioles and veins.
[0021] The glomerular classification subnet takes the glomerular regions segmented by the segmentation subnet as input and selects one of eight preset glomerular type labels to obtain a third label. The eight preset glomerular type labels include normal glomerulus, mesangial proliferative glomerulus, endothelial proliferative glomerulus, segmental sclerotic glomerulus, crescent glomerulus, sclerotic glomerulus, incomplete glomerulus, and other glomerulus.
[0022] The MESC evaluation subnet takes the glomerular regions corresponding to the third label (mesangial proliferative glomeruli, endothelial proliferative glomeruli, segmental sclerotic glomeruli, and crescentic glomeruli) as input, and selects all matching labels from four preset MESC labels to obtain the fourth label. The four preset MESC labels include M, E, S, and C.
[0023] In some possible implementations, after the segmentation subnet segments the glomerular region and before it is input into the glomerular classification subnet, the method further includes:
[0024] The segmented glomerular regions are cropped into squares with a preset resolution.
[0025] In some possible implementations, Oxford typing labels are added to the multimodal kidney pathology images to be processed based on the labels and pathological structures output by a pre-trained second network model, including:
[0026] The fifth label is obtained by selecting one of three preset T-labels based on the area of the tubular atrophy / interstitial fibrosis region segmented by the segmented subnet;
[0027] The fourth and fifth labels output by the pre-trained second network model are combined to obtain the Oxford classification labels for the multimodal kidney pathology images to be processed.
[0028] In some possible implementations, the segmentation subnet adopts a general U-shaped architecture with a symmetrical encoder-decoder structure and skip connections at each layer, and embeds attention gates to highlight lesion-related features and suppress activation of lesion-independent regions.
[0029] The glomerular classification subnet uses ResNet-34 as the backbone and integrates spatial and channel attention units in each residual block. After obtaining features from the convolutional layer, a pyramid pooling layer is introduced to concentrate features at different scales. The probability distribution of different types of glomeruli is also obtained through a fully connected layer.
[0030] The MESC evaluation subnet uses DenseNet-121 as its backbone. The label prediction results are output by the fully connected layer, and the output categories are adjusted to correspond to four preset MESC labels. Each node of the fully connected layer corresponds to one preset MESC label.
[0031] According to a second aspect of the present invention, the present invention also provides a multimodal kidney pathology image processing apparatus, the apparatus comprising:
[0032] The acquisition module is used to acquire multimodal kidney pathology images to be processed, wherein the multimodal kidney pathology images to be processed include immunofluorescence images and PAS-stained pathology images corresponding to the immunofluorescence images;
[0033] A classification module is used to input the immunofluorescence image into a pre-trained first network model, wherein the first network model is used to classify the input immunofluorescence image into kidney disease types to obtain kidney disease types;
[0034] The judgment module is used to determine whether the kidney disease type output by the pre-trained first network model belongs to the preset kidney disease type;
[0035] The identification module is used to input the PAS-stained pathological image corresponding to the immunofluorescence image into a pre-trained second network model when the kidney disease belongs to a preset type. The second network model is used to identify the pathological structure of the input PAS-stained pathological image and output a label for the identified pathological structure.
[0036] The labeling module is used to add Oxford typing labels to the multimodal kidney pathology images to be processed based on the labels and pathological structures output by the pre-trained second network model.
[0037] According to a third aspect of the present invention, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multimodal kidney pathology image processing method as described above.
[0038] According to a fourth aspect of the present invention, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multimodal kidney pathology image processing method as described above.
[0039] This invention provides a multimodal kidney pathology image processing method. It involves inputting immunofluorescence images from the multimodal kidney pathology images to be processed into a pre-trained first network model for anomaly classification. Next, it determines whether the kidney disease type output by the first network model belongs to a preset kidney disease type. If it does, the PAS-stained pathology image corresponding to the immunofluorescence image is input into a pre-trained second network model for pathological structure recognition and labeling. Finally, Oxford typing labels are added to the multimodal kidney pathology images based on the labels and pathological structures output by the second network model. This method automatically identifies multimodal kidney pathology images of kidney disease types and adds Oxford typing labels, providing auxiliary analysis data for kidney examinations, significantly saving human workload, and offering fast processing speed and high accuracy.
[0040] In addition, the multimodal kidney pathology image processing device, electronic device, and non-transitory computer-readable storage medium provided by the present invention can also achieve the above-mentioned technical effects, and will not be described in detail here. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0042] Figure 1 This is one of the flowcharts of the multimodal kidney pathology image processing method provided by the present invention;
[0043] Figure 2 This is a schematic diagram of the structure of the first network model provided by the present invention;
[0044] Figure 3 This is a schematic diagram of the training of the first network model provided by the present invention;
[0045] Figure 4 This is a schematic diagram of the structure of the second network model provided by the present invention;
[0046] Figure 5 This is a schematic diagram of the structure of each sub-model in the second network model provided by the present invention;
[0047] Figure 6 is with Figure 5 Color-coded schematic diagrams of the structure of each sub-model in the corresponding second network model;
[0048] Figure 7 This is a schematic diagram of the structure of the multimodal kidney pathology image processing device provided by the present invention;
[0049] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0051] The following is combined with Figures 1 to 8 This invention describes a multimodal kidney pathology image processing method, apparatus, electronic device, and medium.
[0052] Figure 1 This is a flowchart illustrating the multimodal kidney pathology image processing method provided in this embodiment of the invention. Please refer to it. Figure 1 As shown, this embodiment provides a multimodal kidney pathology image processing method, which can be implemented through steps S101 to S105. Each step will be described in detail below:
[0053] Step 101: Obtain multimodal kidney pathology images to be processed, wherein the multimodal kidney pathology images to be processed include immunofluorescence images and PAS-stained pathology images corresponding to the immunofluorescence images.
[0054] In this embodiment, the multimodal renal pathology images to be processed refer to images that need to be tagged with the Oxford typing. The immunofluorescence images and PAS-stained pathology images included in the multimodal renal pathology images exist in pairs. In practice, the same target kidney can be imaged using both immunofluorescence imaging equipment and PAS-stained pathology images to obtain paired immunofluorescence images and PAS-stained pathology images. It should be noted that the pairing of immunofluorescence images and PAS-stained pathology images means that both originate from the same target kidney disease, not that the number of images is paired.
[0055] Step 102: Input the immunofluorescence image into a pre-trained first network model, wherein the first network model is used to classify the kidney disease type from the input immunofluorescence image.
[0056] In this implementation, the pre-trained first network model is trained using a large number of immunofluorescence images with labeled abnormal categories. After inputting the immunofluorescence images from the multimodal renal pathology images to be processed into the pre-trained first network model, the pre-trained first network model can output the renal disease type. It should be noted that the pre-trained first network model can identify multiple pre-defined basic renal disease types, and the output renal disease type is one of the multiple pre-defined basic renal disease types.
[0057] Step 103: Determine whether the kidney disease type output by the pre-trained first network model belongs to the preset kidney disease type.
[0058] In this embodiment, the preset kidney disease type refers to a specific one among a variety of preset kidney disease types that the first network model can identify.
[0059] Step 104: If the disease belongs to a preset type of kidney disease, the PAS-stained pathological image corresponding to the immunofluorescence image is input into a pre-trained second network model. The second network model is used to identify the pathological structure of the input PAS-stained pathological image and output labels for the identified pathological structures.
[0060] In this embodiment, the pre-trained second network model is trained using a large number of PAS-stained pathological images with labeled pathological structure locations and labels. After the PAS-stained pathological images are input into the pre-trained second network model, the pre-trained second network model can identify the pathological structures and add labels to the corresponding pathological structures.
[0061] In practice, if the kidney disease type output by the pre-trained first network model does not belong to the preset kidney disease type, there is no need to process the PAS-stained pathological image corresponding to the immunofluorescence image.
[0062] Step 105: Add Oxford typing labels to the multimodal kidney pathology images to be processed based on the labels and pathological structures output by the pre-trained second network model.
[0063] This embodiment of the multimodal renal pathology image processing method involves inputting the immunofluorescence image from the multimodal renal pathology image to be processed into a pre-trained first network model to achieve anomaly classification. Then, it determines whether the renal disease type output by the first network model belongs to a preset renal disease type. If it belongs to the preset renal disease type, the PAS-stained pathological image corresponding to the immunofluorescence image is input into a pre-trained second network model to achieve pathological structure recognition and labeling. Finally, based on the labels and pathological structures output by the second network model, Oxford typing labels are added to the multimodal renal pathology image to be processed. This method achieves automatic identification of renal disease types in multimodal renal pathology images and adds Oxford typing labels, providing auxiliary analysis data for renal examinations, significantly saving human workload, and exhibiting fast processing speed and high accuracy.
[0064] In some possible implementations, the first network model includes a label-aware attention subnetwork, a feature extraction subnetwork, and a two-branch contrastive learning subnetwork connected in sequence;
[0065] The labeled attention subnetwork is used to mark key areas of interest and the corresponding order of the marked key areas of interest on immunofluorescence images;
[0066] The feature extraction subnetwork is used to extract features according to the key interest regions and the order corresponding to the marked key interest regions;
[0067] The dual-branch contrastive learning subnetwork includes a coarse-grained branch and a fine-grained branch. The coarse-grained branch is used to select one of seven preset kidney disease labels from the features extracted by the feature extraction subnetwork to obtain a first label for the multimodal kidney pathology image to be processed. The seven preset kidney disease labels include IgA nephropathy, membranous nephropathy, diabetic nephropathy, lupus nephritis, ANCA-associated nephropathy, membranoproliferative glomerulonephritis, and amyloidosis nephritis. The coarse-grained branch is used to select one of two preset etiology labels from the features extracted by the feature extraction subnetwork to obtain a second label for the multimodal kidney pathology image to be processed, based on the first label of IgA nephropathy. The two preset etiology labels include primary and secondary etiology.
[0068] Figure 2 This is a schematic diagram of the structure of the first network model provided by the present invention. Please refer to it. Figure 2As shown, in the specific implementation process, seven immunofluorescence images can be used to construct a contrastive learning model to distinguish between primary and secondary IgA nephropathy. Specifically, a double-branch box consisting of coarse-grained and fine-grained classification branches can be created:
[0069] A coarse-grained branch is used to classify the seven basic types of kidney disease, while a fine-grained branch is specifically designed to extract the differences between primary and secondary IgA nephropathy. Furthermore, a mutual feedback mechanism is designed to optimize the predictions of the coarse-grained branch based on the confidence scores of the fine-grained branch. More importantly, a contrastive learning strategy is employed in each branch to learn more discriminative feature representations between different kidney diseases during training. Before contrastive learning, all kidney diseases were equally separated, resulting in small feature differences between them. However, after contrastive learning, the feature differences between specific categories significantly increased.
[0070] Figure 3 This is a training diagram of the first network model provided by the present invention. To effectively train the dual-branch framework, the network training includes three training phases. Please refer to... Figure 3 As shown, the label-aware attention module is first pre-trained under the supervision of collected regions of interest from nephrologists. In the second training phase, the entire bi-branch model is trained, including coarse-grained and fine-grained branches, as well as a mutual feedback mechanism. In the third training phase, a contrastive learning strategy is employed to further fine-tune the entire model, making the learned feature representations more discriminative. This approach accurately distinguishes different categories of kidney disease, especially primary and secondary nephropathy, providing a powerful tool for immunofluorescence image classification. Finally, a loss function for contrastive learning is added in the third stage, including the overall accuracy and Kappa value for primary / secondary classification. All ablation models are trained under the same settings. All major components contribute to the final classification performance.
[0071] The Kullback-Leibler (KL) divergence loss was used to calculate S. m and The differences are as follows:
[0072]
[0073] in, and These are the prediction and true significance plots for the k-th immunofluorescence label, respectively.
[0074] Regarding the importance of immunofluorescence labeling, r m Predicted label importance and The ranking and mean squared error (MSE) loss are calculated among the real labeled attention pairs as follows:
[0075]
[0076] In the above equation, λ rank and λ mse It is a hyperparameter used to balance the training contributions of the two loss functions. ij It is the preset empirical weight for each pair of tags. and They represent r respectively m and The i-th element. In this way, the developed label-aware attention module can learn from real-world kidney disease classifications, which improves the interpretability of the proposed method. More importantly, the spatial saliency S of the prediction... m Importance of immunofluorescence labeling r m This allows for efficient and accurate classification by highlighting key pathological areas and important immunofluorescence markers. A bibranching nephropathy classification network was developed to better differentiate fine-grained nephropathy. Specifically, the coarse and fine-grained branches tend to classify the basic type of nephropathy, as well as its primary / secondary nature.
[0077]
[0078] In the above equation, θ r and θ s These are hyperparameters that adjust the predictive significance. Furthermore, Mean(·) and Var(·) are the mean and variance operations over all immunofluorescence labels: I is a one-matrix. Then, based on ResNet-18, the participating immunofluorescence sequences X are input into the classification backbone to extract renal disease features. Unlike traditional ResNet, all convolutional structures in the classification backbone are reimplemented as group convolutions to encourage the network to preserve the features of each immunofluorescence label. Specifically, the convolutions are divided into K groups, where K is the number of immunofluorescence labels. Furthermore, spatial and channel attention mechanisms are introduced into each residual block to improve classification efficiency. The extracted features are then simultaneously fed into a two-branch classification structure, where coarse and fine-grained branches develop classifications for seven basic types of kidney disease: IgA nephropathy, membranous nephropathy, diabetic nephropathy, lupus nephritis, ANCA-associated nephropathy, membranoproliferative glomerulonephritis, and amyloidosis nephritis (i.e., MN, IgAN, LN, DN, ANCA, MPGN, and AL), as well as primary / secondary kidney disease. In summary, given the involved immunofluorescence sequence X, the classification process can be presented as follows:
[0079]
[0080] In the above equation, FE(·|φ) represents a feature encoder with learnable parameters p. φ.MLP i (·) and F res (·) represent the node-based multilayer perceptron and the proposed classification backbone network, respectively. Furthermore, and These represent the predictions for coarse and fine-grained branches, respectively. To supervise the predicted kidney disease, the cross-entropy loss is calculated on both branches:
[0081]
[0082] Where y coarse and y fine These are the truth labels for the seven basic types of kidney disease, and the primary / secondary nature of the kidney disease. coarse and p fine is the predicted probability of class j in each branch, and represents the indicator function. Furthermore, w coarse and w fine These are preset weights used to balance the training contribution of each class in the cross-entropy loss, λ. coarse and λ fine It is a hyperparameter used to balance the losses of coarse and fine-grained branching.
[0083] In addition to cross-entropy loss, this embodiment also develops a contrastive learning strategy to encourage the network to learn more discriminative feature representations between predicted categories, employing a contrastive learning and mutual feedback mechanism. In the fine-grained branch, a triplet loss is constructed to increase the feature distance between primary and secondary properties, contrasting it with the distance between basic types of kidney disease. For example, if an immunofluorescence sequence of primary IgA nephropathy is selected as the anchor sample, then immunofluorescence sequences of primary membranous nephropathy and secondary IgA nephropathy are subsequently selected as positive and negative samples, respectively. Let represent the encoded features of the anchor sample, positive sample, and negative sample, respectively.
[0084]
[0085] Where m is a hyperparameter used to adjust the distance boundary. Similarly, in the coarse branch, a triplet for primary IgA nephropathy is constructed by using secondary IgA nephropathy and primary membranous nephropathy as positive and negative samples, respectively. The coarse branch is learned to better distinguish the basic types of kidney disease.
[0086] Furthermore, this embodiment develops a mutual feedback mechanism between the two branches to refine the prediction of the basic type of kidney disease by considering the predictive confidence of the fine-grained branch. Specifically, the confidence score S of the fine-grained branch... fine Defined by the absolute difference between the predicted probabilities of primary and secondary natures. coarse and y fineThese represent the predictions for coarse and fine-grained branches, respectively. The mutual feedback mechanism can be represented as:
[0087]
[0088] In the above equation, θ represents element-wise multiplication; m and θ v This is a hyperparameter used to adjust the scale of the confidence score. In this way, when the fine-grained branch has high classification confidence in the primary and secondary nature of kidney disease, the prediction of the primary / secondary nature of kidney disease with a higher prediction probability will be improved, and vice versa.
[0089] In some possible implementations, determining whether the kidney disease type output by the pre-trained first network model belongs to a preset kidney disease type specifically includes:
[0090] Determine whether the first label output by the pre-trained first network model is IgA nephropathy, and determine whether the second label output by the pre-trained first network model is primary nephropathy;
[0091] If the first label is IgA nephropathy and the second label is primary, then the pre-trained first network model is determined to classify the multimodal kidney pathology images to be processed into a preset kidney disease type.
[0092] If the first label is not IgA nephropathy and / or the second label is not primary, then the pre-trained first network model is determined to classify the multimodal renal pathology images to be processed as not belonging to the preset renal disease type.
[0093] The multimodal renal pathology image processing method in this embodiment utilizes two tags generated at both fine and coarse granular levels to accurately classify the renal disease type in immunofluorescence images. It can accurately distinguish multimodal renal pathology images of primary IgA nephropathy and has high accuracy.
[0094] In some possible implementations, the second network model includes a segmentation subnet, a glomerular classification subnet, and a MESC evaluation subnet connected in sequence;
[0095] The segmentation subnet takes PAS-stained pathological images as input and is used to segment the PAS-stained pathological images according to nine preset semantic categories to obtain regions corresponding to each preset semantic category. The nine preset semantic categories include background, cortex, medulla, glomerulus, tubular atrophy / interstitial fibrosis, interstitial inflammatory cell aggregation, medium and small arteries, arterioles and veins.
[0096] The glomerular classification subnet takes the glomerular regions segmented by the segmentation subnet as input and selects one of eight preset glomerular type labels to obtain a third label. The eight preset glomerular type labels include normal glomerulus, mesangial proliferative glomerulus (M), endothelial proliferative glomerulus (E), segmental sclerotic glomerulus (S), crescentic glomerulus (C), sclerotic glomerulus, incomplete glomerulus and other glomerulus.
[0097] The MESC evaluation subnet takes the glomerular regions corresponding to the third label (mesangial proliferative glomeruli, endothelial proliferative glomeruli, segmental sclerotic glomeruli, and crescentic glomeruli) as input, and selects all matching labels from four preset MESC labels to obtain the fourth label. The four preset MESC labels include M, E, S, and C.
[0098] Figure 4 This is a schematic diagram of the structure of the second network model provided by the present invention. In order to achieve further IgA nephropathy structure identification and Oxford classification of multimodal renal pathology images of primary IgA nephropathy, the second network model can be constructed in the following manner during the specific implementation process:
[0099] It is built upon a three-stage framework, including lesion segmentation, glomerular classification, and a MESC assessment subnetwork. The DeepSNN framework is... Figure 4 The process is described below. Using whole-slice images (WSI) as input, the lesion segmentation subnetwork segments glomeruli and histopathological lesions according to nine semantic categories: background, cortex, medulla, glomerulus, tubular atrophy / interstitial fibrosis, interstitial inflammatory cell aggregation, medium and small arteries, arterioles, and veins. Connectivity component analysis is then performed on the predicted mask output by the segmentation subnetwork. Subsequently, the segmented glomeruli are input into a glomerular classification subnetwork, designed to distinguish eight types of fine-grained glomeruli, including normal glomeruli, MESC glomeruli, sclerotic glomeruli, incomplete glomeruli, and other glomeruli. It is important to note that before being input into the glomerular classification subnetwork, glomeruli are cropped with rectangular bounding boxes and resized to 512x512 to ensure consistency. Subsequently, considering the complexity of MESC glomeruli, a multi-label MESC evaluation subnetwork was specifically developed to simultaneously identify four types of lesions coexisting in MESC glomeruli. Therefore, the constructed second network model can achieve histopathological identification at the pixel level and Oxford classification at the image level.
[0100] In some possible implementations, after the segmentation subnet segments the glomerular region and before it is input into the glomerular classification subnet, the method further includes:
[0101] The segmented glomerular regions are cropped into squares of a preset resolution. For example, before being input into the glomerular classification subnetwork during the specific implementation process, the glomeruli are cropped by rectangular bounding boxes and resized to 512x512 to ensure consistency.
[0102] In some possible implementations, step S105, which involves adding Oxford typing labels to the multimodal kidney pathology image to be processed based on the labels and pathological structures output by the pre-trained second network model, specifically includes:
[0103] The fifth label is obtained by selecting one of three preset T-labels based on the area of the tubular atrophy / interstitial fibrosis region segmented by the segmented subnet;
[0104] Specifically, the three preset T labels include T0, T1, and T2. When the area of the segmented tubular atrophy / interstitial fibrosis region reaches 0 to 25%, the fifth label is T0; when the area of the segmented tubular atrophy / interstitial fibrosis region reaches 26 to 50%, the fifth label is T1; and when the area of the segmented tubular atrophy / interstitial fibrosis region reaches 51 to 100%, the fifth label is T2.
[0105] The fourth and fifth labels output by the pre-trained second network model are combined to obtain the Oxford classification labels for the multimodal kidney pathology images to be processed.
[0106] In some possible implementations, the segmentation subnet adopts a general U-shaped architecture with a symmetrical encoder-decoder structure and skip connections at each layer, and embeds attention gates to highlight lesion-related features and suppress activation of lesion-independent regions.
[0107] The glomerular classification subnet uses ResNet-34 as the backbone and integrates spatial and channel attention units in each residual block. After obtaining features from the convolutional layer, a pyramid pooling layer is introduced to concentrate features at different scales. The probability distribution of different types of glomeruli is also obtained through a fully connected layer.
[0108] The MESC evaluation subnet uses DenseNet-121 as its backbone. The label prediction results are output by the fully connected layer, and the output categories are adjusted to correspond to four preset MESC labels. Each node of the fully connected layer corresponds to one preset MESC label.
[0109] Figure 5 This is a schematic diagram of the structure of each sub-model in the second network model provided by the present invention. Figure 6 is with Figure 5 The corresponding color diagram of the structure of each sub-model in the second network model; and Figure 5 The difference is Figure 6Colored arrows and network layers are used to represent the structure of each part of the second network. The following will combine... Figure 5 and Figure 6 The specific structure of the second network is detailed below: For lesion segmentation, a sub-network was developed based on the state-of-the-art backbone network U-net. Specifically, the segmentation sub-network adopts a general U-shaped architecture with a symmetrical encoder-decoder structure and skip connections in each layer to reduce boundary distortion. Furthermore, a gated attention mechanism is embedded in the segmentation sub-network to highlight lesion-related features and suppress activation of irrelevant regions through attention gating. For the fine-grained glomerular classification sub-network, ResNet-34 is used as the backbone, and spatial and channel attention modules are integrated in each residual block. Pyramid pooling layers are introduced after features are obtained from convolutional layers to concentrate features at different scales. Finally, the probability distribution of different types of glomeruli is obtained through fully connected layers. The attention mechanism in the glomerular classification sub-network facilitates the identification of eight types of fine-grained pathological glomeruli. The MESC glomeruli from the glomerular classification sub-network are then input into the subsequent MESC evaluation sub-network. More specifically, DenseNet-121 can be used as the backbone of the MESC classification sub-network, which encourages feature reuse during propagation. Finally, the prediction results are output by fully connected layers. It is worth noting that the output is adjusted to fit a multi-label setting, where each node in the fully connected layer independently predicts the probability of a specific type of lesion. Therefore, M / E / S / C lesions can be identified simultaneously.
[0110] In the specific implementation, to ensure the performance of the first and second networks, the data can be pre-processed. Specifically, a heuristic similar to that in U-Net is adopted during the training of the lesion segmentation sub-network. For example, during the training phase, the foreground is oversampled to generate training blocks of size 1792x768, ensuring that the training data contains sufficient information. Furthermore, multiple data augmentation and Z-score normalization strategies are performed. During the testing phase, the framework utilizes block expansion and test-time augmentation strategies to significantly reduce boundary distortion. Additionally, inspired by previous work, color adaptive normalization is applied to whole-slice images (WSI) to normalize color intensities from different WSI sources. For the glomerular classification sub-network and the MESC multi-label classification sub-network, connected component analysis is performed on the segmentation mask output by the segmentation sub-network. Then, the segmented glomeruli are cropped into rectangular bounding boxes and resized to 512x512 to ensure consistency before the glomerular blocks are input into the sub-network. During the training of the two sub-networks, a learning rate decay strategy and multiple data augmentations, such as random flipping and brightness transformation, were implemented to alleviate the overfitting problem.
[0111] To optimize network performance, the loss function for each part of the second network model can be set as follows during implementation:
[0112] The loss function for the aforementioned subnetting is as follows:
[0113] Specifically, the output of each level in the decoder is passed to an expansion block to generate prediction masks at different spatial resolutions. Then, given the predictions at a certain scale, a combination of Dice loss and cross-entropy loss is computed, which has proven to be an effective constraint for optimizing the segmentation task. Therefore, the final training objective function is the sum of the corresponding losses at the first four maximum resolutions.
[0114]
[0115] In the formula, Y s These are the predicted segmentation masks and their corresponding ground truth values at different spatial scales. λ s This represents the hyperparameter that balances the contributions of training at different scales. J is the number of classes to be segmented. and y j,p λ represents the true class and the output predicted probability of each pixel in the input image, respectively. dice and λ ce It is a hyperparameter that balances the importance of the two types of loss.
[0116] For the aforementioned glomerular classification subnetwork, a weighted cross-entropy loss is introduced to evaluate the distance between the predicted class and the true label. Mathematically, the classification loss used to train the glomerular classification subnetwork can be expressed as: In the formula, C is the number of glomerular categories to be distinguished, and y is the true label. Let ξ be the predicted probability of the j-th category, and ξ be a preset weight. y It is used to balance the training contributions of different categories.
[0117] For the MESC evaluation subnet, a weighted combination of binary cross-entropy loss is introduced to evaluate the distance between the predicted class and the true label for each subclass. Mathematically, the multi-label classification loss can be expressed as: In the formula, K is the number of Oxford fractal categories to be distinguished, and pl and Let represent the true probability and predicted probability of the l-th subclass, respectively. ηl represents the weights that balance the training contributions of different subclasses. σ(·) represents the sigmoid function.
[0118] In the specific implementation process, in order to ensure the accuracy of segmentation subnet recognition, image post-processing is carried out by using image connected component analysis to filter noise. Regions with fewer than 300 pixels are regarded as noise and uniformly classified as cortex. Due to the special nature of the medulla, it appears in large areas. Regions with fewer than 300,000 pixels are regarded as noise and classified as cortex.
[0119] The multimodal kidney pathology image processing device provided by the present invention will be described below. The multimodal kidney pathology image processing device described below can be referred to in correspondence with the multimodal kidney pathology image processing method described above.
[0120] Please refer to Figure 7 As shown, this embodiment provides a multimodal kidney pathology image processing device, which includes: an acquisition module 710, a classification module 720, a judgment module 730, an identification module 740, and a tag adding module 750. The following is a detailed description of each of these sub-modules:
[0121] The acquisition module 710 is used to acquire multimodal kidney pathology images to be processed, wherein the multimodal kidney pathology images to be processed include immunofluorescence images and PAS-stained pathology images corresponding to the immunofluorescence images;
[0122] The classification module 720 is used to input the immunofluorescence image into a pre-trained first network model, wherein the first network model is used to classify the input immunofluorescence image into kidney disease types to obtain kidney disease types;
[0123] The judgment module 730 is used to determine whether the kidney disease type output by the pre-trained first network model belongs to the preset kidney disease type;
[0124] The identification module 740 is used to input the PAS-stained pathological image corresponding to the immunofluorescence image into a pre-trained second network model when the kidney disease belongs to a preset type. The second network model is used to identify the pathological structure of the input PAS-stained pathological image and output a label for the identified pathological structure.
[0125] The labeling module 750 is used to add Oxford typing labels to the multimodal kidney pathology images to be processed based on the labels and pathological structures output by the pre-trained second network model.
[0126] The multimodal renal pathology image processing device of this embodiment classifies anomalies by inputting immunofluorescence images from the multimodal renal pathology images to be processed into a pre-trained first network model. Then, it determines whether the renal disease type output by the first network model belongs to a preset renal disease type. If it belongs to the preset renal disease type, the PAS-stained pathological image corresponding to the immunofluorescence image is input into a pre-trained second network model to identify the pathological structure and add labels. Finally, Oxford typing labels are added to the multimodal renal pathology images to be processed based on the labels and pathological structures output by the second network model. This realizes the automatic identification of renal disease types in multimodal renal pathology images and the addition of Oxford typing labels, which can provide auxiliary analysis data for renal examination, significantly save human workload, and has a fast processing speed and high accuracy.
[0127] It should be noted that the specific limitations of the multimodal kidney pathology image processing device can be found in the limitations of the multimodal kidney pathology image processing method described above, and will not be repeated here. Each sub-network in the aforementioned multimodal kidney pathology image processing device can be implemented entirely or partially through software, hardware, or a combination thereof. Each sub-network can be embedded in or independent of the processor in the computer device in hardware form, or it can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each sub-network.
[0128] Figure 8 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 8As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute a multimodal kidney pathology image processing method. The method includes: acquiring a multimodal kidney pathology image to be processed, wherein the image includes an immunofluorescence image and a corresponding PAS-stained pathology image; inputting the immunofluorescence image into a pre-trained first network model, wherein the first network model is used to classify the input immunofluorescence image into kidney disease types; determining whether the kidney disease type output by the pre-trained first network model belongs to a preset kidney disease type; if it belongs to the preset kidney disease type, inputting the corresponding PAS-stained pathology image into a pre-trained second network model, wherein the second network model is used to identify pathological structures in the input PAS-stained pathology image and output labels for the identified pathological structures; and adding Oxford typing labels to the multimodal kidney pathology image based on the labels and pathological structures output by the pre-trained second network model.
[0129] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0130] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multimodal renal pathology image processing method provided by the above methods. The method includes: acquiring a multimodal renal pathology image to be processed, wherein the multimodal renal pathology image to be processed includes an immunofluorescence image and a PAS-stained pathology image corresponding to the immunofluorescence image; inputting the immunofluorescence image into a pre-trained first network model, wherein the first network model is used to classify the input immunofluorescence image into renal disease types to obtain renal disease types; determining whether the renal disease type output by the pre-trained first network model belongs to a preset renal disease type; if it belongs to the preset renal disease type, then inputting the PAS-stained pathology image corresponding to the immunofluorescence image into a pre-trained second network model, wherein the second network model is used to identify the pathological structure of the input PAS-stained pathology image and output labels for the identified pathological structures; and adding Oxford typing labels to the multimodal renal pathology image to be processed based on the labels and pathological structures output by the pre-trained second network model.
[0131] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a multimodal renal pathology image processing method provided by the methods described above. The method includes: acquiring a multimodal renal pathology image to be processed, wherein the multimodal renal pathology image to be processed includes an immunofluorescence image and a PAS-stained pathology image corresponding to the immunofluorescence image; inputting the immunofluorescence image into a pre-trained first network model, wherein the first network model is used to classify the input immunofluorescence image into a renal disease type; determining whether the renal disease type output by the pre-trained first network model belongs to a preset renal disease type; if it belongs to the preset renal disease type, inputting the PAS-stained pathology image corresponding to the immunofluorescence image into a pre-trained second network model, wherein the second network model is used to identify the pathological structure of the input PAS-stained pathology image and output labels for the identified pathological structures; and adding Oxford typing labels to the multimodal renal pathology image to be processed based on the labels and pathological structures output by the pre-trained second network model.
[0132] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the subnets can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0133] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for processing multimodal renal pathology images, characterized in that, The method includes: Obtain multimodal kidney pathology images to be processed, wherein the multimodal kidney pathology images to be processed include immunofluorescence images and PAS-stained pathology images corresponding to the immunofluorescence images; The immunofluorescence image is input into a pre-trained first network model, wherein the first network model is used to classify the kidney disease type from the input immunofluorescence image to obtain the kidney disease type; Determine whether the kidney disease type output by the pre-trained first network model belongs to the preset kidney disease type; If the disease belongs to a preset type of kidney disease, the PAS-stained pathological image corresponding to the immunofluorescence image is input into a pre-trained second network model. The second network model is used to identify the pathological structure of the input PAS-stained pathological image and output labels for the identified pathological structure. Oxford typing labels are added to the multimodal kidney pathology images to be processed based on the labels and pathological structures output by the pre-trained second network model.
2. The multimodal kidney pathology image processing method according to claim 1, characterized in that, The first network model includes a label-aware attention subnetwork, a feature extraction subnetwork, and a two-branch contrastive learning subnetwork connected in sequence; The labeled attention subnetwork is used to mark key areas of interest and the corresponding order of the marked key areas of interest on immunofluorescence images; The feature extraction subnetwork is used to extract features according to the key interest regions and the order corresponding to the marked key interest regions; The dual-branch contrastive learning subnetwork includes a coarse-grained branch and a fine-grained branch. The coarse-grained branch is used to select one of seven preset kidney disease labels from the features extracted by the feature extraction subnetwork to obtain a first label for the multimodal kidney pathology image to be processed. The seven preset kidney disease labels include IgA nephropathy, membranous nephropathy, diabetic nephropathy, lupus nephritis, ANCA-associated nephropathy, membranoproliferative glomerulonephritis, and amyloidosis nephritis. The coarse-grained branch is used to select one of two preset etiology labels from the features extracted by the feature extraction subnetwork to obtain a second label for the multimodal kidney pathology image to be processed, based on the first label of IgA nephropathy. The two preset etiology labels include primary and secondary etiology.
3. The multimodal kidney pathology image processing method according to claim 2, characterized in that, Determining whether the kidney disease type output by the pre-trained first network model belongs to the preset kidney disease type includes: Determine whether the first label output by the pre-trained first network model is IgA nephropathy, and determine whether the second label output by the pre-trained first network model is primary nephropathy; If the first label is IgA nephropathy and the second label is primary, then the pre-trained first network model is determined to classify the multimodal kidney pathology images to be processed into a preset kidney disease type. If the first label is not IgA nephropathy and / or the second label is not primary, then the pre-trained first network model is determined to classify the multimodal renal pathology images to be processed as not belonging to the preset renal disease type.
4. The multimodal kidney pathology image processing method according to claim 1, characterized in that, The second network model includes a segmentation subnetwork, a glomerular classification subnetwork, and a MESC evaluation subnetwork connected in sequence; The segmentation subnet takes PAS-stained pathological images as input and is used to segment the PAS-stained pathological images according to nine preset semantic categories to obtain regions corresponding to each preset semantic category. The nine preset semantic categories include background, cortex, medulla, glomerulus, tubular atrophy / interstitial fibrosis, interstitial inflammatory cell aggregation, medium and small arteries, arterioles and veins. The glomerular classification subnet takes the glomerular regions segmented by the segmentation subnet as input and selects one of eight preset glomerular type labels to obtain a third label. The eight preset glomerular type labels include normal glomerulus, mesangial proliferative glomerulus, endothelial proliferative glomerulus, segmental sclerotic glomerulus, crescent glomerulus, sclerotic glomerulus, incomplete glomerulus, and other glomerulus. The MESC evaluation subnet takes the glomerular regions corresponding to the third label (mesangial proliferative glomeruli, endothelial proliferative glomeruli, segmental sclerotic glomeruli, and crescentic glomeruli) as input, and selects all matching labels from four preset MESC labels to obtain the fourth label. The four preset MESC labels include M, E, S, and C.
5. The multimodal kidney pathology image processing method according to claim 4, characterized in that, After the segmentation subnetwork segments the glomerular region and before it is input into the glomerular classification subnetwork, the method further includes: The segmented glomerular regions are cropped into squares with a preset resolution.
6. The multimodal kidney pathology image processing method according to claim 4, characterized in that, Oxford typing labels are added to the multimodal kidney pathology images to be processed based on the labels and pathological structures output by the pre-trained second network model, including: The fifth label is obtained by selecting one of three preset T-labels based on the area of the tubular atrophy / interstitial fibrosis region segmented by the segmented subnet; The fourth and fifth labels output by the pre-trained second network model are combined to obtain the Oxford classification labels for the multimodal kidney pathology images to be processed.
7. The multimodal kidney pathology image processing method according to claim 4, characterized in that, The segmentation subnet adopts a general U-shaped architecture with a symmetrical encoder-decoder structure and skip connections in each layer, and embeds attention gates to highlight lesion-related features and suppress activation in lesion-independent regions. The glomerular classification subnet uses ResNet-34 as the backbone and integrates spatial and channel attention units in each residual block. After obtaining features from the convolutional layer, a pyramid pooling layer is introduced to concentrate features at different scales. The probability distribution of different types of glomeruli is also obtained through a fully connected layer. The MESC evaluation subnet uses DenseNet-121 as its backbone. The label prediction results are output by the fully connected layer, and the output categories are adjusted to correspond to four preset MESC labels. Each node of the fully connected layer corresponds to one preset MESC label.
8. A multimodal kidney pathology image processing device, characterized in that, The device includes: The acquisition module is used to acquire multimodal kidney pathology images to be processed, wherein the multimodal kidney pathology images to be processed include immunofluorescence images and PAS-stained pathology images corresponding to the immunofluorescence images; A classification module is used to input the immunofluorescence image into a pre-trained first network model, wherein the first network model is used to classify the input immunofluorescence image into kidney disease types to obtain kidney disease types; The judgment module is used to determine whether the kidney disease type output by the pre-trained first network model belongs to the preset kidney disease type; The identification module is used to input the PAS-stained pathological image corresponding to the immunofluorescence image into a pre-trained second network model when the kidney disease belongs to a preset type. The second network model is used to identify the pathological structure of the input PAS-stained pathological image and output a label for the identified pathological structure. The labeling module is used to add Oxford typing labels to the multimodal kidney pathology images to be processed based on the labels and pathological structures output by the pre-trained second network model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the multimodal kidney pathology image processing method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multimodal kidney pathology image processing method as described in any one of claims 1 to 7.
Citation Information
Cited By
IgA nephropathy survival analysis method and analysis system based on weak supervised learning
CN121616936A
Early-stage amyloid nephropathy recognition and diagnosis method based on deep learning
CN121686056A
Kidney pathological image processing method integrating multi-tissue segmentation and pathological change quantitative analysis
CN122312633A