A hierarchical multi-scale graph distillation method for multi-organ segmentation

CN122821110APending Publication Date: 2026-09-25CHONGQING NANPENG ARTIFICIAL INTELLIGENCE TECH RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610782969.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

1)构建多分辨率金字塔提取不同尺度特征,并利用图结构建模不同尺度间的关系,实现多尺度信息的高效融合与交互,解决了复杂医学图像中难以兼顾全局语义与局部细节的问题

Benefits of technology

本发明通过构建多分辨率金字塔并利用图结构建模不同尺度间的关系,实现了多尺度信息的高效融合与交互,有效解决了复杂医学图像中难以兼顾全局语义与局部细节的问题,为医学图像分割提供了更全面、更精确的特征表达。同时,基于动态边连接构建的方法优化了特征表示方式,结合实例级动态注意力机制与跨层节点关联,显著提升了图结构信息的表达能力和特征捕捉能力,使模型能够自适应调整不同通道的重要性,增强了对医学图像中不同器官特征的区分能力。在层次化知识蒸馏框架下,通过构建教师网络与学生网络并利用格拉姆矩阵衡量特征分布的相似性,实现了知识迁移的优化,使分割边界更加精确并降低了计算复杂度,同时通过设计总损失函数结合DICE损失用于前景分割以及TGD损失用于提升学生网络性能,进一步提升了模型的泛化能力。采用AdamW优化器进行训练,保证了训练的稳定性和收敛效率。实验结果表明,本发明在多个公开医学图像数据集上均取得了优于现有技术的分割效果,显著提升了医学图像多器官分割的精度与泛化能力,为医学图像分析提供了高效、精确的解决方案,具有重要的临床应用价值和推广意义,为复杂医学图像的分析与处理提供了可靠的技术支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821110A_ABST
    Figure CN122821110A_ABST
Patent Text Reader

Abstract

The application provides a hierarchical multi-scale graph distillation method for multi-organ segmentation. Through the deep integration of multi-scale feature modeling, graph structure interaction and knowledge distillation mechanism, the method aims to improve the accuracy of medical image segmentation. The method first performs multi-resolution pyramid processing on the input medical image, extracts image blocks at different scales as the basic unit, and constructs a teacher network to realize multi-level feature extraction. At the local level, the method processes multi-scale graph structures through graph neural network (GNN) layer by layer. The first layer GNN models the spatial relationship between nodes for a single resolution graph to capture the local morphological features of organs. The second layer GNN realizes multi-scale information interaction through cross-resolution graph connection to solve the problems of organ size difference and boundary ambiguity. To further enhance the representation of key regions, the method uses an instance-level dynamic attention mechanism to assign feature weights to image blocks, and combines residual connection and channel attention modules to realize differentiated focusing of foreground-background features. In the knowledge distillation framework, the method innovatively introduces a hierarchical distillation path. The high-resolution network transfers spatial detail features to the low-resolution network through transform graph distillation, and at the same time, uses a self-distillation mechanism to migrate semantic information in the multi-scale fusion module (MSF) to the main network. This process realizes cross-scale knowledge transfer through feature mapping Gram matrix alignment and attention graph matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of medical image processing and artificial intelligence, specifically relating to a hierarchical multi-scale graph distillation method for multi-organ segmentation. This method significantly improves the accuracy and robustness of multi-organ segmentation in medical images through the deep integration of multi-resolution pyramid feature extraction, graph neural network-driven multi-scale interactive modeling, instance-level dynamic attention mechanisms, and hierarchical knowledge distillation techniques. It is particularly suitable for solving key problems such as significant differences in organ size and blurred boundaries. Background Technology

[0002] Multi-resolution pyramid feature extraction, graph neural network-driven multi-scale interactive modeling, instance-level dynamic attention mechanism, and hierarchical knowledge distillation technology have important applications in fields such as medical image analysis. They can effectively improve the model's feature representation ability, cross-scale information interaction ability, and generalization ability.

[0003] Multi-resolution pyramid feature extraction is an important multi-scale modeling method. Through hierarchical downsampling or feature fusion, it enables the model to simultaneously focus on global semantic information and local detail information. In medical image segmentation tasks, organs vary significantly in scale. Large organs such as the liver require a stronger global receptive field, while small structures like blood vessels rely on high-resolution details. Therefore, constructing multi-resolution pyramid features can effectively enhance the model's ability to perceive targets at different scales. Graph neural network-driven multi-scale interactive modeling maps images or features to graph structures and utilizes methods such as Graph Convolutional Neural Networks (GCNs) and Graph Attention Networks (GATs) to achieve information interaction between features at different scales. Compared to traditional convolutional operations, graph neural networks can better model the relational information in non-Euclidean space, enabling effective aggregation of features across scales. For example, in multi-organ segmentation tasks, organs such as the liver, spleen, and kidneys have clear spatial topological relationships. Traditional CNNs struggle to explicitly model these relationships, while graph neural network-based methods can utilize neighborhood information for structural perception. Instance-level dynamic attention mechanisms are mainly used to adaptively model different instances (such as multiple organs or lesions) in segmentation tasks to improve the model's discriminative ability. Traditional attention mechanisms (such as channel attention and spatial attention) are often globally shared, making it difficult to adaptively adjust for different instances. Instance-level dynamic attention can dynamically adjust the attention allocation strategy based on the morphology, location, and texture features of different targets, enabling the model to focus on key information more accurately. For example, in multi-organ segmentation tasks, the morphology of different organs may vary significantly among different patients. Fixed attention weights cannot adapt to these variations, while instance-level dynamic attention mechanisms can adjust weights according to specific instances, thereby improving the accuracy and robustness of segmentation. Hierarchical knowledge distillation technology, through a teacher-student network framework, transfers knowledge from complex models to student models to improve the model's generalization ability. Traditional knowledge distillation mainly focuses on the information transfer of the output layer or a single feature layer, while hierarchical knowledge distillation achieves finer-grained knowledge transfer through the alignment of information at multiple levels. In multi-organ segmentation tasks, features at different levels contain information at different scales; for example, shallow features contain more edge details, while deep features contain richer semantic information. Hierarchical knowledge distillation can effectively combine information from these different levels, enabling the student model to retain key features while possessing better generalization ability. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and propose a hierarchical multi-scale graph distillation method. Through the synergistic optimization of multi-resolution graph structure modeling, instance-level dynamic attention mechanism and hierarchical distillation path, the accuracy and generalization ability of multi-organ segmentation in medical images are significantly improved.

[0005] The present invention achieves the above objectives through the following technical solutions: This invention includes the following steps: 1) A multi-resolution pyramid is constructed to extract features at different scales, and a graph structure is used to model the relationships between different scales, achieving efficient fusion and interaction of multi-scale information. This solves the problem of balancing global semantics and local details in complex medical images. The specific implementation steps are as follows: a) Representing medical images as ,in The height of the image. For width, This represents the number of channels.

[0006] b) Using the medical image defined in step a) as input to step b), perform bilinear interpolation on... Perform multi-level downsampling to recursively generate multiple image layers with different resolutions: in, It is a scale sequence. It is the original image. It is the first Level downsampled image, Original resolution , express Downsampling, size is , express Downsampling, size is And so on. This step can gradually reduce the resolution while preserving the main structural information of medical images, providing a foundation for multi-scale feature extraction.

[0007] c) The resolution pyramid can be obtained from step b). , where L is the number of layers in the pyramid.

[0008] d) Images of each layer Divided into overlapping image blocks The block size is The overlap rate is 15%. Number of blocks. The calculation formula is: in, and For the first The height and width of the layer graphic.

[0009] e) Employ a pre-trained encoder, remove the global pooling layer, retain the output of the last convolutional layer, and apply this to each image patch. Feature extraction: in For encoder parameters, Indicates the first Layer The feature vector of each node.

[0010] f) Perform L2 normalization on the feature vectors extracted in e) to eliminate dimensional differences: in, is the numerical stability constant.

[0011] 2) The method based on dynamic edge connections optimizes feature representation to achieve differentiated processing of different node attributes, effectively improving the expressive power of structural information. Specifically, it includes the following steps: a) Calculate feature similarity. Define the output of step f) in 1) as a node, and calculate the cosine similarity between nodes: b) Constrain the spatial proximity between nodes: in, For nodes Spatial coordinates ( For row index, (for column indexes), The maximum neighborhood radius, This is the similarity threshold.

[0012] c) For adjacent layers (such as the first) Layer and first The node coordinates of the layer are aligned by bilinear interpolation, and a sparse adjacency matrix is ​​constructed based on the nodes associated across layers.

[0013] d) Use the output of step c) as input to a graph neural network, employing a 4-head attention mechanism, with each head having a dimension of 128. For nodes... and his neighbors Calculate attention weights: in, For learnable parameter vectors, Let || be the projection matrix, and || denotes vector concatenation. Then, feature aggregation is performed to update the node features.

[0014] e) The output of d) is used to perform feature aggregation using a learnable transformation matrix, and finally the local and cross-layer features are concatenated and the dimensionality is reduced.

[0015] f) Channel weights are calculated using a two-layer MLP and normalized using Sigmoid, enabling the model to adaptively adjust the importance of different channels and improve feature representation capabilities.

[0016] 3) Within the hierarchical knowledge distillation framework, knowledge transfer methods can be used to optimize the partition boundary. This method includes the following key steps: a) Construct a teacher network containing a complete pyramid structure, including all graph neural network layers and attention mechanisms. b) Construct a student network that has one less pyramid structure layer than the teacher network in a), removing the highest resolution layer. c) Calculate the Gram matrix of the feature maps of the teacher network and the student network: in, Indicates the first Feature map of the layer The Gram matrix represents the number of channels and is used to measure the similarity of feature distributions.

[0017] d) Design a loss function to optimize the student network by calculating the difference in the Gram matrix between the teacher and student networks: in, and These represent the feature graphs of the teacher and student networks, respectively. This represents the Frobenius norm.

[0018] e) Design the total loss function, incorporating DICE loss. Foreground segmentation is used, along with TGD loss to improve student network performance. The total loss function is shown below: in, and The parameters are learnable, and the AdamW optimizer is used for training with a learning rate of 100%. The weight decays to This ensures training stability.

[0019] The beneficial effects of this patent are as follows: This invention constructs a multi-resolution pyramid and utilizes graph structures to model the relationships between different scales, achieving efficient fusion and interaction of multi-scale information. This effectively solves the problem of balancing global semantics and local details in complex medical images, providing a more comprehensive and accurate feature representation for medical image segmentation. Simultaneously, a method based on dynamic edge connections optimizes feature representation. Combined with instance-level dynamic attention mechanisms and cross-layer node associations, it significantly enhances the expressive power and feature capture capabilities of graph structure information, enabling the model to adaptively adjust the importance of different channels and improve the ability to distinguish features of different organs in medical images. Within a hierarchical knowledge distillation framework, by constructing teacher and student networks and using Gram matrices to measure the similarity of feature distributions, knowledge transfer is optimized, resulting in more accurate segmentation boundaries and reduced computational complexity. Furthermore, by designing a total loss function combined with DICE loss for foreground segmentation and TGD loss to improve student network performance, the model's generalization ability is further enhanced. The AdamW optimizer is used for training, ensuring training stability and convergence efficiency. Experimental results show that the present invention achieves better segmentation results than existing technologies on multiple publicly available medical image datasets, significantly improving the accuracy and generalization ability of multi-organ segmentation in medical images. It provides an efficient and accurate solution for medical image analysis, has important clinical application value and promotion significance, and provides reliable technical support for the analysis and processing of complex medical images. Attached Figure Description

[0020] Figure 1 Flowchart of Multi-Scale Discrete Wavelet Transform (DWT) for Low-Light Road Images Figure 2 Schematic diagram of diffusion model based on frequency analyzer Figure 3 Schematic diagram of a frequency analyzer.

[0021] Figure 4 Flowchart of Conditional Random Fields (CRF) for Improving Segmentation Performance Detailed Implementation

[0022] like Figure 1 , 2 As shown in Figures 1 and 3, the specific implementation details of each part of the present invention are as follows: 1. First, the input medical image is represented as a three-dimensional matrix, containing three dimensions: height, width, and number of channels. To extract multi-scale features, the original image undergoes multi-level downsampling to generate a resolution pyramid. Specifically, the original image is downsampled level by level using bilinear interpolation to generate a series of image layers with different resolutions. For example, the first layer is the original image, the second layer is half the resolution of the original image, the third layer is one-quarter the resolution, and so on. This multi-level downsampling method can gradually reduce the resolution while preserving the main structural information of the image, providing a foundation for subsequent multi-scale feature extraction.

[0023] Each image layer is divided into overlapping image patches, each with a fixed size and a 15% overlap between patches. This partitioning ensures that local details are not missed, while the overlapping regions facilitate smooth transitions between patches. A pre-trained encoder then extracts feature vectors from each patch. The output of the encoder's last convolutional layer is retained, while the global pooling layer is removed to ensure that the extracted features preserve spatial information. Finally, L2 normalization is applied to the extracted feature vectors to eliminate dimensional differences between features, making subsequent graph structure modeling more stable and efficient.

[0024] 2. Using the feature vectors extracted in step 1 as nodes, calculate the cosine similarity between each pair of nodes to measure their similarity in the feature space. To further optimize the graph structure, constrain the spatial proximity between nodes; only nodes within a certain spatial range are considered potential neighbors. This constraint reduces computational complexity while avoiding the introduction of irrelevant node information. Next, bilinear interpolation is used to align the node coordinates of different layers to ensure the correlation between nodes across layers. Based on the aligned node coordinates and feature similarity, a sparse adjacency matrix is ​​constructed to represent the edge connections in the graph structure. Subsequently, the sparse adjacency matrix is ​​used as input to construct a graph neural network. The graph neural network employs a multi-head attention mechanism, with each head having a fixed dimension for calculating attention weights between nodes. Through these attention weights, the features of neighboring nodes are weighted and aggregated to update the feature representation of the current node. Finally, a learnable transformation matrix is ​​used to further fuse the aggregated features, and local features are concatenated with cross-layer features. Dimensionality reduction is then used to obtain the final node feature representation.

[0025] 3. In the hierarchical knowledge distillation framework, a complete teacher network is first constructed, containing all pyramid levels, graph neural network layers, and attention mechanisms. The teacher network learns complex multi-scale feature relationships, providing a high-quality knowledge source for subsequent knowledge transfer. Next, a simplified student network is constructed, which has one less pyramid layer than the teacher network (i.e., the highest resolution layer is removed). This design reduces computational complexity while maintaining performance. To transfer knowledge from the teacher network to the student network, the Gram matrix of the feature maps of both networks is calculated. The Gram matrix measures the similarity of feature distributions and effectively captures high-order relationships between features. Based on the differences in the Gram matrices, a loss function is designed to optimize the student network. Specifically, the Frobenius norm between the Gram matrices of the teacher and student networks is calculated to measure the difference in feature distributions and is used as the optimization objective for knowledge distillation. Furthermore, a total loss function is designed by combining the DICE loss function for foreground segmentation and the TGD loss function based on the Gram matrix. The total loss function combines the DICE and TGD losses in a weighted manner to ensure that the student network achieves optimal segmentation accuracy and knowledge transfer performance. Finally, the AdamW optimizer was used to train the model, with appropriate learning rate and weight decay parameters set to ensure training stability and convergence.

[0026] The foregoing description illustrates the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. Represent medical images as multi-channel matrices and perform multi-level downsampling through bilinear interpolation to generate multi-resolution pyramids, preserving the main structural information of the images and providing a foundation for subsequent multi-scale feature extraction.

2. Based on the multi-resolution pyramid generated in step 1), each layer of the image is divided into overlapping image blocks. A pre-trained encoder is used to extract the feature vector of each image block and L2 normalization is performed to eliminate the difference in dimensions.

3. Based on the normalized feature vectors extracted in step 2), calculate the cosine similarity between nodes, and combine it with spatial proximity constraints to construct a sparse adjacency matrix to optimize the association between nodes.

4. Based on step 3), perform bilinear interpolation to align the node coordinates of adjacent layers, establish cross-layer associations, and further enhance the interaction between multi-scale features.

5. Use the sparse adjacency matrix generated in step 4) as input to the graph neural network, use the multi-head attention mechanism to calculate the attention weights between nodes, and perform feature aggregation to improve feature representation ability.

6. Based on the features aggregated in step 5), local and cross-layer features are concatenated and dimensionality reduced using a learnable transformation matrix to further optimize the feature representation.

7. Construct teacher and student networks. The teacher network contains a complete pyramid structure, and the student network has the highest resolution layer removed. Optimize the performance of the student network through knowledge distillation.

8. Based on the teacher network and student network constructed in step 7), calculate the Gram matrix of their feature maps, and optimize the student network by the difference in Gram matrix to improve the similarity of feature distribution.

9. Design a total loss function, combine the DICE loss and the Gram matrix difference loss in step 8), and use the AdamW optimizer for training to achieve a significant improvement in the accuracy and generalization ability of multi-organ segmentation in medical images.