Spine image segmentation method based on multi-map structure information and semantic description
By combining SAM2 model and multi-map technology, multimodal features of spine images are extracted and fused, the problems of spine horizontal plane image segmentation and classification recognition are solved, and high-precision spine segmentation and classification are achieved.
Patent Information
- Application Number
- CN202510084732.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to achieve accurate segmentation and accurate identification of spinal classification in spinal horizontal images, mainly because the SAM2 model performs poorly in medical image segmentation and the complex anatomical structure of the spinal region.
The spine image segmentation method based on multi-map structural information and semantic description is adopted, combining the multi-modal feature alignment capability of the SAM2 model, CLIP and the morphological structural features of the multi-map, features are extracted through the image encoder, semantic encoder and graph encoder, and these features are fused in the mask decoder to generate the final segmentation result.
The accuracy and recognition ability of the SAM2 model during spinal horizontal segmentation was significantly improved, achieving more accurate spinal area segmentation and correct classification recognition, exceeding the segmentation accuracy of other comparison methods.
Smart Images

Figure CN120014267A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical image technology, and specifically relates to a spinal image segmentation method based on multi-atlas structural information and semantic description. Background Art
[0002] The spine is an indispensable structure in the human body, which provides support for the human body. Accurate spinal segmentation results are very important for the treatment of spinal diseases. It can help doctors better understand the patient's spinal anatomy and better diagnose and treat the patient. However, accurate manual labeling of the spine is a very time-consuming and inefficient task. And due to the different personal experience of the labelers, manual labeling is prone to errors. This is also very difficult for automatic labeling methods based on deep learning. Due to the complex anatomical structure of the spinal region, the noise of medical images, the low contrast of the nerve area in CT, and other reasons, the segmentation results of the spinal region cannot be very accurate.
[0003] Multi-atlas segmentation is stable and accurate. With the development of deep learning, multi-atlas segmentation and deep learning have also been integrated to achieve an advantageous combination. The SAM2 model provides a basic model in the field of image segmentation and video segmentation. Its high accuracy and simplicity provide new ideas for segmentation. However, the training data of the SAM2 model is natural images, while medical images are mostly obtained by X-rays, magnetic resonance imaging, etc. Therefore, medical images and natural images have large differences in attributes such as pixel intensity and texture, so the SAM2 model does not perform well in medical image segmentation. In addition, for spinal images, their anatomical structure is very complex. This is especially reflected in the horizontal plane images of the spine. Due to the great similarity between different cones and different intervertebral discs, SAM2 currently has poor results in accurately segmenting horizontal plane images of the spine and identifying the correct spine classification. Therefore, multimodal information fusion is very important for completing the above tasks. It can provide the network with features from different perspectives, thereby helping accurate segmentation and recognition. CLIP achieves alignment of image modality and text modality through contrastive learning. It can accurately identify the feature alignment of input images and input text. Therefore, many multimodal models involving text features use CLIP's text encoder for text feature extraction. Our method also uses CLIP's pre-trained text encoder for semantic feature extraction to help the network identify which category the segmented spine structure belongs to.
[0004] Therefore, how to combine the prior knowledge of multiple atlases, the multimodal feature alignment capability of CLIP and the image feature extraction capability of the SAM2 model, improve the shortcomings of SAM2 and use it and the high accuracy of multiple atlases to complete the precise segmentation of the spinal region is the key task of the present invention. Summary of the invention
[0005] Purpose of the invention: The present invention proposes a spinal image segmentation method based on multi-atlas structural information and semantic description, which improves the problem that the SAM2 model has poor performance in segmenting and accurately identifying the correct classification when segmenting the horizontal plane of the spine.
[0006] Invention content: The present invention discloses a spinal image segmentation method based on multi-atlas structural information and semantic description, comprising the following steps:
[0007] (1) Use the Hiera model with SAM2 pre-trained weights to extract input image features and use it as an image encoder;
[0008] (2) Generate anatomical description text from multiple atlases and use the CLIP pre-trained text encoder to extract corresponding semantic features. These two parts are used as semantic encoders.
[0009] (3) Use a convolutional layer plus the same Hiera to extract the morphological structural features of multiple atlases and use it as an atlas encoder;
[0010] (4) Use LoRA to fine-tune the image encoder, semantic encoder, and atlas encoder to make them more suitable for the spine segmentation task;
[0011] (5) Use the mask decoder to fuse the features extracted by the three encoders to obtain the final prediction result.
[0012] Furthermore, the implementation process of step (1) is as follows:
[0013] The two-dimensional image is divided into multiple image blocks and position encoding is added. After passing through four stages of multi-scale modules, image features of different scales are extracted. The multi-scale modules of the four stages have 2, 6, 36 and 4 layers respectively. In the last multi-scale module of each stage, the Q matrix in the Transformer module is pooled to generate a feature map with reduced scale. Finally, the number of feature map channels of all scales is adjusted at the neck of the feature pyramid to finally generate image features of different scales.
[0014] The features extracted by the last multi-scale module in each stage are saved as The range of i is 1 to 4, representing the image features of different scales extracted in the four stages; the number of channels of these four features is adjusted to 256 in the neck of the feature pyramid, and the adjusted features are
[0015]
[0016] Among them, FPN is the feature pyramid neck, MSB is the multi-scale module, PatchEmbed is the patch embedding module, and X is the input image; and After upsampling, add:
[0017]
[0018] Among them, UP is the up-sampling operation; finally, give up.
[0019] Furthermore, the implementation process of step (2) is as follows:
[0020] The text generation module is used to generate the corresponding anatomical text description from the multi-atlas image. A corresponding description is generated for each category. The category that does not appear in the multi-atlas image is replaced by an empty string. The category, morphological features and category proportion of the category are described in detail in the text. Then the text description is converted into word embedding using a word segmenter, and the word embedding is converted into a word feature vector using a tag embedding. The final semantic features are generated through a 12-layer residual attention module. Finally, the morphological description features are projected into the same space of the image features through a projection layer. The above is achieved through the following formula:
[0021] text=TextGenerator(Y)
[0022] F T =Proj(RAB 1→12 (TokenEmbed(Tokenizer(text))))
[0023] Among them, Y is a multi-atlas prompt, TextGenerator is a text generation module, Proj is a projection layer, RAB is a residual attention module, TokenEmbed is a token embedding, and Tokenizer is a word segmenter.
[0024] Furthermore, the implementation process of step (3) is as follows:
[0025] The number of channels of multiple multi-atlases is adjusted to 3 through a convolution module to make it suitable for the number of input channels of subsequent Hiera; then the morphological structural features of the multi-atlas are extracted through the above-mentioned Hiera model; this is achieved through the following formula:
[0026]
[0027] Among them, FPN is the feature pyramid neck, MSB is the multi-scale module, PatchEmbed is the patch embedding module, Conv is the convolutional layer, Y is the multi-atlas prompt, and i ranges from 1 to 4.
[0028] Furthermore, the implementation process of step (4) is as follows:
[0029] Learnable LoRA weights are added to the Q and V matrices in the Transformer in the image encoder, semantic encoder, and graph encoder, and then trained, where the rank in LoRA is 8.
[0030] Furthermore, the implementation process of step (5) is as follows:
[0031] The semantic features are concatenated with multi-atlas features as category embedding vectors, and then updated with the input image features in the Transformer module; two layers of self-attention layer, token-to-image attention layer, MLP layer, and image-to-token attention layer are stacked in the Transformer module; the updated image features are The scale is enlarged through two stages of upsampling layers, and the same-scale image features and multi-atlas anatomical structure features are used in each stage, that is, the same-scale and With the same scale and Helps restore scale and finally obtain morphological features Updated category embedding vector After token-to-image attention and multi-layer MLP, category-related features are obtained Finally, after dot multiplication and linear interpolation, the final segmentation result is obtained; the above process is implemented by the following formula:
[0032]
[0033] Among them, concat represents the channel-by-channel concatenation operation, UP represents upsampling, T2IA represents token-to-image attention, seg represents the final segmentation result, and ⊙ represents dot product and linear interpolation operations.
[0034] Beneficial effects: Compared with the prior art, the present invention has the following beneficial effects: the present invention uses the anatomical prior knowledge contained in multiple atlases to make the segmentation result more accurate, and uses semantic features to allow the network to better understand and accurately identify the structure being segmented; the present invention improves the problem that the SAM2 model cannot accurately identify the segmentation category during horizontal plane segmentation of the spine, and uses the image encoder and atlas encoder from the SAM2 pre-trained weights to extract the corresponding features, and uses the text encoder of CLIP to extract the morphological description features, and finally fuses the above features in the mask decoder to generate the final segmentation result; the present invention greatly improves the problem that horizontal plane segmentation of the spine cannot accurately identify the category of the structure being segmented. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1It is a schematic diagram of the spine image segmentation network structure proposed by the present invention. DETAILED DESCRIPTION
[0036] The present invention will be further described below in conjunction with the accompanying drawings.
[0037] The present invention proposes a spine image segmentation method based on multi-atlas structural information and semantic description. Figure 1 The spinal image segmentation network shown in the figure takes as input the spinal image to be segmented and the corresponding multi-atlas image after registration, and outputs the segmentation result. In the network, the image to be segmented is input into the image encoder to extract high-dimensional image features. The anatomical text description is extracted from the multi-atlas prompt, and the semantic features are generated by the semantic encoder. The multi-atlas is input into the atlas encoder to extract the multi-atlas morphological features. Finally, the above features are fused in the mask decoder to produce the final segmentation result. Specifically, the following steps are included:
[0038] Step 1: Use the Hiera model with SAM2 pre-trained weights as the image encoder to extract input image features.
[0039] In the image encoder, the two-dimensional image is divided into multiple image blocks and position encoding is added. After passing through 4 stages of multi-scale modules, image features of different scales are extracted; the 4 stages have 2, 6, 36, and 4 layers of multi-scale modules respectively. In the last multi-scale module of each stage, the Q matrix in the Transformer module is pooled to generate a reduced-scale feature map; finally, the number of feature map channels of all scales is adjusted at the neck of the feature pyramid to finally generate image features of different scales. Specifically, the features extracted by the last multi-scale module of each stage are saved as The range of i is 1 to 4, representing the image features of different scales extracted in the four stages. The number of channels of these four features is adjusted to 256 in the neck of the feature pyramid, and the adjusted features are In addition, the adjusted and After upsampling, add UP is the up-sampling operation), and finally give up.
[0040] The whole process can be described as:
[0041]
[0042] Among them, FPN is the feature pyramid neck, MSB is the multi-scale module, PatchEmbed is the patch embedding module, X is the input image, and i ranges from 1 to 4.
[0043] Step 2: Use the text generation module to generate corresponding anatomical text descriptions from multi-atlas images.
[0044] A corresponding description is generated for each category. For categories that do not appear in the multi-atlas image, an empty string is used instead of the description. The text specifically describes the category, morphological features, and category proportion of the category; then the text description is converted into word embedding using a word segmenter, and the word embedding is converted into a word feature vector using a tag embedding; the final semantic features are generated after passing through a 12-layer residual attention module; finally, the semantic features are projected into the same space of the image features through a projection layer to facilitate subsequent feature fusion.
[0045] The specific anatomy is described as:
[0046] For the background class: In this task, please focus on identifying different classes and segmenting the key areas in the spine image. The background part consists of irrelevant structures surrounding the spine. After excluding the background class, the remaining classes to be segmented are described as follows.
[0047] For Nerves: Identify and segment nerve roots, which are long, thin structures branching out from the spinal cord. Segmentation is challenging due to their thin and fragile nature, especially in areas with close proximity to surrounding tissue or where motion artifacts may obscure their outlines. In this image, the nerve roots account for approximately {ratio}% of the total segmented area.
[0048] For the pyramidal class: Identify and segment the {Pyramidal Name} vertebra, which has a cylindrical appearance and contains significant anatomical landmarks. The main challenge is to accurately delineate the boundaries, especially in the area where the vertebra contacts adjacent tissue, where low contrast and overlapping structures may cause blurred contours. In this image, this vertebra occupies approximately {ratio}% of the total segmented area.
[0049] For the intervertebral disc class: Identify and segment the intervertebral disc located between the {cone name} vertebrae and the {cone name} vertebrae. The disc is round and has uneven thickness, which makes segmentation challenging, especially when distinguishing from the surrounding muscle and fat tissue, which may have similar intensity values in the image. In this image, the disc occupies approximately {ratio}% of the total segmented area.
[0050] The whole process can be described as:
[0051] text=TextGenerator(Y)
[0052] F T =Proj(RAB 1→12 (TokenEmbed(Tokenizer(text))))
[0053] Among them, Y is the multi-atlas prompt, TextGenerator is the text generation module, Proj is the projection layer, RAB is the residual attention module, TokenEmbed is the token embedding, and Tokenizer is the word segmenter.
[0054] Step 3: Use a convolutional layer plus the same Hiera to extract anatomical features of multiple atlases.
[0055] The number of channels of multiple multi-atlases is adjusted to 3 through a convolution module to make it suitable for the number of input channels of the subsequent Hiera; then the multi-atlas features are extracted through the above-mentioned Hiera.
[0056] The whole process can be described as:
[0057]
[0058] Among them, FPN is the feature pyramid neck, MSB is the multi-scale module, PatchEmbed is the patch embedding module, Conv is the convolutional layer, Y is the multi-atlas prompt, and i ranges from 1 to 4.
[0059] Step 4: Use LoRA to fine-tune the above three encoders to make them more suitable for the spine segmentation task. Use LoRA to fine-tune the Transformer in the image encoder, semantic encoder, and atlas encoder. Specifically, add learnable LoRA weights to the Q and V matrices in the Transformer, and then train them with a rank in LoRA of 8.
[0060] Step 5: Use the mask decoder to fuse the features extracted by the three encoders to obtain the final prediction result.
[0061] The semantic features are concatenated with multi-atlas features as category embedding vectors, and then updated with the input image features in the Transformer module. Two layers of self-attention layer, token-to-image attention layer, MLP layer, and image-to-token attention layer are stacked in the Transformer module. The updated image features, i.e. The scale is enlarged through two stages of upsampling layers, and the same-scale image features and multi-atlas anatomical structure features (i.e., the same-scale and With the same scale and ) helps scale recovery and finally obtains the morphological features, namely The updated category embedding vector is After token-to-image attention and multi-layer MLP, category-related features are obtained, namely Finally, after dot multiplication and linear interpolation, the final segmentation result is obtained.
[0062] The whole process can be described as:
[0063]
[0064]
[0065] Among them, concat represents the channel-by-channel concatenation operation, UP represents upsampling, T2IA represents token-to-image attention, seg represents the final segmentation result, and ⊙ represents dot product and linear interpolation operations.
[0066] The present invention is verified on the spine dataset and the neural dataset, wherein the spine dataset is 15 categories: background, neural, S, S / L5, L5, L5 / L4, L4, L4 / L3, L3, L3 / L2, L2, L2 / L1, L1, L1 / T12 and T12, and the size of each 2D picture is 512*512; the neural dataset is two categories: background, neural, and the size of each 2D picture is 160*320. The present invention is compared with U-Net, CE-Net, nnU-Net, TransUNet, Swin-UNet, SAM2_point, SAM2_box and SAM2_mask. The results on the spine dataset are shown in Table 1. The segmentation accuracy of the present invention is better than all other comparison methods (i.e., DICE=0.9298, IoU=0.8739, ASD=1.0979, HD95=3.9261).
[0067] Table 1 Segmentation results on the spine dataset
[0068]
[0069] The results on the neural data set are shown in Table 2. The segmentation accuracy of the present invention is better than all other comparison methods (i.e., DICE=0.8740, IoU=0.7822, ASD=0.8467, HD95=4.1696). From the above results, it can be seen that the present invention fully integrates the morphological structural features and anatomical semantic features extracted from multiple atlases, so the segmentation results exceed all other comparison methods.
[0070] Table 2 Segmentation results on neural datasets
[0071]
[0072]
[0073] The present invention realizes accurate segmentation of horizontal plane images of the spine based on multi-atlas and SAM model. The present invention uses three encoders to extract features from the image to be segmented, text and multi-atlas, and fuses features through a mask decoder to finally produce high-resolution segmentation results. Due to the use of multi-atlas morphological prior knowledge and SAM2 framework, the method has high accuracy, and the integration of anatomical semantic features enables the network to accurately identify the correct category of the segmented area.
[0074] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A spine image segmentation method based on multi-atlas structural information and semantic description, characterized in that: The following steps are involved: (1) Use the Hiera model with SAM2 pre-trained weights to extract input image features and use it as an image encoder; (2) Generate anatomical description text from multiple atlases and use the CLIP pre-trained text encoder to extract corresponding semantic features. These two parts are used as semantic encoders. (3) Use a convolutional layer plus the same Hiera to extract the morphological structural features of multiple atlases and use it as an atlas encoder; (4) Use LoRA to fine-tune the image encoder, semantic encoder, and atlas encoder to make them more suitable for the spine segmentation task; (5) Use the mask decoder to fuse the features extracted by the three encoders to obtain the final prediction result.
2. The spine image segmentation method based on multi-atlas structural information and semantic description according to claim 1, characterized in that: The implementation process of step (1) is as follows: The two-dimensional image is divided into multiple image blocks and position codes are added. After passing through four stages of multi-scale modules, image features of different scales are extracted. In the last multi-scale module of each stage, the Q matrix in the Transformer module is pooled to generate a feature map with reduced scale. Finally, the number of feature map channels of all scales is adjusted at the neck of the feature pyramid to finally generate image features of different scales. The features extracted by the last multi-scale module in each stage are saved as Where i ranges from 1 to 4, representing image features of different scales extracted in the four stages; The number of channels of these four features in the feature pyramid neck is adjusted to 256, and the adjusted features are Among them, FPN is the feature pyramid neck, MSB is the multi-scale module, PatchEmbed is the patch embedding module, and X is the input image; and After upsampling, add: Among them, UP is the up-sampling operation; finally, give up.
3. The spine image segmentation method based on multi-atlas structural information and semantic description according to claim 1, characterized in that: The multi-scale modules of the four stages have 2, 6, 36 and 4 layers respectively.
4. The spine image segmentation method based on multi-atlas structural information and semantic description according to claim 1, characterized in that: The implementation process of step (2) is as follows: The text generation module is used to generate the corresponding anatomical text description from the multi-atlas image. A corresponding description is generated for each category. The category that does not appear in the multi-atlas image is replaced by an empty string. The category, morphological features and category proportion of the category are described in detail in the text. Then the text description is converted into word embedding using a word segmenter, and the word embedding is converted into a word feature vector using a tag embedding. The final semantic features are generated through a 12-layer residual attention module. Finally, the morphological description features are projected into the same space of the image features through a projection layer. The above is achieved through the following formula: text=TextGenerator(Y) F T =Proj(RAB 1→12 (TokenEmbed(Tokenizer(text)))) Among them, Y is a multi-atlas prompt, TextGenerator is a text generation module, Proj is a projection layer, RAB is a residual attention module, TokenEmbed is a token embedding, and Tokenizer is a word segmenter.
5. The spine image segmentation method based on multi-atlas structural information and semantic description according to claim 1, characterized in that: The implementation process of step (3) is as follows: The number of channels of multiple multi-atlases is adjusted through a convolution module to make it suitable for the number of input channels of subsequent Hiera; then the morphological structural features of the multi-atlas are extracted through the above-mentioned Hiera model; This is achieved through the following formula: Among them, FPN is the feature pyramid neck, MSB is the multi-scale module, PatchEmbed is the patch embedding module, Conv is the convolutional layer, Y is the multi-atlas prompt, and i ranges from 1 to 4.
6. The method for segmenting spinal images based on multi-atlas structural information and semantic description according to claim 5, characterized in that: The number of channels is 3.
7. The method for segmenting spinal images based on multi-atlas structural information and semantic description according to claim 1, characterized in that: The implementation process of step (4) is as follows: Learnable LoRA weights are added to the Q and V matrices in the Transformer in the image encoder, semantic encoder, and graph encoder, and then trained, where the rank in LoRA is 8.
8. The method for segmenting spinal images based on multi-atlas structural information and semantic description according to claim 1, characterized in that: The implementation process of step (5) is as follows: The semantic features are concatenated with multi-atlas features as category embedding vectors, and then updated with the input image features in the Transformer module; two layers of self-attention layer, token-to-image attention layer, MLP layer, and image-to-token attention layer are stacked in the Transformer module; the updated image features are The scale is enlarged through two stages of upsampling layers, and the same-scale image features and multi-atlas anatomical structure features are used in each stage, that is, the same-scale and With the same scale and Helps restore scale and finally obtain morphological features Updated category embedding vector After token-to-image attention and multi-layer MLP, category-related features are obtained Finally, after dot multiplication and linear interpolation, the final segmentation result is obtained; the above process is implemented by the following formula: Among them, concat represents the channel-by-channel concatenation operation, UP represents upsampling, T2IA represents token-to-image attention, seg represents the final segmentation result, and ⊙ represents dot product and linear interpolation operations.