A paraspinal muscle muscle map image segmentation method and apparatus
By fine-tuning the MedSAM model and introducing the Unet module, the problem of low accuracy caused by blurred boundaries and large morphological differences in paraspinal muscle MRI image segmentation was solved, achieving high-precision and stable automated quantitative analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY)
- Filing Date
- 2026-01-28
- Publication Date
- 2026-06-23
Smart Images

Figure CN122265307A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of medical image processing technology, specifically relating to a method and apparatus for segmenting paraspinal muscle images. Background Technology
[0002] The paraspinal muscles play a crucial role in maintaining spinal stability and lumbar and back motor function, and their degeneration is closely related to lumbar spine diseases, chronic low back pain, and muscle atrophy. Magnetic resonance imaging (MRI), with its excellent soft tissue resolution, is widely used to assess changes in the structure and composition of the paraspinal muscles. By calculating imaging indicators such as muscle cross-sectional area (CSA), fat cross-sectional area (fCSA), and fat infiltration rate (FI), the degree of muscle degeneration and fat infiltration can be objectively quantified. Traditional methods of manually delineating muscle regions are not only time-consuming and labor-intensive but also highly subjective and inconsistent. While deep learning methods have advantages in segmentation accuracy, in medical imaging, due to difficulties in data acquisition and high annotation costs, datasets are usually limited in size. Furthermore, the paraspinal muscles exhibit significant morphological variations, blurred boundaries, and complex tissue textures. These methods are still prone to overfitting under different muscle groups and imaging conditions, making it difficult to maintain consistency and high accuracy in segmentation results. Summary of the Invention
[0003] This application provides a method and apparatus for segmenting paraspinal muscle images, which can be used to achieve automatic segmentation of paraspinal muscle images.
[0004] A first aspect of this application provides a method for segmenting paraspinal muscle images, the method comprising: The original image segmentation model is trained using sample paraspinal muscle images to obtain the trained original image segmentation model; the original image segmentation model includes at least: an image encoder, a cue encoder, and a mask decoder; the mask decoder includes: a two-layer Transformer module, a two-layer upsampling module, and a multilayer perceptron; The mask decoder in the trained original image segmentation model is improved by replacing the double-layer upsampling module in the mask decoder with the Unet module to obtain the improved image segmentation model. With the parameters of the image encoder and the cue encoder frozen, the improved image segmentation model is trained using the sample paraspinal muscle images to obtain the target image segmentation model. The original paraspinal muscle image to be segmented is input into the target image segmentation model to obtain the segmented paraspinal muscle image.
[0005] A second aspect of this application also provides a paraspinal muscle image segmentation apparatus, used to perform the paraspinal muscle image segmentation method described in the first aspect of this application, the apparatus comprising: The first training module is used to train the original image segmentation model using sample paraspinal muscle images to obtain the trained original image segmentation model; the original image segmentation model includes at least: an image encoder, a cue encoder, and a mask decoder; the mask decoder includes: a two-layer Transformer module, a two-layer upsampling module, and a multilayer perceptron; An improvement module is used to improve the mask decoder in the trained original image segmentation model, so as to replace the double-layer upsampling module in the mask decoder with the Unet module to obtain the improved image segmentation model. The second training module is used to train the improved image segmentation model using the sample paraspinal muscle images under the condition of freezing the parameters of the image encoder and the cue encoder, so as to obtain the target image segmentation model. The segmentation module is used to input the original paraspinal muscle image to be segmented into the target image segmentation model to obtain the segmented paraspinal muscle image.
[0006] A third aspect of this application also provides an electronic device, including a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the paraspinal muscle image segmentation method as described in the first aspect.
[0007] A fourth aspect of this application also provides a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the paraspinal muscle image segmentation method as described in the first aspect.
[0008] The beneficial effects of this application are as follows: The paraspinal muscle image segmentation method proposed in this application is based on a fine-tuned large-scale medical imaging model MedSAM (i.e., the original image segmentation model including an image encoder, a cue encoder, and a mask decoder), and a small Unet module is introduced into its decoder part (replacing the double-layer upsampling module in the mask decoder with the Unet module), realizing targeted feature extraction and fine-grained segmentation. This structure effectively solves the problem of inaccurate segmentation caused by the blurred boundaries, large morphological differences, and complex tissue texture of the paraspinal muscle in MRI images, and effectively improves the accuracy and stability of image segmentation.
[0009] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. It should be noted that the scale in the drawings is for illustration only and does not represent the actual scale.
[0011] Figure 1 This is a flowchart of the steps of a paraspinal muscle image segmentation method in an embodiment of this application; Figure 2 This is a schematic diagram of the model architecture of a primitive image segmentation model in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a mask decoder before the improvement in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of an improved image segmentation model in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a Unet module in one embodiment of this application; Figure 6 This is a schematic diagram of the structure of a paraspinal muscle image segmentation device according to an embodiment of this application; Figure 7 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0012] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or at least two. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0014] The paraspinal muscles play a crucial role in maintaining spinal stability and lumbar and back mobility, and their degeneration is closely related to lumbar spine diseases, chronic low back pain, and muscle atrophy. Magnetic resonance imaging (MRI), with its excellent soft tissue resolution, is widely used to assess changes in the structure and composition of the paraspinal muscles. By calculating imaging parameters such as muscle cross-sectional area (CSA), fat cross-sectional area (fCSA), and fat infiltration rate (FI), the degree of muscle degeneration and fat infiltration can be objectively quantified. Traditional methods of manually delineating muscle regions are not only time-consuming and labor-intensive but also highly subjective and inconsistent. Deep learning-based medical image segmentation methods (such as U-Net and its variants) have achieved significant results in medical image segmentation tasks. These methods can effectively extract multi-scale spatial features, thereby achieving accurate depiction of anatomical structures. However, due to the difficulty in acquiring medical data, the high cost of annotation, and the highly variability in morphology, blurred boundaries, and complex tissue textures of the paraspinal muscles, these methods, such as convolutional neural network (CNN) methods, are still prone to overfitting under different muscle groups and imaging conditions, making it difficult to maintain the consistency and accuracy of segmentation results.
[0015] The Segment Anything Model (SAM) is a large-scale interactive image segmentation model that segments any image based on cues (points, boxes, text). Through large-scale visual data pre-training and a prompt-based segmentation mechanism, it exhibits superior zero-shot and transfer learning capabilities. SAM utilizes a powerful image encoder and prompt encoder to transform user input (points, boxes, or masks) into explicit guiding signals, enabling the model to achieve fast segmentation of arbitrary targets. Its broad generalization performance opens up new directions for general visual segmentation. However, since most of SAM's pre-training data consists of natural scene images, it clearly lacks the ability to understand images with medical backgrounds.
[0016] The Medical Segment Anything Model (MedSAM) transfers the SAM framework to the medical imaging domain. By re-pre-training on large-scale multimodal medical datasets (CT, MRI, ultrasound, etc.), the model gains enhanced medical semantic understanding capabilities. However, the original MedSAM's mask decoder structure is relatively shallow, limiting its ability to recover details. This is especially true in high-resolution Magnetic Resonance Imaging (MRI) images, where the complexity of muscle boundaries and grayscale transitions place higher demands on the decoding module. Furthermore, MedSAM's segmentation mechanism relies on user-provided box prompts as explicit guidance signals. While this feature ensures flexibility and interactivity in segmentation, in medical applications requiring batch processing or fully automated analysis, the prompt-dependent mechanism significantly reduces the system's automation level and processing efficiency.
[0017] In view of the aforementioned problems (existing paraspinal muscle MRI image segmentation methods have low accuracy and difficulty in achieving stable automated quantitative analysis when there are blurred boundaries, large morphological differences, and complex muscle textures), this application provides a paraspinal muscle image segmentation method and apparatus for achieving automatic segmentation of paraspinal muscle images. Specifically, the paraspinal muscle image segmentation method proposed in this application is based on a large medical imaging model MedSAM (i.e., the original image segmentation model including an image encoder, a cue encoder, and a mask decoder) with fine-tuning, and introduces a small Unet module in its decoder part (replacing the double-layer upsampling module in the mask decoder with the Unet module), realizing targeted feature extraction and fine-grained segmentation. This structure effectively solves the problem of inaccurate segmentation caused by blurred boundaries, large morphological differences, and complex tissue textures of the paraspinal muscle in MRI images, and effectively improves the accuracy and stability of image segmentation.
[0018] The first aspect of this application proposes a method for segmenting paraspinal muscle images. (Refer to...) Figure 1 , Figure 1 A flowchart illustrating the steps of a paraspinal muscle image segmentation method is shown, as follows: Figure 1 As shown, the method includes: Step S101: Train the original image segmentation model using sample paraspinal muscle images to obtain the trained original image segmentation model; the original image segmentation model includes at least: an image encoder, a cue encoder, and a mask decoder; the mask decoder includes: a two-layer Transformer module, a two-layer upsampling module, and a multilayer perceptron; Specifically, the sample paraspinal muscle images are magnetic resonance imaging (MRI) images obtained from images of the paraspinal muscles. The untrained initial model (original image segmentation model) can be the MedSAM model, a medical image segmentation algorithm based on the Segment Anything model, which is a SAM model re-pre-trained on large-scale medical image data. The reference... Figure 2 , Figure 2 A schematic diagram of the model architecture of a primitive image segmentation model is shown, such as... Figure 2 As shown, the original image segmentation model MedSAM includes at least: an image encoder, a prompt encoder, and a mask decoder.
[0019] Among them, the image encoder is used to extract the input image (such as...). Figure 2 The image encoder primarily employs a Vision Transformer (ViT)-based architecture, segmenting the image into fixed-size patches and encoding global and local features through multiple Transformer layers. The input to the image encoder is the original image, and the output is a high-dimensional feature map (such as...). Figure 2 The image embbedding shown is used to capture information about the structure, texture, and pathological areas of the image for use by the subsequent decoder.
[0020] A prompt encoder is used to convert user-provided interactive prompts (such as dots, boxes, coarse masks, etc.) into vector representations. The prompt type can be a bounding box, a coarse mask, or a text description. The input is the prompt information; in this embodiment, a prompt box (such as...) is used. Figure 2 The bounding box prompts shown are used as prompt information, and the output is a vector representation obtained by encoding the prompt information. The prompt information and image features are fused in the mask decoder to guide the localization of the segmentation target.
[0021] The mask decoder fuses image features (the output of the image encoder) and cue information (the output of the cue encoder) to generate an accurate segmentation mask for the target region. The mask decoder primarily employs a lightweight Transformer decoder or convolutional module. It integrates the outputs of the image encoder and cue encoder through a cross-attention mechanism. The output is a segmentation result (binary or multi-class mask) with the same resolution as the input image. In this embodiment, the output of the mask decoder is equivalent to the output of the original image segmentation model, i.e., the image segmentation result for the input paraspinal muscle image. For each input cue box, the model generates an accurate segmentation mask for its corresponding target muscle.
[0022] In step S101, this embodiment first preprocesses the obtained paraspinal muscle magnetic resonance images (i.e. sample paraspinal muscle images) and divides them into training sets, validation sets, and test sets; then, the training set and validation set are used to fine-tune the original model (original image segmentation model).
[0023] In some embodiments, step S101, which involves training the original image segmentation model using sample paraspinal muscle images to obtain a trained original image segmentation model, includes: Step S1011: Divide the data into training set, validation set and test set according to the ratio of 8:1:1 of the total number of data; Specifically, the total number of data points refers to the number of paraspinal muscle images obtained. The training set (80%) is used for model parameter learning. The validation set (10%) is used for hyperparameter tuning and monitoring the training process (e.g., early stopping). The test set (10%) is used for final model performance evaluation (not involved in training). Furthermore, the paraspinal muscle images are processed into the input format required by the MedSAM model, resulting in RGB three-channel images with a resolution of 1024×1024.
[0024] Step S1012: Perform data augmentation operations on the sample paraspinal muscle images in the training set. The data augmentation operations include at least random horizontal flipping, random rotation, random brightness adjustment, and random contrast adjustment. Specifically, data augmentation aims to increase data diversity and improve the robustness and generalization ability of the model. At least one of the above data augmentation operations is performed on each sample paraspinal muscle image to obtain newly added sample paraspinal muscle images. Random horizontal flipping simulates human symmetry (paraspinal muscles are typically bilaterally symmetrical). Random rotation involves slightly rotating the sample paraspinal muscle images (e.g., ±15°) to accommodate angular differences during image acquisition. Random brightness / contrast adjustment simulates image differences caused by different devices and scanning parameters.
[0025] Step S1013: Segment the region where the paraspinal muscles are located in the sample paraspinal muscle image to obtain the sample paraspinal muscle mask. Specifically, the sample paraspinal muscle mask is used to provide realistic labels for supervised learning. The paraspinal muscle regions in the image can be manually labeled by a professional physician, or manually corrected using existing automatic segmentation results. The sample paraspinal muscle mask can be a binary mask (background is 0, paraspinal muscle region is 1) or a multi-class mask (such as distinguishing between left and right sides, or different paraspinal muscle groups).
[0026] Step S1014: Obtain the sample prompt box of the sample paraspinal muscle image; Specifically, the sample cue boxes provide weakly supervised cues to the model, guiding it to focus on the target region. The minimum bounding box can be calculated based on the ground truth mask. The cue encoder encodes the box as a positional embedding, serving as prior guidance for the mask decoder (corresponding to the absolute position coordinates of the top-left and bottom-right corners of a 1024*1024 pixel matrix; the cue encoder first adds 0.5 to move it to the pixel center before normalization). Furthermore, based on the sample cue boxes, the image region within the cue boxes is cropped to obtain the paraspinal muscle image of the region to be segmented.
[0027] In some embodiments, step S1014, obtaining a sample prompt box for the sample paraspinal muscle image, includes: Step S1014-1: Use the smallest bounding rectangle of the real muscle segmentation contour in the sample paraspinal muscle image as the real bounding box. Step S1014-2: For the paraspinal muscle images of the samples belonging to the training set, randomly expand outward by 0-20 pixels based on the true annotation box to form the sample prompt box; For example, for a sample paraspinal muscle image 'a' belonging to the training set, the ground truth bounding box A1 of image 'a' is obtained. Based on the ground truth bounding box A1, it is randomly expanded outward by 0-20 pixels to obtain the sample cue box A2. In practical applications, it is considered that the provided cue box is difficult to perfectly match the target contour (i.e., the minimum bounding rectangle), and usually includes a small amount of background. Random expansion simulates this imprecision, allowing the model to learn to recover accurate segmentation from near but imperfect cue boxes, thus enhancing the model's robustness. Furthermore, by introducing randomness (the box seen in each training cycle may be slightly different), the diversity of training data is effectively increased, which helps prevent the model from overfitting to overly ideal cue boxes and improves its generalization ability. In addition, moderately including the background around the target (such as fat and bone) helps the model understand the anatomical context and better define the muscle boundary.
[0028] Step S1014-3: For the sample paraspinal muscle images belonging to the validation set and the test set, expand outward by 10 pixels based on the actual annotation box to form the sample prompt box.
[0029] In this embodiment, the validation set is used to tune hyperparameters and select the best model, while the test set is used for the final performance report. Using a fixed augmentation strategy ensures that all models are compared under identical input conditions on these key sets, avoiding the influence of random factors on the results, eliminating evaluation fluctuations caused by randomness, and making the results reproducible and comparable.
[0030] Step S1015: Using the sample paraspinal muscle image and the sample cue box as model input, and the corresponding sample paraspinal muscle mask as label, the original image segmentation model is trained to obtain the trained original image segmentation model.
[0031] In each round of model training, the input includes: sample paraspinal muscle images from the training set, corresponding sample cue boxes, and sample paraspinal muscle masks as labels. The training objective is to minimize the difference between the predicted mask (the model's output) and the true mask (i.e., the sample paraspinal muscle mask) using a loss function. The optimizer commonly used is AdamW, combined with a learning rate scheduling strategy. Segmentation metrics (such as Dice coefficient and IoU) are monitored on the validation set. Early stopping is employed to prevent overfitting. The trained original image segmentation model is the MedSAM model optimized for paraspinal muscles. MedSAM transfers the SAM framework to the medical imaging domain, enabling the model to possess stronger medical semantic understanding capabilities through re-pre-training on large-scale multimodal medical datasets (CT, MRI, ultrasound, etc.). Figure 2 As shown, the trained original image segmentation model MedSAM includes: an image encoder, a prompt encoder, and a mask decoder. This model is specifically designed for image segmentation of input paraspinal muscle images. The input to the image encoder is the original paraspinal muscle image, and the input to the prompt encoder is the bounding box prompts of the regions where each paraspinal muscle is located in the image. The mask decoder obtains the segmentation mask of each paraspinal muscle in the image.
[0032] The specific training process for step S1015 is as follows: Step S201: Input the sample paraspinal muscle image into the image encoder of the original image segmentation model to obtain the image encoding result; Step S202: Input the sample cue box corresponding to the sample paraspinal muscle image into the cue encoder of the original image segmentation model to obtain the cue encoding result; Step S203: Input the image encoding result and the prompt encoding result into the mask decoder of the original image segmentation model to obtain the prediction mask output by the final model, that is, the mask for the region where the paraspinal muscles are located in the sample paraspinal muscle image.
[0033] Step S204: Calculate the loss function based on the predicted mask and the sample paraspinal muscle mask (label) corresponding to the sample paraspinal muscle image; Step S205: Under the condition of freezing the parameters of the image encoder and the cue encoder, update the parameters of the original image segmentation model according to the loss function; Using the training set, repeat steps S201-205 above. After each training epoch or every N iterations, the model runs inference once on the validation set (forward propagation, without backpropagation or parameter updates). After the entire training process is complete, the best model selected from the validation set is evaluated once and only once on the test set to obtain the original trained image segmentation model.
[0034] Step S102: Improve the mask decoder in the trained original image segmentation model by replacing the double-layer upsampling module in the mask decoder with the Unet module to obtain the improved image segmentation model. The MedSAM model transfers the SAM framework to the medical imaging domain, enhancing its medical semantic understanding capabilities through re-pre-training on large-scale multimodal medical datasets (CT, MRI, ultrasound, etc.). However, the original MedSAM's mask decoder structure is relatively shallow, limiting its detail recovery capabilities; especially in high-resolution MRI images, the complexity of muscle boundaries and grayscale transition features place higher demands on the decoding module. Furthermore, as... Figure 2 As shown, MedSAM's segmentation mechanism relies on user-provided box prompts as explicit guidance signals. While this feature ensures the flexibility and interactivity of segmentation, in medical applications requiring batch processing or fully automated analysis, the prompt-dependent mechanism significantly reduces the system's automation level and processing efficiency.
[0035] Reference Figure 3 , Figure 3 A schematic diagram of the original mask decoder structure is shown, as follows: Figure 3As shown, the original mask decoder includes a two-layer Transformer module, a two-layer upsampling module, and a multilayer perceptron. As described in step S203, in the original mask decoder, the inputs are the image encoding result and the cue encoding result, and the output is the predicted mask.
[0036] Specifically, such as Figure 3 As shown, the two-layer Transformer module consists of two layers of Transformer modules. Each layer of Transformer module includes: self-attention (e.g., ... Figure 3 The diagram shows the self-attention layer, the token-to-image attention layer, the multilayer perceptron (MLP), and the image-to-token attention layer. The image encoding result and the cue encoding result are input into the two-layer Transformer module. During mask decoding, the image embedding output by the image encoder (as shown in Figure 3) and the tokens (the cue token and output token shown in Figure 3) serve as input to the two-layer Transformer module, interacting internally through a bidirectional cross-attention mechanism.
[0037] like Figure 3 As shown, the output of the two-layer Transformer module is then passed through a token to the image attention layer (token to image attn.) to obtain output features. These output features include two types of tokens, including an IoU token (corresponding to...). Figure 3 The IoU token permask shown is accompanied by multiple masked tokens (corresponding to...). Figure 3 The output token permask is shown. The IoU Token has a feature dimension of 256 and is used to evaluate the segmentation quality of the corresponding mask. The number of mask tokens is N+1, where N represents the number of candidate masks in the multi-mask prediction mode (N=3 by default in SAM). The additional mask token is used to represent the master mask in the single-mask prediction mode. Each mask token has a feature dimension of 256 and is used to represent a candidate segmentation hypothesis.
[0038] The output of the two-layer Transformer module also needs to pass through an additional two-layer upsampling module (such as...). Figure 3The 2×conv. Trans. shown. The double-layer upsampling module consists of two cascaded transposed convolutional layers: the first transposed convolution upsamples the feature map from low resolution to medium resolution, and the second transposed convolution upsamples the medium resolution to high resolution, outputting high-resolution image features (not the mask itself).
[0039] like Figure 3 As shown, each mask feature vector ([N+1, 256]) output by the masking module (mlpB) is treated as a set of weights. These weights are dot-producted with the image features output by the double-layer upsampling module to produce N+1 single-channel mask images. Finally, through the sigmoid activation function, N+1 binary mask predictions are obtained.
[0040] During training and inference, the mask decoder selects only the master mask as the effective output mask and generates the segmentation result and updates the model parameters based on the master mask. The other candidate masks are not output as the final segmentation result.
[0041] Reference Figure 4 , Figure 4 A schematic diagram of the structure of an improved image segmentation model is shown, as follows: Figure 4 As shown, the dual-layer upsampling module in the mask decoder is replaced with a Unet module. The improved mask decoder includes: a Unet module, a dual-layer Transformer module, a token-to-image attention layer (token to image attn.) located after the dual-layer Transformer module, and mlp.
[0042] Reference Figure 5 , Figure 5 A schematic diagram of the structure of a Unet module is shown, such as... Figure 5 As shown, the Unet module consists of three coding layers (corresponding to...) Figure 4 The U encoder and three decoding layers (corresponding to) Figure 4 The encoding layer consists of two 3×3 convolutional layers, followed by Batch Normalization (BN) and Rectified Linear Unit (ReLU) activation functions. The first encoding layer (corresponding to the U decoder) is the top layer and the bottom layer is the third layer. Each encoding layer consists of two 3×3 convolutional layers, followed by Batch Normalization (BN) and Rectified Linear Unit (ReLU) activation functions in sequence. Figure 4 In the U encoder layer 1) and the second coding layer (corresponding to Figure 4After the U encoder layer 2, 2×2 max pooling is used for downsampling. Each decoding layer also consists of two 3×3 convolutional layers, each equipped with BN and ReLU. The second decoding layer (corresponding to...) Figure 4 In the U decoder layer 2) and the third decoding layer (corresponding to Figure 4 After the U decoder layer 3, upsampling is performed using the up-conv module, whose structure includes: bilinear upsampling, 3×3 convolution, BN and ReLU.
[0043] The Unet module consists of two inputs. Input 1 is composed of the improved original paraspinal muscle image (the input of the improved image segmentation model) and the paraspinal muscle image of the corresponding muscle region to be segmented (the image block obtained by segmentation according to the prompt box). After being uniformly adjusted to a resolution of 256×256, the two are concatenated along the channel dimension to obtain Input 1. Input 2 is the output of the two-layer Transformer module in the mask decoder of the model (i.e., the output features described in step S1033 below). In the encoding stage, the input data of the first encoding layer is Input 1, and the data of the second and third encoding layers are the downsampling results of the previous encoding layer. In the decoding stage, the input of the second and first decoding layers is the concatenation of the corresponding encoding layer output features and the upsampling results of the previous decoding layer along the channel dimension; the input of the third decoding layer is the concatenation of the third encoding layer output features and Input 2, thereby realizing feature fusion and information completion.
[0044] Step S103: Under the condition of freezing the parameters of the image encoder and the cue encoder, the improved image segmentation model is trained using the sample paraspinal muscle image to obtain the target image segmentation model. Specifically, the sample paraspinal muscle images used in step S103 can be images from the training set obtained in step S101, used to train the improved model. The model input includes: sample paraspinal muscle images, sample cue boxes, and sample image blocks. The model output is a paraspinal muscle segmentation mask. The sample image blocks are obtained by cropping the sample paraspinal muscle images according to the sample cue boxes, resulting in the paraspinal muscle image of the region to be segmented.
[0045] In some embodiments, step S103, under the condition of freezing the parameters of the image encoder and the cue encoder, trains the improved image segmentation model using the sample paraspinal muscle image to obtain the target image segmentation model, including: Step S1031: Input the sample paraspinal muscle image into the image encoder of the improved image segmentation model to obtain the image encoding result; Step S1032: Input the sample cue box corresponding to the sample paraspinal muscle image into the cue encoder of the improved image segmentation model to obtain the cue encoding result; Step S1033: Input the image encoding result and the prompt encoding result into the two-layer Transformer module in the improved mask decoder to obtain the output feature; wherein, the output feature is the final output result of the two-layer Transformer module.
[0046] Step S1034: Input the sample paraspinal muscle image and the output features into the Unet module in the improved mask decoder to obtain the first mask decoding result; Step S1035: Input the output features into the multilayer perceptron in the improved mask decoder to obtain the second mask decoding result; as shown Figure 4 As shown, the output features need to be processed by a token to image attention layer (token to image attn.) before being input into a multilayer perceptron (MLP) to obtain the second mask decoding result.
[0047] Step S1036: Obtain the sample segmentation mask result based on the first mask decoding result and the second mask decoding result; Step S1037: Calculate the loss function based on the sample image segmentation mask result and the sample paraspinal muscle mask corresponding to the sample paraspinal muscle image; Step S1038: Under the condition of freezing the parameters of the image encoder and the cue encoder, update the parameters of the improved image segmentation model according to the loss function to obtain the trained target image segmentation model.
[0048] In some embodiments, step S1037, which calculates a loss function based on the sample image segmentation mask result and the sample paraspinal muscle mask corresponding to the sample paraspinal muscle image, includes: Step S1037-1: Based on the sample image segmentation mask result and the sample paraspinal muscle mask corresponding to the sample paraspinal muscle image, calculate the cross-entropy loss function and the Dice loss function. Step S1037-2: The weighted sum of the cross-entropy loss function and the Dice loss function is used as the final calculated overall loss function; In step S1038, after updating the parameters of the improved image segmentation model according to the loss function, the method further includes: After each training round, the Dice coefficient and IoU index of the validation set are calculated, and the model parameters with the best performance on the validation set are saved as the final weights.
[0049] Specifically, this embodiment uses a weighted combination of Cross Entropy Loss and Dice Loss as the overall loss function, and iteratively optimizes the model weights using the AdamW optimization algorithm. After each training round, the Dice coefficient and IoU performance metrics of the validation set are calculated, and the model parameters with the best performance on the validation set are saved as the final weights. The validation set here can be the validation set obtained in step S1011 above. Specifically, after each training epoch, the model is switched to evaluation mode, and the entire validation set is traversed (without data augmentation, using a fixed 10-pixel expansion of the cue box). For each sample, the model outputs a predicted mask and an IoU score, and the predicted mask with the highest IoU score is selected as the final prediction for that sample. The Dice coefficient and Intersection over Union (IoU) between the predicted mask and the ground truth mask are calculated. The strategy of saving the best model parameters is implemented. Training ends when the training termination condition is met. Training termination conditions can be any of the following: reaching the preset maximum number of epochs; triggering early stop: the verification metric has not improved for several consecutive epochs; performance saturation: the metric improvement is below the threshold.
[0050] In some embodiments, the Unet module consists of three encoding layers and three decoding layers; step S1034, inputting the sample paraspinal muscle image and the output features into the Unet module of the improved mask decoder to obtain a first mask decoding result, includes: Step S1034-1: Crop the sample paraspinal muscle image according to the sample prompt box to obtain a sample image block; for example, for the multifidus muscle in the sample paraspinal muscle image, mark the corresponding sample prompt box b, and then crop the image area selected by the sample prompt box b to obtain sample image block B, which is the image area where the multifidus muscle is located.
[0051] Step S1034-2: The sample paraspinal muscle image and the sample image block are stitched together to obtain a stitched sample image; specifically, the stitched sample image corresponds to Input 1 mentioned above. It is composed of the sample paraspinal muscle image and the paraspinal muscle image (sample image block) of the muscle to be segmented. After both are uniformly adjusted to a resolution of 256×256, they are stitched together along the channel dimension to obtain the stitched sample image.
[0052] Step S1034-3: The stitched sample image is input into the first encoding layer of the Unet module to obtain the first encoding result. Specifically, the first encoding layer consists of two 3×3 convolutional layers, followed by BatchNormalization (BN) and Rectified Linear Unit (ReLU) activation functions, and downsampling is performed using 2×2 max pooling. Therefore, the first encoding result is the final downsampled result output by the first encoding layer.
[0053] Step S1034-4: The first encoding result is input into the second encoding layer of the Unet module to obtain the second encoding result. Specifically, the second encoding layer consists of two 3×3 convolutional layers, followed by BatchNormalization (BN) and Rectified Linear Unit (ReLU) activation functions, and downsampling is performed using 2×2 max pooling. Therefore, the second encoding result is the final downsampled result output by the second encoding layer.
[0054] Step S1034-5: Input the second encoding result into the third encoding layer of the Unet module to obtain the third encoding result; specifically, the third encoding layer consists of two 3×3 convolutional layers, followed by BatchNormalization (BN) and Rectified Linear Unit (ReLU) activation functions in sequence.
[0055] Step S1034-6 involves concatenating the third encoding result with the output feature and inputting the concatenation into the third decoding layer of the Unet module to obtain the third decoding result. Specifically, the input to the third decoding layer is formed by concatenating the output feature of the third encoding layer (i.e., the third encoding result) with input 2, thereby achieving feature fusion and information completion. The third decoding layer consists of two 3×3 convolutional layers, also equipped with BN and ReLU, and then upsampling is performed using the up-conv module. The structure of the third decoding layer includes: bilinear upsampling, 3×3 convolution, BN, and ReLU. Therefore, the third decoding result is the final upsampled output of the third decoding layer.
[0056] Step S1034-7: The third decoding result and the second encoding result are concatenated and input into the second decoding layer of the Unet module to obtain the second decoding result. Specifically, the second decoding layer consists of two 3×3 convolutional layers, also equipped with BN and ReLU, and then upsampling is performed using the up-conv module. The structure of the second decoding layer includes: bilinear upsampling, 3×3 convolution, BN, and ReLU. Therefore, the second decoding result is the final upsampled result output by the second decoding layer.
[0057] Step S1034-8: The second decoding result is concatenated with the first encoding result and then input into the first decoding layer of the Unet module to obtain the first mask decoding result.
[0058] In some embodiments, the improved mask decoder includes multiple Unet modules, each Unet module corresponding to a paraspinal muscle type; the paraspinal muscle types include at least: multifidus, erector spinae, and psoas major. Before step S1034, in which the sample paraspinal muscle image and the output features are input into the Unet module of the improved mask decoder to obtain the first mask decoding result, the method further includes: Based on the sample prompt box, determine the Unet module corresponding to the corresponding muscle type.
[0059] Specifically, there are different types of paraspinal muscles. For the three types of paraspinal muscles (multifidus A1, erector spinae A2, and psoas major A3), three small Unet modules (U1, U2, U3) with the same structure are used to replace the original MedSAM's two-layer upsampling module. The corresponding small Unet module (U1, U2, U3) is selected according to the muscle type selected by the prompt box. For example, based on the muscle type selected by the prompt box, the muscle type is determined to be multifidus A1, and the model automatically selects the corresponding small Unet module U1.
[0060] It is important to note that since there are multiple Unet modules, all of these Unet modules need to be trained during the training of the improved image segmentation model in step S103, and each module is trained separately using images corresponding to its respective muscle type (training is performed according to the process described in step S103). For example, based on the muscle type, the muscle type is determined to be multifidus A1, and the model automatically selects the corresponding small Unet module U1. After selecting the corresponding Unet module U1, step S1034-1 is executed, and the sample paraspinal muscle image is cropped according to the sample prompt box (at least including the prompt box corresponding to multifidus muscle) to obtain sample image blocks (image blocks for multifidus muscle).
[0061] Step S104: Input the original paraspinal muscle image to be segmented into the target image segmentation model to obtain the segmented paraspinal muscle image.
[0062] In some embodiments, step S104, which involves inputting the original paraspinal muscle image to be segmented into the target image segmentation model to obtain a segmented paraspinal muscle image, includes: Obtain the target segmentation hint boxes corresponding to the locations of each paraspinal muscle in the original paraspinal muscle image to be segmented by following these steps: D1. Manually mark the target segmentation prompt box in the original paraspinal muscle image to be segmented; or D2. Input the original paraspinal muscle image to be segmented into a pre-trained cue box generation model to obtain the target segmentation cue box output by the model.
[0063] Specifically, the model adopts a dual-segmentation mode (e.g. Figure 4 As shown in YOLO (manual), there are two modes for obtaining muscle cue boxes: semi-automatic mode D1 and fully automatic mode D2. The semi-automatic mode uses manually labeled cue boxes as input to guide the model to segment muscle regions, improving the segmentation accuracy of complex samples. The fully automatic mode automatically generates cue boxes by training the cue box generation model, achieving end-to-end automatic segmentation, which is suitable for large-scale data processing.
[0064] In some embodiments, the prompt box generation model is trained through the following steps: Step S301: Construct a first training dataset. Each set of first training data in the first training dataset includes: a sample paraspinal muscle image, and rectangular cue boxes indicating the locations of each paraspinal muscle in the image. Specifically, rectangular cue boxes are generated using manually labeled multifidus, erector spinae, and psoas major muscle regions, and a training set (first training dataset) and a validation set are created. The rectangular cue boxes include six categories: left and right multifidus, left and right erector spinae, and left and right psoas major muscles.
[0065] Step S302: Using the paraspinal muscle images from the first training data as input to the initial cue box generation model, and using the rectangular cue boxes from the first training data as labels, the initial cue box generation model is trained. In this embodiment, the initial cue box generation model can be a lightweight YOLOv12 detection network.
[0066] Step S303: During training, save the model weights with the best performance on the validation set to obtain the trained prompt box generation model. Specifically, after each training epoch, switch the model to evaluation mode and evaluate the performance of the current model based on the validation set, calculating the detection performance metric, which includes, but is not limited to, the mean average precision (mAP) on the validation set. When the model's mAP on the validation set in the current training epoch is better than the historical best performance, save the corresponding model parameters as the optimal model weights, thereby obtaining the trained prompt box generation model. The training termination condition can be any of the following: reaching the preset maximum number of epochs; triggering early stopping: the validation metric does not improve for several consecutive epochs; performance saturation: the metric improvement is lower than a threshold.
[0067] Step S104 specifically includes the following steps: Step S401: Obtain the original paraspinal muscle image P to be segmented.
[0068] Step S402: Using semi-automatic mode D1 or fully automatic mode D2, obtain the corresponding prompt boxes for various paraspinal muscles in image P based on the original paraspinal muscle image P. The prompt boxes may include six prompt boxes for the left and right multifidus muscles, the left and right erector spinae muscles, and the left and right psoas major muscles.
[0069] Step S403: Input the original paraspinal muscle image P into the image encoder of the target image segmentation model to obtain the image encoding result; Step S404: Input the cue box corresponding to the original paraspinal muscle image P into the cue encoder of the target image segmentation model to obtain the cue encoding result; Step S405: Input the image encoding result and the prompt encoding result into the two-layer Transformer module in the target image segmentation model to obtain the output features; Step S406: Based on the prompt box, determine the corresponding paraspinal muscle type, and then determine the Unet module to which it belongs. Input the original paraspinal muscle image P and the output features into the corresponding Unet module to obtain the first mask decoding result. Step S407: Input the output features into the multilayer perceptron in the mask decoder to obtain the second mask decoding result; Step S408: Based on the first mask decoding result and the second mask decoding result, the paraspinal muscle segmentation result of the model is obtained as the final output, which is used for subsequent quantitative muscle analysis or clinical auxiliary diagnosis.
[0070] This application's embodiments are based on a large medical imaging model, MedSAM, with fine-tuning. Smaller Unet modules targeting different muscle types (multifidus, erector spinae, and psoas major) are introduced into its decoder section, enabling targeted feature extraction and fine-grained segmentation. This structure effectively solves the segmentation inaccuracies caused by the blurred boundaries, significant morphological differences, and complex tissue textures of the paraspinal muscles in MRI images, significantly improving segmentation accuracy and stability.
[0071] Furthermore, this application trains the lightweight object detection model YOLOv12 (a cue box generation model) to automatically generate cue boxes for the multifidus, erector spinae, and psoas major muscles, achieving an end-to-end automated processing flow from muscle detection to segmentation. It supports both semi-automatic and fully automatic working modes according to application needs, capable of handling accurate segmentation of complex cases and efficiently processing large-scale data, significantly improving its practicality in clinical and research scenarios.
[0072] Furthermore, by combining automatic detection prompts with a segmentation network, the reliance on manually drawing muscle contours is reduced, thus decreasing the workload of manual annotation and subjective bias. Simultaneously, the improved model (target image segmentation model) maintains high performance even with limited labeled data, addressing the critical issue of data scarcity in medical imaging.
[0073] Example 1 This application embodiment tests the segmentation of paraspinal muscle images and provides a suitable test environment and parameter settings. The specific configuration of the experimental environment is as follows: CPU is Intel(R) Xeon(R) Gold 5218 CPU @2.30GHz; GPU is NVIDIA GeForce RTX 3090; operating system is Ubuntu 18.04.6 LTS; programming language is Python 3.8; deep learning framework is PyTorch 1.13.0+cu116.
[0074] The method includes: Step S1 (corresponding to step S1011 in the above embodiment) preprocesses the original paraspinal muscle magnetic resonance images and divides them into training set, validation set and test set.
[0075] The original paraspinal muscle MRI images are in DICOM format. To obtain training images for the model, the raw DICOM data needs to be processed. From the perspective of pixel values, the original pixel values range from 0 to thousands. To facilitate model training, the original pixel values are linearly mapped to 0-255, the image resolution is modified to 1024×1024, and then it is saved as an RGB three-channel PNG image. Then, a professional doctor performs segmentation processing on the paraspinal muscle region in the image to obtain the paraspinal muscle segmentation mask.
[0076] The data is randomly divided into training, validation, and test sets in a ratio of 8:1:1.
[0077] The smallest bounding rectangle of the actual muscle segmentation contour is used as the true bounding box; the training set is randomly expanded outward by 0-20 pixels based on the true bounding box as a cue box; the validation and test sets are fixed at an expansion of 10 pixels as cue boxes to avoid the influence of random factors on the results. The paraspinal muscle image of the region to be segmented is obtained by cropping the paraspinal muscle image within the cue box according to the cue box.
[0078] To improve the model's generalization ability, data augmentation methods based on the Albumentations library were used during the training phase, including random horizontal flipping (p=0.5), random rotation within ±15° (p=0.5), and random brightness and contrast adjustment (p=0.5).
[0079] Step S2: Fine-tune the original model MedSAM using the training and validation sets.
[0080] As described in step S1015 of the above embodiment, the original MedSAM model (i.e., the original image segmentation model) is trained using the paraspinal muscle image, cue box, and paraspinal muscle segmentation mask. After training, the trained original image segmentation model is obtained, and the image encoder and cue encoder of MedSAM are frozen.
[0081] Step S3: Improve the mask decoder of the original MedSAM model to obtain the improved model (i.e., the improved image segmentation model).
[0082] For three types of paraspinal muscles (multifidus, erector spinae, and psoas major), three small Unet modules with the same structure are used to replace the original MedSAM's two-layer upsampling module. The corresponding small Unet module is selected according to the muscle type selected in the prompt box.
[0083] The small Unet module consists of three encoding layers and three decoding layers, with the top layer being the first layer and the bottom layer being the third layer. The small Unet module has two inputs: Input 1 consists of the original paraspinal muscle image and the corresponding muscle region image to be segmented. After being uniformly adjusted to a resolution of 256×256, the two are stitched together in the channel dimension; Input 2 is the output feature of the two-layer decoder (i.e., the two-layer Transformer module) in MedSAM's mask decoder.
[0084] Each coding layer consists of two 3×3 convolutional layers, followed by BatchNormalization (BN) and Rectified Linear Unit (ReLU) activation functions in sequence. The first and second layers are downsampled using 2×2 max pooling. During the encoding phase, the input data for the first coding layer is the small Unet module input 1, and the data for the second and third layers is the downsampled result of the previous coding layer.
[0085] Each decoding layer also consists of two 3×3 convolutional layers, each equipped with Batch Normalization (BN) and ReLU. Upsampling is performed after the second and third layers using an up-conv module, whose structure includes bilinear upsampling, 3×3 convolution, BN, and ReLU. During the decoding stage, the inputs to the second and first layers are the concatenation of the corresponding encoding layer's output features and the upsampling result from the previous decoding layer along the channel dimension. The input to the third decoding layer is the concatenation of the third encoding layer's output features and the input from a small Unet module, thus achieving feature fusion and information completion.
[0086] Step S4: Train and optimize the parameters of the improved model (i.e., the improved image segmentation model) to obtain the target image segmentation model.
[0087] (Corresponding to step S103 in the above embodiment) The improved model is trained using the training set obtained in step S1. The model input includes paraspinal muscle images, a cue box, and paraspinal muscle images of the region to be segmented. The model output is a paraspinal muscle segmentation mask. A weighted combination of the cross-entropy loss function and the Dice loss function is used as the overall loss function, and the model weights are iteratively optimized using the AdamW optimization algorithm. After each round of training, the Dice coefficient and IoU performance metrics of the validation set are calculated, and the model parameters with the best performance on the validation set are saved as the final weights.
[0088] Step S5: Train an object detection model (i.e., a cue box generation model) to obtain muscle cue boxes.
[0089] A paraspinal muscle detection dataset was constructed, and corresponding rectangular bounding boxes were generated using the multifidus, erector spinae, and psoas major muscle regions annotated by doctors. Training and validation sets were also created.
[0090] The lightweight object detection model YOLOv12 was selected for training. During training, the model weights with the best performance on the validation set were saved.
[0091] The trained model is used to detect new MRI images and automatically generate six cue boxes for the left and right multifidus muscles, left and right erector spinae muscles, and left and right psoas major muscles, which are used as input for subsequent segmentation models.
[0092] Step S6: Segment the paraspinal muscle image using the target image segmentation model.
[0093] The model has two modes for obtaining muscle cue boxes: semi-automatic mode and fully automatic mode. In semi-automatic mode, manually labeled cue boxes are used as input to guide the model to segment muscle regions, improving the segmentation accuracy of complex samples. In fully automatic mode, cue boxes are automatically generated through the trained YOLOv12 detection network (i.e., the cue box generation model), achieving end-to-end automatic segmentation, which is suitable for large-scale data processing.
[0094] Based on the muscle type selected in the prompt box, the model automatically selects the corresponding small Unet module, and uses the trained improved model to perform feature extraction and image segmentation on the paraspinal muscle image within the prompt box; the paraspinal muscle segmentation results are output for subsequent quantitative muscle analysis or clinical auxiliary diagnosis.
[0095] To verify the performance improvement of the above method for paraspinal muscle segmentation in images, DIce and IoU were used as evaluation metrics. Under the same configuration environment, this embodiment compared the evaluation metrics of the state-of-the-art model nnUnet and the two segmentation modes (semi-automatic mode D1 and fully automatic mode D2) of the proposed method on the same dataset for six types of paraspinal muscles (LES: left erector spinae, RES: right erector spinae, LPM: left psoas muscle, RPM: right psoas muscle, LMM: left multifidus muscle, RMM: right multifidus muscle)
[0096] Specifically, Table 1 shows the segmentation performance (Dice) of nnUnet and the two segmentation modes in this embodiment.
[0097] Table 1
[0098] Table 2 shows the segmentation performance (IoU) of nnUnet and the two segmentation modes in this embodiment.
[0099] Table 2
[0100] Tables 1 and 2 show the segmentation performance metrics of the two segmentation modes and the state-of-the-art (SOTA) model nnUnet for six muscle classes in this embodiment. In the fully automatic segmentation mode, compared to nnUnet, the Dice improvement for segmenting the six muscle classes is 0.2%–1.4%, and the IoU improvement is 0.4%–2.5%. Particularly for the RMM, LES, and LMM muscle classes, the Dice improvement exceeds 1%. The semi-automatic segmentation mode, due to the use of manually labeled cue boxes, significantly improves segmentation performance, achieving a Dice improvement of 1.6%–4.4% and an IoU improvement of 2.6%–7.7% for segmenting the six muscle classes compared to nnUnet.
[0101] To further visually verify the segmentation effect of the model, the segmentation results of different methods were compared and visualized. Overall, the segmentation results of each model are highly consistent with real muscle data, accurately identifying the morphological regions of the paraspinal muscles, indicating that all models have strong feature extraction and region recognition capabilities. Although the segmentation boundaries of different models are relatively similar, it can be observed that the method proposed in this embodiment performs better in terms of detail preservation and edge continuity, more smoothly fitting the muscle contour boundaries and reducing missegmentation of noisy regions. In particular, the semi-automatic mode using realistic cue boxes shows outstanding performance in terms of the structural integrity of muscle edges and the accuracy of small region recognition.
[0102] In summary, this application proposes a method for segmenting paraspinal muscles from axial magnetic resonance imaging of the spine. While retaining the powerful encoding capabilities of MedSAM, it innovatively introduces a multi-branch Mini U-Net structure, significantly enhancing feature reconstruction and detail recovery capabilities during the decoding stage, thus achieving accurate segmentation of the paraspinal muscle region. Simultaneously, the dual-mode segmentation design allows the model to flexibly adapt to different application scenarios: the automatic mode emphasizes segmentation efficiency and scalable application, while the semi-automatic mode ensures higher model accuracy.
[0103] A second aspect of this application also provides a paraspinal muscle image segmentation apparatus, applied to perform the paraspinal muscle image segmentation method described in the first aspect of this application, with reference to... Figure 6 , Figure 6 A schematic diagram of a paraspinal muscle image segmentation device is shown, as follows: Figure 6 As shown, the device includes: The first training module is used to train the original image segmentation model using sample paraspinal muscle images to obtain the trained original image segmentation model; the original image segmentation model includes at least: an image encoder, a cue encoder, and a mask decoder; the mask decoder includes: a two-layer Transformer module, a two-layer upsampling module, and a multilayer perceptron; An improvement module is used to improve the mask decoder in the trained original image segmentation model, so as to replace the double-layer upsampling module in the mask decoder with the Unet module to obtain the improved image segmentation model. The second training module is used to train the improved image segmentation model using the sample paraspinal muscle images under the condition of freezing the parameters of the image encoder and the cue encoder, so as to obtain the target image segmentation model. The segmentation module is used to input the original paraspinal muscle image to be segmented into the target image segmentation model to obtain the segmented paraspinal muscle image.
[0104] In some embodiments, under the condition of freezing the parameters of the image encoder and the cue encoder, the improved image segmentation model is trained using the sample paraspinal muscle images to obtain a target image segmentation model, including: The sample paraspinal muscle image is input into the image encoder of the improved image segmentation model to obtain the image encoding result; Input the sample cue box corresponding to the sample paraspinal muscle image into the cue encoder of the improved image segmentation model to obtain the cue encoding result; The image encoding result and the prompt encoding result are input into the two-layer Transformer module in the improved mask decoder to obtain the output features; The sample paraspinal muscle image and the output features are input into the Unet module in the improved mask decoder to obtain the first mask decoding result; The output features are input into the multilayer perceptron in the improved mask decoder to obtain the second mask decoding result; Based on the first mask decoding result and the second mask decoding result, the sample segmentation mask result is obtained; Based on the sample image segmentation mask result, and the sample paraspinal muscle mask corresponding to the sample paraspinal muscle image, calculate the loss function; With the parameters of the image encoder and the cue encoder frozen, the parameters of the improved image segmentation model are updated according to the loss function to obtain the trained target image segmentation model.
[0105] In some embodiments, the Unet module consists of three encoding layers and three decoding layers; the step of inputting the sample paraspinal muscle image and the output features into the Unet module of the improved mask decoder to obtain a first mask decoding result includes: The sample paraspinal muscle image is cropped according to the sample prompt box to obtain a sample image block; The sample paraspinal muscle image and the sample image block are stitched together to obtain the stitched sample image; The stitched sample image is input into the first encoding layer of the Unet module to obtain the first encoding result; The first encoding result is input into the second encoding layer of the Unet module to obtain the second encoding result; The second encoding result is input into the third encoding layer of the Unet module to obtain the third encoding result; The third encoding result is concatenated with the output feature and then input into the third decoding layer of the Unet module to obtain the third decoding result; The third decoding result is concatenated with the second encoding result and then input into the second decoding layer of the Unet module to obtain the second decoding result. The second decoding result is concatenated with the first encoding result and then input into the first decoding layer of the Unet module to obtain the first mask decoding result.
[0106] In some embodiments, the improved mask decoder includes multiple Unet modules, each Unet module corresponding to a paraspinal muscle type; the paraspinal muscle types include at least: multifidus, erector spinae, and psoas major. Before inputting the sample paraspinal muscle image and the output features into the Unet module of the improved mask decoder to obtain the first mask decoding result, the device further includes: The determination module is used to determine the Unet module corresponding to the corresponding muscle type based on the sample prompt box.
[0107] In some embodiments, calculating the loss function based on the sample image segmentation mask result and the sample paraspinal muscle mask corresponding to the sample paraspinal muscle image includes: Based on the sample image segmentation mask result, and the sample paraspinal muscle mask corresponding to the sample paraspinal muscle image, calculate the cross-entropy loss function and the Dice loss function; The weighted sum of the cross-entropy loss function and the Dice loss function is used as the final calculated overall loss function; After updating the parameters of the improved image segmentation model according to the loss function, the method further includes: After each training round, the Dice coefficient and IoU index of the validation set are calculated, and the model parameters with the best performance on the validation set are saved as the final weights.
[0108] In some embodiments, inputting the original paraspinal muscle image to be segmented into the target image segmentation model to obtain the segmented paraspinal muscle image includes: Obtain the target segmentation hint boxes corresponding to the locations of each paraspinal muscle in the original paraspinal muscle image to be segmented by following these steps: Manually mark the target segmentation cue box in the original paraspinal muscle image to be segmented; or The original paraspinal muscle image to be segmented is input into a pre-trained cue box generation model to obtain the target segmentation cue box output by the model.
[0109] In some embodiments, the prompt box generation model is trained through the following steps: Construct a first training dataset. Each set of first training data in the first training dataset includes: a sample paraspinal muscle image, and a rectangular cue box indicating the location of each paraspinal muscle in the image. The paraspinal muscle images from the first training data are used as input to the initial prompt generation model, and the rectangular prompts from the first training data are used as labels to train the initial prompt generation model. During training, the model weights with the best performance on the validation set are saved to obtain the trained prompt box generation model.
[0110] In some embodiments, training the original image segmentation model using sample paraspinal muscle images to obtain a trained original image segmentation model includes: The data was divided into training, validation, and test sets in a ratio of 8:1:1 based on the total amount of data. Data augmentation operations are performed on the paraspinal muscle images in the training set, including at least random horizontal flipping, random rotation, random brightness adjustment, and random contrast adjustment. The region containing the paraspinal muscles in the sample paraspinal muscle image is segmented to obtain the sample paraspinal muscle mask. A sample prompt box for obtaining the sample paraspinal muscle image; Using the sample paraspinal muscle image and the sample cue box as model input, and the corresponding sample paraspinal muscle mask as label, the original image segmentation model is trained to obtain the trained original image segmentation model.
[0111] In some embodiments, the sample prompt box for obtaining the sample paraspinal muscle image includes: The smallest bounding rectangle of the real muscle segmentation contour in the sample paraspinal muscle image is used as the real annotation box. For the paraspinal muscle images belonging to the training set, the sample prompt box is randomly expanded outward by 0–20 pixels based on the true annotation box; For sample paraspinal muscle images belonging to the validation set and the test set, the sample cue box is fixedly expanded outward by 10 pixels based on the true annotation box.
[0112] This application also provides an electronic device, see embodiments thereof. Figure 7 , Figure 7 This is a schematic diagram of the electronic device proposed in an embodiment of this application. Figure 7 As shown, the electronic device 100 includes a memory 110 and a processor 120. The memory 110 and the processor 120 are connected via a bus for communication. The memory 110 stores a computer program that can run on the processor 120 to implement the steps in the paraspinal muscle image segmentation method disclosed in the embodiments of this application.
[0113] This application also provides a computer-readable storage medium storing a computer program / instructions thereon, which, when executed by a processor, implements the steps in the paraspinal muscle image segmentation method disclosed in this application.
[0114] This application also provides a computer program product that, when executed on an electronic device, causes a processor to implement the steps of the paraspinal muscle image segmentation method disclosed in this application. The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably.
[0115] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0116] The foregoing provides a detailed description of the paraspinal muscle image segmentation method and apparatus provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
[0117] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0118] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
[0119] The terms "an embodiment," "embodiment," or "one or more embodiments" as used herein mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of this application. Furthermore, please note that the examples of the phrase "in one embodiment" do not necessarily all refer to the same embodiment.
[0120] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0121] In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for segmenting paraspinal muscle images, characterized in that, The method includes: The original image segmentation model is trained using sample paraspinal muscle images to obtain the trained original image segmentation model; the original image segmentation model includes at least: an image encoder, a cue encoder, and a mask decoder; the mask decoder includes: a two-layer Transformer module, a two-layer upsampling module, and a multilayer perceptron; The mask decoder in the trained original image segmentation model is improved by replacing the double-layer upsampling module in the mask decoder with the Unet module to obtain the improved image segmentation model. With the parameters of the image encoder and the cue encoder frozen, the improved image segmentation model is trained using the sample paraspinal muscle images to obtain the target image segmentation model. The original paraspinal muscle image to be segmented is input into the target image segmentation model to obtain the segmented paraspinal muscle image.
2. The paraspinal muscle image segmentation method according to claim 1, characterized in that, With the parameters of the image encoder and the cue encoder frozen, the improved image segmentation model is trained using the sample paraspinal muscle images to obtain the target image segmentation model, including: The sample paraspinal muscle image is input into the image encoder of the improved image segmentation model to obtain the image encoding result; Input the sample cue box corresponding to the sample paraspinal muscle image into the cue encoder of the improved image segmentation model to obtain the cue encoding result; The image encoding result and the prompt encoding result are input into the two-layer Transformer module in the improved mask decoder to obtain the output features; The sample paraspinal muscle image and the output features are input into the Unet module in the improved mask decoder to obtain the first mask decoding result; The output features are input into the multilayer perceptron in the improved mask decoder to obtain the second mask decoding result; Based on the first mask decoding result and the second mask decoding result, the sample segmentation mask result is obtained; Based on the sample image segmentation mask result, and the sample paraspinal muscle mask corresponding to the sample paraspinal muscle image, calculate the loss function; With the parameters of the image encoder and the cue encoder frozen, the parameters of the improved image segmentation model are updated according to the loss function to obtain the trained target image segmentation model.
3. The paraspinal muscle image segmentation method according to claim 2, characterized in that, The Unet module consists of three encoding layers and three decoding layers; the process of inputting the sample paraspinal muscle image and the output features into the Unet module of the improved mask decoder to obtain the first mask decoding result includes: The sample paraspinal muscle image is cropped according to the sample prompt box to obtain a sample image block; The sample paraspinal muscle image and the sample image block are stitched together to obtain the stitched sample image; The stitched sample image is input into the first encoding layer of the Unet module to obtain the first encoding result; The first encoding result is input into the second encoding layer of the Unet module to obtain the second encoding result; The second encoding result is input into the third encoding layer of the Unet module to obtain the third encoding result; The third encoding result is concatenated with the output feature and then input into the third decoding layer of the Unet module to obtain the third decoding result; The third decoding result is concatenated with the second encoding result and then input into the second decoding layer of the Unet module to obtain the second decoding result. The second decoding result is concatenated with the first encoding result and then input into the first decoding layer of the Unet module to obtain the first mask decoding result.
4. The paraspinal muscle image segmentation method according to claim 2, characterized in that, The improved mask decoder includes multiple Unet modules, each Unet module corresponding to a type of paraspinal muscle. The paraspinal muscle types include at least: multifidus, erector spinae, and psoas major. Before inputting the sample paraspinal muscle image and the output features into the Unet module of the improved mask decoder to obtain the first mask decoding result, the method further includes: Based on the sample prompt box, determine the Unet module corresponding to the corresponding muscle type.
5. The paraspinal muscle image segmentation method according to claim 2, characterized in that, The step of calculating the loss function based on the sample image segmentation mask result and the sample paraspinal muscle mask corresponding to the sample paraspinal muscle image includes: Based on the sample image segmentation mask result, and the sample paraspinal muscle mask corresponding to the sample paraspinal muscle image, calculate the cross-entropy loss function and the Dice loss function; The weighted sum of the cross-entropy loss function and the Dice loss function is used as the final calculated overall loss function; After updating the parameters of the improved image segmentation model according to the loss function, the method further includes: After each training round, the Dice coefficient and IoU index of the validation set are calculated, and the model parameters with the best performance on the validation set are saved as the final weights.
6. The paraspinal muscle image segmentation method according to claim 1, characterized in that, The step of inputting the original paraspinal muscle image to be segmented into the target image segmentation model to obtain the segmented paraspinal muscle image includes: Obtain the target segmentation hint boxes corresponding to the locations of each paraspinal muscle in the original paraspinal muscle image to be segmented by following these steps: Manually mark the target segmentation cue box in the original paraspinal muscle image to be segmented; or The original paraspinal muscle image to be segmented is input into a pre-trained cue box generation model to obtain the target segmentation cue box output by the model.
7. The paraspinal muscle image segmentation method according to claim 6, characterized in that, The prompt box generation model is trained through the following steps: Construct a first training dataset. Each set of first training data in the first training dataset includes: a sample paraspinal muscle image, and a rectangular cue box indicating the location of each paraspinal muscle in the image. The paraspinal muscle images from the first training data are used as input to the initial prompt generation model, and the rectangular prompts from the first training data are used as labels to train the initial prompt generation model. During training, the model weights with the best performance on the validation set are saved to obtain the trained prompt box generation model.
8. The paraspinal muscle image segmentation method according to claim 1, characterized in that, The process of training the original image segmentation model using sample paraspinal muscle images to obtain the trained original image segmentation model includes: The data was divided into training, validation, and test sets in a ratio of 8:1:1 based on the total amount of data. Data augmentation operations are performed on the paraspinal muscle images in the training set, including at least random horizontal flipping, random rotation, random brightness adjustment, and random contrast adjustment. The region containing the paraspinal muscles in the sample paraspinal muscle image is segmented to obtain the sample paraspinal muscle mask. A sample prompt box for obtaining the sample paraspinal muscle image; Using the sample paraspinal muscle image and the sample cue box as model input, and the corresponding sample paraspinal muscle mask as label, the original image segmentation model is trained to obtain the trained original image segmentation model.
9. The paraspinal muscle image segmentation method according to claim 8, characterized in that, The sample prompt box for obtaining the sample paraspinal muscle image includes: The smallest bounding rectangle of the real muscle segmentation contour in the sample paraspinal muscle image is used as the real annotation box. For the paraspinal muscle images belonging to the training set, the sample prompt box is randomly expanded outward by 0–20 pixels based on the true annotation box; For sample paraspinal muscle images belonging to the validation set and the test set, the sample cue box is fixedly expanded outward by 10 pixels based on the true annotation box.
10. A paraspinal muscle image segmentation device, characterized in that, The apparatus is used to perform the paraspinal muscle image segmentation method according to any one of claims 1-9, the apparatus comprising: The first training module is used to train the original image segmentation model using sample paraspinal muscle images to obtain the trained original image segmentation model; the original image segmentation model includes at least: an image encoder, a cue encoder, and a mask decoder; the mask decoder includes: a two-layer Transformer module, a two-layer upsampling module, and a multilayer perceptron; An improvement module is used to improve the mask decoder in the trained original image segmentation model, so as to replace the double-layer upsampling module in the mask decoder with the Unet module to obtain the improved image segmentation model. The second training module is used to train the improved image segmentation model using the sample paraspinal muscle images under the condition of freezing the parameters of the image encoder and the cue encoder, so as to obtain the target image segmentation model. The segmentation module is used to input the original paraspinal muscle image to be segmented into the target image segmentation model to obtain the segmented paraspinal muscle image.