This invention discloses a multi-level semantically guided bone tumor
image segmentation method, belonging to the field of medical
image processing technology. The method first completes the spatial alignment and coordinate unification of CT and
MR images through affine registration and spatial
resampling based on
mutual information, and simultaneously performs image preprocessing by enhancing the CT and MR image data. A 3D SwinUNETR
encoder is used to extract multi-scale spatial features from the registered and aligned bone tumor CT and
MR images, and a 3D bidirectional cross-attention
fusion mechanism is used to achieve complementary fusion of multimodal features within the same
lesion region. Simultaneously, a multi-level semantic guidance path is constructed, embedding class-level and case-level textual
semantics into a dynamic controller, which, together with the
multimodal image fusion features, is input into a
multilayer perceptron to predict and generate dynamic convolutional kernel parameters with strong semantic directionality. This enables
adaptive weighting and region focusing of the decoded feature map of the CT image, and finally outputs a semantically guided corrected bone tumor prediction
mask. Experimental results show that this invention effectively improves the model's ability to jointly represent the complex morphological structure and multidimensional clinical
semantics of bone tumor lesions, significantly enhances the segmentation accuracy of bone tumors under conditions of blurred
lesion boundaries and scarce samples, and exhibits excellent generalization performance in independent validation set tests. This method achieves a Dessian similarity coefficient of 0.87 on bone tumor datasets, making it suitable for bone tumor assisted localization and
surgical planning needs.