A bone imaging lesion segmentation method based on deep learning
By using the Encoder module and attention concentration module of the Swin Transformer network in bone imaging lesion segmentation, combined with the Decoder module of stepwise feature fusion, the problems of low segmentation accuracy and poor feature fusion in the existing technology are solved, and higher lesion segmentation accuracy and model generalization ability are achieved.
Patent Information
- Application Number
- CN202310370170.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-04-10
AI Technical Summary
The existing deep learning methods have problems with low segmentation accuracy and poor feature fusion in bone imaging lesions segmentation, especially when facing more lesions information, the model segmentation accuracy is significantly reduced.
The Encoder module based on the Swin Transformer network is used for feature extraction, and the attention concentration module is used to reduce attention distraction. Combined with the Decoder module with step-by-step feature fusion, the accuracy of lesion segmentation is improved.
Through the attention convergence module and step-by-step feature fusion, the accuracy of bone imaging lesions and the generalization ability of the model are improved, the fusion of feature information is enhanced, and the accuracy of segmentation is improved.
Smart Images

Figure CN116433680B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a bone imaging lesion segmentation method based on deep learning. Background Art
[0002] With the advent of single-photon emission computed tomography (SPECT), it has been identified as an effective means of detection for the diagnosis, treatment, evaluation and prevention of a range of serious diseases. Bone imaging, also known as bone scintigraphy, is usually the preferred method for early determination of whether a malignant tumor has metastasized. Compared with traditional medical imaging techniques, although bone imaging has low imaging resolution and is susceptible to noise interference, it can obtain the function and metabolic state of tissues and organs. The manual reading method for determining the location of bone imaging lesions is not only time-consuming and prone to misjudgment and missed judgment, but also has low efficiency and is easily affected by subjective factors, and cannot meet the growing demand for bone imaging lesion segmentation.
[0003] The bone imaging lesion segmentation method based on deep learning can effectively solve the problems of time-consuming, labor-intensive and inefficient manual reading. The literature (Yu Hong; Luo Renze; Chen Chunmeng, et al. Research on SPECT bone imaging lesion segmentation based on improved U-Net [J]. Optoelectronics (Laser); 2022; 33(10): 1110-1120.) proposed to use the improved U-Net network to automatically extract bone imaging lesion information, which not only realizes automatic segmentation but also improves the time-consuming and labor-intensive problem of manual feature extraction; the literature (Luo Mingyang; Lin Qiang; Gao Ruiting, et al. Automatic segmentation of arthritis lesions in SPECT bone imaging [J]. Modern Electronic Technology; 2022; 33(10): 145-152) uses the construction of Mask-based The feature selection method of the CNN model of the R-CNN arthritis lesion segmentation model improves the model's ability to extract the features of arthritis lesions; the literature (Zhang Yi; Li Lin; Pi Yong, et al.; Bone imaging bone lesion segmentation method, system and device based on deep neural network [P]. Chinese Patent: CN115019049A; 2022.09.06) constructs a cascade network model including a refinement network and two neural networks, and in the refinement network, by calculating the DSC index of the first segmentation result of the bone lesion and the second segmentation result of the bone lesion, the bone lesion segmentation result with a DSC index higher than the threshold is defined as a reliable lesion segmentation result of the first stage, and extracts lesions that are difficult to distinguish in the first neural network and the second neural network, and outputs the refined segmentation result. It can be seen that most of the existing deep learning methods use convolutional networks to extract lesion information, which are strong in extracting local information but weak in extracting global and long-distance information. When faced with more lesion information, the model segmentation accuracy is significantly reduced. Summary of the invention
[0004] The task of medical image segmentation is to segment medical images into unconnected areas according to their characteristics, and the features of the same area are similar. With the improvement of living standards, people pay more and more attention to health, and various medical examinations are increasing. Bone imaging is increasingly used as an effective means of early judgment of cancer metastasis. However, medical image processing is a complicated and demanding task, and manual reading greatly occupies medical resources. With the development of deep learning, image segmentation technology has made certain progress, but bone imaging images are still difficult to segment due to large individual differences and difficulty in feature extraction. Existing methods for bone imaging segmentation still have defects such as low segmentation accuracy and poor feature fusion.
[0005] In order to overcome the defects of existing methods, a bone imaging lesion segmentation method based on deep learning is proposed. This method effectively solves the shortcomings of traditional deep learning methods and improves the segmentation accuracy and generalization ability of the model.
[0006] To achieve the above-mentioned purpose of the invention, the technical solution provided is a bone imaging lesion segmentation method based on deep learning, which specifically includes the following steps:
[0007] Step S1: acquiring bone imaging images; annotating the bone imaging images and then assigning the bone imaging images to a bone imaging image training set and a bone imaging image test set; completing the construction of the bone imaging data set;
[0008] Step S2: cutting the bone imaging images in the bone imaging image training set and the bone imaging image test set obtained in step S1 into patch images of the same size; inputting the patch images into the encoder module;
[0009] Step S3: Construct an Encoder module, which encodes the input image and obtains the feature map of the image. The Encoder module uses the feature extraction part of the Swin Transformer network, which includes four stages. The first stage consists of a Linear Embedding module and a Swin Transformer Block module, and the following three stages consist of a Patch Merging module and a Swin Transformer Block module. The Linear Embedding module increases the number of channels of the patch image cut in step S2 through a convolution operation. The Swin Transformer Block mainly consists of two computing units. The first is a window multi-head self-attention unit, and the second is a moving window multi-head self-attention unit.
[0010] Step S4: construct an attention convergence module;
[0011] The attention convergence module re-converges the scattered attention of the feature maps output from the four stages through convolutional layers, ReLU activation functions, and upsampling;
[0012] Step S5: Construct a Decoder module;
[0013] The Decoder module decodes the input feature map; the Decoder first inputs the feature map output by the fourth stage into the attention convergence module to reduce attention dispersion, and then performs step-by-step feature fusion on the feature map output by the attention convergence module. The step-by-step feature fusion splices the feature map output by the fourth stage through the attention convergence module with the feature map output by the third stage through the attention convergence module, and then reduces the number of channels of the spliced feature map by half through the convolution operation, and then splices it with the feature map output by the second stage through the attention convergence module; then uses the convolution operation to reduce the number of channels of the spliced feature map by half to obtain the feature map, and then uses the convolution operation to reduce the number of channels of the spliced feature map by half to obtain the feature map, and then uses the convolution operation to reduce the number of channels of the spliced feature map by half to obtain the feature map; finally, the feature map obtained by the step-by-step feature fusion is input into the linear prediction module to complete the bone imaging lesion segmentation.
[0014] Compared with the prior art, the present invention has the following characteristics:
[0015] The attention convergence module can converge the attention of the feature maps output by the four stages in the Swin Transformer encoder part, reduce the attention distraction caused by the Swin Transformer Block module, and reduce the differences between feature maps with different resolutions through gradual feature fusion, so that the feature map information is fully integrated and the accuracy of lesion segmentation is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a specific implementation flow chart of the present invention.
[0017] Figure 2 It is the overall network model structure diagram of the present invention.
[0018] Figure 3 It is a structural diagram of the attention convergence module in the present invention. DETAILED DESCRIPTION
[0019] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.
[0020] A bone imaging lesion segmentation method based on deep learning; the implementation steps of this method are as follows:
[0021] Step S1: Acquire a bone imaging data set, and divide the bone imaging data set into a bone imaging image training set and a bone imaging image test set.
[0022] Step S2: Cut the bone imaging images in the bone imaging image training set and the bone imaging image test set obtained in step S1 to obtain patch images of the same size, and input them into the Encoder.
[0023] Step S3: Construct an Encoder module, which performs tensor representation on the input image to obtain a feature map of the image; the Encoder module of the present invention adopts the feature extraction part of the Swin Transformer network, which includes four stages; the first stage is composed of a Linear Embedding module and a Swin Transformer Block module, and the following three stages are composed of a Patch Merging module and a Swin Transformer Block module; the Linear Embedding module increases the number of channels of the patch image obtained by cutting in step S2 through a convolution operation, and the Swin Transformer Block is mainly composed of two computing units, the first is a window multi-head self-attention unit, and the second is a moving window multi-head self-attention unit;
[0024] Step S4, constructing an attention convergence module;
[0025] The purpose of this module is to solve the problem of scattered attention caused by multiple Swin Transformer Block modules in the Encoder module. Through convolutional layers, Relu activation functions and upsampling, the scattered attention of the feature maps output from the four stages is re-converged.
[0026] Step S5: Construct a Decoder module;
[0027] The Decoder module decodes the input feature map; the Decoder first inputs the feature map output by the fourth stage into the attention convergence module to reduce attention dispersion, and then performs step-by-step feature fusion on the feature map output by the attention convergence module; the step-by-step feature fusion splices the feature map output by the fourth stage through the attention convergence module with the feature map output by the third stage through the attention convergence module, and then reduces the number of channels of the spliced feature map by half through the convolution operation, and then splices it with the feature map output by the second stage through the attention convergence module; then reduces the number of channels of the spliced feature map by half through the convolution operation to obtain the feature map; then splices it with the feature map output by the first stage through the attention convergence module, and then reduces the number of channels of the spliced feature map by half through the convolution operation to obtain the feature map; finally, the feature map obtained by the step-by-step feature fusion is input into the linear prediction module to complete the bone imaging lesion segmentation.
[0028] The flowchart of the implementation method is as follows Figure 3 As shown; including the following steps:
[0029] Parameter range: Lr = 0.0001, Batch size = 1, Epochs = 400
[0030] Step S10, acquiring bone imaging images and dividing them into a bone imaging image training set and a bone imaging image test set;
[0031] Step S20, performing data preprocessing on the bone imaging image training set images;
[0032] Step S30, constructing an Encoder module to extract semantic information;
[0033] Step S40, constructing an attention convergence module to converge attention again and reduce distraction;
[0034] Step S40, constructing a Decoder module to gradually fuse the feature maps of different resolutions to improve the segmentation accuracy;
[0035] Step S50, segmenting the lesion area.
Claims
1. A bone imaging lesion segmentation method based on deep learning, characterized in that The following steps are involved: Step S1, obtaining bone imaging images, annotating the bone imaging images, and then allocating the bone imaging images into a bone imaging image training set and a bone imaging image test set to complete the construction of the bone imaging data set; Step S2, cutting the bone imaging images in the bone imaging image training set and the bone imaging image test set obtained in step S1 into patch images of the same size, and inputting the patch images into the encoder module; Step S3, construct an Encoder module, which encodes the input image to obtain the feature map of the image; the Encoder module uses the feature extraction part of the Swin Transformer network, which includes four stages; the first stage consists of the Linear Embedding module and the Swin Transformer Block module, and the following three stages are composed of the PatchMerging module and the Swin Transformer Block module; the Linear Embedding module increases the number of channels of the patch image cut in step S2 through convolution operations, and the Swin Transformer Block consists of two computing units, the first is a window multi-head self-attention unit, and the second is a moving window multi-head self-attention unit; Step S4, constructing an attention convergence module; The attention convergence module re-converges the scattered attention of the feature maps output from the four stages through convolutional layers, ReLU activation functions, and upsampling; Step S5, constructing a Decoder module; The Decoder module decodes the input feature map; the Decoder first inputs the feature map output by the fourth stage into the attention convergence module to reduce attention dispersion, and then performs step-by-step feature fusion on the feature map output by the attention convergence module. The step-by-step feature fusion splices the feature map output by the fourth stage through the attention convergence module with the feature map output by the third stage through the attention convergence module, and then reduces the number of channels of the spliced feature map by half through the convolution operation, and then splices it with the feature map output by the second stage through the attention convergence module, and then reduces the number of channels of the spliced feature map by half through the convolution operation to obtain the feature map, and then splices it with the feature map output by the first stage through the attention convergence module, and then reduces the number of channels of the spliced feature map by half through the convolution operation to obtain the feature map; finally, the feature map obtained by the step-by-step feature fusion is input into the linear prediction module to complete the bone imaging lesion segmentation.
Citation Information
Patent Citations
Bone imaging bone focus segmentation method, system and equipment based on deep neural network
CN115019049A