Lumbar cancellous bone CT value automatic calculation method and system based on deep learning
Through deep learning image classification and segmentation model, the CT value of lumbar cancellous bones is automatically calculated, which solves the problem of time-consuming and labor-intensive and susceptible to human factors in traditional methods, and improves the computing efficiency and accuracy.
Patent Information
- Application Number
- CN202510358667.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-08
AI Technical Summary
The traditional lumbar cancellous bone CT value calculation method relies on manual operation, which is time-consuming and labor-intensive and susceptible to human factors, making the accuracy and consistency of the results difficult to guarantee.
Using a deep learning-based method, image classification is performed through the Swin Transformer architecture, combining threshold segmentation and image segmentation models, the CT value of lumbar cancellous bone is automatically calculated, including image preprocessing, threshold segmentation and image segmentation, and finally the average pixel value of the lumbar segment is calculated.
The automated calculation of CT value of lumbar cancellous bone is realized, which improves the efficiency and accuracy of medical imaging classification and segmentation, and reduces artificial errors.
Smart Images

Figure CN120279045A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of medical image classification and segmentation, and particularly to an automatic calculation method and system for lumbar cancellous bone CT values based on deep learning. Background Art
[0002] Clinically, CT examination is a common and important diagnostic method. The CT value of lumbar cancellous bone can help doctors judge the tissue structure and pathological state of the lumbar spine, and can also help doctors judge the bone mass of patients to assist in the diagnosis of osteoporosis. Traditional methods for calculating the CT value of lumbar cancellous bone often rely on manual operations, including manually processing Dicom (Digital Imaging and Communications in Medicine) format data, including selecting different images, manually marking the ROI (region of interest) of each image, and calculating the average CT value of the ROI. This process is not only time-consuming and laborious, but also easily affected by human factors, making it difficult to ensure the accuracy and consistency of the results. Summary of the Invention
[0003] This application provides an automatic calculation method and system for lumbar cancellous bone CT values based on deep learning to at least partially solve the above problems.
[0004] In the first aspect of this application, an automatic calculation method for lumbar cancellous bone CT values based on deep learning is provided. The method includes: Obtain the original Dicom data of the target object, adjust the window width and window level, and convert it into a CT image in PNG format; Input the CT image of the target object into a pre-trained image classification model to respectively determine the lumbar CT images of multiple lumbar segments; Process the lumbar CT image based on a threshold segmentation module to obtain a threshold segmentation image; Input the threshold segmentation image as a prompt message into a pre-trained image segmentation model. The image segmentation model processes the lumbar CT image based on the prompt message to determine the lumbar cancellous bone region in the lumbar CT image; Based on the lumbar cancellous bone regions in the lumbar CT images of multiple lumbar segments, respectively determine the average pixel values of the lumbar cancellous bone regions of multiple lumbar segments to obtain the lumbar cancellous bone CT value results of each lumbar segment.
[0005] Optionally, the image classification model is of the Swin Transformer architecture. The image classification model includes a plurality of Swin Transformer blocks, and each Swin Transformer block includes: a window-based multi-head self-attention operation unit, a shifted window multi-head self-attention operation unit, and a multi-layer perceptron. The multi-head self-attention operation unit divides the input feature map into non-overlapping windows and calculates self-attention within each window. The shifted window multi-head self-attention operation unit performs a window shifting operation based on the multi-head self-attention operation unit. By shifting the windows, information interaction between different windows is achieved to obtain global feature information. The image classification model includes multiple stages. As the stages progress, the size of the feature maps input to the Swin Transformer blocks in each stage gradually decreases, and the number of channels gradually increases. The image classification model is trained based on sample CT images carrying classification labels, and the classification labels include: thoracic vertebra, lumbar vertebra, intervertebral disc, and others.
[0006] Optionally, processing the lumbar CT image by a threshold segmentation module to obtain a threshold segmentation image, including: For each pixel in the lumbar CT image, comparing the gray value of each pixel with a preset threshold; In the case where the gray value of a certain pixel is greater than or equal to the preset threshold, the pixel is determined as a target pixel and assigned a value of 1. In the case where the gray value of a certain pixel is less than the preset threshold, the pixel is determined as a background pixel and assigned a value of 0, obtaining a binary threshold segmentation image in which the target and the background are distinguished.
[0007] Optionally, the image segmentation model includes: an image encoder, a prompt encoder, and a mask decoder. Inputting the threshold segmentation image as prompt information into a pre-trained image segmentation model, and processing the lumbar CT image by the image segmentation model based on the prompt information to determine the lumbar cancellous bone region in the lumbar CT image, including: The image segmentation model extracts the feature representation of the lumbar CT image through the image encoder, extracts the feature representation of the threshold segmentation image through the prompt encoder, and fuses the outputs of the image encoder and the prompt encoder through the mask decoder to obtain the lumbar cancellous bone region in the lumbar CT image.
[0008] Optionally, before inputting the lumbar CT image into the image segmentation model, the method further includes: Adjusting the size of the lumbar CT image to meet the input requirements of the model; Normalizing the image intensity value of the lumbar CT image; Remove the noise or artifacts from the lumbar spine CT images.
[0009] Optionally, based on the lumbar cancellous bone regions in the lumbar spine CT images of each of the multiple lumbar spine segments, respectively determine the average pixel values of the lumbar cancellous bone regions of each of the multiple lumbar spine segments, and obtain the lumbar cancellous bone CT value results for each lumbar spine segment, including: Based on the original CT image data in Dicom format corresponding to the lumbar spine CT images of each of the multiple lumbar spine segments, respectively determine the average pixel values of the lumbar cancellous bone regions of each of the multiple lumbar spine segments, and obtain the lumbar cancellous bone CT value results for each lumbar spine segment.
[0010] The second aspect of the present application provides a deep learning-based automatic lumbar cancellous bone CT value calculation system, and the deep learning-based automatic lumbar cancellous bone CT value calculation system includes: An acquisition module, configured to acquire the original Dicom data of the target object, adjust the window width and window level, and convert it into a CT image in PNG format; A classification module, configured to input the CT image of the target object into a pre-trained image classification model, and respectively determine the lumbar spine CT images of each of the multiple lumbar spine segments; A threshold segmentation module, configured to process the lumbar spine CT image based on the threshold segmentation module to obtain a threshold segmentation image; A lumbar cancellous bone region segmentation module, configured to input the threshold segmentation image as a prompt message into a pre-trained image segmentation model, and process the lumbar spine CT image based on the prompt message through the image segmentation model to determine the lumbar cancellous bone region in the lumbar spine CT image; A calculation module, configured to respectively determine the average pixel values of the lumbar cancellous bone regions of each of the multiple lumbar spine segments based on the lumbar cancellous bone regions in the lumbar spine CT images of each of the multiple lumbar spine segments, and obtain the lumbar cancellous bone CT value results for each lumbar spine segment.
[0011] The third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes, it implements the deep learning-based automatic lumbar cancellous bone CT value calculation method as described in the first aspect of the present invention.
[0012] The fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by the processor, it implements the deep learning-based automatic lumbar cancellous bone CT value calculation method as described in the first aspect of the present invention.
[0013] The fifth aspect of the present application provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by the processor, it implements the steps in the deep learning-based automatic lumbar cancellous bone CT value calculation method as described in the first aspect of the present invention.
[0014] In this application, first, lumbar CT images corresponding to multiple lumbar segments can be classified from a large number of CT images of the target object. Further, threshold segmentation is performed on the lumbar CT images based on the image characteristics of the lumbar part to obtain threshold segmentation images. Then, the lumbar CT images are used as inputs, and the threshold segmentation images are used as prompt information and input into an image segmentation model to segment the lumbar cancellous bone region. Finally, based on the segmentation results, the average pixel values of the lumbar cancellous bone regions of each lumbar segment are calculated and returned to the original lumbar CT images, and the lumbar cancellous bone CT value results of each lumbar segment are finally obtained. Thus, this application can implement an integrated, deep learning-based method for calculating the lumbar cancellous bone CT value, automatically process the original Dicom data, and accurately calculate the lumbar cancellous bone CT value through a deep learning model, so as to improve the efficiency and accuracy of medical image classification and segmentation. Brief Description of the Drawings
[0015] To more clearly illustrate the technical solutions of this application, the drawings required for the description of this application will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 is a flowchart of the steps of the method for automatically calculating the lumbar cancellous bone CT value based on deep learning provided by this application; Figure 2 is a schematic structural diagram of picture classification based on an image classification model in the method for automatically calculating the lumbar cancellous bone CT value based on deep learning provided by this application; Figure 3 is a schematic structural diagram of the image segmentation model in the method for automatically calculating the lumbar cancellous bone CT value based on deep learning provided by this application. Detailed Description of the Embodiments
[0017] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0018] Currently, with the rapid development of computer technology and deep learning, it has gradually become possible to use artificial intelligence algorithms to assist medical image analysis. Deep learning models have demonstrated excellent performance in image recognition, classification, and segmentation. For example, the Transformer model can effectively capture local and global features of images in various visual tasks and has been widely applied in various fields.
[0019] However, there is currently a lack of an integrated, deep learning-based automatic calculation method for the CT value of lumbar cancellous bone. In the related technologies, either the focus is on image preprocessing or on image segmentation, and the entire process from data preprocessing to the final CT value calculation has not been effectively integrated. Therefore, this application proposes a method that can automatically process the original Dicom data and accurately calculate the CT value of lumbar cancellous bone through a deep learning model, so as to improve the efficiency and accuracy of medical image classification and segmentation.
[0020] Specifically, the step process of the deep learning-based automatic calculation method for the CT value of lumbar cancellous bone proposed in this application is as Figure 1 shown, and the method includes the following steps: S101, obtain the original Dicom data of the target object, adjust the window width and window level, and convert it into a CT image in PNG format.
[0021] S102, input the CT image of the target object into a pre-trained image classification model to respectively determine the lumbar CT images of multiple lumbar segments.
[0022] S103, process the lumbar CT image based on the threshold segmentation module to obtain a threshold segmentation image.
[0023] S104, input the threshold segmentation image as a prompt message into a pre-trained image segmentation model, and process the lumbar CT image based on the prompt message through the image segmentation model to determine the lumbar cancellous bone region in the lumbar CT image.
[0024] S105, based on the lumbar cancellous bone regions in the lumbar CT images of multiple lumbar segments, respectively determine the average pixel values of the lumbar cancellous bone regions of multiple lumbar segments to obtain the CT value results of the lumbar cancellous bone of each lumbar segment.
[0025] In this application, in step S101, the CT image of the target object is obtained from the lumbar CT imaging data of the patient (i.e., the target object) collected by the hospital CT device. Generally speaking, there are multiple CT images of the patient obtained, including multiple lumbar CT images and CT images of other parts (such as: thoracic vertebrae, intervertebral discs, etc.).
[0026] In this application, the lumbar CT image is an image of the cross-section of the lumbar vertebrae of the target object.
[0027] In this application, in step S102, the image data can be input into a pre-trained image classification model for image classification, and the category to which each CT image belongs can be obtained therefrom. Specifically, the categories to which the CT images belong include: thoracic vertebra, lumbar vertebra, intervertebral disc, and others. Among them, the CT images of other categories include all other images that do not belong to the above three categories (thoracic vertebra, lumbar vertebra, intervertebral disc).
[0028] Specifically, the image classification model is of the Swin Transformer architecture. The image classification model includes a plurality of Swin Transformer blocks, and each Swin Transformer block includes: a window-based multi-head self-attention operation unit, a shifted-window multi-head self-attention operation unit, and a multi-layer perceptron. The multi-head self-attention operation unit divides the input feature map into non-overlapping windows and calculates self-attention within each window. The shifted-window multi-head self-attention operation unit performs a window shifting operation on the basis of the multi-head self-attention operation unit. By shifting the windows, information interaction between different windows is realized, and global feature information is obtained. The image classification model includes multiple stages. As the stages progress, the size of the feature map input to the Swin Transformer blocks in each stage gradually decreases, and the number of channels gradually increases.
[0029] In this application, the Swin Transformer architecture effectively reduces the computational complexity through the window-based multi-head self-attention mechanism and the shifted-window operation, and at the same time can capture the local and global features of the image. This architecture performs excellently in computer vision tasks, especially having great advantages when dealing with tasks that require fine feature analysis such as medical images. Figure 2 is a schematic diagram of the architecture for picture classification based on the image classification model in this application. As Figure 2 shown, the entire architecture of the image classification model consists of a plurality of Swin Transformer blocks, and these blocks are divided into four stages (Stage 1 - Stage 4).
[0030] In each stage, the size of the feature map gradually decreases, and the number of channels gradually increases. Inside each Swin Transformer block, there are two main operations: W-MSA (Window-based Multi-Head Self-Attention) and SW-MSA (Shifted Window-based Multi-Head Self-Attention), as well as a multi-layer perceptron (MLP). The W-MSA operation performs multi-head self-attention calculation within each window. It divides the input feature map into non-overlapping windows and calculates self-attention within each window. This local attention mechanism reduces the computational complexity while being able to capture local features. Let the input feature map be , where M and N are the height and width of the feature map, and C is the number of channels. W-MSA divides the feature map X into multiple windows and performs self-attention calculation within each window. SW-MSA performs a window shifting operation based on W-MSA. By shifting the windows by a certain amount, information interaction between different windows is achieved, thereby obtaining global feature information. This operation enables the model to capture both local features and consider global information, making it suitable for handling long-range dependencies in images. The input image enters Stage 1 after passing through a linear embedding layer. This stage contains -sized feature maps, where C is the number of channels. This stage mainly performs preliminary processing and feature extraction on the input features. In Stages 2 - 4, the size of the feature maps gradually decreases to , and respectively. As the stages progress, the number of channels of the feature maps increases while the size decreases, which helps the model gradually capture more advanced semantic features.
[0031] The role of the global average pooling layer is to average the spatial features of each channel and output a vector with a length equal to the number of channels, thereby effectively reducing the number of parameters in the model and retaining important global information.
[0032] The classification head is the last layer of the model, which is used to map the feature vector to the category space to achieve the final classification task. In Swin Transformer, the Classification Head is usually a linear layer, and its functions include: (1) Feature mapping: Map the pooled feature vector to the category space, with each dimension corresponding to a category. (2) Output class probability: Convert the output of the linear layer into a probability distribution through the softmax function, representing the probability that the input image belongs to each category.
[0033] In this application, the image classification model is trained based on sample CT images carrying classification labels, and the classification labels include: thoracic vertebra, lumbar vertebra, intervertebral disc, and others.
[0034] Specifically, in this application, CT image data of multiple patients can be obtained to construct sample CT images. Specifically, the sequence of CT image data in Dicom format of multiple patients obtained is converted into sample CT images in PNG file format, and then professional medical image annotation tools are operated by professionals to manually annotate the CT images. All the pictures are divided into four categories: thoracic vertebra, lumbar vertebra, intervertebral disc, and others, and at the same time, pictures of different categories are saved in different folders as the classification data set.
[0035] During the training process of the image classification model, using the already constructed classification data set, fine-tuning is performed on a Swin-transfomer pre-trained on the ImageNet data set to improve the classification effect of the model. In the medical image processing flow of this application, the fine-tuning of the Swin-Transformer model is extremely crucial, and reasonable parameter setting is of utmost importance. In this application, the learning rate is initially set to 1×10 ⁻4 , combined with the strategy of decaying to 0.9 times every 5 Epochs, to make the model converge steadily; the number of training rounds is set between 20 - 50, and training is stopped in a timely manner according to the validation accuracy; the weight decay coefficient is about 0.01 to prevent overfitting.
[0036] In the present application, in step S102, the CT image of the target object includes a plurality of continuous images obtained for the spine of the target object, and for all the continuous CT images of the target object, the image classification model can output the category of each CT image. In particular, in the case where the order of the plurality of CT images corresponds to the order of spine connection, considering the actual situation that different lumbar segments are divided by lumbar discs, a plurality of lumbar segments can be obtained based on the image classification results of the continuous CT images. For example: the image classification results of the 21st to 40th CT images are all lumbar vertebrae, the image classification results of the 41st to 60th CT images are intervertebral discs, and the image classification results of the 61st to 80th CT images are lumbar vertebrae. Then, based on the 41st to 60th intervertebral disc CT images, the 21st to 40th lumbar vertebrae CT images can be used as the lumbar vertebrae CT images corresponding to the first lumbar vertebrae segment, and the 61st to 80th lumbar vertebrae CT images can be used as the lumbar vertebrae CT images corresponding to the second lumbar vertebrae segment. The lumbar vertebrae CT image division method of other lumbar vertebrae segments is similar.
[0037] In the present application, in step S103, considering that in the field of medical image processing, especially for the identification and segmentation of the lumbar region, the threshold segmentation method is often used due to its simplicity and effectiveness. Since the tissue density of the lumbar region is usually significantly different from the surrounding background (such as muscle, fat, etc.), this causes the lumbar region to present a unique grayscale value range in the grayscale image. In the original image, the lumbar region is surrounded by a large number of background details, and the grayscale value of the lumbar region is significantly different from the grayscale value of the background part. Taking advantage of this, the present application proposes a threshold segmentation method, for the lumbar CT image classified by the image classification model, it is processed by the threshold segmentation module to obtain the image after threshold segmentation, as the input of the prompt encoder in the subsequent image segmentation model. The threshold segmentation module preliminarily segments the entire lumbar region, which plays a preliminary segmentation effect, so as to improve the segmentation effect of the subsequent image segmentation model on the cancellous bone region. After threshold segmentation, the image becomes a binary image. The lumbar region is extracted as a whole and appears white, while the rest of the background area is set to black.
[0038] Specifically, step S103 includes: S1031, for each pixel in the lumbar CT image, compare the grayscale value of each pixel with a preset threshold.
[0039] S1032, when the grayscale value of a certain pixel is greater than or equal to the preset threshold, the pixel is determined to be a target pixel and assigned a value of 1. When the grayscale value of a certain pixel is less than the preset threshold, the pixel is determined to be a background pixel and assigned a value of 0. A binary threshold segmentation image is obtained, in which the target and background are distinguished.
[0040] In this application, for each pixel in the lumbar spine CT image, its gray value is compared with a preset threshold. If the gray value of the pixel is greater than or equal to the preset threshold, the pixel is determined to be a target pixel and assigned a value of 1. If the gray value of the pixel is less than the preset threshold, the pixel is determined to be a background pixel and assigned a value of 0. In this way, a binary image is obtained, in which the target and the background are clearly distinguished. Selecting an appropriate threshold is the key to threshold segmentation. In this application, a method with a preset threshold of 90 can be used for segmentation. After obtaining the image after threshold segmentation, it is input as the prompt information of the image segmentation model to assist the segmentation task. This preprocessing step can help the image segmentation model more accurately locate the lumbar spine area, reduce the search space, and thus improve the accuracy and efficiency of segmentation.
[0041] This application uses efficient threshold segmentation and contour extraction techniques, specifically including: using the cv2.threshold function built into the OpenCV library (an open-source library widely used in the field of computer vision) to precisely control the image binarization process. Taking the set threshold (90 in the example code) as the boundary, pixels with pixel values greater than this threshold are classified into one category (set to 255, i.e., white), and pixels less than the threshold are set to another category (set to 0, i.e., black), thereby quickly highlighting the key areas of the image, eliminating unnecessary details, achieving preliminary image simplification and feature extraction, effectively compressing the image information, and improving the efficiency of subsequent processing. Using the cv2.findContours function, combined with the cv2.RETR_TREE retrieval mode, to comprehensively extract all the contours in the image, including nested contours, and construct a complete contour hierarchy; the cv2.CHAIN_APPROX_SIMPLE parameter streamlines the contour storage information, eliminates redundant points, and saves memory resources. Using cv2.moments to calculate the contour center, combined with the position of the center and the contour area calculated by cv2.contourArea, to accurately screen out the contours of the entire lumbar spine area and exclude interfering contours at the edges or that are too small. After screening out the target contours, use the cv2.drawContours function to draw the contours that meet the conditions onto a newly created blank image, fill them white as set, and highlight the key contour information; finally, use cv2.imwrite to save the images before and after processing, which is convenient for backtracking and comparison, and retains key data for subsequent image analysis-based processes, ensuring the repeatability and stability of this application in different application scenarios.
[0042] In this application, in step S104, image segmentation is a crucial component in clinical practice, which helps with accurate diagnosis, treatment, and disease monitoring. In this application, an image segmentation model is trained using the SAM model. The difference between the SAM model and ordinary deep learning models based on neural networks or transformers is that it is a prompt-based image segmentation model. By inputting prompts (which can include various types of prompts such as point, box, mask, etc.) to it, the model can more accurately segment the target area. In this application, by constructing a mask prompt through the threshold segmentation module, the SAM segmentation model can focus more on the lumbar region and will not pay attention to the remaining background areas. This design greatly improves the segmentation effect of lumbar images.
[0043] In this application, the image segmentation model includes: an image encoder, a prompt encoder, and a mask decoder. Specifically, step S104 includes: the image segmentation model extracts the feature representation of the lumbar CT image through the image encoder, extracts the feature representation of the threshold segmentation image through the prompt encoder, and fuses the outputs of the image encoder and the prompt encoder through the mask decoder to obtain the lumbar cancellous bone region in the lumbar CT image.
[0044] Specifically, Image Encoder: In this application, the image classification model uses an image encoder based on Vision Transformer (ViT). This encoder can extract rich feature representations from the input medical images.
[0045] In this application, the ViT-base model can be specifically used, which contains 12 Transformer layers. Each layer consists of a multi-head self-attention (MHSA) block and a multi-layer perceptron (MLP) block, and is optimized through layer normalization technology.
[0046] The image encoder maps the input image to a high-dimensional image embedding space, providing deep-level feature support for subsequent segmentation tasks.
[0047] Prompt Encoder: The function of the prompt encoder is to convert prompt information, such as the entire threshold segmentation image segmented by the threshold in this application, that is, the lumbar region mask, into a feature representation that the model can understand, and assist the mask decoder in decoding the image.
[0048] Mask Decoder: The mask decoder is responsible for fusing the outputs of the image encoder and the prompt encoder to generate the final segmentation result.
[0049] This component contains two Transformer layers for integrating image embeddings and prompt encodings, and two transposed convolutional layers for upsampling the embedding resolution to 256×256.
[0050] After sigmoid activation and bilinear interpolation, the mask decoder outputs a segmentation result with the same size as the input image.
[0051] The design of the SAM model allows it to achieve efficient and accurate segmentation with the assistance of prompt information, and its generalization ability makes it of great value in practical applications. By combining a ViT-based image encoder, a prompt encoder, and a lightweight mask decoder, SAM can efficiently and accurately segment the cancellous bone region in the lumbar spine.
[0052] In this application, before inputting the lumbar spine CT image into the image segmentation model, necessary preprocessing is first performed on the image. This includes resizing the image to meet the input requirements of the model, normalizing the image intensity values, and removing possible noise or artifacts.
[0053] Figure 3 The schematic architecture diagram of the image segmentation model of this application is shown, as Figure 3 shown, in step S104, the preprocessed lumbar spine image is input into the image encoder. The image encoder is based on the Vision Transformer (ViT) architecture and is responsible for extracting high-level feature representations from the image. These features contain the key anatomical information of the lumbar spine region. The image encoder is processed through multiple layers of Transformer, each layer including a self-attention mechanism and a multi-layer perceptron (MLP), to capture the complex structure of the lumbar spine. After being processed by the image encoder, an image embedding is generated, which is a high-dimensional feature vector that encodes the detailed features of the lumbar spine image.
[0054] After receiving the prompt information, that is, the threshold-segmented picture, the prompt encoder converts it into a dense embedding. This embedding is combined with the image embedding to guide the model to focus on the user-specified lumbar spine region. Finally, a good output result is obtained after the model's processing.
[0055] In this application, the training process of the image segmentation model specifically includes: First, obtain sample lumbar spine CT images, and the annotators mark the target region, that is, the cancellous bone region of the lumbar spine, by performing specific operations on the images. For the cancellous bone region of the lumbar spine to be analyzed and its related tissues, use a polygon drawing tool to precisely outline the contour of the target region. This process must strictly follow medical knowledge to ensure an accurate division of the boundary between the target tissue and the surrounding structures.
[0056] Then, the classified sample lumbar CT images are used as the input of the image encoder to obtain the image embedding, and the mask after threshold segmentation is input as prompt information into the prompt encoder to obtain the dense embedding. Then the image embedding and dense embedding are input into the decoder for decoding, and finally the segmentation result is obtained. Based on the final segmentation result and the target area annotation corresponding to the sample lumbar CT image, the image classification model is trained.
[0057] In this application, SAM model training and fine-tuning are crucial, and reasonable parameter setting is the key. The initial learning rate is set to 1×10 ⁻4 , combined with the cosine annealing strategy, decays every 10 epochs to help the model converge stably; the training rounds are set to 30-60, and training is stopped in time according to the verification indicators; the weight decay coefficient is about 0.005-0.01 to prevent overfitting.
[0058] In the present application, in step S101, CT image data of the target object can be acquired, the original CT image data in the Dicom format can be preprocessed, the appropriate window width and window position can be adjusted, and it can be converted into a CT image in the PNG format. Correspondingly, in step S105, based on the lumbar cancellous bone area in the lumbar CT images of each of the multiple lumbar segments, the average pixel value of the lumbar cancellous bone area of each of the multiple lumbar segments is determined respectively, and the lumbar cancellous bone CT value results of each lumbar segment are obtained.
[0059] The method for automatically calculating the lumbar cancellous bone CT value based on deep learning proposed in this application includes: preprocessing the original Dicom format CT image data, including adjusting the appropriate window width and window position, and converting it into a PNG format image. Then the image data is input into the image classification model for image classification, so as to screen out the lumbar CT images of multiple lumbar segments. Next, the lumbar CT image is processed by the threshold segmentation module to obtain the image after threshold segmentation, that is, the entire lumbar region, and use it as the prompt information based on the image segmentation model. Then, the SAM-based image segmentation model is used to segment the lumbar part image under the above prompt to segment the cancellous bone area in the lumbar region. After the segmentation is completed, return to the original Dicom file, use the segmentation result to calculate the average pixel value of the cancellous bone area of each lumbar image, and determine the average pixel value of the lumbar cancellous bone area of each lumbar segment as the output, and finally obtain the lumbar cancellous bone CT value result of the patient.
[0060] Based on the same inventive concept, the present application also provides a deep learning-based automatic lumbar cancellous bone CT value calculation system, and the deep learning-based automatic lumbar cancellous bone CT value calculation system includes: An acquisition module, configured to input the CT image of the target object into a pre-trained image classification model to respectively determine the lumbar CT images of multiple lumbar segments. A classification module, configured to input the CT image of the target object into a pre-trained image classification model to determine the lumbar CT image including the lumbar part. A threshold segmentation module, configured to process the lumbar CT image based on the threshold segmentation module to obtain a threshold segmentation image. A lumbar cancellous bone region segmentation module, configured to input the threshold segmentation image as prompt information into a pre-trained image segmentation model, and process the lumbar CT image based on the prompt information through the image segmentation model to determine the lumbar cancellous bone region in the lumbar CT image. A calculation module, configured to respectively determine the average pixel value of the lumbar cancellous bone region of each lumbar segment based on the lumbar cancellous bone region in the lumbar CT images of multiple lumbar segments, and obtain the lumbar cancellous bone CT value result of each lumbar segment.
[0061] Optionally, the image classification model is of the Swin Transformer architecture, and the image classification model includes multiple Swin Transformer blocks. Each Swin Transformer block includes: a window-based multi-head self-attention operation unit, a shifted window multi-head self-attention operation unit, and a multi-layer perceptron; the multi-head self-attention operation unit divides the input feature map into non-overlapping windows and calculates self-attention within each window; the shifted window multi-head self-attention operation unit performs a window shifting operation on the basis of the multi-head self-attention operation unit; by shifting the windows, information interaction between different windows is realized to obtain global feature information. The image classification model includes multiple stages. As the stages progress, the size of the feature map input to the Swin Transformer blocks of each stage gradually decreases, and the number of channels gradually increases; the image classification model is trained based on sample CT images carrying classification labels, and the classification labels include: thoracic vertebra, lumbar vertebra, intervertebral disc, and others.
[0062] Optionally, the threshold segmentation module is configured to: For each pixel in the lumbar CT image, compare the gray value of each pixel with a preset threshold. When the gray value of a certain pixel is greater than or equal to a preset threshold, the pixel is determined as a target pixel and assigned a value of 1. When the gray value of a certain pixel is less than the preset threshold, the pixel is determined as a background pixel and assigned a value of 0, obtaining a binary threshold segmentation image in which the target and the background are distinguished.
[0063] Optionally, the image segmentation model includes: an image encoder, a prompt encoder, and a mask decoder. The lumbar cancellous bone region segmentation module is used for: The image segmentation model extracts the feature representation of the lumbar CT image through the image encoder, extracts the feature representation of the threshold segmentation image through the prompt encoder, and fuses the outputs of the image encoder and the prompt encoder through the mask decoder to obtain the lumbar cancellous bone region in the lumbar CT image.
[0064] Optionally, before inputting the lumbar CT image into the image segmentation model, the system further includes: A preprocessing module, which is used to adjust the size of the lumbar CT image to meet the input requirements of the model; normalize the image intensity value of the lumbar CT image; remove the noise or artifacts in the lumbar CT image.
[0065] Optionally, the acquisition module is used for: collecting the CT image data of the target object, preprocessing the original CT image data in Dicom format, adjusting the appropriate window width and window level, and converting it into a CT image in PNG format; The calculation module is used for: respectively determining the average pixel value of the lumbar cancellous bone region of each lumbar segment based on the original Dicom format CT image data corresponding to the lumbar CT images of each lumbar segment, and obtaining the lumbar cancellous bone CT value result of each lumbar segment.
[0066] Based on the same inventive concept, the present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the steps in the automatic calculation method of lumbar cancellous bone CT value based on deep learning as described in any one of the above embodiments.
[0067] Based on the same inventive concept, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in the automatic calculation method of lumbar cancellous bone CT value based on deep learning as described in any one of the above embodiments.
[0068] Based on the same inventive concept, the present application provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps in the automatic calculation method of lumbar cancellous bone CT value based on deep learning described in any of the above embodiments.
[0069] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is the difference from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0070] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0071] The present application is described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable terminal devices generate a system for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0072] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0073] These computer program instructions can also be loaded onto a computer or other programmable terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0074] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0075] Finally, it should also be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the element.
[0076] The above has introduced in detail a method for automatically calculating the CT value of lumbar cancellous bone based on deep learning provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An automatic calculation method for CT values of lumbar cancellous bone based on deep learning, characterized in that, The method includes: Obtain the original Dicom data of the target object, adjust the window width and window level, and convert it into a CT image in PNG format; Input the CT image of the target object into a pre-trained image classification model to respectively determine the lumbar CT images of multiple lumbar segments; Process the lumbar CT image based on a threshold segmentation module to obtain a threshold segmentation image; Input the threshold segmentation image as prompt information into a pre-trained image segmentation model, and process the lumbar CT image by the image segmentation model based on the prompt information to determine the lumbar cancellous bone region in the lumbar CT image; Based on the lumbar cancellous bone regions in the lumbar CT images of multiple lumbar segments, respectively determine the average pixel values of the lumbar cancellous bone regions of multiple lumbar segments to obtain the lumbar cancellous bone CT value results of each lumbar segment.
2. The automatic calculation method of lumbar cancellous bone CT value based on deep learning according to claim 1, characterized in that The image classification model is of the Swin Transformer architecture. The image classification model includes multiple Swin Transformer blocks. Each Swin Transformer block includes: a window-based multi-head self-attention operation unit, a shifted window multi-head self-attention operation unit, and a multi-layer perceptron; the multi-head self-attention operation unit divides the input feature map into non-overlapping windows and calculates self-attention within each window; the shifted window multi-head self-attention operation unit performs a window shifting operation on the basis of the multi-head self-attention operation unit; by shifting the windows, information interaction between different windows is achieved to obtain global feature information; The image classification model includes multiple stages. As the stages progress, the size of the feature maps input to the Swin Transformer blocks in each stage gradually decreases, and the number of channels gradually increases; The image classification model is trained based on sample CT images carrying classification labels. The classification labels include: thoracic vertebra, lumbar vertebra, intervertebral disc, and others.
3. The automatic calculation method of lumbar cancellous bone CT value based on deep learning according to claim 1, characterized in that, Processing the lumbar CT image based on a threshold segmentation module to obtain a threshold segmentation image includes: For each pixel in the lumbar CT image, compare the gray value of each pixel with a preset threshold; In the case where the gray value of a certain pixel is greater than or equal to the preset threshold, determine the pixel as a target pixel and assign a value of 1. In the case where the gray value of a certain pixel is less than the preset threshold, determine the pixel as a background pixel and assign a value of 0 to obtain a binary threshold segmentation image, in which the target and the background are distinguished; 4. The automatic calculation method of lumbar cancellous bone CT value based on deep learning according to claim 1, characterized in that, The image segmentation model includes: an image encoder, a prompt encoder, and a mask decoder. Inputting the threshold segmentation image as prompt information into a pre-trained image segmentation model, and processing the lumbar CT image by the image segmentation model based on the prompt information to determine the lumbar cancellous bone region in the lumbar CT image includes: The image segmentation model extracts the feature representation of the lumbar spine CT image through an image encoder, extracts the feature representation of the threshold segmentation image through a prompt encoder, and fuses the outputs of the image encoder and the prompt encoder through a mask decoder to obtain the lumbar cancellous bone region in the lumbar spine CT image.
5. The automatic calculation method of lumbar cancellous bone CT value based on deep learning according to any one of claims 1-4, characterized in that Before inputting the lumbar spine CT image into the image segmentation model, the method further includes: Adjusting the size of the lumbar spine CT image to meet the input requirements of the model; Normalizing the image intensity values of the lumbar spine CT image; Removing the noise or artifacts in the lumbar spine CT image.
6. The automatic calculation method of lumbar cancellous bone CT value based on deep learning according to claim 5, characterized in that Based on the lumbar cancellous bone regions in the lumbar spine CT images of multiple lumbar segments respectively, determining the average pixel values of the lumbar cancellous bone regions of multiple lumbar segments respectively, and obtaining the lumbar cancellous bone CT value results of each lumbar segment, including: Based on the original CT image data in DICOM format corresponding to the lumbar spine CT images of multiple lumbar segments respectively, determining the average pixel values of the lumbar cancellous bone regions of multiple lumbar segments respectively, and obtaining the lumbar cancellous bone CT value results of each lumbar segment.
7. An automatic CT value calculation system for lumbar cancellous bone based on deep learning, characterized in that, The deep learning-based automatic lumbar cancellous bone CT value calculation system includes: An acquisition module, configured to acquire the original DICOM data of the target object, adjust the window width and window level, and convert it into a CT image in PNG format; A classification module, configured to input the CT image of the target object into a pre-trained image classification model, and respectively determine the lumbar spine CT images of multiple lumbar segments; A threshold segmentation module, configured to process the lumbar spine CT image based on the threshold segmentation module to obtain a threshold segmentation image; A lumbar cancellous bone region segmentation module, configured to input the threshold segmentation image as prompt information into a pre-trained image segmentation model, and process the lumbar spine CT image based on the prompt information through the image segmentation model to determine the lumbar cancellous bone region in the lumbar spine CT image; A calculation module, configured to respectively determine the average pixel values of the lumbar cancellous bone regions of multiple lumbar segments based on the lumbar cancellous bone regions in the lumbar spine CT images of multiple lumbar segments, and obtain the lumbar cancellous bone CT value results of each lumbar segment.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the deep learning-based automatic lumbar cancellous bone CT value calculation method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the deep learning-based automatic lumbar cancellous bone CT value calculation method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, it implements the steps in the deep learning-based automatic lumbar cancellous bone CT value calculation method according to any one of claims 1 to 6.
Citation Information
Cited By
CT-based osteoporosis intelligent diagnosis method and system
CN120976221A