A method for detecting vertebral center points in CT images based on deep learning

By using a two-stage center point detection algorithm based on deep learning and utilizing the ISE-Vnet and Spatial-Generalized-3DUnet-DSNT networks, the automation and robustness problems of vertebral center point detection in CT images were solved, and high-precision vertebral center point detection was achieved.

CN117078658BActive Publication Date: 2025-09-23SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311210538.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-19
Publication Date
2025-09-23
Estimated Expiration
2043-09-19

AI Technical Summary

Technical Problem

Existing technologies make it difficult to efficiently and automatically detect the center points of vertebrae in CT images, especially for vertebrae with complex structures. Traditional methods are not robust enough to meet clinical needs.

Method used

A two-stage center point detection algorithm based on deep learning is adopted, including the ISE-Vnet segmentation network and the Spatial-Generalized-3DUnet-DSNT network. Through the coarse segmentation and center point detection stages, automated vertebral center point detection is achieved. ISE-Vnet is used for spine segmentation and sliding window generation, and the DSNT module is combined to improve the detection accuracy.

Benefits of technology

It achieves automated vertebral center point detection at high resolution, improves the robustness and accuracy of detection, is suitable for data containing noisy and deformed vertebrae, and reduces dependence on manual operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117078658B_ABST
    Figure CN117078658B_ABST
Patent Text Reader

Abstract

This paper proposes a deep learning-based method for vertebral center point detection in CT images. This model consists of two stages: a coarse segmentation stage and a center point detection stage. This two-stage model can detect the center points of any vertebra in spinal CT data at high resolution, without relying on manual operations such as selecting vertebral segments, thus achieving automated vertebral center point detection. Compared to other models, the proposed model offers significant advantages in detection accuracy and greater robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting vertebral center points in CT images based on deep learning, and belongs to the field of computer image processing. Background Art

[0002] With the rapid pace of our lives, the advent of computers and mobile phones, and the resulting sedentary work schedules, prolonged use of electronic devices, and lack of exercise, bone diseases are becoming increasingly common among both adolescents and middle-aged adults. Among these, spinal disorders are increasingly becoming a major ailment affecting people's work and daily lives. Statistics show that the prevalence of scoliosis among children in my country is as high as 20%, and 40% of people aged 40 and over experience spinal problems. Spinal deformities can severely impact a patient's future career and marriage. Long-term scoliosis can lead to thoracic and lumbar back pain, spinal stenosis, and nerve root compression, leading to various spinal-related diseases. Doctors use medical imaging technologies such as computed tomography (CT) or magnetic resonance imaging (MRI) to scan the affected areas of the patient's spine. Doctors then analyze these images through experience and manual measurement to diagnose spinal diseases and recommend treatment.

[0003] Computer-aided diagnosis (CAD) was first proposed in 1966. However, due to limitations in equipment and algorithms at the time, the level of CAD remained relatively low. Since the 21st century, rapid advances in computer vision, image processing, and machine learning have provided new computer-aided diagnosis methods for medical imaging, such as X-rays, CT scans, MRIs, and PET scans. This has significantly transformed clinical diagnostic methods, reducing the burden of manual diagnosis while improving diagnostic accuracy.

[0004] Among them, landmark point detection refers to the detection of a group of points with specific semantics in an image, which is of great significance to the fields of human parameter measurement, surgical navigation, etc., and vertebral landmark point detection is one of the most challenging tasks. Clinically, landmark points are usually manually labeled by experienced doctors or experts, which is very time-consuming and labor-intensive, and extremely dependent on the experience of the annotator. In the early days, some traditional algorithms completed landmark point detection through template registration methods, but their robustness was not high enough, and their performance could hardly meet today's clinical needs. Most of the deep learning algorithms that have emerged in recent years are based on two-dimensional images for landmark point detection, which cannot fully utilize the information of three-dimensional images such as CT and MRI, and it is difficult to obtain very good landmark point detection effects for more complex structures such as vertebrae. Summary of the Invention

[0005] This paper takes the vertebral center point detection of all vertebrae in three-dimensional CT images as an example of vertebral landmark point detection, proposes a two-stage center point detection algorithm based on deep learning, and proposes a solution on how to improve the performance of vertebral center point detection.

[0006] The present invention adopts the following technical solutions to solve the above problems:

[0007] The present invention provides a method for detecting vertebral center points in CT images based on deep learning, and the specific steps are as follows:

[0008] Stage 1: Coarse segmentation stage,

[0009] Step 1, data preprocessing;

[0010] The input CT data needs to be uniformly processed to ensure consistency of the data input to the neural network. Data processing includes four aspects. The first is to rotate the CT data to the same direction. The direction used in this paper is the RAS direction. Then the voxels are unified. The voxel size used in the coarse segmentation stage of this paper is 4mm. In addition, in order to keep the same size as the input to the neural network, the original image needs to be cropped or filled to unify the size to [96,96,128]. The last step is to normalize the image, which can effectively promote the convergence of the network.

[0011] Step 2: The ISE-Vnet segmentation network is proposed. After the data processing is completed, the data is input into the ISE-Vnet network, which can output the binary segmentation map of the spine.

[0012] Step 3: After completing the spine segmentation, this paper proposes a sliding window generation algorithm to generate multiple sliding window proposals of the same size based on the segmentation mask. These sliding windows can completely cover all vertebrae and some of the bones and tissues surrounding the vertebrae.

[0013] Step 4: Finally, the data contained in the sliding window suggestion box is input into the two-stage landmark detection algorithm.

[0014] Phase 2: Center point detection phase;

[0015] Step 5: Detect all landmark points in each sliding window. This task is completed by a center point detection network Spatial-Generalized-3DUnet-DSNT.

[0016] The network can be divided into two modules. The first module is the feature extraction network, which is responsible for feature extraction using a 3D-Unet, and a Spatial-Generalized module is used to filter out artifacts of irrelevant landmarks in the feature map output by the 3D-Unet. The second module is the result generation module, which is used to convert the feature map obtained by the feature extraction network into the final result. The introduction of the DSNT module converts the Gaussian heat map into coordinate values ​​through differentiable mathematical calculations, solving the problem of the theoretical lower bound of the error in the traditional Gaussian heat map conversion to coordinate values, thereby improving detection accuracy.

[0017] Step 6: Result Processing. Since not all output landmark points actually exist, the results need to be filtered to remove false positive predictions from the landmark detection network output and retain the results that are truly landmark points. Furthermore, since the coarse segmentation stage finally splits the CT data into multiple sliding windows, landmark detection is also performed within these multiple sliding windows. After removing false positive predictions, the results from these multiple sliding windows need to be merged to obtain the final landmark detection results for the original CT image.

[0018] In the above scheme, since the total number of pixels in the 3D CT data is too large to directly perform landmark detection, a two-stage vertebral center point detection algorithm process is proposed, which decomposes the vertebral center point detection problem of 3D CT data into a coarse segmentation stage and a landmark point detection stage.

[0019] The ISE-Vnet neural network proposed in the coarse segmentation stage is a V-net network improved by the spatial attention mechanism and the channel attention mechanism. This network can well segment the spine. After the segmentation is completed, the original CT image is divided into multiple images of the same size by a sliding window according to the original CT image and the segmentation mask, which serve as the input of the second stage.

[0020] The Spatial-Generalized-3DUnet-DSNT proposed in the center point detection stage learns and generates a spatial mask on the low-resolution feature map. It then multiplies the mask point-by-point with the Gaussian heat map generated by the high-resolution feature map to filter out artifacts in the Gaussian heat map. The DSNT module then generates the coordinates of the center point, solving the problem of inaccurate landmark coordinates obtained indirectly using only the Gaussian heat map.

[0021] The DSNT (Differentiable Spatial to Numerical Transform) module is a differentiable mathematical operation that converts Gaussian heatmaps into coordinate values ​​and is directly integrated after the Gaussian heatmap generation network without the need for a fully connected layer. Using the DSNT module not only retains the strong spatial generalization ability of the Gaussian heatmap, but also has many advantages of directly regressing coordinate points. The first step of the DSNT module is to normalize each Gaussian heatmap to ensure that the sum of the Gaussian heatmaps is 1. Normalization is performed using the softmax function:

[0022]

[0023] Define three 3D matrices X, Y, and Z, whose length, width, and height are consistent with the length, width, and height of the Heatmap input to DSNT. Their values ​​are: X i,j,k =i, Y i,j,k =j, Z i,j,k =k, H represents the output Gaussian heat map. The heat map after softmax normalization is recorded as. Since it is between 0 and 1 and the sum is 1, it satisfies the probability distribution condition and can be written as:

[0024]

[0025] c is the output coordinate of a channel, consisting of three numbers. The above formula is the joint probability distribution of random variables X, Y, and Z. The three coordinate values ​​obtained after DSNT transformation are the expectation of the above joint distribution, which is recorded as:

[0026] μ=E(c).

[0027] The formula for finding expectation from probability theory can be obtained:

[0028]

[0029] <A,B> F The F-norm of matrix A and matrix B, also known as the Frobenius inner product, is the inner product of the matrices and can be expressed as:

[0030]

[0031] To sum up, we can get the DNST expression:

[0032] DNST(H)=μ=[<H,X> F ,<H,Y> F ,<H,Z> F ].

[0033] Compared with the existing technology, the present invention uses the symmetry of brain images to perform feature comparison on the left and right cerebral hemispheres, thereby increasing the sensitivity to abnormal pixel value distribution in the region of interest and facilitating the detection of early ischemic changes. Its beneficial effects are as follows: The present invention can detect any vertebral landmark points containing spinal CT data at a higher resolution, without relying on manual operations such as manually selecting vertebral parts, and can achieve automated vertebral landmark point detection. Compared with traditional algorithms, it has higher robustness, and the proposed model can also perform well for data with noise, metal implants, and deformed vertebrae. The algorithm in the present invention is better than most of the currently excellent vertebral landmark point detection algorithms. The detection effect of landmark points of different vertebrae and different spinal parts is relatively average, and is almost unaffected by height and cross-sectional field of view. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic diagram of the entire process of the present invention;

[0035] Figure 2 This is a flow chart of the coarse segmentation stage;

[0036] Figure 3 This is the ISE-Vnet network structure diagram;

[0037] Figure 4 This is a flow chart of the center point detection phase;

[0038] Figure 5 This is the Spatial-Generalized-3DUnet-DSNT network structure diagram. DETAILED DESCRIPTION

[0039] The present invention is further illustrated below with reference to specific examples. It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0040] Example 1: See Figure 1 A method for assisting in the diagnosis of acute ischemic stroke based on CT plain scan images, the specific steps are as follows:

[0041] Stage 1: Coarse segmentation stage, see Figure 2 .

[0042] Step 1, data preprocessing;

[0043] The input CT data needs to be uniformly processed to ensure consistency of the data input to the neural network. Data processing includes four aspects. The first is to rotate the CT data to the same direction. The direction used in this paper is the RAS direction. Then the voxels are unified. The voxel size used in the coarse segmentation stage of this paper is 4mm. In addition, in order to keep the same size as the input to the neural network, the original image needs to be cropped or filled to unify the size to [96,96,128]. The last step is to normalize the image, which can effectively promote the convergence of the network.

[0044] Step 2, ISE-Vnet segmentation network is proposed, such as Figure 3 ,After the data processing is completed, the data is input into the ,ISE-Vnet network, which can output the binary segmentation map of the ,spine;

[0045] Step 3: After completing the spine segmentation, this paper proposes a sliding window generation algorithm to generate multiple sliding window proposals of the same size based on the segmentation mask. These sliding windows can completely cover all vertebrae and some of the bones and tissues surrounding the vertebrae.

[0046] Step 4: Finally, the data contained in the sliding window suggestion box is input into the two-stage landmark detection algorithm.

[0047] Stage 2: Center point detection stage, such as Figure 4

[0048] Step 5: Detect all landmarks in each sliding window. This task is completed by a center point detection network Spatial-Generalized-3DUnet-DSNT, as shown in Figure 5 .

[0049] The network can be divided into two modules. The first module is the feature extraction network, which is responsible for feature extraction using a 3D-Unet, and a Spatial-Generalized module is used to filter out artifacts of irrelevant landmarks in the feature map output by the 3D-Unet. The second module is the result generation module, which is used to convert the feature map obtained by the feature extraction network into the final result. The introduction of the DSNT module converts the Gaussian heat map into coordinate values ​​through differentiable mathematical calculations, solving the problem of the theoretical lower bound of the error in the traditional Gaussian heat map conversion to coordinate values, thereby improving detection accuracy.

[0050] Step 6: Result Processing. Since not all output landmark points actually exist, the results need to be filtered to remove false positive predictions from the landmark detection network output and retain the results that are truly landmark points. Furthermore, since the coarse segmentation stage finally splits the CT data into multiple sliding windows, landmark detection is also performed within these multiple sliding windows. After removing false positive predictions, the results from these multiple sliding windows need to be merged to obtain the final landmark detection results for the original CT image.

[0051] Effect evaluation:

[0052] This paper takes the vertebral center point detection of all vertebrae in three-dimensional CT images as an example of vertebral landmark point detection, proposes a two-stage center point detection algorithm based on deep learning, and provides an effective solution on how to improve the performance of vertebral center point detection.

[0053] It should be noted that the above embodiments are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Equivalent replacements or substitutions made on the basis of the above technical solutions all fall within the scope of protection of the present invention.

Claims

1. A method for detecting vertebral center points in CT images based on deep learning, characterized in that: The specific steps of the method are as follows: Stage 1: Coarse segmentation stage, as follows Step 1, data preprocessing; The input CT data needs to be uniformly processed to ensure consistency of the data input to the neural network. Data processing includes four aspects. The first is to rotate the CT data to the same direction, which is the RAS direction. Then the voxels are unified. The voxel size used in the coarse segmentation stage is 4mm. In addition, in order to keep the same size as the input to the neural network, the original image needs to be cropped or padded to unify the size to [96,96,128]. The last step is to normalize the image, which can effectively make the neural network converge during the training process. Step 2: The ISE-Vnet segmentation network is proposed. After the data processing is completed, the data is input into the ISE-Vnet network, which can output the binary segmentation map of the spine. Step 3: After completing the spine segmentation, a sliding window generation algorithm is used to generate multiple sliding window proposals of the same size based on the segmentation mask. These sliding windows can completely cover all vertebrae and some of the surrounding bones and tissues. Step 4: Finally, the data contained in the sliding window suggestion box is input into the two-stage landmark detection algorithm; Phase 2: Center point detection phase, details are as follows: Step 5: Detect all landmarks in each sliding window. This task is completed by a landmark detection network. The network is divided into two modules. The first module is the feature extraction network, which is responsible for feature extraction by a 3D-Unet, and a Spatial-Generalized module, which is a spatial overview module, is used to filter out artifacts of irrelevant landmarks in the feature map output by the 3D-Unet. The second module is the result generation module, which is used to convert the feature map obtained by the feature extraction network into the final result. The DSNT module is introduced to convert the Gaussian heat map into coordinate values ​​through differentiable mathematical calculations. Step 6: Result processing: After removing false positive predictions, the results on multiple sliding windows need to be merged to obtain the final landmark point detection results on the original CT image; The Spatial-Generalized-3DUnet-DSNT proposed in the center point detection stage learns and generates a spatial mask on the low-resolution feature map. It then multiplies the mask point-by-point with the Gaussian heat map generated by the high-resolution feature map to filter out artifacts in the Gaussian heat map. The DSNT module then generates the coordinates of the center point, solving the problem of inaccurate landmark coordinates obtained indirectly using only the Gaussian heat map. The DSNT (Differentiable Spatial to Numerical Transform) module is a differentiable mathematical operation that converts Gaussian heatmaps into coordinate values ​​and is directly integrated behind the Gaussian heatmap generation network. The first step of the DSNT module is to normalize each Gaussian heatmap to ensure that the sum of the Gaussian heatmaps is 1. The softmax function is used for normalization: Define three 3D matrices X, Y, and Z, whose length, width, and height are consistent with the length, width, and height of the Heatmap input to DSNT. Their values ​​are: X i,j,k =i, Y i,j,k =j, Z i,j,k =k, H represents the output Gaussian heat map, and the heat map after softmax normalization is recorded as Since it is between 0 and 1 and the sum is 1, the probability distribution condition is satisfied, so it can be written as: c is the output coordinate of a channel, consisting of three numbers. The above formula is the joint probability distribution of random variables X, Y, and Z. The three coordinate values ​​obtained after DSNT transformation are the expectation of the above joint distribution, which is recorded as: μ=E(c), The formula for finding expectation from probability theory is: <A,B> F Represents the F norm of matrix A and matrix B, expressed as: To sum up, we get the DNST expression:

2. The method for detecting vertebral center points in CT images based on deep learning according to claim 1, characterized in that: The ISE-Vnet neural network proposed in the coarse segmentation stage is a V-net network improved by the spatial attention mechanism and the channel attention mechanism. This network can well segment the spine. After the segmentation is completed, the original CT image is divided into multiple images of the same size by a sliding window according to the original CT image and the segmentation mask, which serve as the input of the second stage.

Citation Information

Patent Citations

  • Spine image segmentation method, medium and electronic equipment

    CN113034495A

  • Spinal map segmentation method of 2D convolutional neural network based on mixed attention mechanism

    CN113592794A