Detection of Spinal Vertebrae in Image Data

Through multi-stage detection method, the spinal vertebrae is automatically detected and marked in volume image data, solving the problems of large amounts of calculation and human error in the prior art, and achieving efficient and accurate vertebrae detection and labeling.

CN116457826BActive Publication Date: 2025-08-05KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180072894.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-27
Filing Date
2021-10-25
Publication Date
2025-08-05
Estimated Expiration
2041-10-25

AI Technical Summary

Technical Problem

The prior art is computationally large and prone to human errors when detecting and labeling spinal vertebrae in volumetric image data, and manual annotation by radiologists is time-consuming and tedious.

Method used

Using a multi-stage method, first detecting the vertebrae in the sagittal image and generating a 3-D bounding box, then generating a panoramic image and detecting the vertebrae in the panoramic image, and finally converting the bounding box to the 3-D space for labeling.

Benefits of technology

Reduces calculation requirements, improves detection accuracy and efficiency, reduces human errors, and simplifies the workflow of radiologists.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116457826B_ABST
    Figure CN116457826B_ABST
Patent Text Reader

Abstract

Vertebrae of a spine are detected in a volumetric image using a multi-stage detection algorithm with trained artificial intelligence. In one embodiment, a trained neural network (116) is used to detect individual vertebrae in a sagittal image in a first stage. Two-dimensional bounding boxes surrounding the detected vertebrae are combined to generate a three-dimensional model of the spine. A panoramic image of the spine is generated based on the three-dimensional model to create a straightened view of the spine. The trained neural network is used to detect individual vertebrae in the panoramic image in a second stage. The two-dimensional bounding boxes surrounding the detected vertebrae in the panoramic image are converted to three-dimensional space to create three-dimensional image data having three-dimensional bounding boxes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The following relates generally to medical imaging informatics and, more particularly, to the detection of spinal vertebrae in image data. Background Art

[0002] When a radiologist reads an axial spinal image that is not annotated with the identification of vertebrae, it is difficult for the radiologist to determine which vertebra is represented in the axial spinal image without reference to the sagittal spine and / or other images to cross-reference the position of the axial spinal image relative to the spine. In one instance, the radiologist is assigned to manually annotate and / or count the centers of the vertebrae. Unfortunately, this is tedious, time-consuming, and prone to human error. In "Vertebrae Localization and Segmentation with Spatial Configuration-Net and U-net" (2019) by Payer et al., "An Automatic Multi-stage System for Vertebra Segmentation and Labelling" (2019) by Chen, and "Iterative fully convolutional neural networks" (2019) by Lessmann et al., automatic vertebra detection algorithms are described. Unfortunately, these methods require considerable computing power. At least in view of the above, there is an unresolved need for another (multiple) methods for detecting and labeling spinal vertebrae in volumetric image data. Summary of the Invention

[0003] Some aspects described herein address the above-referenced problems and / or other problems.

[0004] The following describes a multi-stage method for automatic detection and labeling of spinal vertebrae in volumetric image data. In one embodiment, an object detection algorithm is used in a first detection stage to detect individual vertebrae in a sagittal image and generate 2-D bounding boxes for the detected vertebrae. The sagittal image and the 2-D bounding boxes are combined to generate a 3-D model of the spine with 3-D bounding boxes. A panoramic image of the spine is then generated based on the 3-D model to create a straightened view of the spine. In a second stage, an object detection algorithm (or another algorithm) is used to detect individual vertebrae in the panoramic image and generate 2-D bounding boxes for the detected vertebrae. The 2-D bounding boxes of the vertebrae in the panoramic image are converted to 3-D space.

[0005] In one aspect, a system is configured to detect vertebrae of a spine in volumetric image data. The system includes a computing device. The computing device includes a memory having instructions for a vertebra detection module. The computing device also includes a processor configured to execute the instructions to perform a two-stage vertebra detection, wherein a first set of bounding boxes for the vertebrae are detected in a sagittal image and clustered into boxes in a volumetric image in a first stage of the two-stage vertebra detection, a panoramic image of the spine is generated based on the detected first set of bounding boxes, and a second set of bounding boxes for the vertebrae are detected in the panoramic image in a second stage of the two-stage vertebra detection. The computing device also includes a display configured to display a 2-D image of the detected vertebrae. In one example, all detected bounding boxes—both the first and second sets—are labeled as sacrum, C2, or “other” vertebrae. This is at least because both the C2 and sacral vertebrae have unique shapes and serve as anchor points for vertebra annotation after bounding box detection.

[0006] In another aspect, a method is configured for detecting vertebrae of a spine in a volumetric image. The method includes extracting a first set of bounding boxes for the vertebrae in a sagittal image of the spine and, in one example, labeling them as sacrum, C2, or "other" vertebrae. The method then includes clustering the bounding boxes in the sagittal image into boxes in the volumetric image. The method also includes generating a panoramic image of the spine based on the detected first set of bounding boxes. The method also includes extracting a second set of bounding boxes for the vertebrae in the panoramic image using the same labeling as previously described.

[0007] In another aspect, a computer-readable storage medium stores instructions for detecting vertebrae of a spine in volumetric image data. These instructions, when executed by a processor of a computer, cause the processor to extract a first set of bounding boxes for the vertebrae in a sagittal image of the spine and, in one example, label them as sacrum, C2, or “other” vertebrae, cluster them into boxes in the volumetric data, generate a panoramic image of the spine based on the detected first set of volumetric bounding boxes, and extract a second set of bounding boxes for the vertebrae in the panoramic image.

[0008] Still further aspects of the present application will be appreciated to those skilled in the art upon reading and understanding the accompanying description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The invention may take form in various components and arrangements of components, and in various steps and arrangements of steps.The drawings are for purposes of illustrating embodiments only and are not to be construed as limiting the invention.

[0010] Figure 1 An example system having a vertebra detection and labeling module configured to detect and label vertebrae of a spinal column in volumetric image data according to embodiment(s) herein is schematically illustrated.

[0011] Figure 2 Schematically shows a device according to the embodiment(s) herein Figure 1 An example of the system's vertebrae detection and labeling module.

[0012] Figure 3 An example method for detecting and marking vertebrae of a spine in volumetric image data according to embodiment(s) herein is schematically illustrated. DETAILED DESCRIPTION

[0013] Figure 1 Schematically illustrated is an example system 102 according to embodiment(s) herein. System 102 includes at least data repository(s) 104 and computing device 106.

[0014] The data repository(ies) 104 include a physical storage medium configured to store at least digital medical images. In one example, the data repository(ies) 104 is for a healthcare entity(ies), etc., and includes digital medical images of a subject acquired by an imaging modality of the healthcare entity(ies). The physical storage medium is local to the healthcare entity(ies) and / or remote therefrom, such as part of a "cloud" based resource. Examples of imaging modalities include magnetic resonance (MR), computed tomography (CT), single photon emission tomography (SPECT), positron emission tomography (PET), X-ray, etc. The digital medical images include a series of two-dimensional (2-D) images (which collectively provide a three-dimensional (3-D) volumetric image dataset) and / or a 3-D volumetric image dataset. The digital medical images include at least images of vertebrae of the subject's spine.

[0015] The computing device 106 includes a processor 108 (e.g., a central processing unit (CPU), a microprocessor (CPU), a graphics processing unit (GPU), and / or other processors) and a computer-readable storage medium ("memory") 110 (which does not include transient media), such as a physical storage device, such as a hard drive, a solid-state drive, and / or an optical disk. The memory 110 includes at least computer-executable instructions 112 and data 114. The processor 108 is configured to execute the computer-executable instructions 112. In one example, the computing system 106 is configured to provide storage, access, and / or processing of medical information, including digital medical images, electronic reports, and the like. Examples of the computing device 106 include, but are not limited to, a picture archiving and communication system (PACS). When the computing system 106 is a PACS, the digital medical images and / or other electronic information are stored and / or transmitted via the DICOM (Digital Imaging and Communications in Medicine) format and / or other format(s).

[0016] Instructions 112 include at least instructions for a vertebra detection module 116. As described in more detail below, in one embodiment, vertebra detection module 116 is configured to detect individual vertebrae in a spinal column image acquired in a sagittal plane ("sagittal image") and generate 2-D bounding boxes for the detected vertebrae, combine the sagittal image and the 2-D bounding boxes to generate a 3-D model of the spine having 3-D bounding boxes, generate a panoramic image of the detected vertebrae based on the 3-D model to create a straightened view of the spine, detect individual vertebrae in the panoramic image and generate 2-D bounding boxes for the detected vertebrae in the panoramic image, convert the 2-D bounding boxes to 3-D space, and optionally annotate the displayed 2-D image. As used herein, a bounding box delimits or encloses a vertebra in a manner that partially overlaps with or does not partially overlap with one or more adjacent vertebrae. In one example, vertebra detection module 116 reduces computational power used for detection and labeling and / or improves the accuracy of delimiting vertebrae relative to a configuration without vertebra detection module 116.

[0017] Input device(s) 118, such as a keyboard, mouse, touch screen, etc., are in electrical communication with computing system 102. In one example, input device(s) 118 are configured to allow a user to operate computing system 102 via user input, including activating vertebra detection module 116, selecting volumetric image data and / or sagittal images to load, and the like. Human-readable output device(s) 120, such as a display, are also in electrical communication with computing device 106. In one example, output device(s) 120 are configured to display 2-D images of vertebrae, prompt for user input, present instructions, and the like. Input / output ("I / O") 122 is configured to communicate (wired and / or wirelessly) with at least data repository(s) 104, including retrieving / receiving electronic data from and / or transmitting data to data repository(s) 104, input device(s) 118, and / or output device(s) 120.

[0018] Figure 2 Schematically, an example of a vertebra detection module 116 according to an embodiment of the present invention is shown. The vertebra detection module 116 shown includes a data pre-processor 200, a trained vertebra detector 202, a 3-D model generator 204, a panoramic image generator 206, a 2-D to 3-D space converter 208, and an annotator 210.

[0019] Vertebral detection module 116 receives as input image data from a spinal scan of the subject, the image data comprising a series of 2-D images (which collectively provide a 3-D volumetric dataset) and / or a 3-D volumetric dataset, or one or more sets of sagittal images generated from a series of 2-D images (which collectively provide a 3-D volumetric dataset) and / or a 3-D volumetric dataset. As briefly discussed herein, suitable datasets include MR, CT, SPECT, PET, X-ray, and the like. In one embodiment, specific image data is selected via user input from input device(s) 118.

[0020] The data preprocessor 200 is configured to process input image data. In one instance, the result of the processing is one or more sets of sagittal slices. Parameters such as window width, window level, slice thickness, etc. are determined via user input and / or preprogrammed settings. In the case where multiple sets of sagittal slices are generated, in one instance, the window width, window level and / or other parameters are the same for all slices. In another instance, at least two groups have at least one different parameter value. The window width refers to the range of CT numbers (in Hounsfield units (HU)) to be displayed, and the window level refers to the CT number at the midpoint of the range. An example of a window width and window level (W / L) setting for viewing the spine in an image that includes the spine is: W=1800HU, L=400HU. In the case where the input image data includes one or more sets of sagittal images, the data preprocessor 200 is not used to process the input image data to create one or more sets of sagittal images.

[0021] The trained vertebra detector 202 is configured to process one or more sets of sagittal slices. In one example, this includes detecting whether a slice includes a vertebra and generating a 2-D bounding box for the detected vertebra. In one example, vertebra detection begins at the center slice and proceeds slice by slice in both directions to the outermost slices, i.e., the first and last sagittal slices. Vertebral detection ends after all slices have been processed, after a predetermined stopping criterion is met (e.g., after a predetermined number of consecutive slices have been processed without a vertebra being marked, indicating that the spine is outside the data set), or after other predetermined stopping criteria. In one example, the predetermined stopping criterion reduces processing time, for example, by limiting the number of slices processed.

[0022] Examples of suitable detectors include artificial intelligence-based detectors, including neural network-based detectors, such as the Faster Region Convolutional Neural Network described in "Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks" (2015) by Girshick et al., YouOnly Look Once described in "YOLOv3: An Incremental Improvement" (2018) by Redmon et al., and / or other detector(s). For the purposes of brevity and explanation, vertebra detection is described here using the YOLOv3 detector. With the YOLOv3 detector, a neural network is applied to an image, the image is divided into regions, and a bounding box and probability are predicted for each region, where the bounding box is weighted by the predicted probability.

[0023] With the YOLOv3 detector, a predetermined minimum confidence threshold is used to determine whether to generate a 2-D bounding box for an object that is likely detected as a vertebra. In one example, the minimum confidence threshold is 0.55. This means that a 2-D bounding box will only be generated for vertebrae that are detected with a confidence level of 0.55 or higher. In another example, the minimum confidence threshold is 0.50. In yet another example, the minimum confidence threshold is a different value. A minimum confidence threshold of 0 will result in a 2-D bounding box being generated for every object that is likely detected as a vertebra. In general, higher thresholds increase specificity, while lower thresholds increase sensitivity.

[0024] The 3-D model generator 204 is configured to process one or more sets of sagittal images and 2-D bounding boxes. In one instance, this includes combining the 2-D bounding boxes across the sagittal slices for each detected vertebra to generate a 3-D model with a 3-D bounding box. For the purposes of brevity and explanation, the density-based spatial clustering of applied noise (DBSCAN) algorithm is described here, which is described in Ester et al., "Adensity-based algorithm for discovering clusters in large spatial databases with noise. Proceeding of the Second International Conference on Knowledge Discovery and Data Mining (KDD-96)", AAAI Press, pages 226-331. DBSCAN is a density-based non-parametric algorithm for data clustering, in which given a set of points in some space, points that are close to each other are clustered / grouped together, while points whose closest neighbors are far away are considered outliers.

[0025] The panoramic image generator 206 is configured to process the 3-D model and the 3-D bounding box. In one example, this includes generating a panoramic image using a curve that passes through the center of the 3D bounding box. In one example, the curve is extrapolated before the first vertebra and after the last vertebra. This allows the panoramic image generator 206 to add the vertebra(s) lost at the edge. The curve is interpolated and sampled at predetermined intervals (e.g., 0.1, 0.5, 1.0, 2.5, etc. millimeters (mm)). For each point on the curve, a line from the 3-D model is sampled along the projection of a vector from one side of the body to the opposite side (e.g., from the front of the body to the back of the body) at that point on a plane perpendicular to the curve. The result is a panoramic (quasi-sagittal) image of the entire spine that contains vertical alignment.

[0026] The trained vertebra detector 202 is also configured to process panoramic images. In one example, this includes detecting vertebrae in the panoramic image and generating a 2-D bounding box around each detected vertebra. In one example, the minimum confidence threshold is 0.50. Similar to detection in sagittal slices, the minimum confidence threshold can be a different value. Generally, vertebra detection in panoramic images should be more accurate than detection using sagittal slices, at least because the spine is straightened and fully displayed, and any vertebrae that were missed using the sagittal image are now detected here. In one example, this improves sensitivity without reducing specificity. In another embodiment, separate vertebra detectors are used to detect vertebrae in both the sagittal and panoramic images.

[0027] The 2-D to 3-D space converter 208 is configured to process panoramic images having 2-D bounding boxes. In one example, this includes converting the 2-D bounding boxes to 3-D space and adding depth to each bounding box. For example, in one example, each corner of each bounding box is converted to 3-D space, and the bounding boxes of all four corners are set to the bounding box of the vertebra, where the depth is set to the same as the smaller side of the 2-D bounding box (i.e., the smaller of the width or height of the bounding box).

[0028] The trained vertebra detector 202 is also configured to determine the identity of at least a subset of the identified vertebrae. For example, in one example, the trained vertebra detector 202 identifies at least one cervical vertebra, a plurality of sacral vertebrae, and other spinal vertebrae, such as C2, S1, S2, S3, S4, S5, and another symbol or word for all other vertebrae. Vertebrae other than C2 and S1-S5 (i.e., C3-C7, T1-T12, L1-L5, and / or coccyx vertebrae) can be identified by counting backward or forward from a reference identified vertebra. In another example, different combinations of vertebrae subsets (e.g., C2, sacrum, and other vertebrae) are identified, only one vertebra is identified, or all vertebrae are identified.

[0029] The annotator 210 is configured to annotate the displayed 2-D image of the input volumetric image data. In one example, the annotation is a projection of the center of a bounding box onto the current slice. As a non-limiting example, in an off-center axial slice, the annotator 210 projects the center from the central axial slice. Alternatively or additionally, the annotator marks the intersection of the centerline through the vertebra with the current plane. In another example, the detected vertebrae are annotated by displaying a vertebra label next to each vertebra without any bounding box or center projection.

[0030] Next consider the variants.

[0031] In one variation, the entire algorithm is run iteratively, processing sagittal slices, finding bounding boxes, generating a 3-D model, generating a panoramic image, determining bounding boxes for the panoramic image, converting the bounding boxes to 3-D space, and repeating the process until a stopping criterion is met.

[0032] In the above, the trained vertebra detector 202 detects vertebrae in sagittal and panoramic images. In one variation, separate trained vertebra detectors detect vertebrae in sagittal and panoramic images, respectively.

[0033] A non-limiting example of training a vertebra detector to create a trained vertebra detector 202 for a single imaging modality is described below. For the purposes of brevity and explanation, CT is used to describe the training. Sagittal images of the lumbar, thoracic, and cervical spine used in a CT study are annotated. All sagittal images are sampled at the same predetermined resolution with the same predetermined fixed spacing. For example, in one example, the sagittal images are sampled at a resolution of 416×416 with a pixel spacing of 1 mm. For larger images, the image is divided into multiple regions and each region is sampled to cover the entire image.

[0034] Sagittal images with 2-D bounding boxes are fed into the training. These sagittal images with 2-D bounding boxes are also augmented to generate additional training data. Examples of augmented features include one or more of: brightness, contrast, Gaussian noise, shift, scale, rotation, and flip. For each feature, an augmentation probability and an augmentation limit are predetermined. For example, a brightness probability of 0.80 and an augmentation limit of 0.03 will result in brightness being augmented within ±3% of the original value in 80% of the images. Augmenting sagittal images in this way increases data diversity and / or mitigates overfitting of the training data.

[0035] The training dataset is divided into a training subset, a testing subset, and a validation subset of predetermined sizes. The CT training dataset is used to train the trained vertebra detector 202 to detect vertebrae in CT image data. Using a YOLOv3-based network, the vertebra detector is iteratively trained until a stopping criterion is met. In one example, the vertebra detector is iteratively trained until the error between the original bounding box and the generated bounding box meets a predetermined value. The vertebra detector is trained to detect all or a subset of vertebrae.

[0036] In one example, a validation dataset is used during training to determine stopping criteria. In this example, the network weights are not affected during the validation phase, and the quality of training is determined by examining metrics. For example, once training is complete, the network is run on the test dataset, and the metrics are compared to those on the validation dataset. If the metrics deviate within a predetermined tolerance, training is considered effective and terminated.

[0037] Next consider variations on training.

[0038] In a variant, the training images also include panoramic images based on the annotated vertebrae centers and with the same resolution and pixel spacing. Feature enhancement can be different between sagittal and panoramic images, for example, panoramic images do not have shift and rotation enhancement.

[0039] In another variation, a vertebra detector is trained for multiple imaging modalities. For simplicity and explanatory purposes, training is described using CT and MR image data. With this variation, the enhancement can vary across modalities; for example, the contrast limit for CT can be 0.30 and the contrast limit for MR can be 0.40. In one variation, the CT and MR datasets are used to train separate vertebra detectors to detect vertebrae in CT and MR images, respectively. In another variation, the CT and MR datasets are used to train a CT+MR vertebra detector for vertebra detection in either CT or MR image data.

[0040] Figure 3 An example method for detecting and marking vertebrae of a spine in volumetric image data according to embodiment(s) herein is schematically illustrated.

[0041] Should be understood that the action order of one or more methods is not restrictive. Therefore, other sortings are considered herein. In addition, one or more actions can be omitted, and / or one or more additional actions can be included.

[0042] An image loading step 302—as described herein and / or otherwise—loads volumetric image data.

[0043] A 2-D image generation step 304—as described herein and / or otherwise—generates one or more sets of sagittal images from the volumetric image data.

[0044] Alternatively, the loading step—as described herein and / or otherwise—loads one or more sets of sagittal images, and step 304 is omitted.

[0045] Vertebra detection step 306—as described herein and / or otherwise—detects vertebrae in one or more sets of sagittal images and generates 2-D bounding boxes therefor.

[0046] A 3-D model generation step 308—as described herein and / or otherwise—generates a 3-D model of the spine having a 3-D bounding box from the sagittal image and the 2-D bounding box.

[0047] A panoramic image generation step 310—as described herein and / or otherwise—generates a panoramic image based on the 3-D model.

[0048] A vertebra detection step 312—as described herein and / or otherwise—detects vertebrae in the panoramic image and generates 2-D bounding boxes therefor.

[0049] A 2-D to 3-D conversion step 314—as described herein and / or otherwise—converts the 2-D bounding boxes for the vertebrae in the panoramic image to 3-D space.

[0050] An annotation step 316—as described herein and / or otherwise—annotates the displayed 2-D image of the volumetric image data.

[0051] The above methods may be implemented by computer-readable instructions encoded or embedded on a computer-readable storage medium 110, which, when executed by a computer processor(s), cause the processor(s) 108 to perform the described actions. Additionally or alternatively, at least one of the computer-readable instructions is implemented by a signal, carrier wave, or other transient medium, which is not a computer-readable storage medium.

[0052] Although the present invention has been described and illustrated in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary rather than restrictive; the invention is not limited to the disclosed embodiments. Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims.

[0053] The word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single processor or other unit may fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

[0054] The computer program may be stored / distributed on a suitable medium such as an optical storage medium or solid-state medium provided together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems. Any reference signs in the claims should not be construed as limiting the scope.

Claims

1. A system (102) configured to detect vertebrae of a spine in volumetric image data, the system comprising: A computing device (106) comprising: a memory (110) comprising instructions (112) for a vertebra detection module (116); a processor (108) configured to execute the instructions to perform two-stage vertebra detection, wherein: a first set of bounding boxes of the vertebrae are detected in a sagittal image and clustered in the volumetric image data in a first stage of the two-stage vertebra detection, a panoramic image of the spine is generated based on the detected first set of bounding boxes, and a second set of bounding boxes of the vertebrae are detected in the panoramic image in a second stage of the two-stage vertebra detection; and a display (120) configured to display a 2-D image based on the volumetric image data of the detected vertebrae; wherein the vertebra detection module detects the first set of bounding boxes based on a first predetermined confidence level and generates 2-D bounding boxes for the detected vertebrae; wherein the vertebra detection module combines the sagittal image and the 2-D bounding box to generate a 3-D model having a 3-D bounding box; wherein the vertebra detection module generates a curve through the center of the 3-D bounding box; wherein for each point on the curve, the vertebra detection module samples a line from the 3-D model along the projection of a vector from the front of the spine to the back of the spine at that point on a plane perpendicular to the curve to produce the panoramic image, the panoramic image comprising a quasi-sagittal image containing the entire spine in vertical alignment; wherein the vertebra detection module detects the second set of bounding boxes based on a second predetermined confidence level and generates 2-D bounding boxes for the detected vertebrae; and The vertebra detection module converts the 2-D bounding box for the panoramic image into a 3-D space to define a 3-D bounding box for the vertebra.

2. The system of claim 1, wherein the vertebra detection module comprises a neural network trained to detect vertebrae.

3. The system of any one of claims 1 to 2, wherein the vertebra detection module detects the first set of bounding boxes starting with a center image of the sagittal image and moving outward in both directions toward a first image of the sagittal image and a last image of the sagittal image until a stopping criterion is met. 4 . The system of claim 3 , wherein the stopping criterion comprises a predetermined number of consecutive images in the sagittal image in which no vertebrae are detected.

5. The system of any one of claims 1 to 4, wherein the vertebra detection module labels each vertebra of the first set of bounding boxes as a sacrum, C2, or other vertebra.

6. The system of any one of claims 1-5, wherein the vertebra detection module extrapolates the curve before the first vertebra and after the last vertebra to add missing vertebrae.

7. The system of any one of claims 1 to 6, wherein the vertebra detection module labels each vertebra of the second set of bounding boxes as a sacrum, C2, or other vertebra.

8. The system according to any one of claims 1 to 7, wherein the computing device is an image storage and transmission system.

9. A system (102) configured to detect vertebrae of a spine in volumetric image data, the system comprising: A computing device (106) comprising: a memory (110) comprising instructions (112) for a vertebra detection module (116); a processor (108) configured to execute the instructions to perform two-stage vertebra detection, wherein: a first set of bounding boxes of the vertebrae are detected in a sagittal image and clustered in the volumetric image data in a first stage of the two-stage vertebra detection, a panoramic image of the spine is generated based on the detected first set of bounding boxes, and a second set of bounding boxes of the vertebrae are detected in the panoramic image in a second stage of the two-stage vertebra detection; and a display (120) configured to display a 2-D image based on the volumetric image data of the detected vertebrae; wherein the vertebra detection module detects the first set of bounding boxes based on a first predetermined confidence level and generates 2-D bounding boxes for the detected vertebrae; wherein the vertebra detection module combines the sagittal image and the 2-D bounding box to generate a 3-D model having a 3-D bounding box; wherein the vertebra detection module generates a curve passing through the center of the 3-D bounding box; and The vertebra detection module extrapolates the curve before the first vertebra and after the last vertebra to add the missing vertebrae.

10. A computer-implemented method for detecting vertebrae of a spine in volumetric image data, comprising: extracting a first set of bounding boxes of vertebrae in a sagittal image of the spine; generating a panoramic image of the spine based on the detected first set of bounding boxes; as well as extracting a second set of bounding boxes for the vertebrae in the panoramic image; as well as Also includes: detecting the first set of bounding boxes and generating 2-D bounding boxes for the detected vertebrae based on a first predetermined confidence level; detecting a second set of bounding boxes in the panoramic image based on a second predetermined confidence level and generating a second set of 2-D bounding boxes; converting the 2-D bounding box in the panoramic image to 3-D space to define a 3-D bounding box for the vertebra; and The vertebrae are annotated with the 3-D bounding boxes.

11. The computer-implemented method of claim 10 , wherein extracting the first set of bounding boxes comprises: detecting the first set of bounding boxes starting with a center image of the sagittal image and moving outward to a first image of the sagittal image and a last image of the sagittal image; and terminating the detection in response to a predetermined number of consecutive sagittal images having no vertebrae; generating a 2-D bounding box for the detected vertebra; identifying the center of the 2-D bounding box; and Annotate the vertebrae with the 2-D bounding box.

12. A computer-readable storage medium storing computer-executable instructions for detecting vertebrae of a spinal column in volumetric image data, the computer-executable instructions, when executed by a computer processor, causing the processor to: extracting a first set of bounding boxes of vertebrae in a sagittal image of the spine; generating a panoramic image of the spine based on the detected first set of bounding boxes; and extracting a second set of bounding boxes for the vertebrae in the panoramic image; and The computer-executable instructions further cause the processor to: detecting the first set of bounding boxes and generating 2-D bounding boxes for the detected vertebrae based on a first predetermined confidence level; detecting a second set of bounding boxes in the panoramic image based on a second predetermined confidence level and generating a second set of 2-D bounding boxes; as well as converting the 2-D bounding box in the panoramic image to 3-D space to define a 3-D bounding box for the vertebra; as well as The vertebrae are annotated with the 3-D bounding boxes.

13. The computer-readable storage medium of claim 12, wherein the computer-executable instructions further cause the processor to: detecting the first set of bounding boxes starting with a center image of the sagittal image and moving outward to a first image of the sagittal image and a last image of the sagittal image; terminating the detection in response to a predetermined number of consecutive sagittal images having no vertebrae; generating a 2-D bounding box for the detected vertebra; identifying the center of the 2-D bounding box; and Annotate the vertebrae with the 2-D bounding box.