A system for automated and quantitative assessment of metastatic spinal stability
A deep learning system for analyzing 3D medical images enhances the prediction of vertebral compression fractures and spinal stability by automating the assessment of metastatic spine lesions, improving clinical outcomes through accurate quantification and prediction of fracture risk.
Patent Information
- Application Number
- JP2025526244
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-08
- Filing Date
- 2023-11-08
- Publication Date
- 2025-11-06
AI Technical Summary
Current methods for assessing vertebral compression fracture (VCF) risk and spinal mechanical stability in patients with metastatic spine lesions are inadequate, relying on manual image review and the Spinal Instability Tumor Score (SINS) with limited predictive accuracy, and lack computational medical image analysis.
A deep learning-based system that analyzes 3D medical images (CT and MR) to quantify spinal geometry and disease burden, using a machine learning algorithm to calculate mechanical stability and fracture risk by combining deep features with vertebra-specific and patient-specific features, providing an automated Spinal Instability Tumor Scoring (SINS) tool and direct fracture risk estimation.
Improves the accuracy and automation of spinal stability prediction, enhancing clinical decision-making for patients undergoing stereotactic body radiation therapy (SBRT) by quantifying biomarkers from clinical images, reducing the risk of mechanical instability and VCFs.
Smart Images

Figure 2025536474000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates to a system that uses deep learning to automatically quantify spinal structure (geometry / quality) and disease burden from 3D medical images (non-limiting examples are CT and MR) that can be used to predict patient outcomes for spinal stability / fracture risk. More specifically, this disclosure provides an interpretable automated spinal instability tumor scoring (SINS) tool that calculates improved individual image-based parameters used in SINS and calculates stability through a similar scoring system; it also provides an image-based tool that directly estimates vertebral stability / fracture risk. [Background technology]
[0002] Approximately 70% of cancer patients are found to have skeletal metastases during postmortem examination, with the spine being the most common site [1]. These metastases cause pain, mobility problems, mechanical instability, and neurological impairment with some lesions resulting in vertebral compression fractures (VCFs). With advances in image-guided external beam radiation therapy and robotics, stereotactic body radiation therapy (SBRT) has emerged as a highly effective method for treating these tumors locally [1]. However, VCFs are a common complication after SBRT, with postprocedure rates ranging from 10% to 40% [2].
[0003] VCF risk and mechanical stability assessment are currently assessed by manual image review and combining this information with patient factors. The Spinal Instability Tumor Score (SINS) is a standardized method used to assess mechanical instability and has been widely adopted, but its effectiveness for predicting VCF is limited (hazard ratio = 0-5.6) [3]. Existing scientific literature does not address computational medical image analysis, with most studies assessing fracture risk and vertebral mechanical stability using patient-specific factors (e.g., age, primary tumor location, BMI) and / or SINS assessment with limited fidelity. This invention describes an AI-enabled medical image analysis-based tool for the assessment of VCF risk and spinal mechanical stability. Summary of the Invention
[0004] Disclosed herein is a method for assessing mechanical stability and / or fracture risk in a spine with metastasis, the method comprising: a) obtaining 3D imaging data of a patient's spine; b) inputting the image data to a machine learning algorithm trained on a dataset of spine images to determine mechanical stability and / or fracture risk, the machine learning algorithm configured to: i) calculating mechanical stability and / or fracture risk by computing deep features from a feature extractor backbone network; ii) combine deep features with vertebra-specific features derived from convolutional layers, graph networks, and convolutional arms; and / or extract latent deep features for each vertebra; iii) The deep layer features combined with the vertebra-specific features and non-image patient-specific features in step ii) are combined with the high density layer to obtain predictions regarding mechanical stability and / or fracture risk, and training is performed by an optimizer based on 3D image data of the patient's spine.
[0005] The 3D image data of the patient's spine may be CT and / or MR image data.
[0006] The spine image dataset may be composed of CT and / or MR images comprising a clinical cohort of patients with spinal metastases.
[0007] The feature extractor backbone network may be a ResNet network, or a Transformer network, or an Inception network, or one that uses fully connected layers, and uses principal component analysis or radiomics, or linear discriminant analysis, independent component analysis, t-distributed stochastic neighbor embedding, to name a few non-limiting examples.
[0008] The feature extractor backbone network may be a ResNet50 network.
[0009] The present disclosure provides a method for assessing mechanical stability and / or fracture risk in a spine with metastasis, the method comprising: a) obtaining 3D imaging data of a patient's spine; b) inputting the image data into a machine learning algorithm trained on a dataset of spine images to determine mechanical stability and / or fracture risk, the machine learning algorithm being configured to calculate spinal instability oncological score components, the components including bony lesions, spinal alignment, vertebral body collapse, and posterolateral lesions; c) calculating each of the elements of bone lesion, spinal alignment, vertebral body collapse, and posterolateral lesion includes calculating deep features from a feature extractor backbone network, and combining the deep features with vertebra-specific features derived from the convolutional layer and the convolutional arms, thereby segmenting the vertebrae, and combining the deep features with the vertebra-specific features and the non-image patient-specific features with the dense layer to obtain predictions regarding mechanical stability and / or fracture risk, wherein training is performed by an optimizer based on 3D image data of the patient's spine in combination with information regarding patient outcomes; To calculate bone lesion components, the segmented and localized vertebrae are used in a histogram-based analysis to determine the volume and type of tumor tissue present within each metastatic vertebra; Calculate Cobb angles in the coronal and sagittal planes using segmented and localized vertebrae to calculate spinal alignment elements; To calculate vertebral body collapse, the segmented and localized vertebrae are combined with another convolutional network that segments and calculates the volume of the vertebral body in the 3D image, and the volume of the complete vertebral body at the adjacent levels is used to estimate the vertebral body collapse rate; The posterolateral lesions are determined by combining the segmented and localized vertebrae, computing deep features from a feature extraction backbone network, combining the deep features with vertebra-specific features obtained from a graph network, a convolutional layer, and a convolutional arm, and combining the deep features with the vertebra-specific features and a dense layer to determine the posterolateral lesions, and training is performed by an optimizer based on 3D image data of the patient's spine combined with the labeling of the posterolateral lesions.
[0010] The spine image dataset may be composed of CT and / or MR images comprising a clinical cohort of patients with spinal metastases.
[0011] The components of bony lesions, spinal alignment, vertebral collapse, and posterolateral lesions can be combined with clinical pain and vertebral level to derive an automated spinal instability oncologic score.
[0012] The components of bony pathology, spinal alignment, vertebral collapse, and posterolateral pathology are combined with clinical pain and other patient-specific characteristics and vertebral level to obtain an automated SINS score that predicts mechanical stability and / or fracture risk.
[0013] A further understanding of the functional and advantageous aspects of the present disclosure can be realized by reference to the following detailed description and drawings. [Brief explanation of the drawings]
[0014] Embodiments of an AI-enabled medical image analysis-based tool for assessment of VCF risk and spinal mechanical stability are described below, by way of example only, and with reference to the drawings: [Figure 1] 1 shows a schematic diagram of a system for automatic and quantitative assessment of metastatic spinal stability as disclosed herein. [Figure 2] FIG. 2 is a close-up view of the non-limiting VertDetect model 14 of FIG. 1. DETAILED DESCRIPTION OF THE INVENTION
[0015] Various embodiments and aspects of the present disclosure are described with reference to the details discussed below. The following description and drawings are illustrative of the present disclosure and should not be construed as limiting the disclosure. The drawings are not necessarily to scale. Numerous specific details are set forth in order to provide a thorough understanding of various embodiments of the present disclosure. However, in some cases, well-known or conventional details are not described in order to provide a concise discussion of embodiments of the present disclosure.
[0016] As used herein, the terms "comprises" and "comprising" should be interpreted as inclusive and not limited or exclusive. Specifically, when used in this specification, including the claims, the terms "comprises" and "comprising" and variations thereof mean that the specified features, steps, or components are included. These terms should not be interpreted to exclude the presence of other features, steps, or components.
[0017] As used herein, the word "exemplary" means "serving as an example, instance, or illustration," and should not be construed as preferred or advantageous over other configurations disclosed herein.
[0018] As used herein, the terms "about" and "approximately," when used in conjunction with a particle size range, a mixture composition, or other physical attribute or characteristic, are meant to cover slight variations that may exist at the upper and lower limits of the size range, such that, on average, most of the sizes are met, but not to exclude embodiments where statistical sizes may fall outside this range. Unless otherwise specified, the terms "about" and "approximately" mean ±25 percent or less. It is not intended to exclude such embodiments from the present disclosure.
[0019] Unless otherwise stated, any particular range or group should be understood to be a shorthand way of referring to each and every member of the range or group individually, and each and every possible subrange or subgroup encompassed therein, as well as any subranges or subgroups therein. Unless otherwise specified, the present disclosure relates to and expressly incorporates each and every specific member and combination of subranges or subgroups.
[0020] As used herein, the term "on the order of" when used in conjunction with an amount or parameter refers to a range ranging from about one-tenth to ten times the stated amount or parameter.
[0021] definition As used herein, the phrase "machine learning algorithm" refers to an algorithm that "learns" how to perform tasks by being "trained" on a dataset to perform those tasks, optimizing performance by modifying internal factors that affect the algorithm's output. The internal factors are modified so that the output produced matches what is seen in the dataset. Here, machine learning algorithms that learn how to represent imaging data using feature extraction units and how to use these latent representations, in combination with patient-specific factors, are used to make predictions that are consistent with the training dataset.
[0022] As used herein, the phrase "deep features" refers to quantitative values extracted from an image that are the result of a deep learning network architecture. A network architecture with more than one layer is called a deep learning architecture. A deep learning network, as used herein, has multiple layers. A network, as used herein, learns how to represent an image. A representation of an image, including its features, is called a feature representation. In this disclosure, a deep learning network is used to learn feature representations. Because deep learning learns representations, the internal features are called deep features.
[0023] As used herein, the phrase "feature extractor backbone network" refers to a network or portion of a network in which latent features are used or shared across different subsequent tasks. It can refer to a network that, after training, generates a feature representation of the entire image volume input that is useful for the required task. The features extracted by this network are used as a backbone for predictions and decisions for the entire network.
[0024] As used herein, the phrase "convolutional layer" refers to a structure that performs a series of filtering operations on input data to generate an output. A convolutional layer performs discrete convolution operations on several inputs, where the weights are kernels to be learned. A convolutional layer can also have biases, which are learned weights. Many convolutional layers can create filters that derive an output from the entire image or from only a small aspect of the image.
[0025] As used herein, the phrase "convolutional arms" refers to a series of convolutional layers, pooling layers, and fully connected layers. Pooling layers reduce the dimensionality or number of voxel / pixel outputs by combining information from many input voxels / pixels to produce an output with fewer voxels / pixels. Fully connected layers have connections and weights from each input node (pixel / voxel, previous layer output) to all nodes in the fully connected layer. The use of arms indicates that there are several arms that perform calculations in series and are specially trained to perform a specific task.
[0026] As used herein, the phrase "vertebra-specific features" refers to features or aspects that characterize individual vertebrae in imaging data. This is in contrast to features or aspects that describe the entire image, the entire spine, or the patient. The features may be human-understandable, such as vertebra size, vertebral density, or the result of machine learning calculations, such as those by convolutional arms that produce features that describe specific vertebrae but are not intuitively understandable.
[0027] As used herein, the phrase "non-imaging patient-specific characteristics" refers to patient demographic data including age, sex, body mass index, location and tissue of the primary tumor, previous treatment (e.g., systemic medical therapy and / or local treatment (such as radiation received and dose, to give some non-limiting examples), number of vertebral metastases, presence of other metastases in the body, existing fractures, presence of pain).
[0028] As used herein, the phrase "training is performed by an optimizer" refers to adjusting weights within a machine learning algorithm during a phase called training, where input data is labeled to specify its output state. The weights within the machine learning algorithm are adjusted to match the desired output. The optimizer iteratively modifies the weights by changing the weights and then evaluates the output against the desired output from the training dataset. The optimizer continues to change the weights until the output from the machine learning algorithm matches the output in the training dataset.
[0029] This disclosure describes an AI-based tool for the automated and quantitative assessment of metastatic spinal stability. This paper uses deep learning to automatically quantify spinal geometry / quality and disease burden from medical images (CT and MR) that can be used to predict patient outcomes for vertebral stability and fracture risk. Two tools have been developed: 1) an interpretable automated SINS (Spine Instability Tumor Scoring) tool that calculates improved individual image-based parameters used in SINS and calculates stability via a similar scoring system; and 2) a tool that directly estimates spinal stability / fracture risk based on 3D imaging (with or without non-imaging patient-specific clinical data). The automated tool can be used to improve clinical workflow (through automation) and the accuracy, sensitivity, and specificity of spinal stability prediction, which can aid clinicians in better directing treatment to optimize patient outcomes. This tool is particularly useful in patients undergoing planned stereotactic body radiation therapy (SBRT) due to the accessibility of clinical images and the high likelihood (14%) of vertebral compression fractures (VCFs) and associated mechanical instability following this procedure.
[0030] Therefore, we created a dataset of imaging and associated clinical data from patients who underwent SBRT for spinal metastases at Sunnybrook Odette Cancer Centre and trained our algorithm on this retrospective data. We also used an additional clinical research dataset consisting of temporal CT imaging data from patients with spinal metastases, as well as an open dataset (VerSe
[12] ). This technique and the trained algorithm will be useful for clinicians planning treatment for patients with spinal metastases and for companies creating technologies for planning SBRT procedures. This algorithm is novel in that it quantifies biomarkers from clinical images in an automated, quantitative manner using an innovative multitasking architecture. These algorithms replace currently used clinical manual qualitative scoring, which is inadequate for predicting fracture risk, with improved automated predictions.
[0031] Non-limiting exemplary embodiments of the system for automated quantitative assessment of metastatic spinal stability disclosed herein will now be described.
[0032] Reference numbers for Figures 1 and 2 10: Overall flow chart 12, 50: Myelography 14: The VertDetect model includes three main branches: detection, classification, and segmentation. The detection branch detects each vertebra in the 3D CT scan by both determining the vertebral centroid position and placing a bounding box around the entire vertebra. The classification branch utilizes the shared information between each adjacent vertebra to determine which vertebrae are present in the input CT scan. The segmentation branch semantically segments the vertebrae that are positively detected from the classification and detection branches: 16: Vertebral segmentation and localization 18:Vertebral levels to be assessed 20: Bone Lesion Element: Histogram-Based Calculation of Tumor Type and Volume 22: Spinal Alignment Component: Cobb Angle Calculated in Coronal and Sagittal Planes 24: Vertebral crushing element 26: Posterolateral lesion element 28: Deep learning features generated directly by machine learning networks 30: Patient-specific non-imaging factors (i.e., age, sex, BMI, current medical therapy, systemic and local therapy, etc.) 36: AutoSINS score 38: Fracture Risk / Spine Stability Score 52: A backbone architecture consisting of a 3D ResNet-50 and a Feature Pyramid Network (FPN). P1 to P5 are feature maps generated from the convolutional backbone. The feature maps generated from this backbone are then used in further downstream tasks.
[0033] 54, 56, 60-65: 3x3x3 convolution + batch normalization + ReLU 58: 1x1x1 convolution + batch normalization + ReLU 67~69, 128: 1x1x1 convolution 70: Offset calculation 72: Bounding box size 74: Gaussian heatmap 76, 114: Concatenation 80: Bounding box 82: Initial bounding box processing 84: Binary classification block 86: Binary vertebra classification 88: Crop region of interest for each individual vertebra 100,102, 108,110,120,122,124,126: 3x3x3 convolution + instance normalization + ReLU 104, 122: 1x1x1 convolution + instance normalization + ReLU 106: 2x2x2 convolution transpose + instance normalization 130: Segmentation of individual vertebrae in an image 130: Gaussian heatmap used to identify the location of each vertebra.
[0034] A fully 3D end-to-end vertebra instance segmentation model VertDetect is a novel architecture inspired by previous work by Yi et al. [7], Mask R-CNN [3], and FCOS [5][6], modifying layers for 3D use. VertDetect takes a 3D image 50 containing a depiction of a vertebra as input (Figure 2). It refines the first layer of the ResNet backbone 52 in Figure 2 to refine the high-resolution feature map. This architecture uses a Gaussian heatmap 74 shown in Figure 2 for the segmentation branch to help identify which vertebrae to segment. It applies a graph convolutional network (GCN) layer for vertebra classification, allowing vertebra identification to be influenced by neighboring predictions. It uses a linear scheduling aid in gradient descent for the localization stage by combining a modified focal loss from CornerNet with a mean squared error (MSE) loss.
[0035] Discovery Branch The detection branch utilizes an anchorless approach. The full-resolution feature map from the convolutional backbone (P1) is convolved with three convolutions: the first two with 3x3x3 kernels 54 and 56 (Figure 2), and the final kernel 58 with 1x1x1 (Figure 2). The first of these three convolutions, 54 and 56, compress the P1 feature map to 128 features to reduce memory impact. The resulting feature map (after 54, 56, and 58) is sent to three separate convolution blocks: 1) heatmap (62, 65, and 69), as shown in Figure 2; 2) bounding box size (61, 64, and 68); and 3) offset prediction (60, 63, and 67).
[0036] Heatmap 74 in Figure 2 The heatmap has C channels, with each channel corresponding to a vertebra (i.e., channel 0 is C1, channel 1 is C2, etc.). Therefore, each centroid and bounding box is implicitly defined in the heatmap prediction. The maximum value of the heatmap for each channel is used to provide the centroid prediction for each vertebra. A ground truth 3D Gaussian heatmap is generated using the ground truth centroid points in the downsampled space.
[0037] Offset size 70 in Figure 2 The heatmap, i.e., the predicted centroids, are in a downsampled state. Following the work of Yi et al. [7], we shift the predicted centroids using the predicted offset coordinates to compensate for potential differences during upsampling.
[0038] Bounding box size 72 in Figure 2 Both the heatmap output and the offset are used to determine the centroid of the object in the full resolution image, and the bounding box size is used to determine a bounding box around the centroid point.
[0039] Classification Branch Each heatmap channel corresponds to an individual vertebra, and the information is combined with the offset and bounding box size to create a bounding box candidate 80 (FIG. 2) for each vertebra. However, these heatmaps do not consider adjacent vertebrae. The purpose of the classification branch is to leverage information between adjacent vertebrae to improve overall classification and detection. The initial bounding box step, step 82 (FIG. 2), uses predicted centroid locations from the heatmap and combines these predictions with feature maps from the convolutional backbone in the binary classification step 84 of FIG. 2. In the binary classification block step, features from the convolutional backbone are cropped to a region centered on the heatmap predicted centroid. The features from the convolutional backbone are then sent through two convolutional layers and then through a graph convolutional network layer that uses the shared information between vertebrae to perform the best binary classification of the vertebra 88 (see FIG. 2).
[0040] Segmentation Branch The segmentation branch semantically segments the positively detected vertebrae. Prior to extracting the positively detected regions in the previous branch, the unnormalized Gaussian heatmap 131 (Figure 2) is concatenated with the full-resolution input image in step 76 (Figure 2). This Gaussian is centered at the predicted full-resolution centroid location with a standard deviation of 4. Because the vertebra segmentation step is for the complete vertebra, the bounding box includes adjacent vertebrae due to the need to include posterior elements. The bounding box is defined and used to crop the region of interest from the input image in step 86 (Figure 2) and create a feature map from the convolutional backbone. The Gaussian heatmap ensures that the model focuses on the correct vertebra during semantic segmentation. The features derived from the convolutional backbone and Gaussian input image are first convolved as in steps 100, 102, 106, 108, and 110 of Figure 2, desampled as in steps 104 and 112, and merged as in step 114 of Figure 2. The resulting feature maps are then passed through a final convolutional layer as in steps 120, 122, 124, 126, and 128 shown in the figure to generate a segmented prediction 130.
[0041] A tool for direct estimation of vertebral stability / fracture risk based on 3D images provides a method for assessing mechanical stability and fracture risk in the spine with metastases. The method includes: a) obtaining 3D imaging data of the patient's spine 12 as shown in FIG. 1; b) inputting the imaging data into a machine learning algorithm to determine mechanical stability and fracture risk in step 14 of FIG. 1 , wherein the machine learning algorithm is configured to calculate mechanical stability and fracture risk by the following steps: i) Compute deep features from the feature extractor backbone network 28 shown in Figure 1; ii) combining these deep features with vertebra-specific features obtained from convolutional layers, graph networks, and convolutional arms; and / or extracting latent deep features for each vertebra; iii) combining the deep features in step ii) and the non-image patient-specific features 30 in Figure 1 with a high-density layer to obtain predictions on mechanical stability and fracture risk, where training is performed by an optimizer based on 3D image data of the patient's spine.
[0042] Referring again to FIG. 1, the 3D imaging data 12 may be acquired by an imaging system such as a computed tomography (CT) or magnetic resonance imaging (MRI).
[0043] Figure 1 also shows a tool that provides an interpretable automated SINS (Spine Instability Oncological Scoring) tool 36 that calculates improved individual image-based parameters used in SINS and calculates stability through a similar scoring system, providing a method for assessing mechanical stability and fracture risk in the spine with metastases. a) obtaining 3D imaging data of the patient's spine 12; b) inputting imaging data into a machine learning algorithm to determine mechanical stability and fracture risk in step 38 of FIG. 1 , the machine learning algorithm configured to calculate spinal instability tumor score elements, including quantifying bone lesion pathology in step 20 of FIG. 1 , spinal alignment in step 22 of FIG. 1 , vertebral body collapse in step 24 of FIG. 1 , and posterolateral pathology in step 26 of FIG. 1 ; c) calculating each of the elements of bone lesion, vertebral body alignment, vertebral body collapse, and posterolateral lesion includes segmenting the vertebrae (16 in FIG. 1), localizing the vertebrae (18 in FIG. 1), step 14 in FIG. 1, and all of FIG. 2, i.e., calculating deep features for the vertebrae (obtained from the feature extractor backbone network, which may be combined with vertebra-specific features obtained from the convolutional layer and convolutional arms) 14, combining the deep features 28 with the vertebra-specific features and non-image patient-specific features 30 with the dense layer to provide predictions regarding mechanical stability and fracture risk, wherein training is performed by an optimizer based on 3D image data of the patient's spine and information regarding patient outcomes; To calculate the bone lesion components, the segmented and localized vertebrae 16 are used in a histogram-based analysis step 20 to determine the volume and type of tumor tissue present within each metastatic lesion vertebra; The step 22 of calculating spinal alignment elements uses the segmented and localized vertebrae to calculate Cobb angles in the coronal and sagittal planes; The step 24 of calculating vertebral body collapse combines the segmented and localized vertebrae with another convolutional network that segments and calculates the volume of the vertebral bodies in the 3D image and uses the intact vertebral body volume of adjacent levels to estimate the percentage of vertebral body collapse; The posterolateral lesion step 26 is determined by extracting vertebra-specific deep features from a feature extractor backbone network to predict the posterolateral lesion, and training is performed by an optimizer based on the 3D imaging data of the patient's spine and posterolateral lesion markings.
[0044] Dataset used for training Prospective cohort study Over a one-year period, data including patient characteristics and serial CT images (every 4 months) were prospectively collected from 100 patients with spinal metastases using a consistent imaging protocol from T4-L5. 1000 vertebrae within the CT images were semi-automatically segmented (vertebral body, osteoblastic, mixed, and osteolytic tumor burden).
[0045] Sunnybrook SBRT Spine Patients Five hundred patients (1500 lesions) treated with spine SBRT for metastatic vertebral lesions at Sunnybrook Hospital were prospectively added to a clinical database (tumor location, SINS assessment, primary tumor location, receptor status, previous treatment, etc.) and retrospectively collected and linked with clinical images (MRI (T1 and T2), CT) and SBRT-related planning data (RT structure, RT dose, RT plan, MRI and CT registration subjects).
[0046] VerSe Dataset
[12] An open dataset of 141 training, 120 validation, and 113 test samples of 3D CT images. Each scan contains a complete vertebral segmentation, including vertebral level indicators and vertebral center of gravity locations. The dataset consists of C1-L5 vertebrae, including the rare T13 and L6 transitional vertebrae. The dataset does not contain images of implants or disease.
[0047] Step 1: Spine detection, spine segmentation, and feature extraction Input: Computed tomography imaging Output: Deep learning features, vertebral body location, vertebral body ID, vertebral body segmentation.
[0048] The VertDetect model 14 in FIG. 1 used herein can be decomposed into three main branches: detection, classification, and segmentation. The detection branch detects each vertebra in the 3D image by both determining the vertebral centroid position and placing a bounding box around the entire vertebra. The model predicts a position heat map, offset (subvoxel position adjustment relative to the heat map position), and bounding box for each vertebra in the CT volume. The classification branch utilizes shared information between each adjacent vertebra to determine which vertebrae are present in the input CT scan (vertebra ID). The segmentation branch semantically segments vertebrae that are positively detected from the classification and detection branches.
[0049] The overall architecture of VertDetect 14 in Figure 1 is shown in Figure 2. A 3D ResNet-50 [9] and a Feature Pyramid Network (FPN)
[10] , shown as 52 in Figure 2, serve as the backbone architecture. Feature maps from this backbone are then used in further downstream tasks. ResNet-50+FPN is used to extract features from CT images used for vertebra detection, classification, and segmentation tasks in this step, as well as tasks in later steps. The network was trained using data from both the 2019 and 2020 VerSe datasets
[12] . The model was trained with a combined loss.
[0050] Heatmap The heatmap has C channels, with each channel corresponding to a vertebra (i.e., channel 0 is C1, channel 1 is C2, etc.). Therefore, each centroid and bounding box is implicitly defined in the heatmap prediction. The maximum value of the heatmap for each channel is used to provide the centroid prediction for each vertebra. A ground truth 3D Gaussian heatmap is generated using the ground truth centroid points in the downsampled space. The downsampled space was used to match the P1 output of the FPN (2x downsampling from the input image size). The ground truth Gaussian distribution was constructed such that the peak-maximum was 1.0, with i=e^(-{(x-cx) 2 +(y-cy) 2 +(z-cz) 2} / (2σ 2 )), where cx, cy, and cz are the ground truth centroid coordinates, σ is the standard deviation of the distribution, and x, y, and z are points in space.
[0051] Each channel of the heatmap has a single centroid of interest, and the rest is background. To account for the large imbalance between the foreground and background, we used a modified focal loss as in [4][8][7]:
[0052]
number
[0053] where i is the ith index, p is the predicted heatmap, y is the ground truth, and N is the number of centroids. This is expressed as (1-y i ) β term, y i It differs from the original focal loss [4] by reducing the influence of predicted centroids that are close to the ground truth compared to larger predictions in the ∈[0,1] case.
[0054] It was found that a 2x downsampled heatmap performed better than the 4x downsampled heatmap used in [7], likely due to better separation between adjacent vertebrae. However, due to the size of the heatmap and class imbalance, the convergence of the modified focal loss was difficult. To improve convergence, we combined the modified focal loss with a mean squared error (MSE) loss: L heat =aL focal +bL MSE In the formula, a and b are a = (ε / ε')λ and b = ((ε-ε') / ε')λ, respectively, and are the epoch ε and the threshold epoch ε'. This epoch threshold ε' distinguishes between using a combination of MSE and modified focal loss and using only modified focal loss. The constant λ is a scaling term to address the large numerical difference between the modified focal loss and MSE functions, and is = ((1-y) / (ε'-1))ε' + ((yε'-1) / (ε'-1)), with γ = 1e-4. After epoch ε', L heat =L focal The overall loss of the heatmap is:
[0055]
number
[0056] The centroids predicted during training are the maximum points of each channel, which differs from the validation case, as explained in the "Post-processing" section below.
[0057] Offset Size The heatmap, i.e., the predicted centroids, are in a downsampled state. Even if the predicted centroids match the ground truth centroids in the downsampled state, they may not match at full resolution after upsampling due to rounding errors that can occur if the ground truth centroids are not divisible by the downsampling amount. Following the work of Yi et al. [7], we shift the predicted centroids using the predicted offset coordinates to compensate for potential differences during upsampling:
[0058]
number
[0059] i is the ith centroid, cx,i, cy,i, and cz,i is the coordinate of the i-th centroid. The brackets [ ] are the floor operation and n is the downsampling size. A smoothed L1 loss is used to regress the offset:
[0060]
number
[0061] o i and i ^ are the ground truth and prediction, respectively.
[0062] Bounding Box Size Both the heatmap output and the offset are used to determine the centroid of the object in the full resolution image. The bounding box size is used to determine a bounding box around the centroid point. The coordinates of the bounding box for the ith vertebra (bbi) are: bbi=[x0,x1,y0,y1,z0,z1] x0=c x,i -s l x1=c x,i +s r y0=c y,i -s p y1=c y,i +s a z0=c z,i -s i z1=c z,i +s s where ci is the full-scale centroid coordinate of vertebra i, and s is the size. The subscripts for size s correspond to left (l), right (r), posterior (p), anterior (a), inferior (i), and superior (s). All six are required because the centroid defined here is the vertebral body center, not the center of the object's enclosing bounding box, and mirror image sizes (i.e., left and right) are not symmetric.
[0063] The bounding box size is regressed using the log Intersection over Union (IoU) loss function [5][6]
[11] :
[0064]
number
[0065] b ~ where is the predicted bounding box and b is the ground truth. This loss is used in contrast to mean absolute error (MAE) or smoothed L1, which may allow the box size to be slightly modified if the centroid predictions are slightly shifted (i.e., if the centroid predictions are slightly offset, left is larger than right).
[0066] Classification Branch Each heatmap channel corresponds to an individual vertebra. However, these heatmaps do not consider adjacent vertebrae. If the channel corresponding to a T3 vertebra shows a high probability of being present in the scan, the probability for the adjacent T2 and T1 vertebrae should reflect that. The purpose of the classification branch is to exploit the information between adjacent vertebrae to improve overall classification and detection.
[0067] RoiAlign [3] generates C feature maps of size 7x7x7 from P1 cropping and resampling regions focused on the centroid location. The resampled feature maps are then sent through 7x7x7, followed by 1x1x1 convolutions (both with ReLU activation). Graph convolutional network (GCN) layers are used to exploit the shared information of each vertebra. The resulting features are then sent through three GCN layers, resulting in Cx1 logits corresponding to vertebrae present or absent in the scan. Since the class of each vertebra is implicitly defined, a binary classification of the heatmap channel is used to determine whether that channel and predicted centroid correspond to a positive detection. A vertebra is positively detected if sigmoid(qi) > 0.5, where qi is the logit score for the i-th vertebra in the classification branch output. The classification branch is trained using binary cross-entropy, with Lclass = BCE;
[0068]
number
[0069] yi is a binary value specifying whether the ith vertebra is active, pi is the predicted probability that vertebra i is active, and N are all possible vertebrae.
[0070] The classification of vertebrae is further refined by a graph position processing step. The predicted heatmap can have multiple local clusters, which may inaccurately correspond to adjacent vertebrae. To address possible local maxima in the heatmap prediction, a post-processing method is used to determine which local maxima from the local clusters are correct. Non-maximum suppression (NMS) is first used through a max pooling layer to select the top k candidates from each channel of the heatmap prediction.
[0071] These k candidates are then filtered by Euclidean distance to ensure that no channels are from the same local cluster, resulting in k' candidates for each heatmap channel. The logits for each k' candidate from the heatmap prediction are then averaged with the logits from the classification branch for each corresponding vertebra, scaling each corresponding vertebra based on the probability of that particular vertebra being present in the scan. The resulting k' candidates are constructed into a graph, as shown in Figure 5, based on two rules. The first rule is that the axial location of the upper node must be greater than the lower node to ensure correct vertebra ordering. The second rule is that the Euclidean distance between any two connected nodes must be greater than 3 voxels to ensure that two nodes from the same physical location are not used. The weight of each node is taken as the averaged logit. The centroid location of each vertebra is then determined by solving the graph from T (top) to B (bottom) by determining the longest path, i.e., the path with the highest sum of averaged logits; this process is also solved from B to T, and the path with the largest sum is taken as the correct path.
[0072] The model was initially trained on only the heatmap outputs for 500 epochs, referred to as self-initialization. After self-initialization, all outputs were predicted and all loss functions were used. During the first 100 epochs after self-initialization (between epochs 501 and 600 from global initiation), the model was trained using ground truth bounding boxes for the segmentation task. This was done to ensure that the model was trained to accurately segment vertebrae without being adversely affected by inaccurate bounding box predictions as the model converged. After epoch 100 after self-initialization (epoch 600 from global initiation) and when Lheat < 1.0, the model transitioned to using predicted bounding boxes. This continued until the model was complete. Overall, the model was trained for 1500 epochs.
[0073] Step 2: Bone lesions of SINS elements Input: Vertebra segmentation from step 1, computed tomography imaging Output: osteoblastic lesion volume, osteolytic lesion volume, tumor classification (osteolytic, osteoblastic, mixed).
[0074] Bone lesion elements are quantified and automated by calculating the volume of osteolytic and osteoblastic tissue in vertebrae using vertebral body segmentation. Based on the vertebral centroid location prediction from step 1, vertebrae are cropped and vertebral bodies are segmented using a second-order convolutional network. The vertebra detection model described in step 1 was combined with another deep learning network (U-Net, a convolutional neural network (CNN)) to obtain a trabecular-centered vertebral body segmentation model. Histogram-based analysis is used to define osteolytic and osteoblastic tumor lesions. Specifically, the distribution of normal bone tissue is calculated from the trabecular center of the vertebral body in unaffected vertebrae in spine images. Next, a threshold value of mean - standard deviation is used to separate osteolytic disease within the vertebral body, and mean + 2 × standard deviation is used to separate osteoblastic disease. Vertebrae are then classified as osteolytic, osteoblastic, or mixed based on the presence of osteolytic or osteoblastic disease found. The resulting volumes of osteoblastic and osteolytic tissue are used to quantify tumor lesions.
[0075] The algorithm was developed and tested on the Patient Cohort Study dataset and the Sunnybrook Spine SBRT Spine patient dataset.
[0076] Step 3: Spinal alignment of SINS elements Spine localization and segmentation were used from step 1. Cobb angles were calculated in both the coronal and sagittal planes from the slope of a spline curve generated through the vertebral centers of gravity. The calculated angles were validated against manual Cobb angle measurements performed in the sagittal and coronal planes. Manual Cobb angle measurements (local to the target vertebra and spanning the entire spinal region (cervical, thoracic, and lumbar)) were performed on a subset of Sunnybrook spine SBRT-treated patients. Measurements were performed by a colleague-trained spine surgeon based on available CT images from the spine SBRT plan, with consultation from a fellow trained staff spine surgeon.
[0077] Step 4: SINS element vertebral collapse Vertebral body collapse was calculated using a combination of deep learning-based methods. The vertebral body detection model described in Step 2 was combined with another deep learning network (U-Net, a convolutional neural network (CNN)) for vertebral body center segmentation. In this study, the CNN was trained to segment vertebrae with unfractured vertebrae from vertebrae with metastases from a patient cohort study dataset.
[0078] Vertebral collapse was quantified using the CNN for vertebral body segmentation developed in Step 1. This algorithm was used to generate segmentations for each crushed vertebra of interest and for four adjacent non-crushed vertebrae (two proximal and two distal). The locations of all vertebrae were automatically determined using the vertebra detection model described in Step 1. An estimate of the intact volume of each fractured vertebra was estimated using a polynomial fit of the four adjacent vertebral body volumes. Comparing the estimated intact volume to the volume of the fractured vertebra of interest yielded the % collapse. A linear model was used to assess collapse of vertebrae located at the edge of the image; this was most frequently the case for the fifth lumbar vertebra (the lowest vertebra in the spinal column). The automated module was found to produce results consistent with semiquantitative clinical assessment of vertebral collapse.
[0079] Step 5: SINS element posterolateral lesion The neural network from step 1 was retrained to detect posterolateral tumor lesions. Feature maps from a ResNet50+FPN backbone were combined with an additional convolutional layer trained to classify vertebrae. The complete network for this step was pretrained to identify tumor lesions in whole vertebrae and retrained to identify CT scans with tumor lesions in either the posterior or lateral elements of the vertebra. The network was trained using a spine SBRT patient dataset, with CT scans scored by radiation oncology fellows and staff as having no lesion, unilateral lesions, or bilateral lesions in the posterolateral spinal elements.
[0080] Step 7: Automating quantitative calculations of musculoskeletal health biomarkers Spine localization and segmentation were used from step 1. Lumbar (L1-L5) vertebral bodies were segmented using the vertebral body detection model described in step 2, combined with another deep learning network (U-Net, a convolutional neural network (CNN)) for vertebral body center segmentation.
[0081] The psoas muscle region from the midpoint of the L2 / L3 and L4 / L5 intervertebral discs was isolated based on the vertebral position and curvature of the spine found in step 1. The psoas muscle region was segmented by using another deep learning network (U-Net, a convolutional neural network (CNN)) for segmentation derived from the vertebral positions from step 1. The specific biomarkers extracted were bone mineral density, bone mineral density distribution of the lumbar vertebral bodies, and the volume and density of the psoas muscle.
[0082] The muscles segmented for use in biomarker calculation are defined by the region of the spine being imaged and can include: iliocostalis lumborum, iliocostalis thoracis, iliocostalis cervicis, longissimus thoracis, longissimus cervicis, longissimus capitis, spinalis thoracis, spinalis capitis, semispinalis thoracis, semispinalis cervicis, semispinalis capitis, multifidus, rotator brevis, rotator longus, interspinatus, intertransverses, levator ribs, serratus posterior superior, serratus posterior inferior, quadratus lumborum, psoas major, psoas minor, rectus capitis posterior greater, rectus capitis posterior minor, obliquus capitis superior, and obliquus capitis inferior.
[0083] Step 8: Automated Quantitative SINS The current standard for assessing mechanical stability in vertebrae with tumor lesions is the Spinal Instability Oncology Score (SINS). The SINS is a qualitative score based on six criteria: the patient's mechanical pain, vertebral level, lesion type (osteolytic, osteoblastic, or mixed), pre-existing vertebral collapse, malalignment, and the presence of posterior element lesions. The total score is classified into three categories: 0–6 (stable), 7–12 (possibly unstable), and 13–18 (unstable) [6].
[0084] This study automates and enhances SINS by using quantitative aspects of image-based biomarkers that allow for finer discrimination. For example, current SINS scores Cobb angles of 15° and 30° as having deformity, and vertebral collapse of 15% or 45% as equivalent. The approach disclosed herein allows variables to be treated as continuous, potentially improving their predictive power.
[0085] Step 9: Vertebral fracture risk A predictive model (random forest classifier, support vector machine) combined patient factors (gender, age, presence of pain), treatment factors (dose and fractionation), and image-based quantitative biomarkers from previous steps to predict spinal fractures secondary to SBRT. The new predictive model performance was evaluated against SINS, and generalizability was assessed using 5-fold cross-validation, demonstrating that quantitative image-based biomarkers improved the accuracy of predicting fractures.
[0086] We also used the deep learning model developed in Step 1 to predict vertebral fracture risk directly from images without computing intermediate image-based biomarkers. The detection model described above shares imaging features across its different subtasks, and this framework was similarly leveraged to extract features related to metastatic spinal lesions. Furthermore, the implementation of the detection model utilizes information from both the treated vertebrae and the entire spine at once. These common features were used to predict fracture risk using both vertebra- and spine-specific features.
[0087] In summary, in one embodiment, the present disclosure provides a method for assessing mechanical stability and / or fracture risk in a spine with metastasis, comprising: a) acquiring image data of a patient's spine; b) The imaging data is input into A) a computational algorithm, or B) a machine learning algorithm trained on a dataset of spine imaging to determine mechanical stability and / or fracture risk, wherein the algorithm is configured to: i) calculating mechanical stability and / or fracture risk by calculating features from a feature extractor algorithm and / or an image processing algorithm, in particular based on user input; ii) combining the above features within a computational decision-making tool; iii) combining said features in step ii) with non-image patient-specific features to obtain a prediction regarding mechanical stability and / or fracture risk using an optimization scheme based on said image data of the patient's spine.
[0088] In one embodiment, a method is provided for assessing mechanical stability and / or fracture risk of a spine with metastasis, the method comprising: a) acquiring image data of a patient's spine; b) The image data is input to A) a computational algorithm, or B) a machine learning algorithm trained on a dataset of spine images to determine mechanical stability and / or fracture risk, wherein the algorithm is configured to: i) calculating mechanical stability and / or fracture risk by calculating features from a feature extractor algorithm, such as a feature extractor backbone network; ii) combining said features in a computational decision tool including convolutional layers and vertebra-specific features derived from the convolutional arms; and / or extracting latent features for each vertebra; iii) combining the features combined with the vertebra-specific features in step ii) with non-image patient-specific features having a high density layer to obtain predictions regarding mechanical stability and / or fracture risk, wherein training is performed by an optimizer based on the imaging data of the patient's spine.
[0089] In one embodiment, the image data of the patient's spine is CT and / or MR image data.
[0090] In one embodiment, the spinal cord image dataset is comprised of CT and / or MR images comprising a clinical cohort of patients with spinal cord metastases.
[0091] In one embodiment, the feature extractor algorithm is one of the following: ResNet (Residual Network) + feature Pyramid Network (FPN), Transformer, Inception, Convolutional Neural Network (CNN), Fully Connected Neural Network (FCNN), Scale-Invariant Feature Transform (SIFT), Speeded Up Robust Features (SURF), Histogram of Oriented Gradients (HOG), Local Binary Patterns (LBP), GIST, U-Net, AlexNet, Visual Geometry Group Network (VGGNet), GoogLeNet, Radiomics, wavelet transform, Fourier transform, edge detection algorithm, region-based method, filter bank, texture descriptor, color space transformation, geometric descriptor, binary pattern description method, keypoint descriptor, Fast Retinal Keypoints (FREAK), ridge detection, scale space representation, Skeletonization, Principal Component Analysis (PCA), Singular Value Decomposition (SVD).
[0092] In one embodiment, the edge detection algorithm is one of the Sobel algorithm, the Prewitt algorithm, and the Canny algorithm; The region-based method is either the Watershed algorithm or Superpixel Segmentation; The filter bank is a Gabor filter; The texture descriptor is one of the Haralick texture descriptor and the Tamura texture descriptor; Color space conversion is one of Red, Green, Blue (RGB) to Hue, Saturation, Lightness (HSL) and Hue, Saturation, Value (HSV) conversion; The geometric descriptor is one of the Zernike moment geometric descriptor and the Fourier moment geometric descriptor); Binary pattern descriptor and one of Complete Local Binary Pattern (CLBP) and Dominant Local Binary Pattern (DLBP); Keypoint descriptor is one of BRISK (Binary Robust Invariant Scalable Keypoints) and FREAK (Fast Retina Keypoint); Ridge detection uses the eigenvalues of the Hessian matrix; The scale space representation is either a Gaussian pyramid or a Laplacian pyramid.
[0093] In one embodiment, the feature extractor algorithm is a feature extractor backbone network that is a ResNet50+FPN network.
[0094] In one embodiment, the features computed by the feature extraction algorithm and the extracted latent features for each vertebra are deep features.
[0095] In one embodiment, a method for determining neoplastic lesions of the posterolateral element of the spine from images is provided, the method comprising: a) acquiring image data of a patient's spine; b) inputting the image data into a machine learning algorithm trained on a dataset of spine images to determine posterolateral element pathology of the spine; c) Combining the segmented and localized vertebrae and calculating vertebra specific features from a feature extractor backbone network to determine posterolateral pathology, or by retraining a backbone network that is specific to posterolateral pathology classification, use the extracted features to classify using either: Classification branch, machine learning classifiers, A statistical classifier for whether a posterolateral lesion is present or not, and location (facet joint, costovertebral joint, pedicle) if a posterolateral lesion is present, training is performed by an optimizer based on the 3D imaging data of the patient's spine combined with the labeling of the posterolateral lesion.
[0096] In this embodiment, the vertebra-specific features computed from the feature extractor backbone network are deep features.
[0097] In one embodiment, a method is provided for assessing the mechanical stability and / or fracture risk of a spine with metastasis, the method comprising: a) acquiring image data of a patient's spine; b) inputting the imaging data into a machine learning algorithm trained on a dataset of spine imaging to determine mechanical stability and / or fracture risk, the machine learning algorithm configured to calculate spinal instability tumor score factors, the factors including bony lesions, spinal alignment, vertebral body collapse, and posterolateral lesions; c) calculating each of the elements of bone pathology, spinal alignment, vertebral body collapse, and posterolateral pathology includes calculating features from a feature extraction backbone network, combining said features with vertebra-specific features obtained from a convolutional layer, a graph network, and a convolutional arm, and / or segmenting and / or localizing the vertebrae by extracting latent features, and combining said features with said vertebra-specific features and non-image patient-specific features with a dense layer to generate predictions regarding mechanical stability and / or fracture risk, wherein training is performed by an optimizer based on said 3D imaging data of the patient's spine combined with information regarding patient outcomes; Use the segmented and localized vertebrae in a histogram-based analysis to calculate bone lesion elements, or use an automated approach to segment the lesions and determine the volume and type of tumor tissue present within each metastatic lesion vertebra; Calculating Cobb angles in the coronal and sagittal planes using segmented and localized vertebrae to calculate spinal alignment elements; To calculate vertebral collapse, we combine the segmented and localized vertebrae with another convolutional network that segments and calculates the volume of the vertebrae in the 3D image. a) Use intact vertebral volume at adjacent levels, or b) using a statistical model, a machine learning model, or a neural network to estimate the rate of vertebral collapse; The method of claim 8 determines posterolateral lesions.
[0098] In this embodiment, the automated approach to segmenting the lesion includes using any one of thresholding, clustering, region growing, level set methods, active contours, variational methods, graph partitioning methods, simulated annealing, watershed, model-based, classification of features extracted from the image, and neural network-based segmentation.
[0099] In this embodiment, the spinal cord image dataset is composed of CT and / or MR images comprising a clinical cohort of patients with spinal cord metastases.
[0100] In this embodiment, the components of bony lesions, spinal alignment, vertebral body collapse, and posterolateral lesions are combined with clinical pain and vertebral level to derive an automated spinal instability oncological score.
[0101] In this embodiment, the elements of bony pathology, spinal alignment, vertebral collapse, and posterolateral pathology are combined with clinical pain and other patient-specific characteristics and vertebral level to obtain an automated SINS score that predicts mechanical stability and / or fracture risk.
[0102] In this embodiment, elements of bone pathology, spinal alignment, vertebral collapse, and posterolateral pathology are combined with quantification of any one or all of musculoskeletal health, clinical pain, and other patient-specific characteristics, as well as vertebral levels, to produce an automated, enhanced SINS-type assessment that predicts mechanical stability and / or fracture risk better than existing SINSs. Musculoskeletal health is calculated by combining segmented and localized vertebrae with another convolutional network to calculate the volume of vertebral trabecular centers in the 3D image, calculate the density of vertebral trabecular bone, and calculate bone density distribution. Another convolutional network is used to further segment muscle volume and calculate muscle volume, density, and density distribution. Musculoskeletal health includes bone quality, bone mass, bone density, muscle quality, and muscle size.
[0103] The features computed by the feature extraction algorithm and the extracted latent features of each vertebra are deep features.
[0104] References [1] C.-L. Tseng et al., “Spine Stereotactic Body Radiotherapy: Indications, Outcomes, and Points of Caution.,” Glob. spine J., vol. 7, no. 2, pp. 179-197, Apr. 2017, doi: 10.1177 / 2192568217694016. [2] G. Maccauro, M. S. Spinelli, S. Mauro, C. Perisano, C. Graci, and M. A. Rosa, “Physiopathology of Spine Metastasis,” Int. J. Surg. Oncol., vol. 2011, pp. 1-8, 2011, doi: 10.1155 / 2011 / 107969. [3] K. He, G. Gkioxari, P. Dollar, and R. Girshick, “Mask R-CNN,” in Proceedings of the IEEE International Conference on Computer Vision, Oct. 2017, pp. 2980-2988, doi: 10.1109 / ICCV.2017.322. [4] T. Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollar, “Focal Loss for Dense Object Detection,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 42, no. 2, pp. 318-327, 2020, doi: 10.1109 / TPAMI.2018.2858826. [5] Z. Tian, C. Shen, H. Chen, and T. He, “FCOS: Fully convolutional one-stage object detection,” in Proceedings of the IEEE International Conference on Computer Vision, 2019, vol. 2019-Octob, pp. 9626-9635, doi: 10.1109 / ICCV.2019.00972. [6] Z. Tian, C. Shen, H. Chen, and T. He, “FCOS: A Simple and Strong Anchor-free Object Detector,” IEEE Trans. Pattern Anal. Mach. Intell., pp. 1-1, 2020, doi: 10.1109 / tpami.2020.3032166. [7] J. Yi et al., “Object-Guided Instance Segmentation with Auxiliary Feature Refinement for Biological Images,” IEEE Trans. Med. Imaging, vol. 40, no. 9, pp. 2403-2414, 2021, doi: 10.1109 / TMI.2021.3077285. [8] H. Law and J. Deng, “CornerNet: Detecting Objects as Paired Keypoints,” Int. J. Comput. Vis., vol. 128, no. 3, pp. 642-656, 2020, doi: 10.1007 / s11263-019-01204-1. [9] K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770-778, doi: 10.1109 / CVPR.2016.90.
[10] T.-Y. Lin et al., “(FPN) Feature Pyramid Networks for Object Detection,” Cvpr,2017, doi: 10.1109 / CVPR.2017.106.
[11] J. Yu, Y. Jiang, Z. Wang, Z. Cao, and T. Huang, “UnitBox: An advanced object detection network,” in MM 2016 - Proceedings of the 2016 ACM Multimedia Conference, 2016, pp. 516-520, doi: 10.1145 / 2964284.2967274.
[12] A. Sekuboyina et al., “VERSE: A Vertebrae labelling and segmentation benchmark for multi-detector CT images,” Medical Image Analysis, vol. 73. 2021, doi: 10.1016 / j.media.2021.102166.
Claims
1. 1. A method for assessing mechanical stability and / or fracture risk in a spine with metastasis, comprising: a) acquiring image data of a patient's spine; b) inputting said image data to A) a computational algorithm, or B) a machine learning algorithm trained on a dataset of spine images to determine mechanical stability and / or fracture risk; and the algorithm comprises: i) calculating mechanical stability and / or fracture risk by calculating features from a feature extractor algorithm and / or an image processing algorithm, in particular based on user input; ii) combining said features with computational decision-making tools; iii) combining the features in step ii) with non-image patient-specific features to obtain a prediction regarding mechanical stability and / or fracture risk using an optimization scheme based on the image data of the patient's spine; The method is configured to perform the following.
2. 1. A method for assessing mechanical stability and / or fracture risk in a spine with metastasis, comprising: a) acquiring image data of a patient's spine; b) inputting the image data to A) a computational algorithm, or B) a machine learning algorithm trained on a dataset of spine images to determine mechanical stability and / or fracture risk; and the algorithm comprises: i) calculating mechanical stability and / or fracture risk by computing features from a feature extractor algorithm such as a feature extractor backbone network; ii) combining said features in a computational decision tool comprising a convolutional layer and vertebra-specific features obtained from the convolutional arms; and / or extracting latent features for each vertebra; iii) combining the features combined with the vertebra-specific features and non-image patient-specific features in step ii) with a high density layer to obtain a prediction regarding mechanical stability and / or fracture risk, wherein training is performed by an optimizer based on the image data of the patient's spine; The method is configured to perform the following.
3. The method of claim 1 or 2, wherein the image data of the patient's spine is CT and / or MR imaging data.
4. The method of claim 1 or 2, wherein the dataset of spinal cord images consists of CT and / or MR images comprising a clinical cohort of patients with spinal metastases.
5. The feature extractor algorithm: ResNet (Residual Network) + Feature Pyramid Network (FPN), Transformer, Inception, Convolutional Neural Network (CNN), Fully Connected Neural Network (FCN), Scale-Invariant Feature Transform (SIFT), Speed Up Robust Features (SURF), Histogram of Oriented Gradients (HOG), Local Binary Patterns (LBP), GIST, U-Net, AlexNet, Visual Geometry Group Network (VGGNet), GoogLeNet, Radiomics, Wavelet Transform, Fourier Transform, Edge Detection Algorithm, Region-Based Methods, Filter Banks, Texture Descriptors, Color Space Conversion, Geometric Descriptors, Binary Pattern Descriptors, Keypoint Descriptors, Fast Retina Keypoint (FREAK), Ridge Detection, Scale-Space Representation, Skeletonization, Principal Component Analysis (PCA), Singular Value Decomposition (SVD), The feature extractor backbone network is one of 3. The method according to claim 1 or 2.
6. The edge detection algorithm is one of a Sobel algorithm, a Prewitt algorithm, and a Canny algorithm; The Region-based Methods is one of a Watershed algorithm and a Superpixel Segmentation algorithm; the Filter Banks are Gabor filters; The Texture Descriptors are either Haralick Texture Descriptors or Tamura Texture Descriptors; The color space conversion is one of a Red, Green, Blue (RGB) to Hue, Saturation, Lightness (HSL) conversion, a Hue, Saturation, Value (HSV) conversion; The Geometric Descriptors are any one of Zernike moments Geometric Descriptors and Fourier descriptors Geometric Descriptors; the binary pattern descriptor and one of a completed local binary pattern (CLBP) and a dominant local binary pattern (DLBP); The Keypoint Descriptors are any one of Binary Robust Invariant Scalable Keypoints (BRISK) and Fast Retina Keypoints (FREAK); The ridge detection uses eigenvalues of the Hessian matrix; The scale-space representation is one of a Gaussian pyramid and a Laplacian pyramid. The method of claim 5.
7. The method of claim 1 or 2, wherein the feature extractor algorithm is a feature extractor backbone network that is a ResNet50+FPN network.
8. The method of claim 1 or 2, wherein the features calculated by a feature extraction algorithm and the extracted latent features of each vertebra are deep features.
9. 1. A method for determining tumor lesions of the posterolateral element of the spine from images, comprising: a) acquiring image data of a patient's spine; b) inputting the image data into a machine learning algorithm trained on a dataset of spine images to determine pathology of the posterolateral element of the spine; c) combining the segmented and localized vertebrae and calculating vertebra-specific features from a feature extraction backbone network to determine posterolateral pathology, or using the extracted features by retraining a feature extraction backbone network specific to posterolateral pathology classification; Classification branch, Machine learning classifiers, or statistical classifier, Use one of the following to determine the presence, extent (bilateral, unilateral), and location (facet, costovertebral joint, pedicle) of posterolateral lesions. training is performed by an optimizer based on the 3D image data of the patient's spine combined with posterolateral lesion labeling; method.
10. The method of claim 9 , wherein the vertebra-specific features computed from a feature extractor backbone network are deep features.
11. 1. A method for assessing mechanical stability and / or fracture risk in a spine with metastasis, comprising: a) acquiring image data of a patient's spine; b) inputting the image data into a machine learning algorithm trained on a spine image dataset to determine mechanical stability and / or fracture risk, the machine learning algorithm configured to calculate spinal instability tumor score factors, the factors including bony lesions, spinal alignment, vertebral body collapse, and posterolateral lesions; and c) calculating each of the components of bone pathology, spinal alignment, vertebral body collapse, and posterolateral pathology includes calculating features from a feature extraction backbone network and combining the features with vertebra-specific features obtained from a convolutional layer, a graph network, and convolutional arms, and / or segmenting and / or localizing the vertebrae by extracting latent features and combining the features with the vertebra-specific features and non-image patient-specific features with a dense layer to generate a prediction regarding mechanical stability and / or fracture risk; To calculate said components of bone lesions, use the segmented and localized vertebrae in a histogram-based analysis or segment the lesions using an automated approach to determine the volume and type of tumor tissue present within each metastatic lesion vertebra; calculating Cobb angles in the coronal and sagittal planes using the segmented and localized vertebrae to calculate the spinal alignment components; combining the segmented and localized vertebrae with another convolutional network that segments and calculates the volume of the vertebral bodies in the 3D image to calculate the vertebral body collapse; c) using intact vertebral volume at adjacent levels, or d) estimating the rate of vertebral collapse using a statistical model, a machine learning model, or a neural network; The posterolateral lesion is determined by the method of claim 8. method.
12. 12. The method of claim 11 for assessing mechanical stability and / or fracture risk in a spine with metastasis, comprising: The automated approach to segmenting the lesion comprises: thresholding, clustering, region growing, level set methods, active contours, variational methods, graph partitioning, simulated annealing, watershed, model-based methods, classification features extracted from the image and neural network-based segmentation, 12. The method of claim 11, comprising using any one of:
13. 13. The method of claim 10, 11 or 12, wherein said dataset of spinal cord images consists of CT and / or MR images comprising a clinical cohort of patients with spinal metastases.
14. 14. The method of claim 10, 11, 12 or 13, wherein the elements of bone lesion, spinal alignment, vertebral collapse, posterolateral lesion are combined with clinical pain and vertebral level to obtain an automated spinal instability oncological score.
15. 12. The method of claim 11, wherein the elements of bony pathology, spinal alignment, vertebral collapse, posterolateral pathology are combined with clinical pain and other patient-specific characteristics and vertebral level to obtain an automated SINS score that predicts mechanical stability and / or fracture risk.
16. The above elements of bone lesions, spinal alignment, vertebral collapse, and posterolateral lesions are: musculoskeletal health, clinical pain and other patient-specific characteristics, vertebral levels, to obtain an automated, enhanced SINS-type assessment that predicts mechanical stability and / or fracture risk better than existing SINSs; The method of claim 11.
17. 17. The method of claim 16, wherein the musculoskeletal health is calculated by combining the segmented and localized vertebrae with another convolutional network to calculate vertebral trabecular center volume in the 3D image, calculate vertebral trabecular density, calculate bone density distribution, and calculate muscle volume with yet another convolutional network to calculate muscle volume, density, and density distribution.
18. 17. The method of claim 16, wherein the musculoskeletal health includes bone quality, bone mass, bone density, muscle quality, and muscle size.
19. The method of claim 11 , wherein the features calculated by the feature extraction algorithm and the extracted latent features of each vertebra are deep features.