Spinal column centrum compression fracture recognition method and system based on artificial intelligence

By introducing an improved U-Net++ network and a spatiotemporal self-attention fusion mechanism, the problem of automated identification of vertebral compression fractures was solved, achieving efficient and accurate CT image diagnosis, reducing the missed diagnosis rate and improving diagnostic efficiency and transparency.

CN120953189APending Publication Date: 2025-11-14NANJING WANGSHI INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511030111.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies for CT image recognition of vertebral compression fractures suffer from problems such as low diagnostic efficiency, reliance on manual interpretation, easy missed diagnosis of early fractures, limited model recognition accuracy, and lack of interpretability, making it difficult to meet the growing clinical needs.

Method used

An AI-based method for identifying vertebral compression fractures is employed, introducing an improved U-Net++ network with a hybrid channel and spatial attention mechanism. This network is combined with multi-scale residual enhancement and spatiotemporal self-attention fusion mechanisms to achieve deep integration of image features and structured medical knowledge. A structured report is generated through multi-task learning.

Benefits of technology

It significantly improves the accuracy of vertebral segmentation and boundary recognition, realizes fully automated processing from image input to diagnostic report, reduces the rate of missed diagnoses, improves diagnostic efficiency and accuracy, reduces the workload of doctors, and enhances the transparency and consistency of diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953189A_ABST
    Figure CN120953189A_ABST
Patent Text Reader

Abstract

The invention provides a spine centrum compression fracture recognition method and system based on artificial intelligence, and the method comprises the steps: receiving original spine CT sequence data, and outputting a preprocessed standardized image; inputting the preprocessed image into a three-dimensional U-Net + + network to obtain a centrum segmentation probability graph; outputting the segmented centrum set and the position code thereof; for each segmented centrum, extracting a corresponding area based on a centrum mask, and splicing all features to form a comprehensive feature vector of each centrum; inputting the three-dimensional image block of each vertebral body into an image flow network to extract visual semantic features; and automatically generating a structured PDF report. According to the spine centrum compression fracture recognition method and system based on artificial intelligence, the problem of image quality difference caused by different CT devices and scanning parameters is effectively solved through the multi-scale residual enhancement technology and self-adaptive window level adjustment, and standardized processing of images is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical imaging diagnostic technology, specifically to a method and system for identifying vertebral compression fractures based on artificial intelligence. Background Technology

[0002] Osteoporotic vertebral compression fracture (OVCF) is one of the most common complications of osteoporosis, with an extremely high incidence in the elderly. Statistics show that approximately 20-25% of people over 70 years of age have vertebral compression fractures of varying degrees, most of which are asymptomatic. If these fractures are not diagnosed and treated promptly, they can lead to serious complications such as kyphosis, chronic pain, and decreased lung function, significantly impacting patients' quality of life. Therefore, early and accurate identification of vertebral compression fractures is of significant clinical importance for improving patient prognosis.

[0003] CT scans, due to their excellent bone tissue resolution and three-dimensional imaging capabilities, have become an important imaging tool for diagnosing vertebral compression fractures. However, in actual clinical practice, the diagnosis of vertebral compression fractures based on CT images still faces many challenges. First, low diagnostic efficiency is the most prominent problem. A complete spinal CT scan typically contains hundreds of slice images, requiring radiologists to browse layer by layer and carefully observe the morphological changes of each vertebra, with each examination taking an average of 8-10 minutes. Under the pressure of ever-increasing imaging volumes, this traditional manual image interpretation method is no longer sufficient to meet clinical needs.

[0004] The problem of missing early-stage mild compression fractures is particularly serious. When the vertebral body compression is less than 25%, the morphological changes are often subtle and easily overlooked during routine CT scans. Studies have shown that approximately 40-50% of mild compression fractures are missed on initial CT scans. If these early fractures are not treated promptly, they may progress to severe compression or even vertebral collapse, missing the optimal treatment window.

[0005] Existing technologies are poorly adaptable to handling differences in CT equipment, scanning parameters, and image quality, and their performance deteriorates significantly in complex situations such as severe osteoporosis and high image noise. The lack of standardized image preprocessing workflows and robust feature extraction methods limits the model's generalization ability in real-world clinical settings.

[0006] In summary, existing technologies for CT image recognition of vertebral compression fractures suffer from key problems such as low diagnostic efficiency, reliance on manual interpretation, easy missed diagnosis of early fractures, limited model recognition accuracy, and lack of interpretability. There is an urgent need to develop an automatic, accurate, and interpretable intelligent recognition method to improve diagnostic efficiency and accuracy and meet the growing clinical needs. Summary of the Invention

[0007] To overcome the shortcomings of existing technologies, this invention proposes an artificial intelligence-based method and system for identifying vertebral compression fractures. An improved U-Net++ network with a hybrid channel and spatial attention mechanism is introduced, significantly enhancing the accuracy of vertebral segmentation and boundary recognition capabilities. The innovative spatiotemporal self-attention fusion mechanism achieves deep integration of image features and structured medical knowledge, enabling the model to simultaneously utilize visual information and quantitative medical indicators for comprehensive judgment.

[0008] To achieve the above objectives, this invention proposes an artificial intelligence-based method for identifying vertebral compression fractures, comprising the following steps:

[0009] S1: Image preprocessing. The system receives raw spinal CT sequence data and uses an adaptive window width and level algorithm to set the window width to 1500-2000 HU and the window level to 300-500 HU for bone-specific enhancement. A multi-scale residual enhancement module processes the image using Gaussian difference filters with three different scale kernels in parallel. The filtering results at each scale are fused using learnable weights and residually concatenated with the original image. The enhanced image voxels are resampled to an isotropic resolution of 1.0 mm × 1.0 mm × 1.0 mm, and the preprocessed standardized image is output.

[0010] S2: Vertebral body segmentation and localization. The preprocessed image is input into a 3D U-Net++ network with a hybrid channel and spatial attention module to obtain a vertebral body segmentation probability map. A pixel-level 3D mask for each vertebra is obtained through thresholding. Using the anatomical position encoding module, based on the T1-L5 vertebral body sequence determined by the spinal midline extraction algorithm, a unique learnable position embedding vector is generated for each vertebra and fused with the corresponding vertebral body feature map to complete the automatic numbering of vertebral body segments. The segmented vertebral body set and its position encoding are then output.

[0011] S3: Multi-dimensional feature extraction. For each segmented vertebra, extract the corresponding region based on its mask; calculate vertebral morphological parameters, bone density parameters, statistical morphological model anomaly score, and texture features; and concatenate all features to form a comprehensive feature vector for each vertebra.

[0012] S4: Spatiotemporal self-attention fusion diagnosis inputs three-dimensional image blocks of each vertebra into an image stream network to extract visual semantic features, and inputs the comprehensive feature vector into a feature stream network to extract high-order statistical features; intelligent fusion of the two stream features is achieved through a spatiotemporal self-attention mechanism; based on the fused features, fracture probability, risk level and pathological subtype are simultaneously output through multi-task learning;

[0013] S5: Visual report generation. For positive vertebrae with a fracture probability greater than 0.5, gradient-weighted class activation mapping technology is used to generate an attention heatmap; the heatmap is superimposed on the corresponding vertebral region of the original CT image; the diagnostic results, key feature values ​​and visual images of all vertebrae are integrated to automatically generate a structured PDF report.

[0014] Furthermore, the processing procedure of the multi-scale residual enhancement module in step S1 is as follows:

[0015] The normalized image is received as input and filtered using three Gaussian difference filters with σ values ​​of 0.5, 1.0, and 2.0 respectively, to obtain three edge enhancement results with different frequencies.

[0016] The three filtering results are then weighted and fused using learnable weights:

[0017] F enhanced =α1F1+α2F2+α3F3

[0018] F1, F2, and F3 are the Gaussian difference filtering results at three different scales, and α1, α2, and α3 are learnable weight parameters that satisfy normalization constraints, with initial values ​​set to 0.3, 0.4, and 0.3, respectively.

[0019] The enhanced result is fused with the original image using residual connections:

[0020] I enhanced =I norm +λ·F enhanced

[0021] Among them I norm For the standardized input image after window level adjustment, I enhanced For the enhanced output image, λ is the residual coefficient, with a value range of [0.1, 0.5], used to control the enhancement intensity;

[0022] The output enhanced images retain the original bone tissue information while highlighting structural details at different scales.

[0023] Furthermore, the channel and spatial hybrid attention module in step S2 includes:

[0024] Channel attention branch: The input feature map is subjected to global average pooling and global max pooling, processed through a shared two-layer fully connected network, and channel attention weights are generated using the sigmoid activation function;

[0025] Spatial attention branch: The input feature map is subjected to average pooling and max pooling along the channel dimension, processed by a 7×7×7 convolutional layer, and spatial attention weights are generated using the sigmoid activation function;

[0026] The channel attention and spatial attention weights are multiplied element-wise with the input feature map to obtain a refined feature map that simultaneously focuses on important channels and key spatial locations, which is used to enhance the information transmission of skip connections.

[0027] Furthermore, the extraction process of various features in step S3 includes:

[0028] Morphological parameter extraction: The anterior, middle, and posterior heights of the vertebral body were measured in the sagittal plane to calculate key indicators of vertebral body compression.

[0029]

[0030] Where H anterior H middle H posterior The height measurements are for the anterior, middle, and posterior edges of the vertebral body, respectively. aHCR is the anterior edge height compression ratio, and mHCR is the middle edge height compression ratio.

[0031] Statistical morphological model anomaly scoring: Based on a principal component model of vertebral body morphology constructed from healthy individuals, the degree of morphological abnormality of the vertebral body under test is calculated.

[0032]

[0033] Where X i Point cloud of the vertebral body to be tested. The average morphology of a normal vertebral body, v j Let j be the principal component vector. is the projection coefficient of the i-th vertebral body on the j-th principal component, and k is the number of principal components used (k=20). This score reflects the degree to which the vertebral body morphology deviates from the normal. Texture feature extraction: The contrast, correlation, energy and homogeneity of the gray-level co-occurrence matrix and the local binary pattern histogram are calculated in the sagittal plane of the vertebral body center to quantify the changes in trabecular bone structure.

[0034] Furthermore, the spatiotemporal self-attention fusion mechanism in step S4 is as follows:

[0035] The image stream extracts 8×8×8×512-dimensional deep feature maps through the 3D-MobileNetV2 network; the feature stream transforms the 35-dimensional input features into a 64-dimensional semantic representation through a three-layer fully connected network.

[0036] Using the feature stream output as the query vector and the image stream feature map as the key-value pair, the attention weight of the feature stream to different spatial locations in the image stream is calculated through the scaling dot product attention mechanism, thereby realizing selective aggregation of image features guided by structured features.

[0037] The attention-weighted image features are concatenated with the original feature stream output to form a 192-dimensional fused feature representation. This fusion mechanism enables the model to adaptively enhance the feature response of the corresponding lesion area in the image based on structured anomaly information.

[0038] Furthermore, the multi-task learning in step S4 employs a composite loss function:

[0039] The total loss function consists of three parts:

[0040] L total =λ1L cls +λ2L rank +λ3L type

[0041] Where L cls For fracture classification loss, L rank Losses are ranked by risk level, L type Loss due to pathological subtype classification;

[0042] The standard cross-entropy method was used for the binary classification loss of fractures; the multi-category cross-entropy method was used for the pathological subtype classification loss; and the innovative pairwise ranking loss method was used for the risk level ranking loss.

[0043]

[0044] Where N is the number of samples in the batch, s i s j The predicted risk scores for samples i and j are respectively. , where i and j are the true risk levels (1-low risk, 2-medium risk, 3-high risk), and 1 is the indicator function, which takes the value 1 when the condition is true and 0 otherwise. This loss ensures that the predicted score of high-risk samples is strictly greater than that of low-risk samples.

[0045] Set the balance coefficients λ1 = 1.0, λ2 = 0.5, and λ3 = 0.3, and perform end-to-end training using the Adam optimizer.

[0046] An artificial intelligence-based system for identifying vertebral compression fractures includes:

[0047] The image preprocessing module receives spinal CT sequence data, performs bone tissue-specific window enhancement, and improves image quality through multi-scale residual enhancement filtering. The multi-scale residual enhancement uses a formula to achieve multi-frequency information fusion.

[0048]

[0049] Among them I norm For the standardized input image, F j For the filtering result at the j-th scale, αj Here, λ represents the corresponding learnable weights, and λ is the residual coefficient.

[0050] The vertebral segmentation and localization module includes an improved 3D U-Net++ network and an anatomical location encoding submodule. The network integrates a channel-space hybrid attention mechanism in skip connections and outputs a segmentation mask and location encoding for each vertebra.

[0051] The multi-dimensional feature extraction module extracts morphological parameters, bone density parameters, statistical morphological model anomaly scores, and texture features based on the segmentation results. The anomaly score is calculated using a formula... Quantifying the degree of vertebral body morphological abnormalities, X-ray i For the point cloud of the vertebral body to be tested, X reconstructed Principal component reconstruction results, This represents the average morphology of a normal vertebral body.

[0052] The spatiotemporal self-attention fusion diagnostic module includes parallel image stream and feature stream networks. It achieves selective aggregation of the image stream by the feature stream through an attention mechanism and simultaneously outputs fracture probability, risk level and pathological subtype using a multi-task learning strategy.

[0053] The visualization report generation module generates attention heatmaps for positive vertebrae, integrates diagnostic results to generate structured PDF reports, and provides a standard interface with hospital information systems.

[0054] Furthermore, the three-dimensional U-Net++ network in the vertebral segmentation and localization module includes:

[0055] The encoder path consists of four downsampling stages, each containing two 3×3×3 convolutional layers, batch normalization, ReLU activation, and 2×2×2 max pooling, with channel numbers of 64, 128, 256, and 512 respectively. The decoder path consists of four upsampling stages, achieving feature map upsampling through transposed convolutions. Unlike the single connection of traditional U-Net, a dense skip connection strategy is adopted, connecting the features of each layer of the encoder to multiple decoder layers after processing by the attention module, enhancing feature reuse. At the end of the network, a 1×1×1 convolution and sigmoid activation are used to generate a cone segmentation probability map.

[0056] Furthermore, the network structure of the spatiotemporal self-attention fusion diagnostic module includes:

[0057] A lightweight 3D-MobileNetV2 is used as the image stream, containing 17 depthwise separable convolutional blocks. It takes a 64×64×64 vertebral body image block as input and outputs a 512-dimensional feature vector. The feature stream is a three-layer fully connected network with 256, 128, and 64 neurons, respectively. Each layer is followed by ReLU activation and Dropout with a probability of 0.3. The fusion layer reshapes the image stream output into a spatial feature map, and the feature stream output generates a query vector. Dynamic feature fusion is achieved through attention calculation. The output head contains three parallel branches responsible for fracture detection, risk grading, and subtype classification, respectively. They share fused features but are optimized independently.

[0058] Furthermore, it also includes a data augmentation module that performs the following augmentation operations during the training phase:

[0059] Random 3D rotation within ±15 degrees simulates changes in scanning angle; random scaling within 0.8-1.2 times adapts to patients of different body types; random cropping is performed while ensuring the complete vertebral body is included to increase local feature learning; a displacement field is generated using a Gaussian random field with a standard deviation of 4 pixels to simulate elastic deformation and natural morphological changes of the vertebral body; each enhancement is applied independently with a probability of 0.5 to improve the model's generalization ability and robustness.

[0060] Compared with the prior art, the beneficial effects of the present invention are:

[0061] 1. This invention provides an artificial intelligence-based method and system for identifying vertebral compression fractures. Through multi-scale residual enhancement technology and adaptive window level adjustment, it effectively solves the problem of image quality differences caused by different CT equipment and scanning parameters, achieving standardized image processing. An improved U-Net++ network with a hybrid channel and spatial attention mechanism significantly improves the accuracy of vertebral segmentation and boundary recognition capabilities. The innovative spatiotemporal self-attention fusion mechanism achieves deep integration of image features and structured medical knowledge, enabling the model to simultaneously utilize visual information and quantitative medical indicators for comprehensive judgment.

[0062] 2. This invention provides an artificial intelligence-based method and system for identifying vertebral compression fractures, achieving fully automated processing from image input to diagnostic reports, significantly improving diagnostic efficiency and transforming the tedious task of doctors browsing through layers of data into intelligent, rapid screening. The multi-dimensional feature extraction strategy combined with a statistical morphological model significantly enhances the detection capability for mild compression and early-stage fractures, effectively reducing the missed diagnosis rate.

[0063] 3. This invention provides an artificial intelligence-based method and system for identifying vertebral compression fractures. It generates an intuitive attention heatmap using gradient-weighted class activation mapping technology, clearly displaying the key lesion areas the model focuses on, thus enhancing doctors' understanding and trust in the AI ​​diagnostic results. The automatic generation function of structured reports provides comprehensive diagnostic information including quantitative indicators, risk assessments, and visualization results, making AI-assisted diagnosis more transparent and verifiable.

[0064] 4. This invention provides an artificial intelligence-based method and system for identifying vertebral compression fractures, effectively reducing the workload of radiologists and allowing them to focus more on diagnosing complex cases. Standardized diagnostic procedures and quantitative assessment indicators improve diagnostic consistency among different doctors and narrow the diagnostic gap between junior physicians and senior specialists. Early detection of fractures provides patients with optimal treatment opportunities, contributing to improved prognosis and quality of life. Attached Figure Description

[0065] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0066] Figure 1 This is a schematic diagram of the system architecture of the present invention.

[0067] Figure 2 This is a schematic diagram of vertebral body segmentation. Detailed Implementation

[0068] The technical solution of the present invention will be more clearly and completely explained below with reference to the accompanying drawings and through the description of preferred embodiments of the present invention.

[0069] like Figure 1 As shown, this system adopts an end-to-end processing flow, consisting of five core modules connected sequentially. The original spinal CT image is used as the system input and first enters the image preprocessing module. This module standardizes CT images from different sources and parameters into a unified format and enhances the contrast of bone tissue structures through three steps: bone tissue-specific window enhancement, multi-scale residual enhancement filtering, and spatial resampling.

[0070] The preprocessed image is then input into the vertebral segmentation and localization module. This module uses an improved 3D U-Net++ network combined with a channel-space hybrid attention mechanism to achieve accurate segmentation of each vertebra. At the same time, it uses anatomical location coding technology to automatically locate and number the vertebrae within the T1-L5 range.

[0071] After segmentation and localization, the system enters the multi-dimensional feature extraction module, which extracts four types of structured features for each independent vertebra: morphological parameters, bone mineral density parameters, statistical morphological model (SSM) anomaly score, and texture features. The output of this module is divided into two paths: one is a vertebral image patch that preserves spatial information, and the other is the extracted 35-dimensional feature vector.

[0072] Two streams of information are input in parallel into the spatiotemporal self-attention fusion diagnostic module. The vertebral body image blocks are processed through an image stream network, while the feature vectors are processed through a feature stream network. The two streams of features are intelligently aggregated in the fusion layer through a self-attention mechanism, and then simultaneously output three types of diagnostic results through a multi-task learning network: fracture probability (0-1 continuous value), risk level (low / medium / high), and pathological subtype (mild compression / wedge collapse / burst fracture, etc.).

[0073] The dashed box in the lower left corner of the figure represents the visualization report generation module. This module serves as an auxiliary function, receiving intermediate layer features from the fusion diagnosis module, generating an attention heatmap to explain the model's decision-making basis, and integrating all diagnostic information to automatically generate a structured report in PDF format, which is then connected to the hospital information system through the PACS interface.

[0074] As a specific implementation method, this embodiment uses 3500 spinal CT scans collected from a tertiary hospital between 2020 and 2023 for system development and validation. All CT data came from 64-slice or higher spiral CT scanners, with scanning parameters of 120kV tube voltage, automatic tube current adjustment, slice thickness of 1-3mm, and coverage of complete thoracic and lumbar vertebral segments. Of these, 2800 cases were used for model training, 350 for validation, and 350 for testing. Two radiologists with over 10 years of experience independently annotated the data, identifying a total of 1247 vertebral bodies with compression fractures.

[0075] During system implementation, the input raw DICOM format CT data is preprocessed. After reading the DICOM file header information to obtain the scanning parameters, an adaptive algorithm is used to determine the optimal window width and window level parameters. For this batch of data, the system automatically sets the window width to 1800 HU and the window level to 400 HU, creating a clear contrast between bone and soft tissue. Subsequently, a multi-scale residual enhancement module is activated, using Gaussian difference filters with σ values ​​of 0.5, 1.0, and 2.0 to process the images in parallel. In actual operation, the outputs of the three filters are fused using weights α1 = 0.28, α2 = 0.45, and α3 = 0.27 obtained through network learning. The residual coefficient λ is determined to be 0.35 after validation set optimization. After enhancement processing, all image voxels are uniformly resampled to a resolution of 1.0 mm × 1.0 mm × 1.0 mm to ensure consistency in subsequent processing.

[0076] The cone segmentation stage employs an improved 3D U-Net++ network. The network was implemented on an Ubuntu 20.04 system using the PyTorch 1.12 framework, with training hardware consisting of four NVIDIA A100 GPUs. The encoder comprises four downsampling stages, each extracting features through 3×3×3 convolutions and using batch normalization and ReLU activation. In the skip connection part, the channel attention module first performs global average pooling and max pooling on the feature map, followed by two fully connected layers with a dimensionality reduction ratio r=16 to generate channel weights. Spatial attention processes the channel-dimensional pooled feature map using 7×7×7 convolutional kernels. These two attention mechanisms sequentially apply to the original features, enhancing the network's ability to recognize cone boundaries. During training, the batch size was set to 4, the initial learning rate was 0.001, and cosine annealing was used for adjustment. After 150 training epochs, the segmentation network achieved an accuracy of 0.943 Dice coefficient on the test set.

[0077] After vertebral segmentation, the system automatically performs anatomical localization. The midline of the spine is detected using Hough transform, and combined with intervertebral spacing and morphological features, the T1 to L5 vertebrae are labeled sequentially from top to bottom. For cases with incomplete scan ranges, the system can intelligently infer the relative positions of visible vertebrae. Each vertebra is assigned a 128-dimensional learnable position embedding vector, which is continuously optimized during training to ultimately encode the vertebral body's anatomical location information.

[0078] Feature extraction was performed independently for each segmented vertebral body. In a typical case, the processing of the L2 vertebral body was as follows: After extracting the vertebral body region based on the segmentation mask, the superior and inferior endplates were automatically identified on the midsagittal plane. The anterior margin height was measured to be 18.3 mm, the mid-margin height to be 16.7 mm, and the posterior margin height to be 24.5 mm. The calculated aHCR was 0.747 and mHCR was 0.682, indicating moderate compression. The average CT value of the vertebral body cancellous bone region was 142.6 HU, with a standard deviation of 35.2 HU, and the average thickness of the anterior cortical wall was 1.8 mm. In the statistical morphological model analysis, the point cloud of this vertebral body contained 2048 sampling points. After comparison with the normal model, the abnormality score f_SSM was calculated to be 0.312, which was significantly higher than the normal range (<0.1). Gray-level co-occurrence matrix analysis showed increased contrast (2.83) and decreased homogeneity (0.41), reflecting disordered trabecular bone structure. All features were integrated to form a 35-dimensional feature vector.

[0079] Spatiotemporal self-attention fusion diagnosis performed well in the inference stage. After inputting a 64×64×64 vertebral body image block into 3D-MobileNetV2, it was processed through 17 depthwise separable convolutional blocks to extract 512-dimensional deep visual features. In parallel, 35-dimensional structured features were converted into high-level semantic representations through a three-layer fully connected network (256-128-64 neurons). During fusion, the feature stream output served as the query vector Q, which was used as the key K of the image stream feature map to calculate attention weights. Experiments showed that for significantly compressed vertebrae, the attention weights were significantly increased at the anterior vertebral margin and superior endplate region, consistent with clinical concerns. The fused 192-dimensional features, after passing through three output branches, yielded the following diagnostic results for the L2 vertebra: fracture probability 0.89, high risk (probability distribution [0.05, 0.18, 0.77]), and wedge-shaped collapse type (probability distribution [0.08, 0.81, 0.09, 0.02]).

[0080] To verify the effectiveness of multi-task learning, we compared the performance of single-task and multi-task models. During training, the three components of the composite loss function fluctuated significantly in the first 50 rounds, then gradually converged. Ultimately, the fracture classification loss decreased to 0.082, the ranking loss to 0.156, and the subtype classification loss to 0.243. On the test set, the multi-task model achieved a fracture detection sensitivity of 94.2% and a specificity of 91.7%, representing improvements of 2.1% and 3.4% respectively compared to the simple classification task. The Kendall's tau correlation coefficient for risk grading was 0.812, indicating a high degree of consistency between predicted risk and actual severity.

[0081] The visualization module uses Grad-CAM technology to generate interpretive heatmaps. For the L2 vertebra, the system automatically selects the output of the third convolutional block of the network as the target layer, calculates the gradient, and generates an activation map. The heatmap clearly shows that the model focuses on the collapsed area at the anterior superior margin of the vertebra, consistent with the radiologist's interpretation. The generated PDF report includes basic patient information, a list of positive vertebrae (L2 in this case), quantitative indicators (compression ratio, CT value, etc.), risk assessment results, and 2D / 3D visualization images. The report is automatically uploaded to the hospital's PACS system via the HL7 FHIR standard interface.

[0082] In an overall evaluation of 350 test cases, the system's average processing time was 2.3 minutes per case (Intel Xeon Gold 6248R CPU, single RTX 3090 GPU). Compared to the gold standard, the system achieved an overall diagnostic accuracy of 92.8%. Particularly for early-stage cases with mild compression (compression ratio 20-25%), the system's detection rate was 87.3%, significantly higher than the 68.5% achieved by primary care physicians. Even for patients with severe osteoporosis (mean CT value <80 HU), the system maintained an accuracy of 89.2%.

[0083] To evaluate the clinical application value of the system, we conducted a 3-month prospective application study in the radiology department of this hospital. During this period, the system assisted in the diagnosis of 1852 spinal CT scans, detecting 312 positive cases, including 28 cases of mild compression fractures that were initially missed by radiologists during their initial image review. Follow-up confirmed that 23 of these 28 cases were indeed positive, with a false positive rate of only 17.9%. After using the system, the average image review time decreased from 8.2 minutes to 4.7 minutes, and the diagnostic concordance Kappa value improved from 0.72 to 0.86.

[0084] The system demonstrated good generalization ability across different CT devices. In addition to Siemens equipment, which was the primary source of training data, the accuracy rates were 91.2% and 90.8% in tests on GE and Philips equipment, respectively. For images from low-dose CT scans (reduced by 50%), the system maintained a diagnostic accuracy of 88.5% even after enhancement by the preprocessing module.

[0085] This invention achieves intelligent and precise identification of spinal compression fractures by deeply integrating imaging features with structured medical knowledge, thereby improving diagnostic efficiency and accuracy.

[0086] like Figure 2 The diagram illustrates the preprocessing steps from raw CT data to vertebral body separation. The raw CT data is input into the system as multi-slice DICOM data. First, it undergoes a dcm2hu2nii conversion step to transform the DICOM format into the NIfTI format, which is easier for deep learning processing. Simultaneously, the pixel values ​​are converted to standard Hounsfield units.

[0087] The converted data enters the binary segmentation stage, where the system performs preliminary segmentation of the overall spinal structure from multiple perspectives, including axial, sagittal, and coronal planes, separating the spinal region from the background. Subsequently, the vertebral body and pedicle separation algorithm accurately identifies and separates the vertebral body (the area marked in red in the figure), removing accessory structures such as pedicles and spinous processes to obtain a clean vertebral body segmentation result.

[0088] The system first separates individual vertebrae from the segmented overall spine to obtain independent 3D models of the vertebral bodies. For each vertebra, the system selects three slices within the central region of the sagittal plane of the vertebra, and calculates the maximum and minimum values ​​of the anterior and posterior edges by measuring the height of the vertebral body's anterior and posterior edges on these slices.

[0089] In the compression fracture assessment stage, the system analyzes the morphological characteristics of the vertebral body, particularly the height changes of its anterior, middle, and posterior margins (as shown by the dotted lines in the figure), to determine whether compression changes exist. This measurement data will serve as the basis for calculating morphological parameters in the subsequent feature extraction module.

[0090] The entire process automates the process from raw CT images to precise segmentation and preliminary morphological analysis of individual vertebrae, laying the foundation for subsequent feature extraction and intelligent diagnosis.

[0091] The above-described specific embodiments are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Various modifications, substitutions, and improvements made by those skilled in the art to the technical solutions of the present invention based on the provided textual description and drawings, without departing from the design concept and spirit of the present invention, should all fall within the scope of protection of the present invention. The scope of protection of the present invention is determined by the claims.

Claims

1. A method for identifying vertebral compression fractures based on artificial intelligence, characterized in that, Includes the following steps: S1: Receives raw spinal CT sequence data, and uses an adaptive window width and window level algorithm to set the window width to 1500-2000 HU and the window level to 300-500 HU for bone tissue-specific enhancement; processes the image in parallel using three Gaussian difference filters with different scale kernels through a multi-scale residual enhancement module, fuses the filtering results of each scale through learnable weights, and performs residual connection with the original image; resamples the enhanced image voxels to an isotropic resolution of 1.0 mm × 1.0 mm × 1.0 mm, and outputs the preprocessed standardized image; S2: Input the preprocessed image into a 3D U-Net++ network with a channel and spatial hybrid attention module to obtain a vertebral segmentation probability map; obtain a pixel-level 3D mask for each vertebra through thresholding; use the anatomical position encoding module to generate a unique learnable position embedding vector for each vertebra based on the T1-L5 vertebral sequence determined by the spinal midline extraction algorithm and fuse it with the corresponding vertebral feature map to complete the automatic numbering of vertebral segments, and output the segmented vertebral set and its position encoding. S3: For each segmented vertebra, extract the corresponding region based on the vertebral mask; calculate the vertebral morphological parameters, bone density parameters, statistical morphological model anomaly score and texture features; and concatenate all features to form a comprehensive feature vector for each vertebra. S4: Input the 3D image block of each vertebra into the image stream network to extract visual semantic features, and input the comprehensive feature vector into the feature stream network to extract high-order statistical features; realize the intelligent fusion of the two stream features through a spatiotemporal self-attention mechanism; based on the fused features, output the fracture probability, risk level and pathological subtype simultaneously through multi-task learning; S5: For positive vertebral bodies with a fracture probability greater than 0.5, gradient-weighted class activation mapping technology is used to generate an attention heatmap; the heatmap is superimposed on the corresponding vertebral body region of the original CT image; the diagnostic results, key feature values ​​and visualization images of all vertebral bodies are integrated to automatically generate a structured PDF report.

2. The method for identifying vertebral compression fractures based on artificial intelligence according to claim 1, characterized in that, The processing procedure of the multi-scale residual enhancement module in step S1 is as follows: The normalized image is received as input and filtered using three Gaussian difference filters with σ values ​​of 0.5, 1.0, and 2.0 respectively, to obtain three edge enhancement results with different frequencies. The three filtering results are then weighted and fused using learnable weights: F enhanced =α1F1+α2F2+α3F3 F1, F2, and F3 are the Gaussian difference filtering results at three different scales, and α1, α2, and α3 are learnable weight parameters that satisfy normalization constraints, with initial values ​​set to 0.3, 0.4, and 0.3, respectively. The enhanced result is fused with the original image using residual connections: I enhanced =I norm +λ·F enhanced Where I norm For the standardized input image after window level adjustment, I enhanced The enhanced output image is represented by λ, which is the residual coefficient with a value range of [0.1, 0.5], used to control the enhancement intensity.

3. The method for identifying vertebral compression fractures based on artificial intelligence according to claim 1, characterized in that, The channel and spatial hybrid attention module in step S2 includes a channel attention branch and a spatial attention branch; The channel attention branch performs global average pooling and global max pooling on the input feature map, processes it through a shared two-layer fully connected network, and uses the sigmoid activation function to generate channel attention weights. The spatial attention branch performs average pooling and max pooling on the input feature map along the channel dimension, processes it through a 7×7×7 convolutional layer, and uses the sigmoid activation function to generate spatial attention weights. The channel attention and spatial attention weights are then multiplied element-wise with the input feature map to obtain a feature map that simultaneously focuses on important channels and key spatial locations.

4. The method for identifying vertebral compression fractures based on artificial intelligence according to claim 1, characterized in that, The extraction process of various features in step S3 includes: Measure the anterior, mid, and posterior heights of the vertebral body in the sagittal plane to calculate key indicators of vertebral body compression: Where H anterior H middle H posterior The height measurements are for the anterior, middle, and posterior edges of the vertebral body, respectively. aHCR is the anterior edge height compression ratio, and mHCR is the middle edge height compression ratio. Based on a principal component model of vertebral body morphology constructed from healthy individuals, the degree of morphological abnormality of the vertebral body under test is calculated: Where X i Point cloud of the vertebral body to be tested. The average morphology of a normal vertebral body, v j Let j be the principal component vector. is the projection coefficient of the i-th vertebral body on the j-th principal component, k is the number of principal components used, k=20, and the score reflects the degree of deviation of the vertebral body morphology from the normal. The contrast, correlation, energy, and homogeneity of the gray-level co-occurrence matrix, as well as the local binary pattern histogram, are calculated in the sagittal plane at the center of the vertebral body to quantify the changes in trabecular bone structure.

5. The method for identifying vertebral compression fractures based on artificial intelligence according to claim 1, characterized in that, The spatiotemporal self-attention fusion mechanism in step S4 is as follows: The image stream is processed by the 3D-MobileNetV2 network to extract 8×8×8×512-dimensional deep feature maps; The feature stream transforms the 35-dimensional input features into a 64-dimensional semantic representation through a three-layer fully connected network. Using the feature stream output as the query vector and the image stream feature map as the key-value pair, the attention weight of the feature stream to different spatial locations in the image stream is calculated through the scaling dot product attention mechanism, thereby realizing selective aggregation of image features guided by structured features. The attention-weighted image features are concatenated with the original feature stream output to form a 192-dimensional fused feature representation.

6. The method for identifying vertebral compression fractures based on artificial intelligence according to claim 1, characterized in that, The multi-task learning in step S4 employs a composite loss function: The total loss function consists of three parts: L total =λ1L cls +λ2L rank +λ3L type Where L cls For fracture classification loss, L rank Losses are ranked by risk level, L type Loss due to pathological subtype classification; The standard cross-entropy was used for the binary classification loss of fractures. The pathological subtype classification loss uses multi-class cross-entropy; Risk level ranking loss employs an innovative pairwise ranking loss: Where N is the number of samples in the batch, s i s j The predicted risk scores for samples i and j are respectively. , i and j are the true risk levels of samples i and j respectively, 1-low risk, 2-medium risk, 3-high risk, 1 is the indicator function, which takes 1 when the condition is true and 0 otherwise. The loss ensures that the predicted score of high-risk samples is strictly greater than that of low-risk samples. Set the balance coefficients λ1 = 1.0, λ2 = 0.5, and λ3 = 0.3, and perform end-to-end training using the Adam optimizer.

7. An artificial intelligence-based vertebral compression fracture identification system, applicable to the artificial intelligence-based vertebral compression fracture identification method according to any one of claims 1-6, characterized in that, include: The image preprocessing module receives spinal CT sequence data, performs bone tissue-specific window enhancement, and improves image quality through multi-scale residual enhancement filtering. The multi-scale residual enhancement uses a formula to achieve multi-frequency information fusion. Where I norm For the standardized input image, F j For the filtering result at the j-th scale, α j Here, λ represents the corresponding learnable weights, and λ is the residual coefficient. The vertebral segmentation and localization module includes an improved 3D U-Net++ network and an anatomical location encoding submodule. The network integrates a channel-space hybrid attention mechanism in skip connections and outputs a segmentation mask and location encoding for each vertebra. The multi-dimensional feature extraction module extracts morphological parameters, bone density parameters, statistical morphological model anomaly scores, and texture features based on the segmentation results. The anomaly score is calculated using a formula... Quantifying the degree of vertebral body morphological abnormalities, X-ray i For the point cloud of the vertebral body to be tested, X reconstructed Principal component reconstruction results, This represents the average morphology of a normal vertebral body. The spatiotemporal self-attention fusion diagnostic module includes parallel image stream and feature stream networks. It achieves selective aggregation of the image stream by the feature stream through an attention mechanism and simultaneously outputs fracture probability, risk level and pathological subtype using a multi-task learning strategy. The visualization report generation module generates attention heatmaps for positive vertebrae, integrates diagnostic results to generate structured PDF reports, and provides a standard interface with hospital information systems.

8. The artificial intelligence-based spinal vertebral compression fracture identification system according to claim 7, characterized in that, The 3D U-Net++ network in the vertebral body segmentation and localization module includes: The encoder path consists of four downsampling stages, each containing two 3×3×3 convolutional layers, batch normalization, ReLU activation, and 2×2×2 max pooling, with channel numbers of 64, 128, 256, and 512 respectively; the decoder path consists of four upsampling stages, which achieve feature map upsampling through transposed convolution. A dense skip connection strategy is adopted to connect the features of each layer of the encoder to multiple decoder layers after processing by the attention module, thereby enhancing feature reuse; At the end of the network, 1×1×1 convolution and sigmoid activation are used to generate a cone segmentation probability map.

9. The artificial intelligence-based spinal vertebral compression fracture identification system according to claim 7, characterized in that, The network structure of the spatiotemporal self-attention fusion diagnostic module includes: The lightweight 3D-MobileNetV2 is used as the image stream, which contains 17 depthwise separable convolutional blocks. The input is a 64×64×64 cone image block, and the output is a 512-dimensional feature vector. The feature flow is a three-layer fully connected network with 256, 128, and 64 neurons, each layer followed by ReLU activation and Dropout with a probability of 0.3; The fusion layer reshapes the image stream output into a spatial feature map, and the feature stream output generates a query vector. Dynamic feature fusion is achieved through attention calculation. The output head contains three parallel branches responsible for fracture detection, risk grading, and subtype classification, which share fusion features but are optimized independently.

10. The artificial intelligence-based spinal vertebral compression fracture identification system according to claim 7, characterized in that, It also includes a data augmentation module that performs the following augmentation operations during the training phase: Random three-dimensional rotation is performed within a range of ±15 degrees to simulate changes in the scanning angle; Random scaling within the range of 0.8-1.2 times can be applied to accommodate patients of different body types. Random cropping is performed while ensuring the complete vertebral body is included, thereby increasing local feature learning; A displacement field was generated using a Gaussian random field with a standard deviation of 4 pixels to simulate the natural morphological changes of the cone. Each enhancement is applied independently with a probability of 0.5 to improve the model's generalization ability and robustness.

Citation Information

Cited By

  • Medical image lesion segmentation method and system for radiology department

    CN122023446A

  • Osteoporosis assessment method and system based on spine CT image and application method of system

    CN122200080A