Osteoporotic fracture predicting and positioning method based on CT (Computed Tomography) and MRI (Magnetic Resonance Imaging) multimodal fusion

This method for predicting and locating osteoporotic fractures by fusing CT and MRI multimodal data, combined with dual-stream feature extraction and Mamba fusion modules, solves the problems of single-modal dependence and single task in existing technologies. It realizes intelligent diagnosis of fracture risk prediction and location, improves diagnostic accuracy and computational efficiency, and is suitable for real-time diagnosis in primary hospitals.

CN121811067AInactive Publication Date: 2026-04-07JIANGSU TIANYING MEDICAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-04-07
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies for intelligent prediction of osteoporotic fractures suffer from problems such as strong unimodal dependence, lack of multimodal information fusion, single task, insufficient utilization of structural features, and poor model generalization and clinical interpretability, which limit diagnostic accuracy and generalizability.

Method used

An osteoporotic fracture prediction and localization method based on CT and MRI multimodal fusion is adopted. A three-dimensional bimodal fracture localization model is constructed, including a dual-stream feature extraction encoder, a Mamba fusion module and a dual-task inference module, to achieve fracture risk prediction and vertebral level localization. Feature fusion is performed using the Mamba fusion module and end-to-end training is carried out through a joint loss function.

Benefits of technology

It improves the ability to identify early pathological changes in fractures, realizes intelligent diagnosis from fracture risk prediction to specific fracture vertebral body localization, improves the accuracy and interpretability of diagnosis, and reduces computational resource consumption, making it suitable for real-time diagnosis in primary hospitals and routine radiology environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811067A_ABST
    Figure CN121811067A_ABST
Patent Text Reader

Abstract

The invention relates to a CT (Computed Tomography) and MRI (Magnetic Resonance Imaging) multi-modal fusion-based osteoporotic fracture predicting and positioning method, which comprises the following steps of: acquiring and screening a CT image and an MRI image, and marking a fracture region; constructing a three-dimensional bimodal fracture positioning model which comprises a double-flow feature extraction encoder, a Mangban fusion module and a double-task reasoning module, and performing end-to-end training on the constructed model by using the constructed data set to obtain a trained osteoporotic fracture prediction and positioning model; and predicting the CT image and the MRI image to be detected, outputting a fracture probability prediction result, and calibrating on the images. According to the method, structural details of CT and tissue contrast information of MRI are combined, so that the recognizability of fine fracture lines and early bone changes is improved; classification and positioning tasks are executed cooperatively, so that process complexity and error accumulation caused by series connection of multiple models are avoided; the Mangbar fusion module with linear complexity significantly reduces the consumption of computing resources, and improves the reasoning speed and the system stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of medical image analysis, computer-aided diagnosis, and artificial intelligence, specifically to a method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion. Background Technology

[0002] Osteoporotic vertebral compression fracture (OVCF) is one of the most common and disabling chronic bone diseases in aging societies worldwide. According to the World Health Organization (WHO), more than 200 million people globally are affected by osteoporosis, with vertebral compression fractures accounting for more than one-third of these cases. These fractures not only severely impact patients' quality of life but also significantly increase the risk of subsequent fractures. Studies show that the risk of refracture within two years of the initial vertebral fracture can increase by 2–5 times, exhibiting a typical "fracture cascade effect."

[0003] In clinical diagnosis, dual-energy X-ray absorptiometry (DXA) has long been considered the "gold standard" for osteoporosis. However, its two-dimensional projection characteristics cannot reflect the microstructure of trabecular bone and the integrity of cortical bone. More than 80% of patients with osteoporotic fractures do not meet the diagnostic criteria for DXA bone mineral density, resulting in a large number of potentially high-risk individuals not being identified in a timely manner. Meanwhile, routine CT and MRI scans are commonly found in images of thoracic and lumbar spine diseases, lungs, and abdomen, containing rich information on skeletal structure. Therefore, how to utilize these "opportunistic images" to achieve early prediction and localization of fracture risk has become an important direction in the field of artificial intelligence medical imaging. From an industry trend perspective, many domestic and international AI medical companies have begun to develop "intelligent osteoporosis assessment" products. However, most algorithms on the market are currently limited to single-modal (CT or X-ray) images, lacking multimodal information fusion and automated identification of specific vertebral fracture sites. The need for intelligent clinical diagnosis and risk stratification remains unmet.

[0004] Currently, intelligent prediction of osteoporotic fractures mainly focuses on CT image analysis. For example, Chinese patent CN118692613B from Qilu Hospital of Shandong University discloses an artificial intelligence prediction method for spinal refracture in OVCF patients based on CT images. This method combines CT radiomics features with clinical parameters, uses a U-Net network to automatically segment ROIs, and utilizes a 3D DenseNet-121 network for refracture risk prediction. While the feasibility of this method has been validated on multi-center data, it still has significant limitations: First, the model relies entirely on CT single-modality images, failing to integrate soft tissue and bone marrow signal information from MRI images, making it difficult to reflect early pathological changes after fracture. Second, the research scenario is limited to "refracture prediction" after a fracture has already occurred, lacking the ability to identify and locate the initial fracture. Third, the ROI construction method is relatively coarse, failing to distinguish the structural differences between cortical and cancellous bone. Fourth, the model lacks automatic fracture localization functionality, making it impossible to achieve full-process intelligentization from screening to specific lesion identification.

[0005] Furthermore, Chinese patent CN117095817A from Shanghai First People's Hospital discloses a feature-based and CT image-based fracture risk prediction method. Building upon traditional clinical risk models, this method incorporates CT image features, comprehensively analyzes indicators such as bone mass, muscle mass, and bone cross-sectional area, and combines these with clinical characteristics such as age and education level to establish a deep learning model for fracture risk grading. This method is the first to systematically integrate musculoskeletal synergistic features, improving upon the shortcomings of solely relying on bone mineral density indicators. However, this technology is still based only on a single lumbar spine CT image and does not incorporate MRI modal information, resulting in weak cross-site and cross-modal generalization of the model. Simultaneously, its output is limited to risk level and does not achieve automatic identification and localization of specific fractured vertebrae, making it difficult to meet the needs of accurate clinical diagnosis.

[0006] For example, Chinese patent CN115644904B discloses a deep learning-based CT image-based vertebral refracture prediction system for osteoporotic patients. This system utilizes convolutional neural networks to directly extract depth features from CT cross-sectional images to predict the risk of postoperative refracture in osteoporotic patients. The system achieves vertebral-level risk assessment and establishes a relatively standardized image preprocessing and training process, showing promising application prospects. However, this method relies solely on CT image data and does not integrate highly sensitive information from MRI regarding bone marrow edema and early bone metabolism changes. Furthermore, its input is primarily two-dimensional cross-sectional images, resulting in insufficient utilization of spatial structural features. Consequently, the model has limitations in early fracture identification and multimodal generalization.

[0007] In summary, while existing technologies have made some progress in the intelligent prediction of osteoporotic fractures, they generally suffer from the following shortcomings: First, strong unimodal dependence: Current methods mainly rely on CT images, lacking utilization of bone marrow edema, fatty infiltration, and early metabolic changes in MRI, thus failing to achieve multimodal complementarity. Second, limited task scope: Most models can only predict "whether a fracture has occurred" or "risk of refracture," failing to further automate the automatic localization and visualization of the fracture site. Third, insufficient utilization of structural features: Some algorithms only use two-dimensional CT slices for modeling, failing to fully explore three-dimensional spatial features and inter-slice structural information. Fourth, poor model generalization and clinical interpretability: The lack of multi-center standardization processing and cross-modal fusion mechanisms limits the stability and clinical generalizability of the algorithms. Summary of the Invention

[0008] To address the aforementioned problems, this invention provides a method for predicting and locating osteoporotic fractures based on multimodal fusion of CT and MRI. By considering both cortical bone structural information and bone marrow signal characteristics, it achieves intelligent diagnosis from fracture risk prediction to vertebral level localization, providing clinicians with a more accurate and interpretable auxiliary decision-making tool. The specific technical solution is as follows: A method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion includes the following steps: Step 1: Obtain and filter CT and MRI images containing osteoporotic fractures and osteoporotic but unfractured fractures in the thoracic to lumbar spine region, and mark the fracture areas; Step 2: Construct a three-dimensional dual-modal fracture localization model, including: A dual-stream feature extraction encoder is used to extract multi-scale three-dimensional features from input CT and MRI images, respectively. The Mamba fusion module is used to fuse the CT and MRI multi-scale features extracted by the dual-stream feature extraction encoder in stages and output multi-scale three-dimensional fusion features. The dual-task reasoning module, based on the multi-scale three-dimensional fusion features, performs fracture classification and fracture region localization tasks in parallel. Step 3: Using the dataset constructed in Step 1, the model constructed in Step 2 is trained end-to-end by optimizing the joint loss function, which includes classification loss and localization loss, to obtain a trained osteoporotic fracture prediction and localization model. Step 4: Predict the fracture probability of the CT and MRI images to be tested, output the predicted fracture probability, and mark it on the images.

[0009] Preferably, in step 1, fracture cases caused by high-energy or violent injuries are excluded during screening, as well as samples where the time interval between CT and MRI scans exceeds a predetermined threshold or the fracture sites are inconsistent; confirmed osteoporotic fracture samples are retained as positive samples, and osteoporotic non-fracture samples with similar age, gender, and examination equipment are selected as control samples; when marking the fracture area, a professional physician performs three-dimensional marking of the fracture area on the CT images of the positive samples to generate a ROI mask.

[0010] Preferably, the Mamba fusion module in step 2 includes a state space channel exchange module and a dual-state space fusion module; the state space channel exchange module is used to exchange and reorganize the channel dimensions of the input CT and MRI features, and process them through visual state space blocks to achieve shallow feature fusion; the dual-state space fusion module is used to project the shallow fused features to the hidden state space, perform feature interaction through a gated dual attention mechanism, and then project them back to the original feature space and add them to the initial features to achieve deep feature fusion.

[0011] Preferably, the dual-task reasoning module in step 2 includes: a fracture classification reasoning module, which performs three-dimensional global average pooling on the fused global features, and then processes them through a fully connected layer to output the predicted probability that the sample is a fracture; and a fracture localization reasoning module, which predicts the three-dimensional bounding box coordinates and confidence of the fracture region through a three-dimensional detection head based on multi-scale fused features, and outputs the final localization result after non-maximum suppression post-processing.

[0012] Preferably, the joint loss function in step 3 is: ; In the formula, This indicates the total loss resulting from the combined efforts of both tasks. This represents the fracture classification loss, including the binary cross-entropy loss with class weights; This indicates the loss in CT fracture localization, including 3D bounding box regression loss and confidence loss; =1.0、 =5.0 are the weight coefficients for the two types of losses, used to balance the training priority of classification and localization tasks.

[0013] Furthermore, the binary cross-entropy loss with class weights: ; In the formula, N represents the total number of samples in the training batch; and These represent the category weights for "fractured samples" and "non-fractured samples," respectively. =0.7, =0.3; This represents the true classification label of the nth sample. =1 indicates that the nth sample is a fracture sample. =0 indicates that the nth sample is a non-fracture sample; This represents the predicted probability that the nth sample in the network output is a fracture sample.

[0014] Preferably, the 3D bounding box regression loss and confidence loss are: ; In the formula, This represents the fracture bounding box regression loss, used to optimize the spatial matching between the predicted bounding box and the ground truth bounding box. This represents the prediction box confidence loss, used to optimize the accuracy of the prediction box as a fracture area.

[0015] Preferably, when outputting the fracture probability prediction result, the fracture probability output by the model and the voxel coordinates of the three-dimensional positioning box are combined with the physical resolution of the CT image to convert them into physical coordinates, and a visualization report containing quantitative indicators is generated.

[0016] A computer-readable storage medium having a computer program / instructions thereon, which, when executed by a processor, implements the steps of the described method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion.

[0017] A computer program product includes a computer program / instructions that, when executed by a processor, implement the steps of a method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion.

[0018] Compared with the prior art, the present invention has the following beneficial effects: This invention provides a method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion. By combining the structural details of CT with the tissue contrast information of MRI, it improves the identifiability of subtle fracture lines and early bone changes. The classification and localization tasks are executed collaboratively, avoiding the process complexity and error accumulation caused by multiple models being serially linked. The linear complexity Mamba fusion module significantly reduces computational resource consumption and improves inference speed and system stability. The model has real-time inference performance and hardware adaptability, making it suitable for primary hospitals and routine radiology environments, and has high practical application value. Attached Figure Description

[0019] Figure 1 It is a three-dimensional dual-modal fracture localization model diagram; Figure 2 These are model diagrams: (a) Mamba Fusion Module Diagram; (b) State Space Channel Exchange Module Diagram; (c) Visual State Space Module Diagram. Figure 3This is a diagram of a dual-state space fusion module; Figure 4 This is a diagram of a dual-task reasoning module. Detailed Implementation

[0020] The present invention will now be further described with reference to the accompanying drawings.

[0021] like Figures 1 to 4 As shown, a method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion includes the following steps: Step 1: Dataset Construction To establish a multimodal deep learning model for intelligent prediction and localization of osteoporotic fractures, a dataset containing CT and MRI images was constructed. The data comes from real clinical cases and covers the thoracic to lumbar spine regions, which can comprehensively reflect the imaging characteristics of osteoporotic fractures.

[0022] During the data acquisition phase, spinal imaging data from over 900 patients were included. To ensure the accuracy and consistency of model training, the raw data underwent rigorous screening and cleaning: firstly, fracture cases caused by high-energy or violent injuries (such as car accidents, falls, traumatic burst fractures, etc.) were excluded to eliminate interference from non-osteoporotic mechanisms; secondly, samples with CT and MRI scan intervals exceeding 30 days or inconsistent fracture sites were excluded to ensure image alignment and correspondence to the same lesion area. After screening, 385 patient samples confirmed by imaging and clinical diagnosis of osteoporotic fractures were retained. All cases had paired CT and MRI data for the thoracic to lumbar spine regions, covering the T1–L5 vertebral levels. Simultaneously, 300 control samples of osteoporotic patients without fractures were selected. The two groups were similar in age, gender, and examination equipment to reduce model bias. For data annotation, all samples underwent structured processing: two radiologists with over ten years of experience independently determined the presence of fractures and labeled the CT and MRI images with "fracture / non-fracture" categories, respectively. For fracture-positive samples, the fracture area was further precisely delineated on the CT images to form a Region of Interest (ROI) mask for localization supervision during model training. The annotation process was completed using a 3D Slicer. The window width of the CT images was fixed at 1300 HU, and the window level was fixed at 300 HU. All images were 3D cropped to a size of 1×64×256×256. Finally, the complete dataset was randomly divided into a training set (70%), a validation set (15%), and a test set (15%) based on the patients, maintaining a consistent ratio of fracture to non-fracture samples during the partitioning process.

[0023] Through the above steps, a high-quality multimodal paired image dataset containing 385 fracture samples and 300 non-fracture samples was constructed. This dataset not only achieves unified labeling of CT and MRI images, but also provides vertebral fracture location annotation, which can simultaneously support model training and validation for fracture risk prediction and specific fracture vertebral location tasks.

[0024] Step 2, Model Building A three-dimensional bimodal fracture localization model is constructed, whose core architecture includes a two-stream feature extraction encoder, a FusionMamba Block (FMB) module, and a dual-task inference module (classification inference module + localization inference module), such as... Figure 1 As shown, the specific steps are as follows: 2.1 Dual-stream feature extraction encoder The preprocessed input 3D tensor: (CT single-channel density map) (MRI single-channel soft tissue image). A dual-stream feature extraction encoder was constructed using 3DResNet-50 as the basic backbone network to extract 3D features from CT and MRI respectively. CT branch: Input After five 3D convolutional stages (Stage 1 to Stage 5), each stage contains a 3×3×3 3D convolutional layer, a batch normalization layer, a ReLU activation function, and residual connections, outputting CT feature volumes at five scales. : MRI branch: The structure is completely identical to the CT branch, input Output MRI feature volumes at the corresponding scale The dimensions match the CT branch.

[0025] 2.2 Fusion Mamba Block (FMB) Select and As input, each scale corresponds to one FMB. The core includes a State Space Channel Swapping (SSCS) module and a Dual State Space Fusion (DSSF) module to achieve phased feature fusion. 2.2.1 SSCS Module (Shallow Fusion) Preliminary interaction of modal features is achieved through "channel exchange + Visual State Space (VSS) module," such as... Figure 2 As shown: 1. Channel switching: and The channel dimension C of (i=3,4,5) is divided into 4 parts, and reorganized according to "CT1+MRI2+CT3+MRI4" to form Reconstructed according to "MRI1+CT2+MRI3+CT4" Dimensional preservation ; 2. VSS block: will , The slice is unfolded into six 1D sequences along six directions (slice top → bottom, bottom → top; horizontal left → right, right → left; vertical front → back, back → front). After processing by S6 blocks (selective state space blocks), it is reconstructed into a 3D feature volume, and the shallow fusion feature is output. and .

[0026] 2.2.2DSSF Module (Deep Fusion) Deep feature interaction is achieved in the hidden state space through gated dual attention. (See...) Figure 3 : 1. Hidden State Projection: , After normalization, linear layers, depthwise separable convolution, and SiLU activation, the projection is the hidden state feature: in , , where p is the dimension of the hidden state.

[0027] 2. Gating parameter generation: Gating vectors are generated through a linear layer and SiLU activation to control the intensity of modal interactions. 3. Hidden State Dual Attention Interaction: Achieving Complementary Transfer of CT and MRI Features via Gating Vectors: Where · represents the element-wise product of the 3D tensor. 4.3D Projection and Residual Connections: Project the hidden state after interaction back to the original dimension and add residual connections to preserve shallow information. 5.3D Feature Enhancement and Fusion: Deep complementary features are added to the initial features to generate enhanced features, which are then fused element-by-element. The final output is a multi-scale 3D fusion feature. , , .

[0028] 2.3 Dual-Task Reasoning Module (Prediction and Localization) Based on multi-scale fusion features , , To achieve CT fracture classification and regional localization, see Figure 4 .Will downsampling to Dimensions and summation Then downsampling to Dimensions and summation ;Will Upsampling to Dimensions and summation Then Upsampling to Dimensions and summation ; 2.3.1 Fracture Classification Reasoning Module (Binary Classification) Based on fusion features Achieving fracture prediction: 1. To Perform 3D global average pooling to obtain a 1×512 global feature vector. ; 2. After two fully connected layers (512→256→2) and Softmax activation, the fracture probability is output. Judgment rules: A value >0.5 indicates a fracture sample; otherwise, it indicates a non-fracture sample.

[0029] 2.3.2 Fracture Localization Reasoning Module (3D Detection) Based on multi-scale fusion features Achieving CT fracture localization: 3D inspection head: Features are processed through a 3×3×3 convolutional layer to output 3D bounding box coordinates. , , (w,h,d) Confidence and categories (only "fracture" category 1); Post-processing: 3DNMS (IoU threshold 0.3) was used to remove overlapping bounding boxes, retaining... A prediction bounding box with a value >0.5 is output as the detection bounding box for the fracture area on the CT image.

[0030] Step 3: Train the model The network described in this invention was trained using CT-MRI image data annotated by orthopedic experts (including binary labels of "fracture / non-fracture" and 3D true bounding boxes of fracture areas in CT images). The loss function was defined as follows: in, This indicates the total loss resulting from the combined efforts of both missions. This represents the fracture classification loss (binary cross-entropy loss with class weights). This indicates the loss in CT fracture localization (including 3D bounding box regression loss and confidence loss). =1.0、 =5.0 are the weight coefficients for the two types of losses, used to balance the training priority of classification and localization tasks. The definitions of each loss term are as follows: 1. Fracture classification loss: Binary cross-entropy loss with class weights to resolve class imbalance. Where N represents the total number of samples in the training batch. and The categories of "fractured samples" and "non-fractured samples" are respectively represented by their respective weights in this patent. =0.7, =0.3, This represents the true classification label of the nth sample. =1 indicates that the nth sample is a fracture sample. =0 indicates that the nth sample is a non-fracture sample). This represents the predicted probability that the nth sample in the network output is a fracture sample.

[0031] Fracture localization loss: includes 3D bounding box regression loss and confidence loss. in, This represents the fracture bounding box regression loss (3D CIoU loss), used to optimize the spatial matching between the predicted bounding box and the ground truth bounding box; This represents the prediction box confidence loss (binary cross-entropy loss with sample weights), used to optimize the probability accuracy of the prediction box being a fracture region.

[0032] 2.1 Box Regression Loss : 3D CIoU loss is used to measure the box matching degree. in, This represents the intersection-union ratio (IoU) between the predicted bounding boxes output by the network and the ground truth bounding boxes annotated by experts. , To predict the box volume, (the actual bounding box volume) Indicates the center of the prediction box Center of the real frame The squared Euclidean distance is given by , where c represents the length of the spatial diagonal of the smallest cube enclosing both the predicted and ground truth boxes. =0.5 represents the balance coefficient, and v represents the consistency index between the predicted box and the ground truth box in terms of size. Where w, h, and d represent the width, height, and slice depth dimensions of the box, respectively.

[0033] 2.2 Confidence Loss Weighted binary cross-entropy loss to distinguish between positive and negative bounding boxes: Where M represents the total number of scale prediction boxes in the training batch. =5.0、 =1.0 represents the sample weights of "positive sample boxes" (predicted boxes with IoU > 0.5 with the ground truth boxes) and "negative sample boxes" (predicted boxes with IoU < 0.1 with all ground truth boxes), respectively. This represents the true confidence label of the m-th prediction box ( =1 indicates that the m-th predicted box is a positive sample box. =0 indicates that the m-th predicted box is a negative sample box. This represents the confidence probability that the m-th predicted bounding box in the network output is a fracture region.

[0034] Minimize using the AdamW optimization algorithm A joint loss function was used to optimize all network parameters of the dual-stream feature extraction encoder, the Mamba fusion module, and the dual-task head (classification head + localization head). The training batch size was set to 4, and the initial learning rate was set to 1e-4. A cosine annealing learning rate scheduling strategy was adopted (50 training cycles, with the first 20 cycles serving as a learning rate warm-up phase, linearly increasing to 1e-4, and the following 80 cycles decaying to 1e-6 according to a cosine curve), with a total of 100 training cycles. After each training cycle, the fracture classification accuracy and localization mAP index were calculated on the validation set. An early stopping strategy was adopted (training was stopped if there was no improvement in validation set performance for 10 consecutive cycles). The model weights with the best performance on the validation set were saved, and finally, an end-to-end automatic prediction and CT localization model for osteoporotic fractures was established.

[0035] Step 4: Model Reasoning The preprocessed CT and MRI multimodal images to be detected are input into the trained three-dimensional bimodal fracture localization model to automatically predict osteoporotic fractures and locate the fracture area on CT images. The specific process is as follows: Multimodal image preprocessing: The CT and MRI images to be tested are standardized. The CT images are adjusted using bone windowing, denoised by Gaussian filtering (σ=1.0), and then normalized to the [0,1] interval. The CT and MRI images are uniformly cropped to a fixed size (64×256×256), and the preprocessed tensors are output. (Single-channel CT) and (MRI single channel).

[0036] Feature extraction and cross-modal fusion: and The inputs are a dual-stream feature extraction encoder, and the outputs are multi-scale feature volumes. (CT side) and (MRI side); Features at various scales are fused in stages using FMB: first, channel switching and VSS block processing are completed by the SSCS module to generate shallow fused features. , Then, the DSSF module is used to implement gated dual-attention interaction in the hidden state space, and finally, the enhanced fusion feature is output. , , .

[0037] Fracture prediction (binary classification): based on top-level fusion features The global feature vector is obtained by global average pooling. The classification inference module uses a fully connected layer and Softmax activation to output the fracture probability of the sample to be detected. Judgment rule: If If the value is >0.5, it is determined to be a fracture sample and proceeds to the next localization step; otherwise, it is determined to be a non-fracture sample and the result "No fracture detected" is output.

[0038] CT fracture region localization: For samples identified as fractures, based on multi-scale fusion features... , , Perform fracture bounding box prediction; the localization inference module outputs predicted bounding boxes (including coordinates) at various scales. Confidence level Non-maximum suppression was used to remove overlapping boxes, retaining... A prediction bounding box with a value >0.5 is output as the detection bounding box for the fracture area on the CT image.

[0039] Results Output and Quantitative Indicators: The final fracture prediction result (fracture probability) will be output. ) and CT localization results (detection frame coordinates and confidence level) This is linked to the DICOM coordinate system of the original CT image to generate a visualization report (example results are shown below). Figure 1 (As shown). The quantification metrics for the 3D detection bounding box include: Spatial coordinates: Based on the voxel coordinates of the CT image, output the center point of the fracture area. The three-dimensional dimensions (w, h, d) are then converted to physical coordinates (unit: mm). The calculation formula is as follows: in, , , In voxel coordinates, , , These represent the physical resolution of CT images along the x, y, and z axes (unit: mm / voxel).

[0040] Locational reliability: Output The value (0-1) reflects the reliability of the model in locating the fracture area.

[0041] Ultimately, based on intelligent reasoning through the multimodal fusion of CT and MRI, efficient prediction and precise localization of osteoporotic fractures are achieved, assisting clinicians in quickly identifying and diagnosing lesions.

[0042] Compared to existing single-modal fracture detection methods, this invention offers significant comprehensive technical advantages. While traditional CT images can clearly present bone structure and cortical bone morphology, in patients with osteoporosis, due to decreased bone density, fine fracture lines, or soft tissue overlap obscuring the bone, subtle fractures are frequently missed or misdiagnosed. Especially in the early or occult fracture stages, single-CT images lack sufficient sensitivity to changes in bone marrow and soft tissues, limiting diagnostic accuracy. To address these issues, this invention introduces MRI modal information and deeply fuses it with CT images. By setting up a Mamba fusion module, dynamic interaction between CT structural features and MRI soft tissue contrast features is achieved at the model feature level. MRI images can provide indirect soft tissue signs such as bone marrow edema, ligament damage, and local effusion around the fracture, while CT images reflect the continuity and density distribution characteristics of the bone cortex. After linear fusion by the Mamba module, the two form complementary feature representations, improving the model's ability to identify subtle fracture lines and early bone changes. Experimental results show that the proposed multimodal fusion strategy can effectively reduce the risk of misdiagnosis and missed diagnosis caused by the blind zone of a single mode, and significantly improve the sensitivity and stability of the model.

[0043] To achieve automatic fracture identification and localization, this invention designs a dual-task reasoning module. This module consists of a fracture classification reasoning module and a fracture localization reasoning module: the classification reasoning module uses global semantic features to automatically distinguish between "fractured" and "non-fractured" fractures; the localization reasoning module relies on multi-scale spatial features to output a three-dimensional fracture detection box, achieving precise localization of specific vertebrae and fracture areas. Both modules are collaboratively optimized through a joint loss function, allowing the classification and localization tasks to share the underlying semantic representation. This avoids feature bias caused by single-task optimization and achieves an integrated analysis process from "qualitative judgment" to "quantitative localization." Compared to existing schemes that require multiple models to be serially connected for classification and localization, the dual-task design of this invention significantly reduces data transfer and error accumulation, improving reasoning efficiency and diagnostic consistency.

[0044] In terms of model computation efficiency, traditional multimodal fusion methods based on the Transformer structure typically have a time complexity of O(N). 2 (where N is the number of three-dimensional voxels) Processing three-dimensional medical images involves high computational cost and inference latency, and requires high-memory GPUs, hindering real-time clinical deployment. To address this issue, the Mamba fusion module proposed in this invention employs a fusion strategy with linear complexity O(N) to achieve efficient three-dimensional feature interaction. This structure significantly reduces computational costs while maintaining fusion effectiveness, shortening single-instance inference time to a clinically acceptable range, and facilitating deployment in environments with ordinary GPUs and low-to-medium configuration servers. In summary, this invention, through the comprehensive design of "CT-MRI multimodal fusion + dual-task inference architecture + linear complexity fusion mechanism," achieves the following beneficial technical effects: Fusion perception enhancement: Combining the structural details of CT with the tissue contrast information of MRI improves the identifiability of subtle fracture lines and early bone changes; Integrated analysis process: Classification and localization tasks are executed in a coordinated manner, avoiding the process complexity and error accumulation caused by multiple models being linked together; Computational efficiency optimization: The linear complexity Mamba fusion module significantly reduces computational resource consumption and improves inference speed and system stability; High clinical applicability: The model has real-time inference performance and hardware adaptability, making it suitable for primary hospitals and routine radiology environments, and has high practical application value.

[0045] Therefore, this invention not only achieves efficient fusion of multimodal information and dual-task collaborative optimization at the algorithm level, but also takes into account both computational efficiency and clinical usability at the system deployment level, effectively supporting the intelligent and rapid diagnostic process for osteoporotic fractures.

[0046] A computer-readable storage medium having a computer program / instructions thereon, which, when executed by a processor, implements the steps of the described method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion.

[0047] A computer program product includes a computer program / instructions that, when executed by a processor, implement the steps of a method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion.

[0048] The technical principles of the present invention have been described above with reference to specific embodiments. These descriptions are merely for explaining the principles of the invention and should not be construed as limiting the scope of protection of the invention in any way. Based on this explanation, those skilled in the art can readily conceive of other specific embodiments of the invention without inventive effort, and these embodiments will all fall within the scope of protection of the claims of the present invention.

Claims

1. A method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion, characterized in that, Includes the following steps: Step 1: Obtain and filter CT and MRI images containing osteoporotic fractures and osteoporotic but unfractured fractures in the thoracic to lumbar spine region, and mark the fracture areas; Step 2: Construct a three-dimensional dual-modal fracture localization model, including: A dual-stream feature extraction encoder is used to extract multi-scale three-dimensional features from input CT and MRI images, respectively. The Mamba fusion module is used to fuse the CT and MRI multi-scale features extracted by the dual-stream feature extraction encoder in stages and output multi-scale three-dimensional fusion features. The dual-task reasoning module, based on the multi-scale three-dimensional fusion features, performs fracture classification and fracture region localization tasks in parallel. Step 3: Using the dataset constructed in Step 1, the model constructed in Step 2 is trained end-to-end by optimizing the joint loss function, which includes classification loss and localization loss, to obtain a trained osteoporotic fracture prediction and localization model. Step 4: Predict the fracture probability of the CT and MRI images to be tested, output the predicted fracture probability, and mark it on the images.

2. The method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion according to claim 1, characterized in that, In step 1, cases of fractures caused by high energy or violent injuries are excluded during the screening process, as well as samples where the time interval between CT and MRI scans exceeds a predetermined threshold or the fracture sites are inconsistent. Confirmed osteoporotic fracture samples were retained as positive samples, and osteoporotic non-fracture samples with similar distribution in age, sex and examination equipment were selected as control samples. When annotating the fracture area, a professional physician performs three-dimensional annotation of the fracture area on the CT image of the positive sample to generate a ROI mask.

3. The method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion according to claim 1, characterized in that, The Mamba fusion module in step 2 includes a state space channel exchange module and a dual state space fusion module. The state space channel exchange module is used to exchange and reorganize the channel dimensions of the input CT and MRI features, and then process them through the visual state space block to achieve shallow feature fusion. The dual-state space fusion module is used to project shallow fusion features to the hidden state space, perform feature interaction through a gated dual attention mechanism, and then project them back to the original feature space and add them to the initial features for fusion, thereby achieving deep feature fusion.

4. The method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion according to claim 1, characterized in that, The dual-task reasoning module in step 2 includes: The fracture classification reasoning module performs three-dimensional global average pooling on the fused global features, and then processes them through a fully connected layer to output the predicted probability that the sample is a fracture. The fracture localization inference module, based on multi-scale fusion features, predicts the 3D bounding box coordinates and confidence of the fracture area through a 3D detection head, and outputs the final localization result after non-maximum suppression post-processing.

5. The method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion according to claim 1, characterized in that, The joint loss function in step 3 is: ; In the formula, This indicates the total loss resulting from the combined efforts of both tasks. This represents the fracture classification loss, including the binary cross-entropy loss with class weights; This indicates the loss in CT fracture localization, including 3D bounding box regression loss and confidence loss; =1.0、 =5.0 are the weight coefficients for the two types of losses, used to balance the training priority of classification and localization tasks.

6. The method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion according to claim 5, characterized in that, The class-weighted binary cross-entropy loss: ; In the formula, N represents the total number of samples in the training batch; and These represent the category weights for "fractured samples" and "non-fractured samples," respectively. =0.7, =0.3; This represents the true classification label of the nth sample. =1 indicates that the nth sample is a fracture sample. =0 indicates that the nth sample is a non-fracture sample; This represents the predicted probability that the nth sample in the network output is a fracture sample.

7. The method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion according to claim 5, characterized in that, The 3D bounding box regression loss and confidence loss are as follows: ; In the formula, This represents the fracture bounding box regression loss, used to optimize the spatial matching between the predicted bounding box and the ground truth bounding box. This represents the prediction box confidence loss, used to optimize the accuracy of the prediction box as a fracture area.

8. The method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion according to claim 1, characterized in that, When outputting the fracture probability prediction result, the fracture probability output by the model and the voxel coordinates of the three-dimensional positioning box are combined with the physical resolution of the CT image to convert them into physical coordinates, and a visualization report containing quantitative indicators is generated.

9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion as described in claim 1.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the method for predicting and locating osteoporotic fractures based on CT and MRI multimodal fusion as described in claim 1.

Citation Information

Patent Citations

  • Osteoporotic Vertebral Refracture Prediction System Based on Deep Learning of CT Images

    CN115644904B

  • Fracture risk prediction method based on features and CT images

    CN117095817A

  • A CT-based artificial intelligence method for predicting spinal refracture in patients with OVCF

    CN118692613B