Multi-feature fusion diagnosis system and method for L1-L4 lumbar vertebra segments
Through a multi-feature fusion diagnostic system, utilizing a cross-modal attention mechanism and trabecular bone network diagram, the problem of traditional methods being unable to accurately assess the bone status of the L1-L4 lumbar vertebrae segments is solved, achieving high sensitivity and high specificity in the diagnosis of early osteoporosis and providing refined fracture risk assessment.
Patent Information
- Application Number
- CN202511111657.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies make it difficult to accurately diagnose osteoporosis early, especially to effectively assess the bone status and fracture risk of the L1-L4 lumbar vertebrae. Traditional methods ignore the heterogeneity of different regions within the vertebral body and the topological structure of the trabecular network.
A multi-feature fusion diagnostic system is used to learn imaging and clinical data through a cross-modal attention mechanism, construct a trabecular network graph, and combine imaging features and clinical information to generate refined diagnostic results, including vertebral level, sub-region level and patient level assessments.
It significantly improves the sensitivity and specificity of early diagnosis of osteoporosis, can identify local bone loss, provide accurate fracture risk assessment, and support robust diagnosis in a variety of situations with incomplete data.
Smart Images

Figure CN120748692A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence-assisted diagnosis technology, and in particular to a multi-feature fusion diagnosis system and method for the L1-L4 lumbar vertebrae segment. Background Art
[0002] Osteoporosis is a systemic skeletal disease characterized by decreased bone mass and destruction of bone microstructure. The resulting brittle fractures have become one of the leading causes of disability and mortality in the elderly. According to statistics, the prevalence of osteoporosis in people over 60 years old in my country is as high as 36%. The lumbar spine is the most susceptible site for osteoporotic fractures. The L1-L4 segments bear the main axial loads of the human body, and their bone quality directly affects the patient's quality of life. However, osteoporosis often has no obvious clinical symptoms in the early stages, and the condition is already quite serious by the time pain or fractures occur. Therefore, early and accurate diagnosis is of great significance for preventing fractures and improving prognosis.
[0003] The current gold standard for clinical assessment of osteoporosis is bone density measured by dual-energy X-ray absorptiometry (DXA). However, this method only provides two-dimensional projections of areal bone density and cannot reflect the three-dimensional microstructure and spatial distribution characteristics of trabecular bone, resulting in insufficient sensitivity for early bone loss. Although quantitative CT (QCT) can provide three-dimensional volumetric bone density information, traditional QCT analysis typically only calculates the average density of the vertebral body as a whole or at the central level, ignoring the heterogeneity of bone distribution in different anatomical regions within the vertebral body. Biomechanical studies have shown that different regions of the vertebral body, such as the upper and lower endplates, anterior, middle and posterior columns, and cortical bone, have significant differences in load-bearing stress and fracture risk. A single overall density measurement is difficult to accurately assess local bone status and fracture risk.
[0004] In recent years, the development of high-resolution CT technology has made it possible to noninvasively assess trabecular bone microstructure. As the basic structural unit of cancellous bone, the number, thickness, spacing, and connectivity of trabeculae directly determine the mechanical strength of bone. However, existing trabecular bone analysis methods are mostly based on simple morphological parameter statistics and lack in-depth analysis of the topological structure and directionality of the trabecular network. The spatial arrangement of trabeculae follows Wolff's law and is distributed along the principal stress directions, forming a complex three-dimensional network structure. This structural feature contributes to bone strength even more than bone density itself, but traditional methods have difficulty effectively quantifying this network property. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies, this paper proposes a multi-feature fusion diagnostic system and method for the L1-L4 lumbar spine. This system automatically learns the correlation weights between different data sources through a cross-modal attention mechanism, avoiding the subjectivity of manual judgment. The system can discover implicit correlations between imaging findings and biochemical markers.
[0006] To achieve the above objectives, the present invention proposes a multi-feature fusion intelligent diagnosis system for the L1-L4 lumbar vertebrae segment, comprising:
[0007] An image preprocessing module is used to receive a patient's lumbar spine CT image sequence and output standardized three-dimensional voxel data;
[0008] a vertebral anatomical partitioning module, connected to the image preprocessing module, for dividing each vertebra into eight biomechanical stress-related subregions based on an anatomical template, each subregion being defined as Ri (i=1, 2, ..., 8), where R1-R2 are the upper and lower endplate regions, R3-R5 are the anterior, middle, and posterior column regions, R6-R7 are the left and right cortical regions, and R8 is the central cancellous bone region;
[0009] Perform a three-dimensional Hessian matrix analysis on each subregion Ri:
[0010]
[0011] Where H(x,y,z) is the three-dimensional Hessian matrix, I is the grayscale intensity function of the CT image, and x,y,z are the three-dimensional spatial coordinates;
[0012] By eigenvalue decomposition H·v i =λ i ·v i Extract eigenvalues λ1≥λ2≥λ3, where the eigenvector v3 corresponding to λ3 represents the main direction of trabecular bone;
[0013] Identify the local extreme point of the direction field in the 26-neighborhood as the graph node v i , when the angle between the direction vectors of adjacent nodes is less than 30°, an edge connection is established e ij , where e ij ∈E i Represents the connection node v i and v j The undirected edge of
[0014] Construct sub-region bone network graph G i =(V i ,E i ), where V i is the node set, E i is the edge set, node feature vector Contains local density HU i , gradient amplitude and anisotropy index
[0015] The multi-dimensional image feature extraction module is connected in parallel with the trabecular network map construction module to extract four types of quantitative features from each sub-region:
[0016] Density eigenvector F d : Contains the first to fourth order statistical moments of HU values;
[0017] Morphological feature vector F m : Mean and standard deviation of cortical bone thickness calculated by distance transform, and trabecular separation based on connected domain analysis;
[0018] Texture feature vector F t : 3D gray-level co-occurrence matrix features in 13 directions and rotation-invariant 3D-LBP codes;
[0019] Topological eigenvector F s : Figure G i The average degree, clustering coefficient and shortest path length of
[0020] The clinical multimodal data encoding module independently acquires and processes three types of clinical data:
[0021] Structured data encoder: normalizes numerical indicators such as age, gender, and BMI to the range [0,1];
[0022] Biochemical index coder: PTH, 25(OH)D, CTX and other bone metabolism markers were logarithmically transformed and z-score normalized;
[0023] Medical record text encoder: Use the pre-trained BioBERT model to convert free text into 768-dimensional semantic vectors;
[0024] The multimodal graph attention fusion network module receives the outputs of the trabecular network graph construction module, the multidimensional image feature extraction module, and the clinical multimodal data encoding module, including:
[0025] Construct a three-layer heterogeneous graph G'=(V',E'), where V' contains three types of nodes: imaging nodes, clinical nodes, and semantic nodes;
[0026] Design a cross-modal attention mechanism, and the attention weight calculation formula is:
[0027]
[0028] where α ij is the attention weight from node j to node i, h i ∈R d is the hidden layer representation vector of node i, W q ,W k ∈R d ×d is the learnable query and key projection matrix, d = 256 is the hidden layer dimension, and N(i) is the set of neighbor nodes of node i;
[0029] Update node representation via message passing mechanism:
[0030]
[0031] in is the representation of node i in the l-th layer network, W v ∈R d ×d is the learnable value projection matrix, σ is the ReLU activation function;
[0032] The segment-level diagnosis output module is connected to the multimodal graph attention fusion network and outputs three levels of diagnosis results:
[0033] Vertebral level: T-score prediction for each vertebra, osteoporosis grade (normal / osteopenia / osteoporosis), and compression fracture risk score;
[0034] Subregion level: 8×4 risk matrix, indicating the degree of abnormality in 8 subregions of each vertebra from L1 to L4;
[0035] Patient level: comprehensive FRAX score and prediction of future fracture probability.
[0036] Furthermore, in the trabecular network graph construction module, the specific method for constructing edge connections is:
[0037] Calculate the structural similarity between nodes vi and vj:
[0038]
[0039] Where S(v i ,v j ) is the node v i and v j The structural similarity between them is in the range of [0,1]; f i is the feature vector of node i; σ = 0.5 is the scale parameter that controls the similarity decay rate; θ ij is the angle between the main directions of the trabeculae at nodes i and j;
[0040] When S(v i ,v j )>0.7 and the Euclidean distance between nodes is less than 3 voxels, the weight is S(v i ,v j )’s edge connection;
[0041] A spectral clustering algorithm was used to aggregate highly connected nodes into trabecular bone clusters, and the connections between clusters represented the key force transmission paths of the trabecular bone network.
[0042] Furthermore, the multimodal graph attention fusion network adopts a staged training strategy:
[0043] Phase 1: Freeze the clinical data branch and train only with image features. The loss function in is the predicted output based on the image, y is the true label, ||W||2 is the L2 norm of the network weight matrix, and α=0.01 is the regularization coefficient;
[0044] Phase 2: Unfreeze all parameters and introduce multimodal contrastive learning loss:
[0045]
[0046] Among them (z img ∈R 128 is the normalized embedding vector of the imaging modality, z clin ∈R 128 is the normalized embedding vector of the clinical modality, is the cosine similarity, τ = 0.07 is the temperature parameter, and B is the set of all positive and negative sample pairs in the batch;
[0047] Phase 3: Adding temporal consistency constraints: L temporal =||f(x t )-f(x t-Δt )-g(Δt)|| 2
[0048] Where f(·) is the feature extraction network, x t is the CT image at time t, Δt is the time interval between two scans (months),
[0049] g(Δt)=W t ·Δt+b t is a learnable time evolution function, W t ∈R d ×1 is the time decay weight, b t ∈R d is the time bias term;
[0050] The total loss function L = L1 + β·Lcontrast + γ·Ltemporal, where β = 0.1 and γ = 0.05.
[0051] Furthermore, it also includes a confidence calibration module connected to the diagnosis output module:
[0052] Calculate prediction uncertainty: Perform T = 10 forward passes on the same input (with dropout enabled) and calculate the prediction variance:
[0053]
[0054] where p tis the predicted probability of the t-th forward propagation, and p is the mean of T predictions;
[0055] Data integrity score: S complete =0.4·I img +0.3·I clin +0.3·I text , where I img ,I clin ,I text ∈[0,1] is the integrity index of each modal data;
[0056] Calibrated confidence:
[0057]
[0058] Among them C calibrated is the confidence level after calibration, ranging from [0,1], C raw is the confidence of the original output of the model.
[0059] Furthermore, the specific partitioning method used by the vertebral anatomy partitioning module is:
[0060] Construct L1-L4 standard template based on Statistical Shape Model, containing 1024 corresponding points;
[0061] The template is registered to the patient CT image using the CoherentPointDrift algorithm;
[0062] Based on the registered corresponding points, the Voronoi diagram partitioning method is used to automatically generate eight sub-regions with clear anatomical significance;
[0063] Boundary smoothing uses 3D morphological opening operations to eliminate jagged boundaries.
[0064] 6. The system according to claim 1, wherein the multi-dimensional image feature extraction module further comprises:
[0065] Adaptive feature selection unit, which uses LASSO regularization to automatically select the feature subset most relevant to the diagnostic task. The regularization parameter is determined by 5-fold cross-validation.
[0066] The feature normalization unit uses domain adaptation technology to eliminate device-related bias for different CT devices and scanning protocols, and adopts the maximum mean difference (MMD) as the inter-domain distance metric.
[0067] 7. The system according to claim 1, wherein the segment-level diagnostic output module further comprises:
[0068] The natural language generation unit, based on the pre-trained GPT-2 model, converts structured diagnostic results into natural language reports that conform to clinical habits;
[0069] Interactive 3D visualization unit, using Web graphics library technology to achieve real-time rendering of bone distribution inside the vertebral body, supporting arbitrary angle rotation and cross-section viewing;
[0070] The personalized risk assessment unit combines patients' lifestyle factors and uses the Cox proportional hazards model to predict 10-year fracture risk.
[0071] A method for diagnosing lumbar vertebrae bone health using the system of claim 1, comprising the following steps:
[0072] S1: Acquire thin-slice CT scan data of the patient's lumbar spine, with a slice thickness of ≤1 mm, and obtain a DICOM format image sequence;
[0073] S2: Image preprocessing module performs:
[0074] S2.1: Apply a deep learning-based metal artifact removal algorithm;
[0075] S2.2: Use bilateral filter to reduce noise and preserve edge information;
[0076] S2.3: Intensity window normalized to [-1000, 2000] HU;
[0077] S3: The vertebral anatomy partitioning module automatically locates the L1-L4 vertebral centers and completes the 8-region division;
[0078] S4: Parallel execution of feature extraction:
[0079] S4.1: The trabecular bone network map construction module generates 32 sub-region maps (4 vertebrae × 8 regions) through Hessian matrix analysis;
[0080] S4.2: Multi-dimensional image feature extraction module calculates the 43-dimensional feature vector of each sub-region;
[0081] S5: Clinical multimodal data encoding module processes patient clinical information and generates standardized feature representations;
[0082] S6: The multimodal graph attention fusion network integrates the outputs of S4 and S5 and performs 4-layer graph neural network processing through the cross-modal attention mechanism;
[0083] S7: Segment-level diagnostic output module generates hierarchical diagnostic results and visualization reports;
[0084] S8: The confidence calibration module evaluates diagnostic reliability based on prediction uncertainty and data integrity, and prompts manual review when the calibration confidence is <0.7.
[0085] Furthermore, in step S6, the specific process of the multimodal fusion is as follows:
[0086] S6.1: Combine 32 sub-region graph nodes, 15 clinical feature nodes, and 1 semantic node into a 48-node heterogeneous graph;
[0087] S6.2: Initialize the cross-modal edge weight matrix W∈R48×48 using Xavier initialization.
[0088] S6.3: Perform 4 rounds of message passing and feature updates. In each round, the attention mechanism is used to calculate the weight of information propagation between nodes. LayerNormalization is applied after the update.
[0089] S6.4: Use global average pooling to aggregate node features and output the diagnosis prediction vector through the fully connected layer.
[0090] Furthermore, it also includes:
[0091] S9: When historical CT data of patients are available, perform temporal comparative analysis to evaluate the rate and trend of bone degradation using temporal consistency constraints;
[0092] S10: Generate personalized treatment recommendations and follow-up plans based on diagnostic results and patient characteristics in combination with clinical guidelines.
[0093] Compared with the prior art, the present invention has the following beneficial effects:
[0094] 1. This invention provides a multi-feature fusion diagnostic system and method for the L1-L4 lumbar spine. By meticulously dividing each vertebra into eight biomechanical stress-related subregions and independently analyzing the characteristics of each subregion, this system can capture the spatial heterogeneity of bone distribution within the vertebral body, overcoming the limitations of traditional methods that only measure average density at the overall or central level. This refined zoning strategy enables the system to identify early localized bone loss, particularly subtle changes in key weight-bearing areas such as the anterior column and upper and lower endplates, significantly improving the sensitivity and specificity of early diagnosis of osteoporosis.
[0095] 2. The present invention provides a multi-feature fusion diagnostic system and method for the L1-L4 lumbar vertebral segments. The trabecular structure is modeled as a three-dimensional network graph, the main directions of the trabeculae are extracted through Hessian matrix analysis, and a graph structure reflecting the actual spatial connection relationship of the trabeculae is constructed. This method makes full use of the network topological characteristics of the trabeculae and can more accurately evaluate the mechanical strength of the bone than traditional morphological parameters.
[0096] 3. The present invention provides a multi-feature fusion diagnosis system and method for the L1-L4 lumbar vertebrae segment, which automatically learns the association weights between different data sources through a cross-modal attention mechanism, avoiding the subjectivity of manual experience judgment.
[0097] 4. This paper provides a multi-feature fusion diagnostic system and method for the L1-L4 lumbar spine. This system supports robust diagnosis in a variety of data-incomplete scenarios, providing reliable results based on available data even when some biochemical markers or medical history information are missing. A phased training strategy and domain adaptation technology enable the system to adapt to differences in equipment and scanning protocols across hospitals, demonstrating good generalization capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0099] Figure 1 This is a schematic diagram of the system flow of the present invention DETAILED DESCRIPTION
[0100] The technical solutions of the present invention will be more clearly and completely explained below through description of preferred embodiments of the present invention in conjunction with the accompanying drawings.
[0101] like Figure 1 As shown, the present invention provides a multi-feature fusion intelligent diagnostic system for the L1-L4 lumbar spine. This system integrates medical imaging data and multimodal clinical information to comprehensively assess the health of lumbar vertebrae. The system first receives the patient's lumbar CT image sequence through an image preprocessing module, standardizes the raw DICOM format data, and outputs high-quality 3D voxel data, laying the foundation for subsequent analysis.
[0102] The preprocessed image data then enters the vertebral anatomical zoning module, which automatically identifies the positions of the L1-L4 vertebrae based on the anatomical template and accurately divides each vertebra into eight biomechanical stress-related sub-regions, including the upper and lower endplate regions, the anterior, middle and posterior column regions, the left and right cortical regions, and the central cancellous bone region. This refined zoning strategy fully considers the differences in biomechanical properties of different parts of the vertebrae.
[0103] After completing the anatomical partitioning, the system uses a parallel processing architecture to simultaneously perform two key feature extraction tasks. The trabecular network graph construction module performs a three-dimensional Hessian matrix analysis on each sub-region, extracts the trabecular direction field information, identifies local extreme points as graph nodes, and establishes edge connections based on the angles between the direction vectors, ultimately constructing a network graph G = (V, E) reflecting the trabecular microstructure. Simultaneously, the multidimensional imaging feature extraction module comprehensively quantifies the imaging features of each sub-region from four dimensions: density, morphology, texture, and topology. This includes 43 quantitative indicators such as HU value statistical moments, cortical bone thickness, 3D grayscale co-occurrence matrix features, and graph topological properties.
[0104] While processing imaging data, the clinical multimodal data encoding module independently acquires and processes patients' clinical information. This module includes three specialized encoders: a structured data encoder processes numerical indicators such as age, gender, and BMI; a biochemical marker encoder standardizes bone metabolism markers such as PTH, 25(OH)D, and CTX; and a medical record text encoder utilizes the pre-trained BioBERT model to convert free-text medical records into semantic vector representations.
[0105] All extracted imaging features and encoded clinical data are integrated into a multimodal graph attention fusion network. This network constructs a three-layer heterogeneous graph consisting of imaging nodes, clinical nodes, and semantic nodes. It uses a cross-modal attention mechanism to calculate the association weights between different modalities and utilizes a message passing mechanism to propagate and update node representations within the graph structure, achieving a deep fusion of imaging features and clinical information.
[0106] The comprehensive features after multimodal fusion processing are sent to the segment-level diagnostic output module, which generates three levels of diagnostic results: the vertebral level output includes the T-score prediction, osteoporosis grade and compression fracture risk score of each vertebra; the sub-region level provides an 8×4 risk matrix, which details the degree of abnormality in the eight sub-regions of each vertebra from L1 to L4; and the patient level comprehensively evaluates the FRAX score and predicts the probability of future fractures.
[0107] To ensure the reliability of diagnostic results, the system is equipped with a confidence calibration module. This module calculates prediction uncertainty and assesses data integrity to provide a calibrated confidence score for each diagnostic result. When the confidence score falls below a set threshold, the system prompts the clinician for manual review, thereby ensuring diagnostic accuracy while avoiding the risk of misdiagnosis.
[0108] The data flow of the entire system is clear and unambiguous: imaging data starts with CT image input, undergoes preprocessing, anatomical partitioning, feature extraction and other steps, and is finally integrated with clinical data in a fusion network to generate multi-level diagnostic results and confidence assessments, providing clinicians with comprehensive, accurate and interpretable lumbar spine bone health assessment reports.
[0109] As a specific example, thin-slice CT scan data of the lumbar spine of a 65-year-old female patient were acquired using the following scanning parameters: tube voltage 120 kV, tube current 280 mA, slice thickness 0.625 mm, and reconstruction matrix 512 × 512. This yielded a 320-slice DICOM-formatted image sequence covering the L1-L4 segments. The image preprocessing module first applied a deep learning-based metal artifact removal algorithm to address potential metal implant artifacts. A bilateral filter was then used for noise reduction, with a spatial domain standard deviation of 5 and an intensity domain standard deviation of 30. This effectively suppressed image noise while preserving trabecular bone edge information. Finally, the CT value window was normalized to the [-1000, 2000] HU range to generate standardized 3D voxel data.
[0110] The vertebral anatomical partitioning module uses the Coherent Point Drift algorithm for registration based on the Statistical Shape Model standard template containing 1024 anatomical correspondence points. The algorithm uses a Gaussian kernel width of 2, a regularization weight of 3, and a maximum number of 150 iterations. After registration, eight anatomically defined subregions are automatically generated based on the corresponding points using the Voronoi diagram method. Three-dimensional morphological opening is performed using a 3×3×3 structuring element to eliminate jagged irregularities at the partition boundaries. In this example, the L1 vertebra was successfully partitioned into eight subregions, each containing an average of approximately 15,000 voxels.
[0111] The trabecular bone network graph construction module performs a three-dimensional Hessian matrix analysis on each subregion, with a Gaussian kernel standard deviation set to 1.5 voxels. Within a 26-neighborhood, points with a Frobenius norm greater than 0.3 and a local maximum are marked as graph nodes. In this example, 287 nodes were identified in the anterior column of the L1 vertebra, 342 nodes in the middle column, and 256 nodes in the posterior column, reflecting the differences in trabecular bone density across these regions. The threshold for establishing edge connections between nodes was set to 0.7, and the Euclidean distance was restricted to 3 voxels. Finally, a spectral clustering algorithm was used to cluster the nodes into trabecular bone clusters, resulting in 23 clusters in the L1 vertebra, with an average of 4.7 inter-cluster connections.
[0112] The multidimensional image feature extraction module extracts 43-dimensional quantitative features from each subregion. In this example, the density features of the anterior column of the L1 vertebra showed a mean bone density (HU) value of 125.3, a standard deviation of 42.7, a skewness of -0.8, and a kurtosis of 3.2, indicating low and uneven bone density. Morphological analysis revealed that the mean cortical thickness in this region was only 1.8 mm, 2.5 mm below the normal value, and the trabecular separation reached 1.2 mm, 0.6 mm above the normal value. Texture features, calculated from 13 directions, showed a 3D gray-level co-occurrence matrix with an energy value of 0.023, a contrast of 156.8, a correlation of 0.42, and a uniformity of 0.087, indicating disordered trabecular structure. Topological analysis revealed that the average degree of the bone network in this region was 3.1, the clustering coefficient was 0.52, and the average shortest path length was 5.2, all of which deviated from the normal range.
[0113] The clinical multimodal data encoding module processes the patient's clinical information. The patient was 65 years old, female, with a BMI of 22.5 kg / m², no history of fractures, 15 years of menopause, no history of glucocorticoid use, and a history of mild rheumatoid arthritis. Biochemical tests showed a serum PTH level of 45 pg / mL, within the normal range; a 25(OH)D level of 18 ng / mL, below the normal value of 30 ng / mL; a bone resorption marker CTX of 0.58 ng / mL, above the upper limit of 0.45 ng / mL for postmenopausal women; and a bone formation marker P1NP of 35 ng / mL, within the normal range. The medical record documented the patient's chief complaint of "lower back pain for three months, limited bending, and worsening symptoms in the morning." The semantic vector extracted by the BioBERT model successfully identified key clinical features such as pain and limited mobility.
[0114] The multimodal graph attention fusion network constructs a heterogeneous graph consisting of 48 nodes, including 32 imaging nodes corresponding to eight subregions of four vertebrae, 15 clinical feature nodes including age, gender, BMI, and bone metabolism markers, and one semantic node representing medical record text. The network uses a four-layer graph neural network architecture with a hidden layer dimension of 256. In this example, the cross-modal attention mechanism learned an attention weight of 0.78 between the L1 vertebral anterior column imaging node and the CTX biochemical indicator node, and a weight of 0.65 with the 25(OH)D node, indicating that the model successfully captures the association between bone destruction, vitamin D deficiency, and enhanced bone resorption.
[0115] The training process employed a phased strategy, using a dataset of 3,200 patients. In the first phase, training was performed for 50 epochs using only imaging features, with an initial learning rate of 1e-4 and a batch size of 32. In the second phase, all parameters were unfrozen and trained for 30 epochs, with the learning rate reduced to 5e-5. In the third phase, temporal consistency constraints were added for patients with historical CT scan data. In this case, the CT scan data was from one year ago. The time-decay weights learned by the model indicated an annual bone density decline of approximately 3.2%. The entire training process took approximately 18 hours on a server equipped with an NVIDIA A100 GPU.
[0116] Results generated by the segment-level diagnostic output module showed that the L1 vertebral T-score was -2.8, indicating osteoporosis, with a compression fracture risk score of 78 / 100. The L2 vertebral T-score was -2.3, indicating osteopenia, with a risk score of 65 / 100. The L3 vertebral T-score was -1.8, and the L4 vertebral T-score was -1.5, both near the lower limit of normal. An 8×4 subregion-level risk matrix clearly showed that the anterior column and superior endplate regions of the L1 vertebral body had the highest risk, with values of 0.85 and 0.82, respectively, indicating that these two regions are most susceptible to compression fractures. Patient-level FRAX scores revealed a 10-year probability of major osteoporotic fracture of 18.5% and a hip fracture of 7.2%, both exceeding the average for the patient's age group.
[0117] The confidence calibration module performed 10 forward propagations using the Monte Carlo Dropout method, with a dropout rate set to 0.1. The predicted variance for the diagnosis of L1 vertebral osteoporosis was 0.023, indicating a high degree of confidence in the model's diagnosis. The data integrity score was 0.91, with imaging data integrity rated at 1.0, biochemical parameters rated at 0.9 due to the lack of osteocalcin, and medical record text completeness rated at 0.8. The final calibration confidence score was 0.83, exceeding the threshold of 0.7, indicating a reliable diagnosis.
[0118] The system automatically generated a diagnostic report: "A 65-year-old female patient with severe osteoporosis at the L1 vertebra, a T-score of -2.8. Radiographic findings include sparse trabeculae and thinning of the bone cortex in the anterior column and superior endplate, indicating a high risk of compression fracture. Bone mass was reduced at the L2 vertebra, and near the lower limit of normal at the L3-L4 vertebrae. Biochemical examinations revealed vitamin D deficiency and active bone resorption. A comprehensive assessment indicated an elevated 10-year fracture risk. Recommendations include: 1) Vitamin D supplementation of 800-1000 IU / day; 2) Consider initiating bisphosphonate anti-osteoporosis therapy; 3) Moderate weight-bearing exercise; and 4) Recheck bone metabolism markers in 3-6 months."
[0119] The entire diagnostic process, from data input to report generation, takes about 45 seconds, meeting the needs of clinical real-time diagnosis.
[0120] The above-described specific embodiments merely describe preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, substitutions, and improvements made to the technical solutions of the present invention by persons of ordinary skill in the art based on the textual description and accompanying drawings provided herein, without departing from the design concept and spirit of the present invention, shall fall within the scope of protection of the present invention. The scope of protection of the present invention is determined by the claims.
Claims
1. A multi-feature fusion diagnostic system for the L1-L4 lumbar vertebrae segment, characterized by: include: An image preprocessing module is used to receive a patient's lumbar spine CT image sequence and output standardized three-dimensional voxel data; a vertebral anatomical partitioning module, connected to the image preprocessing module, for dividing each vertebra into eight biomechanical stress-related subregions based on an anatomical template, each subregion being defined as Ri, i = 1, 2, ..., 8, where R1-R2 are the upper and lower endplate regions, R3-R5 are the anterior, middle, and posterior column regions, R6-R7 are the left and right cortical regions, and R8 is the central cancellous bone region; Perform a three-dimensional Hessian matrix analysis on each subregion Ri: Where H(x,y,z) is the three-dimensional Hessian matrix, I is the grayscale intensity function of the CT image, and x,y,z are the three-dimensional spatial coordinates; By eigenvalue decomposition H·v i =λ i ·v i Extract eigenvalues λ1≥λ2≥λ3, where the eigenvector v3 corresponding to λ3 represents the main direction of trabecular bone; Identify the local extreme point of the direction field in the 26-neighborhood as the graph node v i , when the angle between the direction vectors of adjacent nodes is less than 30°, an edge connection is established e ij , where e ij ∈E i Represents the connection node v i and v j The undirected edge of Construct sub-region bone network graph G i =(V i , E i ), where V i is the node set, E i is the edge set, node feature vector Contains local density HU i , gradient amplitude and anisotropy index The multi-dimensional image feature extraction module is connected in parallel with the trabecular network map construction module to extract four types of quantitative features from each sub-region: Density eigenvector F d : Contains the first to fourth order statistical moments of HU values; Morphological feature vector F m : Mean and standard deviation of cortical bone thickness calculated by distance transform, and trabecular separation based on connected domain analysis; Texture feature vector Ft: 3D gray-level co-occurrence matrix features in 13 directions and rotation-invariant 3D-LBP code; Topological eigenvector F s : Figure G i The average degree, clustering coefficient and shortest path length of The clinical multimodal data encoding module independently acquires and processes three types of clinical data: Structured data encoder: normalizes numerical indicators such as age, gender, and BMI to the range [0, 1]; Biochemical index coder: PTH, 25(OH)D, CTX and other bone metabolism markers were logarithmically transformed and z-score normalized; Medical record text encoder: Use the pre-trained BioBERT model to convert free text into 768-dimensional semantic vectors; The multimodal graph attention fusion network receives the outputs of the trabecular network graph construction module, the multidimensional image feature extraction module, and the clinical multimodal data encoding module, including: Construct a three-layer heterogeneous graph G′=(V′, E′), where V′ contains three types of nodes: imaging nodes, clinical nodes, and semantic nodes; Design a cross-modal attention mechanism, and the attention weight calculation formula is: where α ij is the attention weight from node j to node i, h i ∈R d is the hidden layer representation vector of node i, W q , W k ∈R d ×d is the learnable query and key projection matrix, d = 256 is the hidden layer dimension, and N(i) is the set of neighbor nodes of node i; Update node representation via message passing mechanism: in is the representation of node i in the l-th layer network, W v ∈R d ×d is the learnable value projection matrix, σ is the ReLU activation function; The segment-level diagnosis output module is connected to the multimodal graph attention fusion network and outputs three levels of diagnosis results: Vertebral level: T-score prediction for each vertebra, osteoporosis grade (normal / osteopenia / osteoporosis), and compression fracture risk score; Subregion level: 8×4 risk matrix, indicating the degree of abnormality in 8 subregions of each vertebra from L1 to L4; Patient level: comprehensive FRAX score and prediction of future fracture probability.
2. The multi-feature fusion diagnosis system for the L1-L4 lumbar segment according to claim 1 is characterized in that: In the trabecular network graph construction module, the specific method for constructing edge connections is: Calculate the structural similarity between nodes vi and vj: Where S(v i ,v j ) is the node v i and v j The structural similarity between them is in the range of [0, 1]; f i is the feature vector of node i; σ = 0.5 is the scale parameter that controls the similarity decay rate; θ ij is the angle between the main directions of the trabeculae at nodes i and j; When S(v i ,v j )>0.7 and the Euclidean distance between nodes is less than 3 voxels, the weight is S(v i ,v j )’s edge connection; A spectral clustering algorithm was used to aggregate highly connected nodes into trabecular bone clusters, and the connections between clusters represented the key force transmission paths of the trabecular bone network.
3. The multi-feature fusion diagnosis system for the L1-L4 lumbar vertebrae according to claim 1 is characterized in that: The multimodal graph attention fusion network module adopts a staged training strategy: Phase 1: Freeze the clinical data branch and train only with image features. The loss function in is the predicted output based on the image, y is the true label, ||W|| 2 is the L2 norm of the network weight matrix, α = 0.01 is the regularization coefficient; Phase 2: Unfreeze all parameters and introduce multimodal contrastive learning loss: Among them (z img ∈R 128 is the normalized embedding vector of the imaging modality, Z clin ∈R 128 is the normalized embedding vector of the clinical modality, is the cosine similarity, τ = 0.07 is the temperature parameter, and B is the set of all positive and negative sample pairs in the batch; Phase 3: Adding temporal consistency constraints: L temporal =||f(x t )-f(x t-Δt )-g(Δt)|| 2 Where f(·) is the feature extraction network, x t is the CT image at time t, Δt is the time interval between two scans (months), g(Δt)=W t ·Δt+b t is a learnable time evolution function, W t ∈R d ×1 is the time decay weight, b t ∈R d is the time bias term; The total loss function L = L1 + β·Lcontrast + γ·Ltemporal, where β = 0.1 and γ = 0.
05.
4. The multi-feature fusion diagnosis system for the L1-L4 lumbar vertebrae according to claim 1 is characterized in that: It also includes a confidence calibration module, which is connected to the diagnostic output module: Calculate prediction uncertainty: Perform T = 10 forward passes on the same input (with dropout enabled) and calculate the prediction variance: where p t is the predicted probability of the t-th forward propagation, and p is the mean of T predictions; Data integrity score: S complete =0.4·I img +0.3·I clin +0.3·I text , where I img , I clin , I text ∈[0, 1] is the integrity index of each modal data; Calibrated confidence: Among them C calibrated is the confidence after calibration, ranging from [0, 1], C raw is the confidence of the original output of the model.
5. The multi-feature fusion diagnosis system for the L1-L4 lumbar vertebrae according to any one of claims 1-4, characterized in that: The specific partitioning method used by the vertebral anatomy partitioning module is: Construct L1-L4 standard template based on Statistical Shape Model, containing 1024 corresponding points; The template is registered to the patient CT image using the Coherent Point Drift algorithm; Based on the registered corresponding points, the Voronoi diagram partitioning method is used to automatically generate eight sub-regions with clear anatomical significance; Boundary smoothing uses 3D morphological opening operations to eliminate jagged boundaries.
6. The multi-feature fusion diagnosis system for the L1-L4 lumbar vertebrae according to claim 1 is characterized in that: The multi-dimensional image feature extraction module also includes: Adaptive feature selection unit, which uses LASSO regularization to automatically select the feature subset most relevant to the diagnostic task. The regularization parameter is determined by 5-fold cross-validation. The feature normalization unit uses domain adaptation technology to eliminate device-related bias for different CT devices and scanning protocols, and adopts the maximum mean difference (MMD) as the inter-domain distance metric.
7. The multi-feature fusion diagnosis system for the L1-L4 lumbar segment according to claim 1 is characterized in that: The segment-level diagnostic output module further includes: The natural language generation unit, based on the pre-trained GPT-2 model, converts structured diagnostic results into natural language reports that conform to clinical habits; Interactive 3D visualization unit, using Web graphics library technology to achieve real-time rendering of bone distribution inside the vertebral body, supporting arbitrary angle rotation and cross-section viewing; The personalized risk assessment unit combines patients' lifestyle factors and uses the Cox proportional hazards model to predict 10-year fracture risk.
8. A multi-feature fusion diagnosis method for the L1-L4 lumbar vertebrae segment, applicable to the multi-feature fusion diagnosis system for the L1-L4 lumbar vertebrae segment according to any one of claims 1-7, characterized in that: The following steps are involved: S1: Acquire thin-slice CT scan data of the patient's lumbar spine, with a slice thickness of ≤1 mm, and obtain a DICOM format image sequence; S2: Image preprocessing module performs: S2.1: Apply a deep learning-based metal artifact removal algorithm; S2.2: Use bilateral filter to reduce noise and preserve edge information; S2.3: Intensity window normalized to [-1000, 2000] HU; S3: The vertebral anatomy partitioning module automatically locates the L1-L4 vertebral centers and completes the 8-region division; S4: Parallel execution of feature extraction: S4.1: The trabecular bone network map construction module generates 32 sub-region maps (4 vertebrae × 8 regions) through Hessian matrix analysis; S4.2: Multi-dimensional image feature extraction module calculates the 43-dimensional feature vector of each sub-region; S5: Clinical multimodal data encoding module processes patient clinical information and generates standardized feature representations; S6: The multimodal graph attention fusion network integrates the outputs of S4 and S5 and performs 4-layer graph neural network processing through the cross-modal attention mechanism; S7: Segment-level diagnostic output module generates hierarchical diagnostic results and visualization reports; S8: The confidence calibration module evaluates diagnostic reliability based on prediction uncertainty and data integrity, and prompts manual review when the calibration confidence is less than 0.
7.
9. The multi-feature fusion diagnosis method for the L1-L4 lumbar vertebrae segment according to claim 8, characterized in that: In step S6, the specific process of the multimodal fusion is: S6.1: Combine 32 sub-region graph nodes, 15 clinical feature nodes, and 1 semantic node into a 48-node heterogeneous graph; S6.2: Initialize the cross-modal edge weight matrix W∈R48×48 using Xavier initialization. S6.3: Perform 4 rounds of message passing and feature updates. In each round, use the attention mechanism to calculate the weight of information propagation between nodes, and apply Layer Normalization after the update. S6.4: Use global average pooling to aggregate node features and output the diagnosis prediction vector through the fully connected layer.
10. The multi-feature fusion diagnosis method for the L1-L4 lumbar vertebrae segment according to claim 8 or 9, characterized in that: Also includes: S9: When historical CT data of patients are available, perform temporal comparative analysis to evaluate the rate and trend of bone degradation using temporal consistency constraints; S10: Generate personalized treatment recommendations and follow-up plans based on diagnostic results and patient characteristics in combination with clinical guidelines.
Citation Information
Cited By
Multi-modal orthopedic infection diagnosis method and system based on double-flow encoder fusion
CN121122681A
A Multimodal Diagnostic Method and System for Orthopedic Infections Based on Dual-Stream Encoder Fusion
CN121122681B
Rheumatoid arthritis multi-dimensional intelligent data analysis method and system
CN121439170A
Concealed femoral neck fracture AI detection system based on two-stage artifact contrast learning
CN122335871A
AI-based detection system for occult femoral neck fractures using two-stage artifact contrast learning
CN122335871B