Osteoporosis auxiliary diagnosis method fusing CT image features and semantic knowledge graph

By integrating CT image features with semantic knowledge graphs to assist in the diagnosis of osteoporosis, we have achieved collaborative analysis of multimodal data and anatomical perception segmentation, which solves the shortcomings of existing technologies in terms of diagnostic accuracy and interpretability, and improves the precision and reliability of osteoporosis diagnosis.

CN120977534APending Publication Date: 2025-11-18NANJING WANGSHI INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510921413.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing osteoporosis diagnostic technologies have shortcomings in multimodal information fusion, refined image analysis, knowledge reasoning, and clinical adaptability, resulting in limited diagnostic accuracy and interpretability, making them difficult to apply widely in clinical settings.

Method used

An osteoporosis-assisted diagnostic method integrating CT image features and semantic knowledge graphs achieves collaborative analysis of images, electronic medical record text, and structured clinical data through a multimodal deep fusion architecture and an anatomically perceptive adaptive segmentation network. It automatically identifies and eliminates interfering factors and generates interpretable diagnostic pathways.

Benefits of technology

It improves the accuracy and interpretability of osteoporosis diagnosis, with a segmentation accuracy of 0.89. The diagnostic results include detailed reasoning, which significantly increases doctors' trust in the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977534A_ABST
    Figure CN120977534A_ABST
Patent Text Reader

Abstract

The invention provides an osteoporosis auxiliary diagnosis method fusing CT image features and a semantic knowledge graph, and the method is characterized in that the method comprises the steps: multi-source heterogeneous data collection and standardization processing; constructing a modal exclusive depth coding network; mapping and aligning a unified semantic space; carrying out multi-level cross-modal attention fusion; self-adaptive segmentation and feature enhancement of anatomical perception are carried out; semantic reasoning guided by the knowledge graph; performing multi-task collaborative diagnosis output; and training optimization and interference elimination. Through automatic multi-modal analysis and intelligent diagnosis, the workload of doctors in the imaging department is remarkably relieved, and the diagnosis time of a single example is shortened to be within 3 minutes from the average 15-20 minutes. The standardized diagnosis process of the system improves the diagnosis consistency among different doctors, reduces the diagnosis deviation caused by experience difference, and helps primary hospitals to improve the osteoporosis diagnosis and treatment level.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence and medical image analysis, and particularly relates to an osteoporosis auxiliary diagnosis method fusing CT image features and semantic knowledge graph. BACKGROUND

[0002] Osteoporosis is a systemic bone disease characterized by bone mass loss and bone microstructure destruction, leading to increased bone fragility and increased risk of fracture. With the acceleration of population aging, osteoporosis has become a chronic disease that seriously threatens the health of the middle-aged and elderly, affecting about 200 million people worldwide. Early and accurate diagnosis is of great significance for preventing osteoporotic fractures and improving the quality of life of patients.

[0003] Currently, the diagnosis of osteoporosis in clinical practice mainly relies on dual-energy X-ray absorptiometry (DXA) to measure bone density, but this method can only provide two-dimensional projection area bone density and cannot reflect the three-dimensional structure and bone mass information of the skeleton. Although quantitative CT (QCT) can provide volume bone density and bone structure information, traditional QCT analysis methods have many limitations: first, existing methods mainly rely on single image features such as bone density values or simple texture parameters, making it difficult to fully capture the complex pathological changes of osteoporosis; second, image analysis is disconnected from clinical information, failing to effectively integrate patient history, laboratory tests, medication history, and other multidimensional information, resulting in limited diagnostic accuracy; third, there is a lack of systematic medical knowledge reasoning mechanism, which cannot provide intelligent decision support based on clinical experience and medical guidelines.

[0004] Insufficient granularity of image feature modeling is a key problem. Existing methods usually perform global analysis on the entire region of interest, ignoring the specificity of different anatomical structures. For example, there are differences in osteoporosis manifestations in different parts such as vertebral cortical bone, cancellous bone, and endplate, which require targeted analysis strategies. In addition, common interference factors such as bone cement filling, bone tumors, and aortic calcification can seriously affect the accuracy of bone density measurement, but existing systems lack effective identification and exclusion mechanisms.

[0005] In clinical practice, doctors' diagnostic decisions are based not only on image findings, but also on a variety of factors such as patient age, gender, medical history, medication, lifestyle, and other factors. However, existing systems are difficult to simulate this complex clinical reasoning process, and the output results often lack explainability, making it difficult for doctors to understand the basis of the system's judgment, limiting its acceptance and application value in clinical practice. At the same time, data quality in actual clinical environment is uneven, and there may be image blurring, incomplete medical records, missing examination items, and other situations, and existing methods have high requirements for data integrity, and their performance decreases significantly when faced with incomplete data.

[0006] In summary, the existing osteoporosis auxiliary diagnosis technology has obvious deficiencies in multi-modal information fusion, refined image analysis, knowledge reasoning, clinical adaptability, etc. It is urgent to develop new technical solutions to improve the accuracy, interpretability and practicality of diagnosis, and better serve the clinical diagnosis and treatment needs. SUMMARY

[0007] To overcome the deficiencies of the prior art, the present application proposes an osteoporosis auxiliary diagnosis method that fuses CT image features and semantic knowledge graph, which can maintain stable performance even when facing incomplete data. Even if some modal data is missing, the system can maintain diagnostic accuracy through compensation mechanisms of other available modalities, with a performance decline of no more than 5%. At the same time, the system can automatically identify and exclude common interference factors such as bone cement filling, bone tumors, and vascular calcification, avoiding the risk of misdiagnosis in clinical practice.

[0008] To achieve the above-mentioned purpose, the present application proposes an osteoporosis auxiliary diagnosis method that fuses CT image features and semantic knowledge graph, comprising the following steps:

[0009] Step S1: Multi-source heterogeneous data acquisition and standardization processing, specifically as follows:

[0010] S1.1: Collect CT scan sequences of patient's lumbar vertebrae L1-L4 and hip, obtain three-dimensional body data in DICOM format, and standardize HU value to the range of [-1000, 3000];

[0011] S1.2: Extract electronic medical record text within the corresponding time window, including radiology report, bone density assessment record and clinical diagnosis description, and perform word segmentation, stop word removal and medical terminology standardization processing;

[0012] S1.3: Collect structured clinical indicators, including age, gender, BMI, blood calcium, 25(OH)D level, history of previous fractures, etc., and perform missing value filling and normalization;

[0013] S1.4: Time align the collected multi-modal data, and establish a multi-modal data set associated with the patient's unified identifier;

[0014] Step S2: Construct a modal-specific deep encoding network, specifically as follows:

[0015] S2.1: For CT image data, construct an image encoder based on 3D ResNet-50, set the input size to 128x128x64 voxels, extract spatial features layer by layer through 5 residual blocks, and output a 512-dimensional image embedding vector V_img;

[0016] S2.2: For textual data, the ClinicalBERT model pre-trained on the MIMIC-III dataset is adopted to tokenize the medical record text and input the 12-layer Transformer encoder to extract the 768-dimensional text embedding vector V_text corresponding to the [CLS] token;

[0017] S2.3: For structured data, a 3-layer fully connected network (input dimension-256-128-64) is designed with ReLU activation and Dropout (p=0.3) regularization, outputting a 64-dimensional structured feature vector V struct ;

[0018] S2.4: A modal quality evaluation mechanism is established to calculate the prediction uncertainty of each encoder output through Monte Carlo Dropout (T=10 times of forward propagation) to generate the confidence score Conf m ;

[0019] Step S3: Unified semantic space mapping and alignment, specifically as follows:

[0020] S3.1: A shared projection layer is designed to map the three heterogeneous embedding vectors {V img ,V text ,V struct} to a 256-dimensional shared space through linear transformation and layer normalization to obtain z img ,z text ,z struct ;

[0021] S3.2: A contrastive learning loss is introduced to maximize the cosine similarity of different modal embeddings for the same patient and minimize the similarity between different patients, with a temperature parameter τ=0.07;

[0022] S3.3: Based on the confidence score of S2.4, the representation strength of each modality in the shared space is dynamically adjusted;

[0023] Step S4: Multi-level cross-modal attention fusion, specifically as follows:

[0024] S4.1: A three-modal cyclic attention mechanism is constructed to calculate the cross-attention of image→text, text→structured, and structured→image in turn, using 8 heads of attention with a dimension of 32;

[0025] S4.2: A residual gated unit is designed to fuse the original embedding and attention output through a learnable gating parameter α:

[0026]

[0027] Where:

[0028] The feature vector after fusion of the first layer, dimension 256

[0029] α (l) : gating parameter of the first layer, value range [0,1], dynamically calculated by the linear layer activated by sigmoid

[0030] Output feature of the cross-modal attention mechanism of the first layer

[0031] Input feature of the first layer (original modality embedding or output of the previous layer)

[0032] S4.3: Stack 3 fusion Transformer layers, output 512-dimensional deep fusion feature F multi ;

[0033] Step S5: Adaptive segmentation and feature enhancement of anatomical perception, as follows:

[0034] S5.1: Construct a dual-branch A3-Net network:

[0035] Global branch: divide the image into 16x16 patches using Vision Transformer, extract global dependency F through 12-layer Transformer global ;

[0036] Local branch: generate spatial attention masks for 15 key structures using the anatomical probability atlas constructed from 1000 normal bone CTs, extract local features F local ;

[0037] S5.2: Fuse the features of the two branches through a gating mechanism:

[0038] F seg = σ(W g ×F global )⊙F global +(1-σ(W g ×F global ))⊙F local

[0039] Where:

[0040] F seg : fusion feature output of the segmentation network, dimension same as input feature

[0041] F global : feature map extracted by the global branch

[0042] F local : feature map extracted by the local branch based on anatomical prior

[0043] W g : gating weight matrix, dimension dxd (d is feature dimension)

[0044] σ(): sigmoid activation function, limits output in [0, 1]

[0045] : element-wise product operation (Hadamard product)

[0046] S5.3: Fine segmentation of key structures such as vertebral body, femoral neck, etc., extraction of Haralick texture features (14 dimensions) and morphological features, formation of 128-dimensional anatomical feature vector F anatomy ;

[0047] S5.4: Concatenate anatomical features and multi-modal fusion features to get 640-dimensional enhanced features F enhanced ;

[0048] Step S6: Knowledge graph guided semantic reasoning, as follows:

[0049] S6.1: Build an osteoporosis medical knowledge graph, containing 116 entities (3 types of disease states, 47 symptoms, 28 risk factors, 15 drugs, 23 imaging manifestations) and 8 semantic relationships;

[0050] S6.2: Map enhanced features to graph query vectors and perform 2-hop reasoning through 2-layer graph attention network:

[0051]

[0052] Where:

[0053] Feature representation of node i at the l+1 layer

[0054] Feature representation of node j at the l layer

[0055] Attention coefficient of node j to node i in the l layer

[0056] Neighbor node set of node i

[0057] W (l) : learnable weight matrix of the l layer;

[0058] S6.3: Generate interpretable diagnostic basis based on reasoning path;

[0059] Step S7: Multi-task collaborative diagnosis output, as follows:

[0060] S7.1: Output the three-level diagnosis (normal: T≥-1.0, osteopenia: -2.5<T<-1.0, osteoporosis: T≤-2.5) through the classification head;

[0061] S7.2: Predict the BMD value through the regression head;

[0062] S7.3: Generate a structured diagnosis report containing risk level, key imaging findings, and recommended follow-up period;

[0063] S7.4: Output an attention visualization heat map marking high-risk anatomical regions;

[0064] Step S8: Training optimization and interference exclusion, specifically as follows:

[0065] S8.1: Use multi-task learning to dynamically adjust the weights of each task through gradient normalization;

[0066] S8.2: Identify and exclude interference factors such as bone cement (HU>3000 and volume>5mm 3 ), bone tumors, and vascular calcification;

[0067] S8.3: Support robust processing for missing modalities, using a learnable [MASK] embedding to replace missing modalities.

[0068] Further, in the residual gate fusion mechanism of S4.2, the calculation method of the gate parameter a is: a = σ(W α × [A output ; z input ]) where [·;·] represents feature concatenation operation, W α is a learnable weight matrix, and σ is a sigmoid activation function, ensuring that a ∈ [0,1].

[0069] Further, the dual-branch fusion of S5.2 uses a spatial attention-guided adaptive mechanism to dynamically adjust the fusion ratio of global and local features based on the importance of anatomical structures.

[0070] Further, the graph attention reasoning network of S6.2 learns the attention weights between nodes to achieve semantic mapping and logical reasoning from patient features to medical knowledge.

[0071] Further, the method's robust processing strategy when modalities are missing:

[0072]

[0073] Where:

[0074] F tobust : Robustly fused feature representation

[0075] Set of available modalities

[0076] Conf m : confidence score of modality m (calculated by S2.4)

[0077] z m : embedding representation of modality m in unified semantic space The formula dynamically adjusts the fusion weight according to the confidence of the available modalities.

[0078] Further, the system performance indicators are: classification accuracy ≥ 92%, BMD prediction error < 0.05 g / cm 2 , single instance reasoning time < 3 seconds.

[0079] Compared with the prior art, the beneficial effects of the present application are:

[0080] 1. The present application provides an osteoporosis auxiliary diagnosis method fusing CT image features and semantic knowledge graph, which realizes the collaborative analysis of CT images, electronic medical record texts and structured clinical data through a multi-modal deep fusion architecture, and fully utilizes the complementary advantages of different information sources.

[0081] 2. The present application provides an osteoporosis auxiliary diagnosis method fusing CT image features and semantic knowledge graph, which adopts an anatomically-aware adaptive segmentation network to accurately identify and analyze 15 key anatomical structures such as L1-L4 vertebrae, acetabulum, femoral head, femoral neck and Ward's triangle. The segmentation accuracy reaches an average Dice coefficient of 0.89, and the vertebrae segmentation accuracy is as high as 0.92. By extracting 14-dimensional Haralick texture features and morphological features of each anatomical region, quantitative evaluation of the trabecular microstructure is realized, providing accurate quantitative indicators for personalized risk assessment.

[0082] 3. The present application provides an osteoporosis auxiliary diagnosis method fusing CT image features and semantic knowledge graph, which can generate an interpretable diagnosis path based on 2-hop graph reasoning, clearly indicating the key risk factors and image manifestations leading to the diagnosis conclusion. The output structured report not only contains the diagnosis result, but also provides complete reasoning basis and personalized suggestions, significantly improving the doctor's trust in the system diagnosis. BRIEF DESCRIPTION OF DRAWINGS

[0083] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0084] Figure 1is a schematic diagram of the method flow of the present application

[0085] Figure 2 is a schematic diagram of an adaptive segmentation network DETAILED DESCRIPTION

[0086] The technical solutions of the present application will be described more clearly and completely below by combining the drawings and through the description of the preferred embodiments of the present application.

[0087] As Figure 1 shown, the present application is:

[0088] Step S1: Multi-source heterogeneous data acquisition and standardization processing;

[0089] Step S2: Construction of modal exclusive deep coding network;

[0090] Step S3: Unified semantic space mapping and alignment;

[0091] Step S4: Multi-level cross-modal attention fusion;

[0092] Step S5: Anatomical perception adaptive segmentation and feature enhancement;

[0093] Step S6: Knowledge graph guided semantic reasoning;

[0094] Step S7: Multi-task collaborative diagnosis output;

[0095] Step S8: Training optimization and interference elimination

[0096] As Figure 2To achieve accurate identification of key bone structures and personalized risk analysis, this study proposes an Anatomy-Aware Adaptive Attention Network (A3-Net) guided by anatomical priors. This network effectively improves the segmentation performance in high-risk areas of osteoporosis by modeling global context and local anatomical region features through a dual-branch architecture. In the specific process, one branch introduces a structure-guided module that generates a structure attention map using pre-built anatomical atlases or automatically extracted landmark points. This map serves as a saliency guide to impose spatial constraints on sensitive structures such as vertebral bodies, acetabula, and femoral necks in CT images, enhancing the model's focus on key areas. The other branch uses a global Transformer encoder to model long-range dependencies and semantic context in the complete image, capturing spatial relationships between bone structures. The two features are integrated through a gated spatial attention mechanism (GSA) in the fusion stage, dynamically regulating the fusion ratio between local structure-guided information and global semantic information. This enhances the perception of texture and structural changes while maintaining boundary accuracy, ultimately outputting fine segmentation results with anatomical consistency and clinical sensitivity.

[0097] As a specific implementation, first, multi-source heterogeneous data collection and preprocessing are performed. 5000 patient medical data are obtained from cooperative hospitals, including CT scan images of lumbar vertebrae L1-L4 and hip, electronic medical record texts, and clinical test indicators. The CT images use a unified scanning protocol: tube voltage 120kV, tube current 200mA, layer thickness 1mm, reconstruction matrix 512x512, all image data is desensitized and stored in DICOM format. The original HU value is standardized to the range of-1000 to 3000 through linear transformation, ensuring the consistency of data collected by different devices. For electronic medical record texts, regular expressions are used to extract key descriptions in radiology reports, including "bone density reduction", "bone trabecula sparse", "compression fracture" and other characteristic expressions, and synonyms are unified to standard terms through a medical terminology dictionary. Structured clinical data includes patient basic information such as age, gender, height, weight, biochemical indicators such as serum calcium 2.2-2.6mmol / L, 25-hydroxyvitamin D, alkaline phosphatase, and medical history information. Numerical features are standardized by z-score, and categorical features are one-hot encoded.

[0098] In the construction of the modality-specific encoding network, an improved 3D ResNet-50 architecture is designed for CT image data. The network input is a 128x128x64 voxel block. First, an initial feature is extracted through a 7x7x7 3D convolution layer, outputting 64 feature channels. Then there are five residual block groups, each containing a different number of residual units: the first group has three units, maintaining 64 channels; the second group has four units, expanding to 128 channels; the third group has six units, reaching 256 channels; the fourth and fifth groups each have three units, with the channel number increasing to 512. Each residual unit uses a bottleneck structure, sequentially passing through a 1x1x1 convolution to reduce dimension, a 3x3x3 convolution to extract features, and a 1x1x1 convolution to increase dimension, and a skip connection is used to alleviate the gradient vanishing problem. Finally, the spatial features are compressed into a 512-dimensional vector through global average pooling. The text encoding uses the ClinicalBERT model, which is pre-trained on the MIMIC-III dataset containing about 2 million clinical records. The input text is first tokenized using WordPiece, with a maximum length of 512 tokens. After processing through a 12-layer Transformer encoder, the 768-dimensional feature at the CLS position is extracted as the text representation. Structured data is encoded through a 3-layer fully connected network, with neuron counts of input dimension, 256, 128, and 64, respectively. Each layer is followed by BatchNorm and ReLU activation, and a Dropout rate of 0.3 is used to prevent overfitting.

[0099] To evaluate the quality of each modality data, the Monte Carlo Dropout method is implemented to calculate the prediction uncertainty. In the inference phase, the Dropout activation is kept, and 10 forward propagations are performed on the same input to calculate the output variance as the uncertainty measure, and then the confidence score is obtained. The embedding vectors of the three modalities are mapped to a 256-dimensional shared semantic space through independent linear projection layers. The projection matrix is initialized using the Xavier method to ensure stable gradient propagation. The mapped features are normalized through LayerNorm to eliminate the scale difference between different modalities.

[0100] Cross-modal fusion adopts a recurrent attention mechanism to realize deep semantic interaction. First, the attention between images and text is calculated. The image features are used as queries, and the text features are used as keys and values. Through 8-head attention mechanism, the association between each spatial position in the image and the text description is captured. The dimension of each attention head is 32, and the scaling dot product attention is used to ensure numerical stability. The obtained image-text fusion features are then used for the second round of attention calculation with the structured data features to form the preliminary three-modal fusion representation. To preserve the original modal information, a residual gating unit is designed to dynamically adjust the fusion ratio through learnable parameters. The gating parameters are obtained by concatenating the attention output and the original input, sending it to a linear layer, and then activating it with a sigmoid function to ensure the value is between 0 and 1. The entire fusion network is stacked with 3 layers, each containing multi-head self-attention, cross-modal attention, residual connection, and feedforward network, and finally outputs a 512-dimensional fusion feature.

[0101] The implementation of the anatomy-aware segmentation network A3-Net is divided into two parallel branches. The global branch adopts the VisionTransformer architecture, which divides the input CT slices into 16x16 image blocks. Each block is projected linearly to obtain a 768-dimensional embedding, which is then sent to a 12-layer Transformer encoder after adding learnable position encoding. Each layer contains multi-head self-attention and MLP modules to capture long-range dependencies between different anatomical structures. The local branch uses a pre-constructed anatomical atlas to guide attention focusing. The anatomical atlas is constructed by collecting CT data from 1000 normal people, aligning them to the standard space using ANTs registration tools, and calculating the probability distribution of each voxel belonging to a specific anatomical structure, including L1-L4 vertebrae, acetabulum, femoral head, femoral neck, Ward's triangle, and other 15 key parts. Based on this probability map, a spatial attention mask is generated, which is multiplied element-wise with the input features to strengthen the feature expression of key areas. The features of the two branches are adaptively fused through a gating mechanism, and the gating weight is calculated based on the global features, realizing the complementary advantages of global semantics and local details.

[0102] After segmentation, quantitative features are extracted for each anatomical structure. Texture features use the gray level co-occurrence matrix method to calculate Haralick features in 0, 45, 90, and 135 degrees, including 14 indicators such as energy, contrast, correlation, homogeneity, and entropy, to comprehensively describe the microstructure characteristics of bone trabeculae. Morphological features include volume calculated by voxel counting multiplied by voxel volume, surface area extracted using the marching cubes algorithm, sphericity reflecting the regularity of the structure, and compactness describing the internal filling situation. These features are concatenated with the multi-modal fusion features to form a 640-dimensional enhanced feature vector.

[0103] The knowledge graph is constructed based on medical literature and clinical guidelines in the field of osteoporosis. Entities include 3 disease states, i.e., normal, osteopenia, and osteoporosis, 47 clinical symptoms such as low back pain, height loss, and humpback, 28 risk factors including age, gender, smoking, alcohol consumption, and glucocorticoid use, and 15 drug treatments such as bisphosphonates, calcitonin, and estrogen. There are 8 types of relationship types, including "cause", "manifest", "treatment", and "prevention". The TransE algorithm is used to learn 100-dimensional entity and relationship embeddings, and the graph representation is optimized by minimizing the positive sample distance and maximizing the negative sample distance. The enhanced features of the patient are mapped to the graph query vector through MLP, and a 2-layer graph attention network is used for reasoning, each layer aggregating 2-hop neighbor information, and the attention coefficient being calculated through a learnable attention vector and a LeakyReLU activation.

[0104] The multi-task output module includes three task heads. The classification head is a 3-layer MLP structure with 512-256-3 neurons, and outputs three classification probabilities normalized by softmax. The regression head is also a 3-layer MLP with 512-256-1 neurons, which directly outputs the BMD prediction value in grams per square centimeter. The T-score is calculated according to the WHO standard, using the average BMD value of 1.200 and the standard deviation of 0.120 of healthy adults aged 20-30 as a reference. The segmentation head adopts a U-Net decoder structure, which fuses multi-scale features through skip connections and outputs a segmentation mask with the same size as the input.

[0105] The training adopts a multi-task learning strategy, and the total loss includes classification loss, regression loss, segmentation loss, and contrastive learning loss. The classification loss uses cross-entropy, the regression loss uses Huber loss to improve the robustness to outliers, and the segmentation loss uses Dice loss. The weights of each task are dynamically adjusted through gradient normalization to balance the learning speed of different tasks. The optimizer is AdamW with an initial learning rate of 0.0001 and a weight decay of 0.01, using cosine annealing scheduling with a minimum learning rate of 0.000001. The batch size is 16, and the training is performed for 100 rounds, with the best model saved after each round on the validation set.

[0106] The data augmentation strategy includes random rotation of positive and negative 15 degrees to simulate scan angle changes, random translation of positive and negative 10 pixels to simulate positioning deviations, intensity perturbation of positive and negative 10% to simulate scan parameter differences, and elastic deformation to simulate respiratory motion. During training, a modality is randomly discarded with a probability of 0.2 to force the model to learn the compensation mechanism between modalities. The missing modality is replaced by a learnable MASK embedding with the same dimension as the corresponding modality embedding.

[0107] The interference factor identification is realized by threshold and morphological analysis. The bone cement appears as extremely high density, HU value greater than 3000, and the region with a volume greater than 5 cubic millimeters is identified and marked for exclusion by connected domain analysis. Bone tumors are identified by detecting abnormal expansive lesions, and the main features include cortical destruction, soft tissue mass, etc. Vascular calcification adopts a tubular structure detection algorithm based on the Hessian matrix to identify linear or curved high-density shadows. These interference areas are automatically excluded during BMD calculation to ensure measurement accuracy.

[0108] The system is deployed on a server equipped with NVIDIA V100 GPU, with model parameter amount of about 85 million and memory occupancy of 6.2 GB. The single instance complete inference process takes 2.8 seconds, including data loading 0.3 seconds, feature extraction 1.2 seconds, fusion inference 0.8 seconds, and result generation 0.5 seconds. On an independent test set containing 1000 cases, the three-classification accuracy is 93.2%, the sensitivity is 91.5%, and the specificity is 94.1%. The average absolute error of BMD prediction is 0.042, and the correlation coefficient is 0.94. The average Dice coefficient of the segmentation task is 0.89, and the segmentation accuracy of the vertebral body is as high as 0.92, and the femoral neck region is slightly lower at 0.86.

[0109] In practical application, the diagnostic report generated by the system includes osteoporosis risk level and confidence, BMD value and T-score of each anatomical site, textual description of key imaging findings, risk factor analysis based on knowledge graph reasoning, personalized follow-up recommendations, 12 months for normal patients, 6 months for osteopenia, and 3 months for osteoporosis, as well as visualized results with heat map annotations, which intuitively display high-risk areas. The entire system has been in trial operation in 3 first-class hospitals for 6 months, and has assisted in the diagnosis of more than 15000 cases, significantly improving the early detection rate and diagnostic consistency of osteoporosis.

[0110] The above specific embodiments only describe the preferred embodiments of the present application, and do not limit the protection scope of the present application. Without departing from the design concept and spirit of the present application, various modifications, substitutions and improvements of the technical solutions of the present application made by those skilled in the art according to the description and drawings of the present application should belong to the protection scope of the present application. The protection scope of the present application is determined by the claims.

Claims

1. An osteoporosis auxiliary diagnosis method fusing CT image features and semantic knowledge graphs, characterized in that, Comprise the following steps: Step S1: Multi-source heterogeneous data acquisition and standardization processing, specifically as follows: S1.1: Collect CT scan sequences of patient lumbar L1-L4 and hip, obtain three-dimensional body data in DICOM format, and standardize HU value to [-1000, 3000] range; S1.2: Extract electronic medical record text in the corresponding time window, including radiology report, bone density assessment record and clinical diagnosis description, and perform word segmentation, stop word removal and medical terminology standardization processing; S1.3: Collect structured clinical indicators and perform missing value filling and normalization; S1.4: Time alignment of the collected multi-modal data, establishment of multi-modal data set associated with patient unified identifier; Step S2: Constructing a modality-specific deep encoding network, specifically as follows: S2.1: For CT image data, construct an image encoder based on 3D ResNet-50, set the input size to 128x128x64 voxels, extract spatial features layer by layer through 5 residual blocks, and output a 512-dimensional image embedding vector V img ; S2.2: For text data, the ClinicalBERT model pre-trained on the MIMIC-III dataset is adopted. After tokenizing the medical record text, the 12-layer Transformer encoder is inputted, and the 768-dimensional text embedding vector V corresponding to the [CLS] label is extracted text ; S2.3: For structured data, a 3-layer fully connected network is designed with ReLU activation and Dropout regularization, outputting a 64-dimensional structured feature vector V struct ; S2.4: Establish a modal quality evaluation mechanism, calculate the prediction uncertainty of each encoder output by Monte Carlo Dropout, and generate a confidence score Conf m ; Step S3: Unified semantic space mapping and alignment, specifically as follows: S3.1: Design shared projection layer, three heterogeneous embedding vectors {V img ,V text ,V struct} are uniformly mapped to 256-dimensional shared space through linear transformation and layer normalization, obtaining z img ,z text ,z struct ; S3.2: Introduce contrastive learning loss, maximize the cosine similarity of different modalities of the same patient, and minimize the similarity between different patients, with temperature parameter τ = 0.07; S3.3: Based on the confidence score of S2.4, dynamically adjust the representation strength of each modality in the shared space; Step S4: Multi-level cross-modal attention fusion, specifically as follows: S4.1: Construct a three-modal recurrent attention mechanism, and calculate the cross-modal attention of image→text, text→structure, and structure→image in turn, using 8 attention heads with dimension 32; S4.2: Design a residual gated unit to fuse the original embedding and attention output through a learnable gating parameter α: Wherein: the feature vector of the first layer after fusion, dimension 256 alpha (l) : gating parameter of the first layer, value range [0, 1], dynamically calculated by the linear layer with sigmoid activation Output features of the l-th layer cross-modal attention mechanism Input features of layer l S4.3: Stack 3 fused Transformer layers, output 512-dimensional deep fused features F multi ; Step S5: Anatomically-aware adaptive segmentation and feature enhancement, specifically as follows: S5.1: Construct a dual-branch A3-Net network: Global branch: divide the image into 16x16 patches using Vision Transformer, extract global dependencies by 12-layer Transformer F global ; Local branch: using the anatomical probability atlas constructed by 1000 normal bone CT, generate spatial attention masks of 15 key structures, extract local features F local ; S5.2: Fuse the features of the two branches through a gating mechanism: F seg = σ(W g × F global ) ⊙ F global + (1 - σ(W g × F global )) ⊙ F local Wherein: F seg : fused feature output of the segmentation network, same dimension as the input features F global : feature map of global branch extraction F local : local branch based on features extracted from anatomical priors W g : a gating weight matrix of dimension d x d σ(): sigmoid activation function, limiting the output to the range [0, 1] ⊙: element-wise product operation S5.3: Fine segmentation is performed on the key structure, Haralick texture features and morphological features are extracted, and a 128-dimensional anatomical feature vector F is formed anatomy ; S5.4: concatenate the anatomical features with the multi-modal fusion features to obtain 640-dimensional enhanced features F enhanced ; Step S6: Knowledge graph guided semantic reasoning, specifically as follows: S6.1: Construct an osteoporosis medical knowledge graph containing 116 entities and 8 semantic relationships; S6.2: Map the enhanced features to graph query vectors and perform 2-hop reasoning through a 2-layer graph attention network: Wherein: Feature representation of node i at layer l+1 characteristic representation of node j at layer l attention coefficient of node j to node i in the first layer The set of neighboring nodes of node i W (l) : a matrix of learnable weights for the first layer; S6.3: Generate interpretable diagnostic basis based on the reasoning path; Step S7: Multi-task collaborative diagnosis output, specifically as follows: S7.1: Output three-level diagnosis (normal: T≥-1.0, osteopenia: -2.5<T<-1.0, osteoporosis: T≤-2.5) through the classification head; S7.2: Predict BMD value through the regression head; S7.3: Generate a structured diagnostic report containing risk level, key image findings and recommended follow-up period; S7.4: Output attention visualization heat map, marking high-risk anatomical regions; Step S8: Training optimization and interference exclusion, specifically as follows: S8.1: Use multi-task learning to dynamically adjust the task weights through gradient normalization; S8.2: Identify and exclude interference factors; S8.3: Support robust processing for missing modalities, using a learnable [MASK] embedding to replace the missing modalities. 2.The method of claim 1, wherein, In the residual gating fusion mechanism of S4.2, the gating parameter α is calculated as follows: α = σ(W α ×[A output ;z input ])in[·; • represents a feature concatenation operation, W α is a learnable weight matrix, and σ is a sigmoid activation function, ensuring that a e [0, 1]. 3.The method of claim 1, wherein, The double-branch fusion of S5.2 adopts a spatial attention guided adaptive mechanism to dynamically adjust the fusion proportion of global and local features according to the importance of the anatomical structure. 4.The method of claim 1, wherein, The graph attention reasoning network of S6.2 realizes semantic mapping and logical reasoning from patient features to medical knowledge by learning the attention weights between nodes. 5.The method of claim 1, wherein, Robust processing strategy of the method in the case of missing modalities: wherein: F robust : robust fused feature representation Set of available modes Conf m : Confidence score for modality m, computed by S2.4; z m : The embedding of modality m in the unified semantic space dynamically adjusts the fusion weight of the formula according to the confidence of the available modality. 6.The method of claim 1, wherein, System performance indicators: classification accuracy ≥ 92%, BMD prediction error <0.05 g / cm 2 , singleton inference time <3 seconds.

Citation Information

Cited By

  • Multi-source heterogeneous medical data fusion and intelligent diagnosis method

    CN121545724A

  • Multi-modal mapping knowledge domain embedding and complementing method for power grid

    CN121638414A

  • Osteoporosis assessment method and system based on spine CT image and application method of system

    CN122200080A