Temporomandibular joint disease DC diagnosis system based on CBCT image

By designing a temporomandibular joint disease diagnosis system based on CBCT images, multimodal fusion analysis of image characteristics and medical record information is realized, and the problem of insufficient diagnosis in the prior art is solved, which significantly improves the accuracy and efficiency of diagnosis.

CN120072259APending Publication Date: 2025-05-30FOURTH MILITARY MEDICAL UNIVERSITY
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510094480.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art cannot effectively integrate CBCT imaging features and medical record text information, resulting in the inaccurate and comprehensive diagnosis of temporomandibular joint diseases.

Method used

A temporomandibular joint disease diagnosis system based on CBCT images was designed, and multimodal fusion analysis of image features and medical record information was realized through data acquisition and preprocessing modules, feature extraction and analysis modules, multimodal data fusion modules and DC classification and diagnosis modules.

Benefits of technology

The system can realize the precise classification, lesion localization and severity grading of temporomandibular joint diseases, significantly improving the accuracy and efficiency of diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072259A_ABST
    Figure CN120072259A_ABST
Patent Text Reader

Abstract

The invention relates to the field of temporomandibular joint disease diagnosis, and discloses a temporomandibular joint disease DC diagnosis system based on a CBCT image, and the system comprises a data collection and preprocessing module which is used for collecting CBCT image data and medical record information of a patient, carrying out denoising processing, window width and window level adjustment, normalization processing and multi-plane slicing on the CBCT image data to generate multi-channel two-dimensional data; and a feature extraction and analysis module. By using the three-dimensional anatomical features of the CBCT image and the semantic features of the medical record text information, multi-modal fusion and analysis are performed through the deep learning model, and the problem of feature loss possibly caused by a single data source is effectively solved. Through the feature extraction and analysis module and the multi-modal data fusion module, the system can accurately identify the lesion area and the disease type of the temporomandibular joint from multiple dimensions, it is ensured that the diagnosis result meets the DC classification standard, and the diagnosis accuracy is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of temporomandibular joint disease diagnosis, and particularly to a DC diagnosis system for temporomandibular joint diseases based on CBCT images. Background Art

[0002] Temporomandibular joint disease (TMD) is a common and complex disease type, involving multiple anatomical structures such as muscles, joints, and jaws. The main symptoms of TMD include masticatory muscle pain, joint clicking, and joint movement dysfunction, and its diagnosis and treatment directly affect the quality of life of patients. Due to the unclear etiology of temporomandibular joint diseases and the influence of multiple pathogenic factors, traditional diagnostic methods rely highly on the experience of clinicians and medical equipment, making it difficult to achieve efficient and accurate diagnosis. Therefore, exploring diagnostic systems based on advanced medical imaging technology and artificial intelligence algorithms has become an important direction for improving the diagnostic quality and efficiency of TMD.

[0003] In the prior art, MRI and CBCT (cone beam computed tomography) have been widely used in the auxiliary diagnosis of temporomandibular joint diseases due to their high-resolution and three-dimensional imaging capabilities for hard and soft tissues. However, MRI is expensive and cannot be widely promoted. The analysis of traditional CBCT images mostly relies on manual judgment, and doctors need to view the images frame by frame and combine medical record information for comprehensive judgment. These two methods are not only time-consuming and laborious, but also easily limited by the experience and judgment ability of doctors. In addition, some studies have attempted to introduce artificial intelligence technology for automated analysis of CBCT images, but existing systems are mostly limited to single-modal data (such as imaging data) and cannot fully utilize the semantic clues contained in medical record information, resulting in insufficient comprehensiveness and accuracy of diagnosis.

[0004] However, the core problem existing in the prior art is the lack of a system that can perform fusion analysis on three-dimensional image features and medical record text information, and cannot effectively solve the problems of feature loss and incomplete diagnosis caused by a single data source. This deficiency directly limits the system's ability to classify temporomandibular joint diseases, locate lesions, and grade the severity, ultimately affecting the accuracy and efficiency of clinical diagnosis. Therefore, there is an urgent need for an intelligent system that can achieve multi-modal data fusion and comprehensively diagnose by integrating image features and medical record information to improve the quality and efficiency of TMD diagnosis. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a DC diagnosis system for temporomandibular joint diseases based on CBCT images, which solves the problem in the prior art that CBCT image features and medical record text information cannot be fused and analyzed, thereby achieving the accurate classification, lesion location, and severity grading of temporomandibular joint diseases.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: A DC diagnosis system for temporomandibular joint diseases based on CBCT images, comprising:

[0007] A data acquisition and preprocessing module, configured to acquire CBCT image data and medical record information of patients, and perform denoising processing, window width and window level adjustment, normalization processing, and multi-plane slicing on the CBCT image data to generate multi-channel two-dimensional data;

[0008] Among them, the denoising processing removes noise in the image through Gaussian filtering or median filtering to improve the quality of subsequent image processing;

[0009] A feature extraction and analysis module, configured to perform object detection on the preprocessed CBCT image data based on the optimized YOLOv8 object detection algorithm, extract multi-scale features through a feature pyramid network, and enhance the feature extraction ability for key regions of the temporomandibular joint through a channel attention mechanism and a spatial attention mechanism;

[0010] A multi-modal data fusion module, configured to extract anatomical features from CBCT image data, extract text semantic features from medical record information, and generate unified fusion features through cross-modal feature alignment;

[0011] A DC classification and diagnosis module, configured to classify the types of temporomandibular joint diseases based on the fusion features, locate the lesion sites, grade the severity of the diseases, and output the classification, location, and grading results;

[0012] A model optimization and deployment module, configured to optimize the model through transfer learning technology, reduce the model calculation complexity through model lightweight technology, and deploy the optimized model to edge devices.

[0013] Preferably, the data acquisition and preprocessing module includes:

[0014] An image reading unit, configured to extract CBCT image data from DICOM format files;

[0015] An image denoising unit, configured to remove noise in the image through Gaussian filtering or median filtering;

[0016] An image adjustment unit, configured to adjust the window width and window level of the CBCT image data based on diagnostic requirements;

[0017] An image slicing unit, configured to slice the three-dimensional CBCT image data into two-dimensional data in the coronal plane, sagittal plane, and transverse plane;

[0018] A data normalization unit, configured to normalize the pixel values of the image data and adjust the image size;

[0019] A data augmentation unit for performing random rotation, flipping, scaling, and adding random noise to the image data.

[0020] Preferably, the feature extraction and analysis module includes:

[0021] A detection unit for performing object detection on the condyle, articular disc, and glenoid fossa in the CBCT image data based on the optimized YOLOv8 algorithm;

[0022] A multi-scale feature extraction unit for performing multi-scale feature extraction on the CBCT image data through a feature pyramid network;

[0023] An attention enhancement unit for enhancing the detection ability of the feature extraction model for the key regions of the temporomandibular joint through a channel attention mechanism and a spatial attention mechanism;

[0024] A feature fusion unit for fusing the features generated by the upsampling and downsampling paths to improve the detection accuracy of the target region;

[0025] Among them, the channel attention mechanism enhances the representation ability of important features by weighting the feature channels; while the spatial attention mechanism improves the recognition accuracy of key regions by weighting the feature spatial regions.

[0026] Preferably, the multi-modal data fusion module includes:

[0027] An image feature extraction unit for extracting anatomical features related to the temporomandibular joint from the CBCT image data;

[0028] An image feature extraction unit for extracting anatomical features related to the temporomandibular joint from the CBCT image data;

[0029] A text feature extraction unit for extracting semantic features of the patient from the medical record text information;

[0030] A cross-modal alignment unit for embedding the image features and text features into a unified semantic space through a deep learning model, or achieving feature alignment through a specific algorithm;

[0031] A fusion feature generation unit for generating fusion features through an optimized algorithm for subsequent classification and diagnosis.

[0032] Preferably, the DC classification and diagnosis module includes:

[0033] A disease classification unit for classifying the types of temporomandibular joint diseases according to the DC classification criteria;

[0034] A lesion grading unit for grading the diseases as mild, moderate, and severe according to the severity of the disease lesions;

[0035] A lesion localization unit for localizing the lesion site of the temporomandibular joint through an optimized IoU algorithm;

[0036] A diagnostic report generation unit for outputting classification results, grading results, the location of the lesion site, and treatment suggestions.

[0037] Preferably, the model optimization and deployment module includes:

[0038] A transfer learning sub-module for pre-training the model on a publicly available medical image dataset and fine-tuning the model weights through a local CBCT dataset;

[0039] A model lightweighting sub-module for reducing the computational complexity of the model through model pruning, distillation, and quantization techniques;

[0040] A parameter optimization sub-module for optimizing the model training strategy by adopting learning rate warm-up and cosine annealing scheduling techniques;

[0041] A deployment sub-module for deploying the optimized model to a portable edge device and supporting real-time inference.

[0042] Preferably, the data acquisition and preprocessing module automatically adjusts the slice direction and thickness of the CBCT image to ensure that the image data contains the feature information of the condyle, articular disc, and glenoid fossa.

[0043] Preferably, the feature extraction and analysis module shares the feature weights of the object detection task and the classification task through a joint learning strategy and optimizes the model performance through dynamic parameter adjustment.

[0044] Preferably, the diagnostic results include the disease type, disease severity, the location of the lesion site, and treatment suggestions, and the diagnostic results can be output to the electronic medical record system through the system interface.

[0045] Preferably, the system realizes the collaborative diagnosis of image data and text data through a multi-modal data fusion module and supports real-time analysis on low-computing-power devices.

[0046] The present invention provides a DC diagnostic system for temporomandibular joint diseases based on CBCT images.

[0047] It has the following beneficial effects:

[0048] 1. The present invention effectively solves the problem of feature loss that may be caused by a single data source by using the three-dimensional anatomical features of CBCT images and the semantic features of medical record text information, and performing multi-modal fusion and analysis through a deep learning model. Through the feature extraction and analysis module and the multi-modal data fusion module, the system can accurately identify the lesion area and disease type of the temporomandibular joint from multiple dimensions, ensure that the diagnosis results meet the DC classification standard, and significantly improve the accuracy of diagnosis.

[0049] 2. The present invention uses an optimized YOLOv8 model combined with the IoU improvement algorithm to accurately locate key anatomical structures such as the condyle, articular disc, and glenoid fossa, and can clearly label the spatial position and scope of the lesion site. The positioning results have high precision and high resolution, providing detailed anatomical information for doctors and facilitating subsequent treatment planning.

[0050] 3. By carefully analyzing the area, volume, and texture features of the lesion area, the present invention can perform a graded diagnosis of the severity of the disease according to the DC classification standard. The system grading is not limited to a simple classification of mild, moderate, and severe, but can also output specific scores of the disease by combining multi-dimensional features, providing a more refined reference basis for the treatment decisions of clinicians.

[0051] 4. By using deep learning technology and the real-time inference ability of edge computing devices, the present invention realizes the full-process automation from data input to diagnosis result output. The optimized deep learning model combined with quantization and pruning techniques significantly reduces the inference time, and the diagnosis process only takes a few seconds, greatly improving the diagnosis efficiency and meeting the clinical real-time requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a flowchart of the present invention;

[0053] Figure 2 is a schematic diagram of model optimization and deployment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] Please refer to the attached Figure 1 - attached Figure 2 , the embodiments of the present invention provide a DC diagnosis system for temporomandibular joint diseases based on CBCT images, including:

[0056] 1. Data acquisition and preprocessing module

[0057] For the data acquisition and preprocessing module, this module plays a key role in the DC diagnosis system of temporomandibular joint diseases and is the basis for realizing subsequent diagnosis tasks. Through this module, CBCT image data and relevant medical record information can be obtained from patients, and the image data can be subjected to multi-level cleaning, standardization, and slicing processing to ensure the quality and consistency of the image data in subsequent analysis and modeling. The medical record information is also processed through structuring and standardization to complement the image data and provide good input conditions for multi-modal fusion. The design of this module focuses on solving technical problems such as inconsistent quality, insufficient diversity, and noise interference of CBCT image data.

[0058] 1.1 Data Acquisition

[0059] In this embodiment, the data acquisition part includes the acquisition of CBCT image data and the collection of patient medical record information.

[0060] CBCT image data: The CBCT scanning device is used to collect images of the patient's temporomandibular joint area, generating an image data file containing three-dimensional structure information, and the data is stored in DICOM format. The DICOM file contains the patient's anatomical data, imaging parameters (such as slice thickness, resolution, scanning angle), and metadata information.

[0061] Medical record information: The clinical doctor records text information such as the patient's chief complaint, symptom description, medical history information, and previous treatment records. These data need to be processed through structuring to be converted into a standard data format that can be processed by machines.

[0062] 1.2 Data Preprocessing

[0063] In this embodiment, the main task of the data preprocessing part is to perform standardization and enhancement processing on the collected CBCT image data and medical record information to meet the input requirements of subsequent deep learning models. The specific contents are as follows:

[0064] Preprocessing of image data:

[0065] Noise removal: The image data is denoised through Gaussian filtering or median filtering to remove random noise and artifacts generated during the imaging process and retain the anatomical detail features contained in the image. The Gaussian filter is used to smooth the background area, while the median filter preserves the edge features well.

[0066] Window width and window level adjustment: According to the diagnostic requirements, the window width and window level of the CBCT image are adjusted to highlight the target anatomical structure of the temporomandibular joint. For bone tissue, bone window parameters are selected; for the soft tissue area, the window width is adjusted to make the articular disc more clearly visible.

[0067] 3D Slicing: The CBCT image data is sliced from three-dimensional space into two-dimensional multi-channel data. The slicing directions include the coronal plane, sagittal plane, and transverse plane, and the slice thickness is controlled between 0.5 mm and 1.0 mm to ensure clear anatomical features of the target area. After slicing, the data is stored in the form of a sequence of images for subsequent model input.

[0068] Data Standardization:

[0069] Normalization Processing: Adjust the range of image pixel values to [0, 1] to eliminate the problem of inconsistent gray values caused by differences in imaging devices and parameters.

[0070] Size Adjustment: Adjust all slice data to a unified size (such as 256×256 pixels) to adapt it to the input format of the deep learning model, while keeping the image ratio unchanged to avoid loss of anatomical features caused by deformation.

[0071] Data Augmentation:

[0072] Geometric Transformation: Perform random rotation (angle range ±15°), mirror flipping, and scaling (scale range ±10%) on the image data to simulate different angle and scale changes that may occur during the image acquisition process and enhance the generalization ability of the model.

[0073] Random Noise Addition: Introduce slight random noise into some data to simulate interference in the actual imaging scenario and improve the robustness of the model on noisy data.

[0074] Medical Record Information Preprocessing:

[0075] Word Segmentation Processing: Perform word segmentation and annotation on the medical record text to decompose the unstructured text information into structured data with semantic labels.

[0076] Standardized Coding: Convert the symptom descriptions in the medical record (such as "joint pain", "limited mouth opening") into a fixed semantic coding format to ensure that the subsequent model can efficiently parse and utilize them.

[0077] Multi-channel Data Generation:

[0078] Fusion of Image and Label Information: Overlay the annotation data from the medical record information on the image data to generate multi-channel input. Each channel contains data information from different sources (such as anatomical structures, medical record labels, semantic encodings) to enable the subsequent model to fully learn the fusion features.

[0079] Through the data acquisition and preprocessing module, the standardized and normalized processing of CBCT image data and medical record information can be achieved, ensuring data consistency and quality. The specific implementation of this module includes technical details such as noise removal, image slicing, normalization, data augmentation, and text semantic processing, aiming to provide high-quality input data for the subsequent feature extraction and analysis module, and solving the problems of uneven data quality, large noise interference, and non-fusion of multi-source data in traditional methods.

[0080] 2. Feature Extraction and Analysis Module

[0081] For the feature extraction and analysis module, this module is the core component of the DC diagnosis system for temporomandibular joint diseases based on CBCT images. Its role is to extract the anatomical features of the temporomandibular joint from the preprocessed CBCT image data, providing basic support for disease classification and diagnosis. Through the optimized YOLOv8 object detection model and multi-scale feature extraction technology, this module can effectively capture the geometric features and positional relationships of the key structures of the temporomandibular joint (such as the condyle, articular disc, and glenoid fossa), and enhance the feature extraction ability of the key areas of the temporomandibular joint through the channel attention mechanism and spatial attention mechanism. This module is directly related to the subsequent multi-modal data fusion module, and the feature extraction results will be used as input for feature fusion with medical record information.

[0082] 2.1 Core Technologies of the Feature Extraction and Analysis Module

[0083] In this embodiment, the feature extraction and analysis module includes four parts: object detection, feature extraction, multi-scale analysis, and feature fusion, which respectively realize the localization of key areas in CBCT images, the capture of fine-grained features, and the integration of multi-level features.

[0084] 2.2 Object Detection

[0085] In this embodiment, the object detection task is completed through the optimized YOLOv8 object detection algorithm. The specific optimization points are as follows:

[0086] Network Structure Optimization:

[0087] Introduce the CSP (Cross Stage Partial) structure to separate shallow and deep features, and improve the feature extraction ability of the model through partial residual connections while reducing computational overhead.

[0088] Use PANet (Path Aggregation Network) to enhance the fusion ability of features at different levels for more accurate detection of complex anatomical structures.

[0089] Training Optimization:

[0090] Adopt a multi-scale anchor box strategy, set the anchor box sizes suitable for small target detection of the temporomandibular joint, and ensure that small target areas such as the condyle and articular disc can be fully detected.

[0091] Use an improved IoU (Intersection over Union) loss function to optimize the positioning accuracy of the target box, especially the spatial position relationship between the condyle and the articular disc.

[0092] Detection targets:

[0093] The three-dimensional geometric shape of the condyle;

[0094] The displacement direction of the articular disc;

[0095] The integrity and morphological characteristics of the glenoid fossa.

[0096] 2.3 Feature extraction

[0097] In this embodiment, the feature extraction task is completed by the backbone network in YOLOv8, and the following types of key features are extracted:

[0098] Global anatomical features: Extract the overall geometric structure of the temporomandibular joint region through the main convolutional layer, including the global spatial morphology of the bony joint and the articular disc.

[0099] Local detail features: Capture the edge features in the image through the shallow convolutional layer, and enhance the detection ability of fine-grained information such as the contour of the articular disc and the surface morphology of the condyle.

[0100] Semantic features: Extract high-level features through the deep convolutional layer, distinguish between normal joints and diseased areas, and generate semantic information suitable for DC classification.

[0101] 2.4 Multi-scale feature analysis

[0102] In this embodiment, the multi-scale feature analysis task is implemented through the Feature Pyramid Network (FPN), specifically including:

[0103] High-resolution feature extraction: Enhance the shallow feature map for detecting small target structures (such as the articular disc) in the temporomandibular joint.

[0104] Low-resolution feature extraction: Integrate the deep feature map for capturing the overall anatomical information of the image and assisting in global analysis.

[0105] Feature enhancement: Use the Laplacian pyramid structure to fuse features of different resolutions, and at the same time combine the upsampling and downsampling paths to enhance the model's understanding ability of complex anatomical structures.

[0106] 2.5 Attention mechanism

[0107] In this embodiment, to improve the model's attention to the target area, a channel attention mechanism and a spatial attention mechanism are introduced:

[0108] Channel attention mechanism: By weighting the channel features, it enhances the feature extraction ability for key areas (such as the condyle and articular disc), and suppresses the interference of background noise on detection.

[0109] Spatial attention mechanism: Through the spatial weight distribution, it highlights the spatial information of the key parts of the temporomandibular joint, ensuring clear feature expression of the lesion area.

[0110] 2.6 Feature fusion

[0111] In this embodiment, the extracted multi-scale features are fused, and the specific steps are as follows:

[0112] Feature integration: The Path Aggregation technology is adopted to fuse the high-resolution detailed features with the low-resolution semantic features.

[0113] Feature output: Generate a comprehensive feature map for disease classification and localization, providing high-quality input for the subsequent multi-modal data fusion module.

[0114] The feature extraction and analysis module completes the detection and analysis of the key structures of the temporomandibular joint in CBCT images through the optimized YOLOv8 algorithm. It uses the multi-scale feature extraction technology to comprehensively capture the global anatomical information and local detailed information, and strengthens the model's attention to small targets through the attention mechanism. The features output by the module contain complete anatomical information and lesion area identification, laying a foundation for disease classification and diagnosis.

[0115] 3. Multi-modal data fusion module

[0116] For the multi-modal data fusion module, which is an important part of the temporomandibular joint disease DC diagnosis system based on CBCT images, it is responsible for fusing the anatomical features in CBCT image data with the semantic features in the medical record text information. The module aims to comprehensively utilize the information advantages of different modal data to provide a more comprehensive feature expression for subsequent disease classification and lesion site localization. Through feature extraction, alignment, and fusion of image and text data, the module can effectively integrate multi-source information, solve the problem of feature loss that may be caused by single-modal data, and provide support for accurate disease diagnosis.

[0117] 3.1 Feature extraction

[0118] In this embodiment, the multi-modal data fusion module first extracts features from the image data and text data respectively:

[0119] Image feature extraction

[0120] Input: The multi-scale feature maps output from the feature extraction and analysis module, including geometric features, texture features, and spatial relationships of the key regions of the temporomandibular joint (such as the condyle, articular disc, and glenoid fossa).

[0121] Processing: Extract image features through the deep network of the YOLOv8 model, encode the anatomical structure information into a multi-dimensional vector representation for subsequent fusion steps.

[0122] Output: The image features include the geometric shape, texture distribution, and spatial location of the lesion site, providing a physical basis for disease classification and localization.

[0123] Text Feature Extraction

[0124] Input: Structured medical record text data, including patient chief complaints, symptom descriptions, and previous diagnosis records.

[0125] Processing: Use a bidirectional long short-term memory network (Bi LSTM) to perform semantic analysis on the text, extract high-level semantic features related to diseases. At the same time, perform semantic encoding on the medical record text to convert natural language into a fixed-length vector representation.

[0126] Output: The text features include the patient's symptom information, disease course description, and subjective diagnostic clues from doctors, providing complementary semantic information for the image data.

[0127] 3.2 Feature Alignment

[0128] In this embodiment, the image features and text features are fused through cross-modal alignment technology. Specifically, it includes the following steps:

[0129] Feature Standardization

[0130] To achieve the alignment of image features and text features, first perform standardization processing on the two types of features. The image features are unified in scale through a normalization method, and the text features are transformed into the same dimension as the image features through a semantic embedding layer.

[0131] Alignment Mechanism

[0132] Use the Transformer network to align the image and text features. The Transformer analyzes the correlation between the two types of features through the attention mechanism and generates fusion weights.

[0133] During the alignment process, use semantic clues in the medical record text (such as "difficulty in opening the mouth", "joint noise") to establish a correspondence with the anatomical structure lesion areas detected in the image features.

[0134] 3.3 Feature Fusion

[0135] In this embodiment, the aligned features are integrated into unified diagnostic features through a fusion strategy, which is specifically as follows:

[0136] Weighted fusion

[0137] In the fusion stage, the weights of the image features and text features are dynamically adjusted through reinforcement learning technology to ensure that the fused features contain the most relevant information. The weight of the image features is determined by the importance of its lesion area, and the weight of the text features is determined by the strong correlation of semantic clues.

[0138] Fusion feature generation

[0139] The weighted features are further processed through a fully connected layer to generate a unified fused feature vector. The fused feature vector synthesizes the spatial features in the image and the semantic information of the medical record text, providing multi-dimensional input for disease classification and lesion localization.

[0140] 3.4 Output and subsequent processing

[0141] In this embodiment, the fused feature vector is output to the DC classification and diagnosis module for disease type classification, lesion area localization, and severity grading. The fused features can make up for the deficiencies of single-modal data and improve the accuracy and robustness of classification and localization.

[0142] The multi-modal data fusion module fully utilizes the information advantages of multi-modal data through the extraction, alignment, and fusion of CBCT image features and medical record text features. The output of the module provides high-quality feature expressions for subsequent disease classification and localization, solves the limitations of single-modal data in complex disease diagnosis scenarios, and provides reliable technical support for the accurate diagnosis of temporomandibular joint diseases.

[0143] 4. DC classification and diagnosis module

[0144] For the DC classification and diagnosis module, this module is the core part of the DC diagnosis system for temporomandibular joint diseases based on CBCT images, and is used to input the data after feature extraction and multi-modal fusion for disease type classification, lesion location, and grading diagnosis of disease severity. The module is designed based on the DC classification criteria (Diagnostic Criteria for Temporomandibular Disorders), combines deep learning models to complete the automatic classification and grading of diseases, and generates a standardized diagnostic report for doctors to refer to. Through this module, the system realizes the accurate distinction of disease types and the efficient identification of lesion areas, and further provides personalized treatment suggestions for clinical practice.

[0145] 4.1 Disease type classification

[0146] In this embodiment, the disease type classification classifies the fused features through a deep learning model, which specifically includes the following steps:

[0147] Classification criteria

[0148] According to the DC classification criteria of temporomandibular joint diseases, the diseases are divided into the following main types:

[0149] Muscle diseases: including masticatory muscle pain and myofascial pain.

[0150] Disc displacement: including disc displacement with reduction and disc displacement without reduction.

[0151] Osteoarthritis: including osteoarthritis, osteopathic lesions and other osteoarthropathy-related problems.

[0152] Model design

[0153] Use a multi-layer fully connected neural network (MLP) to complete the classification task of the fused features. The input layer receives the feature vector output from the multi-modal fusion module, the hidden layer performs non-linear transformation through the ReLU activation function, and the output layer uses the Softmax function to achieve multi-class probability prediction.

[0154] Optimization strategy

[0155] Adopt the cross-entropy loss function to measure the deviation between the model classification result and the actual class label.

[0156] Introduce a classification weight balancing strategy to solve the class imbalance problem caused by insufficient data in some classes.

[0157] Output result

[0158] The output of the classification model is the probability value of each disease type, and the class with the highest probability is selected as the final classification result.

[0159] 4.2 Lesion area localization

[0160] In this embodiment, the lesion area localization completes the identification of the lesion site of the temporomandibular joint anatomical structure based on the fused features, and the specific implementation is as follows:

[0161] Target site

[0162] Condyle: Locate whether there is bone destruction or deformation of the condyle.

[0163] Disc: Detect the displacement direction and position change of the disc.

[0164] Fossa: Evaluate whether there is a structural change in the fossa.

[0165] Localization method

[0166] Use a deep learning model based on improved IoU (Intersection over Union) to localize the lesion area.

[0167] The model outputs the coordinate frame of the lesion area, and the coordinate frame contains the boundary information of the lesion area (such as the center point, width, and height).

[0168] Optimization details

[0169] Add an IoU optimization term to the loss function to ensure that the model pays more attention to the localization accuracy during training.

[0170] Use data augmentation techniques to improve the adaptability of the model to localize the lesion area under different imaging conditions.

[0171] Output results

[0172] Output the spatial position and related structural features of the lesion site, and the results are represented in the form of a coordinate frame and transmitted to the diagnostic report generation unit.

[0173] 4.3 Severity grading

[0174] In this embodiment, the severity grading is based on the classification and localization results, and the grading diagnosis is completed by combining the lesion area characteristics, as follows:

[0175] Grading criteria

[0176] According to the area, volume, and texture characteristics of the lesion area, the disease severity is divided into mild, moderate, and severe.

[0177] Mild: The lesion range is limited to a small area, and there is no significant structural damage;

[0178] Moderate: The lesion range expands, and there are certain structural changes;

[0179] Severe: The lesion range is extensive, accompanied by obvious anatomical structure damage.

[0180] Model implementation

[0181] Use a support vector regression (SVR) model to perform regression analysis on the features, generate the score of the disease severity, and complete the grading by combining the preset threshold.

[0182] Data input

[0183] The input data of the grading model includes the area, volume, density, and texture characteristics of the lesion area.

[0184] Output results

[0185] Output the severity level of the disease and the corresponding score, providing a reference for the treatment plan.

[0186] 4.4 Diagnostic Report Generation

[0187] In this embodiment, the diagnostic report generation unit automatically generates a standardized report based on the classification, localization, and grading results. The specific steps are as follows:

[0188] Report Content

[0189] Disease type: Based on the output of the classification model, record the specific disease type.

[0190] Lesion location: Mark the spatial location and characteristics of the lesion according to the localization result.

[0191] Severity: Mark the severity of the disease according to the grading result.

[0192] Recommended treatment plan: Generate preliminary treatment suggestions (such as non-surgical treatment or surgical intervention) by combining the disease type and severity.

[0193] Output Format

[0194] The diagnostic report is stored in the form of an electronic document, supporting the output formats of PDF and HL7 interface for electronic medical records, which is convenient for docking with the hospital information system.

[0195] The DC classification and diagnosis module realizes the comprehensive diagnosis of temporomandibular joint diseases, covering disease type classification, lesion location, and severity grading. Based on the deep learning model and DC classification criteria, the module can generate high-precision classification and localization results, and at the same time provide standardized diagnostic reports and treatment suggestions for clinicians. Through this module, the system significantly improves the diagnostic efficiency and accuracy of temporomandibular joint diseases.

[0196] 5. Model Optimization and Deployment Module

[0197] For the model optimization and deployment module, this module is a key technical component of the DC diagnostic system for temporomandibular joint diseases based on CBCT images, aiming to improve the performance of the deep learning model and meet the actual needs in clinical applications. Through technologies such as transfer learning, model lightweighting, and optimization of training strategies, the module effectively solves the efficiency problems caused by insufficient data scale, high model complexity, and computational resource limitations. At the same time, the module supports the deployment of the model on embedded devices and edge computing devices to meet the clinical requirements of low latency and efficient diagnosis. This module is directly related to the aforementioned data processing and classification diagnosis modules, and its optimization process ensures the accuracy of diagnosis and the real-time performance of the system.

[0198] 5.1 Model Optimization

[0199] In this embodiment, model optimization includes three parts: transfer learning, model lightweighting processing, and optimization of training strategies:

[0200] Transfer learning

[0201] Base model selection: Select the YOLOv8 model pre-trained on public medical image datasets (such as ChestX-ray or NIH Lung Image Dataset) as the base model, and utilize the general visual features it has learned for transfer learning.

[0202] Local data fine-tuning: Apply the pre-trained model to CBCT image data, and further fine-tune it in the local training set to adjust the model parameters to adapt to the specific detection task of temporomandibular joint diseases.

[0203] Few-shot optimization: Adopt the strategy of freezing the shallow weights of the pre-trained model and only training the deep feature extraction part to reduce the overfitting risk in the few-shot case.

[0204] Model lightweight processing

[0205] Pruning: Perform structural pruning on the trained model, remove redundant network connections and unimportant feature channels, reduce the model's computational complexity, and at the same time maintain the core feature extraction ability.

[0206] Distillation: Adopt model distillation technology to transfer the knowledge of a complex model with high performance (teacher model) to a lightweight model (student model) to maintain the diagnostic accuracy while simplifying the model.

[0207] Quantization: Quantize the model parameters, convert the floating-point weights to a low bit width (such as INT8 format), significantly reduce the model's storage requirements and computational complexity, and adapt to the resource limitations of embedded devices.

[0208] Training strategy optimization

[0209] Learning rate scheduling: Adopt the learning rate warm-up and cosine annealing scheduling strategies to improve the convergence speed at the beginning of training and stabilize the model performance by gradually reducing the learning rate in the later stage.

[0210] Mini-batch training: Use the strategy of dynamic batch size, and utilize a larger batch for training within the allowable range of resources to improve the stability and accuracy of the model.

[0211] Regularization processing: Suppress overfitting through L2 regularization and Dropout techniques to enhance the generalization ability of the model.

[0212] 5.2 Model deployment

[0213] In this embodiment, the model deployment part is responsible for integrating the optimized deep learning model into clinical applications, specifically including the following steps:

[0214] Deployment environment preparation

[0215] Edge computing device: Select an edge device with high-performance computing capabilities (such as NVIDIA Jetson Nano or Raspberry Pi 4) to meet the low-latency diagnosis requirements.

[0216] Containerized deployment: Use Docker container technology to encapsulate the model into a container image that is convenient for migration and expansion, ensuring the consistency of the deployment environment.

[0217] Model compression and export

[0218] Export format: Convert the optimized model to ONNX (Open Neural Network Exchange) or TensorRT format for efficient operation on embedded and edge computing devices.

[0219] Compression and packaging: Combine weight quantization and weight chunking techniques to further compress the model file size.

[0220] Real-time inference

[0221] Inference framework integration: Integrate a deep learning inference framework (such as TensorFlow Lite or ONNX Runtime) on the edge device to achieve real-time inference of CBCT images.

[0222] Multithreaded processing: Enable a multithreaded inference mechanism to improve the inference efficiency and response speed of CBCT images.

[0223] System integration and testing

[0224] Interface with the hospital information system: Integrate with the hospital informatization system through standardized interfaces (such as HL7 or FHIR protocols) to achieve real-time synchronization and storage of diagnosis results.

[0225] Stability testing: Conduct load testing and stability testing in the edge device environment on the deployed system to ensure that the system can operate stably under high concurrency.

[0226] The model optimization and deployment module optimizes the model performance through techniques such as transfer learning, pruning, distillation, and quantization, and uses containerized deployment and inference framework integration to achieve real-time inference capabilities. The module design ensures the efficient operation of the system in resource-constrained scenarios, provides reliable support for the intelligent diagnosis of temporomandibular joint diseases, and meets the needs of clinical practical applications.

[0227] Working principle: The input data is standardized and enhanced by the data acquisition and preprocessing module to ensure the quality of CBCT images and the structured expression of medical record texts. Then, the feature extraction and analysis module uses the optimized YOLOv8 object detection algorithm and multi-scale feature extraction technology to extract the geometric, texture, and spatial features of the temporomandibular joint from the image data. Next, the multi-modal data fusion module aligns and weights the fusion of image features and semantic features in the medical record text to generate a unified feature representation. Further, the DC classification and diagnosis module classifies the disease types according to the DC classification standard based on the fusion features, locates the lesion area, grades the disease severity, and generates a standardized diagnosis report. Finally, the model optimization and deployment module optimizes the model performance through techniques such as transfer learning, pruning, and quantization, and deploys the model to edge devices to achieve real-time diagnosis. The entire system forms a closed loop from data acquisition to diagnosis report generation, making full use of the information advantages of multi-modal data and realizing the efficient and accurate diagnosis of temporomandibular joint diseases.

[0228] Example 1:

[0229] Scenario description

[0230] This example describes a temporomandibular joint disease (DC) diagnosis system based on CBCT (cone beam computed tomography) images, which is deployed in the clinical diagnosis process of the stomatology department of a hospital. The stomatology department of the hospital has received a large number of patients, especially those with symptoms of temporomandibular joint (TMJ) diseases, such as joint pain and limited mouth opening. The goal of the system is to improve the disease diagnosis efficiency through automated technology, reduce the workload of doctors, and improve the accuracy of diagnosis.

[0231] 1. Data acquisition and preprocessing

[0232] In the stomatology department of the hospital, Mr. Zhang came to see a doctor, complaining of pain in the right temporomandibular joint and limited mouth opening. The doctor arranged a CBCT scan for him to obtain detailed joint image data.

[0233] Step 1: Image acquisition

[0234] The image acquisition module obtains the three-dimensional image data of Mr. Zhang's temporomandibular joint from the hospital's CBCT scanner. After scanning, the image data is transmitted to this diagnosis system through the hospital information system (HIS) for processing.

[0235] Step 2: Image preprocessing

[0236] Through the data preprocessing module in the system, the imaging data is first denoised to remove random noise during the scanning process, ensuring image clarity. Then, window width and window level adjustments are performed to ensure that the gray-scale differences between bone tissue and soft tissue are suitable for doctors' diagnostic needs.

[0237] In this step, the system also standardizes the imaging data, including unifying the image size and normalizing the pixel values to ensure the consistency and comparability of all imaging data.

[0238] 2. Feature Extraction and Target Detection

[0239] Step 3: Target Region Detection

[0240] The system uses an optimized version of the YOLOv8 target detection algorithm to automatically detect key structures within the temporomandibular joint region, including the condyle, articular disc, and glenoid fossa, etc. Through multi-scale feature extraction, the system can simultaneously detect different parts of the joint in images of different sizes, regardless of whether the diseased area of the joint is significant.

[0241] Step 4: Feature Analysis and Localization

[0242] During the feature extraction process, the system combines spatial and channel attention mechanisms. By focusing on the joint region in the image, the system enhances its ability to detect diseased areas. For Mr. Zhang's image, the system accurately identified the deformed areas of the condyle and articular disc, suggesting that he may have arthritis or joint degenerative diseases.

[0243] 3. Multi-modal Data Fusion and Diagnosis

[0244] Step 5: Medical Record Information Fusion

[0245] The system also extracts relevant clinical information of Mr. Zhang from the hospital's medical record information system, including data such as gender, age, chief complaint, and past medical history. Through the multi-modal data fusion module, the system combines these text data with the imaging data to generate a unified feature representation for further disease diagnosis.

[0246] Step 6: Disease Classification and Diagnosis

[0247] The system classifies and grades temporomandibular joint diseases based on the fused features. In Mr. Zhang's case, the system, according to the imaging features and medical record information, determines that he may have degenerative joint disease (e.g., degenerative changes of the articular disc and osteophytes of the condyle), and grades the lesion, marking the severity as moderate.

[0248] The system also automatically locates the diseased area and generates the following diagnostic report:

[0249] Disease Type: Temporomandibular Joint Degenerative Lesion

[0250] Lesion location: right condyle, articular disc

[0251] Lesion grade: moderate degenerative changes

[0252] Suggested treatment: Physical therapy is recommended, and arthroscopic examination or surgical treatment may be considered if necessary.

[0253] 4. Model Optimization and Real-time Deployment

[0254] Step 7: Transfer Learning and Fine-tuning

[0255] Since the hospital has accumulated a large amount of CBCT image data, the system uses transfer learning technology to fine-tune the model pre-trained on the public dataset. By training on the local dataset, the performance of the model on a specific patient population is optimized.

[0256] Step 8: Edge Device Deployment

[0257] The optimized model is deployed to the portable edge devices within the hospital. Doctors can quickly load the imaging data of patients through this device for real-time diagnosis. For Mr. Zhang, the doctor only needs to scan his CBCT images with a handheld device, and the system can complete the automatic diagnosis and generate a report within a few minutes.

[0258] 5. Clinical Feedback and Optimization

[0259] Mr. Zhang's diagnostic report was quickly sent to the stomatologist. The doctor carried out clinical treatment according to the treatment suggestions provided by the system and decided to perform a series of physical therapy programs. In the subsequent follow-up visits, the doctor highly evaluated the diagnostic accuracy and efficiency of the system, believing that the system has greatly improved the diagnostic efficiency and reduced the workload of doctors.

[0260] Through the description of this embodiment, the temporomandibular joint disease diagnosis system based on CBCT images has been successfully applied in the clinical environment of the hospital. The system can not only automatically and quickly complete image data processing and disease diagnosis, but also combine the clinical medical record information of patients to provide more accurate diagnostic results. The application of this system has significantly improved the diagnostic efficiency, reduced the burden on doctors, and at the same time improved the diagnostic accuracy, providing more personalized treatment plans for patients.

[0261] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A DC diagnosis system for temporomandibular joint disease based on CBCT images, characterized in that: include: The data acquisition and preprocessing module is used to collect the patient's CBCT image data and medical history information, and to perform denoising, window width and window position adjustment, normalization and multi-plane slicing on the CBCT image data to generate multi-channel two-dimensional data; Wherein, the denoising process removes noise in the image by Gaussian filtering or median filtering to improve the quality of subsequent image processing; The feature extraction and analysis module is used to perform target detection on the preprocessed CBCT image data based on the optimized YOLOv8 target detection algorithm, extract multi-scale features through the feature pyramid network, and enhance the feature extraction capability of the key area of ​​the temporomandibular joint through the channel attention mechanism and the spatial attention mechanism; Multimodal data fusion module, used to extract anatomical features from CBCT image data, extract text semantic features from medical record information, and generate unified fusion features through cross-modal feature alignment; DC classification and diagnosis module, which is used to classify the types of temporomandibular joint diseases based on fusion features, locate the lesions, grade the severity of the diseases, and output classification, location and grading results; The model optimization and deployment module is used to optimize the model through transfer learning technology, reduce the model calculation complexity through model lightweight technology, and deploy the optimized model to edge devices.

2. A CBCT image-based DC diagnosis system for temporomandibular joint disease according to claim 1, characterized in that: The data acquisition and preprocessing module includes: An image reading unit, used for extracting CBCT image data from DICOM format files; An image denoising unit, used to remove noise in the image by Gaussian filtering or median filtering; An image adjustment unit, used to adjust the window width and window level of CBCT image data based on diagnostic requirements; An image slicing unit, used for slicing the three-dimensional CBCT image data into two-dimensional data of the coronal plane, sagittal plane and transverse plane; A data normalization unit, used to normalize the pixel values ​​of the image data and adjust the image size; The data enhancement unit is used to randomly rotate, flip, scale and add random noise to the image data.

3. The DC diagnosis system for temporomandibular joint disease based on CBCT images according to claim 1, characterized in that: The feature extraction and analysis module includes: A detection unit, used for detecting the condyle, articular disc and fossa in the CBCT image data based on an optimized YOLOv8 algorithm; A multi-scale feature extraction unit, used for extracting multi-scale features from CBCT image data through a feature pyramid network; Attention enhancement unit, used to enhance the detection capability of the feature extraction model for the key areas of the temporomandibular joint through channel attention mechanism and spatial attention mechanism; A feature fusion unit is used to fuse the features generated by the up-sampling and down-sampling paths to improve the detection accuracy of the target area; Among them, the channel attention mechanism improves the representation ability of important features by weighting feature channels; and the spatial attention mechanism improves the recognition accuracy of key areas by weighting feature space areas.

4. The DC diagnosis system for temporomandibular joint disease based on CBCT images according to claim 1, characterized in that: The multimodal data fusion module includes: An image feature extraction unit, used for extracting anatomical features related to the temporomandibular joint from the CBCT image data; The text feature extraction unit is used to extract the semantic features of patients from the medical record text information; the cross-modal alignment unit is used to embed image features and text features into a unified semantic space through a deep learning model, or to achieve feature alignment through a specific algorithm; The fusion feature generation unit is used to generate fusion features through an optimization algorithm for subsequent classification and diagnosis.

5. The DC diagnosis system for temporomandibular joint disease based on CBCT images according to claim 1, characterized in that: The DC classification and diagnosis module includes: Disease classification unit, used to classify the types of temporomandibular joint diseases according to the DC classification criteria; Lesion grading unit, used to grade the disease into mild, moderate and severe according to the severity of disease lesions; A lesion localization unit is used to locate the lesion site of the temporomandibular joint by using an optimized IoU algorithm; The diagnostic report generation unit is used to output classification results, grading results, lesion location and treatment recommendations.

6. The DC diagnosis system for temporomandibular joint disease based on CBCT images according to claim 1, characterized in that: The model optimization and deployment module includes: The transfer learning submodule is used to pre-train the model on public medical imaging datasets and fine-tune the model weights through local CBCT datasets; Model lightweight submodule, which is used to reduce the computational complexity of the model through model pruning, distillation and quantization techniques; Parameter optimization submodule, which is used to optimize the model training strategy using learning rate warm-up and cosine annealing scheduling techniques; The deployment submodule is used to deploy the optimized model to portable edge devices and support real-time inference.

7. The DC diagnosis system for temporomandibular joint disease based on CBCT images according to claim 1, characterized in that: The data acquisition and preprocessing module automatically adjusts the slice direction and thickness of the CBCT image to ensure that the image data contains the characteristic information of the condyle, articular disc and fossa.

8. The DC diagnosis system for temporomandibular joint disease based on CBCT images according to claim 1, characterized in that: The feature extraction and analysis module shares the feature weights of the target detection task and the classification task through a joint learning strategy, and optimizes the model performance through dynamic parameter adjustment.

9. The DC diagnosis system for temporomandibular joint disease based on CBCT images according to claim 1, characterized in that: The diagnosis results include disease type, disease severity, lesion location and treatment recommendations. The diagnosis results can be output to the electronic medical record system through the system interface.

10. The DC diagnosis system for temporomandibular joint disease based on CBCT images according to claim 1, characterized in that: The system achieves collaborative diagnosis of image data and text data through a multimodal data fusion module, and supports real-time analysis in low-computing power devices.

Citation Information

Cited By

  • Oral and maxillofacial surgery image recognition and diagnosis method and system based on deep learning

    CN120280132A

  • Joint cavity positioning method and device based on deep learning model

    CN120318329A

  • A joint cavity positioning method and device based on deep learning model

    CN120318329B

  • Temporomandibular joint three-dimensional reconstruction system based on multi-modal image fusion

    CN120876747A

  • Pulmonary nodule benign and malignant identification and prediction system based on multi-modal feature fusion

    CN121617603A