Orthodontic abnormity diagnosis intelligent prompting system based on deep learning

Through deep learning technology, the multimodal data is integrated, unified feature vectors are generated and diagnostic tips are provided, which solves the problem of difficult multimodal data in existing systems, and realizes efficient and accurate orthodontic diagnosis and doctor-assisted diagnosis.

CN120299685APending Publication Date: 2025-07-11AFFILIATED STOMATOLOGICAL HOSPITAL OF NANJING MEDICAL UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510461106.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing orthodontic diagnostic system cannot effectively integrate multimodal data, resulting in high diagnostic complexity, high misdiagnosis rate, long accumulation of doctor experience, uneven distribution of medical resources, and difficult to ensure the diagnosis quality of young doctors.

Method used

The intelligent prompt system for orthodontic abnormality diagnosis based on deep learning is adopted. Through multimodal data acquisition, preprocessing, fusion and intelligent diagnosis modules, a unified feature vector is generated, and combined with the decision support module, the precise alignment and joint modeling of multimodal data are realized, providing diagnostic results and difference prompts.

Benefits of technology

It improves the comprehensiveness and accuracy of diagnosis, reduces misdiagnosis and missed diagnosis, reduces the work burden of doctors, promotes coordinated diagnosis between doctors and systems, and improves diagnosis consistency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299685A_ABST
    Figure CN120299685A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of orthodontic abnormality diagnosis, in particular to an orthodontic abnormality diagnosis intelligent prompt system based on deep learning, which comprises a multi-modal data acquisition module used for transmitting a patient ID (Identity) through an API (Application Program Interface), automatically importing image data, image data, facial three-dimensional data and dentition three-dimensional data of a patient, and transmitting the image data to the multi-modal data acquisition module; meanwhile, medical record text information is obtained in real time through an EMR interface; according to the method, collected data are diversified in form and comprehensive in content, anatomical structure information is provided for patient photos, skull side position X-rays, CBCT, 3dMD, IOS, medical record reports, image data, image data, three-dimensional data and the like, symptom description and medical history are supplemented through text data, multi-person data fusion is achieved, potential association which is difficult to discover through single-mode data can be captured, facial features of a patient are comprehensively evaluated, and the method is suitable for clinical application. The comprehensiveness and the accuracy of diagnosis are improved, more abundant diagnosis basis can be provided for complex mismatching deformity and multi-modal data fusion, and missed diagnosis and misdiagnosis are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of orthodontic anomaly diagnosis, and particularly to an intelligent prompt system for orthodontic anomaly diagnosis based on deep learning. Background Art

[0002] In 1956, a number of scientists studied and discussed a series of issues of using machines to simulate human intelligence, and at the same time proposed the concept of artificial intelligence. Thus, artificial intelligence was formally born as a new discipline. The application of artificial intelligence in the field of orthodontics started relatively early. For example, multiple linear regression can be considered a generalized artificial intelligence method. Kim and Vietas, and Kim used statistical regression algorithms to judge the degree of vertical and sagittal dysplasia in patients. It was not until after 2010 that methods such as deep learning algorithms began to be applied to the field of artificial intelligence in orthodontics. At present, the research on artificial intelligence in the field of orthodontics basically covers the processes of diagnostic analysis and treatment decision-making. However, in the field of orthodontics, relatively few high-quality labeled data are available, which limits the performance and generalization ability of deep learning models. Insufficient sample size may lead to overfitting of the model, reducing its accuracy and reliability in actual clinical applications. Deep learning is mainly used in aspects such as image analysis, diagnostic assistance, and preliminary design of treatment plans. However, in complex clinical decision-making and the formulation of personalized treatment plans, the professional knowledge and experience of doctors are still required for comprehensive judgment.

[0003] The prototype of multimodal fusion appeared in the military and industrial fields, mainly for the integration of sensor data such as radar and infrared. With the rise of deep learning, multimodal fusion entered the algorithm innovation period. The combination of convolutional neural network (CNN) and recurrent neural network (RNN) made it possible to jointly model images and texts. In the field of orthodontic diagnosis and treatment, multimodal data fusion technology is gradually moving from the laboratory to the clinic. Its core lies in integrating multi-source heterogeneous data such as images and texts. By fusing data of different modalities, the occlusal conditions of patients can be more comprehensively reflected, providing more accurate information for the formulation of diagnostic and treatment plans. In the process of multimodal data fusion, partial information loss may occur. Data of different modalities have different spatial and temporal resolutions, making it difficult to perform precise alignment, such as the problem of spatio-temporal dimension mismatch. The anatomical landmark points on two-dimensional cephalometric radiographs need to be non-rigidly registered with three-dimensional CBCT data. However, due to the different imaging angles and ranges of two-dimensional X-ray films and three-dimensional CT images, traditional affine transformation will cause local deformation errors.

[0004] In the current era of rapid development of medical informatization, various systems within a hospital should enhance the quality and efficiency of medical services through data sharing and integration. However, in many hospital systems, imaging systems (such as lateral cephalometric X-rays and CBCTs) are independent of the electronic medical record system, resulting in poor integration of patient information and data isolation. Traditional orthodontic diagnosis relies on doctors manually comparing images and medical record materials. When conducting orthodontic diagnosis, doctors need to switch between multiple interfaces, manually extract and compare data, which is time-consuming and prone to missing key information, leading to fragmented analysis of image features and text descriptions and increasing the complexity of diagnosis.

[0005] The current challenges in the field of orthodontic diagnosis and treatment lie in the dual dilemmas of the long cycle for doctors to accumulate clinical experience and the uneven distribution of medical resources. In China, there are fewer than 0.7 orthodontic doctors per 100,000 people, which results in a long case accumulation cycle for doctors, heavy medical burdens. At the same time, young doctors need to undergo long-term systematic training to independently handle routine cases, and it takes even longer case accumulation to master the diagnosis and treatment capabilities of complex cases. In the current industry situation of high diagnostic requirements and a large patient base, young doctors may have subjective errors due to various reasons such as fatigue and lack of experience, affecting the quality of diagnosis and a high incidence of misdiagnosis.

[0006] In orthodontic clinical diagnosis and treatment, a patient's teeth, jaws, joints, and soft tissues together form a dynamic and complex biomechanical system. Single-dimensional data analysis often fails to fully reveal its internal pathological mechanism. Diagnosis based solely on single-modal data (such as images or text) lacks comprehensive analysis of multi-modal data and is difficult to comprehensively evaluate the coordination of hard and soft tissues in a patient's dentofacial deformity, prone to misdiagnosis or missed diagnosis. Currently, orthodontic intelligent diagnosis systems only support the processing of a single data type, such as image analysis or text mining, and cannot achieve fusion analysis of multi-modal data. However, the analysis of one data type cannot meet the needs of orthodontic diagnosis. The formats and dimensions of image data, graphic data, IOS data, and text data vary greatly, and it is difficult to align heterogeneous data. Existing technologies are difficult to achieve precise alignment and joint modeling. The lack of multi-modal fusion analysis limits the comprehensiveness and accuracy of orthodontic intelligent diagnosis systems in diagnosis.

[0007] To solve the above problems, the present invention proposes an intelligent prompt system for orthodontic anomaly diagnosis based on deep learning. Summary of the Invention

[0008] To overcome the problem of data isolation caused by poor integration of patient information, the present invention proposes an intelligent prompt system for orthodontic anomaly diagnosis based on deep learning.

[0009] The technical solution of the present invention is: an intelligent prompt system for orthodontic anomaly diagnosis based on deep learning, including:

[0010] Multimodal Data Acquisition Module: Used to automatically import the patient's imaging data, image data, three-dimensional facial data, and three-dimensional dental arch data by passing in the patient ID through the API interface, and at the same time, obtain the medical record text information in real time through the EMR interface;

[0011] Data Preprocessing Module: Used to perform modality-specific standardization, noise removal, and feature extraction on multimodal data to generate hard tissue feature vectors, soft tissue feature vectors, dental arch feature vectors, and text feature vectors;

[0012] Multimodal Data Fusion Module: Used to perform cross-modal fusion on feature vectors of different modalities through a multi-scale feature pyramid network to generate a unified 768-dimensional comprehensive feature vector;

[0013] Intelligent Diagnosis Module: Used to analyze the fused feature vectors based on a deep neural network and output diagnostic results including bony probability, Angle classification, and sagittal plane anomaly categories;

[0014] Decision Support Module: Used to quantify the differences between the intelligent diagnostic results and the physician's diagnostic results, and implement a hierarchical prompting strategy according to the severity of the differences.

[0015] Preferably, the multimodal data acquisition module specifically includes: docking with the hospital Dolphin system through the RESTful API to obtain lateral cephalometric X-ray films and CBCT data in DICOM format in real time; supporting the standardized import of 3dMD facial scan data and IOS intraoral scan data; collecting structured medical record texts from the hospital EMR system based on the HL7 standard, including medical history, symptom descriptions, and previous treatment records.

[0016] Preferably, the data preprocessing module includes:

[0017] Imaging Data Processing Unit: Using wavelet transform or 3D UNet for denoising, and locating landmark points on the lateral cephalometric film through a graph convolutional network, with the error controlled within 2 mm;

[0018] Image Data Processing Unit: Segmenting the facial area through Mask R-CNN and extracting geometric features in combination with 3D point cloud data;

[0019] Text Data Processing Unit: Extracting keywords and semantic features through the CRF model and the BERT pre-training model.

[0020] Preferably, the imaging data processing unit converts the pixel values into standard HU values by parsing DICOM metadata, and realizes 1 mm through trilinear interpolation 3Isotropic resolution, using wavelet threshold denoising to process mild artifacts, and a pre-trained 3D UNet model to process severe artifacts. Automatic landmark localization of lateral cephalograms is achieved based on a multi-scale graph convolutional network, with an average localization error within the clinically acceptable range of 2 mm. Finally, tooth candidate boxes are generated through a 3D ResNet-RPN network, and high-precision tooth instance segmentation is realized by combining Sobel edge enhancement and 3D UNet.

[0021] Preferably, the image data processing unit segments the facial region by adopting a Mask R-CNN model, calculates the facial proportion and symmetry index, optimizes the mesh quality through a Laplacian smoothing algorithm, calculates the soft tissue thickness, facial volume, and curvature features, and finally establishes the correspondence of anatomical landmark points between the 2D photo and the 3D model based on the PnP algorithm, and realizes the fusion of features through Cross-Modal Transformer.

[0022] Preferably, the text data processing unit filters non-standard characters through regular expressions, realizes the conversion from colloquial expressions to standard medical terms based on a term mapping table, combines a CRF model and rule-based word segmentation for entity recognition, extracts 500-dimensional sparse features through TF-IDF, also adopts Word2Vec to generate 300-dimensional word vectors, and combines a clinically dedicated BERT model to output 768-dimensional context semantic features.

[0023] Preferably, the multi-modal data fusion module includes: projecting text features to 768 dimensions through a fully connected layer and mapping them to spatial dimensions, performing 1×1×1 convolutions on high-level, middle-level, and low-level features respectively to unify the dimensions, realizing cross-level feature fusion through three linear interpolations and 3×3×3 convolutions, and finally splicing multi-scale features into a 768-dimensional final representation through global average pooling.

[0024] Preferably, the intelligent diagnosis module includes: mapping 768-dimensional features to a skeletal malocclusion probability of 0-1 through a Sigmoid function, and then parallelly outputting the logits of Angle classification and sagittal abnormality. Finally, multi-task outputs are dynamically weighted and fused according to the skeletal probability value, and the formula is: final_logits = (bone_prob × angle_logits) + ((1 - bone_prob) × sagittal_logits).

[0025] Preferably, the decision support module includes:

[0026] Difference detection unit: Quantify the difference by calculating the overlap degree of the confidence intervals between the intelligent diagnosis and the physician diagnosis.

[0027] Three - level prompting strategy: Trigger yellow, orange, or red warnings according to the CIO value, corresponding to non - blocking reminders, forced filling of reasons for differences, and superior review respectively.

[0028] Preferably, the three - level prompting strategy is specifically as follows:

[0029] CIO < 0.5: Yellow warning, remind the physician through a non - blocking pop - up window to review the key parameters;

[0030] CIO < 0.3: Orange warning, require the physician to fill in the reasons for differences and record remarks to optimize subsequent model training;

[0031] CIO < 0.1: Red warning, for high - risk differences, require the physician to verify the diagnosis result again and push it to the superior physician for review synchronously.

[0032] Advantages of the present invention:

[0033] 1. The forms of collected data are diverse and the content is comprehensive, including patient photos, lateral cephalometric X - rays, CBCT, 3dMD, IOS, medical record reports, image data, image data, three - dimensional data, etc., which provide anatomical structure information. Text data supplements symptom descriptions and medical histories. The fusion of multiple data sources can capture potential associations that are difficult to detect by single - modality data, comprehensively evaluate the patient's facial features, improve the comprehensiveness and accuracy of diagnosis. For complex malocclusions, the fusion of multi - modality data can provide richer diagnostic basis and reduce missed diagnoses and misdiagnoses.

[0034] 2. When there are significant differences between the intelligent diagnosis result and the physician's diagnosis, the system automatically triggers a difference alarm, generates a difference analysis report, and visually displays the divergence points through a heat map. The system dynamically adjusts the difference detection threshold according to the case complexity to ensure the accuracy and timeliness of the prompt. By quantitatively evaluating the diagnostic differences, the system can effectively reduce misjudgments caused by the physician's lack of experience or fatigue. The intelligent prompt function helps the physician quickly locate the divergence points, promotes the consistency of the diagnosis results, greatly reduces the subjective error of the physician's treatment decision - making. At the same time, it promotes the physician's autonomous learning and cultivates the orthodontic diagnosis and treatment thinking.

[0035] 3. Clinical physicians are the main body of the diagnosis and treatment process and need to exert professional autonomy in the diagnosis and treatment process. As an intelligent assistant for physicians, this system prompts abnormal diagnosis results, and the final decision - making power is still in the hands of the physicians, fully respecting the clinical experience and subjective judgment of the physicians, avoiding rigid decision - making caused by complete reliance on AI. Through automated processing and data integration, the system significantly reduces the physician's work burden, enabling them to focus on key decisions. The system gradually optimizes the diagnosis model by continuously learning the physician's feedback and corrections, realizing the co - evolution of man - machine, and improving the accuracy and practicality of diagnosis. Description of the Drawings

[0036] Figure 1 The figure shows a schematic diagram of the system architecture of an intelligent prompt system for orthodontic abnormality diagnosis based on deep learning according to the present invention. Detailed implementation manners

[0037] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0038] Please refer to Figure 1 , the present invention provides an embodiment: an intelligent prompt system for orthodontic abnormality diagnosis based on deep learning, including:

[0039] Multimodal data acquisition module: used to input the patient ID through the API interface, automatically import the patient's imaging data, image data, three-dimensional facial data, and three-dimensional dentition data, and at the same time obtain the medical record text information in real time through the EMR interface;

[0040] Data preprocessing module: used to perform modality-specific standardization, noise removal, and feature extraction on multimodal data to generate hard tissue feature vectors, soft tissue feature vectors, dentition feature vectors, and text feature vectors;

[0041] Multimodal data fusion module: used to perform cross-modal fusion of feature vectors of different modalities through a multi-scale feature pyramid network to generate a unified 768-dimensional comprehensive feature vector;

[0042] Intelligent diagnosis module: used to analyze the fused feature vectors based on a deep neural network and output a diagnosis result including the bony probability, Angle classification, and sagittal abnormality category;

[0043] Decision support module: used to quantify the difference between the intelligent diagnosis result and the physician's diagnosis result, and implement a hierarchical prompt strategy according to the severity of the difference.

[0044] The multimodal data acquisition module specifically includes:

[0045] Imaging data interface: dock with the hospital Dolphin system through the RESTful API to obtain the lateral cephalometric X-ray film and CBCT data in DICOM format in real time; support the standardized import of 3dMD facial scan data and IOS intraoral scan data; collect structured medical records from the hospital EMR system based on the HL7 standard, including medical history, symptom description, and previous treatment records.

[0046] The data preprocessing module includes:

[0047] Imaging data processing unit: adopt wavelet transform or 3D UNet for denoising, and locate the landmark points of the lateral cephalometric film through a graph convolutional network, with the error controlled within 2 mm;

[0048] Image data processing unit: Segment the facial area through Mask R-CNN and extract geometric features by combining 3D point cloud data;

[0049] Text data processing unit: Extract keywords and semantic features through the CRF model and the BERT pre-trained model.

[0050] The described imaging data processing unit converts pixel values to standard HU values by parsing DICOM metadata, and realizes 1mm 3 isotropic resolution, uses wavelet threshold denoising to process mild artifacts, uses a pre-trained 3D UNet model to process severe artifacts, realizes automatic localization of cephalometric landmarks based on a multi-scale graph convolutional network, with an average localization error within the clinically acceptable range of 2 millimeters, and finally generates tooth candidate boxes through a 3D ResNet-RPN network, and combines Sobel edge enhancement and 3D UNet to achieve high-precision tooth instance segmentation.

[0051] The described image data processing unit segments the facial area by using the Mask R-CNN model, calculates the facial proportion and symmetry index, optimizes the mesh quality through the Laplacian smoothing algorithm, calculates soft tissue thickness, facial volume and curvature features, and finally establishes the correspondence of anatomical landmarks between the 2D photo and the 3D model based on the PnP algorithm, and realizes the fusion of features through Cross-ModalTransformer.

[0052] The text data processing unit filters non-standard characters through regular expressions, realizes the conversion from colloquial expressions to standard medical terms based on a term mapping table, also combines the CRF model and rule-based word segmentation for entity recognition, extracts 500-dimensional sparse features through TF-IDF, also uses Word2Vec to generate 300-dimensional word vectors, and combines a clinical-specific BERT model to output 768-dimensional context semantic features.

[0053] The described multi-modal data fusion module includes: projecting text features to 768 dimensions through a fully connected layer and mapping them to spatial dimensions, performing 1×1×1 convolution on high-level, middle-level, and low-level features respectively to unify the dimensions, realizing cross-level feature fusion through trilinear interpolation and 3×3×3 convolution, and finally splicing multi-scale features into a 768-dimensional final representation through global average pooling.

[0054] The intelligent diagnosis module includes: mapping 768-dimensional features to a skeletal malocclusion probability between 0 and 1 through the Sigmoid function, then parallelly outputting the logits of Angle classification and sagittal abnormality, and finally dynamically weighted fusion of multi-task outputs according to the skeletal probability value. The formula is: final_logits = (bone_prob × angle_logits) + ((1 - bone_prob) × sagittal_logits).

[0055] The decision support module includes:

[0056] Difference detection unit: quantifying the difference by calculating the overlap degree of the confidence intervals between the intelligent diagnosis and the physician diagnosis;

[0057] Three-level prompt strategy: triggering yellow, orange, or red warnings according to the CIO value, corresponding to non-blocking reminders, forcing the filling of the difference reasons, and superior review respectively.

[0058] Preferably, the three-level prompt strategy is specifically as follows:

[0059] CIO < 0.5: yellow warning, reminding the physician through a non-blocking pop-up window to review the key parameters;

[0060] CIO < 0.3: orange warning, requiring the physician to fill in the difference reasons and record remarks to optimize subsequent model training;

[0061] CIO < 0.1: red warning, for high-risk differences, requiring the physician to verify the diagnosis result again and synchronously push it to the superior physician for review.

[0062] For the multi-modal data acquisition layer:

[0063] The multi-modal data acquisition layer of the present invention mainly includes an API interface and an EMR interface. The API interface is a dedicated interface for medical images, used to collect multi-modal data of patients' imaging from the hospital's Dolphin software; the EMR interface is an electronic medical record docking module, used to collect the medical record data of patients from the hospital's electronic medical record system. The specific implementation method is as follows:

[0064] Among them, the API interface interacts with the real-time data of the Dolphin software through the RESTful API. After inputting the patient ID, the Dolphin software imports photo files, 3DMD data, and IOS URL data, as well as the binary data of lateral cephalometric X-rays and CTBT. This system downloads and parses the relevant data. The API interface automatically synchronizes the latest data from the Dolphin software at regular intervals, and at the same time displays the progress bar and log information of data import, which is convenient for statistical verification of patient information.

[0065] Among them, the EMR interface is docked with the hospital electronic medical record system through the HL7 (Health Level Seven) standard to collect the medical record information of patients, including medical history, diagnosis results, etc. The EMR interface supports real-time data synchronization to ensure the timeliness and accuracy of data. After the data is collected, the system will perform normalization processing on the data, including data cleaning, format conversion, and standardized storage, to ensure the accuracy and consistency of subsequent analysis.

[0066] For the intelligent analysis layer:

[0067] Among them, the implementation method of the image data processing module is as follows:

[0068] The system first reads the DICOM file and parses the metadata in it, such as device type, pixel spacing, slice thickness, etc. If necessary tags are missing, such as Modality, PixelSpacing, etc., an error will be thrown and the processing will be terminated. According to the device types of CBCT and panoramic films, corresponding parameters such as HU value conversion slope and intercept are selected to convert the original pixel values into standard HU values. Physical value conversion: HU = pixel × RescaleSlope + RescaleIntercept. Use trilinear interpolation to resample the image to an isotropic resolution of 1mm 3 to achieve spatial standardization.

[0069] def_resample_isotropic(self,volume,original_spacing):

[0070] from scipy.ndimage import zoom

[0071] zoom_factors = [original_spacing[i] / 1.0 for i in range(3)]

[0072] return zoom(volume,zoom_factors,order = 3)

[0073] Subsequently, image denoising processing is carried out. The image is decomposed by wavelet to obtain coefficients at multiple scales. The threshold is calculated according to the standard deviation of the highest frequency coefficients, and the coefficients are processed by soft thresholding, and then the image is reconstructed to achieve wavelet threshold denoising. For severe artifacts, a pre-trained 3D UNet model is used for denoising. The encoder part of UNet gradually extracts features through convolutional layers, the decoder part restores the resolution through deconvolutional layers, and the middle fuses low-level and high-level features through skip connections. The final output is normalized to the range of [0,1] through the Sigmoid function to achieve deep denoising.

[0074] Perform multi-scale processing on the input lateral cephalometric radiograph: One branch processes the original resolution image, and the other branch processes the image after 2x downsampling. Each branch extracts 64-dimensional features through convolutional layers, constructs a graph structure based on spatial proximity, takes image pixels as nodes, and the similarity between pixels as edge weights. Gradually aggregate neighborhood information through a two-layer graph convolutional network (GCN) to generate high-level features of 128 and 256 dimensions. Perform global average pooling on the features output by the graph convolution to obtain a 256-dimensional global feature vector. Map the features to two-dimensional coordinates (x, y) through a fully connected layer to predict the positions of landmark points.

[0075] Use 3D ResNet as the backbone network to extract 256-dimensional feature maps, generate candidate boxes through the Region Proposal Network (RPN). Each candidate box corresponds to a tooth region. The RPN outputs 18 channels, representing the object probabilities and bounding box offsets of 9 anchor boxes respectively. Perform 3D non-maximum suppression (NMS) on the generated candidate boxes to remove redundant boxes with high overlap. Set the IoU threshold to 0.7 and retain the most likely candidate regions. For each candidate region, use the Sobel operator to calculate the edge map to enhance the contrast of the tooth boundary. Concatenate the original image and the edge map as a 2-channel input and send it into 3D UNet for fine segmentation. The output of UNet generates a tooth mask through the Sigmoid function. Finally, for the segmented tooth mask, calculate the following geometric features:

[0076] Volume: The total number of voxels in the mask;

[0077] Eccentricity: The degree of flatness of the tooth shape;

[0078] Orientation: The direction angle of the long axis of the tooth;

[0079] Long axis length: The longest diameter of the tooth;

[0080] Short axis length: The shortest diameter of the tooth.

[0081] Use the pre-trained 3D ResNet to extract the depth features of the tooth region, perform global average pooling on the feature maps, thereby obtaining a 2048-dimensional feature vector. Concatenate the geometric features and the deep learning features as the final feature vector, and achieve feature fusion through weighted summation. Feature fusion strategy: Feature = α * Geometric + β * Deep, where α and β are determined by cross-validation.

[0082] Among them, the implementation method of the image data processing module is as follows:

[0083] Read the patient's photos in JPEG or PNG format using the PIL library, convert the patient's photos to a standard format, normalize the pixel values to the range [0, 1], read the 3D model file using the trimesh library, obtain the vertex and face data, normalize the point cloud data within the unit sphere, and convert the 3DMD to a standard format.

[0084] In 2D image feature extraction, use the pre-trained Dlib library to detect 68 key facial points, calculate facial ratios, such as: face height / face width: distance from the mental point to the hairline / distance between the two zygomatic points; nasal width / face width: distance between the alae nasi / distance between the two zygomatic points, mirror the key points on the left and right sides of the face symmetrically, and calculate the root mean square error of the distances between the corresponding points.

[0085] Calculate the soft tissue thickness through the distance between the 3D model and the skeletal reference points, calculate the overall facial volume using the Convex Hull algorithm, and calculate the point cloud curvature using Open3D.

[0086] Design a Cross-Modal Transformer module to fuse 2D image features and 3D model features through cross-attention mechanism, and generate the final feature vector through a fully connected layer.

[0087] Among them, the implementation method of the IOS data processing module is as follows:

[0088] Adopt bilateral filtering to remove scanning noise and retain the sharp edges of the teeth. Establish a global coordinate system through three-point fitting: the incisal edge of the maxillary central incisor and the mesial buccal cusps of the bilateral first molars.

[0089] Extraction of dental arch morphological features:

[0090] Calculate the ratio of dental arch length / width:

[0091] Measurement of dental arch length: Take the line connecting the distal contact points of the second permanent molars on the left and right sides as the baseline, and the perpendicular line drawn from the mesial contact point of the central incisor to the baseline is the total length of the dental arch, which can be divided into three segments: the vertical distance from the mesial contact point of the incisor to the line connecting the cusp tips of the canines is the anterior segment length of the dental arch; the vertical distance from the line connecting the canines to the line connecting the mesial contact points of the first molars is the middle segment length of the dental arch; the vertical distance between the mesial surfaces of the first molars and the distal surfaces of the second molars is the posterior segment length of the dental arch.

[0092] Measurement of dental arch width: The width of the anterior segment of the dental arch (the width between the cusp tips of the canines on the left and right sides), the width of the middle segment of the dental arch (the width between the central fossae of the first premolars on the left and right sides), and the width of the posterior segment of the dental arch (the width between the central fossae of the first molars on the left and right sides).

[0093] Calculate the height of the palatal vault: By fitting the dental arch curve, calculate the vertical distance from the vertex of the palatal vault to the occlusal plane.

[0094] Extract the dental arch symmetry index: Calculate the three-dimensional coordinate mirror symmetry error of the homologous teeth on the left and right sides.

[0095] Occlusal feature extraction:

[0096] Quantify the overjet and overbite, and calculate the horizontal distance from the incisal edge of the upper central incisor to the labial surface of the lower incisors and the vertical distance from the incisal edge of the upper central incisor to the labial surface of the lower incisors.

[0097] Among them, the implementation method of the text data processing module is as follows:

[0098] The system connects to the hospital's EMR system through a standard interface to obtain structured medical record data (such as diagnosis conclusions, treatment records). These data are structured and sorted according to a standardized orthodontic medical record template. The template pre-sets required fields such as Angle malocclusion classification, dentition crowding degree grading, and description of the relationship between the upper and lower jaws. Use the regular expression [^\u4e00-\u9fa5a-zA-Z0-9%,.()] to filter non-Chinese, English characters and non-common symbols, such as consecutive extra spaces, mixed English punctuation marks, and occasional garbled characters in the medical record. Convert colloquial expressions into standard medical terms through a predefined term mapping table. For fields with missing information, the system calls a completion interface based on a generative pre-trained model to automatically fill in reasonable content according to the context semantics. For example, deduce and complete the "orthodontic treatment plan for closing the gap" from "there are obvious gaps in the anterior tooth area".

[0099] The term standardization module relies on a professional orthodontic term library to convert daily colloquial expressions or non-standard medical terms into unified standard vocabulary. For example, convert the patient's main complaint of "front teeth protruding outwards" into "labial inclination of the upper anterior teeth", and correct the commonly used "class II skeletal" to "Angle class II skeletal malocclusion". This process is realized through the construction of a bidirectional mapping dictionary for real-time replacement to ensure the consistency of subsequent analysis.

[0100] In the basic natural language processing stage, the system adopts a hybrid word segmentation strategy that combines a rule engine and a statistical model. For complex professional vocabulary unique to the orthodontic field, a rule dictionary is predefined to ensure accurate segmentation; for general vocabulary and new terms related to the context, the conditional random field model is used to dynamically analyze the word boundary probability to identify out-of-vocabulary words. After word segmentation, the CRF sequence annotation model further extracts key medical entities, and the annotation types include diagnosis conclusions, anatomical locations, symptom descriptions, treatment measures, etc.

[0101] In the traditional feature extraction stage, the TF-IDF algorithm is used to quantify the importance of words. The algorithm is optimized for the characteristics of medical texts: the minimum document frequency is set to 5, and only words that appear in at least 5 medical records are retained; the maximum document frequency is set to 85%, excluding commonly used high-frequency words such as "patient" and "examination"; at the same time, the N-gram window is extended to the 3rd order to capture cross-word association patterns such as "anterior open bite - speech disorder". Finally, 500 most discriminative keywords are selected as sparse features.

[0102] The deep semantic representation part is carried out synchronously. The Word2Vec model is trained on a million-level medical corpus, and 300-dimensional word vectors are learned through the Skip-gram architecture, so that words with similar semantics (such as "deep overbite" and "vertical excess") are close in the vector space. At the same time, the clinical-specific BERT model dynamically encodes the complete medical record text to extract 768-dimensional sentence vectors containing context information. Output: text feature vectors, with the shape of (B, L, D_text), where: B: batch size, L: text sequence length, D_text: text feature dimension, 1068 dimensions, concatenated by Word2Vec and BERT.

[0103] Among them, the implementation method of the multimodal data fusion module is as follows:

[0104] Use the torch.cat function to concatenate the feature vectors of image data, picture data, and IOS data.

[0105] multimodal_features = torch.cat([photo_features, headfilm_features, cbct_features, face_scan_features, ios_features], dim = 1)

[0106] Convert the PyTorch tensor to a Numpy array and store the multimodal feature vectors in a standardized format for subsequent models to use.

[0107] multimodal_features_np = multimodal_features.detach().numpy()

[0108] Project the text feature vectors (B, L, D_text) to 768 dimensions through a fully connected layer, and expand the sequence length L of the text feature vectors to the spatial dimensions (H, W, D) to map the text features to the same spatial dimension.

[0109] Input data: feature maps [f1, f2, f3] from different levels, corresponding to high-level, middle-level, and low-level features respectively:

[0110] f1: High-level features, with a shape of (B, 768, H1, W1, D1), containing global semantic information.

[0111] f2: Middle-level features, with a shape of (B, 384, H2, W2, D2), capturing medium-scale anatomical structures.

[0112] f3: Low-level features, with a shape of (B, 192, H3, W3, D3), retaining local details.

[0113] H is the size of the feature map in the height direction; W is the size of the feature map in the width direction; D is the size of the feature map in the depth direction.

[0114] Apply 1×1×1 convolution to the features of each level respectively to unify the number of channels to 256, and unify the features of different levels to the same dimension (256-dimensional) for subsequent fusion. Starting from the high-level features, double the resolution through three linear interpolations, then add them element-wise to the low-level features to transfer the high-level semantic information to the low-level and enhance the semantic expression ability of the low-level features. Apply 3×3×3 convolution to the fused features of each level to further refine the features. Finally, perform global average pooling on the features of each level to obtain a vector of (B, 256, 1, 1, 1), and concatenate the features of all levels to obtain a final representation of (B, 768).

[0115] Among them, the implementation method of intelligent diagnosis is as follows:

[0116] Input the features after multi-modal fusion, with a shape of (B, 768), into the diagnostic decision tree.

[0117] Use a fully connected layer to map the 768-dimensional features to 1 dimension, output the bony probability through the Sigmoid function to judge whether the case belongs to skeletal malocclusion, use a fully connected layer to map the 768-dimensional features to 4 dimensions to classify the malocclusion type and output the original logits of Angle classification, and use a fully connected layer to map the 768-dimensional features to 3 dimensions to judge the sagittal anomaly type and output the original logits of sagittal classification.

[0118] Dynamically adjust the weights of Angle classification and sagittal classification according to the bony probability, and the formula is as follows:

[0119] final_logits = (bone_prob * angle_logits) + ((1 - bone_prob) * sagittal_logits)

[0120] When the bony probability is high, the Angle classification result dominates the diagnosis; when the bony probability is low, the sagittal classification result dominates the diagnosis, and finally the category with the highest probability is selected as the diagnosis result.

[0121] For the decision support layer:

[0122] Among them, the implementation method of the diagnostic difference detection module is as follows:

[0123] The intelligent analysis layer inputs the probability distribution vector of the diagnosis result to the decision support layer: the shape is (B, num_classes), representing the prediction probability of each diagnosis category. The physician's diagnosis result inputs the diagnosis result to the decision support layer, and the diagnostic difference detection module converts the category label marked by the physician into a one-hot encoding.

[0124] Calculate the confidence interval of the intelligent diagnosis, extract the top two high probability values and their corresponding categories predicted by the model. The confidence interval is defined as [the second highest probability value, the highest probability value]. The confidence interval of the physician's diagnosis is a fixed range [0.9, 1.0]. If the confidence interval of the intelligent diagnosis has no overlap with the physician's interval or CIO < the set threshold, an alarm will be triggered, where CIO = the length of the overlapping interval / the length of the intelligent diagnosis interval.

[0125] Generate a comparison table to show the differences in key features concerned by the intelligent diagnosis and the physician's diagnosis. Through visual prompts, the system will display prompt information on the physician's working interface in the form of different colored marks and pop-up windows, highlighting the areas with high confidence in the diagnostic difference detection module, as well as highlighting the areas concerned by the physician.

[0126] Among them, the implementation method of the prompt strategy module is as follows:

[0127] For the severity of the difference, the system implements a three-level dynamic prompt strategy:

[0128] CIO < 0.5: Level 1 prompt (yellow warning), indicating a small degree of difference. Remind the physician through a non-blocking pop-up window to review the key parameters. The physician can choose to ignore the prompt or conduct a review.

[0129] CIO < 0.3: Level 2 prompt (orange warning), indicating a medium difference. Pop up a blocking pop-up window in the center of the screen. The system provides a drop-down menu or a text box for the physician to select or fill in the reason for the difference, such as "poor image quality" or "special patient history", and record the remarks to optimize the subsequent model training.

[0130] CIO < 0.1: Tertiary intervention (red warning), indicating a large difference. The system forces the physician to verify the diagnosis result again and automatically pushes the medical record data and the difference analysis report to the workstation of the superior physician. The superior physician can view the diagnostic basis of the intelligent diagnosis and the primary diagnosing physician and make a final ruling. The ruling result of the superior physician is incorporated into the feedback loop for optimizing the model parameters and prompting strategies.

[0131] In addition, the system compares the population data of similar cases through the cross-validation module, dynamically adjusts the difference threshold, and uses federated learning technology to integrate multi-center data while protecting privacy, improving the robustness of difference detection. Finally, the system incorporates the ruling result of the physician on the difference into the feedback loop, iteratively optimizes the model parameters and prompting strategies, and forms a full-cycle quality control system of "detection - explanation - intervention - learning" to ensure the dynamic balance between intelligent diagnosis and clinical practice.

Claims

1. An intelligent prompt system for orthodontic abnormality diagnosis based on deep learning, characterized in that, It includes: Multimodal data acquisition module: used to automatically import the patient's imaging data, image data, three-dimensional facial data, and three-dimensional dentition data by passing in the patient ID through the API interface, and at the same time, obtain the medical record text information in real time through the EMR interface; Data preprocessing module: used to perform modality-specific standardization, noise removal, and feature extraction on multimodal data to generate hard tissue feature vectors, soft tissue feature vectors, dentition feature vectors, and text feature vectors; Multimodal data fusion module: used to perform cross-modal fusion on feature vectors of different modalities through a multi-scale feature pyramid network to generate a unified 768-dimensional comprehensive feature vector; Intelligent diagnosis module: used to analyze the fused feature vectors based on a deep neural network and output a diagnosis result including the probability of bone, Angle classification, and sagittal plane anomaly category; Decision support module: used to quantify the difference between the intelligent diagnosis result and the physician's diagnosis result, and implement a hierarchical prompt strategy according to the severity of the difference.

2. The intelligent prompt system for orthodontic anomaly diagnosis based on deep learning according to claim 1, wherein The multimodal data acquisition module specifically includes: docking with the hospital Dolphin system through the RESTful API to obtain lateral cephalometric X-ray films and CBCT data in DICOM format in real time; supporting the standardized import of 3dMD facial scan data and IOS intraoral scan data; collecting structured medical record texts from the hospital EMR system based on the HL7 standard, including medical history, symptom description, and previous treatment records.

3. An intelligent prompt system for orthodontic abnormality diagnosis based on deep learning according to claim 1, characterized in that The data preprocessing module includes: Imaging data processing unit: uses wavelet transform or 3D UNet for denoising, and locates the landmark points on the lateral cephalometric film through a graph convolutional network, with the error controlled within 2mm; Image data processing unit: segments the facial area through Mask R-CNN and extracts geometric features in combination with 3D point cloud data; Text data processing unit: extracts keywords and semantic features through the CRF model and the BERT pre-training model.

4. The image data processing unit according to claim 3, wherein: Parse DICOM metadata, convert pixel values to standard HU values, and achieve 1mm isotropic resolution through trilinear interpolation 3 For mild artifacts, wavelet threshold denoising is used. For severe artifacts, a pre-trained 3D UNet model is used. Automatic landmark localization on lateral cephalograms is achieved based on a multi-scale graph convolutional network, with an average localization error within the clinically acceptable range of 2 mm. Finally, tooth candidate boxes are generated through a 3D ResNet-RPN network, and high-precision tooth instance segmentation is achieved by combining Sobel edge enhancement and 3D UNet.

5. The image data processing unit according to claim 3, wherein: Use the Mask R-CNN model to segment the facial area, calculate the facial ratio and symmetry index, optimize the mesh quality through the Laplacian smoothing algorithm, calculate the soft tissue thickness, facial volume, and curvature features, and finally establish the corresponding relationship between the anatomical landmark points of the 2D photo and the 3D model based on the PnP algorithm, and realize the fusion of features through Cross-ModalTransformer.

6. The text data processing unit according to claim 3, characterized in that: Filter non-standard characters through regular expressions, realize the conversion from colloquial expressions to standard medical terms based on the term mapping table, also combine the CRF model with rule-based word segmentation for entity recognition, extract 500-dimensional sparse features through TF-IDF, also use Word2Vec to generate 300-dimensional word vectors, and combine with the clinical-specific BERT model to output 768-dimensional context semantic features.

7. An intelligent prompting system for orthodontic anomaly diagnosis based on deep learning according to claim 1, characterized in that, The multimodal data fusion module includes: projecting the text features to 768 dimensions through a fully connected layer and mapping them to the spatial dimension, performing 1×1×1 convolution on the high-level, middle-level, and low-level features respectively to unify the dimensions, realizing cross-level feature fusion through three linear interpolations and 3×3×3 convolution, and finally splicing the multi-scale features into a 768-dimensional final representation through global average pooling.

8. An intelligent prompt system for orthodontic anomaly diagnosis based on deep learning according to claim 1, characterized in that, The intelligent diagnosis module includes: mapping 768-dimensional features to the skeletal malocclusion probability of 0-1 through the Sigmoid function, then parallelly outputting the logits of Angle classification and sagittal abnormality, and finally dynamically weighted fusion of multi-task outputs according to the skeletal probability value. The formula is: final_logits = (bone_prob × angle_logits) + ((1 - bone_prob) × sagittal_logits).

9. An intelligent prompt system for orthodontic abnormality diagnosis based on deep learning according to claim 1, characterized in that, The decision support module includes: Difference detection unit: quantifying the difference by calculating the overlap degree of the confidence intervals between the intelligent diagnosis and the physician diagnosis; Three-level prompt strategy: triggering yellow, orange or red warnings according to the CIO value, corresponding to non-blocking reminder, forcing to fill in the difference reason and superior review respectively.

10. The intelligent prompting system for orthodontic abnormality diagnosis based on deep learning according to claim 1, characterized in that, The specific content of the three-level prompt strategy is: CIO < 0.5: yellow warning, reminding the physician through a non-blocking pop-up window to review the key parameters; CIO < 0.3: orange warning, requiring the physician to fill in the difference reason and record remarks to optimize subsequent model training; CIO < 0.1: red warning, for high-risk differences, requiring the physician to verify the diagnosis result again and synchronously push it to the superior physician for review.

Citation Information

Cited By

  • Knee osteoarthritis dynamic grading prediction and intervention system and method based on large model

    CN121122719A

  • Intelligent allergic rhinitis diagnosis system based on multi-dimensional data fusion

    CN121191733A

  • Clinical information acquisition and synchronization system for digestive system department

    CN121601129A