Intelligent auxiliary diagnosis system for oral diseases fusing multi-modal images
The intelligent assisted diagnostic system based on multimodal imaging has achieved precise segmentation and identification of teeth and soft tissues, as well as three-dimensional reconstruction. It has solved the problem of spatial position deviation in image fusion in existing technologies and improved the comprehensiveness and relevance of oral disease diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FOURTH MILITARY MEDICAL UNIVERSITY
- Filing Date
- 2026-05-13
- Publication Date
- 2026-06-26
AI Technical Summary
Existing technologies have limitations in the fusion processing of multimodal images in the diagnosis of oral diseases. They cannot achieve precise segmentation and identification of teeth and soft tissues, or three-dimensional reconstruction. They are also difficult to accurately identify the boundaries of inflammatory lesions, and the spatial positional deviation between images has not been resolved.
An intelligent auxiliary diagnostic system for oral diseases using multimodal imaging is employed. Through data reception, image analysis, multimodal fusion, and intelligent diagnostic modules, it achieves automatic delineation and segmentation of teeth and soft tissues, three-dimensional reconstruction and feature fusion, generating a unified multimodal fused oral digital model, and performing disease type listing and quantitative assessment.
It achieves a clear correspondence between tooth location and soft tissue boundaries, detailed reconstruction of bone density distribution and root canal structure, accurate identification of inflammatory lesion boundaries, and construction of a unified multimodal fusion oral digital model, thereby improving the comprehensiveness and relevance of disease diagnosis.
Smart Images

Figure CN122289258A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent diagnostic technology for oral medical images, and in particular to an intelligent auxiliary diagnostic system for oral diseases that integrates multimodal images. Background Technology
[0002] In the diagnosis of oral diseases, the application of multimodal imaging is becoming increasingly widespread, but existing processing techniques have limitations. For intraoral optical images, conventional methods often only achieve simple segmentation of teeth and soft tissues without associating them with specific tooth location markers, and soft tissue areas also lack detailed markings. For cone-beam computed tomography (CBCT) images, the focus is often on overall three-dimensional reconstruction of the jawbone, neglecting the extraction of root canal structures and quantitative analysis of bone density distribution. For oral and maxillofacial magnetic resonance imaging (MRI), the focus is usually on displaying soft tissue morphology, making it difficult to accurately identify the boundaries of inflammatory lesions and simultaneously present key structures such as muscles, blood vessels, and nerves. In the multimodal fusion stage, existing technologies mostly involve simple image overlay, failing to address the spatial positional discrepancies between two-dimensional and three-dimensional heterogeneous models, and feature integration remains superficial, failing to form a unified panoramic view of anatomy and pathology.
[0003] Differentiated depth processing for images of different modalities is required. This includes automatically delineating and segmenting intraoral optical images to generate 2D structural maps with tooth position markers and soft tissue region markings; performing 3D alveolar bone reconstruction and root canal structure extraction on cone-beam computed tomography (CBCT) images to generate a 3D jawbone model containing bone density distribution and root canal morphology parameters; and enhancing soft tissue structures and identifying inflammatory regions on oral and maxillofacial magnetic resonance imaging (MRI) images to generate a soft tissue model displaying the boundaries of muscles, blood vessels, nerves, and suspected inflammatory lesions. Simultaneously, spatial registration and feature fusion of the aforementioned 2D structural maps, 3D jawbone model, and soft tissue model are necessary to construct a unified multimodal fused oral digital model, resolving the issues of positional misalignment and feature fragmentation in existing fusion methods. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and propose an intelligent auxiliary diagnostic system for oral diseases that integrates multimodal imaging.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: an intelligent auxiliary diagnostic system for oral diseases integrating multimodal imaging, comprising: The data receiving module receives multimodal oral imaging data uploaded by users, including intraoral optical photographs, cone-beam computed tomography images, and oral and maxillofacial magnetic resonance images. The image analysis module automatically delineates and segments the contours of teeth and soft tissues in the intraoral optical photographs, generating a two-dimensional structural map with tooth position markers and soft tissue area markers. It performs three-dimensional alveolar bone reconstruction and root canal structure extraction processing on the cone-beam computed tomography images, generating a three-dimensional jawbone model containing bone density distribution and root canal morphology parameters. It also performs soft tissue structure enhancement and inflammatory area identification processing on the oral and maxillofacial magnetic resonance images, generating a soft tissue model displaying the boundaries of muscles, blood vessels, nerves, and suspected inflammatory lesions. The multimodal fusion module performs spatial registration and feature fusion on the two-dimensional structural diagram, the three-dimensional jawbone model, and the soft tissue model to construct a unified multimodal fusion oral digital model. The intelligent diagnostic module calls a pre-trained multimodal oral disease diagnostic model to analyze and reason about the unified multimodal fusion oral digital model, generating a list of potential disease types and quantitative evaluation parameters for specific tooth positions or regions.
[0006] As a further aspect of the present invention, the intraoral optical photograph is subjected to automatic delineation and segmentation of the tooth and soft tissue contours to generate a two-dimensional structural diagram with tooth position markers and soft tissue region markings, including: A semantic segmentation model based on a deep neural network was used to perform pixel-level classification on the intraoral optical photographs, separating the tooth pixel region, gingival pixel region, buccal mucosa pixel region, and background pixel region. The isolated tooth pixel regions are subjected to connected component analysis based on morphology and edge detection to mark the independent contour boundary of each tooth. According to the International Dental Federation's tooth position recording method, each tooth with a marked independent outline boundary is automatically numbered and assigned a unique tooth position identifier; Texture and color feature analysis were performed on the gingival pixel region and the buccal mucosa pixel region. Combined with a predefined anatomical atlas, the gingival papilla, attached gingiva, and vestibular sulcus soft tissue anatomical regions were marked. By integrating the independent contour boundaries, the tooth position identifiers, and the marking information of the soft tissue anatomical regions, a two-dimensional structural map with tooth position identifiers and soft tissue region markings is generated.
[0007] As a further aspect of the present invention, the cone-beam computed tomography (CBCT) images are subjected to three-dimensional alveolar bone reconstruction and root canal structure extraction processing to generate a three-dimensional jawbone model containing bone density distribution and root canal morphology parameters, including: The cone-beam computed tomography (CBCT) images are subjected to denoising, artifact correction, and grayscale normalization preprocessing to obtain standardized three-dimensional volume data. In the standardized three-dimensional volume data, threshold segmentation and region growing algorithms are applied to separate high-density tooth tissue and relatively low-density bone tissue, and to initially reconstruct the three-dimensional surface mesh of alveolar bone and tooth. Based on the three-dimensional surface mesh, linear structure enhancement filtering based on the Hessian matrix is performed on the internal region of the tooth to enhance the visualization of the root canal structure; The root canal centerline is extracted from the enhanced three-dimensional volume data, and the geometric features of the root canal cross-section are calculated along the root canal centerline to obtain the root canal morphology parameters including diameter, curvature, and branching morphology. The bone density distribution data is estimated and assigned by calculating the correspondence between the voxel gray values of the bone tissue region and the known bone density template. The three-dimensional surface mesh with bone density attributes is integrated with the root canal centerline model containing geometric features to generate the three-dimensional jawbone model containing bone density distribution and root canal morphology parameters.
[0008] As a further aspect of the present invention, the magnetic resonance imaging of the oral and maxillofacial region is subjected to soft tissue structure enhancement and inflammatory area identification processing to generate a soft tissue model displaying the boundaries of muscles, blood vessels, nerves, and suspected inflammatory lesions, including: Multiplanar reconstruction of the oral and maxillofacial magnetic resonance images is performed to generate coronal, sagittal, and axial soft tissue contrast-enhanced image sequences. A multi-scale filter-based blood vessel and nerve bundle enhancement algorithm is applied to the soft tissue contrast enhancement image sequence to highlight the tubular structures of the blood vessels and nerves. The enhanced image was segmented at the pixel level using a tissue classification model to distinguish muscle tissue, adipose tissue, blood vessels and nerve bundles, and their three-dimensional spatial contours were extracted. By analyzing the abnormal signal intensity and regional distribution of T2-weighted and diffusion-weighted sequences, and combining them with adjacency analysis, the suspected inflammatory lesion regions with high signal or limited diffusion were identified, and their boundaries were delineated. The three-dimensional spatial contours of the segmented muscle tissue, adipose tissue, blood vessels, and nerve bundles are combined with the delineated boundaries of the suspected inflammatory lesion area through three-dimensional surface rendering and fusion to generate the soft tissue model.
[0009] As a further aspect of the present invention, the two-dimensional structural diagram, the three-dimensional jawbone model, and the soft tissue model are spatially registered and feature-fused to construct a unified multimodal fused oral digital model, including: The three-dimensional jawbone model containing bone density distribution and root canal morphology parameters is used as the reference spatial coordinate system; Based on the landmarks or global geometric features of the surface of the three-dimensional jawbone model, the two-dimensional structural map with tooth position markers and soft tissue region markers is subjected to two-dimensional to three-dimensional projection registration based on homography transformation or thin plate spline transformation, so that the two-dimensional tooth contour is aligned with the three-dimensional tooth surface. A voxel registration algorithm based on maximizing mutual information is used to rigidly or non-rigidly register the soft tissue model displaying the boundaries of muscles, blood vessels, nerves and suspected inflammatory lesions with the three-dimensional jawbone model in the reference spatial coordinate system. Establish a mapping table of correspondences between the registered models so that the same anatomical point has consistent spatial coordinates or indices in the two-dimensional structural diagram, the three-dimensional jawbone model and the soft tissue model; The registered and aligned two-dimensional contour information, three-dimensional jawbone geometry and density information, and three-dimensional soft tissue and lesion information are integrated into the same data structure to form the unified multimodal fusion oral digital model. The multimodal fusion oral digital model supports the simultaneous display and query of multiple image information from any viewpoint.
[0010] As a further aspect of the present invention, the step of calling a pre-trained multimodal oral disease diagnostic model to analyze and reason about the unified multimodal fusion oral digital model, generating a list of potential disease types and quantitative evaluation parameters for specific tooth positions or regions, includes: From the unified multimodal fusion oral digital model, multimodal feature blocks centered on specific tooth positions or the suspected inflammatory lesion areas are extracted. The multimodal feature blocks include at least local tooth geometry, adjacent bone density, soft tissue texture, and signal intensity changes. The extracted multimodal feature blocks are input into the feature encoding layer of the pre-trained multimodal oral disease diagnostic model, and the feature data of different modalities are processed through multiple parallel feature extraction pathways. In the fusion inference layer of the multimodal oral disease diagnosis model, cross-modal attention weighted fusion is performed on abstract features from different feature extraction pathways to generate a fused high-level semantic feature vector. The high-level semantic feature vector is input into the classification and regression head of the multimodal oral disease diagnostic model; The classification and regression header outputs a list of potential disease types, with each item in the list associated with a disease name and its predicted probability. The classification and regression head simultaneously outputs the quantitative assessment parameters, which include lesion volume, bone defect depth, root canal filling density score, and inflammatory infiltration area size.
[0011] As a further aspect of the present invention, from the unified multimodal fusion oral digital model, multimodal feature blocks centered on specific tooth positions or the suspected inflammatory lesion area are extracted, including: Determine the target analysis center point based on the focus selected by the user interaction or the abnormal area automatically detected by the system; With the target analysis center point as the center of the sphere, a three-dimensional spherical region of interest is defined in the unified multimodal fusion oral digital model; Within the three-dimensional spherical region of interest, surface point cloud data and normal directions are sampled from the tooth geometry sub-model of the unified multimodal fusion oral digital model to form tooth geometry feature components; From the bone density distribution sub-model of the unified multimodal fusion oral digital model, the bone density values of each voxel in the three-dimensional spherical region of interest are extracted, and their statistical distribution histogram is calculated to form the bone density feature components. From the soft tissue texture sub-model of the unified multimodal fusion oral digital model, the intensity values of multi-sequence magnetic resonance signals and their spatial gradients of the soft tissue in the three-dimensional spherical region of interest are extracted to form soft tissue texture feature components. The tooth geometric feature components, the bone density feature components, and the soft tissue texture feature components are spatially aligned and dimensionally packaged to form the multimodal feature block.
[0012] As a further aspect of the present invention, the pre-trained multimodal oral disease diagnostic model is trained through the following process: Historical oral multimodal image data and corresponding disease diagnosis labels confirmed by experts were collected to construct a training dataset. Each sample in the training dataset contains registered multimodal image data and corresponding tooth position-level disease labels and quantitative parameter labels. Construct an initial multimodal neural network model that includes the feature encoding layer, the fusion inference layer, and the classification and regression head; The initial multimodal neural network model is trained end-to-end using the training dataset. During training, a cross-modal consistency loss function is designed to constrain the consistency of feature representations of the same anatomical structure learned by the model from different modalities within the latent space; Design a multi-task joint loss function to simultaneously optimize the accuracy of disease classification and the precision of quantitative parameter regression; When the overall performance of the model on an independent validation dataset reaches a preset standard, training is stopped, and the pre-trained multimodal oral disease diagnostic model is obtained.
[0013] As a further aspect of the present invention, it also includes: The report generation module associates and binds the list of potential disease types and the quantitative assessment parameters with the corresponding spatial locations in the unified multimodal fusion oral digital model; In the interactive view of the unified multimodal fusion oral digital model, preset color coding, highlighted outlines, or three-dimensional annotations are used to visually mark the tooth positions predicted to have diseases or the suspected inflammatory lesion areas. When a user interacts by clicking or hovering over a marked area, an information panel dynamically pops up, displaying a list of the corresponding potential disease types and their predicted probabilities, as well as detailed values of the quantitative assessment parameters. It provides a sectioning tool that allows users to perform virtual sections at any location in the unified multimodal fused oral digital model, and simultaneously display fusion information from different original modal images and lesion boundary coverage maps derived from model inference on the sectioning surface; The system automatically integrates interactive content, the list of potential disease types, the quantitative assessment parameters, and key image screenshots to generate a structured, richly illustrated, and interactive diagnostic report document.
[0014] As a further aspect of the present invention, in the interactive view of the unified multimodal fusion oral digital model, preset color coding, highlighted outlines, or three-dimensional annotations are used to visually mark the tooth positions predicted to have disease or the suspected inflammatory lesion areas, including: Different color coding schemes are assigned to different diseases according to the list of potential disease types. For example, caries is represented by red, periodontitis by orange, and periapical lesions by purple. For teeth predicted to have disease, the tooth surface in the unified multimodal fusion oral digital model is covered with a semi-transparent stain using the color coding scheme. For the suspected inflammatory lesion area, within the soft tissue model or jawbone model of the unified multimodal fusion oral digital model, the isosurface extraction technique is used to generate the three-dimensional boundary surface of the lesion area, and the three-dimensional boundary surface is rendered using the color coding scheme corresponding to the disease. Overlay a bright glowing outline at the edge of all marked areas to enhance visual differentiation; The correspondence between the color coding scheme and the highlighted luminous outline is displayed in a diagram at a fixed position in the interactive view.
[0015] Compared with the prior art, the advantages and positive effects of the present invention are as follows: Differential depth processing technology automatically delineates and segments intraoral optical images to generate two-dimensional structural maps with tooth position markers and soft tissue region markings, clearly corresponding the tooth positions to the boundaries of soft tissues such as the gingiva and mucosa. It performs three-dimensional alveolar bone reconstruction and root canal structure extraction on cone-beam computed tomography (CBCT) images, generating a three-dimensional jawbone model containing bone density distribution and root canal morphology parameters, refining details such as bone density changes and root canal length and curvature. Furthermore, it enhances soft tissue structures and identifies inflammatory areas on oral and maxillofacial magnetic resonance imaging (MRI) images, generating soft tissue models displaying the boundaries of muscles, blood vessels, nerves, and suspected inflammatory lesions, accurately capturing inflammatory signals and key anatomical structures. This enables the output of structured and parametric information from various image modalities, providing more refined anatomical localization and pathological evidence for disease analysis.
[0016] The multimodal fusion module eliminates positional discrepancies between 2D structural diagrams, 3D jawbone models, and soft tissue models through spatial registration algorithms. It then integrates multi-source features such as tooth position markers, bone density, root canal parameters, and inflammatory boundaries through feature fusion to construct a unified multimodal fusion oral digital model. This model integrates planar, bone, soft tissue, and pathological information into a single coordinate system, presenting a panoramic view of oral structure and lesions. This enables subsequent diagnostic models to comprehensively infer features from different dimensions, improving the comprehensiveness and correlation judgment ability of identifying potential diseases in specific tooth positions or regions. Attached Figure Description
[0017] Figure 1 This is a timing diagram of the intelligent auxiliary diagnostic system for oral diseases that integrates multimodal images, as described in this invention. Figure 2 Flowchart generated for a 3D jawbone model; Figure 3 Line graphs showing the attention weights of different modalities under different disease types; Figure 4 A graph showing the training process of a multimodal oral disease diagnostic model; Figure 5 Radar charts are used to evaluate the effectiveness of visual labeling. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0019] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0020] See Figure 1 The overall implementation scheme of the intelligent auxiliary diagnostic system for oral diseases integrating multimodal images described in this invention is as follows: During system operation, the data receiving module first receives oral multimodal image data uploaded by the user. This data includes at least intraoral optical photographs, cone-beam computed tomography (CBCT) images, and oral and maxillofacial magnetic resonance imaging (MRI) images. Subsequently, the image analysis module processes this data in parallel. It performs automatic delineation and segmentation of tooth and soft tissue contours on the intraoral optical photographs, outputting a two-dimensional structural map with tooth position markers and soft tissue region markings. It performs three-dimensional alveolar bone reconstruction and root canal structure extraction on the CBCT images, outputting a three-dimensional jawbone model containing bone density distribution and root canal morphology parameters. It performs soft tissue structure enhancement and inflammatory region identification on the oral and maxillofacial MRI images, outputting a soft tissue model displaying the boundaries of muscles, blood vessels, nerves, and suspected inflammatory lesions. The multimodal fusion module receives the above processing results and performs spatial registration and feature fusion of the two-dimensional structural map, the three-dimensional jawbone model, and the soft tissue model to construct a unified multimodal fused oral digital model. Finally, the intelligent diagnosis module calls the pre-trained multimodal oral disease diagnosis model to analyze and reason about the unified multimodal fusion oral digital model, and outputs a list of potential disease types and quantitative assessment parameters for specific tooth positions or regions, thus completing the assisted diagnosis process.
[0021] In one embodiment of the present invention, when processing intraoral optical photographs in the image analysis module, a semantic segmentation model based on a deep neural network is used to perform pixel-level classification of the intraoral optical photographs, separating tooth pixel regions, gingival pixel regions, buccal mucosa pixel regions, and background pixel regions. Connectivity analysis based on morphology and edge detection is performed on the separated tooth pixel regions to mark the independent contour boundaries of each tooth. According to the International Dental Federation's tooth position recording method, each tooth with marked independent contour boundaries is automatically numbered and assigned a unique tooth position identifier. Texture and color feature analysis is performed on the gingival and buccal mucosa pixel regions, and combined with predefined anatomical atlases, the gingival papilla, attached gingiva, and vestibular sulcus soft tissue anatomical regions are marked. The independent contour boundaries, tooth position identifiers, and soft tissue anatomical region marking information are integrated to render and generate a two-dimensional structural map with tooth position markings and soft tissue region markings.
[0022] In practice, the automatic delineation and segmentation of intraoral optical photographs is accomplished through a deep neural network semantic segmentation model deployed in the image analysis module. This deep neural network semantic segmentation model accepts one or more intraoral optical photographs uploaded by the user as input. In practice, the architecture of the deep neural network semantic segmentation model can adopt an encoder-decoder form. The encoder part consists of multiple convolutional layers and pooling layers to extract multi-level features. The decoder part gradually restores the spatial resolution and classifies each pixel through upsampling and skip connections, outputting a pixel-level classification map with the same size as the input image. In this classification map, each pixel is assigned a label value, which corresponds to a preset category, including tooth pixel region, gingival pixel region, buccal mucosa pixel region, and background pixel region. For example, when processing an intraoral optical photograph containing the maxillary anterior teeth region, the deep neural network semantic segmentation model will judge pixel by pixel and classify it into one of the above four pixel regions, thereby completing the initial pixel region separation.
[0023] In some embodiments, for the separated tooth pixel region, the system performs connected component analysis based on morphological operations and edge detection algorithms. The morphological operations include a dilation-erosion closing operation to fill the small pores inside the tooth region and smooth the boundaries. Then, the edge detection algorithm is used to identify the contour edges of each tooth object. By scanning the binarized tooth pixel region and marking the interconnected pixel sets, the system can distinguish different connected components. Each independent connected component represents the pixel set of a tooth. The system marks the outer contour boundary of each connected component as the independent contour boundary of the tooth. For example, for the separated tooth pixel region containing six upper anterior teeth, connected component analysis can identify six independent pixel sets and accurately delineate the crown contour of each tooth.
[0024] In some embodiments, after obtaining the independent contour boundary of each tooth, the system automatically numbers the tooth position according to the International Dental Federation tooth position recording method. The system first determines the dental arch morphology and midline position by analyzing the overall spatial distribution relationship of all independent contour boundaries. For the maxillary dental arch, the two independent contour boundaries located on both sides of the midline and adjacent to the midline are identified as central incisors and assigned tooth position identifiers "11" and "21". According to the predefined rules of tooth type and relative position, the system assigns corresponding two-digit tooth position identifiers to lateral incisors, canines, etc. in sequence. For example, the identified right maxillary first premolar will be assigned a unique tooth position identifier "14".
[0025] In practice, the processing of gingival pixel regions and buccal mucosa pixel regions involves texture and color feature analysis. The system extracts local binary pattern texture features and color statistical features in a specific color space from the pixel set classified as gingival pixel regions and buccal mucosa pixel regions. These features are then input into a soft tissue region classifier based on a predefined anatomical atlas. This classifier further classifies pixels into more refined soft tissue anatomical regions, such as the gingival papilla region, attached gingiva region, and vestibular sulcus region. The predefined anatomical atlas contains prior knowledge of the typical relative positional relationships and morphological features of these regions.
[0026] It is understandable that the process of rendering and generating the final two-dimensional structural map is a step of information integration and visualization. The system overlays the independent contour boundaries, associated tooth position identifiers, and soft tissue anatomical region marking information onto the original intraoral optical photograph or a simplified background. The tooth position identifiers are displayed as text labels in the center or near the corresponding tooth contour. Different soft tissue anatomical regions can be distinguished by using semi-transparent color filling or different boundary line types. For example, the attached gingiva region can be filled with light pink, the gingival papilla region with dark pink, and the vestibule region with dashed boundary lines. Finally, a two-dimensional structural map integrating all analysis results with clear tooth position identifiers and soft tissue region markings is generated.
[0027] Optionally, during the connected component analysis, for adjacent teeth in the tooth pixel region that may be stuck together due to close contact, the system can apply the watershed algorithm for segmentation. The watershed algorithm treats each local gray-level minimum region as a "catchment basin" and the gradient magnitude of the pixel as the terrain height. By simulating the flooding process, the stuck objects are finally segmented according to the gradient ridge. This addresses the case where two mandibular molars appear as a connected component in the binary image due to close contact between their adjacent surfaces. By calculating the distance transformation map of the connected component and finding the ridge, it can effectively segment them into two independent tooth object contours.
[0028] In practice, training the deep neural network semantic segmentation model for pixel-level classification requires a large amount of labeled data. Each sample in the training dataset consists of an original intraoral optical photograph and a corresponding pixel-level label image. The pixel-level label images are generated by professionals based on anatomical knowledge. A loss function is defined to measure the difference between the pixel classification image predicted by the model and the true label image. The loss function adopts cross-entropy loss, and its formula is expressed as: in: This represents the cross-entropy loss value of the semantic segmentation model. This represents the total number of pixels in an input image. The total number of categories is represented (e.g., teeth, gums, buccal mucosa, and background). `i` represents the pixel index, iterating through all pixels from 1 to N, and `c` represents the category index, iterating through all categories from 1 to C. It is an indicator function, when pixel The true category is The value is 1 if it is true, and 0 otherwise. It is a deep neural network semantic segmentation model that predicts pixels. Category The probability is used to iteratively optimize the model parameters through backpropagation to minimize this loss function, thereby enabling the deep neural network semantic segmentation model to learn the complex relationship from intraoral optical photographs to pixel category mappings.
[0029] It is understandable that the accuracy of automatic tooth position numbering depends on the correct identification of the dental arch morphology and tooth sequence. After marking the independent contour boundaries, the system calculates the geometric center point of each boundary contour and fits a dental arch curve based on these center points. According to the arrangement order of teeth along the dental arch curve and the shape characteristics of their contours (such as aspect ratio and convex hull shape), it matches a predefined tooth type template, thereby converting the general sequence position into a specific tooth position identifier that conforms to the International Dental Federation's tooth position recording method.
[0030] See Figure 2In one embodiment of the present invention, when processing cone-beam computed tomography (CBCT) images in the image analysis module, the CBCT images are first preprocessed with denoising, artifact correction, and grayscale normalization to obtain standardized three-dimensional volume data. In the standardized three-dimensional volume data, threshold segmentation and region growing algorithms are applied to separate high-density dental tissue from relatively low-density bone tissue, and a preliminary three-dimensional surface mesh of the alveolar bone and tooth structure is reconstructed. Based on the three-dimensional surface mesh, linear structure enhancement filtering based on the Hessian matrix is performed on the internal regions of the tooth to enhance the visualization of the root canal structure. The root canal centerline is extracted from the enhanced three-dimensional volume data, and the geometric features of the root canal cross-section are calculated along the root canal centerline to obtain root canal morphology parameters including diameter, curvature, and branching morphology. By calculating the correspondence between the voxel grayscale values of the bone tissue region and a known bone density template, bone density distribution data is estimated and assigned. The three-dimensional surface mesh with bone density attributes is integrated with the root canal centerline model containing geometric features to generate a three-dimensional jawbone model containing bone density distribution and root canal morphology parameters.
[0031] In the image analysis module, when processing MRI images of the oral and maxillofacial region, multiplanar reconstruction is performed to generate coronal, sagittal, and axial soft tissue contrast-enhanced image sequences. A multi-scale filter-based vascular and nerve bundle enhancement algorithm is applied to these sequences to highlight the tubular structures of blood vessels and nerves. A tissue classification model is used to segment the enhanced images at the pixel level, distinguishing muscle tissue, adipose tissue, vascular cavities, and nerve bundles, and extracting their three-dimensional spatial contours. By analyzing the abnormal signal intensity and regional distribution of T2-weighted and diffusion-weighted sequences, combined with adjacency analysis, suspected inflammatory lesion areas with high signal or limited diffusion are identified, and their boundaries are delineated. The three-dimensional spatial contours of the segmented muscle tissue, adipose tissue, vascular cavities, and nerve bundles are then combined with the delineated boundaries of the suspected inflammatory lesion areas through three-dimensional surface rendering and fusion to generate a soft tissue model.
[0032] In practice, the preprocessing of cone-beam computed tomography (CBCT) images includes denoising, artifact correction, and grayscale normalization. The original CBCT image data consists of a series of two-dimensional projection images. Initial three-dimensional volume data is generated through a back-projection reconstruction algorithm. This initial three-dimensional volume data often contains noise and wire hardening artifacts caused by metal fillers or restorations. In practice, a three-dimensional nonlocal mean filtering algorithm is applied to denoise the initial three-dimensional volume data. The nonlocal mean filtering algorithm suppresses noise by weighted averaging of pixels with similar neighborhood structures in the image. For metal artifact correction, a method combining sine curve completion and iterative reconstruction is used. First, projection data severely attenuated by metal objects is detected and interpolated in the projection domain. Then, algebraic reconstruction techniques are used for a finite number of iterations to generate artifact-reduced three-dimensional volume data. Finally, grayscale normalization is performed to map the grayscale values of each voxel in the three-dimensional volume data to a unified normalization range, ensuring that the grayscale values of the same tissue type are comparable in different scans, and outputting standardized three-dimensional volume data.
[0033] In some embodiments, threshold segmentation and region growing algorithms are used to separate high-density dental tissue from relatively low-density bone tissue in standardized three-dimensional volume data. Setting a high grayscale threshold can initially segment the dental tissue region containing enamel, dentin, and restorative materials. However, since the grayscale of the tooth root and the surrounding bone tissue overlaps, simple threshold segmentation may lead to inaccurate boundaries. Therefore, a region growing algorithm is combined. Seed points are selected from the crown portion, and three-dimensional region growing is performed based on the grayscale similarity and spatial adjacency between voxels until the boundary between cementum and bone tissue is reached, thereby accurately separating the three-dimensional region of the entire tooth structure. For bone tissue, a lower grayscale threshold is used in combination with morphological opening operations to separate the alveolar bone region. Based on these segmented binary regions, a moving cube algorithm or a Poisson surface reconstruction algorithm is used to generate a three-dimensional surface mesh model of the alveolar bone and tooth.
[0034] In some embodiments, Hessian matrix-based linear structure enhancement filtering is performed on the internal regions of the tooth to enhance the visualization of root canal structures. Root canals appear as elongated, curved, low-density tubular structures in cone-beam computed tomography (CBCT) images. Hessian matrix-based linear structure enhancement filtering identifies tubular structures by analyzing the second derivative characteristics of the gray-level distribution in the neighborhood of each voxel. For each voxel in the 3D volumetric data, the second partial derivatives of its gray-level value in three orthogonal directions are calculated to form a Hessian matrix. Eigenvalue decomposition of the Hessian matrix yields three eigenvalues. If the eigenvalues satisfy a certain condition, it indicates that the voxel is located in a tubular structure. A tubular similarity metric is calculated based on the eigenvalues and used to enhance the original image. The formula for calculating this metric involves the eigenvalue calculation of the Hessian matrix. in: It is an eigenvalue of the Hessian matrix and satisfies , This represents the first eigenvalue of the sorted Hessian matrix, expressed in grayscale values per square millimeter. It reflects the curvature along the direction of slowest grayscale change within the local neighborhood of a voxel. This represents the second eigenvalue of the sorted Hessian matrix, expressed in grayscale values per square millimeter. It reflects the curvature along the midpoint of grayscale changes within the local neighborhood of a voxel. This represents the third eigenvalue of the sorted Hessian matrix, expressed in grayscale values per square millimeter. It reflects the curvature along the direction of the fastest grayscale change within the local neighborhood of a voxel. and It is a parameter that controls sensitivity. The closer the value is to 1, the more voxels it represents. The higher the probability that it belongs to a linear structure such as a root canal, the more likely it is to be. By combining this response value with the original image, three-dimensional volumetric data of the root canal structure can be generated.
[0035] In practice, the root canal centerline is extracted from the enhanced 3D volume data and its geometric features are calculated. The centerline is extracted from the segmented binary region of the root canal using a minimum path extraction or topology refinement algorithm. The centerline consists of a series of ordered 3D spatial points. Planes perpendicular to the tangent direction of the centerline are taken at fixed intervals along the centerline. The contour of the root canal cross section is calculated on each plane. The equivalent diameter, area, perimeter and other geometric parameters are calculated for each cross section contour. The local bending angle of the root canal is calculated by analyzing the directional changes of the center points of adjacent cross sections. The branching morphology of the root canal is identified by detecting the branching points of the centerline. Finally, a set of root canal morphology parameters including diameter, curvature and branching morphology is obtained.
[0036] It is understandable that the estimation of bone density distribution data is achieved by calculating the correspondence between the voxel gray values of bone tissue regions and known bone density templates. The system has a built-in gray-density lookup table or calibration curve established by scanning known bone density standard phantoms. For each voxel belonging to bone tissue in the three-dimensional jawbone model, its gray value in the standardized three-dimensional volume data is input into the lookup table or calibration curve to map the corresponding bone density value, usually in milligrams of hydroxyapatite per cubic centimeter. The estimated bone density value is used as attribute data and associated with each vertex or facet of the three-dimensional surface mesh model. Finally, the three-dimensional surface mesh model with bone density attribute data and the root canal centerline model containing geometric features are integrated in the same three-dimensional spatial coordinate system to generate a three-dimensional jawbone model containing bone density distribution and root canal morphology parameters.
[0037] In practice, the processing of oral and maxillofacial magnetic resonance images begins with multiplanar reconstruction. The original oral and maxillofacial magnetic resonance images may be mainly axial sequences. The multiplanar reconstruction algorithm generates coronal and sagittal image sequences by interpolation based on the three-dimensional spatial information of the original axial images, thereby obtaining coronal, sagittal and axial image sequences with enhanced soft tissue contrast. These image sequences provide a multi-view data foundation for subsequent tubular structure enhancement and tissue segmentation.
[0038] Optionally, a multi-scale filter-based enhancement algorithm for blood vessels and nerve bundles is applied to the soft tissue contrast enhancement image sequence. Blood vessels and nerve bundles also appear as tubular structures in magnetic resonance imaging. A Hessian matrix-based filter, similar to that used for root canal enhancement in cone-beam computed tomography (CBCT) images, is employed, but in a multi-scale space. By convolving the images with Gaussian kernel functions of different scales, blood vessels and nerve bundles of different diameters are detected. The tubular structure response at each scale is calculated, and the multi-scale response results are fused to finally generate an image in which blood vessels and nerve bundles are significantly enhanced, highlighting the tubular structures of blood vessels and nerves.
[0039] In some embodiments, a tissue classification model is used to perform pixel-level segmentation on the enhanced image to distinguish different soft tissue types. The tissue classification model can be a semantic segmentation model based on a three-dimensional convolutional neural network. The training data of this model consists of pixel-level labels of muscle tissue, adipose tissue, blood vessels, and nerve bundles annotated by experts. The model accepts multi-sequence magnetic resonance images after multi-plane reconstruction as input and outputs the probability that each voxel belongs to the above four types of tissues or background. The segmentation label of each voxel is obtained by taking the category with the highest probability. All voxel sets belonging to the same category are extracted according to the segmentation label, and a three-dimensional contour extraction algorithm is used to generate the three-dimensional spatial contours of muscle tissue, adipose tissue, blood vessels, and nerve bundles.
[0040] It is understandable that identifying suspected inflammatory lesion areas is achieved by analyzing abnormal signals in T2-weighted and diffusion-weighted sequences. In T2-weighted sequence images, inflammatory lesion areas usually appear as abnormally high signals, while in diffusion-weighted sequence images, inflammatory lesion areas may appear as diffusion-restricted, i.e., high signals. The system first registers the images of these two sequences to ensure spatial consistency. Then, by setting a signal intensity threshold higher than that of the surrounding normal tissue and combining it with three-dimensional connected component analysis, candidate lesion areas are initially screened. Further adjacency analysis is then performed, such as analyzing whether the candidate areas are adjacent to the root apex, root bifurcation, or known root canal infection areas, to improve the specificity of identification. For the finally identified suspected inflammatory lesion areas, three-dimensional region growth or active contour model algorithms are used to delineate their continuous and smooth three-dimensional spatial boundaries.
[0041] Optionally, the process of generating the soft tissue model involves 3D surface rendering and fusion. The 3D spatial contours of the segmented muscle tissue, adipose tissue, blood vessels, and nerve bundles are converted into 3D triangular mesh surfaces. The 3D boundaries of the identified suspected inflammatory lesion areas are also converted into mesh surfaces. Using a 3D graphics rendering engine, different color and transparency attributes are assigned to different types of mesh surfaces. For example, muscle tissue is rendered as red and semi-transparent, adipose tissue as yellow and semi-transparent, blood vessels as blue, nerve bundles as green, and suspected inflammatory lesion areas as opaque red. Finally, all mesh surfaces are fused and displayed in the same 3D scene to generate a comprehensive soft tissue model that displays the boundaries of muscles, blood vessels, nerves, and suspected inflammatory lesions.
[0042] In one embodiment of the present invention, the multimodal fusion module uses a three-dimensional jawbone model containing bone density distribution and root canal morphology parameters as the reference spatial coordinate system when performing spatial registration and feature fusion. Based on the landmark points or global geometric features on the surface of the three-dimensional jawbone model, a two-dimensional to three-dimensional projection registration based on homography transformation or thin-plate spline transformation is performed on the two-dimensional structural map with tooth position markers and soft tissue region markers, aligning the two-dimensional tooth contour with the three-dimensional tooth surface. A voxel registration algorithm based on maximizing mutual information is used to perform rigid or non-rigid registration between the soft tissue model displaying the boundaries of muscles, blood vessels, nerves, and suspected inflammatory lesions and the three-dimensional jawbone model in the reference spatial coordinate system. A mapping table of correspondences between the registered models is established, so that the same anatomical point has consistent spatial coordinates or indices in the two-dimensional structural map, the three-dimensional jawbone model, and the soft tissue model. The registered and aligned two-dimensional contour information, three-dimensional jawbone geometry and density information, and three-dimensional soft tissue and lesion information are integrated into the same data structure to form a unified multimodal fused oral digital model, which supports the simultaneous display and query of multiple image information from any viewpoint.
[0043] The intelligent diagnostic module calls a pre-trained multimodal oral disease diagnostic model for analysis and inference. This process extracts multimodal feature blocks centered on specific tooth positions or suspected inflammatory lesion areas from a unified multimodal fused oral digital model. These multimodal feature blocks include at least local tooth geometry, adjacent bone density, soft tissue texture, and signal intensity variations. The extracted multimodal feature blocks are input into the feature encoding layer of the pre-trained multimodal oral disease diagnostic model, where feature data from different modalities are processed through multiple parallel feature extraction pathways. In the fusion inference layer of the multimodal oral disease diagnostic model, cross-modal attention-weighted fusion is performed on the abstract features from different feature extraction pathways to generate a fused high-level semantic feature vector. This high-level semantic feature vector is input into the classification and regression head of the multimodal oral disease diagnostic model. The classification and regression head outputs a list of potential disease types, with each item associated with a disease name and its predicted probability. The classification and regression head also outputs quantitative evaluation parameters, including lesion volume, bone defect depth, root canal filling density score, and the size of the inflammatory infiltration area.
[0044] In practice, the multimodal fusion module performs spatial registration and feature fusion to construct a unified multimodal fused oral digital model. This process uses a three-dimensional jawbone model containing bone density distribution and root canal morphology parameters as the reference spatial coordinate system. The three-dimensional jawbone model itself originates from cone-beam computed tomography (CBCT) images and has a clear three-dimensional physical size and spatial orientation. In practice, aligning the two-dimensional structural map with tooth position markers and soft tissue region markings with the three-dimensional jawbone model involves two-dimensional to three-dimensional projection registration based on homography transformation or thin-plate spline transformation. During the operation, a [missing information - likely a reference point] is selected on the tooth surface of the three-dimensional jawbone model. The three-dimensional landmarks corresponding to the visible feature points of the tooth contour in the two-dimensional structural image are grouped together, such as the cusp and adjacent contact points. At the same time, the two-dimensional pixel coordinates of these feature points are identified in the two-dimensional structural image. A homography transformation matrix or a more flexible thin plate spline transformation function is calculated to map the two-dimensional pixel coordinates to the landmark points on the surface of the three-dimensional model. Through this transformation function, the entire two-dimensional structural image can be "wrapped" or projected onto the surface of the three-dimensional jawbone model, so that the two-dimensional tooth contour is aligned with the three-dimensional tooth surface. For the contours of soft tissues such as the gingiva, the same transformation is used to map them to the corresponding approximate anatomical positions in the three-dimensional model.
[0045] In some embodiments, a voxel registration algorithm based on maximizing mutual information is used to register a soft tissue model displaying the boundaries of muscles, blood vessels, nerves, and suspected inflammatory lesions with a three-dimensional jawbone model in a reference spatial coordinate system. The soft tissue model is derived from oral and maxillofacial magnetic resonance imaging (MRI) images, which have different imaging principles and contrasts from cone-beam computed tomography (CBCT) images. The voxel registration algorithm based on maximizing mutual information does not rely on identical grayscale features, but instead optimizes spatial transformation parameters to maximize the mutual information of the joint grayscale distribution of the two images. Mutual information is an indicator that measures the statistical dependence between two images. The algorithm first converts the surface model of the three-dimensional jawbone model into volume data, where the voxel values represent the signed distance to the model surface or simple tissue category identifiers. The soft tissue model itself is already volume data containing different tissue labels. By iteratively adjusting the parameters of rigid or non-rigid transformation, the mutual information value between the transformed soft tissue model volume data and the derived volume data of the three-dimensional jawbone model is calculated until the transformation parameters that maximize mutual information are found, thus completing rigid or non-rigid registration, so that, for example, the cortical bone boundary of the mandible is aligned in the two models.
[0046] It is understandable that establishing a mapping table of correspondences between models after registration is a data association process. The mapping table is a data structure that records the spatial coordinates or a unique global index of each pixel in the 2D structural image, each triangle or vertex in the 3D jawbone model, and each voxel in the soft tissue model in the unified world coordinate system after registration transformation. When a user clicks on a 3D location in the unified multimodal fusion oral digital model, the system can quickly retrieve the corresponding original 2D structural image pixels, 3D jawbone model geometric attributes, and soft tissue model tissue labels by querying the mapping table, thus achieving consistent mapping of anatomical points.
[0047] In practice, the registered and aligned information is integrated into the same data structure to form a unified multimodal fusion oral digital model. This data structure can be a layered scene graph or a custom hybrid data container. The container stores the registered two-dimensional contour polygon data, three-dimensional jawbone triangular mesh data and its associated bone density attributes, three-dimensional soft tissue label data, and surface mesh data extracted from the soft tissue model. The data structure also maintains the spatial transformation matrix and mapping relationship between different data layers. The multimodal fusion oral digital model supports rotation, translation, and scaling from any perspective in the graphical interface. When the user switches the display mode, image information from different original modalities can be displayed synchronously or overlaid. For example, in the three-dimensional rendering view, the two-dimensional contour projection of a tooth, its internal three-dimensional root canal model, and the pseudo-color bone density map of the surrounding jawbone can be highlighted simultaneously.
[0048] In practice, the intelligent diagnosis module calls a pre-trained multimodal oral disease diagnosis model for analysis and reasoning. The analysis and reasoning process extracts multimodal feature blocks centered on specific tooth positions or suspected inflammatory lesion areas from the unified multimodal fusion oral digital model. A multimodal feature block is a data package containing local tooth geometry point clouds extracted from the fusion model, bone density matrix of adjacent areas, multi-sequence magnetic resonance signal intensity patches of soft tissue texture, and spatial gradient change patches of signal intensity. For example, for a tooth position suspected of having periapical lesions, the feature block will contain all relevant multimodal data in a spherical region around the root apex of that tooth.
[0049] In some embodiments, the extracted multimodal feature blocks are input into the feature encoding layer of a pre-trained multimodal oral disease diagnostic model. The feature encoding layer contains multiple parallel feature extraction pathways, each consisting of a convolutional neural network or a transformer module, specifically designed to process feature data of a particular modality. Local tooth geometry point cloud data is processed through a point cloud feature extraction pathway, bone density matrix through a two-dimensional or three-dimensional convolutional neural network pathway, and soft tissue texture image patch through another image convolutional neural network pathway. Each pathway works independently, encoding the original high-dimensional, redundant feature data into low-dimensional, abstract semantic feature vectors.
[0050] In practical implementation, cross-modal attention-weighted fusion is achieved in the fusion inference layer of the multimodal oral disease diagnostic model. The fusion inference layer receives abstract feature vectors from different feature extraction pathways. First, these feature vectors are projected into a common latent space and aligned. Then, weighted fusion based on a cross-attention mechanism is adopted. The cross-attention mechanism allows a feature vector from one modality to be used as a query to focus on and aggregate relevant information from feature vectors from other modalities. The weighted fusion process can be implemented through the following form of calculation: in: This represents a query matrix derived from a dominant modality feature vector. and Representing respectively from the The key matrix and value matrix of each modal eigenvector. It is the total number of modes. It is the dimension of the key vector. The weights calculated by the function determine the contribution of different modal features in the fusion, ultimately generating a high-level semantic feature vector that integrates information from all modalities. .
[0051] It is understandable that the high-level semantic feature vector is input into the classification and regression head of the multimodal oral disease diagnostic model to complete the final output. The classification and regression head is usually composed of fully connected layers. The classification head outputs a probability vector, and each element in the vector corresponds to the predicted probability of a potential disease type, such as dental caries, pulpitis, periapical periodontitis, periodontitis, etc. The system sorts the potential disease types from high to low according to the probability values and generates a list of potential disease types. Each item in the list contains the disease name and its predicted probability. The regression head outputs a real number vector, and each element in the vector corresponds to the predicted value of a quantitative evaluation parameter. These quantitative evaluation parameters include lesion volume, bone defect depth, root canal filling compactness score, and inflammatory infiltration range size. The output of the regression head is appropriately scaled to correspond to the real physical units or scoring standards.
[0052] Optionally, the extraction of multimodal feature blocks can be based on abnormal regions automatically detected by the system. Anomaly detection is performed on a unified multimodal fusion oral digital model through pre-set rules or a lightweight anomaly screening model. The screening model quickly scans the model and marks regions with bone density significantly lower than the threshold, abnormal soft tissue signals, or geometric shapes significantly deviating from the normal statistical range. These regions are identified as candidate analysis center points, thereby triggering the extraction of multimodal feature blocks centered on these points and subsequent deep analysis and inference processes.
[0053] See Figure 3 This is a line graph showing the attention weights of different modalities across various disease types. It intuitively demonstrates that the system does not simply concatenate multimodal data, but rather dynamically allocates the weights of each modality based on the pathological characteristics of different diseases through a cross-modal attention mechanism. For example, bone mineral density features have the highest weight in diagnosing periapical periodontitis, while soft tissue texture features dominate in soft tissue inflammation. This directly reflects the intelligence and specificity of the technology, rather than general multimodal fusion. The graph shows that certain modalities have extremely low weights in specific diseases, providing a clear direction for subsequent model lightweighting and feature selection. Modal pathways with low contribution to specific diseases can be specifically optimized or pruned, improving model efficiency without sacrificing performance.
[0054] In one embodiment of the present invention, the process of extracting multimodal feature blocks from a unified multimodal fusion oral digital model specifically includes: determining a target analysis center point based on the focus selected by the user interaction or an abnormal region automatically detected by the system; defining a three-dimensional spherical region of interest (ROI) in the unified multimodal fusion oral digital model with the target analysis center point as the center; sampling surface point cloud data and normal directions from the tooth geometry sub-model of the unified multimodal fusion oral digital model to form tooth geometry feature components; extracting bone density values of each voxel within the three-dimensional spherical ROI from the bone density distribution sub-model of the unified multimodal fusion oral digital model, calculating its statistical distribution histogram, and forming bone density feature components; and extracting multi-sequence magnetic resonance signal intensity values and their spatial gradients of the soft tissue within the three-dimensional spherical ROI from the soft tissue texture sub-model of the unified multimodal fusion oral digital model, forming soft tissue texture feature components. Spatially aligning and dimensionally packaging the tooth geometry feature components, bone density feature components, and soft tissue texture feature components to form multimodal feature blocks.
[0055] The training process of the pre-trained multimodal oral disease diagnostic model includes: collecting historical multimodal oral imaging data and corresponding expert-confirmed disease diagnostic annotations to construct a training dataset. Each sample in the training dataset contains registered multimodal imaging data, corresponding tooth position-level disease labels, and quantization parameter labels. An initial multimodal neural network model is constructed, comprising a feature encoding layer, a fusion inference layer, and a classification and regression head. End-to-end supervised training of the initial multimodal neural network model is performed using the training dataset. During training, a cross-modal consistency loss function is designed to constrain the consistency of feature representations learned by the model from different modalities for the same anatomical structure within the latent space; a multi-task joint loss function is designed to simultaneously optimize the accuracy of disease classification and the precision of quantization parameter regression. Training stops when the model's overall performance on an independent validation dataset reaches a preset standard, resulting in the pre-trained multimodal oral disease diagnostic model.
[0056] In practice, the process of extracting multimodal feature blocks centered on specific tooth positions or suspected inflammatory lesion areas from a unified multimodal fusion oral digital model begins with determining the target analysis center point. The determination of the target analysis center point depends on two methods: the user directly interacts with the selected three-dimensional spatial point on the unified multimodal fusion oral digital model through a graphical interface; or the system scans the entire model through a built-in automatic detection algorithm, marks areas with abnormal bone density, abnormal soft tissue signal intensity, or geometric morphology that significantly deviates from the statistical model, and uses the centroid of these abnormal areas as the target analysis center point. For example, when analyzing a case of suspected periapical lesion, the user can directly click on the low-density area of the upper tooth root apex in the three-dimensional jawbone model, and the three-dimensional coordinates of the clicked position are set as the target analysis center point.
[0057] In some embodiments, a three-dimensional spherical region of interest (ROI) is defined in a unified multimodal fusion oral digital model with the target analysis center point as the center. The radius of the three-dimensional spherical ROI is a preset fixed value, such as 5 mm or 8 mm, and can also be adaptively adjusted according to the type of target anatomical structure. When the target analysis center point is located on the crown, the radius of the three-dimensional spherical ROI can be set smaller to focus on the tooth itself; when the target analysis center point is located on a large soft tissue lesion, the radius of the three-dimensional spherical ROI is increased accordingly to cover the entire lesion and its surrounding tissues. The three-dimensional spherical ROI serves as a spatial constraint, used to extract relevant local data from the large multimodal fusion oral digital model for subsequent feature extraction.
[0058] In practice, the geometric feature components of the tooth are extracted from the three-dimensional spherical region of interest. The operation samples surface point cloud data and normal directions from the tooth geometric sub-model of the unified multimodal fusion oral digital model. The tooth geometric sub-model is a data structure containing triangular meshes of the tooth surface. The system performs uniform point sampling at the intersection of the three-dimensional spherical region of interest and the tooth geometric sub-model to obtain the coordinates of a series of three-dimensional spatial points. At the same time, the normal vector of the triangular mesh surface at each sampling point is calculated. The set of coordinates of these points and their corresponding normal direction vectors together constitute the geometric feature components of the tooth, which represent the three-dimensional shape and orientation of the local tooth surface.
[0059] In practice, bone density feature components are extracted from a three-dimensional spherical region of interest. The operation extracts the bone density values of each voxel within the three-dimensional spherical region of interest from the bone density distribution sub-model of the unified multimodal fusion oral digital model. The bone density distribution sub-model is a three-dimensional volume data, with each voxel storing a bone density value. The system traverses all voxels within the three-dimensional spherical region of interest, reads their bone density values, and then calculates a statistical distribution histogram of these bone density values. The histogram divides the range of bone density values into several continuous intervals and counts the number of voxels in each interval. This statistical distribution histogram vector constitutes the bone density feature components, reflecting the density distribution of bone in a local area.
[0060] In practice, soft tissue texture feature components are extracted within a three-dimensional spherical region of interest. The operation extracts data from the soft tissue texture sub-model of a unified multimodal fusion oral digital model. The soft tissue texture sub-model contains registered multi-sequence magnetic resonance signal intensity volume data, such as T1-weighted, T2-weighted, and proton density-weighted sequences. Within the three-dimensional spherical region of interest, the system extracts the signal intensity values of each magnetic resonance sequence to form a three-dimensional array. Simultaneously, it calculates the gradient amplitude of the signal intensity in three spatial directions. The three-dimensional array of signal intensity from the multiple sequences and its corresponding spatial gradient amplitude array are then concatenated to form the soft tissue texture feature component. This component characterizes the signal characteristics and spatial variations of local soft tissue in multi-contrast images.
[0061] As can be understood, referring to Table 1, the spatial alignment and dimensional packaging of the tooth geometry feature components, bone density feature components, and soft tissue texture feature components are used to form the final multimodal feature block. Spatial alignment ensures that corresponding spatial points from different sub-models have consistent coordinate indices within the three-dimensional spherical region of interest. Dimensional packaging converts different feature components (which may be point clouds, histogram vectors, or three-dimensional image blocks) into fixed-length feature vectors through flattening, interpolation, or pooling operations. These feature vectors are then concatenated along the feature dimensions to form a unified multimodal feature representation, i.e., a multimodal feature block, which is used as input to the subsequent multimodal oral disease diagnostic model.
[0062] Table 1: Schematic diagram of multimodal feature block composition In practice, the training process of the pre-trained multimodal oral disease diagnostic model begins with data collection and annotation. Historical oral multimodal imaging data and corresponding disease diagnosis annotations confirmed by experts are collected to construct a training dataset. Each sample in the training dataset contains intraoral optical photographs, cone-beam computed tomography images, and oral and maxillofacial magnetic resonance imaging data that have been spatially registered, as well as tooth position-level disease labels and quantitative parameter labels provided by oral experts after review. For example, the disease label is "distal proximal maxillofacial caries of tooth 46", and the quantitative parameter label is "bone defect depth: 2.1 mm".
[0063] In some embodiments, the initial multimodal neural network model includes a feature encoding layer, a fusion inference layer, and a classification and regression head. The feature encoding layer consists of multiple parallel sub-networks, each adapted to process feature inputs of a specific modality. For example, a point cloud convolutional network processes dental geometric features, a three-dimensional convolutional neural network processes bone density feature blocks, and a two-dimensional convolutional neural network processes soft tissue texture image blocks. The fusion inference layer receives the feature vectors encoded by each sub-network and uses an attention mechanism or cross-network for feature interaction and fusion. The classification and regression head contains two branches: one branch is a fully connected layer with a softmax activation function for disease classification, and the other branch is a linear fully connected layer for quantized parameter regression.
[0064] It is understandable that the initial multimodal neural network model is trained end-to-end in a supervised manner using a training dataset. The training process uses stochastic gradient descent or its variant optimization algorithm. In each iteration, a small batch of samples is sampled from the training dataset. The multimodal feature blocks of each sample are input into the initial multimodal neural network model. The model forward propagates to generate the disease type prediction probability and the quantization parameter prediction value. The loss between the model prediction result and the true label is calculated. The loss function value is used to update all parameters of the initial multimodal neural network model through the backpropagation algorithm.
[0065] In practical implementation, a cross-modal consistency loss function is designed to constrain the feature representations learned by the model from different modalities. This function encourages that, for the same anatomical site, the features extracted from different imaging modalities (such as cone-beam computed tomography and magnetic resonance imaging) have representations in the latent space that are as close as possible. One implementation method is to calculate the cosine similarity or mean square error between feature vectors from different modal encoders, and to incorporate maximizing similarity or minimizing error as part of the optimization objective. The formula can be expressed as: in: and Representing respectively from the The and the first The feature vector output by the modal encoder This represents the cross-modal consistency loss, which, by minimizing the loss, facilitates the alignment of feature representations from different modalities in the latent space.
[0066] In practical implementation, a multi-task joint loss function is designed to simultaneously optimize the accuracy of disease classification and the precision of quantitative parameter regression. The multi-task joint loss function is a weighted sum of classification loss and regression loss. Classification loss typically uses the cross-entropy loss function, while regression loss uses either the smoothing L1 loss or the mean squared error loss function. The form of the multi-task joint loss function is as follows: in: It is classification loss. It is a regression loss. This refers to the aforementioned cross-modal consistency loss. , , These are hyperparameters used to balance the weights of different loss terms. During training, they are minimized by the joint loss function of multiple tasks. To update the model parameters.
[0067] Optionally, the criterion for stopping training is that the model's overall performance on an independent validation dataset reaches a preset standard. The overall performance is evaluated through multiple indicators, including the accuracy, precision, and recall of disease classification, as well as the mean absolute error and coefficient of determination of the regression of quantified parameters. The preset standard can be that the above indicators no longer improve within several consecutive training cycles, or reach a pre-set threshold. When the stopping criterion is met, the training process terminates, the model parameters at this time are saved, and a pre-trained multimodal oral disease diagnostic model that can be used for inference is obtained.
[0068] See Figure 4 This is a graph showing the training process of a multimodal oral disease diagnostic model, visually illustrating the changes in loss and accuracy over 100 training rounds. Its core value lies in the following aspects: Both training and validation losses rapidly decreased from close to 1.0, then stabilized after approximately 60 rounds, eventually converging in the 0.05-0.1 range. This indicates that the model effectively learned data features, and the training process was stable without drastic fluctuations. Training and validation accuracy continuously increased, eventually reaching approximately 0.95 and 0.90 respectively, indicating that the model's performance on both the training and validation sets was effectively improved, demonstrating good generalization ability. Training accuracy was consistently slightly higher than validation accuracy, but the difference remained within a reasonable range (approximately 5%), and the typical overfitting phenomenon of "continuously increasing training accuracy while stagnant or decreasing validation accuracy" was not observed, indicating a reasonable model design and effective regularization strategies.
[0069] In one embodiment of the present invention, the system further includes a report generation module, which associates and binds the list of potential disease types and quantitative assessment parameters with the corresponding spatial locations in a unified multimodal fusion oral digital model. In the interactive view of the unified multimodal fusion oral digital model, preset color coding, highlighted outlines, or 3D annotations are used to visually mark the teeth predicted to have diseases or suspected inflammatory lesion areas. When the user interacts by clicking or hovering over the marked area, an information panel dynamically pops up, displaying the corresponding list of potential disease types and their predicted probabilities, as well as detailed quantitative assessment parameter values. A cross-section tool is provided, allowing users to perform virtual cross-sections at any location in the unified multimodal fusion oral digital model, and simultaneously displaying fusion information from different original modal images and a lesion boundary coverage map derived from model inference on the cross-section surface. The system automatically integrates interactive content, the list of potential disease types, quantitative assessment parameters, and key image screenshots to generate a structured, richly illustrated, and interactively visualized diagnostic report document.
[0070] When visually marking areas in the interactive view of the unified multimodal fusion oral digital model, different color coding schemes are assigned based on the different disease types in the potential disease type list. For teeth predicted to have disease, a semi-transparent staining is applied to the tooth surface using the color coding scheme. For suspected inflammatory lesion areas, isosurface extraction technology is used to generate a 3D boundary surface of the lesion area within the soft tissue or jawbone model of the unified multimodal fusion oral digital model, and the 3D boundary surface is rendered using the corresponding disease color coding scheme. Highlighted luminous outlines are overlaid at the edges of all marked areas to enhance visual differentiation. The correspondence between the color coding scheme and the highlighted luminous outlines is displayed graphically at fixed positions in the interactive view.
[0071] In practice, after the report generation module is activated, it associates and binds the list of potential disease types and quantitative assessment parameters output by the intelligent diagnosis module with the corresponding spatial locations in the unified multimodal fusion oral digital model. The association and binding are achieved by establishing a data structure mapping table. The mapping table records the correspondence between each disease prediction item and its quantitative parameters and one or more three-dimensional spatial coordinate ranges or anatomical structure identifiers within the model. For example, the prediction item for "periapical periodontitis of distal root of tooth 36" will be associated with the set of triangular facets or voxel blocks representing the periapical region of distal root of tooth 36 in the unified multimodal fusion oral digital model. At the same time, the quantitative assessment parameters associated with this item, such as "lesion volume: 78 cubic millimeters", are also stored in the same mapping record.
[0072] In practical implementation, preset color coding, highlighted outlines, or 3D annotations are used in the interactive view of the unified multimodal fusion oral digital model to visually mark the tooth positions predicted to have disease or suspected inflammatory lesions. The system assigns different color coding schemes according to different disease types in the potential disease type list. The color coding schemes are stored in the form of lookup tables. For example, caries uses red, periodontitis uses orange, and periapical lesions use purple. The specific RGB color values corresponding to each disease are predefined in the lookup table. For the tooth position predicted to have disease, the system finds the corresponding tooth surface triangular facet by querying the association binding mapping table and uses the color coding scheme corresponding to the disease to mark this part of the triangular facet. The corner facets are covered with semi-transparent staining. This staining is achieved by assigning specific semi-transparent material colors to these facets in the graphics rendering pipeline. For suspected inflammatory lesion areas, the system uses isosurface extraction techniques, such as the moving cube algorithm, to generate 3D boundary surfaces of the lesion areas from associated voxel blocks within the soft tissue model or jawbone model of the unified multimodal fusion oral digital model. The generated 3D boundary surface mesh is then rendered using a color coding scheme corresponding to the disease. At the edges of all marked areas, the system overlays and draws bright luminous contour lines to enhance visual differentiation. These bright luminous contour lines are achieved by applying specific intensity of luminous coloring to the edge pixels of the marked areas during the post-processing stage.
[0073] In some embodiments, the correspondence between the color coding scheme and the highlighted glowing outline is displayed in a legend at a fixed position in the interactive view. The legend is a semi-transparent rectangular panel, usually placed in the corner of the view. The legend lists all the disease type names that appear in the current view, and each name is accompanied by a color block corresponding to the disease and a style example of the highlighted glowing outline. When the user adds or removes the displayed disease types through interaction, the legend content is dynamically updated to keep consistent with the current view's label content.
[0074] It is understandable that when a user interacts by clicking or hovering over a marked area, the system dynamically pops up an information panel to display detailed information. The interaction click or hover event triggers the system to perform a spatial picking calculation, which calculates the intersection points of the ray emitted from the viewpoint through the screen cursor position with all visible objects in the unified multimodal fusion oral digital model. When the ray intersects with a marked tooth surface or the boundary surface of a lesion area, the system queries the association binding mapping table based on the spatial coordinates of the intersection point or the surface identifier it is located on, retrieves all disease prediction entries and quantitative assessment parameters associated with that location, and then dynamically generates a non-modal information panel window near the screen cursor. The information panel window displays the corresponding potential disease type list and its predicted probability in a structured list format, and displays detailed quantitative assessment parameter values in a table format.
[0075] In practical implementation, the report generation module provides a sectioning tool that allows users to perform virtual sections at any location on the unified multimodal fused oral digital model. The sectioning tool allows users to define a sectioning plane on the 3D model. The sectioning plane is interactively defined by the user by dragging a movable and rotatable plane control. The system calculates the intersection of all geometric and volume data in the unified multimodal fused oral digital model with the plane according to the equation of the sectioning plane. For 2D structural diagram information, the system calculates the intersection of the sectioning plane with the projected 2D contour line and draws the intersection line on the sectioning plane. For 3D jawbone models, the system calculates the intersection of the sectioning plane with the triangular mesh and displays it as the section contour line. At the same time, it performs color mapping and filling according to the bone density value of the voxels inside the intersection line. For soft tissue models, the system calculates the interface between the sectioning plane and the volume data of each tissue label and displays different tissues in pseudo-color form. Finally, the fused information from different original modal images is displayed on the sectioning plane simultaneously.
[0076] It is understandable that the lesion boundary coverage map derived from model reasoning is displayed on the cross-sectional plane at the same time. The lesion boundary coverage map is obtained by intersecting the three-dimensional boundary surface of the lesion region output by the intelligent diagnostic module with the cross-sectional plane. The intersection line is the boundary line of the lesion on the cross-sectional plane. The system overlays this boundary line with a striking color and line width on the fused information displayed on the cross-sectional plane, clearly showing the relationship between the lesion range and the surrounding anatomical structures.
[0077] In practice, the specific allocation of the color coding scheme follows a configurable mapping function, which maps disease type identifiers to a three-dimensional color vector. A mathematical expression of this mapping relationship can be: in: Indicates the disease type identifier. It is a uniquely encoded vector or feature vector corresponding to this disease type. It is a predefined color mapping matrix, where each row represents the basis vector of a color channel. It is the calculated RGB color vector used for rendering, which is adjusted by the color mapping matrix. It can change the overall visual effect of the color coding scheme.
[0078] In some embodiments, the automatic generation of a structured, visually interactive diagnostic report document is an integrated process. The report generation module captures the user's operation log in the interactive view, including the specific viewing angle, activated marked areas, performed virtual sectioning operations, and corresponding sectional images. At the same time, it extracts the complete list of potential disease types and quantitative assessment parameters output by the intelligent diagnostic module. The system automatically extracts image screenshots of key perspectives from a unified multimodal fusion oral digital model, such as a three-dimensional view of the entire dentition, a close-up view of teeth with diseases, and a sectional view of the lesion area. These elements—interactive content, list of potential disease types, quantitative assessment parameters, and key image screenshots—are integrated according to a predefined report template. The report template defines the document's chapter structure, title format, and text and image layout, ultimately generating a visually interactive diagnostic report document with both text and images. The document format can be PDF or HTML.
[0079] Optionally, the dynamically pop-up information panel supports user interaction. The disease name or quantitative parameter value in the information panel can be set as a hyperlink. When the user clicks on a quantitative parameter such as "bone defect depth", the system can automatically adjust the perspective and zoom of the 3D view, and tile and highlight the bone defect area most related to the parameter in the center of the view, realizing deep linkage between the report data and the 3D model.
[0080] Optionally, the generated visual interactive diagnostic report document can contain interactive elements. When the report document is generated in HTML format, the key image screenshots embedded in the report can be associated with the original 3D model data. When a user clicks on a static image in the report, a lightweight web-based 3D viewer can be launched. The viewer reloads the corresponding 3D scene state and allows the user to perform basic rotation and zoom operations to verify the diagnostic findings from different angles.
[0081] See Figure 5 This is a radar chart evaluating the effectiveness of the visualization feature. The chart shows that report generation completeness is currently a relative weakness, providing a clear direction for future technological iterations. Targeted improvements can be made by optimizing information integration logic, adding automatic screenshot functionality, and enriching report templates. High scores in other dimensions also provide a baseline reference for the system's core competitiveness, allowing for further enhancement of its differentiated advantages. When presented to medical institutions or regulatory agencies, this radar chart visually demonstrates that the system not only possesses high diagnostic accuracy but also excellent visual interaction and report generation capabilities, making it a complete, efficient, and easy-to-use clinical support tool. It transforms abstract technical performance into intuitive visual advantages, significantly improving the technology's clinical acceptance and promotion potential.
[0082] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. An intelligent auxiliary diagnostic system for oral diseases integrating multimodal imaging, characterized in that, The system includes: The data receiving module receives multimodal oral imaging data uploaded by users, including intraoral optical photographs, cone-beam computed tomography images, and oral and maxillofacial magnetic resonance images. The image analysis module automatically delineates and segments the contours of teeth and soft tissues in the intraoral optical photographs, generating a two-dimensional structural map with tooth position markers and soft tissue area markers. It performs three-dimensional alveolar bone reconstruction and root canal structure extraction processing on the cone-beam computed tomography images, generating a three-dimensional jawbone model containing bone density distribution and root canal morphology parameters. It also performs soft tissue structure enhancement and inflammatory area identification processing on the oral and maxillofacial magnetic resonance images, generating a soft tissue model displaying the boundaries of muscles, blood vessels, nerves, and suspected inflammatory lesions. The multimodal fusion module performs spatial registration and feature fusion on the two-dimensional structural diagram, the three-dimensional jawbone model, and the soft tissue model to construct a unified multimodal fusion oral digital model. The intelligent diagnostic module calls a pre-trained multimodal oral disease diagnostic model to analyze and reason about the unified multimodal fusion oral digital model, generating a list of potential disease types and quantitative evaluation parameters for specific tooth positions or regions.
2. The intelligent auxiliary diagnostic system for oral diseases integrating multimodal imaging as described in claim 1, characterized in that, The intraoral optical photographs are automatically delineated and segmented to generate a two-dimensional structural map with tooth position markers and soft tissue region markings, including: A semantic segmentation model based on a deep neural network was used to perform pixel-level classification on the intraoral optical photographs, separating the tooth pixel region, gingival pixel region, buccal mucosa pixel region, and background pixel region. The isolated tooth pixel regions are subjected to connected component analysis based on morphology and edge detection to mark the independent contour boundary of each tooth. According to the International Dental Federation's tooth position recording method, each tooth with a marked independent outline boundary is automatically numbered and assigned a unique tooth position identifier; Texture and color feature analysis were performed on the gingival pixel region and the buccal mucosa pixel region. Combined with a predefined anatomical atlas, the gingival papilla, attached gingiva, and vestibular sulcus soft tissue anatomical regions were marked. By integrating the independent contour boundaries, the tooth position identifiers, and the marking information of the soft tissue anatomical regions, a two-dimensional structural map with tooth position identifiers and soft tissue region markings is generated.
3. The intelligent auxiliary diagnostic system for oral diseases integrating multimodal imaging according to claim 2, characterized in that, The cone-beam computed tomography (CBCT) images are processed to perform three-dimensional alveolar bone reconstruction and root canal structure extraction, generating a three-dimensional jawbone model containing bone density distribution and root canal morphology parameters, including: The cone-beam computed tomography (CBCT) images are subjected to denoising, artifact correction, and grayscale normalization preprocessing to obtain standardized three-dimensional volume data. In the standardized three-dimensional volume data, threshold segmentation and region growing algorithms are applied to separate high-density tooth tissue and relatively low-density bone tissue, and to initially reconstruct the three-dimensional surface mesh of alveolar bone and tooth. Based on the three-dimensional surface mesh, linear structure enhancement filtering based on the Hessian matrix is performed on the internal region of the tooth to enhance the visualization of the root canal structure; The root canal centerline is extracted from the enhanced three-dimensional volume data, and the geometric features of the root canal cross-section are calculated along the root canal centerline to obtain the root canal morphology parameters including diameter, curvature, and branching morphology. The bone density distribution data is estimated and assigned by calculating the correspondence between the voxel gray values of the bone tissue region and the known bone density template. The three-dimensional surface mesh with bone density attributes is integrated with the root canal centerline model containing geometric features to generate the three-dimensional jawbone model containing bone density distribution and root canal morphology parameters.
4. The intelligent auxiliary diagnostic system for oral diseases integrating multimodal imaging according to claim 3, characterized in that, The oral and maxillofacial magnetic resonance images are processed with soft tissue structure enhancement and inflammatory area identification to generate a soft tissue model displaying the boundaries of muscles, blood vessels, nerves, and suspected inflammatory lesions, including: Multiplanar reconstruction of the oral and maxillofacial magnetic resonance images is performed to generate coronal, sagittal, and axial soft tissue contrast-enhanced image sequences. A multi-scale filter-based blood vessel and nerve bundle enhancement algorithm is applied to the soft tissue contrast enhancement image sequence to highlight the tubular structures of the blood vessels and nerves. The enhanced image was segmented at the pixel level using a tissue classification model to distinguish muscle tissue, adipose tissue, blood vessels and nerve bundles, and their three-dimensional spatial contours were extracted. By analyzing the abnormal signal intensity and regional distribution of T2-weighted and diffusion-weighted sequences, and combining them with adjacency analysis, the suspected inflammatory lesion regions with high signal or limited diffusion were identified, and their boundaries were delineated. The three-dimensional spatial contours of the segmented muscle tissue, adipose tissue, blood vessels, and nerve bundles are combined with the delineated boundaries of the suspected inflammatory lesion area through three-dimensional surface rendering and fusion to generate the soft tissue model.
5. The intelligent auxiliary diagnostic system for oral diseases integrating multimodal imaging according to claim 4, characterized in that, The two-dimensional structural diagram, three-dimensional jawbone model, and soft tissue model are spatially registered and feature-fused to construct a unified multimodal fused oral digital model, including: The three-dimensional jawbone model containing bone density distribution and root canal morphology parameters is used as the reference spatial coordinate system; Based on the landmarks or global geometric features of the surface of the three-dimensional jawbone model, the two-dimensional structural map with tooth position markers and soft tissue region markers is subjected to two-dimensional to three-dimensional projection registration based on homography transformation or thin plate spline transformation, so that the two-dimensional tooth contour is aligned with the three-dimensional tooth surface. A voxel registration algorithm based on maximizing mutual information is used to rigidly or non-rigidly register the soft tissue model displaying the boundaries of muscles, blood vessels, nerves and suspected inflammatory lesions with the three-dimensional jawbone model in the reference spatial coordinate system. Establish a mapping table of correspondences between the registered models so that the same anatomical point has consistent spatial coordinates or indices in the two-dimensional structural diagram, the three-dimensional jawbone model and the soft tissue model; The registered and aligned two-dimensional contour information, three-dimensional jawbone geometry and density information, and three-dimensional soft tissue and lesion information are integrated into the same data structure to form the unified multimodal fusion oral digital model. The multimodal fusion oral digital model supports the simultaneous display and query of multiple image information from any viewpoint.
6. The intelligent auxiliary diagnostic system for oral diseases integrating multimodal imaging according to claim 5, characterized in that, The process involves calling a pre-trained multimodal oral disease diagnostic model to analyze and reason about the unified multimodal fusion oral digital model, generating a list of potential disease types and quantitative assessment parameters for specific tooth positions or regions, including: From the unified multimodal fusion oral digital model, multimodal feature blocks centered on specific tooth positions or the suspected inflammatory lesion areas are extracted. The multimodal feature blocks include at least local tooth geometry, adjacent bone density, soft tissue texture, and signal intensity changes. The extracted multimodal feature blocks are input into the feature encoding layer of the pre-trained multimodal oral disease diagnostic model, and the feature data of different modalities are processed through multiple parallel feature extraction pathways. In the fusion inference layer of the multimodal oral disease diagnosis model, cross-modal attention weighted fusion is performed on abstract features from different feature extraction pathways to generate a fused high-level semantic feature vector. The high-level semantic feature vector is input into the classification and regression head of the multimodal oral disease diagnostic model; The classification and regression header outputs a list of potential disease types, with each item in the list associated with a disease name and its predicted probability. The classification and regression head simultaneously outputs the quantitative assessment parameters, which include lesion volume, bone defect depth, root canal filling density score, and inflammatory infiltration area size.
7. The intelligent auxiliary diagnostic system for oral diseases integrating multimodal imaging according to claim 6, characterized in that, From the unified multimodal fusion oral digital model, multimodal feature blocks centered on specific tooth positions or the suspected inflammatory lesion areas are extracted, including: Determine the target analysis center point based on the focus selected by the user interaction or the abnormal area automatically detected by the system; With the target analysis center point as the center of the sphere, a three-dimensional spherical region of interest is defined in the unified multimodal fusion oral digital model; Within the three-dimensional spherical region of interest, surface point cloud data and normal directions are sampled from the tooth geometry sub-model of the unified multimodal fusion oral digital model to form tooth geometry feature components; From the bone density distribution sub-model of the unified multimodal fusion oral digital model, the bone density values of each voxel in the three-dimensional spherical region of interest are extracted, and their statistical distribution histogram is calculated to form the bone density feature components. From the soft tissue texture sub-model of the unified multimodal fusion oral digital model, the intensity values of multi-sequence magnetic resonance signals and their spatial gradients of the soft tissue in the three-dimensional spherical region of interest are extracted to form soft tissue texture feature components. The tooth geometric feature components, the bone density feature components, and the soft tissue texture feature components are spatially aligned and dimensionally packaged to form the multimodal feature block.
8. The intelligent auxiliary diagnostic system for oral diseases integrating multimodal imaging according to claim 7, characterized in that, The pre-trained multimodal oral disease diagnostic model is trained through the following process: Historical oral multimodal image data and corresponding disease diagnosis labels confirmed by experts were collected to construct a training dataset. Each sample in the training dataset contains registered multimodal image data and corresponding tooth position-level disease labels and quantitative parameter labels. Construct an initial multimodal neural network model that includes the feature encoding layer, the fusion inference layer, and the classification and regression head; The initial multimodal neural network model is trained end-to-end using the training dataset. During training, a cross-modal consistency loss function is designed to constrain the consistency of feature representations of the same anatomical structure learned by the model from different modalities within the latent space; Design a multi-task joint loss function to simultaneously optimize the accuracy of disease classification and the precision of quantitative parameter regression; Training stops when the overall performance of the model on an independent validation dataset reaches a preset standard, thus obtaining the pre-trained multimodal oral disease diagnostic model.
9. The intelligent auxiliary diagnostic system for oral diseases integrating multimodal imaging according to claim 8, characterized in that, Also includes: The report generation module associates and binds the list of potential disease types and the quantitative assessment parameters with the corresponding spatial locations in the unified multimodal fusion oral digital model; In the interactive view of the unified multimodal fusion oral digital model, preset color coding, highlighted outlines, or three-dimensional annotations are used to visually mark the tooth positions predicted to have diseases or the suspected inflammatory lesion areas. When a user interacts by clicking or hovering over a marked area, an information panel dynamically pops up, displaying a list of the corresponding potential disease types and their predicted probabilities, as well as detailed values of the quantitative assessment parameters. It provides a sectioning tool that allows users to perform virtual sections at any location in the unified multimodal fused oral digital model, and simultaneously display fusion information from different original modal images and lesion boundary coverage maps derived from model inference on the sectioning surface; The system automatically integrates interactive content, the list of potential disease types, the quantitative assessment parameters, and key image screenshots to generate a structured, richly illustrated, and interactive diagnostic report document.
10. The intelligent auxiliary diagnostic system for oral diseases integrating multimodal imaging according to claim 9, characterized in that, In the interactive view of the unified multimodal fusion oral digital model, preset color coding, highlighted outlines, or 3D annotations are used to visually mark the teeth predicted to have disease or the suspected inflammatory lesion areas, including: Different color coding schemes are assigned to different diseases according to the list of potential disease types. For example, caries is represented by red, periodontitis by orange, and periapical lesions by purple. For teeth predicted to have disease, the tooth surface in the unified multimodal fusion oral digital model is covered with a semi-transparent stain using the color coding scheme. For the suspected inflammatory lesion area, within the soft tissue model or jawbone model of the unified multimodal fusion oral digital model, the isosurface extraction technique is used to generate the three-dimensional boundary surface of the lesion area, and the three-dimensional boundary surface is rendered using the color coding scheme corresponding to the disease. Overlay a bright glowing outline at the edge of all marked areas to enhance visual differentiation; The correspondence between the color coding scheme and the highlighted luminous outline is displayed in a diagram at a fixed position in the interactive view.