Oral panoramic image processing system

By combining the MMC-YOLO network model with a multi-task information integration and analysis module, the problem of the existing technology being difficult to simultaneously detect multiple dental complications and process multi-scale pathological features is solved. This enables accurate classification of the severity of periodontitis and accurate identification of multi-scale lesions, improving detection and segmentation performance.

CN120278989BActive Publication Date: 2025-10-10GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510431918.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-10-10
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

Existing oral panoramic imaging technology has difficulty in simultaneously detecting multiple oral diseases concurrent with teeth, cannot effectively integrate multi-dimensional pathological information, and performs poorly when processing pathological features at different scales, especially in identifying lesion areas with subtle density changes.

Method used

The MMC-YOLO network model is used, combined with target detection and instance segmentation technology, through the multi-scale progressive feature aggregation module, coordinate-guided fusion module and edge-enhanced spatial fusion module, to achieve simultaneous detection of dental complications, comprehensive analysis of multi-dimensional pathological information and accurate identification of multi-scale lesions.

Benefits of technology

It improves the ability to classify the severity of periodontitis, enhances the perception of caries and periradicular lesions in early small-scale scenarios, improves the model's ability to effectively segment tooth edges, and reduces computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278989B_ABST
    Figure CN120278989B_ABST
Patent Text Reader

Abstract

The application discloses an oral cavity panoramic image processing system, comprising: an image acquisition module, which is used for collecting initial oral cavity panoramic images containing various oral diseases; an image processing module, which is connected with the image acquisition module and is used for target detection and instance segmentation labeling on the initial oral cavity panoramic images, and expanding the training data set by using a data enhancement technology to obtain target oral cavity panoramic images; a model construction module, which is connected with the image processing module and is used for training the MMC-YOLO network model through the target oral cavity panoramic images to obtain a target MMC-YOLO network model; and an analysis and detection module, which is connected with the model construction module and is used for synchronous detection of teeth and concurrent oral diseases, comprehensive analysis of multi-dimensional pathological information and accurate identification of multi-scale lesions through the target MMC-YOLO network model to obtain analysis results. The system realizes synchronous detection and accurate grading of multiple types of lesions in the oral cavity panoramic image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and in particular relates to an oral panoramic image processing system. Background Art

[0002] Panoramic oral imaging is crucial for the early detection, diagnosis, and prevention of oral diseases. Two main panoramic imaging techniques are clinically used: conventional panoramic X-ray imaging and CBCT panoramic imaging. Conventional panoramic X-ray imaging uses a rotating X-ray source and a synchronously moving photoreceptor, collimated through slits, to obtain continuous planar projections of the oral and maxillofacial regions. This technique is simple to use, requires minimal examination time, and provides clinicians with a comprehensive overview of anatomical structures such as the teeth, jaws, and temporomandibular joint at a relatively low radiation dose and cost. CBCT panoramic imaging utilizes a cone-shaped X-ray beam to acquire multi-angle two-dimensional projection data, which are then reconstructed into high-resolution images through computer processing. CBCT images generate high-contrast images that accurately depict tooth roots and periodontal structures. Geometric distortion is significantly reduced, ensuring accurate anatomical proportions. Tissue overlap is also significantly minimized, providing more detailed information about oral physiology. Both panoramic imaging techniques provide rich, panoramic oral information and are invaluable for the diagnosis of various oral diseases, including caries, wisdom teeth, periradiculopathy, and periodontitis.

[0003] Although oral panoramic imaging technology can provide a wealth of oral physiological information, accurate diagnosis requires the effective extraction of pathological information. Traditional manual interpretation methods face numerous limitations in this regard. Clinicians rely primarily on professional experience to assess the morphological and radiographic density characteristics of teeth, alveolar bone, and gingival tissue, a process that inevitably involves subjective variability. Furthermore, manual interpretation lacks sensitivity to subtle pathological changes such as early caries or latent root lesions, which can easily lead to clinical misdiagnosis. More critically, traditional methods lack effective means to quantify the severity of certain oral diseases (such as periodontitis), making it difficult to provide objective and consistent disease grading assessments, increasing the risk of missed or inaccurate diagnoses. Although computer-aided diagnosis (CAD) systems based on traditional image processing techniques (including edge detection and morphological operations) have achieved partial automation, these systems rely on manually set feature extraction rules and have limited generalization capabilities. They cannot fully cope with the complex and diverse lesion morphologies in panoramic images and struggle to meet the precise diagnostic needs of modern dentistry.

[0004] In recent years, deep learning technology has made significant progress in the field of medical image analysis. Methods based on convolutional neural network architectures have demonstrated superior performance compared to traditional methods in specific applications such as caries detection and tooth segmentation, significantly improving lesion localization and quantification accuracy. However, current deep learning-based dental panoramic image analysis systems still have several limitations: First, most existing studies adopt a single-task framework, which limits the ability to simultaneously detect multiple coexisting pathologies on a single tooth (e.g., simultaneously detecting caries and confirming the presence of concurrent periodontitis). Second, existing systems struggle to provide quantitative assessments of complex diseases such as periodontitis because they cannot effectively integrate key measurements such as alveolar bone level assessment, enamel-cement junction localization, and tooth position characteristics. Accurately quantifying periodontitis severity requires models that can simultaneously perform object detection and instance segmentation tasks and integrate this information for comprehensive analysis. Third, existing methods perform poorly when processing pathological features at different scales, especially for lesion regions with subtle density variations. For example, deep convolutional feature extraction often misinterprets adjacent low-density radiographic features as interdental background when processing caries on proximal surfaces.

[0005] In summary, for the pathology-assisted diagnosis system based on oral panoramic images, the core challenge of current research is to develop a comprehensive framework that can simultaneously detect the characteristics of oral diseases concurrent with teeth, detect oral diseases that require processing multi-dimensional pathological information, and adapt to multi-scale scenarios. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides an oral panoramic image processing system, comprising:

[0007] An image acquisition module, used for collecting initial oral panoramic images containing various oral diseases;

[0008] An image processing module, connected to the image acquisition module, is used to perform target detection and instance segmentation annotation on the initial oral panoramic image, and to expand the training data set using data enhancement technology to obtain a target oral panoramic image;

[0009] A model building module, connected to the image processing module, is used to train the MMC-YOLO network model using the target oral panoramic image to obtain a target MMC-YOLO network model;

[0010] The analysis and detection module is connected to the model building module and is used to perform synchronous detection of dental complications and oral diseases, comprehensive analysis of multi-dimensional pathological information, and accurate identification of multi-scale lesions through the target MMC-YOLO network model to obtain analysis results.

[0011] Preferably, the target detection and instance segmentation annotation include target detection and segmentation annotation of lesion type, location and boundary information;

[0012] The data augmentation includes rotation, scaling and flipping.

[0013] Preferably, the target MMC-YOLO network model includes a shared backbone network, a target detection neck network, a target detection detection head, an instance segmentation neck network, an instance segmentation segmentation head, and a multi-task information comprehensive calculation and analysis module;

[0014] Wherein, the shared backbone network is used to extract common features of oral panoramic images;

[0015] The target detection neck network is used to process the features of the target detection task;

[0016] The target detection head is used to identify caries, wisdom teeth and periradicular lesions;

[0017] The instance segmentation neck network is used to process the features of the instance segmentation task;

[0018] The example segmentation head is used to segment the alveolar bone, the enamel-cementum junction and the tooth contour;

[0019] The multi-task information comprehensive calculation and analysis module is used to integrate target detection and instance segmentation results to obtain periodontitis grading results and multi-disease detection results.

[0020] Preferably, the shared backbone network and the target detection neck network include a multi-scale progressive feature aggregation module;

[0021] The multi-scale progressive feature aggregation module is used to integrate local and global features through a progressive multi-scale feature extraction strategy and partial convolution operations; and adopts channel segmentation and progressive channel reduction methods to retain key features while reducing computational complexity.

[0022] Preferably, the multi-scale progressive feature aggregation module integrates local and global features through a progressive multi-scale feature extraction strategy and a partial convolution operation, comprising:

[0023] Given an input feature, a standard 3×3 convolution operation is first performed, and then the feature is divided into two parts in the channel dimension. The first part extracts medium-scale features through 5×5 convolution, and a quarter of the channel extracts global features through 7×7 convolution; finally, the multi-scale features are fused through channel splicing and transformation.

[0024] Preferably, a coordinate-guided fusion module is introduced into the process of identifying caries, wisdom teeth and periradicular lesions by the target detection head;

[0025] The coordinate-guided fusion module is used to combine the edge and texture information of shallow features with the semantic information of deep features, dynamically adjust the feature expression through the coordinate attention mechanism, and enhance the feature modeling ability of the target area when small-scale lesions are confused with background structures.

[0026] Preferably, the coordinate-guided fusion module combines edge and texture information of shallow features with semantic information of deep features, and dynamically adjusts feature expression through a coordinate attention mechanism, including:

[0027] First, the shallow features and deep features are dimensionally adjusted and preliminarily fused. The mapping function of the 1×1 convolution transformation is defined to align the number of feature channels. Then, the adjusted shallow features and deep features are concatenated along the channel dimension to generate fused features. Through the local-global interaction in the spatial dimension, an adaptive attention map is generated to optimize the feature expression.

[0028] Preferably, the process of generating an adaptive attention map to optimize feature expression through local-global interaction in the spatial dimension includes:

[0029] First, the input features are locally aggregated, and then the spatial dependencies in the horizontal and vertical directions are modeled using separable directional convolution through initial feature transformation. Finally, the attention map is generated through feature transformation and normalization.

[0030] Preferably, the instance segmentation head segments the alveolar bone, CEJ and tooth contours through feature stitching, edge enhancement and spatial attention mechanism while retaining fine-grained structural information.

[0031] Compared with the prior art, the present invention has the following advantages and technical effects:

[0032] The oral panoramic image processing system of the present invention combines target detection with instance segmentation, and aggregates the output information of multiple tasks into the designed pathological feature integration analysis module, thereby effectively classifying the severity of periodontitis. In addition, the combination of multi-task information can also specifically mark other oral diseases occurring concurrently in the same tooth.

[0033] The present invention uses a multi-scale progressive feature aggregation strategy to achieve cross-scale feature fusion while reducing computational complexity and retaining basic features for detecting and segmenting multi-scale lesions that are difficult to distinguish from surrounding tissues.

[0034] This method dynamically enhances the representation of shallow and deep features, accurately capturing subtle pathological features that are easily confused with background structures, and enhances the perception of early-stage caries and periradicular lesions in small-scale scenes.

[0035] The present invention retains fine-grained structural information through feature splicing, edge enhancement and spatial attention mechanism, thereby improving the effective segmentation of tooth edges and enhancing the ability of the model to distinguish between fuzzy boundaries and background. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0037] Figure 1 Schematic diagram of the system workflow of an embodiment of the present invention;

[0038] Figure 2 This is a diagram of the MMC-YOLO framework according to an embodiment of the present invention;

[0039] Figure 3 Schematic diagram of each module in the target MMC-YOLO network model of an embodiment of the present invention;

[0040] Figure 4 This is a periodontitis grading result diagram output by the MMC-YOLO framework according to an embodiment of the present invention;

[0041] Figure 5 This is a graph of the comprehensive labeling results of multi-task information fusion and concurrent diseases output by the MMC-YOLO framework in an embodiment of the present invention. DETAILED DESCRIPTION

[0042] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0043] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0044] This embodiment proposes an innovative multi-task collaborative framework to address the complex detection needs of multiple types of lesions in oral panoramic images, aiming to solve three key challenges: simultaneous detection of dental complications and oral diseases, comprehensive analysis of multi-dimensional pathological information, and accurate identification of multi-scale lesions.

[0045] The framework integrates target detection and instance segmentation technology to build an end-to-end oral lesion detection system. Under this architecture, the embodiment designs a pathological feature integration analysis module to simultaneously complete feature extraction of tooth concurrent diseases and accurate classification of periodontitis severity. To enhance the detection capability of the model in multi-scale lesion scenarios, three function modules that work synergistically are further proposed: a multi-scale progressive feature aggregation module, a coordinate-guided fusion module, and an edge-enhanced spatial fusion module. These modules improve the performance of the network in the tasks of variant scale lesion detection and fine segmentation through hierarchical feature extraction and fusion strategies.

[0046] The system aims to solve the following problems:

[0047] Simultaneous detection of tooth concurrent oral diseases: By integrating target detection and instance segmentation into a multi-task network, the output multi-task information is processed through a designed pathological feature integration analysis module to classify periodontitis according to the required pathological information (alveolar bone position feature, enamel-dentin junction position feature, tooth position feature), and simultaneously detect other concurrent oral diseases on the corresponding teeth (caries, wisdom teeth, periapical lesions).

[0048] Multi-scale scene oral disease lesion feature extraction enhancement: For the complex and variable pathological features in panoramic images, a multi-scale progressive feature aggregation module is developed to effectively utilize the feature information of small-scale lesions. A coordinate-guided fusion module is designed to enhance the discrimination between small-scale lesions and the background. An edge-enhanced spatial fusion module is also proposed to improve the feature fusion effect of objects with fuzzy edge features.

[0049] As shown in Figures 1-5 , the embodiment provides an oral panoramic image processing system, which includes:

[0050] An image acquisition module is used to collect initial oral panoramic images containing multiple oral diseases.

[0051] An image processing module is connected to the image acquisition module and is used for target detection and instance segmentation labeling of the initial oral panoramic images, and data augmentation techniques are used to expand the training data set to obtain target oral panoramic images.

[0052] A model construction module is connected to the image processing module and is used to train the MMC-YOLO network model through the target oral panoramic images to obtain the target MMC-YOLO network model.

[0053] The analysis and detection module is connected to the model building module and is used to perform simultaneous detection of dental complications and oral diseases, comprehensive analysis of multi-dimensional pathological information, and accurate identification of multi-scale lesions through the target MMC-YOLO network model to obtain analysis results.

[0054] Furthermore, target detection and instance segmentation annotation include target detection and segmentation annotation of lesion type, location and boundary information;

[0055] Data augmentation includes rotation, scaling, and flipping.

[0056] Furthermore, the target MMC-YOLO network model includes a shared backbone network, a target detection neck network, a target detection detection head, an instance segmentation neck network, an instance segmentation segmentation head, and a multi-task information comprehensive calculation and analysis module;

[0057] Among them, the shared backbone network is used to extract common features of oral panoramic images;

[0058] The target detection neck network is used to process the features of the target detection task;

[0059] The target detection head is used to identify caries, wisdom teeth and periradicular lesions;

[0060] The instance segmentation neck network is used to process the features of the instance segmentation task;

[0061] The instance segmentation head is used to segment the alveolar bone, CEJ, and tooth contours;

[0062] The multi-task information comprehensive calculation and analysis module is used to integrate target detection and instance segmentation results to obtain periodontitis grading results and multi-disease detection results.

[0063] Specifically, the target MMC-YOLO network model consists of six core parts with clear functions: shared backbone network (Backbone), target detection neck network (Detect neck), target detection detection head (Detect head), instance segmentation neck network (Segment neck), instance segmentation segmentation head (Segment head) and multi-task information comprehensive calculation and analysis module. The overall architecture is as follows Figure 2As shown. In the shared backbone network and the target detection neck network, this embodiment designs a multi-scale progressive feature aggregation module (MPFAM), which is specially designed to deal with multi-task detection scenarios with significant scale differences in oral images, such as tiny caries to large-scale alveolar bone boundaries. The coordinate guided fusion module (CGFM) is introduced in the target detection feature processing link. By fusing multi-scale features and combining the global context perception mechanism, the feature expression ability and discrimination accuracy of the target area are significantly improved. In response to the feature processing requirements of instance segmentation, this embodiment develops an edge enhanced spatial fusion module (ESFM), which effectively improves the utilization efficiency of common fuzzy boundary features in oral images through the organic combination of feature splicing, edge enhancement and spatial attention mechanism. In the output part of the multi-task network, the target detection task detection head focuses on accurately identifying three types of key oral lesion features: caries, wisdom teeth and periradicular lesions, which are different in scale and form. The instance segmentation task is completed by three specially designed segmentation heads, which are responsible for the segmentation of alveolar bone contour, enamel-cementum junction and complete tooth contour respectively. The final step in the entire analysis process is the pathological feature integration analysis module, which systematically integrates the results of target detection and instance segmentation. This module can not only grade the severity of periodontitis, but also comprehensively evaluate the pathological condition of individual teeth and specially mark teeth with multiple lesions to obtain marking results.

[0064] The multi-task framework proposed in this embodiment is based on target detection and instance segmentation, integrated pathological information, and the PCA tooth growth axis positioning method to classify the severity of periodontitis and label other concurrent oral diseases.

[0065] Furthermore, a multi-scale progressive feature aggregation module is included in the shared backbone network and the object detection neck network;

[0066] The multi-scale progressive feature aggregation module is used to integrate local and global features through a progressive multi-scale feature extraction strategy and partial convolution operations; and adopts channel segmentation and progressive channel reduction methods to retain key features while reducing computational complexity.

[0067] Furthermore, the multi-scale progressive feature aggregation module integrates local and global features through a progressive multi-scale feature extraction strategy and partial convolution operations. The process includes:

[0068] Given an input feature, a standard 3×3 convolution operation is first performed, and then the feature is divided into two parts in the channel dimension. The first part extracts medium-scale features through 5×5 convolution, and a quarter of the channel extracts global features through 7×7 convolution; finally, the multi-scale features are fused through channel splicing and transformation.

[0069] Specifically, it should be noted that Figure 3In the figure, (a) is the grouped convolution of MPFAM. (b) is the Multi-Scale Progressive Feature Aggregation Module (MPFAM), where c_in represents the input channel. (c) is the Coordinate Attention Module of CGFM. (d) is the Coordinate Guided Fusion Module (CGFM). (e) is the Edge Enhanced Spatial Fusion Module (ESFM), where σ(w) is the weight, τ is the weight threshold, and α is the edge feature scaling factor.

[0070] In order to effectively deal with the multi-scale feature integration problem in dental panoramic image analysis, this embodiment proposes a multi-scale progressive feature aggregation model (MPFAM). Figure 3 (b) This module achieves a balance between computational efficiency and feature expression capability through a progressive multi-scale feature extraction strategy and partial convolution operations.

[0071] Specifically, given the input features Where C represents the number of channels, H and W represent the height and width of the feature map respectively.

[0072] First, perform a standard 3×3 convolution operation:

[0073]

[0074] The features are then divided into two parts along the channel dimension:

[0075]

[0076] Next, the first part extracts medium-scale features through 5×5 convolution:

[0077]

[0078] Furthermore, one-quarter of the channels are subjected to 7×7 convolution to extract global features:

[0079]

[0080] Finally, these multi-scale features are fused through channel concatenation and transformation:

[0081] X fused =Concat(X3,X 2,2 ,X 1,2 )

[0082]

[0083] X out =X compressed +X

[0084] Among them, φ1, φ2, φ3, φ4 represent convolution operations with ReLU activation, W1, W2, W3 are learnable weight matrices of convolutions of different scales, and b1, b2, b3 are the corresponding bias terms X. i,j Represents the feature map after segmentation, i represents the number of layers, j represents the segmentation part, Concat represents the feature concatenation operation in the channel dimension, X fused , X compressed , and X out They represent the feature map after feature fusion, the feature map after feature compression, and the output feature map respectively.

[0085] The multi-scale progressive feature aggregation module integrates local-global hierarchical convolution with a selective channel computation strategy, using a progressive channel reduction approach to efficiently extract and fuse features from different receptive fields. This multi-path architecture enables adaptive characterization of fine-grained lesion contours and comprehensive morphological analysis, improving multi-scale feature representation capabilities while reducing computational costs.

[0086] The multi-scale progressive feature aggregation module of this embodiment achieves cross-scale feature fusion while reducing computational complexity and preserving essential features. This is used to detect and segment multi-scale lesions that are difficult to distinguish from surrounding tissue. The proposed cross-stage feature fusion module balances model accuracy with reduced model complexity.

[0087] Furthermore, a coordinate-guided fusion module is introduced in the object detection head’s process of identifying caries, wisdom teeth, and periradicular lesions;

[0088] The coordinate-guided fusion module is used to combine the edge and texture information of shallow features with the semantic information of deep features, and dynamically adjust the feature expression through the coordinate attention mechanism; it enhances the feature modeling ability of the target area when small-scale lesions are confused with background structures.

[0089] Furthermore, the coordinate-guided fusion module combines the edge and texture information of shallow features with the semantic information of deep features, and dynamically adjusts the feature expression through the coordinate attention mechanism. The process includes:

[0090] First, the shallow features and deep features are dimensionally adjusted and preliminarily fused. The mapping function of the 1×1 convolution transformation is defined to align the number of feature channels. Then, the adjusted shallow features and deep features are concatenated along the channel dimension to generate fused features. Through the local-global interaction in the spatial dimension, an adaptive attention map is generated to optimize the feature expression.

[0091] Furthermore, the process of generating an adaptive attention map to optimize feature expression through local-global interaction in the spatial dimension includes:

[0092] First, the input features are locally aggregated, and then the spatial dependencies in the horizontal and vertical directions are modeled using separable directional convolution through initial feature transformation. Finally, the attention map is generated through feature transformation and normalization.

[0093] Specifically, in multi-task dental image analysis, the features of caries, wisdom teeth and periradicular lesions are distributed in different scale spaces. Their feature expression is not only affected by background noise, but also suffers from the problem of feature weakening. The coordinate guided fusion module proposed in this embodiment is as follows: Figure 3 As shown in (c), through the improved multi-scale feature fusion and coordinate attention optimization mechanism, the feature expression is dynamically adjusted to enhance the feature modeling ability of the target area.

[0094] The coordinate-guided fusion module first fusions shallow features and deep features Perform dimension adjustment and preliminary fusion. Shallow features contain rich edge and texture information, while deep features contain stronger semantic representation.

[0095] In order to align the number of feature channels, a mapping function φ is defined, which is a 1×1 convolution transformation:

[0096] X l′ =φ(X l )=W a *X l +b a

[0097] in, is the convolution kernel weight, is the bias, and the symbol * represents the convolution operation. The shallow features after adjustment and deep features X h Generate fusion features by splicing along the channel dimension:

[0098]

[0099] This fused feature provides a unified feature space input for subsequent feature modeling.

[0100] The coordinate-aware attention module generates an adaptive attention map through local-global interaction in the spatial dimension to optimize feature expression. First, local information aggregation is performed on the input features:

[0101]

[0102] Where k = 7 is the pooling kernel size. Then, through the initial feature transformation: F1 = Conv(P), the spatial dependencies in the horizontal and vertical directions are modeled using separable directional convolution:

[0103] F h =Convh (F1,K h ),F v =Conv v (F h ,K v )

[0104] where K h , K v are the convolution kernel sizes in the horizontal and vertical directions, respectively.

[0105] Finally, the attention map is generated by feature transformation and normalization: A = σ(Conv(F v ));

[0106] Among them, σ is the Sigmoid activation function.

[0107] Generated attention map Applied to the original feature: X a =A⊙X f ;

[0108] where ⊙ represents element-wise multiplication.

[0109] Weighted feature X a The key area features are highlighted while suppressing the interference of redundant features. In order to further improve the interaction between shallow and deep features, the recursive enhancement of multi-scale features is achieved through feature segmentation and fusion operations. The weighted fusion feature X a Divide along the channel dimension into feature parts corresponding to the shallow and deep layers:

[0110] X l,a =X a [1:C h ,:,:],X h,a =X a [C h :2C h ,:,:]

[0111] The segmented features are multiplied with the original features by element-wise multiplication to generate interactive enhancement features:

[0112] X l′ =X l′ ⊙X l,a ,X h′ =X h ⊙X h,a

[0113] The final module output is generated by cross-adding shallow and deep features:

[0114] X out =Concat(X l′ +Xh ,X h′ +X l′ )

[0115] This output characteristic It combines the fine-grained spatial information of shallow features and the semantic information of deep features.

[0116] The coordinate-guided fusion module employs a directional feature enhancement algorithm to amplify signal representations in low-density areas. By combining a coordinate-aware mechanism with a spatial attention matrix, the module effectively models the complex positional relationships between adjacent structures. Its technical architecture implements multi-scale feature aggregation with position-encoding convolutions, enabling the differentiation of lesions in specific locations from surrounding background tissue. In detection tasks, CGFM's gradient-sensitive feature extraction technology effectively enhances the extraction of subtle structural features of the periodontal space, and its coordinate-guided feature amplification mechanism improves lesion detection in small-scale scenarios.

[0117] The coordinate-guided fusion module of this embodiment accurately captures subtle pathological features that are easily confused with background structures by fusing low-level and high-level features and dynamically enhancing feature expression.

[0118] Furthermore, the instance segmentation head segments the alveolar bone, CEJ, and tooth contours through feature stitching, edge enhancement, and spatial attention mechanism, while preserving fine-grained structural information.

[0119] Specifically, in dental image segmentation tasks such as teeth, alveolar bone boundaries, and enamel-cementum junction (CEJ), feature maps at different network layers contain rich multi-scale information. Figure 3 As shown in Figure (e), this embodiment develops the Edge Enhanced Spatial Fusion (ESFM) module as a lightweight solution to effectively fuse these features. ESFM combines feature concatenation, edge enhancement, and a spatial attention mechanism to achieve efficient fusion and dynamic weighting of multiple layers of features, improving the model's ability to express target region and boundary features.

[0120] Given a set of input features from different layers in a multi-task segmentation network Each of these features First, the ESFM module concatenates these feature maps along the channel dimension to form fused features:

[0121]

[0122] in This operation can aggregate semantic information and detail information at different levels, providing rich feature expressions for subsequent processing.

[0123] In order to enhance the perception of boundary information (such as alveolar bone, CEJ and other anatomical structures), ESFM introduces an edge enhancement mechanism on the fusion feature. Specifically, a fixed-weight Sobel convolution kernel is used to detect edges on the fusion feature. The Sobel convolution kernel is defined as:

[0124]

[0125] Extract edge features through depth convolution operation:

[0126] X edge =φ edge (X cat ),φ edge (X cat )=K sobel *X cat

[0127] Then, the edge feature and the fusion feature are fused according to the weighting coefficient α = 0.1 to generate the edge enhancement feature:

[0128] X enhanced =X cat +α·X edge

[0129] This operation can highlight the boundary information of the target area, especially significantly improving the accuracy in large-scale alveolar bone and CEJ segmentation.

[0130] Feature map after edge enhancement Usually contains a lot of redundant information, so ESFM introduces a lightweight spatial attention mechanism to dynamically weight important target areas. Specifically, first calculate the global average pooling and maximum pooling features for each channel:

[0131]

[0132] in Then, the two are concatenated into a single feature map along the channel dimension:

[0133]

[0134] The concatenated feature is input into a lightweight 3×3 convolution to generate spatial attention weights:

[0135] W spatial (h,w)=σ(φ spatial (X spatial ))

[0136] Among them, φ spatial (·) is a 3×3 convolution operation, and σ(·) is a Sigmoid activation function used to normalize the weights. Finally, spatial dynamic weighting is achieved through point-by-point weighting operations:

[0137] X weighted (c,h,w)=X enhanced (c,h,w)·W spatial (h,w),

[0138] This mechanism can effectively focus on the target area while suppressing background noise.

[0139] Finally, in order to reduce the channel dimension and enhance the expressibility of features, ESFM performs channel compression on the weighted features through a 1×1 convolution layer:

[0140]

[0141] in is the convolution kernel, is the bias term, Module output.

[0142] ESFM captures tooth boundary features while preserving fine structural details, excelling in tasks requiring precise delineation. This module employs the Sobel operator to extract gradient information and combines this with an adaptive spatial attention weighting mechanism to enhance boundary contrast while suppressing irrelevant background noise. This two-stream architecture fuses edge-aware features with spatial context via residual connections, enabling precise identification of morphological transformations, which is crucial for accurate segmentation of anatomical structures with complex boundaries.

[0143] The edge-enhanced spatial fusion module of this embodiment retains fine-grained structural information through feature splicing, edge enhancement, and spatial attention mechanisms, thereby improving the effective segmentation of periodontitis staging pathological indicators and other oral disease pathological indicators.

[0144] Furthermore, the pathological feature integration analysis module of this embodiment integrates target detection and instance segmentation information to achieve comprehensive oral multi-disease detection. It performs dual functions: computational periodontitis grading (ICPS) and multi-disease detection (MDD), and can identify individual teeth affected by multiple diseases at the same time, including caries, wisdom teeth, periradicular lesions, and periodontitis. It is particularly noteworthy that this module uses special markings for teeth showing multiple pathologies to intuitively highlight specific lesion areas, such as Figure 5 demonstrated, promoting more targeted and effective clinical decision-making.

[0145] In assessing periodontitis grading, this module applies the clinical knowledge that healthy alveolar bone typically lies 1-2 mm below the cementoenamel junction (CEJ). Any deviation from this anatomical relationship indicates possible bone resorption. Based on established periodontitis classification criteria, the system systematically calculates the radiographic bone position (RBL) ratio to objectively categorize disease severity.

[0146] The computational process begins by precisely localizing the alveolar bone, CEJ, and overall tooth structure for each tooth through instance segmentation. By correlating the CEJ position with the image's calibrated scale information, the system establishes an expected normal range for alveolar bone positioning. For teeth that appear outside this normal range, the system applies the following quantitative assessments:

[0147]

[0148] where d CEJ-AB represents the distance from CEJ to alveolar bone (in millimeters), R L Indicates the root length of the tooth.

[0149] In these measurements, the system calculates and distances based on the growth direction axis of the teeth. To accurately determine the growth direction of each tooth, such as Figure 4 As shown, this embodiment uses a complex multi-step process:

[0150] (1) First, the segmented tooth data is filtered to exclude wisdom teeth and impacted teeth that may have abnormal directions due to atypical positioning.

[0151] (2) Cubic spline interpolation is used to smooth tooth boundary points to ensure uniform distribution and prevent density irregularities from affecting subsequent calculations.

[0152] (3) Principal component analysis (PCA) was then applied to extract the principal direction vectors, and the tooth growth axis was determined by eigendecomposition of the covariance matrix.

[0153] After calculating the relevant indicators along the tooth growth axis, the system uses the root bone loss (RBL) ratio to stage periodontitis according to the criteria defined by Berglundh et al.

[12] :

[0154] Stage 1 (mild): RBL <15%;

[0155] Stage 2 (moderate): RBL between 15% and 33%;

[0156] Stage 3 (severe): RBL≥33%.

[0157] like Figure 4 As shown in the figure, different colors are used to visually display the periodontitis stage of each tooth on a tooth-by-tooth basis. Figure 5As shown, since teeth with multiple types of lesions usually require different treatment methods than single lesions, this embodiment uses unique color markings for teeth affected by multiple oral pathologies in the visualization results to enhance diagnostic support. The hyperparameter settings are shown in Table 1.

[0158] Table 1

[0159]

[0160]

[0161] Performance evaluation indicators:

[0162] In this embodiment, for the target detection task, mAP50, mAP50:95, parameters, and computational complexity (GFLOPs) are selected as detection indicators for positioning performance; for the instance segmentation task, mIoU, parameters, and computational complexity (GFLOPs) are selected as indicators for evaluating instance segmentation performance.

[0163] Model Validation: Finally, the model's performance was verified on the collected dataset. The experimental results for the object detection task are shown in Table 2. Compared to the baseline model, the MMC-YOLO method improved mAP50 for caries detection from 80.6% to 81.8%, an increase of 1.2 percentage points, and mAP50:95 from 47.9% to 48.6%. For wisdom tooth detection, mAP50 increased from 96.2% to 97.0%, and mAP50:95 increased from 80.7% to 81.2%. For periradicular lesion detection, mAP50 increased from 84.1% to 84.6%, and mAP50:95 increased from 61.7% to 61.9%. Furthermore, as shown in Tables 3-4, experimental results on instance segmentation tasks demonstrate that, compared to A-YOLO-M, MMC-YOLO improves mIoU (millisecond intersection over union) in alveolar bone segmentation from 96.8% to 97.4%, mIoU (millisecond intersection over union) in enamel-cementum junction segmentation from 95.3% to 96.1%, and mIoU (millisecond intersection over union) in tooth segmentation from 92.9% to 93.6%. MMC-YOLO also achieved significant improvements in periodontitis classification, with accuracy increasing from 77.3% to 86.7% for the no periodontitis category, from 86.7% to 91.3% for mild periodontitis, from 72.6% to 77.2% for moderate periodontitis, and from 79.0% to 84.7% for severe periodontitis. At the same time, the computational complexity of MMC-YOLO was reduced from 78GFLOPs to 76GFLOPs, and the number of parameters was reduced from 14.9M to 13.3M, achieving the goal of reducing resource consumption while improving detection and segmentation performance, proving that this method has excellent performance in practical applications.

[0164] Table 2

[0165]

[0166]

[0167] Table 3

[0168]

[0169] Table 4

[0170]

[0171] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. An oral panoramic image processing system, characterized in that: include: An image acquisition module, used for collecting initial oral panoramic images containing various oral diseases; An image processing module, connected to the image acquisition module, is used to perform target detection and instance segmentation annotation on the initial oral panoramic image, and to expand the training data set using data enhancement technology to obtain a target oral panoramic image; A model building module, connected to the image processing module, is used to train the MMC-YOLO network model using the target oral panoramic image to obtain a target MMC-YOLO network model; An analysis and detection module, connected to the model building module, is used to perform simultaneous detection of dental complications and oral diseases, comprehensive analysis of multi-dimensional pathological information, and accurate identification of multi-scale lesions through the target MMC-YOLO network model to obtain analysis results; The target MMC-YOLO network model includes a shared backbone network, a target detection neck network, a target detection detection head, an instance segmentation neck network, an instance segmentation segmentation head, and a multi-task information comprehensive calculation and analysis module; Wherein, the shared backbone network is used to extract common features of oral panoramic images; The target detection neck network is used to process the features of the target detection task; The target detection head is used to identify caries, wisdom teeth and periradicular lesions; The instance segmentation neck network is used to process the features of the instance segmentation task; The segmentation head of the example segmentation is used to segment the alveolar bone, the enamel-cementum junction, and the tooth contour; The multi-task information comprehensive calculation and analysis module is used to integrate target detection and instance segmentation results to obtain periodontitis grading results and multi-disease detection results; The shared backbone network and the target detection neck network include a multi-scale progressive feature aggregation module; the multi-scale progressive feature aggregation module is used to integrate local and global features through a progressive multi-scale feature extraction strategy and partial convolution operation; and adopts channel segmentation and progressive channel reduction methods to retain key features and reduce computational complexity; The target detection head introduces a coordinate-guided fusion module during the process of identifying caries, wisdom teeth and periradicular lesions; The coordinate-guided fusion module is used to combine the edge and texture information of shallow features with the semantic information of deep features, and dynamically adjust the feature expression through the coordinate attention mechanism, including: First, the shallow features and deep features are dimensionally adjusted and preliminarily fused, and the mapping function of the 1×1 convolution transformation is defined to align the number of feature channels; Then the adjusted shallow features and deep features are spliced ​​along the channel dimension to generate fusion features. Through the local-global interaction of the spatial dimension, an adaptive attention map is generated to optimize the feature expression, including: first, the fusion features of the input Perform local information aggregation and express it as P; then through the initial feature transformation, it is ; Separate directional convolution is then used to model the spatial dependencies in the horizontal and vertical directions, which are expressed as: ,in 、 are the convolution kernel sizes in the horizontal and vertical directions respectively; finally, the attention map is generated by feature transformation and normalization, which is expressed as ,in, is the Sigmoid activation function; the generated attention map A is used to obtain the weighted feature Xa, ,in, Represents element-by-element multiplication, generating the feature expression of the final output based on Xa.

2. The system according to claim 1, wherein: The target detection and instance segmentation annotation include target detection and segmentation annotation of lesion type, location and boundary information; The data augmentation includes rotation, scaling and flipping.

3. The system according to claim 1, wherein: The multi-scale progressive feature aggregation module integrates local and global features through a progressive multi-scale feature extraction strategy and partial convolution operation, including the following steps: Given an input feature, a standard 3×3 convolution operation is first performed, and then the feature is divided into two parts in the channel dimension. The first part extracts medium-scale features through 5×5 convolution, and a quarter of the channel extracts global features through 7×7 convolution; finally, the multi-scale features are fused through channel splicing and transformation.

4. The system according to claim 1, wherein: The instance segmentation head segments the alveolar bone, CEJ, and tooth contours through feature stitching, edge enhancement, and spatial attention mechanism, while preserving fine-grained structural information.