Oral panoramic image processing system

Through the MMC-YOLO network model combined with object detection and instance segmentation technology, the subjective differences in pathological information extraction and multi-scale lesion recognition in oral panoramic imaging are solved, periodontitis grading and synchronous detection of multiple concurrent diseases are realized, and the accuracy and efficiency of diagnosis are improved.

CN120278989AActive Publication Date: 2025-07-08GUANGDONG UNIV OF TECH

Patent Information

Application Number
CN202510431918.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-08
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The existing oral panoramic imaging technology has subjective differences in pathological information extraction and diagnosis, difficulty in quantifying the severity of periodontitis, and inability to detect multiple complications of pathology and deal with multi-scale lesion characteristics at the same time.

Method used

The MMC-YOLO network model is adopted, combined with object detection and instance segmentation technology, and through the multi-task information comprehensive analysis module, the alveolar bone, enamel-cephalon boundary and tooth position characteristics are integrated to realize periodontitis grading and synchronous detection of multiple concurrent diseases.

Benefits of technology

It improves the precise classification of the severity of periodontitis and the detection ability of multiple concurrent diseases, enhances the recognition ability of multi-scale lesions, and reduces the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278989A_ABST
    Figure CN120278989A_ABST
Patent Text Reader

Abstract

The invention discloses an oral cavity panoramic image processing system, and the system comprises an image collection module which is used for collecting an initial oral cavity panoramic image containing a plurality of oral cavity diseases; the image processing module is connected with the image acquisition module and is used for performing target detection and instance segmentation labeling on the initial oral panoramic image and expanding a training data set by adopting a data enhancement technology to obtain a target oral panoramic image; a model construction module connected with the image processing module and used for training an MMC-YOLO network model through the target oral panoramic image to obtain a target MMC-YOLO network model; and the analysis and detection module is connected with the model construction module, and is used for performing synchronous detection of tooth concurrent oral diseases, comprehensive analysis of multi-dimensional pathological information and accurate recognition of multi-scale pathological changes through the target MMC-YOLO network model to obtain an analysis result. The system provided by the invention realizes synchronous detection and accurate grading of multiple types of focuses in the oral panoramic image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and particularly relates to an oral panoramic image processing system. Background Art

[0002] Oral panoramic imaging technology is crucial for the early detection, diagnosis, and prevention of oral diseases. Clinically, two main panoramic imaging techniques are applied: traditional X-ray panoramic imaging and CBCT panoramic imaging. Traditional X-ray panoramic imaging uses a rotating X-ray source and a synchronously moving sensor to obtain continuous planar projections of the oral and maxillofacial region through slit collimation. This technique is easy to operate, has a short examination time, and provides clinicians with an overall overview of anatomical structures such as teeth, jaws, and temporomandibular joints at a relatively low radiation dose and economic cost. CBCT panoramic imaging, on the other hand, uses a cone-shaped X-ray beam to collect multi-angle two-dimensional projection data, which is then reconstructed into high-resolution images by computer processing. The images generated by CBCT technology have high contrast, can accurately display the root and periodontal structures; geometric distortion is significantly reduced, ensuring the accuracy of anatomical proportions; at the same time, tissue overlap is greatly reduced, providing more detailed oral physiological information. Both of these panoramic imaging techniques can provide rich oral panoramic information and are of great value for the diagnosis of various oral diseases such as dental caries, wisdom teeth, periapical lesions, periodontitis, etc.

[0003] Although oral panoramic imaging technology can provide rich oral physiological information, effective extraction of pathological information from it is required to achieve accurate diagnosis. Traditional manual interpretation methods face many limitations in this regard. Clinicians mainly rely on professional experience to evaluate the morphological and radiographic density characteristics of teeth, alveolar bone, and gingival tissues. Inevitably, there are subjective differences in this process. At the same time, manual interpretation is not sensitive enough to subtle pathological changes such as early dental caries or occult root lesions, which is prone to clinical misdiagnosis. More critically, traditional methods lack effective means to quantify the severity of certain oral diseases (such as periodontitis), making it difficult to provide objective and consistent disease grading evaluations, increasing the risk of diagnostic omissions or inaccuracies. Although computer-aided diagnosis (CAD) systems based on traditional image processing techniques (including edge detection and morphological operations) have achieved partial automation, these systems rely on artificially set feature extraction rules, have limited generalization ability, cannot fully handle the complex and diverse lesion morphologies in panoramic images, and are difficult to meet the accurate diagnosis requirements of modern oral medicine.

[0004] In recent years, deep learning technology has made remarkable progress in the field of medical image analysis. Methods based on the convolutional neural network architecture have shown superior performance compared to traditional methods in specific applications such as dental caries detection and tooth segmentation, significantly improving the accuracy of lesion localization and quantification. However, current deep learning-based dental panoramic image analysis systems still have limitations in many aspects: First, most existing studies adopt a single-task framework, which restricts the ability to simultaneously detect multiple coexisting pathologies on a single tooth (such as confirming the presence of concurrent periodontitis while detecting dental caries); Second, existing systems are difficult to provide a quantitative assessment of complex diseases such as periodontitis because they cannot effectively integrate key measurement indicators such as alveolar bone level assessment, cementoenamel junction localization, and tooth position characteristics. Accurately quantifying the severity of periodontitis requires the model to be able to perform object detection and instance segmentation tasks simultaneously and integrate this information for comprehensive analysis; Third, existing methods perform poorly in dealing with pathological features of different scales, especially with limited recognition ability for lesion areas with subtle density changes. For example, when deep convolutional feature extraction processes interproximal dental caries, adjacent low-density radiolucent features are often misinterpreted as interdental space background.

[0005] In summary, for the current pathological assistant diagnosis system based on oral panoramic images, developing a comprehensive framework that can simultaneously detect the oral disease characteristics concurrent in teeth, detect oral diseases that require processing multi-dimensional pathological information, and adapt to multi-scale scenarios has become the core challenge in current research. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides an oral panoramic image processing system, including:

[0007] An image acquisition module for collecting initial oral panoramic images containing multiple oral diseases;

[0008] An image processing module connected to the image acquisition module for performing object detection and instance segmentation annotation on the initial oral panoramic images, and using data augmentation technology to expand the training data set to obtain target oral panoramic images;

[0009] A model construction module connected to the image processing module for training the MMC-YOLO network model through the target oral panoramic images to obtain a target MMC-YOLO network model;

[0010] An analysis and detection module connected to the model construction module for synchronously detecting oral diseases concurrent in teeth, comprehensively analyzing multi-dimensional pathological information, and accurately identifying multi-scale lesions through the target MMC-YOLO network model to obtain analysis results.

[0011] Preferably, the object detection and instance segmentation annotation includes object detection and segmentation annotation of lesion type, location, and boundary information;

[0012] The data augmentation includes rotation, scaling, and flipping.

[0013] Preferably, the target MMC-YOLO network model includes a shared backbone network, an object detection neck network, an object detection head, an instance segmentation neck network, an instance segmentation head, and a multi-task information comprehensive calculation and analysis module;

[0014] Among them, the shared backbone network is used to extract general features of the oral panoramic image;

[0015] The object detection neck network is used to process the features of the object detection task;

[0016] The object detection head is used to identify dental caries, wisdom teeth, and periapical lesions;

[0017] The instance segmentation neck network is used to process the features of the instance segmentation task;

[0018] The instance segmentation head is used to segment the alveolar bone, cementoenamel junction, and tooth contour;

[0019] The multi-task information comprehensive calculation and analysis module is used to integrate the object detection and instance segmentation results to obtain the periodontitis grading result and the multi-disease detection result.

[0020] Preferably, the shared backbone network and the object detection neck network include a multi-scale progressive feature aggregation module;

[0021] The multi-scale progressive feature aggregation module is used to integrate local and global features through a progressive multi-scale feature extraction strategy and partial convolution operations; and adopts a channel splitting and progressive channel reduction method to reduce the computational complexity while retaining key features.

[0022] Preferably, the process of integrating local and global features by the multi-scale progressive feature aggregation module through a progressive multi-scale feature extraction strategy and partial convolution operations includes:

[0023] Given the input features, first perform a standard 3×3 convolution operation, then divide the features into two parts in the channel dimension. The first part extracts medium-scale features through a 5×5 convolution, and one-fourth of the channels extract global features through a 7×7 convolution; finally, fuse the multi-scale features through channel concatenation and transformation.

[0024] Preferably, a coordinate-guided fusion module is introduced in the process of the object detection head identifying dental caries, wisdom teeth, and periapical lesions;

[0025] The coordinate-guided fusion module is used to combine the edge and texture information of shallow features with the semantic information of deep features, and dynamically adjust the feature representation through the coordinate attention mechanism; enhance the feature modeling ability of the target area in the case of confusion between small-scale lesions and background structures.

[0026] Preferably, the process of the coordinate-guided fusion module combining the edge and texture information of shallow features with the semantic information of deep features and dynamically adjusting the feature representation through the coordinate attention mechanism includes:

[0027] First, perform dimensionality adjustment and preliminary fusion on the shallow features and deep features, define the mapping function of the 1×1 convolution transformation to align the number of feature channels, then splice the adjusted shallow features and deep features along the channel dimension to generate fused features, and generate an adaptive attention map through local-global interaction in the spatial dimension to optimize the feature representation.

[0028] Preferably, the process of generating an adaptive attention map through local-global interaction in the spatial dimension to optimize the feature representation includes:

[0029] First, perform local information aggregation on the input features, then through the initial feature transformation, use separable directional convolution to model the spatial dependencies in the horizontal and vertical directions, and finally generate the attention map through feature transformation and normalization.

[0030] Preferably, the instance segmentation head segments the alveolar bone, CEJ, and tooth contour through feature splicing, edge enhancement, and spatial attention mechanism, while retaining fine-grained structural information.

[0031] Compared with the prior art, the present invention has the following advantages and technical effects:

[0032] The oral panoramic image processing system of the present invention combines object detection and instance segmentation, aggregates the output information of multiple tasks into the designed pathological feature integration and analysis module, thereby realizing effective classification of the severity of periodontitis, and can also perform special annotation on other oral diseases concurrent in the same tooth by combining multi-task information.

[0033] Through the strategy of multi-scale progressive feature aggregation, the present invention realizes cross-scale feature fusion, reduces the computational complexity, and retains the basic features for detecting and segmenting multi-scale lesions that are difficult to distinguish from surrounding tissues.

[0034] The present invention dynamically enhances the feature representation of shallow and deep features, while accurately capturing subtle pathological features that are easily confused with background structures. It enhances the perception of dental caries and periapical lesion foci in the early small-scale scenario.

[0035] The present invention preserves fine-grained structural information through feature splicing, edge enhancement, and spatial attention mechanisms, thereby improving the effective segmentation of tooth edges and enhancing the model's ability to distinguish between blurred boundaries and the background. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0037] Figure 1 It is a schematic diagram of the system working process of the embodiment of the present invention;

[0038] Figure 2 It is a framework diagram of MMC-YOLO of the embodiment of the present invention;

[0039] Figure 3 It is a schematic diagram of the principles of each module in the target MMC-YOLO network model of the embodiment of the present invention;

[0040] Figure 4 It is a diagram of the periodontitis grading results output by the MMC-YOLO framework of the embodiment of the present invention;

[0041] Figure 5 It is a diagram of the comprehensive marker results of multi-task information fusion and concurrent diseases output by the MMC-YOLO framework of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0043] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0044] In view of the complex detection requirements of multiple types of lesions in oral panoramic images in this embodiment, an innovative multi-task collaboration framework is proposed, aiming to solve three key challenges: the synchronous detection of concurrent oral diseases in teeth, the comprehensive analysis of multi-dimensional pathological information, and the accurate identification of multi-scale lesions.

[0045] This framework constructs an end-to-end oral lesion detection system by organically integrating object detection and instance segmentation technologies. Under this architecture, in this embodiment, a pathological feature integration and analysis module is designed, which can simultaneously complete the feature extraction of dental concurrent diseases and the accurate grading of the severity of periodontitis. To enhance the detection ability of the model in multi-scale lesion scenarios, three synergistic functional modules are further proposed: a multi-scale progressive feature aggregation module, a coordinate-guided fusion module, and an edge-enhanced spatial fusion module. These modules improve the performance of the network in the detection of variant-scale lesions and fine segmentation tasks through a hierarchical feature extraction and fusion strategy.

[0046] This system aims to solve the following problems:

[0047] Simultaneous detection of dental concurrent oral diseases: By integrating object detection and instance segmentation into a multi-task network, the multi-task information output is passed through the designed pathological feature integration and analysis module, and the grading of periodontitis is achieved according to the pathological information required for periodontitis grading (alveolar bone position feature, enamel-cementum junction position feature, tooth position feature), and at the same time, the conditions of other concurrent oral diseases (dental caries, wisdom teeth, periapical lesions) on the corresponding teeth are detected.

[0048] Enhanced extraction of oral disease lesion features in multi-scale scenarios: For the pathological features with complex positions and scales in panoramic images, we developed a multi-scale progressive feature aggregation module to effectively utilize the feature information of small-scale lesions. We designed a coordinate-guided fusion module to enhance the discrimination ability between small-scale lesions and the background. We also proposed an edge-enhanced spatial fusion module to improve the feature fusion effect of segmented objects with fuzzy edge features.

[0049] As Figures 1-5 shown, in this embodiment, an oral panoramic image processing system is provided, including:

[0050] An image acquisition module for collecting initial oral panoramic images containing various oral diseases;

[0051] An image processing module, connected to the image acquisition module, for performing object detection and instance segmentation annotation on the initial oral panoramic image, and using data augmentation technology to expand the training data set to obtain a target oral panoramic image;

[0052] A model construction module, connected to the image processing module, for training the MMC-YOLO network model through the target oral panoramic image to obtain a target MMC-YOLO network model;

[0053] The analysis and detection module, connected to the model construction module, is used to synchronously detect dental concurrent oral diseases, comprehensively analyze multi-dimensional pathological information, and accurately identify multi-scale lesions through the target MMC-YOLO network model to obtain analysis results.

[0054] Furthermore, the object detection and instance segmentation annotation include object detection and segmentation annotation of lesion type, location, and boundary information;

[0055] Data augmentation includes rotation, scaling, and flipping.

[0056] Furthermore, the target MMC-YOLO network model includes a shared backbone network, an object detection neck network, an object detection head, an instance segmentation neck network, an instance segmentation head, and a multi-task information comprehensive calculation and analysis module;

[0057] Among them, the shared backbone network is used to extract general features of oral panoramic images;

[0058] The object detection neck network is used to process the features of the object detection task;

[0059] The object detection head is used to identify dental caries, wisdom teeth, and periapical lesions;

[0060] The instance segmentation neck network is used to process the features of the instance segmentation task;

[0061] The instance segmentation head is used to segment the alveolar bone, CEJ, and tooth contour;

[0062] The multi-task information comprehensive calculation and analysis module is used to integrate the object detection and instance segmentation results to obtain the periodontitis grading result and multi-disease detection result.

[0063] Specifically, the target MMC-YOLO network model consists of six core parts with clear functions: a shared backbone network (Backbone), an object detection neck network (Detect neck), an object detection head (Detect head), an instance segmentation neck network (Segment neck), an instance segmentation head (Segment head), and a multi-task information comprehensive calculation and analysis module. The overall architecture is as Figure 2As shown in the figure. In the shared backbone network and the object detection neck network, in this embodiment, a multi-scale progressive feature aggregation module (MPFAM) is designed. This module is specifically designed to handle multi-task detection scenarios with significant scale differences in oral images, such as from tiny dental caries to large-scale alveolar bone boundaries. In the object detection feature processing link, a coordinate-guided fusion module (CGFM) is introduced. By fusing multi-scale features and combining a global context awareness mechanism, the feature expression ability and discrimination accuracy of the target area are significantly improved. For the feature processing requirements of instance segmentation, in this embodiment, an edge enhancement spatial fusion module (ESFM) is developed. This module effectively improves the utilization efficiency of the common fuzzy boundary features in oral images through the organic combination of feature splicing, edge enhancement, and spatial attention mechanisms. In the output part of the multi-task network, the object detection task detection head focuses on accurately identifying three types of key oral lesion features: dental caries, wisdom teeth, and periapical lesions, which are different in scale and manifestation forms. The instance segmentation task is completed by three specially designed segmentation heads, which are responsible for segmenting the alveolar bone contour, the enamel-cementum junction, and the complete tooth contour respectively. The final link of the entire analysis process is the pathological feature integration and analysis module. This module systematically fuses the results of object detection and instance segmentation, can not only grade the severity of periodontitis, but also comprehensively evaluate the lesion status of a single tooth, and specially mark the teeth with multiple lesions to obtain the marking results.

[0064] The multi-task framework proposed in this embodiment grades the severity of periodontitis and labels other concurrent oral diseases based on the comprehensive pathological information of object detection and instance segmentation combined with the PCA tooth growth axis positioning method.

[0065] Furthermore, the shared backbone network and the object detection neck network include a multi-scale progressive feature aggregation module;

[0066] The multi-scale progressive feature aggregation module is used to integrate local and global features through a progressive multi-scale feature extraction strategy and partial convolution operations; and adopts a channel segmentation and progressive channel reduction method to reduce the computational complexity while retaining key features.

[0067] Furthermore, the process of integrating local and global features by the multi-scale progressive feature aggregation module through a progressive multi-scale feature extraction strategy and partial convolution operations includes:

[0068] Given the input features, first perform a standard 3×3 convolution operation, then divide the features into two parts in the channel dimension. The first part extracts medium-scale features through a 5×5 convolution, and one-fourth of the channels extract global features through a 7×7 convolution; finally, fuse the multi-scale features through channel splicing and transformation.

[0069] Specifically, it should be noted that Figure 3Among them, (a) is the grouped convolution of MPFAM. (b) is the multi-scale progressive feature aggregation module (MPFAM), where c_in represents the input channels. (c) is the coordinate attention module of CGFM. (d) is the coordinate-guided fusion module (CGFM). (e) is the edge-enhanced spatial fusion module (ESFM), where σ(w) is the weight, τ is the weight threshold, and α is the edge feature scaling factor.

[0070] To effectively handle the multi-scale feature integration problem in dental panoramic image analysis, this embodiment proposes a multi-scale progressive feature aggregation module (MPFAM) as shown in Figure 3 (b). This module achieves a balance between computational efficiency and feature expression ability through a progressive multi-scale feature extraction strategy and partial convolution operations.

[0071] Specifically, given the input feature where C represents the number of channels, and H and W represent the height and width of the feature map respectively.

[0072] First, a standard 3×3 convolution operation is performed:

[0073]

[0074] Then, the feature is divided into two parts in the channel dimension:

[0075]

[0076] Next, the first part extracts medium-scale features through a 5×5 convolution:

[0077]

[0078] And, one-fourth of the channels extract global features through a 7×7 convolution:

[0079]

[0080] Finally, these multi-scale features are fused through channel concatenation and transformation:

[0081] X fused = Concat(X3, X 2,2 , X 1,2 )

[0082]

[0083] X out = X compressed + X

[0084] Among them, φ1, φ2, φ3, φ4 represent convolutional operations with ReLU activation, W1, W2, W3 are learnable weight matrices for convolutions of different scales, b1, b2, b3 are corresponding bias terms, and X i,j represents the segmented feature map, i represents the layer number, j represents the segmented part, Concat represents the feature concatenation operation in the channel dimension, and X fused , X compressed , and X out represent the feature map after feature fusion, the feature map after feature compression, and the output feature map respectively.

[0085] The multi-scale progressive feature aggregation module integrates local-global hierarchical convolution and selective channel calculation strategies, and uses a progressive channel reduction method to efficiently extract and fuse features with different receptive fields. This multi-path architecture can adaptively represent fine-grained lesion contours and comprehensive morphological analysis, while reducing the computational cost and improving the multi-scale feature representation ability.

[0086] The multi-scale progressive feature aggregation module of this embodiment reduces the computational complexity while realizing cross-scale feature fusion, and retains the basic features for detecting and segmenting multi-scale lesions that are difficult to distinguish from surrounding tissues. The proposed cross-stage feature fusion module reduces the complexity of the model while taking into account the model accuracy.

[0087] Furthermore, a coordinate-guided fusion module is introduced during the process of the object detection head identifying dental caries, wisdom teeth, and periapical lesions;

[0088] The coordinate-guided fusion module is used to combine the edge and texture information of shallow features with the semantic information of deep features, and dynamically adjust the feature expression through the coordinate attention mechanism; enhance the feature modeling ability of the target area in the case of confusion between small-scale lesions and background structures.

[0089] Furthermore, the process of the coordinate-guided fusion module combining the edge and texture information of shallow features with the semantic information of deep features and dynamically adjusting the feature expression through the coordinate attention mechanism includes:

[0090] First, perform dimensional adjustment and preliminary fusion on the shallow features and deep features, define the mapping function of the 1×1 convolution transformation to align the number of feature channels, then concatenate the adjusted shallow features and deep features along the channel dimension to generate fused features, and generate an adaptive attention map through local-global interaction in the spatial dimension to optimize the feature expression.

[0091] Furthermore, the process of generating an adaptive attention map through local-global interaction in the spatial dimension to optimize the feature expression includes:

[0092] First, local information aggregation is performed on the input features. Then, through an initial feature transformation, separable directional convolutions are used to model the spatial dependencies in the horizontal and vertical directions. Finally, an attention map is generated through feature transformation and normalization.

[0093] Specifically, in multi-task dental image analysis, the lesion features of dental caries, wisdom teeth, and periapical lesions are distributed in different-scale spaces. Their feature expressions are not only interfered by background noise but also suffer from feature weakening problems. The coordinate-guided fusion module proposed in this embodiment, as shown in Figure 3 (c), dynamically adjusts the feature expression and enhances the feature modeling ability for the target region through improved multi-scale feature fusion and coordinate attention optimization mechanisms.

[0094] The coordinate-guided fusion module first performs dimensionality adjustment and preliminary fusion on shallow features and deep features . Shallow features contain rich edge and texture information, while deep features contain stronger semantic representations.

[0095] To align the number of feature channels, a mapping function φ is defined, which is a 1×1 convolutional transformation:

[0096] X l′ = φ(X l ) = W a * X l + b a

[0097] where is the convolutional kernel weight, is the bias, and the symbol * represents the convolutional operation. The adjusted shallow features and the deep feature X h are concatenated along the channel dimension to generate a fused feature:

[0098]

[0099] This fused feature provides a unified feature space input for subsequent feature modeling.

[0100] The coordinate-aware attention module generates an adaptive attention map to optimize the feature expression through local-global interaction in the spatial dimension. First, local information aggregation is performed on the input features:

[0101]

[0102] where k = 7 is the pooling kernel size. Then, through an initial feature transformation: F1 = Conv(P). Next, separable directional convolutions are used to model the spatial dependencies in the horizontal and vertical directions:

[0103] F h = Convh (F1, K h ), F v = Conv v (F h , K v )

[0104] where K h , K v are the convolution kernel sizes in the horizontal and vertical directions respectively.

[0105] Finally, the attention map is generated through feature transformation and normalization: A = σ(Conv(F v ));

[0106] where σ is the Sigmoid activation function.

[0107] The generated attention map is applied to the original feature: X a = A ⊙ X f ;

[0108] where ⊙ represents element-wise multiplication.

[0109] The weighted feature X a highlights the key region features and at the same time suppresses the interference of redundant features. To further enhance the interaction ability between shallow and deep features, recursive enhancement of multi-scale features is achieved through feature segmentation and fusion operations. The weighted fusion feature X a is split along the channel dimension into the corresponding feature parts of the shallow and deep layers:

[0110] X l,a = X a [1:C h , :, :], X h,a = X a [C h :2C h , :, :]

[0111] The segmented features are multiplied by the original features element-wise to generate the interaction-enhanced features:

[0112] X l′ = X l′ ⊙ X l,a , X h′ = X h ⊙ X h,a

[0113] The final module output is generated through the cross-adding operation of the shallow and deep features:

[0114] X out = Concat(X l′ + Xh ,X h′ +X l′ )

[0115] This output feature combines the fine-grained spatial information of shallow features and the semantic information of deep features.

[0116] The coordinate-guided fusion module adopts a directional feature enhancement algorithm to amplify the signal representation in low-density regions. By combining the coordinate-aware mechanism with the spatial attention matrix, this module effectively models the complex positional relationships between adjacent structures. Its technical architecture realizes multi-scale feature aggregation with position-encoded convolution, enabling the differentiation of lesions at special positions from the surrounding background tissues. In the detection task, the gradient-sensitive feature extraction technology of CGFM effectively enhances the extraction of subtle structural features of the periodontal space, and improves the detection of lesions in small-scale scenarios through its coordinate-guided feature amplification mechanism.

[0117] The coordinate-guided fusion module of this embodiment accurately captures subtle pathological features that are easily confused with background structures through underlying and high-level feature fusion and dynamic enhanced feature expression.

[0118] Furthermore, the instance segmentation head segments the alveolar bone, CEJ, and tooth contour through feature concatenation, edge enhancement, and spatial attention mechanisms, while retaining fine-grained structural information.

[0119] Specifically, in dental image segmentation tasks such as teeth, alveolar bone boundaries, and enamel-cementum junction (CEJ), the feature maps of different network layers contain rich multi-scale information. As Figure 3 (e) shows, this embodiment develops an Edge-enhanced Spatial Fusion (ESFM) module as a lightweight solution to effectively fuse these features. ESFM achieves efficient fusion and dynamic weighting of multi-layer features by combining feature concatenation, edge enhancement, and spatial attention mechanisms, improving the model's ability to express target regions and boundary features.

[0120] Given a set of input features from different layers in a multi-task segmentation network where each feature First, the ESFM module concatenates these feature maps along the channel dimension to form a fused feature:

[0121]

[0122] where This operation can aggregate semantic information and detail information at different levels, providing rich feature expressions for subsequent processing.

[0123] To enhance the perception ability of boundary information (such as anatomical structures like alveolar bone, CEJ, etc.), an edge enhancement mechanism is introduced in the fused features by ESFM. Specifically, a Sobel convolutional kernel with fixed weights is used to detect edges in the fused features. The Sobel convolutional kernel is defined as:

[0124]

[0125] Edge features are extracted through depth convolution operations:

[0126] X edge =φ edge (X cat ), φ edge (X cat )=K sobel *X cat

[0127] Then, the edge features and the fused features are fused according to the weighting coefficient α = 0.1 to generate edge-enhanced features:

[0128] X enhanced =X cat +α·X edge

[0129] This operation can highlight the boundary information of the target area, especially significantly improving the accuracy in large-scale alveolar bone and CEJ segmentation.

[0130] The feature map after edge enhancement usually contains a large amount of redundant information. Therefore, ESFM introduces a lightweight spatial attention mechanism to dynamically weight important target areas. Specifically, first, the global average pooling and max pooling features are calculated for each channel:

[0131]

[0132] where Subsequently, the two are concatenated along the channel dimension into a single feature map:

[0133]

[0134] This concatenated feature is input into a lightweight 3×3 convolution to generate spatial attention weights:

[0135] W spatial (h,w)=σ(φ spatial (X spatial ))

[0136] where, φ spatial (·) is a 3×3 convolution operation, and σ(·) is the Sigmoid activation function used to normalize the weights Finally, spatial dynamic weighting is achieved through point-by-point weighting operations:

[0137] X weighted (c, h, w) = X enhanced (c, h, w) · W spatial (h, w),

[0138] This mechanism can effectively focus on the target area while suppressing background noise.

[0139] Finally, to reduce the channel dimension and enhance the expressiveness of features, ESFM compresses the weighted features in channels through a 1×1 convolutional layer:

[0140]

[0141] where is the convolutional kernel, is the bias term, is the module output.

[0142] ESFM captures the tooth boundary features while preserving fine structural details, especially performing well in tasks that require precise delineation. This module uses the Sobel operator to extract gradient information and combines it with an adaptive spatial attention weighting mechanism to enhance the boundary contrast while suppressing irrelevant background noise. This dual-stream architecture fuses edge-aware features with the spatial background through residual connections, enabling precise identification of morphological transitions, which is crucial for the accurate segmentation of anatomical structures with complex boundaries.

[0143] The edge enhancement spatial fusion module of this embodiment preserves fine-grained structural information through feature stitching, edge enhancement, and spatial attention mechanisms, thereby improving the effective segmentation of periodontitis staging pathological indicators and other oral disease pathological indicators.

[0144] Furthermore, the pathological feature integration and analysis module of this embodiment integrates object detection and instance segmentation information to achieve comprehensive detection of multiple oral diseases. It performs dual functions: computational periodontitis staging (ICPS) and multiple disease detection (MDD), and can identify individual teeth affected by multiple diseases simultaneously, including dental caries, wisdom teeth, periapical lesions, and periodontitis. Notably, this module uses a special marking for teeth showing multiple pathologies, intuitively highlighting specific lesion areas, as Figure 5 shown, facilitating more targeted and effective clinical decision-making.

[0145] In the grading assessment of periodontitis, this module applies clinical knowledge that healthy alveolar bone is usually located 1-2 millimeters below the cementoenamel junction (CEJ). Any deviation from this anatomical relationship indicates possible bone resorption. According to the established periodontitis classification criteria, the system calculates the radiographic bone level (RBL) ratio to objectively classify the disease severity.

[0146] The calculation process first precisely locates the alveolar bone, CEJ, and overall tooth structure of each tooth through instance segmentation. By correlating the CEJ position with the calibration scale information of the image, the system establishes the expected normal range for alveolar bone localization. For teeth showing beyond this normal range, the system applies the following quantitative assessment:

[0147]

[0148] where d CEJ-AB represents the distance from the CEJ to the alveolar bone (in millimeters), and R L represents the root length of the tooth.

[0149] In these measurements, the system calculates distances based on the growth direction axis of the tooth. To accurately determine the growth direction of each tooth, as Figure 4 shown, this embodiment employs a complex multi-step process:

[0150] (1) First, filter the segmented tooth data to exclude wisdom teeth and impacted teeth that may cause abnormal directions due to atypical positioning.

[0151] (2) Use cubic spline interpolation to smooth the tooth boundary points to ensure uniform distribution and prevent density irregularities from affecting subsequent calculations.

[0152] (3) Then apply principal component analysis (PCA) to extract the principal direction vector and determine the tooth growth axis through eigenvalue decomposition of the covariance matrix.

[0153] After calculating the relevant metrics along the tooth growth direction axis, the system stages periodontitis using the root bone loss (RBL) ratio according to the criteria defined by Berglundh et al.

[12] :

[0154] Stage 1 (mild): RBL < 15%;

[0155] Stage 2 (moderate): RBL between 15% and 33%;

[0156] Stage 3 (severe): RBL ≥ 33%.

[0157] As Figure 4 shown, the periodontitis stage of each tooth is visually displayed in different colors on a per-tooth basis. Additionally, as Figure 5As shown, since teeth with multiple lesion types usually require different treatment methods from single lesions, in this embodiment, unique color markings are used for teeth affected by multiple oral pathologies in the visualization results to enhance diagnostic support. The hyperparameter settings are shown in Table 1.

[0158] Table 1

[0159]

[0160]

[0161] Performance evaluation metrics:

[0162] In this embodiment, for the object detection task part, mAP50, mAP50:95, number of parameters (Parameters), and computational complexity (GFLOPs) are selected as the detection metrics for localization performance; for the instance segmentation part, mIoU, number of parameters (Parameters), and computational complexity (GFLOPs) are selected as the metrics for evaluating instance segmentation performance.

[0163] Model verification: Finally, the performance of the model was verified on the collected dataset. The experimental results of the object detection task are shown in Table 2. Compared with the baseline model (Baseline), the mAP50 of the MMC-YOLO method in caries detection increased from 80.6% to 81.8%, an increase of 1.2 percentage points, and the mAP50:95 increased from 47.9% to 48.6%; in wisdom tooth detection, the mAP50 increased from 96.2% to 97.0%, and the mAP50:95 increased from 80.7% to 81.2%; for periapical lesion detection, the mAP50 increased from 84.1% to 84.6%, and the mAP50:95 increased from 61.7% to 61.9%. In addition, as shown in Tables 3-4, the experimental results of the instance segmentation task show that compared with A-YOLO-M, the mIoU of MMC-YOLO in alveolar bone segmentation increased from 96.8% to 97.4%, the mIoU of enamel-cementum junction segmentation increased from 95.3% to 96.1%, and the mIoU of tooth segmentation increased from 92.9% to 93.6%. In the periodontitis condition grading task, MMC-YOLO also made significant progress. The accuracy of the non-periodontitis category increased from 77.3% to 86.7%, the mild periodontitis increased from 86.7% to 91.3%, the moderate periodontitis increased from 72.6% to 77.2%, and the severe periodontitis increased from 79.0% to 84.7%. At the same time, the computational complexity of MMC-YOLO decreased from 78 GFLOPs to 76 GFLOPs, and the number of parameters decreased from 14.9M to 13.3M, achieving a reduction in resource consumption while improving detection and segmentation performance, proving that this method has excellent performance in practical applications.

[0164] Table 2

[0165]

[0166]

[0167] Table 3

[0168]

[0169] Table 4

[0170]

[0171] The above is only the preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the technical field within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An oral panoramic imaging processing system, characterized in that, Including: An image acquisition module, which is used to collect initial oral panoramic images containing various oral diseases; An image processing module, connected to the image acquisition module, which is used to perform object detection and instance segmentation annotation on the initial oral panoramic images, and use data augmentation technology to expand the training data set to obtain target oral panoramic images; A model construction module, connected to the image processing module, which is used to train the MMC-YOLO network model through the target oral panoramic images to obtain a target MMC-YOLO network model; An analysis and detection module, connected to the model construction module, which is used to perform synchronous detection of teeth concurrent oral diseases, comprehensive analysis of multi-dimensional pathological information, and precise identification of multi-scale lesions through the target MMC-YOLO network model to obtain analysis results.

2. The system according to claim 1, wherein The object detection and instance segmentation annotation include object detection and segmentation annotation of lesion type, location, and boundary information; The data augmentation includes rotation, scaling, and flipping.

3. The system according to claim 1, wherein The target MMC-YOLO network model includes a shared backbone network, an object detection neck network, an object detection head, an instance segmentation neck network, an instance segmentation head, and a multi-task information comprehensive calculation and analysis module; Among them, the shared backbone network is used to extract general features of oral panoramic images; The object detection neck network is used to process features for object detection tasks; The object detection head is used to identify dental caries, wisdom teeth, and periapical lesions; The instance segmentation neck network is used to process features for instance segmentation tasks; The instance segmentation head is used to segment the alveolar bone, cementoenamel junction, and tooth contour; The multi-task information comprehensive calculation and analysis module is used to integrate the object detection and instance segmentation results to obtain periodontitis grading results and multi-disease detection results.

4. The system according to claim 3, wherein The shared backbone network and the object detection neck network include a multi-scale progressive feature aggregation module; The multi-scale progressive feature aggregation module is used to integrate local and global features through a progressive multi-scale feature extraction strategy and partial convolution operations; and use channel segmentation and progressive channel reduction methods to reduce the computational complexity while retaining key features.

5. The system according to claim 4, wherein The process of the multi-scale progressive feature aggregation module integrating local and global features through a progressive multi-scale feature extraction strategy and partial convolution operations includes: Given the input features, first perform a standard 3×3 convolution operation, then divide the features into two parts in the channel dimension. The first part extracts medium-scale features through a 5×5 convolution, and one-fourth of the channels extract global features through a 7×7 convolution; finally, the multi-scale features are fused through channel splicing and transformation.

6. The system according to claim 3, wherein A coordinate-guided fusion module is introduced in the process of the object detection head identifying dental caries, wisdom teeth, and periapical lesions; The coordinate-guided fusion module is used to combine the edge and texture information of shallow features with the semantic information of deep features, and dynamically adjust the feature representation through the coordinate attention mechanism; enhance the feature modeling ability of the target area in the case of confusion between small-scale lesions and background structures.

7. The system according to claim 6, wherein The process of the coordinate-guided fusion module combining the edge and texture information of shallow features with the semantic information of deep features and dynamically adjusting the feature representation through the coordinate attention mechanism includes: First, perform dimensionality adjustment and preliminary fusion on the shallow features and deep features, define the mapping function of the 1×1 convolution transformation to align the number of feature channels, and then splice the adjusted shallow features and deep features along the channel dimension to generate fused features, and generate an adaptive attention map through local-global interaction in the spatial dimension to optimize the feature representation.

8. The system according to claim 7, wherein The process of generating an adaptive attention map through local-global interaction in the spatial dimension to optimize the feature representation includes: First, perform local information aggregation on the input features, then through the initial feature transformation, use separable directional convolutions to model the spatial dependencies in the horizontal and vertical directions, and finally generate the attention map through feature transformation and normalization.

9. The system according to claim 3, wherein The instance segmentation head segments the alveolar bone, CEJ, and tooth contour through feature splicing, edge enhancement, and spatial attention mechanism, while retaining fine-grained structural information.

Citation Information

Patent Citations

  • Deep learning-based panoramic dental film disease identification and auxiliary diagnosis method and system

    CN117934433A

  • Oral cavity panoramic X-ray image multi-pathological instance segmentation method based on deep learning

    CN118864838A

  • Oral cavity panoramic X-ray image tooth segmentation method based on deep learning

    CN119273919A

  • Rapid pathological image analysis method and apparatus based on magnification-aligned transformer

    WO2025065803A1

Cited By

  • Oral periodontal image analysis method and system based on multi-scale feature fusion

    CN120953746A

  • A Multi-Scale Feature Fusion Method and System for Oral and Periodontal Image Analysis

    CN120953746B

  • Bottle body packaging defect detection system based on machine vision

    CN121073965A

  • Lightweight AI-based distribution line unmanned aerial vehicle edge end real-time visual identification and target detection method and system

    CN121459227A

  • Method for automatically extracting features around tooth boundary and method for identifying oral cavity lesion

    CN122115887A