A Multi-Scale Feature Fusion Method and System for Oral and Periodontal Image Analysis

By employing a multi-scale feature fusion method for oral and periodontal image analysis, combined with multi-branch feature extraction and attention mechanisms, the problem of incomplete capture of lesion features in single-scale image analysis is solved, thus achieving high-precision diagnosis of periodontal lesions.

CN120953746BActive Publication Date: 2026-03-06长沙市口腔医院
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511112248.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2026-03-06
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

In existing technologies, single-scale oral and periodontal image analysis cannot accurately capture subtle lesion features, such as periodontal pocket depth and alveolar bone resorption, leading to missed diagnoses or misdiagnoses.

Method used

A multi-scale feature fusion approach is adopted, which acquires panoramic images, local apical images, and microscopic periodontal pocket images. SIFT feature point matching is used for alignment, and preprocessing techniques such as nonlocal mean denoising, CLAHE, guided filtering, Gamma correction, and morphological cap transformation are combined. Multi-branch feature extraction network is used for multi-scale feature extraction, and feature fusion is performed through cross-scale attention modules and attention mechanisms. Finally, a Softmax classifier is used to locate and classify lesion areas.

Benefits of technology

It enables comprehensive capture of lesion information at different levels, improves the diagnostic accuracy of periodontal lesions, and avoids missed diagnoses and misdiagnoses caused by a single scale.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953746B_ABST
    Figure CN120953746B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for multi-scale feature fusion in oral and periodontal image analysis, relating to the field of image data processing technology. The method includes: acquiring multi-scale oral and periodontal images; aligning the multi-scale oral and periodontal images; preprocessing the multi-scale oral and periodontal images; extracting multi-scale features from the multi-scale oral and periodontal images using a multi-branch feature extraction network; fusing the extracted multi-scale features; locating and classifying periodontal lesion areas based on the fused features; and outputting periodontal image analysis results including lesion location and type. This invention, through multi-branch feature extraction, can process details in images at different scales, fully utilize features at each scale, and perform multi-scale feature fusion, avoiding missed diagnoses caused by the inability to detect small or deep lesions at a single scale.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image data processing technology, and in particular to a method and system for multi-scale feature fusion in oral and periodontal image analysis. Background Technology

[0002] Periodontal diseases (such as gingivitis and periodontitis) may not have obvious symptoms in their early stages, and traditional oral examinations may not be able to detect potential problems in time. However, periodontal image analysis can help doctors detect these lesions early, and even identify the affected areas before symptoms appear.

[0003] With the rapid development of science and technology, image analysis techniques, represented by convolutional neural networks, have been applied in various fields. However, when current image analysis techniques are applied to oral and periodontal image analysis, they often rely solely on single-scale images. Many subtle pathological features (such as periodontal pocket depth and alveolar bone resorption) in these images are typically localized. Single-scale images often fail to accurately capture these details, especially in smaller lesion areas or at a microscopic level. This lack of attention to local features can lead to missed or misdiagnosed lesions. Summary of the Invention

[0004] In view of the shortcomings of the prior art, the purpose of this invention is to provide a multi-scale feature fusion method for oral and periodontal image analysis, which can solve the problem that the prior art relies on a single-scale oral and periodontal image. There are many subtle pathological features (such as periodontal pocket depth, alveolar bone resorption, etc.) in the oral and periodontal image, which are usually located in local areas. Single-scale images often cannot accurately capture these details, especially in smaller lesion areas or at the microscopic level. The lack of attention to local features may lead to missed or misdiagnosed lesions.

[0005] A first aspect of this invention proposes a method for analyzing oral and periodontal images using multi-scale feature fusion, comprising:

[0006] S1: Acquire multi-scale oral and periodontal images;

[0007] S2: Align the multi-scale oral and periodontal images;

[0008] S3: Preprocess the multi-scale oral and periodontal images;

[0009] S4: Multi-scale feature extraction is performed on the multi-scale oral and periodontal images using a multi-branch feature extraction network;

[0010] S5: Fusion of the extracted multi-scale features;

[0011] S6: Based on fusion characteristics, locate and classify periodontal lesion areas;

[0012] S7: Outputs periodontal image analysis results including the location and type of lesions.

[0013] A second aspect of this invention provides a multi-scale feature fusion oral and periodontal image analysis system, comprising: a processor and a memory;

[0014] The memory stores programs or instructions that can run on the processor, which, when executed by the processor, implement the steps of the oral and periodontal image analysis method with multi-scale feature fusion as described in the first aspect.

[0015] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0016] In this embodiment of the invention, by acquiring oral and periodontal images at different scales (such as panoramic images, local periapical images, and microscopic periodontal pocket images), lesion information at different levels can be comprehensively captured. Panoramic images provide global structural information, local images capture details of specific areas, and microscopic images can reveal subtle lesions, ensuring that information from all levels is included in the analysis. By using a multi-branch feature extraction network to extract features at multiple scales, details in the images can be processed at different scales, making full use of features at each scale, and multi-scale feature fusion can be performed to avoid missed diagnoses caused by the inability to detect small-scale or deep lesions at a single scale. Attached Figure Description

[0017] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0018] Figure 1 This is a flowchart illustrating a multi-scale feature fusion method for oral and periodontal image analysis provided in an embodiment of the present invention.

[0019] Figure 2 This is a schematic diagram of the structure of a multi-scale feature fusion method for oral and periodontal image analysis provided in an embodiment of the present invention.

[0020] Figure 3 This is a schematic diagram of the structure of a multi-scale feature fusion oral and periodontal image analysis system provided in an embodiment of the present invention. Detailed Implementation

[0021] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0022] The following description, in conjunction with the accompanying drawings, details the multi-scale feature fusion method for oral and periodontal image analysis provided by the present invention through specific embodiments and application scenarios.

[0023] Reference manual attached Figure 1 The diagram illustrates a flowchart of a multi-scale feature fusion method for oral and periodontal image analysis provided by an embodiment of the present invention.

[0024] Reference manual attached Figure 2 The diagram shows a structural schematic of a multi-scale feature fusion method for oral and periodontal image analysis provided by an embodiment of the present invention.

[0025] This invention provides a method for multi-scale feature fusion in oral and periodontal image analysis, which may include the following steps:

[0026] S1: Acquire multi-scale oral and periodontal images.

[0027] Optionally, multi-scale oral and periodontal images include: global panoramic images, local periapical images, and microscopic periodontal pocket images.

[0028] S2: Align the oral and periodontal images.

[0029] In one possible implementation, S2 specifically involves: using a registration algorithm based on SIFT (Scale Invariant Feature Transform) feature point matching to unify the global panoramic image, the local apical image, and the microscopic periodontal pocket image into the same coordinate system and adjust them to the same spatial resolution.

[0030] It should be noted that SIFT-based registration algorithm is a commonly used image alignment method, aiming to align the same object or scene in different images to a unified coordinate system. This algorithm detects and extracts local feature points in the image, which remain invariant under conditions of scale, rotation, and illumination changes. The SIFT algorithm first constructs an image pyramid in a multi-scale space and extracts stable feature points at each scale. Next, by calculating local region descriptors around the feature points, reliable matching is ensured even with image rotation, scaling, and perspective changes. Finally, by matching feature points and using geometric transformations (such as homography matrices) to register the images, alignment of different images is achieved. This method is widely used in multimodal imaging, medical imaging, and computer vision, exhibiting strong robustness, especially when processing images with deformation, noise, or different resolutions. SIFT-based registration algorithm is a mature existing technology and will not be elaborated upon further in this invention.

[0031] In this embodiment of the invention, unifying the global panoramic image, the local periapical image, and the microscopic periodontal pocket image into the same coordinate system and adjusting them to the same spatial resolution ensures the spatial consistency of images from different scales and modalities. This registration method helps to accurately align key structures and lesion areas in each image, thereby improving the accuracy of subsequent feature extraction and fusion, and avoiding errors or missed diagnoses caused by inconsistent image positions. Furthermore, by adjusting the images to the same spatial resolution, the comparability of information between different images and the fusion effect can be guaranteed, effectively improving the reliability and diagnostic accuracy of multi-scale analysis.

[0032] S3: Preprocess the oral and periodontal images.

[0033] In one possible implementation, S3 specifically includes sub-steps S301 to S303:

[0034] S301: For global panoramic images, a nonlocal mean denoising algorithm is used to suppress device noise, and combined with the CLAHE algorithm, the overall contour of the alveolar bone is enhanced.

[0035] Nonlocal means denoising is a denoising method based on the similarity between image pixels rather than local neighborhoods. Unlike traditional denoising methods (such as mean filtering and Gaussian filtering), which only consider the mean of local neighboring pixels, nonlocal means algorithms calculate a weighted average based on the similarity of all pixels in the image to achieve a more refined denoising effect. In this algorithm, each pixel value in the image is weighted and averaged with other pixels in the entire image based on its similarity (such as texture or grayscale mode), rather than relying solely on neighboring pixels. This effectively suppresses noise while preserving details and edge information in the image, making it particularly suitable for removing device noise and avoiding image blurring caused by excessive smoothing. Nonlocal means denoising is a mature existing technology, and will not be elaborated further in this invention.

[0036] CLAHE (Adaptive Histogram Equalization) is an image enhancement technique primarily used to improve image contrast, especially in low-contrast images, making details clearer. In CLAHE, the image is divided into multiple small regions (called "grids" or "blocks"), and histogram equalization is performed within each region, thereby enhancing the contrast of local areas. To prevent noise amplification due to over-enhancement, CLAHE introduces a contrast limiting mechanism to restrict the histogram gain within each small region. This effectively enhances the alveolar bone contour in the image, making it more prominent, while avoiding artifacts or noise caused by over-enhancement. The CLAHE algorithm is a mature existing technology and will not be elaborated upon further in this invention.

[0037] S302: For local root apex images, a guided filtering algorithm is used to preserve the details of the periodontal ligament while smoothing the background, and a Gamma correction algorithm is used to optimize the contrast of the cementum-periodontal ligament interface.

[0038] Guided filtering is an image-guided filtering method that smooths the background while preserving image details. It uses a guide image to guide the filtering process; the guide image is usually the same as the image to be filtered, but it can also be a different image. Within a local region, guided filtering effectively preserves edge and texture details without affecting the smoothing effect of the entire image by performing a weighted average of the image to be filtered. The weights are determined by the local structure of the guide image. This allows guided filtering to remove noise or background interference while preserving the clarity of details such as periodontal ligaments, making it ideal for detail enhancement in medical images. Guided filtering is a mature existing technology and will not be elaborated upon further in this invention.

[0039] Gamma correction is an image brightness adjustment technique used to improve image contrast and brightness distribution. By applying a nonlinear transformation to the pixel values ​​of an image, gamma correction can adjust the overall perceived brightness of the image. Specifically, it changes the brightness ratio of the image by exponentially mapping each pixel value according to a specific gamma value (usually a constant greater than 1). For the cementum-periodontal ligament interface, gamma correction can enhance the contrast of these areas, making the interface clearer and helping dentists better identify and analyze subtle structural differences, especially in complex local periapical images. Gamma correction is a mature existing technology and will not be elaborated upon in this invention.

[0040] S303: For microscopic periodontal pocket images, morphological top-cap transformation is used to remove surface reflections of gingival tissue, and Gaussian blur filtering algorithm is combined to weaken texture interference in non-lesion areas.

[0041] Morphological top-hat transformation is an image processing technique primarily used to extract detailed structures from an image whose brightness is less than that of the surrounding background. This transformation is based on morphological operations. First, it removes highlighted areas from the image through erosion, and then restores the image structure through dilation, thereby highlighting small, bright areas. Specifically, in microscopic periodontal pocket images, top-hat transformation can effectively remove reflective or overly bright areas on the gingival tissue surface, areas that may be artifacts caused by changes in lighting. This operation better preserves the details of the actual lesion area, improving image quality and analyzability. Morphological top-hat transformation is a mature existing technology and will not be elaborated upon further in this invention.

[0042] Gaussian blur filtering is a commonly used image smoothing technique. It blurs the image by applying a weighted average of a Gaussian function, thereby reducing noise and details. Gaussian filtering effectively preserves smooth areas of the image while reducing prominent noise or irrelevant details. In microscopic periodontal pocket images, combining Gaussian blur filtering can weaken texture interference in non-lesion areas, making the features of lesion areas more prominent. This allows doctors to more clearly identify lesion areas without being affected by subtle interference from the background or non-lesion areas. Gaussian blur filtering is a mature existing technology and will not be elaborated upon in this invention.

[0043] S4: Multi-scale feature extraction is performed on oral periodontal images through a multi-branch feature extraction network.

[0044] Optionally, the multi-branch feature extraction network includes global branches, local branches, and micro branches.

[0045] Optionally, global branches, local branches, and micro branches are all set with four stages.

[0046] It's important to note that the image features processed at each stage are progressively abstracted from low to high levels to capture semantic information at different scales and levels within the image. In the early stages of the network, basic, low-level image features are extracted, such as edges, textures, and shapes. As the network progresses to deeper stages, it can abstract higher-level features with richer semantic information from these low-level features. Examples include global structure and deep lesion patterns. This step-by-step extraction approach helps the network gain a more comprehensive understanding of the image, gradually enhancing its understanding of the target from details to the global level.

[0047] Optionally, each stage of the global branch uses the first 3 convolutional blocks of ResNet-18.

[0048] It's worth noting that using the first three convolutional blocks of ResNet-18 as the feature extraction part of the global branch effectively extracts global features from the image. ResNet-18 employs residual connections, which effectively alleviates the vanishing gradient problem during deep network training, thereby improving the network's training stability and convergence speed. By using the first three convolutional blocks of ResNet, the network can capture macroscopic structural information in the image, such as global shape and location distribution, which helps extract holistic features, making it particularly suitable for tasks requiring global semantic understanding.

[0049] Optionally, each stage of a local branch uses the first two dense blocks of DenseNet-121.

[0050] It's worth noting that DenseNet-121 uses dense connections to ensure that the output of each layer can be combined with features from all preceding layers, enhancing feature propagation efficiency and information utilization. Using the first two dense blocks of DenseNet-121 for local branches helps extract local detail features of the image (such as local lesions and edge information). This dense connection design effectively enhances the expressive power of local information while avoiding information loss or gradient vanishing, thereby improving the network's performance in extracting features from local regions.

[0051] Optionally, each stage of the micro-branch is downsampled three times using the downsampling path of U-Net.

[0052] It's worth noting that U-Net's downsampling path focuses on capturing detailed image information, particularly excelling in image segmentation and microscopic feature extraction. By applying U-Net's downsampling path to the microscopic branch and performing three downsampling iterations, it effectively extracts microscopic details from images, such as minute lesions or texture information. The downsampling path progressively compresses the image space while enhancing the semantic information of the image, enabling the network to recognize finer-grained features, making it particularly suitable for capturing microscopic lesion areas in periodontal images.

[0053] While ResNet-18, DenseNet-121, and U-Net are mature existing technologies, innovatively building a multi-branch feature extraction network using some of their network structures is a significant innovation. Through this innovative combination of multi-branch feature extraction networks, these three mature network structures can complement each other's strengths and weaknesses, resulting in more refined extraction and fusion of global, local, and microscopic features. This innovative design breaks through the limitations of a single network architecture, fully leveraging the advantages of each network at different feature levels, ultimately improving the performance and accuracy of the entire model when dealing with complex medical images. This innovation not only improves model performance but also provides new ideas for the design of multi-branch network architectures, further promoting the development of the field of medical image analysis.

[0054] In one possible implementation, S4 specifically includes sub-steps S401 to S403:

[0055] S401: Extract global features from the global panoramic image through a global branch. Extracting global features from the panoramic image through the global branch allows for the acquisition of overall structural information, such as alveolar bone morphology and large-scale lesions.

[0056] S402: Local branching extracts local features from local root apex images. Local branching extracts local features from local root apex images, focusing on lesions or details within a small area, thus enhancing the ability to identify local lesions.

[0057] S403: Micro-branching extracts microscopic features from micro-periodontal pocket images. Micro-branching helps capture more detailed lesion features or tissue textures from micro-periodontal pocket images.

[0058] In this embodiment of the invention, by dividing the image into three levels—global, local, and micro—and extracting features at different levels through global, local, and micro branches respectively, comprehensive capture of multi-scale information can be achieved.

[0059] S5: Fusion of the extracted multi-scale features.

[0060] In one possible implementation, S5 specifically includes:

[0061] S50A1: Input the extracted global features, local features, and micro features into the cross-scale attention module.

[0062] It's worth noting that inputting the extracted global, local, and micro-level features into the cross-scale attention module helps enhance the model's ability to integrate information at various scales. The cross-scale attention module effectively captures the interdependencies between different scales, thus highlighting key information at each scale during feature fusion and suppressing redundant or unimportant features. This helps improve the accuracy of feature fusion, resulting in a more comprehensive and detailed final feature representation, thereby improving classification and localization accuracy.

[0063] S50A2: The global features are upsampled to 256×256 through bilinear interpolation, concatenated with local features, and then convolved with 1×1 to generate a global guide map.

[0064] It's important to note that generating a global guide map ensures spatial consistency between global information and local details. This process ensures that global information can work collaboratively with local features at higher resolutions, enhancing the connection between global structure and local details in the image, thus making subsequent feature fusion more accurate. 1×1 convolution further reduces the number of channels, avoiding computational redundancy and improving fusion efficiency.

[0065] S50A3: Performs a dot product operation between micro-features and the global guidance map to generate an attention mask for the micro-features.

[0066] It should be noted that the attention mask used to generate micro-features can effectively guide information from global features to the micro-level. This operation helps the network increase its focus on micro-features, making key information within these features (such as minute lesions and tissue details) stand out more, and avoiding the impact of local noise or irrelevant areas on the accurate identification of lesion areas. This step strengthens the global guidance of micro-information and enhances the expressive power of micro-features.

[0067] S50A4: Weighted fusion of attention masks for local and micro features to generate fused features.

[0068] It should be noted that by weighting and fusing local and micro-features using attention masks to generate fused features, precise feature weighting can be achieved between local details and micro-features. This weighted fusion process, by considering the importance of micro-features, allows the network to reasonably allocate the contributions of local and micro-information according to specific task requirements, thereby generating more discriminative feature representations. This step optimizes the feature fusion effect, ensuring that the fused features can better support subsequent classification and lesion detection.

[0069] In one possible implementation, S5 specifically includes:

[0070] S50B1: Inputs global features into a channel attention mechanism to enhance semantic-specific feature representations through the interdependencies between channel graphs.

[0071]

[0072] in, G represents the feature generated after the global feature in the i-th stage is processed by the channel attention mechanism, where CA represents the channel attention mechanism, and G represents the feature generated after the global feature in the i-th stage is processed by the channel attention mechanism. i This represents the global feature of the i-th stage. This indicates element-wise multiplication.

[0073] It's important to note that inputting global features into the channel attention (CA) mechanism helps the network adaptively adjust based on the feature importance of each channel. By modeling the interdependencies between different channels, the CA mechanism can highlight important semantic information while suppressing less important channels. This approach helps improve the expressive power of global features, ensuring the model focuses more on features relevant to the classification task, thus improving classification accuracy. This is especially beneficial when dealing with complex medical images, as it can better capture the overall features of lesion regions.

[0074] S50B2: Inputs local features into a spatial attention mechanism to enhance local details and suppress irrelevant regions.

[0075]

[0076] in, SA represents the feature generated after the local features of the i-th stage are processed by the spatial attention mechanism, and L represents the spatial attention mechanism. i This represents the local features of the i-th stage.

[0077] It's worth noting that inputting local features into a spatial attention (SA) mechanism can help the model identify important regions in an image, enhance local details, and suppress irrelevant areas. The spatial attention mechanism focuses attention on key information in the image, such as lesions or texture features, by assigning different weights to each spatial location, while minimizing the influence of background and irrelevant regions. This helps improve the model's sensitivity to details, especially in the recognition of local features, and avoids interference from background noise.

[0078] S50B3: Inputs microscopic features into a dual attention mechanism to jointly enhance tissue texture and spatial detail.

[0079]

[0080] in, M represents the feature generated after the micro-features of the i-th stage are processed by the dual attention mechanism. i This represents the microscopic features of the i-th stage.

[0081] It's important to note that by inputting microscopic features into a dual-attention mechanism—combining channel attention and spatial attention—tissue texture and spatial detail can be jointly enhanced. This dual-attention mechanism not only improves the focus on microscopic features but also ensures that these features can better capture minute lesions, textures, and details, thus enhancing the expression of features at the microscopic level. This process helps to capture subtle lesions or pathological changes that are difficult to detect, improving the representational power of microscopic features, especially excelling in fine-grained analysis of medical images.

[0082] S50B4: Integrate the results generated by each fusion path:

[0083]

[0084]

[0085] in, This represents the fusion result of fusing global-local-micro features with the fusion features of the i-th stage. This indicates a 3×3 convolution operation; Concat indicates concatenation. This represents the features generated by downsampling from the (i-1)th stage feature fusion module, and Avgpool represents the average pooling operation. This represents a 1×1 convolution operation. This represents the fusion feature map of the (i-1)th stage.

[0086] It's important to note that fusing global, local, and micro-level feature results (through concatenation and convolution operations) effectively integrates feature information from different branches, forming a more discriminative comprehensive feature set. Operations such as 3×3 convolution and average pooling can effectively extract useful information from features at various scales while reducing redundancy and improving computational efficiency. This step ensures that features at different scales are well-fused, strengthening the model's ability to integrate information at different levels and improving the final feature performance.

[0087] S50B5: Generates fused features using a residual inverted multilayer perceptron.

[0088]

[0089]

[0090] Among them, Fi Let represent the fusion feature map of the i-th stage, IRMLP represent the residual inverted multilayer perceptron, LN represent the layer normalization process, and x represent the input of the residual inverted multilayer perceptron.

[0091] It's worth noting that further processing the fused features using the Inverted Residual Multilayer Perceptron (IRMLP) enhances the network's expressive power and avoids gradient vanishing and information loss. The IRMLP structure, by introducing multiple 1×1 and 3×3 convolutions along with layer normalization (LN) operations, ensures a more stable deep learning process for the features, reduces the risk of overfitting, and improves the model's robustness. The introduction of the residual structure ensures smoother information flow within the network, preventing information degradation. Ultimately, IRMLP further enhances the expressive power of the fused features, enabling the final features to be used more effectively for classification tasks.

[0092] In their research on feature fusion, the applicant proposed an independent multi-scale feature fusion approach, with each scale offering its own advantages. For the S50A1 to S50A4 feature fusion methods, through progressive attention mechanisms and weighted fusion, information from different scales is gradually integrated, making the features at each stage more refined and richer. The input and output of each stage can be refined layer by layer, improving the accuracy of the fusion. However, compared to the S50B1 to S50B5 feature fusion method, while the use of cross-scale attention mechanisms can accurately model the dependencies between features, the computational cost is high, potentially affecting the model's training speed and inference efficiency. The S50B1 to S50B5 feature fusion methods employ channel attention (CA), spatial attention (SA), and dual attention mechanisms. These mechanisms can precisely control the weighted fusion of features, and when combined with the Inverted Residual Multilayer Perceptron (IRMLP), further optimization of feature flow is achieved, ensuring smooth information transfer and preventing gradient vanishing or information degradation. However, compared to the S50A1 to S50A4 feature fusion methods, the weighting process of channel attention and spatial attention may introduce interdependence between features, especially when the features are semantically similar. In this case, the weighting of the attention mechanism may not be able to effectively distinguish subtle differences between different features, resulting in unsatisfactory fusion results. In practical applications, the detection accuracy of the second feature fusion method is slightly higher than that of the first.

[0093] S6: Based on fusion characteristics, locate and classify periodontal lesion areas.

[0094] In one possible implementation, the fused features can be directly input into the Softmax classifier to locate and classify periodontal lesion areas.

[0095] The Softmax classifier is a commonly used activation function for multi-class classification problems. It calculates an exponential function for each class and normalizes it, ensuring that the sum of all output values ​​is 1, thus representing a probability distribution. Each class's output value represents its probability. The Softmax classifier maps the network's output to a class probability, ultimately selecting the class with the highest probability as the prediction.

[0096] In one possible implementation, S6 specifically includes:

[0097] S601: Calculate the variance of the weights of the first 3 convolutional kernels as an indicator of feature fusion stability.

[0098] The first three convolutional kernels refer to the kernels used in the initial three convolutional layers of the network. By calculating the variance of the kernel weights of these convolutional layers, the stability of feature fusion can be evaluated. This evaluation helps determine which mode (lightweight, standard, or expert integration mode) to choose for the localization and classification of periodontal lesions.

[0099] Furthermore, the variance of the convolutional kernel can be used as an indicator to reflect the stability of feature learning. If the weights of the convolutional kernels in the first three layers vary greatly (large variance), it may indicate significant instability or overfitting during feature extraction, leading to fluctuations in performance on subsequent classification tasks.

[0100] S602: When the feature fusion stability index is in When selecting the lightweight mode, the fused features are input into the Softmax classifier for the localization and classification of periodontal lesion areas, and st represents the feature fusion stability index.

[0101] It should be noted that when feature fusion is stable and there is little redundant information, the lightweight Softmax classifier can be used to achieve fast and low-resource-consumption inference.

[0102] S603: When the feature fusion stability index is in When selecting the standard mode, the fused features are input into ResNet34 for the localization and classification of periodontal lesion areas.

[0103] ResNet34 is a variant of ResNet, which introduces residual connections to allow for deeper networks without the vanishing gradient problem. ResNet34 employs 34 convolutional layers, each adding to the output of the previous layer via residual connections, ensuring smooth information propagation through the network and solving the training problems of traditional deep networks. ResNet34 exhibits strong representational capabilities in image classification tasks, effectively capturing complex features, and providing high-accuracy classification and localization, particularly when dealing with highly complex and variable medical images.

[0104] It should be noted that when the feature fusion is moderately stable, choosing ResNet34 as the standard mode provides a good balance between feature extraction capability and computational efficiency.

[0105] S604: When the feature fusion stability index is in When selecting the expert integration mode, the fused features are input into a multi-expert system formed by integrating three sub-networks: Softmax classifier, MobileNetV3, and ResNet34. The system outputs the results through gating weights to locate and classify periodontal lesion areas.

[0106] MobileNetV3 is a lightweight convolutional neural network. It reduces computation and the number of parameters by employing depthwise separable convolution, thereby improving inference speed and computational efficiency. MobileNetV3 further optimizes model performance through a carefully designed modular structure (such as the hardware-friendly ReLU activation function and the Squeeze-and-Excitation module). In periodontal lesion detection tasks, MobileNetV3 can serve as a highly efficient model, rapidly locating and classifying lesion areas in environments with low computational resources.

[0107] Among them, the gated weighting mechanism is a technology that uses dynamic adjustment of the weights of different sub-networks or different feature channels to weight and fuse multiple information sources.

[0108] Specifically, three classifiers, Softmax, MobileNetV3, and ResNet34, can be used to output the probability of the current image being in various lesion types. Then, a weighted summation method is used to calculate the final probability, and the category with the highest final probability is determined as the final detection result.

[0109] It should be noted that when the feature fusion stability index is poor, a single model may not be able to fully cover the feature differences at different scales and modalities. In this case, introducing an expert ensemble system, through multi-network complementarity, enables adaptive expert collaborative judgment, which greatly improves the model's performance on complex and heterogeneous images, making it suitable for the analysis of complex and difficult cases.

[0110] S7: Outputs periodontal image analysis results including the location and type of lesions.

[0111] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0112] In this embodiment of the invention, by acquiring oral and periodontal images at different scales (such as panoramic images, local periapical images, and microscopic periodontal pocket images), lesion information at different levels can be comprehensively captured. Panoramic images provide global structural information, local images capture details of specific areas, and microscopic images can reveal subtle lesions, ensuring that information from all levels is included in the analysis. By using a multi-branch feature extraction network to extract features at multiple scales, details in the images can be processed at different scales, making full use of features at each scale, and multi-scale feature fusion can be performed to avoid missed diagnoses caused by the inability to detect small-scale or deep lesions at a single scale.

[0113] Reference manual attached Figure 3 The diagram shows a schematic of the structure of a multi-scale feature fusion oral and periodontal image analysis system provided by an embodiment of the present invention.

[0114] This invention provides a multi-scale feature fusion oral and periodontal image analysis system 20, including a processor 201 and a memory 202.

[0115] The memory 202 stores programs or instructions that can run on the processor 201. When the program or instructions are executed by the processor 201, they implement the steps of the above-described multi-scale feature fusion oral and periodontal image analysis method and achieve the same technical effect. To avoid repetition, the present invention will not elaborate further.

[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the protection scope of the present invention.

Claims

1. A multi-scale feature fusion periodontal image analysis method, characterized in that, The method comprises the following steps: S1: acquiring multi-scale periodontal images of the oral cavity; S2: aligning the multi-scale periodontal images of the oral cavity; S3: preprocessing the multi-scale periodontal images of the oral cavity; S4: extracting multi-scale features of the multi-scale periodontal images of the oral cavity through a multi-branch feature extraction network; S5: fusing the extracted multi-scale features; S6: positioning and classifying periodontal lesion regions according to the fused features; S7: outputting periodontal image analysis results containing lesion positions and types; The multi-branch feature extraction network comprises a global branch, a local branch and a microscopic branch; the global branch, the local branch and the microscopic branch are all provided with four stages; each stage of the global branch adopts the first three convolution blocks of ResNet-18; each stage of the local branch adopts the first two dense blocks of DenseNet-121; Each stage of the microscopic branch adopts the downsampling path of U-Net for three times of downsampling; The S6 specifically comprises: S601: calculating the variance of the weights of the first three convolution kernels as a feature fusion stability index; S602: When the feature fusion stability index is in lightweight mode, input the fusion features to the Softmax classifier for periodontal lesion region positioning and classification, and st represents the feature fusion stability index. S603: When the feature fusion stability index is in standard mode, input the fusion features to ResNet34 for periodontal lesion region positioning and classification; S604: When the feature fusion stability index is in the expert integration mode is selected, the fusion features are input into a multi-expert system formed by the integration of three sub-networks of Softmax classifier, MobileNetV3, and ResNet34, the output is weighted through a gating weight, and periodontal lesion region positioning and classification are performed.

2. The multi-scale feature fused periodontal image analysis method according to claim 1, wherein, The multi-scale periodontal images of the oral cavity comprise a global panoramic image, a local apical image and a microscopic periodontal pocket image.

3. The multi-scale feature fused periodontal image analysis method according to claim 2, characterized in that, The S2 specifically comprises: The global panoramic image, the local apical image and the microscopic periodontal pocket image are unified to the same coordinate system and adjusted to the same spatial resolution through a registration algorithm based on SIFT feature point matching.

4. The multi-scale feature fused periodontal image analysis method according to claim 2, wherein, S3 specifically comprises: S301: for the global panoramic image, using a non-local mean denoising algorithm to suppress device noise, combining a CLAHE algorithm to enhance the overall profile of the alveolar bone; S302: for the local apical image, using a guided filter algorithm to retain periodontal ligament details while smoothing the background, using a Gamma correction algorithm to optimize the contrast of the cementum-periodontal membrane interface; S303: for the microscopic periodontal pocket image, using morphological top-hat transformation to remove the reflection of the gum tissue surface, combining a Gaussian blur filter algorithm to weaken the texture interference of non-lesion regions.

5. The multi-scale feature fused periodontal image analysis method according to claim 2, wherein, The S4 specifically comprises: S401: extracting global features in the global panoramic image through the global branch; S402: extracting local features in the local apical image through the local branch; S403: extracting microscopic features in the microscopic periodontal pocket image through the microscopic branch.

6. The multi-scale feature fused periodontal image analysis method according to claim 5, wherein, The S5 specifically comprises: S50A1: inputting the extracted global features, local features and microscopic features into a cross-scale attention module; S50A2: performing bilinear interpolation upsampling on the global features to 256x256, splicing the global features with the local features, and generating a global guide map through 1x1 convolution; S50A3: performing dot product operation on the microscopic features and the global guide map to generate an attention mask of the microscopic features; S50A4: weighting and fusing the local features and the attention mask of the microscopic features to generate the fused features.

7. The multi-scale feature fused periodontal image analysis method according to claim 5, wherein, The S5 specifically comprises: S50B1: input the global feature into a channel attention mechanism to enhance the feature representation of specific semantics through mutual dependency between channel graphs; S50B2: input the local feature into a spatial attention mechanism to enhance local details and suppress irrelevant areas; S50B3: input the microscopic feature into a dual attention mechanism to jointly enhance tissue texture and spatial details; S50B4: integrate the results generated by each fusion path; S50B5: generate the fusion feature through a residual inverted multi-layer perception.

8. A multi-scale feature fusion oral periodontal image analysis system, characterized by, Comprise: a processor; a memory, the memory having computer readable instructions stored thereon, the computer readable instructions, when executed by the processor, implementing the multi-scale feature fusion periodontal image analysis method of any one of claims 1-7.

Citation Information

Patent Citations

  • Tooth lesion detection method based on global feature optimization Transform

    CN119067950A

  • Oral panoramic image processing system

    CN120278989A