Dental image segmentation method and system based on frequency domain enhancement and dynamic scanning mechanism

The dental image segmentation method using frequency domain enhancement and dynamic scanning mechanism solves the problems of blurring and missegmentation at the junction of teeth and gums, improves the real-time performance and accuracy of dental image segmentation, and is suitable for clinical diagnosis and treatment planning.

CN121708030APending Publication Date: 2026-03-20ZHEJIANG GONGSHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511834717.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing dental image segmentation methods suffer from blurring and missegmentation at the tooth-gingival junction, have high computational complexity, lack robustness in noisy environments, and fail to fully utilize frequency domain features, leading to unstable boundary details and making them difficult to apply in real-time clinical settings.

Method used

A method based on frequency domain enhancement and dynamic scanning mechanism is adopted. The sampling position and scanning order are adaptively adjusted through dynamic scanning mechanism. Combined with frequency domain wavelet decomposition and reconstruction, the frequency domain information of the feature map is enhanced, and multi-scale fusion is performed in the decoder to generate a segmentation mask.

Benefits of technology

While maintaining linear computational overhead, it improves the spatial continuity representation of tooth edges and gingival interfaces, enhances segmentation accuracy under complex lighting and saliva reflection conditions, solves the problem of global morphological distortion caused by structural continuity disruption and edge enhancement, and achieves real-time high-precision dental image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121708030A_ABST
    Figure CN121708030A_ABST
Patent Text Reader

Abstract

The invention provides a dental image segmentation method and system based on a frequency domain enhancement and dynamic scanning mechanism, the method is constructed on a visual state space model framework, and the core comprises a dynamic scanning part and a frequency domain enhancement part, the dynamic scanning block adaptively adjusts a sampling position and a scanning sequence through a trainable offset prediction network, dynamic scanning based on image content is realized, spatial continuity is kept, and the structural characterization capability is improved; frequency domain enhancement is combined with wavelet decomposition and spectrum pooling technologies, high and low frequency characteristics are balanced, and intermediate frequency components are enhanced, so that accurate boundary positioning and structure identification are still kept under unfavorable imaging conditions such as noise, light reflection and uneven illumination. Compared with the existing dental image segmentation method, the method provided by the invention can generate high-quality tooth, gingival and oral cavity tissue segmentation masks. The method can be widely applied to the fields of digital dental diagnosis, treatment planning and intelligent medical image analysis, and has relatively high practical value and popularization prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision, and in particular to a dental image segmentation method and system based on frequency domain enhancement and dynamic scanning mechanism. Background Technology

[0002] In dental medical image analysis, segmentation technology can not only perform precise separation of oral tissues such as teeth, gums, tongue, lips and cheeks, but also plays an irreplaceable role in lesion identification, clinical diagnosis, treatment planning and surgical navigation.

[0003] Existing dental image segmentation methods still have significant limitations. While early convolutional neural networks (CNNs) achieved some success in medical image segmentation, their limited local receptive field of the convolutional kernel resulted in insufficient modeling of long-range dependencies and complex boundaries, easily leading to blurring and misclassification at the tooth-gingival junction. In recent years, the Transformer architecture has been introduced, achieving global dependency modeling through self-attention mechanisms. However, its computational complexity increases quadratically with the input length, creating a significant computational and storage burden on high-resolution dental images, limiting its real-time clinical applications. Furthermore, oral images often contain light reflections, interference from metal restorations, and motion blur; existing methods lack robustness in noisy environments and struggle to accurately segment boundaries. In addition, existing methods have limited utilization of frequency domain features, often neglecting the modeling of mid-frequency features, leading to unstable boundary details. While recently proposed state-space models such as Mamba replace self-attention with linear complexity, improving efficiency, they still suffer from impaired spatial continuity and insufficient utilization of frequency domain information in dental image segmentation. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a dental image segmentation method and system based on frequency domain enhancement and dynamic scanning mechanisms.

[0005] The objective of this invention is achieved through the following technical solution: a dental image segmentation method and system based on frequency domain enhancement and dynamic scanning mechanism, the method comprising the following steps:

[0006] S1. Acquire the input dental image and perform preprocessing operations to ensure the consistency and robustness of the input data;

[0007] S2. A dynamic scanning mechanism is introduced for feature extraction. The sampling position and scanning order are adaptively adjusted by the offset prediction network to obtain a feature sequence that maintains spatial continuity. The feature map is obtained after sampling.

[0008] S3. After the feature extraction stage, a frequency domain enhancement mechanism is introduced to perform wavelet decomposition, band reconstruction and frequency domain separation and fusion on the feature map of the input image. The enhanced frequency domain information is mapped back to the spatial domain and used in conjunction with the wavelet reconstruction results to generate the final enhanced features.

[0009] S4. Input the final enhanced features into the decoder to restore the spatial resolution and generate a segmentation mask;

[0010] S5. Perform channel compression on the decoding results and output the segmentation results of the oral cavity region.

[0011] Furthermore, the dynamic scanning mechanism includes:

[0012] The two-dimensional feature map is expanded into a one-dimensional sequence in row-major order. The one-dimensional sequence is then input into the offset prediction network to predict the offset vector of each sampling point. The updated sampling position is obtained by adding the offset position to the original coordinates, thereby adjusting the scanning order.

[0013] Furthermore, the offset prediction network consists of multiple sequentially connected sub-modules. First, it extracts local spatial information of the input feature sequence through depthwise separable convolutional layers, reducing computational complexity while preserving the independence between channels. Then, layer normalization is introduced to ensure the numerical stability of features during transmission and avoid gradient anomalies. On this basis, the GELU activation function is used to perform nonlinear transformation on the features to enhance the network's ability to fit complex patterns. Finally, a linear fully connected layer maps the high-dimensional features to offset prediction results with the same spatial resolution as the input feature map, thereby forming a complete offset vector field output.

[0014] Furthermore, the feature map obtained after sampling is specifically obtained by extracting information from the feature map using bilinear interpolation, calculated as follows:

[0015] in, For interpolation functions, The updated sampling point coordinates, This is the original feature map.

[0016] Furthermore, the wavelet decomposition process includes: performing wavelet transform on the feature map after dynamic scanning to decompose it into low-frequency global structure information, horizontal high-frequency detail information, vertical high-frequency detail information, and diagonal high-frequency detail information;

[0017] Wavelet transform is used to decompose the features of the input image to obtain low-frequency and high-frequency sub-bands, as shown in the following formula:

[0018] in, Indicates input features, These represent the low-frequency and high-frequency sub-bands, respectively.

[0019] Furthermore, the band reconstruction specifically includes:

[0020] The frequency domain subbands are reconstructed through convolution and inverse wavelet transform to obtain the enhanced feature map:

[0021] in, This represents a one-dimensional convolution operation, and IDWT represents inverse wavelet transform. This step enhances the high-frequency information at the boundary while preserving the low-frequency structure, thereby improving the segmentation accuracy of the model under complex imaging conditions.

[0022] Furthermore, the frequency domain separation and fusion specifically includes: separating high and low frequency information in the reconstructed feature map Y in the frequency domain and fusing them, as shown in the following formula:

[0023]

[0024] in, Indicates Fourier transform, Indicates inverse transformation, It is the Fourier transform centering function, where, This is a balancing parameter used to control the high- and low-frequency fusion ratio.

[0025] Furthermore, both the decoder and the encoder used for feature extraction include several layers. During the decoding process, a skip connection structure is adopted to fuse the frequency domain enhanced features with the upsampled decoded features, thereby improving the segmentation accuracy of the tooth boundary region and finally obtaining the predicted segmentation mask.

[0026] On the other hand, this specification also provides a dental image segmentation system based on frequency domain enhancement and dynamic scanning mechanisms. This system includes a dynamic scanning module, a frequency domain enhancement module, a decoding module, and an output interface. The dynamic scanning module receives dental images and performs adaptive sampling through an offset prediction network. The frequency domain enhancement module decomposes and fuses multi-scale features in the frequency domain. The decoding module is responsible for restoring the enhanced features into a segmentation mask. The output interface provides the final segmentation results to a dental auxiliary diagnosis and treatment system, enabling automatic identification and localization of teeth, gums, and other oral tissues.

[0027] On the other hand, this specification also provides a dental image segmentation device based on frequency domain enhancement and dynamic scanning mechanisms. This device includes a memory and one or more processors. The memory stores executable code, and the processors, when executing the code, can implement the segmentation process as described above. This device can be deployed in dental clinical image acquisition equipment or embedded as an independent intelligent segmentation module into a digital diagnostic platform, offering advantages in real-time performance and high precision.

[0028] The beneficial effects of this invention are as follows: 1. By introducing a dynamic scanning mechanism, structural adaptive constraints are applied to the scanning direction and step size, making the dynamic scanning compatible with the tooth arch distribution and soft tissue continuity. This significantly improves the spatial continuity representation of the tooth edge and gingival interface while maintaining linear computational overhead, solving the problem of "structural continuity disruption" in existing technologies; 2. Based on the unique characteristics of uneven brightness, wet reflection, and mucosal texture in dental images, a frequency domain enhancement module is proposed to emphasize gingival texture, alveolar boundary, and tooth information in the scale domain. This allows the mid-frequency features and the structural continuity preserved by the dynamic scanning to complement each other, manifested as: the scanning mechanism ensuring macroscopic contour consistency, and frequency domain enhancement strengthening local texture recognition. The combined effect is not replaceable in existing technologies, and it can significantly improve mIoU under complex lighting and saliva reflection conditions; 3. The cross-scale fusion in the decoding stage of this invention is not the conventional U-Net layer-by-layer stitching, but is based on the alignment of dual-domain features of "dynamic scanning (spatial domain) + mid-frequency enhancement (frequency domain)" before multi-scale reconstruction, so that global structural information and high-frequency edge details can be synchronously transmitted back. This design solves the problem in existing methods that it is difficult to achieve both "edge enhancement leading to global morphological distortion" and "global structure preservation leading to local detail loss", and achieves a stable improvement in mIoU; 4. The model structure proposed in this invention has both real-time inference speed suitable for clinical use and robustness in complex oral environments (wet surface, reflection, occlusion), and can be directly deployed in various digital dental terminals such as oral scanning systems, oral X-ray, and intraoral camera systems, providing reliable support for clinical diagnosis, gingival margin assessment, restoration boundary recognition, and orthodontic scheme design. Attached Figure Description

[0029] Figure 1 A schematic diagram of a dental image segmentation method and system provided in an embodiment of the present invention;

[0030] Figure 2 This is a schematic diagram illustrating the implementation of the dynamic scanning module provided in an embodiment of the present invention;

[0031] Figure 3 Comparison of segmentation results of the present invention and the U-Mamba and HQ-SAM methods in four different dental scenarios, provided for embodiments of the present invention;

[0032] Figure 4 This is a schematic diagram of a dental image segmentation device based on frequency domain enhancement and dynamic scanning mechanism, provided as an embodiment of the present invention. Detailed Implementation

[0033] Numerous specific details are set forth in the following description to provide a thorough understanding of the implementation of the invention. However, it should be noted that the invention can also be implemented in other different ways, and those skilled in the art can make extensions and improvements based on it without departing from the spirit and essence of the invention. Therefore, the following embodiments should not be considered as limiting the scope of protection of the invention. The present invention provides a dental image segmentation method and system based on a hybrid frame, capable of image segmentation of multiple oral regions such as teeth, gums, tongue, lips, and cheeks. The method unfolds sequentially through steps such as input preprocessing, dynamic scanning mechanism, frequency domain enhancement mechanism, decoding and reconstruction, and final result output, as follows: Figure 1 As shown, the method includes the following steps:

[0034] S1. Acquire the input dental image and perform preprocessing operations such as normalization and size adjustment to ensure the uniformity and robustness of the input data;

[0035] The input dental image of this invention can be obtained from intraoral endoscopy, 3D dental scanning, or other image acquisition devices. The image is in color RGB format. For the input image, its size is first adjusted to a uniform standard resolution; in this embodiment, a standard resolution of 512×512 pixels is selected. Assume the original input image is... ,in , These represent the original height and width, respectively, and C represents the number of channels, typically set to 3. The image is scaled using bilinear interpolation. The image pixel values ​​are then normalized, mapping the pixel intensity range [0, 255] to the interval [0, 1], using the following formula:

[0036]

[0037] Further zero-mean standardization was performed to obtain:

[0038]

[0039] Where μ is the average pixel value of the entire image. The standard deviation is the pixel value. If the input is a single-channel grayscale image, it is copied three times to form three-channel data, ensuring consistency in subsequent model processing. The final result is the standardized input image. The input image is then fed into a primary feature extraction network for convolution processing. A 3×3 convolution kernel and a downsampling method with a stride of 2 are used, combined with batch normalization and the GELU activation function. After multiple convolution iterations, a primary feature map is obtained. Where h and w are the spatial dimensions after downsampling, and in this embodiment, we take... d is set to 256. This feature map serves as the input for the subsequent dynamic scanning mechanism.

[0040] S2. In the sequence modeling process, a dynamic scanning mechanism is introduced, such as... Figure 2 As shown, the offset prediction network adaptively adjusts the sampling position and scanning order to obtain a feature sequence that maintains spatial continuity.

[0041] First, the two-dimensional feature map Expand into a one-dimensional sequence according to row order. ,in Each element Subsequently, an offset prediction network (OPN) is constructed to predict the offset at each sampling point. The OPN consists of 3×3 depthwise separable convolutions, layer normalization, GELU activation, and 1×1 convolutional layers, with an output dimension of... The two-dimensional coordinate offset corresponding to each sampling point Let the original sampling coordinates be... The updated coordinates are:

[0042]

[0043] in This is the predicted offset. The new set of sampling points. A new feature representation is obtained by sampling from the original feature map using bilinear interpolation:

[0044]

[0045] in The bilinear interpolation operator is defined as follows:

[0046] here, and The component representing the predicted location, and and This corresponds to the index of the feature map location. Take the first [item] in the feature map line, number The feature vector of the entire channel at the column position. Function Bilinear weights are calculated by measuring the distances between the predicted point and its neighboring grid locations. Crucially, these weights are non-zero only with respect to the four nearest grid points, ensuring smooth interpolation in the feature space. The feature map obtained after interpolation sampling is... .

[0047] Based on this, we introduce, for example Figure 2 The state-space model shown performs temporal modeling of dynamically sampled sequence features. SSM can map two-dimensional image structures to one-dimensional sequence representations with state memory capabilities, capturing long-range dependencies across regions using a linear recursive structure while maintaining a computational complexity of O(n). Compared to traditional self-attention mechanisms that require explicit construction of a global attention matrix, SSM can maintain the continuous representation of regions without increasing computational cost.

[0048] Therefore, the dynamic scanning mechanism provides a spatially continuous sampling sequence, while SSM provides global sequence state modeling. The two work together to enable the model to maintain the oral cavity structure morphology without fragmentation even when the sequence is unfolded, significantly improving the ability to preserve tooth margins, gingival interfaces, and tongue segmentation.

[0049] S3. In the feature extraction stage, a frequency domain enhancement mechanism is introduced to perform wavelet decomposition and frequency domain separation and fusion on the input image features, balance the distribution of high and low frequency features, enhance the mid-frequency components, and suppress noise and illumination interference.

[0050] This invention introduces a frequency domain enhancement mechanism in the feature extraction stage, performing wavelet decomposition and spectral pooling on the features of the input image to balance the representation of different frequency components. In specific implementation, the dynamically scanned feature map... The input wavelet decomposition module uses the two-dimensional discrete wavelet transform (DWT) to decompose the features into four sub-bands: the low-frequency component LL and the high-frequency components LH, HL, and HH. The wavelet decomposition process is represented as follows:

[0051]

[0052] Among them, LL contains low-frequency global structural information of the image, while LH, HL, and HH contain high-frequency detail information in the horizontal, vertical, and diagonal directions.

[0053] The low-frequency subband preserves the overall structure and tissue contour, while the high-frequency subband contains detailed features such as tooth edges and soft tissue texture. Subsequently, a 1×1 convolution is applied to each subband to enhance the representational power between subbands, while simultaneously achieving channel compression and noise suppression. After weighted fusion, the enhanced feature map Y is reconstructed through inverse wavelet transform (IDWT).

[0054]

[0055] To further improve the segmentation degradation problem caused by the imbalance between high and low frequencies in dental images, this invention introduces frequency domain separation and fusion operations after band reconstruction.

[0056] Frequency domain separation and fusion are employed to maintain consistency with the above formula. A Fourier transform Z=F(Y) is performed on the reconstructed feature map Y, and the low-frequency region is selected using the spectrum centering function G(⋅). :

[0057]

[0058] The high-frequency region is defined as:

[0059] Then, using the equilibrium parameters Integrating high and low frequency components:

[0060] Finally, the enhanced frequency domain information features are mapped back to the spatial domain and used in conjunction with the wavelet reconstruction results to generate the final enhanced features suitable for subsequent decoding.

[0061] S4. Input the features after dynamic scanning and frequency domain enhancement into the decoder module to restore the spatial resolution and generate a segmentation mask;

[0062] The decoder structure consists of symmetrical upsampling and convolutional modules, with the input being the enhancement features. Specifically, each decoder layer includes either a deconvolution operation or an upsampling interpolation combined with a convolution operation. Let the output of the l-th layer be:

[0063]

[0064] Where Up represents the upsampling operation, which magnifies the spatial size by a factor of 2; Conv is the convolution operation; BN is normalization; and σ is the activation function. This invention introduces a skip connection mechanism during the decoding process, concatenating features from the corresponding layers of the encoder to the decoder layer to maintain the fusion of low-level spatial information and high-level semantic information. Finally, the decoder outputs a feature map as follows: ,in This indicates the number of categories. In this embodiment, the value is 5, corresponding to the five categories of teeth, gums, tongue, lips, and cheeks.

[0065] After upsampling, the decoder fuses the encoder output features at the corresponding level with the upsampled decoded features using a skip connection mechanism. Fusion methods include channel-wise concatenation and pixel-wise addition; this embodiment uses concatenation to preserve information from both types of features to the greatest extent possible. Subsequently, a 3×3 convolutional kernel is used to further compress the number of channels and perform feature fusion, thereby achieving an effective combination of shallow edge features and deep semantic features. In this process, low-level features mainly contain local details such as tooth boundaries and gingival margins, while high-level features encode the overall oral structure and semantic relationships between tissues. Through concatenation and convolutional fusion operations, the clarity of the tooth contour and the consistency of the overall structure can be effectively preserved in the prediction results.

[0066] In each decoding stage, to enhance the expressive power of multi-scale features, this embodiment introduces a dilated convolution module on the fused features. This module employs a multi-branch dilated convolution structure with dilation rates of 1, 3, 5, and 7, and performs channel-by-channel concatenation and convolutional compression after the branch outputs to ensure that the network can simultaneously perceive feature patterns under different receptive fields. In dental image segmentation tasks, the introduction of this multi-scale convolution can strengthen the model's ability to handle both fine tissues and large-scale structures, avoiding boundary breaks and missed detections caused by a single convolution scale. Subsequently, the fused features are numerically constrained through layer normalization and activation functions before being input into the next stage's upsampling module until the spatial resolution is restored to the same level as the input image.

[0067] S5 outputs the segmentation results of oral cavity areas such as teeth, gums, tongue, lips, and cheeks.

[0068] The output layer uses 1×1 convolution to compress the channels of the decoded result, obtaining the class prediction probability for each pixel location. Let the output mask be:

[0069]

[0070] The Softmax operation ensures that the sum of the predicted probabilities for all classes is 1. This ultimately yields the segmentation mask. Each pixel is assigned the label with the highest probability among five categories, enabling automatic segmentation of teeth, gums, tongue, lips, and cheeks. The output is visualized as an image with different colored masks, clearly distinguishing different oral tissue regions.

[0071] To verify the effectiveness of the dental image segmentation method based on frequency domain enhancement and dynamic scanning mechanism proposed in this invention, the entire network was trained end-to-end. During training, preprocessed oral images were used as input, and a weighted combination of cross-entropy loss and Dice loss was adopted as the optimization objective to simultaneously improve the overall accuracy of region recognition and the fine-grained segmentation capability of boundaries. The optimizer employed AdamW or SGD, and a learning rate decay strategy was used to stabilize the training process. Model training was performed on a fully labeled dental dataset, and training, validation, and test sets were defined to ensure the reliability of the evaluation results.

[0072] To verify the effectiveness of the dental image segmentation method based on frequency domain enhancement and dynamic scanning mechanism proposed in this invention, we compared its segmentation results with current mainstream methods (U-Mamba and HQ-SAM). As shown in Figure 3, in multiple targets under four different scenarios, the segmentation results of this invention are superior to U-Mamba and HQ-SAM, demonstrating better segmentation performance. This invention performs better in terms of boundary preservation, detail restoration, and segmentation continuity at complex tissue boundaries. For example, in complex scenarios including tooth-gingival transitions, tongue coverage, and cheek occlusion, this invention can more accurately locate tissue edges and reduce pseudo-segmentation areas.

[0073] Based on the embodiments of the above method, a system for implementing the method is provided. The system includes: a preprocessing module, a dynamic scanning module, a frequency domain enhancement module, and a decoding module.

[0074] The preprocessing module is used to preprocess the acquired dental images;

[0075] The dynamic scanning module is used to receive dental images and achieve adaptive sampling and sequence generation through offset prediction network and bilinear interpolation; the frequency domain enhancement module is used to perform wavelet decomposition and spectral pooling on the output of the dynamic scanning module to balance high and low frequency information and enhance mid-frequency features; the decoding module is used to receive the frequency domain enhanced features, decode them, and generate the final segmentation mask.

[0076] Corresponding to the aforementioned embodiment of a dental image segmentation method based on frequency domain enhancement and dynamic scanning mechanism, the present invention also provides an embodiment of a dental image segmentation device based on frequency domain enhancement and dynamic scanning mechanism.

[0077] See Figure 4 The present invention provides a dental image segmentation device based on frequency domain enhancement and dynamic scanning mechanism, comprising a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it is used to implement a dental image segmentation method based on frequency domain enhancement and dynamic scanning mechanism in the above embodiment.

[0078] The dental image segmentation device based on frequency domain enhancement and dynamic scanning mechanism provided by this invention can be applied to any device with data processing capabilities, such as a computer. The device embodiment can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, such as... Figure 4 The diagram shown is a hardware structure diagram of any device with data processing capabilities, including the dental image segmentation device based on frequency domain enhancement and dynamic scanning mechanism provided by this invention. (Except for...) Figure 4 In addition to the processor, memory, network interface, and non-volatile memory shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.

[0079] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0080] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0081] This invention also provides a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements a dental image segmentation method based on frequency domain enhancement and dynamic scanning mechanism as described in the above embodiments.

[0082] The computer-readable storage medium can be an internal storage unit of any data processing device described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device of any data processing device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of any data processing device. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.

[0083] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the dental image segmentation method based on frequency domain enhancement and dynamic scanning mechanism.

[0084] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.

[0085] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. This application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A dental image segmentation method and system based on frequency domain enhancement and dynamic scanning mechanism, characterized in that, The method includes the following steps: S1. Acquire the input dental image and perform preprocessing operations to ensure the consistency and robustness of the input data; S2. A dynamic scanning mechanism is introduced for feature extraction. The sampling position and scanning order are adaptively adjusted by the offset prediction network to obtain a feature sequence that maintains spatial continuity. The feature map is obtained after sampling. S3. After the feature extraction stage, a frequency domain enhancement mechanism is introduced to perform wavelet decomposition, band reconstruction and frequency domain separation and fusion on the feature map of the input image. The enhanced frequency domain information is mapped back to the spatial domain and used in conjunction with the wavelet reconstruction results to generate the final enhanced features. S4. Input the final enhanced features into the decoder to restore the spatial resolution and generate a segmentation mask; S5. Perform channel compression on the decoding results and output the segmentation results of the oral cavity region.

2. The dental image segmentation method based on frequency domain enhancement and dynamic scanning mechanism according to claim 1, characterized in that, The dynamic scanning mechanism includes: The two-dimensional feature map is expanded into a one-dimensional sequence in row-major order. The one-dimensional sequence is then input into the offset prediction network to predict the offset vector of each sampling point. The updated sampling position is obtained by adding the offset position to the original coordinates, thereby adjusting the scanning order.

3. The dental image segmentation method based on frequency domain enhancement and dynamic scanning mechanism according to claim 1, characterized in that, The offset prediction network consists of multiple sequentially connected sub-modules. First, it extracts local spatial information of the input feature sequence through depthwise separable convolutional layers, reducing computational complexity while preserving the independence between channels. Then, layer normalization is introduced to ensure the numerical stability of features during transmission and avoid gradient anomalies. On this basis, the GELU activation function is used to perform nonlinear transformation on the features to enhance the network's ability to fit complex patterns. Finally, a linear fully connected layer maps the high-dimensional features to offset prediction results with the same spatial resolution as the input feature map, thereby forming a complete offset vector field output.

4. The dental image segmentation method based on frequency domain enhancement and dynamic scanning mechanism according to claim 1, characterized in that, The feature map obtained after sampling is specifically obtained by extracting information from the feature map using bilinear interpolation, calculated as follows: in, For interpolation functions, The updated sampling point coordinates, This is the original feature map.

5. The dental image segmentation method based on frequency domain enhancement and dynamic scanning mechanism according to claim 1, characterized in that, The wavelet decomposition process includes: performing wavelet transform on the feature map after dynamic scanning to decompose it into low-frequency global structure information, horizontal high-frequency detail information, vertical high-frequency detail information, and diagonal high-frequency detail information; Wavelet transform is used to decompose the features of the input image to obtain low-frequency and high-frequency sub-bands, as shown in the following formula: in, Indicates input features, These represent the low-frequency and high-frequency sub-bands, respectively.

6. The dental image segmentation method based on frequency domain enhancement and dynamic scanning mechanism according to claim 5, characterized in that, The band reconstruction specifically refers to: The frequency domain subbands are reconstructed through convolution and inverse wavelet transform to obtain the enhanced feature map: in, This represents a one-dimensional convolution operation, and IDWT represents inverse wavelet transform. This step enhances the high-frequency information at the boundary while preserving the low-frequency structure, thereby improving the segmentation accuracy of the model under complex imaging conditions.

7. The dental image segmentation method based on frequency domain enhancement and dynamic scanning mechanism according to claim 1, characterized in that, The frequency domain separation and fusion specifically includes: separating high and low frequency information from the reconstructed feature map Y in the frequency domain and then fusing them, as shown in the following formula: in, Indicates Fourier transform, Indicates inverse transformation, It is the Fourier transform centering function, where, This is a balancing parameter used to control the high- and low-frequency fusion ratio.

8. The dental image segmentation method based on frequency domain enhancement and dynamic scanning mechanism according to claim 1, characterized in that, The decoder and the encoder used for feature extraction both consist of several layers. During the decoding process, a skip connection structure is adopted to fuse the frequency domain enhanced features with the upsampled decoded features, thereby improving the segmentation accuracy of the tooth boundary region and finally obtaining the predicted segmentation mask.

9. A system for implementing the method according to any one of claims 1-8, characterized in that, The system includes: a preprocessing module, a dynamic scanning module, a frequency domain enhancement module, and a decoding module; The preprocessing module is used to preprocess the acquired dental images; The dynamic scanning module is used to receive dental images and achieve adaptive sampling and sequence generation through offset prediction network and bilinear interpolation; the frequency domain enhancement module is used to perform wavelet decomposition and spectral pooling on the output of the dynamic scanning module to balance high and low frequency information and enhance mid-frequency features; the decoding module is used to receive the frequency domain enhanced features, decode them, and generate the final segmentation mask.

10. A dental image segmentation device based on frequency domain enhancement and dynamic scanning mechanism, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that... When the processor executes the executable code, it implements the method as described in any one of claims 1-8.