Tooth image intelligent analysis system and method for oral surgery

By using technical means such as backbone network and spatial adaptive modules in the intelligent dental image analysis system, the problem of insufficient extraction of complex structures and deep features in dental image analysis is solved, and higher analysis accuracy and robustness are achieved, providing more reliable support for oral surgery.

CN120107254AActive Publication Date: 2025-06-06JILIN UNIVERSITY

Patent Information

Application Number
CN202510585504.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

Existing dental image analysis techniques are difficult to accurately deal with complex tooth structures, blurred boundaries and insufficient extraction of deep features, resulting in limitations in the accuracy and consistency of diagnostic results.

Method used

An intelligent dental image analysis system for oral surgery is proposed. Multi-scale feature maps are extracted through the backbone network, and combined with spatial adaptive modules and edge detection branches, middle- and deep feature maps and edge information are fused to generate multi-scale significant fusion feature maps, and finally the tooth voxel probability map and discrete segmentation label map are obtained through the classification layer.

Benefits of technology

It improves the accuracy and robustness of dental image analysis, can capture details and structures in dental images more accurately, enhances the positioning ability of the lesion site, and provides more reliable support for oral surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107254A_ABST
    Figure CN120107254A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of tooth image analysis, and discloses an intelligent tooth image analysis system and method for oral surgery, and the method comprises the steps: firstly obtaining tooth image data, and extracting a multi-scale feature map through a backbone network, and the multi-scale feature map comprises a shallow-layer feature map, a middle-layer feature map and a deep-layer feature map; then, strengthening the spatial characteristics of the middle-layer and deep-layer feature maps by adopting a spatial self-adaptive module, and refining the shallow-layer feature map by utilizing edge detection branches so as to capture tooth edge details; then, the feature maps are fused to generate a multi-scale significant fusion feature map, and the multi-scale significant fusion feature map is processed by a classification layer to obtain a tooth voxel probability map. And finally, converting into a discrete segmentation label graph as a semantic segmentation result. According to the method, multi-scale features and spatial context information are integrated, so that the accuracy and robustness of tooth image analysis are improved, and more reliable support is provided for oral surgery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of dental image analysis, and more specifically, to a dental image intelligent analysis system and method for oral surgery. Background Art

[0002] In the field of oral surgery, accurate dental image analysis is essential for diagnosis and treatment planning. Traditional dental image analysis methods often rely on the doctor's experience and expertise, which is not only time-consuming but also easily affected by subjective factors, resulting in limitations in the consistency and accuracy of diagnostic results. With the development of medical imaging technology, digital image processing has become one of the important means to improve diagnostic efficiency and accuracy. However, existing digital image processing technology still faces many challenges when faced with complex tooth structures, blurred boundaries, and insufficient feature extraction.

[0003] First, the anatomical structure of teeth and their surrounding tissues is very complex, and there are huge differences between different individuals. Traditional image analysis methods are difficult to capture these detailed information comprehensively and accurately, especially when it comes to detecting subtle lesions. Secondly, due to the influence of factors such as shooting angle and equipment resolution, the boundaries in tooth images are often not clear enough, which brings difficulties to subsequent segmentation and recognition. Furthermore, how to effectively extract representative feature information from tooth images is the key to achieving automated analysis. However, most current methods focus on the extraction of shallow features and pay insufficient attention to deep features, resulting in inaccurate positioning of lesions.

[0004] Therefore, an intelligent dental imaging analysis solution for oral surgery is desired. Summary of the invention

[0005] In view of the problems that the existing technology has in processing dental images, such as being sensitive to noise, difficulty in accurately segmenting complex structures, and ignoring spatial context information, the present invention proposes a dental image intelligent analysis system and method for oral surgery.

[0006] According to one aspect of the present application, a method for intelligent analysis of dental images for oral surgery is provided, comprising: acquiring dental image data for oral surgery; inputting the dental image data into a backbone network to obtain a dental image shallow feature map, a dental image middle feature map and a dental image deep feature map; inputting the dental image middle feature map and the dental image deep feature map into a spatial adaptive module to obtain a dental image middle spatial enhancement feature map and a dental image deep spatial enhancement feature map; inputting the dental image shallow feature map into an edge detection branch to obtain a dental edge information feature map; fusing the dental image middle spatial enhancement feature map, the dental image deep spatial enhancement feature map and the dental edge information feature map to obtain a dental image multi-scale significant fusion feature map; inputting the dental image multi-scale significant fusion feature map into a classification layer to obtain a dental voxel probability map; and converting the dental voxel probability map into a dental voxel discrete segmentation label map as a dental image semantic segmentation result.

[0007] In the above-mentioned intelligent analysis method of dental images for oral surgery, the shallow feature map of the dental image is input into the edge detection branch to obtain the tooth edge information feature map, including: inputting the shallow feature map of the dental image into the convolution layer of the edge detection branch to obtain the tooth image edge feature focus feature map; inputting the tooth image edge feature focus feature map into the activation layer of the edge detection branch to obtain the tooth edge feature mask feature map, and the activation layer uses the Sigmoid activation function; calculating the position point multiplication between the tooth edge feature mask feature map and the shallow feature map of the dental image to obtain the tooth edge information feature map.

[0008] In the above-mentioned intelligent analysis method of dental images for oral surgery, the backbone network is a SwinTransformer network.

[0009] In the above-mentioned intelligent analysis method of dental images for oral surgery, the middle-layer feature map of the dental image and the deep-layer feature map of the dental image are input into a spatial adaptive module to obtain a middle-layer spatial enhancement feature map of the dental image and a deep-layer spatial enhancement feature map of the dental image, including: inputting the middle-layer feature map of the dental image into an offset learning module based on a convolutional layer to obtain a middle-layer feature offset map of the dental image; inputting the middle-layer feature map of the dental image into a modulation scalar learning module based on a convolutional layer to obtain a middle-layer feature modulation map of the dental image; combining the middle-layer feature offset map of the dental image and the middle-layer feature modulation map of the dental image, and inputting the middle-layer feature map of the dental image into a spatial adaptive enhancement component based on a deformable convolutional layer to obtain the middle-layer spatial enhancement feature map of the dental image.

[0010] In the above-mentioned intelligent analysis method of dental images for oral surgery, the dental image mid-layer feature offset map and the dental image mid-layer feature modulation map are combined, and the dental image mid-layer feature map is input into a spatial adaptive enhancement component based on a deformable convolutional layer to obtain the dental image mid-layer spatial enhancement feature map, including: performing bilinear interpolation sampling on the dental image mid-layer feature map based on the dental image mid-layer feature offset map to obtain the dental image mid-layer sampling feature map; calculating the position point multiplication between the dental image mid-layer sampling feature map and the dental image mid-layer feature modulation map to obtain the dental image mid-layer spatial enhancement feature map.

[0011] In the above-mentioned intelligent analysis method of dental images for oral surgery, the middle-layer spatial enhancement feature map of the dental image, the deep-layer spatial enhancement feature map of the dental image and the tooth edge information feature map are fused to obtain a multi-scale significant fusion feature map of the dental image, including: inputting the middle-layer spatial enhancement feature map of the dental image, the deep-layer spatial enhancement feature map of the dental image and the tooth edge information feature map into a jump connection layer to obtain an initial multi-scale significant fusion down-sampling feature map of the dental image; performing complementary fusion optimization of visual semantic features on the initial multi-scale significant fusion down-sampling feature map of the dental image to obtain a multi-scale significant fusion down-sampling feature map of the dental image; and upsampling the multi-scale significant fusion down-sampling feature map of the dental image to obtain the multi-scale significant fusion feature map of the dental image.

[0012] In the above-mentioned intelligent analysis method of dental images for oral surgery, the complementary fusion optimization of visual semantic features is performed on the initial multi-scale significant fusion down-sampling feature map of the dental image to obtain the multi-scale significant fusion down-sampling feature map of the dental image, including: performing void convolution operations with different void rates on the middle-layer space enhancement feature map of the dental image, the deep-layer space enhancement feature map of the dental image and the tooth edge information feature map to obtain the middle-layer space enhancement void convolution feature map of the dental image, the deep-layer space enhancement void convolution feature map of the dental image and the tooth edge information void convolution feature map; taking the middle-layer space enhancement void convolution feature map of the dental image as a benchmark, calculating The cross-layer geometric feature responses of the deep spatial enhanced cavity convolution feature map of the dental image and the tooth edge information cavity convolution feature map are obtained to obtain a first dental image cross-layer geometric feature response feature map and a second dental image cross-layer geometric feature response feature map; based on the shift window mechanism and the relative position encoding mechanism, the first dental image cross-layer geometric feature response feature map and the second dental image cross-layer geometric feature response feature map are fused and corrected to obtain a dental image multi-scale corrected feature map; the dental image multi-scale corrected feature map is point-multiplied with the initial dental image multi-scale significantly fused down-sampled feature map to obtain the dental image multi-scale significantly fused down-sampled feature map.

[0013] In the above-mentioned intelligent analysis method of dental images for oral surgery, the cavity rates of the cavity convolution operation are 3, 5, and 1 respectively.

[0014] In the above-mentioned intelligent analysis method of dental images for oral surgery, the classification layer includes a point convolution layer and a Softmax classification unit.

[0015] According to another aspect of the present application, a dental image intelligent analysis system for oral surgery is also provided, comprising: a dental image data acquisition module, used to acquire dental image data for oral surgery; a dental image multi-scale feature extraction module, used to input the dental image data into a backbone network to obtain a dental image shallow feature map, a dental image middle feature map and a dental image deep feature map; a dental image spatial enhancement module, used to input the dental image middle feature map and the dental image deep feature map into a spatial adaptation module to obtain a dental image middle spatial enhancement feature map and a dental image deep spatial enhancement feature map; a dental image The edge detection module is used to input the shallow feature map of the tooth image into the edge detection branch to obtain the tooth edge information feature map; the tooth image multi-scale feature fusion module is used to fuse the middle-layer spatial enhancement feature map of the tooth image, the deep-layer spatial enhancement feature map of the tooth image and the tooth edge information feature map to obtain the multi-scale significant fusion feature map of the tooth image; the tooth image classification processing module is used to input the multi-scale significant fusion feature map of the tooth image into the classification layer to obtain the tooth voxel probability map; the tooth image semantic segmentation module is used to convert the tooth voxel probability map into a tooth voxel discrete segmentation label map as the tooth image semantic segmentation result.

[0016] Compared with the prior art, the dental image intelligent analysis system and method for oral surgery provided by the present application first obtains dental image data and extracts multi-scale feature maps, including shallow, middle and deep feature maps, through a backbone network. Subsequently, a spatial adaptive module is used to enhance the spatial characteristics of the middle and deep feature maps, while an edge detection branch is used to refine the shallow feature map to capture the details of the tooth edges. Next, these feature maps are fused to generate a multi-scale significant fusion feature map, which is then processed by the classification layer to obtain a tooth voxel probability map. Finally, it is converted into a discrete segmentation label map as a semantic segmentation result. This method aims to improve the accuracy and robustness of dental image analysis by integrating multi-scale features with spatial context information, and to provide more reliable support for oral surgery. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0018] Figure 1 The figure shows a schematic flow chart of a method for intelligent analysis of dental images for oral surgery according to an embodiment of the present application.

[0019] Figure 2 The figure illustrates a schematic flow chart of S3 in the intelligent analysis method of dental images for oral surgery according to an embodiment of the present application.

[0020] Figure 3 The figure illustrates a schematic flow chart of S33 in the intelligent analysis method of dental images for oral surgery according to an embodiment of the present application.

[0021] Figure 4 The figure illustrates a schematic flow chart of S4 in the intelligent analysis method of dental images for oral surgery according to an embodiment of the present application.

[0022] Figure 5 The figure illustrates a schematic flow chart of S5 in the intelligent analysis method of dental images for oral surgery according to an embodiment of the present application.

[0023] Figure 6 The figure shows a schematic block diagram of a dental image intelligent analysis system for oral surgery according to an embodiment of the present application. DETAILED DESCRIPTION

[0024] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.

[0025] Figure 1 FIG. 1 is a schematic flow chart of a method for intelligent analysis of dental images for oral surgery according to an embodiment of the present application. Figure 1As shown, the present application provides a dental image intelligent analysis method for oral surgery, including: S1, obtaining dental image data for oral surgery; S2, inputting the dental image data into a backbone network to obtain a dental image shallow feature map, a dental image middle feature map and a dental image deep feature map; S3, inputting the dental image middle feature map and the dental image deep feature map into a spatial adaptive module to obtain a dental image middle spatial enhancement feature map and a dental image deep spatial enhancement feature map; S4, inputting the dental image shallow feature map into an edge detection branch to obtain a dental edge information feature map; S5, fusing the dental image middle spatial enhancement feature map, the dental image deep spatial enhancement feature map and the dental edge information feature map to obtain a dental image multi-scale significant fusion feature map; S6, inputting the dental image multi-scale significant fusion feature map into a classification layer to obtain a dental voxel probability map; S7, converting the dental voxel probability map into a dental voxel discrete segmentation label map as a dental image semantic segmentation result.

[0026] Specifically, in step S1, dental image data for oral surgery is obtained. It should be understood that CBCT has been widely used in the field of oral surgery, considering its advantages of being able to provide three-dimensional views, low radiation dose and high image resolution. Therefore, by scanning the patient with a CBCT device, a three-dimensional image data set containing detailed information of the teeth and their surrounding structures can be obtained. These data include not only the morphological characteristics of the teeth themselves, but also the adjacent bone structures and soft tissue conditions, providing a rich information basis for subsequent analysis.

[0027] In a specific embodiment, the doctor or technician will select appropriate scanning parameters according to clinical needs, such as field of view (FOV), scanning time and resolution, to ensure that high-quality images that meet diagnostic requirements and minimize the amount of radiation received by the patient are obtained. Once the settings are completed, the patient will be guided to the correct position and instructed to remain still to complete the scanning process smoothly. Subsequently, the generated raw image data will be transmitted to a dedicated image processing workstation, where technicians use specific software tools to pre-process the data, including noise reduction, artifact correction and other steps, to improve image quality and prepare for further analysis.

[0028] Specifically, in step S2, the dental image data is input into the backbone network to obtain a shallow feature map of the dental image, a middle feature map of the dental image, and a deep feature map of the dental image. It should be understood that although the traditional convolutional neural network performs well in many visual tasks, it has limitations in processing long-distance dependencies and global information. In contrast, the backbone network uses the self-attention mechanism within the local window, combined with the moving window strategy to capture information exchange across windows, which can maintain computational efficiency and effectively capture global dependencies. This is crucial for parsing subtle structural differences in dental images. In addition, by constructing feature maps hierarchically, it can ensure that the model can not only recognize obvious anatomical structures, but also deeply understand the deep semantic information hidden behind them. This method helps to improve the working effect of subsequent spatial adaptive enhancement processing and edge detection branches, and lays a solid foundation for the ultimate realization of accurate tooth segmentation and classification.

[0029] Specifically, in one embodiment, the backbone network is a Swin Transformer network. As an improved Transformer architecture, Swin Transformer introduces a hierarchical feature representation method that can effectively capture feature information at different scales. In practical applications, Swin Transformer gradually extracts shallow feature maps, middle feature maps, and deep feature maps of dental images from the input dental image data through a series of stages, each of which contains multiple levels of self-attention mechanisms.

[0030] In a specific embodiment, first, Swin Transformer performs preliminary feature extraction on the input dental image data to generate relatively basic shallow feature maps of dental images. These feature maps contain basic information such as edges and color contrast. As the depth of the network increases, in the intermediate stage, the model begins to focus on more abstract structural features, such as the relative position relationship between teeth or the morphological characteristics of certain specific areas, thereby generating mid-level feature maps of dental images. In the last one or several stages, the model will further refine more complex and representative deep feature maps of dental images. These feature maps often correspond to higher-level concepts, such as the three-dimensional shape of the entire tooth or its relationship with other tissues.

[0031] Specifically, in step S3, the middle-layer feature map of the dental image and the deep-layer feature map of the dental image are input into the spatial adaptive module to obtain the middle-layer spatial enhanced feature map of the dental image and the deep-layer spatial enhanced feature map of the dental image. It should be understood that although the traditional convolutional neural network performs well in many visual tasks, it has limitations in processing medical images with significant scale changes and non-uniform distribution characteristics. Especially in the field of oral surgery, the morphology of teeth and their surrounding tissues varies and is intertwined, which places higher requirements on feature extraction. By introducing the spatial adaptive module, the sensitivity and robustness of the model to detail features can be greatly improved without significantly increasing the computational cost. In addition, this method also allows the model to automatically adjust its internal parameters according to the characteristics of the input data, thereby improving the adaptability to various types of image data.

[0032] Specifically, the processing of the spatial adaptive module is described by taking the middle layer feature map of the tooth image as an example. Figure 2 As shown, in step S3, the middle-layer feature map of the tooth image and the deep-layer feature map of the tooth image are input into the spatial adaptive module to obtain the middle-layer spatial enhancement feature map of the tooth image and the deep-layer spatial enhancement feature map of the tooth image, including: S31, inputting the middle-layer feature map of the tooth image into the offset learning module based on the convolution layer to obtain the middle-layer feature offset map of the tooth image; S32, inputting the middle-layer feature map of the tooth image into the modulation scalar learning module based on the convolution layer to obtain the middle-layer feature modulation map of the tooth image; S33, combining the middle-layer feature offset map of the tooth image and the middle-layer feature modulation map of the tooth image, inputting the middle-layer feature map of the tooth image into the spatial adaptive enhancement component based on the deformable convolution layer to obtain the middle-layer spatial enhancement feature map of the tooth image. The processing process of inputting the deep-layer feature map of the tooth image into the spatial adaptive module can also refer to the above-mentioned processing process for the middle-layer feature map of the tooth image.

[0033] Specifically, the spatial adaptation module consists of two parts: a convolutional layer-based offset learning module and a convolutional layer-based modulation scalar learning module. The offset learning module is usually composed of a series of convolutional layers, and its function is to learn an offset vector for each pixel position to indicate the adjustment direction of the center position of the local area that should be paid attention to at that position. The core of this step is that it allows the model to dynamically adjust the receptive field to better adapt to irregular shapes or boundaries in the image. For example, when processing dental images, this mechanism helps to more accurately locate the junction between teeth and gums or other subtle structures. At the same time, the modulation scalar learning module also uses a convolutional layer structure, but its goal is to generate a scalar value for each pixel position, which is used to adjust the feature response intensity at the corresponding position. This means that in addition to changing the position of the receptive field, the information weights of each part of the feature map can also be dynamically adjusted according to the importance of the image content. For example, when emphasizing certain key anatomical landmarks, the response intensity of the relevant area can be increased, while it can be appropriately weakened in the background area, so as to improve the effectiveness and pertinence of the overall feature expression.

[0034] In a specific embodiment, for a convolutional layer-based offset learning module, for example, there is an input feature map F with a size (height, width, number of channels). In order to generate the mid-level feature offset map of the dental image, a network consisting of several convolutional layers is designed in this application. First, a standard convolutional layer (including ReLU activation function) is used, and the number of output channels is set to 2C (because each position requires two values ​​to represent the offset on the x-axis and y-axis). Then, another 1x1 convolutional layer is used to reduce the number of channels to 2, so that the final offset map O is obtained, with a size of . Each element here represents the offset that should be applied to the corresponding position.

[0035] For the convolutional layer-based modulation scalar learning module, continuing to use the above input feature map F as an example, in order to generate the mid-level feature modulation map of the dental image, this application constructs a similar network structure. First, a convolutional layer (which may also include ReLU activation) is used to process the input feature map, but this time the number of output channels is set to C. Then, a 1x1 convolutional layer with a Sigmoid activation function is passed to ensure that the output value is between 0 and 1, and a modulation scalar map M with a size of Each value in this scalar map M represents the enlargement or reduction ratio of the feature map at the corresponding position.

[0036] Subsequently, the offset map and modulation map generated by the above two modules are fed into the spatial adaptive enhancement component based on the deformable convolution layer. Figure 3As shown, in step S33, the middle-layer feature offset map of the dental image and the middle-layer feature modulation map of the dental image are combined, and the middle-layer feature map of the dental image is input into the spatial adaptive enhancement component based on the deformable convolution layer to obtain the middle-layer spatial enhancement feature map of the dental image, including: S331, based on the middle-layer feature offset map of the dental image, bilinear interpolation sampling is performed on the middle-layer feature map of the dental image to obtain the middle-layer sampling feature map of the dental image; S332, the position point multiplication between the middle-layer sampling feature map of the dental image and the middle-layer feature modulation map of the dental image is calculated to obtain the middle-layer spatial enhancement feature map of the dental image.

[0037] Specifically, in step S4, the tooth image shallow feature map is input into the edge detection branch to obtain a tooth edge information feature map. The edge detection branch is a network used for edge detection. It should be understood that the edge detection branch can significantly improve the accuracy of edge detection without sacrificing details, and provide strong support for subsequent multi-scale feature fusion and tooth voxel segmentation tasks.

[0038] In one embodiment, Figure 4 As shown, in step S4, the shallow feature map of the tooth image is input into the edge detection branch to obtain the tooth edge information feature map, including: S41, inputting the shallow feature map of the tooth image into the convolution layer of the edge detection branch to obtain the tooth image edge feature focus feature map; S42, inputting the tooth image edge feature focus feature map into the activation layer of the edge detection branch to obtain the tooth edge feature mask feature map, and the activation layer uses the Sigmoid activation function; S43, calculating the position point multiplication between the tooth edge feature mask feature map and the shallow feature map of the tooth image to obtain the tooth edge information feature map.

[0039] Specifically, in step S5, the middle-layer spatial enhancement feature map of the dental image, the deep-layer spatial enhancement feature map of the dental image, and the tooth edge information feature map are fused to obtain a multi-scale significant fusion feature map of the dental image. It is often difficult to fully capture all important information in the dental image by relying solely on a feature map of a certain level. For example, a shallow feature map may contain rich detail information, but lacks understanding of the overall structure; on the contrary, although a deep feature map is good at identifying a large range of structural patterns, it may have defects in detail description. Therefore, by fusing feature maps at different levels and using jump connections, these problems can be effectively overcome to ensure that the generated multi-scale significant fusion feature map has sufficient detail clarity and can accurately reflect global structural characteristics.

[0040] In one embodiment, Figure 5As shown, in step S5, the middle-layer spatial enhancement feature map of the dental image, the deep-layer spatial enhancement feature map of the dental image and the tooth edge information feature map are fused to obtain a multi-scale significant fusion feature map of the dental image, including: S51, inputting the middle-layer spatial enhancement feature map of the dental image, the deep-layer spatial enhancement feature map of the dental image and the tooth edge information feature map into a jump connection layer to obtain an initial multi-scale significant fusion down-sampling feature map of the dental image; S52, performing complementary fusion optimization of visual semantic features on the initial multi-scale significant fusion down-sampling feature map of the dental image to obtain a multi-scale significant fusion down-sampling feature map of the dental image; S53, upsampling the multi-scale significant fusion down-sampling feature map of the dental image to obtain the multi-scale significant fusion feature map of the dental image.

[0041] Specifically, the feature maps from the encoder part (corresponding to the middle and deep feature maps output by the backbone network) are directly passed to the decoder part (i.e., the feature fusion stage) through the jump connection. This connection method not only helps to alleviate the gradient vanishing problem, but also ensures that important details of the original image are retained in the process of restoring the image resolution. Inside the jump connection layer, in order to fuse feature maps from different sources, addition or splicing may be used. For example, the middle-layer spatial enhancement feature map and the deep-layer spatial enhancement feature map can be spliced ​​to form a new feature map, and then added to the tooth edge information feature map to obtain the initial tooth image multi-scale significant fusion down-sampled feature map.

[0042] Preferably, considering that when fusing the mid-layer spatial enhancement feature map of the dental image, the deep-layer spatial enhancement feature map of the dental image and the dental edge information feature map, since the mid-layer spatial enhancement feature map of the dental image and the deep-layer spatial enhancement feature map of the dental image are respectively enhanced by offset and modulation scalar based on the mid-layer visual feature space representation and the deep-layer visual feature space representation of the dental image data, and the dental edge information feature map is also enhanced by edge feature activation in the shallow-layer visual feature space representation of the dental image data, it is expected that complementary fusion of visual semantic features can be achieved while maintaining the integrity of spatial details during fusion.

[0043] Based on this, in the present application, the complementary fusion optimization of visual semantic features is performed on the initial multi-scale significant fusion down-sampling feature map of the tooth image to obtain the multi-scale significant fusion down-sampling feature map of the tooth image, including: first, for the mid-layer spatial enhancement feature map of the tooth image , the deep spatial enhancement feature map of the tooth image and the tooth edge information feature map , respectively perform dilated convolution operations with dilation rates of {3, 5, 1} , and , to obtain the convolutional feature map of the middle-layer space enhancement cavity of the tooth image, the convolutional feature map of the deep-layer space enhancement cavity of the tooth image, and the convolutional feature map of the cavity of the tooth edge information, expressed as: ; ; ; in, Represents the tooth edge information cavity convolution feature map, Represents the spatial enhancement of the hollow convolution feature map in the middle layer of the tooth image, Represents the deep spatial enhancement of the dental image cavity convolution feature map. Here, the above cavity convolution operation is used to achieve multi-scale cross-layer structure decomposition association with spatial invariance.

[0044] Then, the cross-layer geometric feature response is established with the middle-layer feature as the benchmark, that is, the cross-layer geometric feature response of the deep-layer spatial enhancement cavity convolution feature map of the tooth image and the tooth edge information cavity convolution feature map is calculated to obtain the first tooth image cross-layer geometric feature response feature map and the second tooth image cross-layer geometric feature response feature map, which are expressed as: ; ; in, It means subtracting by position. It means point multiplication by position. It represents the inverse of the feature value of each position of the convolution feature map of the tooth edge information cavity. It represents the inverse of the feature value of each position of the convolution feature map of the deep spatial enhancement cavity of the tooth image. Represents the first tooth image cross-layer geometric feature response feature map, It represents the cross-layer geometric feature response feature map of the second tooth image. That is, the depth semantic confidence change compensation is performed through the depth reference geometric response, which is equivalent to the depth confidence semantic space alignment for the middle layer features.

[0045] In this way, based on the shift window mechanism and relative position encoding mechanism of the Swin Transformer network, the first tooth image cross-layer geometric feature response feature map and the second tooth image cross-layer geometric feature response feature map can be fused and corrected to obtain a multi-scale corrected feature map of the tooth image, which is expressed as: ; here, Indicates that the feature map The feature matrix of is shifted along the width and height directions , Indicates that the feature map The feature matrix of is shifted along the width and height directions , and specifically, circular shift can be used, for example, the first row is shifted up one position to fill the last row, It means adding by position. Represents the multi-scale corrected feature map of dental images.

[0046] Finally, the multi-scale correction feature map of the tooth image is point-multiplied with the multi-scale significant fusion down-sampled feature map of the initial tooth image to perform fusion optimization to obtain the multi-scale significant fusion down-sampled feature map of the tooth image. In this way, on the basis of the more essential cross-layer interactive representation of the feature set space by the deep semantic confidence, the multi-scale decomposition response coordination and deep semantic space alignment are used to perform complementary fusion for the visual feature space representation of different deep features.

[0047] Then, after completing the fusion optimization, it is necessary to perform an upsampling operation on the multi-scale significant fusion down-sampled feature map of the dental image to restore the spatial resolution of the original image. In the present application, this can be achieved through methods such as deconvolution or bilinear interpolation. Deconvolution is a commonly used upsampling technique that increases the spatial size of the feature map by learning an inverse convolution kernel; while bilinear interpolation performs interpolation calculations based on adjacent pixel values, which is simple but can also provide a good smoothing effect. Which method to choose depends on the specific task requirements and the desired level of accuracy.

[0048] Specifically, in step S6, the multi-scale significant fusion feature map of the tooth image is input into the classification layer to obtain a tooth voxel probability map, and the classification layer includes a point convolution layer and a Softmax classification unit. It should be understood that by using the point convolution layer, subtle anatomical features can be effectively captured while maintaining a high resolution; and with the help of the Softmax classification unit, these features can be mapped to specific category labels, providing doctors with intuitive and easy-to-interpret results.

[0049] In a specific embodiment, the multi-scale significant fusion feature map of the dental image is first input into the point convolution layer. The point convolution layer extracts a more abstract feature representation by applying a set of learnable filters to perform convolution operations on small areas around each voxel. Since these filters can reduce the number of parameters without losing spatial information, they help improve computational efficiency and reduce the risk of overfitting. Next, the feature map processed by the point convolution layer is sent to the Softmax classification unit. The Softmax function can convert the original output into a probability distribution so that each voxel has a probability value corresponding to a specific category. Specifically, the Softmax function is used to determine which voxels belong to the tooth structure and which do not.

[0050] Specifically, in step S7, the tooth voxel probability map is converted into a tooth voxel discrete segmentation label map as a tooth image semantic segmentation result. It should be understood that it is often difficult to directly obtain an ideal segmentation effect by relying solely on the original tooth voxel probability map, because the probability value itself does not always perfectly reflect the actual anatomical structure boundary. Therefore, further converting the tooth voxel probability map into a tooth voxel discrete segmentation label map as a tooth image semantic segmentation result can not only significantly improve the quality of the segmentation result, but also enhance the model's adaptability to various complex situations.

[0051] Specifically, after obtaining the tooth voxel probability map, first, a threshold is applied to each voxel to determine which category it belongs to. Specifically, a probability threshold (for example, 0.5, which can be adjusted according to actual conditions and is not specifically limited in this embodiment) can be set. If the probability value of a certain category corresponding to a certain voxel exceeds this threshold, the voxel is marked as this category; otherwise, it may be regarded as part of the background or other categories. This method is simple and direct, but sometimes it may not be sufficient to deal with complex boundary conditions or unevenly distributed probability values. Therefore, in practical applications, more sophisticated strategies may be necessary, such as dynamic threshold adjustment or combining multiple conditions for judgment to improve segmentation accuracy. Further, in order to ensure the consistency and accuracy of the segmentation results, connected component analysis can also be performed. Connected component analysis can help identify and separate different objects or regions in the image, which is particularly important for distinguishing different teeth that are closely adjacent or teeth and surrounding tissues. Through connected component analysis, isolated small noise points can be effectively removed, and voxels with similar attributes can be classified into one category, thereby generating clearer and more accurate segmentation results. Finally, morphological operations can also be performed. Morphological operations include dilation, erosion, opening, and closing, which can help correct holes, breaks, or other irregularities that may occur during the segmentation process. For example, appropriate dilation operations can fill segmentation gaps caused by noise or low probability values, while erosion can be used to remove unnecessary edge extensions to ensure that the final segmentation contour is as close to the actual anatomical structure as possible.

[0052] In summary, the intelligent analysis method of dental images for oral surgery provided in this application first obtains dental image data and extracts multi-scale feature maps, including shallow, middle and deep feature maps, through a backbone network. Subsequently, a spatial adaptive module is used to enhance the spatial characteristics of the middle and deep feature maps, while the edge detection branch is used to refine the shallow feature map to capture the details of the tooth edges. Next, these feature maps are fused to generate a multi-scale significant fusion feature map, which is then processed by the classification layer to obtain a tooth voxel probability map. Finally, it is converted into a discrete segmentation label map as a semantic segmentation result. This method aims to improve the accuracy and robustness of dental image analysis by integrating multi-scale features with spatial context information, and to provide more reliable support for oral surgery.

[0053] The present application also provides a dental image intelligent analysis system for oral surgery, such as Figure 6As shown, the dental image intelligent analysis system 100 for oral surgery includes: a dental image data acquisition module 11, used to acquire dental image data for oral surgery; a dental image multi-scale feature extraction module 12, used to input the dental image data into a backbone network to obtain a dental image shallow feature map, a dental image middle feature map and a dental image deep feature map; a dental image space enhancement module 13, used to input the dental image middle feature map and the dental image deep feature map into a space adaptation module to obtain a dental image middle space enhancement feature map and a dental image deep space enhancement feature map; a dental image edge detection module 14, used for inputting the shallow feature map of the tooth image into the edge detection branch to obtain the tooth edge information feature map; the tooth image multi-scale feature fusion module 15, used for fusing the middle-layer spatial enhancement feature map of the tooth image, the deep-layer spatial enhancement feature map of the tooth image and the tooth edge information feature map to obtain the multi-scale significant fusion feature map of the tooth image; the tooth image classification processing module 16, used for inputting the multi-scale significant fusion feature map of the tooth image into the classification layer to obtain the tooth voxel probability map; the tooth image semantic segmentation module 17, used for converting the tooth voxel probability map into a tooth voxel discrete segmentation label map as the tooth image semantic segmentation result.

[0054] The basic principles of the present application are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present application are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present application. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, not for limitation, and the above details do not limit the present application to being implemented by adopting the above specific details.

[0055] The block diagrams of the devices, apparatuses, equipment, and systems involved in this application are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagram. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open words, referring to "including but not limited to", and can be used interchangeably with them. The words "or" and "and" used here refer to the words "and / or" and can be used interchangeably with them, unless the context clearly indicates otherwise. The words "such as" used here refer to the phrase "such as but not limited to", and can be used interchangeably with them.

[0056] It should also be noted that in the apparatus, device and method of the present application, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present application.

[0057] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

[0058] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.

Claims

1. A dental image intelligent analysis method for oral surgery, characterized in that: include: Acquire dental imaging data for oral surgery; Inputting the tooth image data into a backbone network to obtain a tooth image shallow feature map, a tooth image middle feature map and a tooth image deep feature map; Inputting the middle-layer feature map of the tooth image and the deep-layer feature map of the tooth image into a spatial adaptive module to obtain a middle-layer spatial enhancement feature map of the tooth image and a deep-layer spatial enhancement feature map of the tooth image; Inputting the shallow feature map of the tooth image into the edge detection branch to obtain a tooth edge information feature map; Fusing the middle-layer spatial enhancement feature map of the tooth image, the deep-layer spatial enhancement feature map of the tooth image, and the tooth edge information feature map to obtain a multi-scale significant fusion feature map of the tooth image; Inputting the multi-scale significant fusion feature map of the tooth image into the classification layer to obtain a tooth voxel probability map; The tooth voxel probability map is converted into a tooth voxel discrete segmentation label map as a tooth image semantic segmentation result.

2. The method for intelligent analysis of dental images for oral surgery according to claim 1, characterized in that: Inputting the shallow feature map of the tooth image into the edge detection branch to obtain a tooth edge information feature map, including: Inputting the shallow feature map of the tooth image into the convolution layer of the edge detection branch to obtain a focused feature map of the edge feature of the tooth image; Input the tooth image edge feature focus feature map into the activation layer of the edge detection branch to obtain a tooth edge feature mask feature map, wherein the activation layer uses a Sigmoid activation function; The tooth edge information feature map is obtained by calculating the multiplication of the tooth edge feature mask feature map and the tooth image shallow feature map by position points.

3. The method for intelligent analysis of dental images for oral surgery according to claim 1, characterized in that: The backbone network is a Swin Transformer network.

4. The method for intelligent analysis of dental images for oral surgery according to claim 1, characterized in that: Inputting the middle-layer feature map of the tooth image and the deep-layer feature map of the tooth image into a spatial adaptive module to obtain a middle-layer spatial enhancement feature map of the tooth image and a deep-layer spatial enhancement feature map of the tooth image, including: Inputting the middle-layer feature map of the tooth image into an offset learning module based on a convolutional layer to obtain a middle-layer feature offset map of the tooth image; Inputting the middle-layer feature map of the tooth image into a modulation scalar learning module based on a convolutional layer to obtain a middle-layer feature modulation map of the tooth image; Combined with the tooth image mid-layer feature offset map and the tooth image mid-layer feature modulation map, the tooth image mid-layer feature map is input into a spatial adaptive enhancement component based on a deformable convolutional layer to obtain the tooth image mid-layer spatial enhancement feature map.

5. The method for intelligent analysis of dental images for oral surgery according to claim 4, characterized in that: Combining the tooth image mid-layer feature offset map and the tooth image mid-layer feature modulation map, inputting the tooth image mid-layer feature map into a spatial adaptive enhancement component based on a deformable convolutional layer to obtain the tooth image mid-layer spatial enhancement feature map, including: Based on the tooth image middle layer feature offset map, bilinear interpolation sampling is performed on the tooth image middle layer feature map to obtain a tooth image middle layer sampling feature map; The spatial enhancement feature map of the middle layer of the tooth image is obtained by calculating the multiplication of the sampling feature map of the middle layer of the tooth image and the feature modulation map of the middle layer of the tooth image according to the position point.

6. The method for intelligent analysis of dental images for oral surgery according to claim 1, characterized in that: The tooth image middle-layer spatial enhancement feature map, the tooth image deep-layer spatial enhancement feature map and the tooth edge information feature map are fused to obtain a multi-scale significant fusion feature map of the tooth image, including: Input the middle-layer spatial enhancement feature map of the tooth image, the deep-layer spatial enhancement feature map of the tooth image, and the tooth edge information feature map into a skip connection layer to obtain an initial tooth image multi-scale significant fusion down-sampling feature map; Performing complementary fusion optimization of visual semantic features on the initial multi-scale significant fusion down-sampled feature map of the tooth image to obtain a multi-scale significant fusion down-sampled feature map of the tooth image; The multi-scale significant fusion down-sampled feature map of the tooth image is up-sampled to obtain the multi-scale significant fusion feature map of the tooth image.

7. The method for intelligent analysis of dental images for oral surgery according to claim 6, characterized in that: The method of performing complementary fusion optimization of visual semantic features on the initial multi-scale significant fusion down-sampled feature map of the tooth image to obtain the multi-scale significant fusion down-sampled feature map of the tooth image comprises: Performing cavity convolution operations with different cavity rates on the middle-layer spatial enhancement feature map of the tooth image, the deep-layer spatial enhancement feature map of the tooth image, and the tooth edge information feature map, respectively, to obtain a cavity convolution feature map of the middle-layer spatial enhancement of the tooth image, a cavity convolution feature map of the deep-layer spatial enhancement of the tooth image, and a cavity convolution feature map of the tooth edge information; Taking the spatial enhancement cavity convolution feature map of the middle layer of the tooth image as a reference, calculating the cross-layer geometric feature response of the deep spatial enhancement cavity convolution feature map of the tooth image and the tooth edge information cavity convolution feature map to obtain a first tooth image cross-layer geometric feature response feature map and a second tooth image cross-layer geometric feature response feature map; Based on a shift window mechanism and a relative position encoding mechanism, the first tooth image cross-layer geometric feature response feature map and the second tooth image cross-layer geometric feature response feature map are fused and corrected to obtain a tooth image multi-scale correction feature map; The multi-scale corrected feature map of the tooth image is point-multiplied with the multi-scale significant fusion down-sampling feature map of the initial tooth image to obtain the multi-scale significant fusion down-sampling feature map of the tooth image.

8. The method for intelligent analysis of dental images for oral surgery according to claim 7, characterized in that: The dilation rates of the dilated convolution operation are 3, 5, and 1 respectively.

9. The method for intelligent analysis of dental images for oral surgery according to claim 1, characterized in that: The classification layer includes a point convolution layer and a Softmax classification unit.

10. A dental image intelligent analysis system for oral surgery, used to execute the dental image intelligent analysis method for oral surgery according to any one of claims 1 to 9, characterized in that: include: A tooth image data acquisition module, used to acquire tooth image data for oral surgery; A tooth image multi-scale feature extraction module, used for inputting the tooth image data into a backbone network to obtain a tooth image shallow feature map, a tooth image middle feature map and a tooth image deep feature map; A tooth image spatial enhancement module, used for inputting the tooth image middle layer feature map and the tooth image deep layer feature map into a spatial adaptive module to obtain a tooth image middle layer spatial enhancement feature map and a tooth image deep layer spatial enhancement feature map; A tooth image edge detection module, used for inputting the tooth image shallow feature map into the edge detection branch to obtain a tooth edge information feature map; A tooth image multi-scale feature fusion module, used to fuse the tooth image middle-layer spatial enhancement feature map, the tooth image deep-layer spatial enhancement feature map and the tooth edge information feature map to obtain a tooth image multi-scale significant fusion feature map; A tooth image classification processing module, used for inputting the multi-scale significant fusion feature map of the tooth image into the classification layer to obtain a tooth voxel probability map; The tooth image semantic segmentation module is used to convert the tooth voxel probability map into a tooth voxel discrete segmentation label map as the tooth image semantic segmentation result.

Citation Information

Patent Citations

  • Medical image processing method and device, electronic equipment and storage medium

    CN113506310A

  • Tooth segmentation and reconstruction method based on CBCT image and storage medium

    CN114757960A

  • Tooth instance segmentation based on collaborative learning

    CN117372356A

  • Fusion convolutional adaptive network skin lesion segmentation method

    CN118072024A

  • Medical image segmentation method and system based on natural language processing technology

    CN118429369A

Cited By

  • Tooth health preliminary screening method and device based on image recognition, equipment and medium

    CN121504870A