Oral and maxillofacial surgery image recognition and diagnosis method and system based on deep learning
By acquiring and processing oral and maxillofacial three-dimensional image data and using a multi-scale feature fusion network to identify lesion areas, the problem of insufficient accuracy in traditional imaging diagnosis methods is solved, and efficient and accurate imaging diagnosis is achieved.
Patent Information
- Application Number
- CN202510766267.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Traditional oral and maxillofacial imaging diagnostic methods rely on two-dimensional images, which make it difficult to fully display complex three-dimensional anatomical structures and lesion characteristics, resulting in insufficient diagnostic accuracy and reliability. Existing computer-aided diagnosis methods have difficulty fully extracting and fusing multi-scale features when processing three-dimensional imaging data, and are unable to accurately identify lesion areas.
By acquiring oral and maxillofacial three-dimensional image data, performing image feature extraction processing, calling the pre-trained multi-scale feature fusion network for feature fusion, generating a fusion feature map, identifying the location and morphology of the lesion area, and generating a diagnostic report.
It improves the accuracy and efficiency of oral and maxillofacial imaging diagnosis, enables precise identification and detailed diagnosis of lesion areas, provides detailed diagnostic basis, and supports rapid transmission and intuitive display.
Smart Images

Figure CN120280132B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to a deep learning-based oral and maxillofacial surgery image recognition and diagnosis method and system. Background Art
[0002] In the clinical diagnosis and treatment of oral and maxillofacial surgery, accurate image recognition and diagnosis are crucial for early disease detection, treatment plan development, and prognosis assessment. Traditional oral and maxillofacial imaging diagnostic methods rely primarily on physicians' experience in observing and analyzing two-dimensional images. However, two-dimensional images have information limitations and cannot fully display the complex three-dimensional anatomical structures and pathological characteristics of the oral and maxillofacial region. This can lead to inaccurate judgments about lesions, compromising the accuracy and reliability of the diagnosis.
[0003] With the continuous development of medical imaging technology, oral and maxillofacial three-dimensional imaging data has gradually become an important basis for clinical diagnosis. However, three-dimensional imaging data contains a large amount of continuously scanned slice image data, which is huge and complex. Manual processing and analysis are inefficient and cannot meet the needs of rapid clinical diagnosis. At the same time, existing computer-aided diagnosis methods often have difficulty in fully extracting and fusing multi-scale features in the images when processing three-dimensional imaging data, and are unable to accurately identify the location and morphology of the lesion area, resulting in the accuracy and reliability of the diagnostic results still needing to be improved. Therefore, the development of an efficient and accurate deep learning-based oral and maxillofacial surgical image recognition and diagnosis method is of great practical significance. Summary of the Invention
[0004] In view of this, an embodiment of the present invention provides an oral and maxillofacial surgery image recognition and diagnosis method and system based on deep learning.
[0005] In a first aspect, an embodiment of the present invention provides a method for oral and maxillofacial surgery image recognition and diagnosis based on deep learning, which is applied to an oral and maxillofacial surgery image recognition and diagnosis system based on deep learning, comprising:
[0006] Acquire a target patient's oral and maxillofacial three-dimensional image data set, wherein the three-dimensional image data set includes multiple sets of continuous scan layer image data;
[0007] Performing image feature extraction processing on the three-dimensional image data set to obtain a hierarchical image feature set of the continuously scanned slice image data; the hierarchical image feature set includes local anatomical structure features and global spatial distribution features;
[0008] Calling a pre-trained multi-scale feature fusion network to perform multi-scale feature fusion on the hierarchical image feature set to generate a fusion feature map;
[0009] Performing lesion area recognition processing based on the fused feature map to determine the location information and morphological description information of the oral and maxillofacial abnormality area of the target patient;
[0010] A diagnosis report is generated based on the location information and morphological description information, and the diagnosis report is transmitted to a medical terminal device for display.
[0011] In a second aspect, an embodiment of the present invention provides an oral and maxillofacial surgery image recognition and diagnosis system based on deep learning, comprising:
[0012] processor;
[0013] A memory, wherein a computer program is stored in the memory, and when the computer program is executed, the oral and maxillofacial surgery image recognition and diagnosis method based on deep learning described in the first aspect is implemented.
[0014] As described above, in an embodiment of the present invention, by obtaining a three-dimensional oral and maxillofacial image data set of a target patient and performing image feature extraction processing on the three-dimensional image data set, the local anatomical structure features and global spatial distribution features of the continuously scanned layer image data can be accurately obtained, and the pre-trained multi-scale feature fusion network is called to perform multi-scale feature fusion on the hierarchical image feature set, effectively integrating feature information of different scales to generate a more representative fusion feature map, which greatly improves the ability to identify the lesion area. The lesion area recognition processing is performed based on the fusion feature map, which can accurately determine the position information and morphological description information of the abnormal oral and maxillofacial area of the target patient, providing doctors with detailed and accurate diagnostic basis, and finally generating a diagnostic report based on the position information and morphological description information, and transmitting it to the medical terminal device for display, thereby realizing the rapid transmission and intuitive display of the diagnostic results, and effectively improving the accuracy and efficiency of oral and maxillofacial surgical imaging diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 A schematic flow chart of the steps of a deep learning-based oral and maxillofacial surgery image recognition and diagnosis method provided by an embodiment of the present invention;
[0016] Figure 2 The embodiment of the present invention provides a method for executing Figure 1 A structural schematic diagram of the oral and maxillofacial surgery image recognition and diagnosis system based on deep learning in the oral and maxillofacial surgery image recognition and diagnosis method based on deep learning. DETAILED DESCRIPTION
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field without making creative efforts based on the embodiments of the present invention are within the scope of protection of the present invention.
[0018] See also Figure 1 As shown:
[0019] Step S110: Acquire a three-dimensional image data set of the oral and maxillofacial region of a target patient, wherein the three-dimensional image data set includes multiple sets of continuous scan layer image data.
[0020] During the imaging recognition and diagnosis process in oral and maxillofacial surgery, the target patient's oral and maxillofacial three-dimensional image data set must be scanned using relevant medical imaging equipment, such as an oral and maxillofacial CT scanner. Before scanning, appropriate scanning parameters can be set based on the patient's specific situation. Scanning parameters include scanning range, layer spacing, etc. The scanning range must fully cover the key areas of the target patient's oral and maxillofacial area, such as the teeth, jaw, and surrounding soft tissues, to ensure that complete and effective image information is obtained. The layer spacing determines the distance between the image data of two adjacent scan layers, and its size affects the spatial resolution of the image.
[0021] During the scanning process, the oral and maxillofacial CT scanner rotates around the patient's head. Each time the scanner rotates to a specific angle, a set of slice image data is acquired. As the scanner continues to rotate, multiple sets of continuous slice image data are obtained. These slice image data are spatially continuous, displaying structural information of the target patient's oral and maxillofacial structures from different angles and levels. For example, scanning in the sagittal, coronal, and transverse planes can produce slice image data in different directions, which together form a three-dimensional image data set.
[0022] From a data structure perspective, each slice of image data can be viewed as a two-dimensional matrix, with each element representing the grayscale value at that location. A three-dimensional image data set is a three-dimensional matrix formed by stacking multiple such two-dimensional matrices in the order in which they were scanned. Each element in this three-dimensional matrix has a specific spatial coordinate, corresponding to a specific location within the oral and maxillofacial region.
[0023] Step S120: performing image feature extraction processing on the three-dimensional image data set to obtain a hierarchical image feature set of the continuously scanned slice image data; the hierarchical image feature set includes local anatomical structure features and global spatial distribution features.
[0024] After acquiring a 3D image data set, in order to extract valuable diagnostic information from it, it is necessary to perform image feature extraction processing on it. Image feature extraction processing is a complex and critical process that can transform the original image data into a feature set with specific meaning.
[0025] Step S121: performing spatial normalization processing on the continuously scanned slice image data to obtain a standardized three-dimensional image sequence.
[0026] Because oral and maxillofacial structures vary in spatial position, size, and orientation among different patients, spatial normalization of the continuously scanned image slices is necessary to facilitate subsequent unified processing and analysis. Spatial normalization primarily involves two core steps: spatial registration and resampling.
[0027] The purpose of spatial registration is to spatially align the oral and maxillofacial imaging data of different patients so that they have the same spatial coordinate system. Before performing spatial registration, a set of reference points needs to be determined. These reference points can be some landmark anatomical structures in the oral and maxillofacial region, such as the vertices of the teeth, the boundary points of the jaw, etc. For example, two sets of feature points can be extracted from the three-dimensional image data set, which are respectively recorded as feature point set A and feature point set B. Feature point set A comes from the image data of the target patient, and feature point set B comes from a pre-set standard template image data.
[0028] Next, we need to find a suitable transformation matrix T so that feature point set A, after being transformed by transformation matrix T, can coincide as closely as possible with feature point set B. This transformation matrix T can be determined by minimizing the distance error between feature point sets A and B. Specifically, the iterative closest point (ICP) algorithm can be used to solve for transformation matrix T. The basic idea of the ICP algorithm is to iterate continuously, finding the nearest neighbor of each point in feature point set A in feature point set B with each iteration, and then calculating the optimal transformation between these two sets of corresponding points until convergence conditions are met.
[0029] After obtaining the transformation matrix T, it is applied to the three-dimensional image data set of the target patient to transform the spatial coordinates of each voxel, thereby achieving spatial registration of the image data.
[0030] Resampling is performed after spatial registration to unify the resolution of the image data. Different scanning devices and scanning parameters may result in different resolutions of the image data, which can cause difficulties in subsequent processing. Therefore, it is necessary to resample the registered image data according to the preset resolution requirements.
[0031] Assume that the preset resolution is a specific voxel size, denoted as Δx, Δy, and Δz. During the resampling process, the value of each voxel is recalculated on the new spatial grid. The value of the new voxel can be calculated using linear interpolation. Specifically, for each voxel on the new spatial grid, its neighboring voxels in the original image data are found, and then the value of the voxel is calculated using a linear weighted method based on the values of these neighboring voxels and their distances from the voxel. After resampling, a standardized three-dimensional image sequence is obtained, whose resolution and spatial coordinate system are unified.
[0032] Step S122: calling a three-dimensional convolutional neural network to perform shallow feature extraction processing on the standardized three-dimensional image sequence to obtain a primary image feature set; the primary image feature set includes edge contour features and texture distribution features.
[0033] After obtaining the standardized 3D image sequence, a 3D convolutional neural network (3D CNN) is used to perform shallow feature extraction. A 3D convolutional neural network is a deep learning model specifically designed for processing 3D data, which can automatically learn the feature information in the image data.
[0034] Step S1221: performing convolution kernel sliding processing on the standardized three-dimensional image sequence to generate an initial convolution feature map, and performing maximum pooling processing on the initial convolution feature map to obtain a dimensionality reduction feature map.
[0035] When performing shallow feature extraction, the first step is to perform a sliding convolution kernel on the standardized 3D image sequence. The convolution kernel is a 3D matrix with a specific size and weight, denoted as the convolution kernel C. The size of the convolution kernel is usually selected based on the specific task and data characteristics, for example, 3×3×3, 5×5×5, etc.
[0036] The convolution kernel C is slid across the normalized 3D image sequence. Each time it slides, the kernel C performs a convolution operation with the corresponding region of the image sequence. The convolution operation multiplies the kernel C element-by-element with the corresponding region of the image sequence, then sums all the products to produce a convolution result. As the convolution kernel C continues to slide, a series of convolution results are generated, which are combined to form the initial convolution feature map.
[0037] Assume that the dimensions of a standardized 3D image sequence are H×W×D×C1 (H represents height, W represents width, D represents depth, and C1 represents the number of channels), and the dimensions of the convolution kernel C are h×w×d×C1×C2 (h, w, and d represent the kernel's dimensions in height, width, and depth, respectively, and C2 represents the number of output channels). During the convolution operation, parameters such as stride and padding need to be considered. The stride determines the distance the convolution kernel slides each time, while padding adds a set number of zero elements around the edges of the image sequence to ensure that the feature map output after the convolution operation meets the required size.
[0038] After generating the initial convolutional feature map, we need to perform max pooling on it to reduce the data dimension and improve computational efficiency. Max pooling is a downsampling operation that divides the initial convolutional feature map into fixed-size regions, denoted as pooling windows W. For each pooling window W, the maximum value within it is taken as the pooling result for that region.
[0039] Assume that the pooling window W is of size p×p×p, with a stride of s. During max pooling, the pooling window W is slid across the initial convolutional feature map, sliding s units at a time. The maximum value of each element within the pooling window is taken to obtain a reduced-dimensional feature value. As the pooling window continues to slide, a reduced-dimensional feature map is generated. While smaller than the initial convolutional feature map, the reduced-dimensional feature map retains the key feature information from the initial convolutional feature map.
[0040] Step S1222: inputting the reduced-dimensionality feature map into the residual connection module for feature compensation processing to obtain a compensated feature map; the feature compensation processing includes performing an element-by-element addition operation on the reduced-dimensionality feature map and the original feature map transferred through the jump connection to generate a compensated feature map with detail-preserving characteristics.
[0041] After obtaining the reduced-dimensional feature map, to avoid losing important details during feature extraction, it needs to be fed into the residual connection module for feature compensation. The residual connection module is a key component of the 3D convolutional neural network. It can pass feature information from the original standardized 3D image sequence to subsequent processing stages through skip connections.
[0042] Assume the reduced-dimensionality feature map is F1, with dimensions H1×W1×D1×C2. When performing a skip connection, the original normalized 3D image sequence undergoes a series of convolution and pooling operations until its size matches that of the reduced-dimensionality feature map F1, denoted as F0. Then, the reduced-dimensionality feature map F1 is element-wise added to the original feature map F0, transferred via the skip connection. Specifically, the elements at corresponding positions in F1 and F0 are added together to produce a new element, which is then combined to form the compensated feature map F2.
[0043] The formula for calculating the compensated feature map F2 is F2 = F1 + F0. This element-by-element addition restores some of the details lost during the dimensionality reduction process, resulting in a more detailed compensated feature map. This allows the model to utilize this retained detail in subsequent processing to more accurately identify and analyze oral and maxillofacial structural features.
[0044] Step S1223: performing channel attention weight allocation processing on the compensated feature map to generate a weighted feature map; the channel attention weight allocation processing includes performing a global average pooling operation on the channel dimension of the compensated feature map to generate a channel description vector; inputting the channel description vector into the fully connected layer to generate channel attention weights, and recalibrating each channel of the compensated feature map based on the channel attention weights to generate a weighted feature map.
[0045] After obtaining the compensated feature map, in order to highlight the importance of different channels in the compensated feature map, it is necessary to perform channel attention weight allocation processing. The channel attention weight allocation process mainly includes two steps: global average pooling and fully connected layer mapping.
[0046] First, a global average pooling operation is performed on the channel dimension of the compensated feature map. This operation averages the compensated feature map across the spatial dimensions to produce a channel description vector. Assuming the dimensions of the compensated feature map are H2×W2×D2×C2, the global average pooling operation sums all elements in the height, width, and depth directions of each channel, then divides by the total number of elements to obtain the average value for that channel. The average values of all channels are combined to form a channel description vector of length C2.
[0047] The channel description vector is then input into the fully connected layer for mapping. The fully connected layer is a simple linear transformation layer that outputs a channel attention weight vector of length C2 based on the input of the channel description vector. The channel attention weight vector represents the importance of each channel in the compensation feature map.
[0048] Specifically, the calculation formula for the fully connected layer is: W = f(W1*V+b1), where V is the channel description vector, W1 is the weight matrix of the fully connected layer, b1 is the bias vector, and f is the activation function, such as the Sigmoid function. Through the activation function, the output of the fully connected layer is mapped to the interval [0, 1], resulting in the channel attention weight vector W.
[0049] Finally, each channel of the compensated feature map is recalibrated based on the channel attention weight vector W. The channel attention weight vector W is element-wise multiplied by each channel of the compensated feature map to obtain a weighted feature map. Assuming the compensated feature map is F2 and the weighted feature map is F3, the formula for calculating F3 is F3 = F2 * W. This method enhances the feature information of important channels and suppresses the feature information of unimportant channels, allowing the weighted feature map to better highlight key features.
[0050] Step S1224: extract edge contour features and texture distribution features from the weighted feature map to form a primary image feature set; the edge contour features are used to perform gradient feature extraction on the weighted feature map through a learnable edge convolution kernel, and the texture distribution features are used to perform adaptive texture pattern extraction on the weighted feature map through a multi-scale convolution kernel group.
[0051] After obtaining the weighted feature map, in order to obtain the primary image feature set, it is necessary to extract edge contour features and texture distribution features from the weighted feature map.
[0052] Edge contour features are extracted by applying gradient feature extraction to weighted feature maps using a learnable edge convolution kernel. A learnable edge convolution kernel is a convolution kernel with a specific structure and weights that automatically learns edge information in images. The weights of the learnable edge convolution kernel are continuously updated during training to adapt to different image data and task requirements.
[0053] A learnable edge convolution kernel is applied to the weighted feature map, and its gradient information is obtained through convolution. Gradient information reflects the rate of change of grayscale values in the image. Areas with large grayscale changes typically correspond to edge contours. Specifically, the learnable edge convolution kernel slides across the weighted feature map. Each time it slides, it performs a convolution operation with the corresponding area of the weighted feature map to obtain a gradient value. As the convolution kernel continues to slide, a gradient feature map is eventually generated, which represents the edge contour characteristics of the image.
[0054] The extraction of texture distribution features is achieved by adaptively extracting texture patterns from weighted feature maps using a multi-scale convolution kernel group. The multi-scale convolution kernel group contains multiple convolution kernels of different sizes and scales, which can capture texture information at different scales.
[0055] Assume that the multi-scale convolution kernel group contains three convolution kernels of different sizes, denoted as convolution kernel K1, convolution kernel K2, and convolution kernel K3, with their sizes increasing in sequence. These three convolution kernels are applied to the weighted feature map, and the convolution operation yields three texture feature maps at different scales, denoted as texture feature map F4, texture feature map F5, and texture feature map F6.
[0056] Convolution kernels of different sizes can capture texture information of varying coarseness and complexity. Smaller kernels can capture detailed texture information, while larger kernels can capture macroscopic texture patterns. Combining these three texture feature maps forms a texture distribution feature.
[0057] Finally, the extracted edge contour features and texture distribution features are combined to form a primary image feature set, which contains important feature information such as edge contour and texture distribution in oral and maxillofacial images.
[0058] Step S123: performing cross-layer feature enhancement processing on the primary image feature set to obtain an enhanced image feature set; the enhanced image feature set includes local anatomical structure features after spatial resolution enhancement.
[0059] After obtaining the primary image feature set, in order to further enhance the feature information and improve the feature expression ability, it is necessary to perform cross-layer feature enhancement on the primary image feature set. Cross-layer feature enhancement can fuse feature information at different levels to obtain local anatomical structure features with enhanced spatial resolution.
[0060] Step S1231: performing feature stitching processing on the edge contour features and texture distribution features in the primary image feature set to generate a stitching feature map.
[0061] In order to fully utilize the edge contour features and texture distribution features in the primary image feature set, they first need to be spliced. Feature splicing is to splice the edge contour features and texture distribution features in the channel dimension.
[0062] Assume that the size of the edge contour feature is H3×W3×D3×C3, and the size of the texture distribution feature is H3×W3×D3×C4. When performing feature splicing, the edge contour feature and the texture distribution feature are concatenated in the channel dimension, and the size of the spliced feature map obtained is H3×W3×D3×(C3 + C4). Through feature splicing processing, different types of feature information can be integrated together, providing a richer feature representation for subsequent processing. In this way, the spliced feature map contains information on both the edge contour and the texture distribution, and can more comprehensively describe the local anatomical structure of the oral and maxillofacial region.
[0063] Step S1232: Perform dilated convolution processing on the spliced feature map to generate a multi-scale receptive field feature map; the dilated convolution processing uses convolution kernels with different dilation rates to act on the spliced feature map in parallel, generating feature sub-maps with different receptive field ranges; the feature sub-maps are merged in the channel dimension to generate a multi-scale receptive field feature map, and feature pyramid fusion processing is performed on the multi-scale receptive field feature map to obtain a multi-scale fusion feature map; the feature pyramid fusion processing includes upsampling or downsampling the feature sub-maps with different scales to a unified resolution and performing an element-wise addition operation.
[0064] After obtaining the spliced feature map, in order to obtain feature information at different scales, dilated convolution processing needs to be performed on it. The dilated convolution processing uses convolution kernels with different dilation rates to act on the spliced feature map in parallel. The dilation rate represents the distance between elements in the convolution kernel, and different dilation rates can make the convolution kernel have different receptive field ranges.
[0065] Assume that three convolution kernels with different dilation rates are used, denoted as convolution kernel K4, convolution kernel K5, and convolution kernel K6, and their dilation rates are r1, r2, and r3 (r1 < r2 < r3) respectively. Apply these three convolution kernels to the spliced feature map respectively, and through convolution operations, three feature sub-maps with different receptive field ranges are obtained, denoted as feature sub-map F7, feature sub-map F8, and feature sub-map F9.
[0066] Convolution kernel K4 has a smaller dilation rate r1, and its receptive field range is relatively small, capable of capturing the detailed information in the spliced feature map. Convolution kernel K5 has a medium dilation rate r2, and its receptive field range is moderate, capable of capturing some local structural information. Convolution kernel K6 has a larger dilation rate r3, and its receptive field range is large, capable of capturing more macroscopic feature information.
[0067] Then, these three feature submaps are merged in the channel dimension to generate a multi-scale receptive field feature map. Assuming that the size of feature submap F7 is H4×W4×D4×C5, the size of feature submap F8 is H4×W4×D4×C6, and the size of feature submap F9 is H4×W4×D4×C7, the size of the merged multi-scale receptive field feature map is H4×W4×D4×(C5+C6+C7).
[0068] Next, feature pyramid fusion is performed on the multi-scale receptive field feature maps. The purpose of feature pyramid fusion is to fuse feature submaps of different scales to achieve a uniform resolution. First, feature submaps of different scales need to be upsampled or downsampled. For smaller feature submaps, methods such as bilinear interpolation can be used to upsample them to the same size as other feature submaps. For larger feature submaps, methods such as average pooling can be used to downsample them to the same size as other feature submaps.
[0069] Assume that after upsampling or downsampling, feature subgraphs F7, F8, and F9 all have the same size, H5×W5×D5×C8. These three feature subgraphs are then element-wise added together to produce a multi-scale fused feature graph F10. This element-wise addition operation sums the elements at corresponding positions in the three feature subgraphs to obtain the value of the element at that position in the multi-scale fused feature graph. This feature pyramid fusion process effectively fuses feature information at different scales. The multi-scale fused feature graph comprehensively reflects feature information at different scales, providing a more comprehensive feature representation for subsequent processing.
[0070] Step S1233: performing nonlinear activation processing on the multi-scale fusion feature map to generate an activation feature map; the nonlinear activation processing uses an activation function with a gating mechanism to dynamically activate each channel of the multi-scale fusion feature map.
[0071] After obtaining the multi-scale fusion feature map, in order to introduce nonlinear factors and improve the expressiveness of the model, it is necessary to perform nonlinear activation processing on it. Here, an activation function with a gating mechanism is used to dynamically activate each channel of the multi-scale fusion feature map.
[0072] An activation function with a gating mechanism can dynamically adjust the activation level of each channel based on the input feature information. Assume that the multi-scale fusion feature map is F10, with dimensions of H5×W5×D5×C8. An activation function with a gating mechanism calculates a gating value for each channel, indicating whether the channel should be activated and to what extent.
[0073] Specifically, an activation function with a gating mechanism first performs a linear transformation on each channel of the multi-scale fused feature map, generating an intermediate result. This intermediate result is then mapped to the [0, 1] interval using a nonlinear function, such as the sigmoid function, to generate a gating value. Finally, the gating value is multiplied by the original value of the channel to obtain the activated channel value.
[0074] Assume that for the i-th channel of the multi-scale fusion feature map F10, its original value is F10(i), the intermediate result obtained after linear transformation is M(i), the gated value is G(i), and the channel value after activation processing is A(i). The calculation process is as follows: First, M(i) is obtained through linear transformation. The linear transformation can be expressed as M(i) = W2 * F10(i) + b2, where W2 is the weight matrix of the linear transformation and b2 is the bias vector. Then, the gated value G(i) = Sigmoid(M(i)) is calculated using the Sigmoid function. Finally, the activated channel value A(i) = G(i) * F10(i). This processing is performed on all channels of the multi-scale fusion feature map F10 to obtain the activated feature map F11. Through this nonlinear activation process, the model can learn more complex feature representations, enhancing its ability to express oral and maxillofacial image features.
[0075] Step S1234: performing spatial attention weight allocation processing on the activation feature map to obtain an enhanced image feature set; the spatial attention weight allocation processing includes performing a dual-path operation of maximum pooling and average pooling on the spatial dimension of the activation feature map to generate a dual-path pooled feature map; performing channel splicing on the dual-path pooled feature map and then generating a spatial attention weight map through a convolution layer, and performing spatial dimension weighting on the activation feature map based on the spatial attention weight map to generate an enhanced image feature set.
[0076] After obtaining the activation feature map, in order to highlight the importance of different spatial locations in the activation feature map, it is necessary to perform spatial attention weight allocation processing. The spatial attention weight allocation process mainly includes three steps: dual-path pooling, channel splicing, and convolution to generate weight maps.
[0077] First, a dual-path operation of maximum pooling and average pooling is performed on the spatial dimensions of the activation feature map. The maximum pooling operation divides the spatial dimensions of the activation feature map into fixed-size regions, denoted as pooling windows W1. For each pooling window W1, the maximum value within it is taken as the pooling result for that region. The average pooling operation takes the average value within that region as the pooling result.
[0078] Assume that the size of the activation feature map F11 is H5×W5×D5×C8, the size of the pooling window W1 is p1×p1×p1, and the step size is s1. When performing the maximum pooling operation, the pooling window W1 is slid on the activation feature map F11, each sliding s1 units, and the maximum value of the elements in each pooling window is taken to obtain the maximum pooling feature map F12. When performing the average pooling operation, the pooling window W1 is also slid on the activation feature map F11, each sliding s1 units, and the average of the elements in each pooling window is calculated to obtain the average pooling feature map F13.
[0079] Then, the maximum pooling feature map F12 and the average pooling feature map F13 are spliced in the channel dimension to obtain the dual-path pooling feature map F14. Assuming that the size of the maximum pooling feature map F12 is H6×W6×D6×C8 and the size of the average pooling feature map F13 is H6×W6×D6×C8, the size of the spliced dual-path pooling feature map F14 is H6×W6×D6×(2*C8).
[0080] Next, the dual-pooled feature map F14 is processed through a convolutional layer to generate a spatial attention weight map F15. The convolutional layer performs a convolution operation on the dual-pooled feature map F14 to learn the importance of different spatial locations. Assume that the convolution kernel of the convolutional layer is K7, and its size is h1×w1×d1×(2*C8)×1. Through the convolution operation, the dual-pooled feature map F14 is converted into a spatial attention weight map F15 of size H6×W6×D6×1.
[0081] Finally, the activation feature map F11 is spatially weighted based on the spatial attention weight map F15. The spatial attention weight map F15 is element-wise multiplied with the activation feature map F11 to obtain the enhanced image feature set F16. For each element in the activation feature map F11, it is multiplied with the element at the corresponding position in the spatial attention weight map F15 to obtain the element value at that position in the enhanced image feature set F16. This spatial attention weight distribution process enhances the feature information of important spatial locations in the activation feature map and suppresses the feature information of unimportant spatial locations, allowing the enhanced image feature set to better highlight key local anatomical features.
[0082] Step S124: performing global pooling processing on the enhanced image feature set to obtain global spatial distribution features.
[0083] After obtaining the enhanced image feature set, global pooling is performed on it to obtain the global information of the image data. Global pooling can compress the enhanced image feature set in the spatial dimension to obtain a global feature that can represent the entire image data.
[0084] Step S1241: Perform channel - dimension average pooling on the local anatomical structure features in the enhanced image feature set to generate a channel - statistical feature vector.
[0085] First, perform channel - dimension average pooling on the local anatomical structure features in the enhanced image feature set. Assume the size of the enhanced image feature set F16 is H7×W7×D7×C9. The channel - dimension average pooling process sums all the elements in each channel in the height, width, and depth directions, and then divides by the total number of elements to obtain the average value of that channel.
[0086] Specifically, for the j - th channel of the enhanced image feature set F16, the sum of all its elements is Sum(j), and the total number of elements is H7*W7*D7. Then the average value of this channel is Avg(j)=Sum(j) / (H7*W7*D7). Combining the average values of all channels together forms a channel - statistical feature vector V1 with a length of C9. The channel - statistical feature vector V1 reflects the average feature information of each channel in the enhanced image feature set, providing a concise feature representation for subsequent processing.
[0087] Step S1242: Perform a fully - connected layer mapping on the channel - statistical feature vector to generate a dimensionality - reduced feature vector.
[0088] After obtaining the channel - statistical feature vector, in order to further reduce the dimensionality of the data, a fully - connected layer mapping needs to be performed on it. The fully - connected layer is a simple linear transformation layer that can output a dimensionality - reduced feature vector based on the input of the channel - statistical feature vector.
[0089] Assume the length of the channel - statistical feature vector V1 is C9, the weight matrix of the fully - connected layer is W3 with a size of C9×C10 (C10 < C9), and the bias vector is b3. The mapping calculation of the fully - connected layer can be expressed as V2 = W3*V1 + b3, where V2 is the dimensionality - reduced feature vector. Through the fully - connected layer mapping process, the information in the channel - statistical feature vector is compressed and integrated to obtain a dimensionality - lower dimensionality - reduced feature vector V2. The dimensionality - reduced feature vector V2 retains the main information in the channel - statistical feature vector, reduces data redundancy, and improves computational efficiency.
[0090] Step S1243: Map the dimensionality - reduced feature vector through a fully - connected layer to be consistent with the channel dimension of the enhanced image feature set, and perform feature broadcasting processing to generate a spatial weight distribution map; the feature broadcasting processing includes copying and expanding the mapped feature vector to the same spatial size as the enhanced image feature set.
[0091] After obtaining the reduced-dimensionality feature vector, it needs to be mapped through a fully connected layer to match the channel dimensions of the enhanced image feature set. Assume that the length of the reduced-dimensionality feature vector V2 is C10, and the number of channels in the enhanced image feature set F16 is C9. The weight matrix of the fully connected layer is W4, with dimensions C10 × C9, and the bias vector is b4. Through the mapping calculation of the fully connected layer, the mapped feature vector V3 = W4 * V2 + b4 is obtained, with a length of C9.
[0092] Feature broadcasting is then performed. Feature broadcasting involves replicating and expanding the mapped feature vector V3 to the same spatial dimensions as the enhanced image feature set F16. Assuming the dimensions of the enhanced image feature set F16 are H7×W7×D7×C9, each element in the mapped feature vector V3 is replicated H7*W7*D7 times, filling a space of H7×W7×D7×C9, resulting in a spatial weight distribution map F17. Spatial weight distribution map F17 assigns a weight value to each position in the enhanced image feature set F16, reflecting its importance in global space.
[0093] Step S1244: performing channel dimension weighting and spatial dimension aggregation processing on the enhanced image feature set based on the spatial weight distribution map to obtain global spatial distribution features.
[0094] After obtaining the spatial weight distribution map, the enhanced image feature set is then weighted in the channel dimension and aggregated in the spatial dimension. First, channel weighting is performed by element-by-element multiplication of the spatial weight distribution map F17 and the enhanced image feature set F16. For each element in the enhanced image feature set F16, it is multiplied with the element at the corresponding position in the spatial weight distribution map F17 to obtain the weighted enhanced image feature set F18.
[0095] Then spatial dimension aggregation processing is performed. Spatial dimension aggregation processing can be performed by average pooling or summation. Taking average pooling as an example, average pooling operations are performed on the height, width and depth directions of the weighted enhanced image feature set F18. Assuming that the size of the weighted enhanced image feature set F18 is H7×W7×D7×C9, all elements of each channel in the height, width and depth directions are summed, and then divided by the total number of elements to obtain the average value of the channel. The average values of all channels are combined to form the global spatial distribution feature V4. The global spatial distribution feature V4 is a vector of length C9, which integrates the information of the enhanced image feature set in the global space, and provides an important global feature representation for subsequent feature fusion and lesion area identification.
[0096] Step S125: performing feature alignment processing on the local anatomical structure features and the global spatial distribution features to form a hierarchical image feature set; the feature alignment processing includes performing a spatial coordinate mapping operation on the local anatomical structure features to generate a spatial alignment feature vector corresponding to the global spatial distribution features; performing a feature broadcasting operation on the global spatial distribution features to expand their spatial dimensions to be consistent with the spatial alignment feature vector, and adjusting the channel dimension matching; performing channel dimension splicing on the expanded global spatial distribution features and the spatial alignment feature vector to generate a hierarchical image feature set with a unified spatial dimension.
[0097] After obtaining the local anatomical structure features (i.e., enhanced image feature set F16) and the global spatial distribution features V4, in order to effectively fuse them together, feature alignment processing is required to form a hierarchical image feature set.
[0098] First, a spatial coordinate mapping operation is performed on the local anatomical structure feature. The spatial coordinate mapping operation maps the spatial coordinates of each element in the local anatomical structure feature to a unified coordinate system, so that its spatial coordinates are consistent with the spatial coordinates corresponding to the global spatial distribution feature. Assume that the spatial coordinate system of the local anatomical structure feature F16 is coordinate system A, and the spatial coordinate system corresponding to the global spatial distribution feature V4 is coordinate system B. Through a coordinate transformation matrix T1, the spatial coordinates of each element in the local anatomical structure feature F16 are converted from coordinate system A to coordinate system B to obtain a spatial alignment feature vector F19. The spatial coordinates of the spatial alignment feature vector F19 are consistent with the spatial coordinates corresponding to the global spatial distribution feature V4.
[0099] Then, a feature broadcast operation is performed on the global spatial distribution features. The feature broadcast operation expands the spatial dimensions of the global spatial distribution features V4 to be consistent with the spatial alignment feature vector F19 and adjusts the channel dimensions to match. Assume that the global spatial distribution features V4 is a vector of length C9 and the dimensions of the spatial alignment feature vector F19 are H7×W7×D7×C9. Each element in the global spatial distribution features V4 is replicated H7*W7*D7 times and filled into a space of size H7×W7×D7×C9, resulting in the expanded global spatial distribution features F20.
[0100] Finally, the expanded global spatial distribution features F20 are concatenated with the spatially aligned feature vector F19 in the channel dimension. Channel-dimensional concatenation involves concatenating the expanded global spatial distribution features F20 and the spatially aligned feature vector F19 in the channel dimension. Assuming the dimensions of the expanded global spatial distribution features F20 are H7×W7×D7×C9, and the dimensions of the spatially aligned feature vector F19 are H7×W7×D7×C9, the dimensions of the concatenated hierarchical image feature set F21 are H7×W7×D7×(2*C9). Through this channel-dimensional concatenation, local anatomical structural features and global spatial distribution features are effectively fused together to form a hierarchical image feature set F21 with a unified spatial dimension. The hierarchical image feature set F21 integrates local and global feature information, providing a more comprehensive and richer feature representation for subsequent multi-scale feature fusion and lesion region identification.
[0101] Step S130: calling a pre-trained multi-scale feature fusion network to perform multi-scale feature fusion on the hierarchical image feature set to generate a fusion feature map.
[0102] After obtaining the hierarchical image feature set, a pre-trained multi-scale feature fusion network is used to further mine and integrate the feature information, generating a fused feature map. The multi-scale feature fusion network analyzes and fuses the hierarchical image feature set from different scales and angles, enabling the fused feature map to more comprehensively and accurately reflect the structure and pathology of the oral and maxillofacial region.
[0103] Step S131: performing feature splicing processing on the local anatomical structure feature and the global spatial distribution feature to generate an initial splicing feature.
[0104] To fully utilize the local anatomical features and global spatial distribution features in the hierarchical image feature set, we first perform feature concatenation. Assume that the local anatomical feature is F16 (size H7 × W7 × D7 × C9), and the global spatial distribution feature is F20 (size H7 × W7 × D7 × C9) after feature broadcasting. During feature concatenation, the local anatomical feature F16 and the global spatial distribution feature F20 are concatenated along the channel dimension to obtain the initial concatenated feature F22, size H7 × W7 × D7 × (2 * C9). This feature concatenation integrates local and global feature information, providing richer feature input for subsequent multi-scale feature fusion.
[0105] Step S132: performing grouped axial self-attention mechanism processing on the initial spliced features to generate a self-attention weight distribution map; the grouped axial self-attention mechanism processing includes decomposing the features into axial sub-regions along the spatial dimension, respectively calculating the correlation matrix between the spatial positions in each axial sub-region, and normalizing the correlation matrix to generate a local self-attention weight distribution map.
[0106] After obtaining the initial concatenated features, a grouped axial self-attention mechanism is applied to them to mine the association information between different spatial locations in the features. This grouped axial self-attention mechanism mainly includes three steps: feature decomposition, correlation matrix calculation, and normalization.
[0107] First, decompose the initial stitching feature F22 into axial sub-regions along the spatial dimension. Assuming the size of the initial stitching feature F22 is H7×W7×D7×(2*C9), it can be divided in the height, width, and depth directions to obtain multiple axial sub-regions. For example, it is divided into m intervals in the height direction, n intervals in the width direction, and p intervals in the depth direction, thus obtaining m*n*p axial sub-regions. The size of each axial sub-region is (H7 / m)×(W7 / n)×(D7 / p)×(2*C9).
[0108] Then, the correlation matrix between the spatial positions in each axial sub-region is calculated separately. For each axial sub-region, it is expanded into a two-dimensional matrix, the number of rows of the matrix is the number of spatial positions in the axial sub-region, and the number of columns is the number of feature channels. Then, by calculating the similarity between different rows in the two-dimensional matrix, the correlation matrix is obtained. The similarity can be calculated using methods such as dot product and cosine similarity. Assume that the two-dimensional matrix after expansion of a certain axial sub-region is M, the number of rows is r, and the number of columns is (2*C9). The element R(i, j) of the correlation matrix R represents the similarity between the i-th row and the j-th row.
[0109] Finally, the correlation matrix is normalized to generate a local self-attention weight distribution map. The purpose of normalization is to map the elements in the correlation matrix to the interval [0, 1] so that each element represents the importance weight of the position in the axial sub-region. The Softmax function can be used to normalize the correlation matrix. For each row of the correlation matrix R, the normalized value of each element in the row is calculated by the Softmax function to obtain the local self-attention weight distribution map. By combining the local self-attention weight distribution maps of all axial sub-regions, the self-attention weight distribution map F23 of the entire initial splicing feature is obtained. The size of the self-attention weight distribution map F23 is the same as that of the initial splicing feature F22, which reflects the importance association information between different spatial positions in the initial splicing feature.
[0110] Step S133: Dynamically weighting the local anatomical structure features based on the self-attention weight distribution map to generate weighted local features; the dynamic weighting processing includes performing an element-by-element multiplication operation on the self-attention weight distribution map and the local anatomical structure features.
[0111] After obtaining the self-attention weight distribution map, we dynamically weight the local anatomical features based on it. The local anatomical features were obtained in the previous step and are denoted as F16. The self-attention weight distribution map is denoted as F23. Dynamic weighting is performed by element-wise multiplication of the self-attention weight distribution map F23 and the local anatomical features F16.
[0112] For each element in the local anatomical structure feature F16, multiply it by the element at the corresponding position in the self-attention weight distribution map F23 to obtain a weighted element value. For example, assuming that the element value at position (x, y, z, c) in the local anatomical structure feature F16 is F16(x, y, z, c), and the element value at position (x, y, z, c) in the self-attention weight distribution map F23 is F23(x, y, z, c), then the weighted element value is F16(x, y, z, c)*F23(x, y, z, c). This process is performed on the elements at all positions in the local anatomical structure feature F16 to obtain the weighted local feature F24.
[0113] This dynamic weighting process enhances the feature information of important spatial locations in the local anatomical structure features and suppresses the feature information of unimportant spatial locations. Because the self-attention weight distribution map reflects the importance of the association between different spatial locations, the weighted local features can better focus on key local anatomical structure features.
[0114] Step S134: performing feature fusion processing on the weighted local features and the global spatial distribution features to generate intermediate fusion features.
[0115] After obtaining the weighted local features, they are fused with the global spatial distribution features. After feature broadcasting, the global spatial distribution features become F20, and the weighted local features become F24. Feature fusion can be performed in a variety of ways; here, a combination of concatenation and weighted addition is used.
[0116] First, the weighted local feature F24 and the global spatial distribution feature F20 are concatenated in the channel dimension to obtain the concatenated feature F25. Assuming that the size of the weighted local feature F24 is H7×W7×D7×C9 and the size of the global spatial distribution feature F20 is H7×W7×D7×C9, the size of the concatenated feature F25 is H7×W7×D7×(2*C9).
[0117] Then, perform a weighted summation operation on the concatenated feature F25. To balance the contributions of the weighted local features and the global spatial distribution features, weights w1 and w2 are assigned to them respectively, and w1 + w2 = 1. For each element in each channel of the concatenated feature F25, multiply it by the corresponding weight according to whether it belongs to the weighted local feature or the global spatial distribution feature part, and then add the results.
[0118] For example, for the elements in the first C9 channels (corresponding to the weighted local feature F24) of the concatenated feature F25, multiply by the weight w1; for the elements in the last C9 channels (corresponding to the global spatial distribution feature F20), multiply by the weight w2. Assume that the element value at position (x, y, z, c) in the concatenated feature F25 is F25(x, y, z, c). If c < C9, the weighted element value is w1 * F25(x, y, z, c); if c >= C9, the weighted element value is w2 * F25(x, y, z, c). Process all the elements in this way to obtain the intermediate fusion feature F26.
[0119] Through this feature fusion process, the information of the weighted local features and the global spatial distribution features is effectively integrated. The intermediate fusion feature F26 contains both the key local anatomical structure information and the global spatial distribution information, providing a more comprehensive feature representation for subsequent cross-channel feature interaction processing.
[0120] Step S135: Perform cross-channel feature interaction processing on the intermediate fusion feature to generate a fused feature map; the cross-channel feature interaction processing includes decomposing the intermediate fusion feature into multiple channel groups, performing independent convolution operations on each channel group, and then rearranging the channel order to generate a fused feature map with channel interaction characteristics.
[0121] After obtaining the intermediate fusion feature, to further promote information interaction between different channels, perform cross-channel feature interaction processing on it to generate a fused feature map. The cross-channel feature interaction processing mainly includes three steps: channel grouping, independent convolution, and channel rearrangement.
[0122] Step S1351: Divide the intermediate fusion feature into multiple sub-feature groups along the channel dimension; where each sub-feature group contains a predetermined number of consecutive channels.
[0123] The size of the intermediate fusion feature F26 is H7×W7×D7×(2*C9). Divide it along the channel dimension, and each sub-feature group contains a predetermined number of consecutive channels. Assume that the channel dimension of the intermediate fusion feature F26 is evenly divided into k sub-feature groups, and each sub-feature group contains (2*C9) / k consecutive channels.
[0124] For example, the first sub-feature group contains channels 1 to (2*C9) / k, the second sub-feature group contains channels (2*C9) / k+1 to 2*(2*C9) / k, and so on. This results in k sub-feature groups, denoted as F26_1, F26_2, ..., F26_k. Each sub-feature group has a size of H7×W7×D7×((2*C9) / k). Channel segmentation groups the channel information of the intermediate fusion features.
[0125] Step S1352: performing group convolution processing on each of the sub-feature groups to generate a group feature map; the group convolution processing uses an independent convolution kernel to act on each sub-feature group separately.
[0126] After obtaining the sub-feature groups, group convolution is performed on each sub-feature group. Group convolution uses a separate convolution kernel to apply to each sub-feature group. Assume that for each sub-feature group F26_i (i=1, 2, ..., k), a separate convolution kernel K_i is used for the convolution operation.
[0127] The size of the convolution kernel K_i is h2×w2×d2×((2*C9) / k)×C11, where h2, w2, and d2 represent the kernel's height, width, and depth, respectively, and C11 represents the number of output channels. The convolution kernel K_i is slid across the sub-feature group F26_i. Each time it slides, the kernel convolves with the corresponding region of the sub-feature group F26_i, producing a convolution result. As the convolution kernel K_i continues to slide, a series of convolution results are generated, which are combined to form the grouped feature map F27_i.
[0128] The size of the grouped feature map F27_i is H8×W8×D8×C11, where H8, W8, and D8 are determined by the convolution operation's stride, padding, and other parameters. Through grouped convolution, each sub-feature group undergoes independent feature extraction. Information between different sub-feature groups does not interact at this stage, but channel information within each sub-feature group is further mined and integrated.
[0129] Step S1353: performing channel shuffling processing on the grouped feature graph to generate a shuffled feature graph; the channel shuffling processing rearranges the channels in each sub-feature group according to a predetermined rule to promote cross-group information interaction.
[0130] After obtaining the group feature map, in order to promote information interaction between different sub-feature groups, the group feature map is subjected to channel shuffling processing. Channel shuffling processing is to rearrange the channels in each sub-feature group according to a predetermined rule.
[0131] First, k grouped feature maps F27_1, F27_2, ..., F27_k are concatenated in the channel dimension to obtain a concatenated feature map F28 with a size of H8×W8×D8×(k*C11). Then, the channels of the concatenated feature map F28 are rearranged according to a predetermined rule.
[0132] For example, the channels are arranged so that every k channels are extracted and reassembled. Assuming that the channels of the spliced feature map F28 are numbered 1 to k*C11, channels 1, k+1, 2*k+1, ... are used as a new group of channels, channels 2, k+2, 2*k+2, ... are used as another group of channels, and so on. After this rearrangement, a shuffled feature map F29 is obtained.
[0133] Channel shuffling breaks the channel boundaries of the original sub-feature groups, allowing the channel information between different sub-feature groups to interact. In this way, the information between the sub-feature groups that were originally processed independently is integrated.
[0134] Step S1354: performing point-by-point convolution processing on the shuffled feature map to generate cross-channel interaction features; the point-by-point convolution processing uses a 1×1 convolution kernel to adjust the channel dimension.
[0135] After obtaining the shuffled feature map, we perform point-by-point convolution on it to generate cross-channel interaction features. This point-by-point convolution is performed using a 1×1 convolution kernel. The 1×1 convolution kernel adjusts the channel dimension and promotes information interaction between different channels.
[0136] Assume that the size of the shuffled feature map F29 is H8×W8×D8×(k*C11). Use a 1×1 convolution kernel K8 with a size of 1×1×1×(k*C11)×C12 for point-by-point convolution. Slide the convolution kernel K8 over the shuffled feature map F29. Each time it slides, the convolution kernel K8 performs a convolution operation with the corresponding area of the shuffled feature map F29.
[0137] Since the convolution kernel K8 has a size of 1×1×1, it only operates on the channel dimension of the shuffled feature map F29. Through the convolution operation, the channel dimension of the shuffled feature map F29 is adjusted from (k*C11) to C12, and the cross-channel interaction feature F30 is obtained, whose size is H8×W8×D8×C12.
[0138] Point-by-point convolution allows for linear combination of information across channels, further facilitating cross-channel information interaction. By adjusting channel dimensions, feature redundancy is reduced while retaining important information across channels.
[0139] Step S1355: performing residual connection processing on the cross-channel interaction features and the intermediate fusion features to generate a fusion feature map; the residual connection processing adds the cross-channel interaction features and the original intermediate fusion features element by element.
[0140] After obtaining the cross-channel interaction features, they are residually connected with the intermediate fusion features to generate a fused feature map. The purpose of the residual connection process is to integrate the information of the cross-channel interaction features and the intermediate fusion features while avoiding the loss of the original intermediate fusion feature information during the feature fusion process.
[0141] First, the intermediate fusion feature F26 needs to be processed through a series of operations to make its size the same as the cross-channel interaction feature F30. The intermediate fusion feature F26 can be adjusted through operations such as convolution and pooling to obtain the adjusted intermediate fusion feature F31, whose size is H8×W8×D8×C12.
[0142] Then, the cross-channel interaction feature F30 is added to the adjusted intermediate fusion feature F31 element by element. For the element value F30(x, y, z, c) at position (x, y, z, c) in the cross-channel interaction feature F30 and the element value F31(x, y, z, c) at position (x, y, z, c) in the adjusted intermediate fusion feature F31, they are added to obtain the element value F32(x, y, z, c) at position (x, y, z, c) in the fusion feature map F32 = F30(x, y, z, c) + F31(x, y, z, c).
[0143] Through residual connection processing, the fused feature map F32 not only incorporates new cross-channel information brought by cross-channel interaction features, but also retains the original information of the intermediate fused features. This integration of information enables the fused feature map to more comprehensively and accurately reflect the characteristic information of oral and maxillofacial images, providing high-quality feature input for subsequent lesion area identification processing.
[0144] Step S140: performing lesion area recognition processing based on the fused feature map to determine the position information and morphological description information of the oral and maxillofacial abnormality area of the target patient.
[0145] After obtaining the fused feature map, lesion region recognition processing is performed based on it to determine the location and morphological description of the target patient's oral and maxillofacial abnormalities. Lesion region recognition processing is a key step in the entire image recognition diagnosis method. It can identify areas of possible lesions from the fused feature map and accurately describe the location and morphology of these areas.
[0146] Step S141: performing candidate region generation processing on the fused feature map to obtain a plurality of candidate region bounding boxes; the candidate region generation processing uses a sliding window mechanism to generate initial bounding boxes at different scales and aspect ratios on the fused feature map.
[0147] To identify areas in the fused feature map where lesions may be present, we first perform candidate region generation. This process uses a sliding window mechanism. A sliding window is a rectangular frame of a set size and shape that slides across the fused feature map.
[0148] The dimensions of the fused feature map F32 are H8×W8×D8×C12. In the sliding window mechanism, sliding windows of different scales and aspect ratios are used to slide across the fused feature map. For example, different window sizes can be set, such as small, medium, and large, with each window size having a different aspect ratio, such as 1:1, 2:1, or 1:2.
[0149] Each sliding window slides across the fused feature map at a set step size. With each slide, the window's position on the fused feature map is recorded, and this position defines an initial bounding box. As the sliding window continues to slide, multiple initial bounding boxes are generated. These initial bounding boxes cover different areas of the fused feature map, and each represents a candidate region where a lesion may be located. By combining all generated initial bounding boxes, multiple candidate region bounding boxes are generated.
[0150] Step S142: performing feature cropping processing on each candidate region bounding box to generate a candidate region feature map; the feature cropping processing intercepts feature data of the corresponding region from the fused feature map according to the bounding box coordinates.
[0151] After obtaining multiple candidate region bounding boxes, feature clipping is performed on each candidate region bounding box. The purpose of feature clipping is to extract the feature data corresponding to each candidate region from the fused feature map and generate a candidate region feature map.
[0152] For each candidate region bounding box, its coordinate information determines its position on the fused feature map. Based on the coordinates of the bounding box, the feature data of the corresponding region is intercepted from the fused feature map F32. For example, assuming that the coordinates of the upper left corner of a candidate region bounding box are (x1, y1, z1) and the coordinates of the lower right corner are (x2, y2, z2), then the feature data within the range of x1 to x2, y1 to y2, and z1 to z2 are intercepted from the fused feature map F32 to obtain a sub-feature map.
[0153] This sub-feature map is used as the candidate region feature map for the candidate region. This feature cropping process is performed on all candidate region bounding boxes to obtain multiple candidate region feature maps. Each candidate region feature map contains feature information for the corresponding candidate region, providing specific feature input for subsequent anomaly probability prediction.
[0154] Step S143: Input the candidate region feature map into a pre-trained lesion classification network for abnormality probability prediction processing to obtain an abnormality confidence score for each candidate region; the lesion classification network includes multiple fully connected layers for mapping the candidate region feature map into an abnormality probability value.
[0155] After obtaining the candidate region feature maps, they are fed into a pre-trained lesion classification network for abnormality probability prediction. The lesion classification network is a trained deep learning model that can predict the presence and probability of an abnormality in a region based on the input candidate region feature maps.
[0156] The lesion classification network consists of multiple fully connected layers. A fully connected layer is a simple linear transformation layer that maps the input feature vector to a new feature space. For each candidate region feature map, it is first expanded into a one-dimensional vector. Assuming the size of the candidate region feature map is H9×W9×D9×C12, it is expanded into a one-dimensional vector with a length of H9*W9*D9*C12.
[0157] This one-dimensional vector is then fed into the first fully connected layer of the lesion classification network. The weight matrix of this first fully connected layer linearly transforms the input one-dimensional vector, generating a new feature vector. This new feature vector is then fed into the next fully connected layer, and so on, through multiple fully connected layers.
[0158] The output of the last fully connected layer is a vector of length 1. This vector is mapped to the interval [0, 1] using a sigmoid function, resulting in the anomaly probability value for the candidate region. This anomaly probability value is the anomaly confidence score for the candidate region. This anomaly probability prediction process is performed on all candidate region feature maps to obtain an anomaly confidence score for each candidate region.
[0159] Step S144: Filter the candidate regions corresponding to the abnormal confidence scores according to a preset confidence threshold to generate an abnormal region candidate set.
[0160] After obtaining the anomaly confidence score for each candidate region, these candidate regions are screened according to a preset confidence threshold to generate a candidate set of anomaly regions. The preset confidence threshold is a pre-set value used to determine whether a candidate region is likely to be an anomaly region.
[0161] For the anomaly confidence score of each candidate region, if the score is greater than the preset confidence threshold, the candidate region is considered to be a possible anomaly region and is added to the anomaly region candidate set; if the score is less than or equal to the preset confidence threshold, the candidate region is considered unlikely to be an anomaly region and is excluded.
[0162] Through this screening operation, candidate regions with high anomaly confidence scores are selected from all candidate regions to form a candidate set of abnormal regions. The candidate set of abnormal regions contains candidate regions that may contain lesions, reducing the workload of subsequent processing and improving the efficiency of lesion region identification.
[0163] Step S145: sorting each candidate region in the abnormal region candidate set in descending order according to the abnormality confidence score to generate a sorted candidate list.
[0164] After obtaining the abnormal region candidate set, in order to further screen out the areas most likely to be lesions, it is necessary to sort each candidate region in the abnormal region candidate set in descending order by abnormality confidence score to generate a sorted candidate list. Sorting helps prioritize candidate regions with high abnormality probability, improving the efficiency and accuracy of subsequent processing.
[0165] The sorting process is based on the anomaly confidence score. For each candidate region in the anomaly region candidate set, its anomaly confidence score is a quantitative indicator representing the likelihood of an anomaly in that region. These candidate regions are sorted from high to low according to their anomaly confidence score to form a sorted candidate list.
[0166] In actual operation, common sorting algorithms can be used, such as quick sort, merge sort, etc. Taking quick sort as an example, its basic idea is to select a base element and divide the list into two parts, so that the elements in the left part are less than or equal to the base element, and the elements in the right part are greater than or equal to the base element, and then recursively sort the left and right parts respectively.
[0167] Assume that the set of candidate anomaly regions is set A, where each candidate region has a corresponding anomaly confidence score. Select the first candidate region in set A as the base element and divide set A into two parts: set B, which contains candidate regions with anomaly confidence scores less than or equal to the base element, and set C, which contains candidate regions with anomaly confidence scores greater than the base element. Then, recursively sort set B and set C, respectively, to obtain a sorted candidate list in descending order of anomaly confidence score.
[0168] The first candidate region in the ranked candidate list has the highest anomaly confidence score, meaning it is most likely to be the lesion region, while the likelihood of anomaly for subsequent candidate regions decreases. This ranked candidate list provides an ordered set of candidate regions for subsequent deduplication and final abnormal region determination.
[0169] Step S146: Select the first candidate region from the sorted candidate list as the base region, and calculate the spatial overlap between the remaining candidate regions and the base region; the spatial overlap is calculated by using an intersection-over-union algorithm to calculate the overlapping area ratio between the base region bounding box and the other candidate region bounding boxes.
[0170] After obtaining the sorted candidate list, the first candidate region in the list is selected as the base region. Since this candidate region has the highest anomaly confidence score, it is prioritized as a reference to determine whether other candidate regions overlap with it.
[0171] Next, the spatial overlap between the remaining candidate regions and the benchmark region is calculated. Spatial overlap is a measure of the degree of spatial overlap between two candidate regions and is calculated using the Intersection over Union (IoU) algorithm.
[0172] The core of the intersection-in-union algorithm is to calculate the intersection area and union area of the bounding box of the reference region and the bounding boxes of other candidate regions, and then divide the intersection area by the union area to obtain the overlapping area ratio.
[0173] Assume that the bounding box of the reference area is a rectangular box R1, whose upper left corner coordinates are (x1, y1, z1) and lower right corner coordinates are (x2, y2, z2); the bounding box of another candidate area is a rectangular box R2, whose upper left corner coordinates are (x3, y3, z3) and lower right corner coordinates are (x4, y4, z4).
[0174] First, calculate the intersection area of the two bounding boxes. The upper left corner coordinates of the intersection area are (max(x1, x3), max(y1, y3), max(z1, z3)), and the lower right corner coordinates are (min(x2, x4), min(y2, y4), min(z2, z4)). If the upper left corner coordinates of the intersection area are greater than the lower right corner coordinates, it means that the two bounding boxes do not intersect, and the intersection area is 0.
[0175] Then, calculate the volume of the intersection region, V_intersection. Assuming the length, width, and height of the intersection region are l, w, and h respectively, then V_intersection = l*w*h, where l = max(0, min(x2, x4) - max(x1, x3)), w = max(0, min(y2, y4) - max(y1, y3)), and h = max(0, min(z2, z4) - max(z1, z3)).
[0176] Next, calculate the union area of the two bounding boxes. The union area is equal to the sum of the volumes of the two bounding boxes minus the intersection area. The volume of the reference region bounding box V1 = (x2-x1)*(y2-y1)*(z2-z1), and the volume of the other candidate region bounding box V2 = (x4-x3)*(y4-y3)*(z4-z3). Therefore, the union area V_union = V1+V2-V_intersection.
[0177] Finally, calculate the intersection-over-union (IoU) = V_intersection / V_union. The value of IoU is the spatial overlap of the two candidate regions, which reflects the degree of spatial overlap between the two candidate regions. A larger value indicates a higher degree of overlap.
[0178] For each candidate region in the sorted candidate list, excluding the baseline region, the spatial overlap between the candidate region and the baseline region is calculated using the above method. These spatial overlap values will be used in subsequent screening operations to remove candidate regions that highly overlap with the baseline region, thus avoiding repeated identification of the same lesion region.
[0179] Step S147: screening the candidate areas whose spatial overlap satisfies the conditions according to a preset overlap threshold, removing them, and updating the sorted candidate list.
[0180] After calculating the spatial overlap between the remaining candidate regions and the reference region, these candidate regions are screened based on a preset overlap threshold. The preset overlap threshold is a pre-set value used to determine whether two candidate regions overlap too much and need to be removed.
[0181] If the spatial overlap between a candidate region and the reference region is greater than a preset overlap threshold, it means that the two candidate regions are likely to describe the same lesion region. In order to avoid repeated identification, the candidate region needs to be removed from the sorted candidate list.
[0182] The specific operation is to traverse each candidate area in the sorted candidate list except the base area and compare its spatial overlap with the base area with a preset overlap threshold. If the spatial overlap exceeds the preset overlap threshold, the candidate area is removed from the sorted candidate list; if the spatial overlap is less than or equal to the preset overlap threshold, the candidate area is retained.
[0183] For example, assuming that the preset overlap threshold is a specific value T, for the candidate area R in the sorted candidate list, if its spatial overlap with the reference area IoU(R)>T, the candidate area R is removed from the sorted candidate list; if IoU(R)<=T, the candidate area R continues to remain in the sorted candidate list.
[0184] After this screening and removal process, the sorted candidate list is updated. In the updated sorted candidate list, the overlap between each candidate area and the reference area is below the preset overlap threshold, which reduces the possibility of duplicate identification and makes the subsequent determination of abnormal areas more accurate.
[0185] Step S148: Repeat the following operations until the sorting candidate list is an empty set: select the candidate area with the highest anomaly confidence score in the current sorting candidate list as the base area, and calculate the spatial overlap between the remaining candidate areas and the base area; if the spatial overlap between the candidate area and the base area exceeds a preset overlap threshold, remove the candidate area from the sorting candidate list; add the base areas that are not removed to the final abnormal area set, and record their position information and morphological description information; the morphological description information includes the maximum diameter, shape irregularity and edge clarity index of the abnormal area.
[0186] After completing a screening and removal process, there may still be multiple candidate regions in the sorted candidate list. To ensure that all possible abnormal regions are accurately identified, a series of operations need to be repeated until the sorted candidate list is empty.
[0187] Each time the operation is repeated, the candidate region with the highest anomaly confidence score is first selected from the current sorted candidate list as the new base region. Since the sorted candidate list is sorted in descending order of anomaly confidence score, the first candidate region is the region with the highest anomaly confidence score.
[0188] Then, the spatial overlap between the remaining candidate regions and the new reference region is calculated. The calculation method is the same as before, using the intersection-over-union algorithm to calculate the overlap ratio between the reference region bounding box and the bounding boxes of the other candidate regions.
[0189] Then, the spatial overlap of these regions is determined based on a preset overlap threshold. If the spatial overlap between a candidate region and the new reference region exceeds the preset overlap threshold, it indicates that the two candidate regions are likely to describe the same lesion region, and the candidate region is removed from the sorted candidate list.
[0190] After removing candidate regions with high overlap, the remaining reference regions are added to the final set of abnormal regions. The location and morphological description of the reference regions are also recorded. The location information can be represented by the coordinates of the bounding box, such as the upper left and lower right corners. The morphological description includes the maximum diameter of the abnormal region, shape irregularity, and edge clarity.
[0191] The maximum diameter can be calculated by calculating the length of the longest line segment within the bounding box of the abnormal region. Shape irregularity can be measured by comparing the actual shape of the abnormal region with a regular shape (such as a circle or rectangle). For example, the ratio of the perimeter of the abnormal region to the perimeter of a regular shape of the same area can be calculated. Edge clarity can be determined by analyzing the grayscale gradient at the edge of the abnormal region. A larger grayscale gradient indicates a sharper edge.
[0192] Repeat the above steps, each time selecting the candidate region with the highest anomaly confidence score from the updated sorted candidate list as the reference region, performing spatial overlap calculation, screening, and removal operations until the sorted candidate list is empty. At this point, the final set of anomaly regions contains all the screened and confirmed anomaly regions. The location and morphological description of these regions will be used in the subsequent diagnosis report generation.
[0193] Step S150: Generate a diagnosis report based on the location information and morphological description information, and transmit the diagnosis report to a medical terminal device for display.
[0194] After determining the location information and morphological description information of the target patient's oral and maxillofacial abnormalities, a diagnostic report needs to be generated based on this information and transmitted to the medical terminal device for display so that the doctor can intuitively understand the patient's condition.
[0195] Step S151: Matching the position information with a preset anatomical structure database to determine the anatomical part name corresponding to the abnormal area; the matching process spatially aligns the center coordinates of the abnormal area with the standard coordinates of the anatomical structure, and combines the topological relationship constraints between the anatomical structures to determine the anatomical part name corresponding to the minimum distance.
[0196] In order to accurately describe the location of the abnormal area, the location information needs to be matched with the preset anatomical structure database. The preset anatomical structure database contains the standard coordinate information of various anatomical parts of the oral and maxillofacial region and the topological relationships between them.
[0197] First, calculate the center coordinates of the abnormal region. For each abnormal region's bounding box, its center coordinates can be calculated using the coordinates of the upper left corner and lower right corner of the bounding box. Assuming the coordinates of the upper left corner of the bounding box are (x1, y1, z1) and the coordinates of the lower right corner are (x2, y2, z2), the center coordinates are ((x1+x2) / 2, (y1+y2) / 2, (z1+z2) / 2).
[0198] Then, the center coordinates of the abnormal area are spatially registered with the standard coordinates of the anatomical structure in the preset anatomical structure database. The purpose of spatial registration is to map the center coordinates of the abnormal area to the coordinate system of the anatomical structure database for accurate matching.
[0199] Next, the distance between the center coordinates of the abnormal region and the standard coordinates of each anatomical structure is calculated, incorporating topological constraints between the anatomical structures. Topological constraints can help eliminate matching results that do not conform to the anatomical structure logic. For example, an abnormal region cannot be located in a non-anatomical region between two anatomical sites.
[0200] Finally, the anatomical part corresponding to the minimum distance is determined. The anatomical part with the smallest distance between the center coordinates of the abnormal region and the standard coordinates of each anatomical structure is selected as the anatomical part corresponding to the abnormal region. This allows the location of the abnormal region within the oral and maxillofacial anatomy to be accurately determined.
[0201] Step S152: Compare the morphological description information with a preset lesion feature database to determine the pathological type prediction result of the abnormal area; the comparison process uses a nearest neighbor algorithm to perform similarity matching between the morphological description feature vector and the standard lesion feature in the database.
[0202] After determining the anatomical part corresponding to the abnormal region, the morphological description information needs to be compared with the preset lesion feature database to determine the predicted pathological type of the abnormal region. The preset lesion feature database contains standard lesion features for various common pathological types, such as the characteristic values of the maximum diameter, shape irregularity, and edge clarity of lesions of different pathological types.
[0203] First, the morphological description information is converted into a morphological description feature vector. The morphological description feature vector is a multidimensional vector, with each dimension corresponding to a morphological description indicator, such as maximum diameter, shape irregularity, edge clarity, etc. For example, the morphological description feature vector can be expressed as (d, s, e), where d represents the maximum diameter, s represents the shape irregularity, and e represents the edge clarity.
[0204] Then, a nearest neighbor algorithm is used to perform a similarity match between the morphological description feature vector and the standard lesion features in a pre-set lesion feature database. The basic idea of the nearest neighbor algorithm is to calculate the similarity between the morphological description feature vector and each standard lesion feature vector in the database and select the pathology type corresponding to the standard lesion feature vector with the highest similarity as the pathology type prediction result of the abnormal area.
[0205] The similarity can be calculated using methods such as Euclidean distance and cosine similarity.
[0206] The Euclidean distance between the morphological description feature vector and all standard lesion feature vectors in the database is calculated, and the pathological type corresponding to the standard lesion feature vector with the smallest distance is selected as the pathological type prediction result of the abnormal area.
[0207] Step S153: generating a diagnosis suggestion text according to the anatomical part name and the pathological type prediction result; the diagnosis suggestion text is generated by semantically combining the anatomical part name and the pathological type through a predefined template.
[0208] After determining the anatomical location and pathology prediction for the abnormal region, a diagnostic recommendation is generated based on this information. This recommendation serves as a crucial reference for doctors in diagnosis and treatment, and must accurately and clearly describe the abnormal region and provide recommendations.
[0209] A predefined template is a pre-set text format that semantically combines the anatomical site name and pathology type. For example, a predefined template might read "A [pathology type] lesion was found at [anatomical site name]; further [specific examination or treatment recommendations] are recommended."
[0210] Substitute the anatomical site name and pathology type prediction result into a predefined template to generate a diagnostic recommendation text. For example, if the anatomical site name is "maxilla" and the pathology type prediction result is "cyst", the diagnostic recommendation text may be "A cyst lesion was found in the maxilla. Further enhanced CT scan is recommended to clarify the nature of the lesion."
[0211] Step S154: formatting and combining the diagnosis suggestion text with the location information and morphological description information to generate a diagnosis report; the formatting and combining process formats and integrates the text information and coordinate data according to a standard medical report format.
[0212] After the diagnostic recommendation text is generated, it is formatted and combined with the location information and morphological description information to generate a diagnostic report. The purpose of this formatting and combination is to typeset and integrate the text information and coordinate data according to the standard medical report format, making the diagnostic report standardized and readable.
[0213] First, the location information is presented in a clear manner, such as by listing the coordinates of the abnormal region's bounding box or center. Next, the morphological description information, such as maximum diameter, shape irregularity, and edge clarity, is described in detail. Finally, the diagnostic recommendation text is added to the report, correlating it with the location and morphological description information.
[0214] In terms of layout, follow the standard format of medical reports, setting the format of the title, body, paragraphs, etc. For example, set the report title at the beginning, such as "Oral and Maxillofacial Imaging Diagnostic Report", and then list the patient information, abnormal area location information, morphological description information, pathology type prediction results, and diagnostic recommendations.
[0215] Finally, the formatted and combined content is saved as a diagnostic report file, such as PDF or Word format. The generated diagnostic report is transmitted to a medical terminal device for display. Doctors can view the diagnostic report on the medical terminal device to understand the patient's oral and maxillofacial abnormalities, providing a basis for subsequent diagnosis and treatment.
[0216] In the above embodiments, privacy-sensitive data such as the patient's oral and maxillofacial three-dimensional image data set is involved. In order to protect the patient's privacy and prevent data leakage, a series of privacy protection and anti-leakage technical means are adopted. For example, during the data acquisition process, the patient's identity information is encrypted. When collecting oral and maxillofacial three-dimensional image data, the patient's real name, ID number and other sensitive identity information are not directly recorded. Instead, an encryption algorithm is used to convert this information into an encryption code. For example, a symmetric encryption algorithm, such as the AES algorithm, is used to encrypt the patient's identity information, and the encryption key is managed by a dedicated security system. At the same time, strict security management is performed on the acquisition equipment. Ensure that the software and hardware of the acquisition equipment have undergone security testing to prevent data from being illegally obtained during the acquisition process. Strict authority control is implemented for access to the acquisition equipment, and only authorized personnel can operate the acquisition equipment.
[0217] This method involves multiple pre-trained models, such as three-dimensional convolutional neural networks, lesion classification networks, and multi-scale feature fusion networks. The general steps of model training include data preparation, model construction, model training, and model evaluation. For details, please refer to the above model application process to execute the corresponding training process, which will not be repeated here.
[0218] Based on the above description, in another embodiment, the present invention further provides an oral and maxillofacial surgery image recognition and diagnosis system based on deep learning, see Figure 2 , Figure 2 This is a structural diagram of a deep learning-based oral and maxillofacial surgery image recognition and diagnosis system 100 provided in an embodiment of the present invention. The deep learning-based oral and maxillofacial surgery image recognition and diagnosis system 100 may vary significantly due to different configurations or performance, and may include one or more central processing units (CPUs) 112 (for example, one or more processors) and a memory 111. The memory 111 may be either short-term storage or persistent storage. The program stored in the memory 111 may include one or more modules, each of which may include a series of instruction operations in the deep learning-based oral and maxillofacial surgery image recognition and diagnosis system 100. Furthermore, the CPU 112 may be configured to communicate with the memory 111 to execute the series of instruction operations in the memory 111 on the deep learning-based oral and maxillofacial surgery image recognition and diagnosis system 100.
[0219] The deep learning-based oral and maxillofacial surgery image recognition and diagnosis system 100 may also include one or more power supplies, one or more communication units 113, one or more transmission output interfaces, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0220] The steps performed by the deep learning-based oral and maxillofacial surgery image recognition and diagnosis system in the above embodiment can be combined with Figure 2 The structure of the oral and maxillofacial surgery image recognition and diagnosis system based on deep learning is shown.
[0221] In addition, an embodiment of the present invention further provides a storage medium, which is used to store a computer program, and the computer program is used to execute the method provided by the above embodiment.
[0222] An embodiment of the present invention further provides a computer program product including instructions, which, when executed on a computer, enables the computer to execute the method provided in the above embodiment.
[0223] Those skilled in the art will understand that all or part of the steps of the above-mentioned method embodiment can be completed by hardware related to program instructions, and the above-mentioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the above-mentioned storage medium can be at least one of the following media: read-only memory (English: Readonly Memory, abbreviated: ROM), RAM, magnetic disk or optical disk, etc., various media that can store program codes.
[0224] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. The device and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. A person of ordinary skill in the art can understand and implement it without making any creative effort.
[0225] The above is merely one specific implementation step of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A deep learning-based oral and maxillofacial surgery image recognition and diagnosis method, characterized in that: include: Acquire a set of three-dimensional oral and maxillofacial image data of a target patient, the three-dimensional image data set comprising multiple sets of continuously scanned slice image data, wherein the scanning parameters include a scanning range and a slice spacing, the scanning range covering key areas of the target patient's oral and maxillofacial region, including teeth, jawbone, and surrounding soft tissue, and the slice spacing determining the distance between two adjacent layers of scanned slice image data. During the scanning process, the oral and maxillofacial CT scanner rotates around the patient's head and scans. Each time it rotates to a specific angle, a set of slice image data is acquired to obtain multiple sets of continuous slice image data. The slice image data are spatially continuous and present structural information of the target patient's oral and maxillofacial region from different angles and levels, including sagittal, coronal, and transverse planes. Each set of slice image data is a two-dimensional matrix, wherein each element in the two-dimensional matrix represents the image grayscale value of the corresponding position. The three-dimensional image data set is a three-dimensional matrix formed by stacking multiple two-dimensional matrices in a scanning order, and each element in the three-dimensional matrix has its own specific spatial coordinates, corresponding to a specific position in the oral and maxillofacial region. Performing image feature extraction processing on the three-dimensional image data set to obtain a hierarchical image feature set of the continuously scanned slice image data; the hierarchical image feature set includes local anatomical structure features and global spatial distribution features; Calling a pre-trained multi-scale feature fusion network to perform multi-scale feature fusion on the hierarchical image feature set to generate a fusion feature map; Performing lesion area recognition processing based on the fused feature map to determine the location information and morphological description information of the oral and maxillofacial abnormality area of the target patient; generating a diagnosis report based on the location information and morphological description information, and transmitting the diagnosis report to a medical terminal device for display; The calling of a pre-trained multi-scale feature fusion network to perform multi-scale feature fusion on the hierarchical image feature set to generate a fusion feature map includes: Performing feature splicing processing on the local anatomical structure feature and the global spatial distribution feature to generate an initial splicing feature; Performing a grouped axial self-attention mechanism on the initial spliced features to generate a self-attention weight distribution map; the grouped axial self-attention mechanism includes decomposing the features into axial sub-regions along the spatial dimension, calculating the correlation matrix between spatial positions in each axial sub-region, and normalizing the correlation matrix to generate a local self-attention weight distribution map; Dynamically weighting the local anatomical structure features based on the self-attention weight distribution map to generate weighted local features; the dynamic weighting processing includes performing an element-by-element multiplication operation on the self-attention weight distribution map and the local anatomical structure features; Performing feature fusion processing on the weighted local features and the global spatial distribution features to generate intermediate fusion features; Performing cross-channel feature interaction processing on the intermediate fusion features to generate a fusion feature map; the cross-channel feature interaction processing includes decomposing the intermediate fusion features into multiple channel groups, performing independent convolution operations on each channel group, and rearranging the channel order to generate a fusion feature map with channel interaction characteristics; The performing cross-channel feature interaction processing on the intermediate fusion features to generate a fusion feature map includes: Dividing the intermediate fusion features into a plurality of sub-feature groups according to the channel dimension; wherein each sub-feature group includes a predetermined number of continuous channels; Performing group convolution processing on each of the sub-feature groups to generate a group feature map; the group convolution processing uses an independent convolution kernel to act on each sub-feature group respectively; Performing channel shuffling processing on the grouped feature graph to generate a shuffled feature graph; the channel shuffling processing rearranges the channels in each sub-feature group according to a predetermined rule to promote cross-group information interaction; Performing point-by-point convolution processing on the shuffled feature map to generate cross-channel interaction features; the point-by-point convolution processing uses a 1×1 convolution kernel to adjust the channel dimension; Performing residual connection processing on the cross-channel interaction feature and the intermediate fusion feature to generate a fusion feature map; the residual connection processing adds the cross-channel interaction feature and the original intermediate fusion feature element by element; The performing of lesion area identification processing based on the fused feature map to determine the location information and morphological description information of the oral and maxillofacial abnormality area of the target patient includes: Performing a candidate region generation process on the fused feature map to obtain a plurality of candidate region bounding boxes; the candidate region generation process uses a sliding window mechanism to generate initial bounding boxes at different scales and aspect ratios on the fused feature map; Performing feature clipping processing on each candidate region bounding box to generate a candidate region feature map; the feature clipping processing intercepts feature data of the corresponding region from the fused feature map according to the bounding box coordinates; Inputting the candidate region feature map into a pre-trained lesion classification network for abnormality probability prediction processing to obtain an abnormality confidence score for each candidate region; the lesion classification network includes multiple fully connected layers for mapping the candidate region feature map into an abnormality probability value; Filtering candidate regions corresponding to the abnormal confidence scores according to a preset confidence threshold to generate an abnormal region candidate set; Sorting each candidate region in the abnormal region candidate set in descending order according to the abnormality confidence score to generate a sorted candidate list; Selecting the first candidate region from the sorted candidate list as the reference region, and calculating the spatial overlap between the remaining candidate regions and the reference region; the spatial overlap is calculated using an intersection-over-union algorithm to calculate the overlap area ratio between the reference region bounding box and the other candidate region bounding boxes; Screening the candidate areas whose spatial overlap satisfies the conditions according to a preset overlap threshold value for removal, and updating the sorted candidate list; Repeat the following operations until the sorted candidate list is an empty set: select the candidate area with the highest anomaly confidence score in the current sorted candidate list as the base area, and calculate the spatial overlap between the remaining candidate areas and the base area; if the spatial overlap between the candidate area and the base area exceeds a preset overlap threshold, remove the candidate area from the sorted candidate list; add the base areas that are not removed to the final abnormal area set, and record their position information and morphological description information; the morphological description information includes the maximum diameter, shape irregularity and edge clarity index of the abnormal area.
2. The deep learning-based oral and maxillofacial surgery image recognition and diagnosis method according to claim 1, characterized in that: The performing image feature extraction processing on the three-dimensional image data set to obtain the hierarchical image feature set of the continuously scanned layer image data includes: Performing spatial normalization processing on the continuously scanned slice image data to obtain a standardized three-dimensional image sequence; Calling a three-dimensional convolutional neural network to perform shallow feature extraction processing on the standardized three-dimensional image sequence to obtain a primary image feature set; the primary image feature set includes edge contour features and texture distribution features; Performing cross-layer feature enhancement processing on the primary image feature set to obtain an enhanced image feature set; the enhanced image feature set includes local anatomical structure features after spatial resolution enhancement; Performing global pooling processing on the enhanced image feature set to obtain global spatial distribution features; Performing feature alignment processing on the local anatomical structure features and the global spatial distribution features to form a hierarchical image feature set; the feature alignment processing includes performing a spatial coordinate mapping operation on the local anatomical structure features to generate a spatial alignment feature vector corresponding to the global spatial distribution features; Performing a feature broadcasting operation on the global spatial distribution feature, expanding its spatial dimension to be consistent with the spatial alignment feature vector, and adjusting the channel dimension matching; The expanded global spatial distribution features are concatenated with the spatial alignment feature vector in channel dimension to generate a hierarchical image feature set with a unified spatial dimension.
3. The deep learning-based oral and maxillofacial surgery image recognition and diagnosis method according to claim 2, characterized in that: The calling of the three-dimensional convolutional neural network to perform shallow feature extraction processing on the standardized three-dimensional image sequence to obtain a primary image feature set includes: Performing convolution kernel sliding processing on the standardized three-dimensional image sequence to generate an initial convolution feature map, and performing maximum pooling processing on the initial convolution feature map to obtain a dimensionality reduction feature map; Inputting the reduced-dimensionality feature map into a residual connection module for feature compensation processing to obtain a compensated feature map; the feature compensation processing includes performing an element-by-element addition operation on the reduced-dimensionality feature map and the original feature map transferred through the jump connection to generate a compensated feature map with detail preservation characteristics; Performing a channel attention weight allocation process on the compensated feature map to generate a weighted feature map; the channel attention weight allocation process includes performing a global average pooling operation on the channel dimension of the compensated feature map to generate a channel description vector; Inputting the channel description vector into a fully connected layer to generate a channel attention weight, and recalibrating each channel of the compensated feature map based on the channel attention weight to generate a weighted feature map; The edge contour features and texture distribution features in the weighted feature map are extracted to form a primary image feature set; the edge contour features are used to extract gradient features from the weighted feature map through a learnable edge convolution kernel, and the texture distribution features are used to extract adaptive texture patterns from the weighted feature map through a multi-scale convolution kernel group.
4. The deep learning-based oral and maxillofacial surgery image recognition and diagnosis method according to claim 2, characterized in that: The performing cross-layer feature enhancement processing on the primary image feature set to obtain an enhanced image feature set includes: Performing feature splicing processing on the edge contour features and texture distribution features in the primary image feature set to generate a splicing feature map; Performing a dilated convolution process on the spliced feature map to generate a multi-scale receptive field feature map; the dilated convolution process uses convolution kernels with different dilation rates to act in parallel on the spliced feature map to generate feature submaps with different receptive field ranges; Merging the feature subgraphs in channel dimensions to generate a multi-scale receptive field feature map, and performing feature pyramid fusion processing on the multi-scale receptive field feature map to obtain a multi-scale fused feature map; the feature pyramid fusion processing includes upsampling or downsampling feature subgraphs of different scales to a uniform resolution and performing element-by-element addition operation; Performing nonlinear activation processing on the multi-scale fusion feature map to generate an activation feature map; the nonlinear activation processing uses an activation function with a gating mechanism to dynamically activate each channel of the multi-scale fusion feature map; Performing a spatial attention weight allocation process on the activation feature map to obtain an enhanced image feature set; the spatial attention weight allocation process includes performing a dual-path operation of maximum pooling and average pooling on the spatial dimension of the activation feature map to generate a dual-path pooling feature map; The dual-channel pooling feature map is channel-joined and then a spatial attention weight map is generated through a convolutional layer. The activation feature map is spatially weighted based on the spatial attention weight map to generate an enhanced image feature set.
5. The deep learning-based oral and maxillofacial surgery image recognition and diagnosis method according to claim 2, characterized in that: The performing global pooling processing on the enhanced image feature set to obtain global spatial distribution features includes: Performing channel dimension average pooling processing on the local anatomical structure features in the enhanced image feature set to generate a channel statistical feature vector; Performing a fully connected layer mapping process on the channel statistical feature vector to generate a dimension-reduced feature vector; Mapping the reduced-dimensional feature vector to a channel dimension consistent with the enhanced image feature set through a fully connected layer, and performing feature broadcasting processing to generate a spatial weight distribution map; the feature broadcasting processing includes copying and expanding the mapped feature vector to the same spatial size as the enhanced image feature set; Based on the spatial weight distribution map, the enhanced image feature set is subjected to channel dimension weighting and spatial dimension aggregation processing to obtain a global spatial distribution feature.
6. The deep learning-based oral and maxillofacial surgery image recognition and diagnosis method according to claim 1, characterized in that: Generating a diagnosis report according to the location information and morphological description information includes: Matching the position information with a preset anatomical structure database to determine the anatomical part name corresponding to the abnormal area; the matching process spatially aligns the center coordinates of the abnormal area with the standard coordinates of the anatomical structure, and combines the topological relationship constraints between the anatomical structures to determine the anatomical part name corresponding to the minimum distance; Comparing the morphological description information with a preset lesion feature database to determine a pathological type prediction result of the abnormal area; the comparison process uses a nearest neighbor algorithm to perform similarity matching between the morphological description feature vector and the standard lesion feature in the database; Generate a diagnosis suggestion text based on the anatomical part name and the pathological type prediction result; the diagnosis suggestion text is generated by semantically combining the anatomical part name and the pathological type through a predefined template; The diagnosis suggestion text is formatted and combined with the position information and morphological description information to generate a diagnosis report; the formatting and combining process formats and integrates the text information and coordinate data according to a standard medical report format.
7. A deep learning-based oral and maxillofacial surgery image recognition and diagnosis system, characterized by: include: processor; A memory having a computer program stored therein, wherein the computer program, when executed, implements the oral and maxillofacial surgery image recognition and diagnosis method based on deep learning according to any one of claims 1 to 6.