Oral and maxillofacial surgery image recognition and diagnosis method and system based on deep learning

Through the multi-scale feature fusion of oral and maxillofacial image data based on deep learning, the problems of low efficiency and insufficient accuracy of three-dimensional image data processing in traditional imaging diagnostic methods are solved, and efficient and accurate lesion area identification and diagnostic report generation are achieved.

CN120280132AActive Publication Date: 2025-07-08JILIN UNIVERSITY

Patent Information

Application Number
CN202510766267.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

Traditional oral and maxillofacial imaging diagnostic methods rely on two-dimensional images and are difficult to fully display the three-dimensional anatomical structure and lesion characteristics, resulting in insufficient diagnostic accuracy and reliability. The existing computer-assisted diagnostic methods are difficult to fully extract and fuse multi-scale features when processing three-dimensional image data, and cannot accurately identify the lesion area.

Method used

Using a deep learning-based method, image feature extraction and multi-scale feature fusion are performed by acquiring oral and maxillofacial three-dimensional image data, pre-trained multi-scale feature fusion networks are used to generate a fusion feature map, identify lesion areas and generate diagnostic reports.

Benefits of technology

It improves the accuracy and efficiency of oral and maxillofacial imaging diagnosis, can accurately identify the location and morphology of the lesion area, provide detailed diagnostic basis, and achieve rapid transmission and intuitive display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120280132A_ABST
    Figure CN120280132A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an oral and maxillofacial surgical image recognition and diagnosis method and system based on deep learning, and the method comprises the steps: firstly obtaining an oral and maxillofacial three-dimensional image data set of a target patient, then carrying out the image feature extraction processing of the three-dimensional image data set, and obtaining a hierarchical image feature set; comprising local anatomical structure features and global spatial distribution features, and then calling a pre-trained multi-scale feature fusion network to perform multi-scale feature fusion on the hierarchical image feature set to generate a fusion feature map. And performing focus area identification processing based on the fusion characteristic spectrum, determining position information and form description information of an oral and maxillofacial abnormal area of the target patient, generating a diagnosis report according to the position information and form description information of the oral and maxillofacial abnormal area, and transmitting the diagnosis report to medical terminal equipment for display. Therefore, the accuracy and efficiency of oral and maxillofacial surgery image diagnosis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and in particular, to an oral and maxillofacial surgical image recognition and diagnosis method and system based on deep learning. Background Art

[0002] In the clinical diagnosis and treatment process of oral and maxillofacial surgery, accurate image recognition and diagnosis are of crucial significance for early disease detection, treatment plan formulation, and prognosis evaluation. Traditional oral and maxillofacial image diagnosis methods mainly rely on doctors' observation and analysis of two-dimensional image films based on their own experience. However, two-dimensional images have information limitations and are difficult to comprehensively display the complex three-dimensional anatomical structures and lesion characteristics of the oral and maxillofacial region, which may lead to deviations in doctors' judgments of lesions and affect the accuracy and reliability of diagnosis.

[0003] With the continuous development of medical imaging technology, oral and maxillofacial three-dimensional image data has gradually become an important basis for clinical diagnosis. However, three-dimensional image data contains a large amount of continuous scanned slice image data, with a huge and complex data volume. Manual processing and analysis are inefficient and difficult to meet the needs of rapid clinical diagnosis. At the same time, existing computer-aided diagnosis methods often have difficulty fully extracting and fusing multi-scale features in images when processing three-dimensional image data, and are unable to accurately identify the location and morphology of lesion areas, resulting in the accuracy and reliability of diagnosis results still needing to be improved. Therefore, it is of great practical significance to develop an efficient and accurate oral and maxillofacial surgical image recognition and diagnosis method based on deep learning. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide an oral and maxillofacial surgical image recognition and diagnosis method and system based on deep learning.

[0005] In a first aspect, embodiments of the present invention provide an oral and maxillofacial surgical image recognition and diagnosis method based on deep learning, which is applied to an oral and maxillofacial surgical image recognition and diagnosis system based on deep learning, and includes:

[0006] Obtain a set of oral and maxillofacial three-dimensional image data of a target patient, where the three-dimensional image data set contains multiple groups of continuous scanned slice image data;

[0007] Perform image feature extraction processing on the three-dimensional image data set to obtain a hierarchical image feature set of the continuous scanned slice image data; the hierarchical image feature set includes local anatomical structure features and global spatial distribution features;

[0008] Call a pre-trained multi-scale feature fusion network to perform multi-scale feature fusion on the hierarchical image feature set to generate a fused feature map;

[0009] Perform lesion area recognition processing based on the fused feature map to determine the location information and morphological description information of the oral and maxillofacial abnormal area of the target patient;

[0010] Generate a diagnostic report according to the location information and morphological description information, and transmit the diagnostic report to a medical terminal device for display.

[0011] In a second aspect, an embodiment of the present invention provides an oral and maxillofacial surgical image recognition and diagnosis system based on deep learning, including:

[0012] A processor;

[0013] A memory, in which a computer program is stored, and when the computer program is executed, it implements the oral and maxillofacial surgical image recognition and diagnosis method based on deep learning described in the first aspect.

[0014] As above, in the embodiment of the present invention, by acquiring the oral and maxillofacial three-dimensional image data set of the target patient and performing image feature extraction processing on the three-dimensional image data set, the local anatomical structure features and global spatial distribution features of the continuous scanned layer image data can be accurately obtained. Call the pre-trained multi-scale feature fusion network to perform multi-scale feature fusion on the hierarchical image feature set, effectively integrate the feature information of different scales, generate a more representative fused feature map, greatly improve the recognition ability of the lesion area, and perform lesion area recognition processing based on the fused feature map, which can accurately determine the location information and morphological description information of the oral and maxillofacial abnormal area of the target patient, provide detailed and accurate diagnostic basis for doctors, and finally generate a diagnostic report according to the location information and morphological description information and transmit it to a medical terminal device for display, realizing the rapid transmission and intuitive display of the diagnostic results, and effectively improving the accuracy and efficiency of oral and maxillofacial surgical image diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic flowchart of the steps of an oral and maxillofacial surgical image recognition and diagnosis method provided by an embodiment of the present invention;

[0016] Figure 2 For performing Figure 1 in the oral and maxillofacial surgical image recognition and diagnosis method based on deep learning provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts in the embodiments of the present invention belong to the scope of protection of the present invention.

[0018] See Figure 1 as shown in:

[0019] Step S110: Obtain a three-dimensional oral and maxillofacial image data set of a target patient, where the three-dimensional image data set includes multiple groups of continuous scanned slice image data.

[0020] In the process of image recognition and diagnosis in oral and maxillofacial surgery, the three-dimensional oral and maxillofacial image data set of the target patient needs to be scanned using relevant medical imaging equipment, such as an oral and maxillofacial CT scanner. Before scanning, appropriate scanning parameters can be set according to the specific situation of the patient. The scanning parameters cover the scanning range, slice interval, etc. The scanning range needs to fully cover the key areas of the target patient's oral and maxillofacial region, such as teeth, jawbones, surrounding soft tissues, etc., so as to ensure that complete and effective image information is obtained. The slice interval determines the distance between adjacent scanned slice image data, and its size affects the spatial resolution of the image.

[0021] During the scanning process, the oral and maxillofacial CT scanner rotates around the patient's head for scanning. Every time it rotates to a specific angle, a group of slice image data is collected. As the scanner continues to rotate, multiple groups of continuous slice image data are obtained. These slice image data are spatially continuous and show the structural information of the target patient's oral and maxillofacial region from different angles and levels. For example, by scanning in the sagittal plane, coronal plane, and transverse plane respectively, slice image data in different directions can be obtained, which together constitute the three-dimensional image data set.

[0022] From the perspective of data structure, each group of slice image data can be regarded as a two-dimensional matrix, and each element in the matrix represents the image gray value at that position. The three-dimensional image data set is a three-dimensional matrix stacked by multiple such two-dimensional matrices in the scanning order. Each element in this three-dimensional matrix has its specific spatial coordinates, corresponding to a specific position in the oral and maxillofacial region.

[0023] Step S120: Perform image feature extraction processing on the three-dimensional image data set to obtain a hierarchical image feature set of the continuous scanned slice image data; the hierarchical image feature set includes local anatomical structure features and global spatial distribution features.

[0024] After obtaining the three-dimensional image data set, in order to extract valuable information for diagnosis from it, it is necessary to perform image feature extraction processing on it. Image feature extraction processing is a complex and crucial process that can transform the original image data into a set of features with specific meanings.

[0025] Step S121: Perform spatial normalization processing on the continuous scanned slice image data to obtain a normalized three-dimensional image sequence.

[0026] Since there are differences in the spatial position, size, and orientation of the oral and maxillofacial structures of different patients, in order to facilitate subsequent unified processing and analysis, it is necessary to perform spatial normalization processing on the continuous scanned slice image data. Spatial normalization processing mainly includes two core steps: spatial registration and resampling.

[0027] The purpose of spatial registration is to align the oral and maxillofacial image data of different patients in space so that they have the same spatial coordinate system. Before performing spatial registration, a set of reference points needs to be determined first. These reference points can be some landmark anatomical structures in the oral and maxillofacial region, such as the vertices of teeth, the boundary points of the jawbone, etc. Exemplarily, two sets of feature points can be extracted from the three-dimensional image data set, denoted as feature point set A and feature point set B respectively. Feature point set A comes from the image data of the target patient, and feature point set B comes from a pre-set standard template image data.

[0028] Next, a suitable transformation matrix T needs to be found such that after the feature point set A is transformed by the transformation matrix T, it can coincide with the feature point set B as much as possible. The transformation matrix T can be determined by minimizing the distance error between the feature point set A and the feature point set B. Specifically, the iterative closest point (ICP) algorithm can be used to solve the transformation matrix T. The basic idea of the ICP algorithm is to iterate continuously. In each iteration, the nearest neighbor point of each point in the feature point set A in the feature point set B is found, and then the optimal transformation between these two sets of corresponding points is calculated until the convergence condition is met.

[0029] After obtaining the transformation matrix T, apply it to the three-dimensional image data set of the target patient to transform the spatial coordinates of each voxel, thereby realizing the spatial registration of the image data.

[0030] Resampling is an operation performed after spatial registration, and its purpose is to unify the resolution of the image data. Different scanning devices and scanning parameters may result in different resolutions of the image data, which will bring difficulties to subsequent processing. Therefore, it is necessary to resample the registered image data according to the preset resolution requirements.

[0031] Assume that the preset resolution is a specific voxel size, denoted as Δx, Δy, and Δz. During the resampling process, the value of each voxel is recalculated on the new spatial grid. The method of linear interpolation can be used to calculate the value of the new voxel. Specifically, for each voxel on the new spatial grid, find its adjacent voxels in the original image data, and then calculate the value of this voxel by linearly weighting according to the values of these adjacent voxels and their distances from this voxel. After resampling, a standardized three-dimensional image sequence is obtained, and its resolution and spatial coordinate system are unified.

[0032] Step S122: Invoke a three-dimensional convolutional neural network to perform shallow feature extraction processing on the standardized three-dimensional image sequence to obtain a set of primary image features; the set of primary image features includes edge contour features and texture distribution features.

[0033] After obtaining the standardized three-dimensional image sequence, invoke a three-dimensional convolutional neural network (3D CNN) to perform shallow feature extraction processing on it. A three-dimensional convolutional neural network is a deep learning model specifically used to process three-dimensional data, which can automatically learn the feature information in the image data.

[0034] Step S1221: Perform convolutional kernel sliding processing on the standardized three-dimensional image sequence to generate an initial convolutional feature map, and perform max-pooling processing on the initial convolutional feature map to obtain a downsampled feature map.

[0035] When performing shallow feature extraction, first perform convolutional kernel sliding processing on the standardized three-dimensional image sequence. The convolutional kernel is a three-dimensional matrix with a specific size and weights, denoted as convolutional kernel C. The size of the convolutional kernel is usually selected according to specific tasks and data characteristics. For example, 3×3×3, 5×5×5, etc. can be selected.

[0036] Slide the convolutional kernel C on the standardized three-dimensional image sequence. Each time it slides, the convolutional kernel C performs a convolutional operation with the corresponding region of the image sequence. The process of the convolutional operation is to multiply the elements of the convolutional kernel C and the corresponding region of the image sequence element by element, and then add up all the products to obtain a convolutional result. As the convolutional kernel C slides continuously, a series of convolutional results are generated, and these convolutional results are combined together to form the initial convolutional feature map.

[0037] Assume that the size of the standardized three-dimensional image sequence is H×W×D×C1 (H represents height, W represents width, D represents depth, and C1 represents the number of channels), and the size of the convolutional kernel C is h×w×d×C1×C2 (h, w, and d respectively represent the sizes of the convolutional kernel in the height, width, and depth directions, and C2 represents the number of output channels). During the convolution operation, parameters such as the stride and padding need to be considered. The stride determines the distance by which the convolutional kernel slides each time, and padding is to add a set number of zero elements around the boundaries of the image sequence to ensure that the size of the output feature map after the convolution operation meets the requirements.

[0038] After generating the initial convolutional feature map, in order to reduce the dimension of the data and improve the computational efficiency, it is necessary to perform max-pooling processing on it. Max-pooling processing is a downsampling operation that divides the initial convolutional feature map into fixed-size regions, denoted as pooling window W. For each pooling window W, the maximum value among them is taken as the pooling result of this region.

[0039] Assume that the size of the pooling window W is p×p×p and the stride is s. When performing max-pooling processing, the pooling window W is slid on the initial convolutional feature map, sliding s units each time, and the maximum value of the elements within each pooling window is taken to obtain a reduced-dimensional feature value. As the pooling window slides continuously, a reduced-dimensional feature map is generated. The size of the reduced-dimensional feature map is smaller than that of the initial convolutional feature map, but it retains the main feature information in the initial convolutional feature map.

[0040] Step S1222: Input the reduced-dimensional feature map into the residual connection module for feature compensation processing to obtain a compensated feature map; the feature compensation processing includes performing an element-wise addition operation on the reduced-dimensional feature map and the original feature map transmitted through the skip connection to generate a compensated feature map with the characteristic of retaining details.

[0041] After obtaining the reduced-dimensional feature map, in order to avoid losing some important detailed information during the feature extraction process, it is necessary to input the reduced-dimensional feature map into the residual connection module for feature compensation processing. The residual connection module is an important part of the three-dimensional convolutional neural network, which can transmit the feature information in the original standardized three-dimensional image sequence to the subsequent processing stage through the skip connection.

[0042] Suppose the dimensionality-reduced feature map is F1, with a size of H1×W1×D1×C2. When performing skip connections, the original normalized three-dimensional image sequence needs to go through a series of convolution and pooling operations to make its size the same as that of the dimensionality-reduced feature map F1, denoted as F0. Then, the dimensionality-reduced feature map F1 and the original feature map F0 passed through the skip connection are subjected to an element-wise addition operation. Specifically, for the corresponding elements in F1 and F0, they are added together to obtain a new element, and these new elements are combined to form the compensated feature map F2.

[0043] The calculation formula for the compensated feature map F2 is F2 = F1 + F0. Through this element-wise addition operation, some detailed information lost during the dimensionality reduction process is re-supplemented, making the compensated feature map have better detail retention characteristics. In this way, in subsequent processing, the model can utilize this retained detailed information to more accurately identify and analyze the structural features of the oral and maxillofacial region.

[0044] Step S1223: Perform channel attention weight assignment processing on the compensated feature map to generate a weighted feature map; the channel attention weight assignment processing includes performing global average pooling operation on the channel dimension of the compensated feature map to generate a channel description vector; inputting the channel description vector into a fully connected layer to generate channel attention weights, and re-calibrating each channel of the compensated feature map based on the channel attention weights to generate a weighted feature map.

[0045] After obtaining the compensated feature map, in order to highlight the importance of different channels in the compensated feature map, it is necessary to perform channel attention weight assignment processing on it. The channel attention weight assignment processing mainly includes two steps: global average pooling and fully connected layer mapping.

[0046] First, perform global average pooling operation on the channel dimension of the compensated feature map. The global average pooling operation will average the compensated feature map in the spatial dimension to obtain a channel description vector. Suppose the size of the compensated feature map is H2×W2×D2×C2, and the global average pooling operation will sum all the elements of each channel in the height, width, and depth directions, and then divide by the total number of elements to obtain the average value of the channel. Combining the average values of all channels together forms a channel description vector with a length of C2.

[0047] Then, input the channel description vector into a fully connected layer for mapping. The fully connected layer is a simple linear transformation layer, which can output a channel attention weight vector with a length of C2 according to the input of the channel description vector. This channel attention weight vector represents the importance degree of each channel in the compensated feature map.

[0048] Specifically, the calculation formula of the fully connected layer is: W = f(W1 * V + b1), where V is the channel description vector, W1 is the weight matrix of the fully connected layer, b1 is the bias vector, and f is the activation function, such as the Sigmoid function. Through the action of the activation function, the output of the fully connected layer is mapped to the interval [0, 1] to obtain the channel attention weight vector W.

[0049] Finally, based on the channel attention weight vector W, each channel of the compensated feature map is recalibrated. The channel attention weight vector W is multiplied element-wise with each channel of the compensated feature map to obtain the weighted feature map. Assuming the compensated feature map is F2 and the weighted feature map is F3, the calculation formula of F3 is F3 = F2 * W. In this way, the feature information of important channels is enhanced, and the feature information of unimportant channels is suppressed, making the weighted feature map more prominent in key features.

[0050] Step S1224: Extract the edge contour features and texture distribution features in the weighted feature map to form a set of primary image features; the edge contour features are extracted by a learnable edge convolution kernel for gradient feature extraction of the weighted feature map, and the texture distribution features are extracted by a multi-scale convolution kernel group for adaptive texture pattern extraction of the weighted feature map.

[0051] After obtaining the weighted feature map, in order to obtain a set of primary image features, it is necessary to extract the edge contour features and texture distribution features from the weighted feature map.

[0052] The extraction of edge contour features is achieved by using a learnable edge convolution kernel to perform gradient feature extraction on the weighted feature map. The learnable edge convolution kernel is a convolution kernel with a specific structure and weights, which can automatically learn the edge information in the image. The weights of the learnable edge convolution kernel will be continuously updated during the training process to adapt to different image data and task requirements.

[0053] Apply the learnable edge convolution kernel to the weighted feature map, and the gradient information of the weighted feature map is obtained through convolution operation. The gradient information reflects the change rate of gray values in the image, and the places with large gray value changes usually correspond to the edge contours. Specifically, the learnable edge convolution kernel slides on the weighted feature map, and each time it slides, it performs a convolution operation with the corresponding area of the weighted feature map to obtain a gradient value. As the convolution kernel slides continuously, a gradient feature map is finally generated, and this gradient feature map represents the edge contour features in the image.

[0054] The extraction of texture distribution features is achieved by using a multi-scale convolution kernel group to perform adaptive texture pattern extraction on the weighted feature map. The multi-scale convolution kernel group contains multiple convolution kernels of different sizes and scales, which can capture texture information at different scales.

[0055] Suppose the multi-scale convolution kernel group contains three convolution kernels of different sizes, denoted as convolution kernel K1, convolution kernel K2, and convolution kernel K3 respectively, and their sizes increase in sequence. Apply these three convolution kernels to the weighted feature map respectively, and through convolution operations, texture feature maps at three different scales are obtained, denoted as texture feature map F4, texture feature map F5, and texture feature map F6 respectively.

[0056] Convolution kernels of different sizes can capture texture information of different thicknesses and complexities. Smaller convolution kernels can capture texture information with rich details, while larger convolution kernels can capture macroscopic texture patterns. Combining these three texture feature maps together forms a texture distribution feature.

[0057] Finally, combine the extracted edge contour features and texture distribution features together to form a primary image feature set, and the primary image feature set contains important feature information such as edge contours and texture distributions in the oral and maxillofacial image.

[0058] Step S123: Perform cross-layer feature enhancement processing on the primary image feature set to obtain an enhanced image feature set; the enhanced image feature set contains local anatomical structure features with enhanced spatial resolution.

[0059] After obtaining the primary image feature set, in order to further enhance the feature information therein and improve the expression ability of the features, it is necessary to perform cross-layer feature enhancement processing on the primary image feature set. Cross-layer feature enhancement processing can fuse feature information at different levels, so as to obtain local anatomical structure features with enhanced spatial resolution.

[0060] Step S1231: Perform feature splicing processing on the edge contour features and texture distribution features in the primary image feature set to generate a spliced feature map.

[0061] In order to make full use of the edge contour features and texture distribution features in the primary image feature set, it is first necessary to perform feature splicing processing on them. Feature splicing processing is to splice the edge contour features and texture distribution features in the channel dimension.

[0062] Suppose the size of the edge contour feature is H3×W3×D3×C3, and the size of the texture distribution feature is H3×W3×D3×C4. When performing feature splicing, the edge contour feature and the texture distribution feature are concatenated in the channel dimension, and the size of the spliced feature map obtained is H3×W3×D3×(C3 + C4). Through feature splicing processing, different types of feature information can be integrated together, providing a richer feature representation for subsequent processing. In this way, the spliced feature map contains information on both the edge contour and the texture distribution, and can more comprehensively describe the local anatomical structure of the oral and maxillofacial region.

[0063] Step S1232: Perform dilated convolution processing on the spliced feature map to generate a multi-scale receptive field feature map; the dilated convolution processing uses convolution kernels with different dilation rates to act on the spliced feature map in parallel, generating feature sub-maps with different receptive field ranges; the feature sub-maps are merged in the channel dimension to generate a multi-scale receptive field feature map, and the multi-scale receptive field feature map is subjected to feature pyramid fusion processing to obtain a multi-scale fusion feature map; the feature pyramid fusion processing includes upsampling or downsampling the feature sub-maps with different scales to a unified resolution and performing element-wise addition operations.

[0064] After obtaining the spliced feature map, in order to obtain feature information at different scales, it is necessary to perform dilated convolution processing on it. The dilated convolution processing uses convolution kernels with different dilation rates to act on the spliced feature map in parallel. The dilation rate represents the distance between elements in the convolution kernel, and different dilation rates can make the convolution kernel have different receptive field ranges.

[0065] Suppose three convolution kernels with different dilation rates are used, denoted as convolution kernel K4, convolution kernel K5, and convolution kernel K6 respectively, and their dilation rates are r1, r2, and r3 (r1 < r2 < r3). These three convolution kernels are respectively applied to the spliced feature map, and three feature sub-maps with different receptive field ranges are obtained through convolution operations, denoted as feature sub-map F7, feature sub-map F8, and feature sub-map F9 respectively.

[0066] Convolution kernel K4 has a smaller dilation rate r1, and its receptive field range is relatively small, which can capture the detailed information in the spliced feature map. Convolution kernel K5 has a medium dilation rate r2, and its receptive field range is moderate, which can capture some local structural information. Convolution kernel K6 has a larger dilation rate r3, and its receptive field range is large, which can capture more macroscopic feature information.

[0067] Then, these three feature sub - graphs are merged in the channel dimension to generate a multi - scale receptive field feature map. Assume that the size of feature sub - graph F7 is H4×W4×D4×C5, the size of feature sub - graph F8 is H4×W4×D4×C6, and the size of feature sub - graph F9 is H4×W4×D4×C7. The size of the merged multi - scale receptive field feature map is H4×W4×D4×(C5 + C6 + C7).

[0068] Next, perform feature pyramid fusion processing on the multi - scale receptive field feature map. The purpose of feature pyramid fusion processing is to fuse feature sub - graphs of different scales so that they have a unified resolution. First, up - sampling or down - sampling operations need to be performed on feature sub - graphs of different scales. For feature sub - graphs with a smaller size, methods such as bilinear interpolation can be used for up - sampling to make their size the same as that of other feature sub - graphs; for feature sub - graphs with a larger size, methods such as average pooling can be used for down - sampling to make their size the same as that of other feature sub - graphs.

[0069] Assume that after the up - sampling or down - sampling operation, feature sub - graph F7, feature sub - graph F8, and feature sub - graph F9 all have the same size H5×W5×D5×C8. Then, perform an element - wise addition operation on these three feature sub - graphs to obtain a multi - scale fusion feature map F10. The element - wise addition operation adds the elements at the corresponding positions in the three feature sub - graphs to obtain the element value at that position in the multi - scale fusion feature map. Through this feature pyramid fusion processing, feature information of different scales is effectively fused, and the multi - scale fusion feature map can comprehensively reflect feature information at different scales, providing a more comprehensive feature representation for subsequent processing.

[0070] Step S1233: Perform non - linear activation processing on the multi - scale fusion feature map to generate an activation feature map; the non - linear activation processing uses an activation function with a gating mechanism to dynamically activate each channel of the multi - scale fusion feature map.

[0071] After obtaining the multi - scale fusion feature map, in order to introduce non - linear factors and improve the expression ability of the model, non - linear activation processing needs to be performed on it. Here, an activation function with a gating mechanism is used to dynamically activate each channel of the multi - scale fusion feature map.

[0072] The activation function with a gating mechanism can dynamically adjust the activation degree of each channel according to the input feature information. Assume that the multi - scale fusion feature map is F10, and its size is H5×W5×D5×C8. The activation function with a gating mechanism will calculate a gating value for each channel, and this gating value indicates whether the channel should be activated and the degree of activation.

[0073] Specifically, the activation function with a gating mechanism first performs a linear transformation on each channel of the multi-scale fusion feature map to obtain an intermediate result. Then, through a non-linear function, such as the Sigmoid function, the intermediate result is mapped to the interval [0, 1] to obtain the gating value. Finally, the gating value is multiplied by the original value of the channel to obtain the channel value after activation processing.

[0074] Assume that for the i-th channel of the multi-scale fusion feature map F10, its original value is F10(i), the intermediate result obtained through linear transformation is M(i), the gating value is G(i), and the channel value after activation processing is A(i). Then the calculation process is as follows: First, M(i) is obtained through linear transformation, and the linear transformation can be expressed as M(i) = W2 * F10(i) + b2, where W2 is the weight matrix of the linear transformation and b2 is the bias vector. Then, the gating value G(i) = Sigmoid(M(i)) is calculated through the Sigmoid function. Finally, the activated channel value A(i) = G(i) * F10(i) is obtained. Performing such processing on all channels of the multi-scale fusion feature map F10 yields the activated feature map F11. Through this non-linear activation processing, the model can learn more complex feature representations and enhance the model's ability to express oral and maxillofacial imaging features.

[0075] Step S1234: Perform spatial attention weight assignment processing on the activated feature map to obtain an enhanced image feature set; the spatial attention weight assignment processing includes performing a dual-path operation of max pooling and average pooling on the spatial dimension of the activated feature map to generate a dual-path pooled feature map; concatenating the channels of the dual-path pooled feature map and then generating a spatial attention weight map through a convolutional layer, and weighting the spatial dimension of the activated feature map based on the spatial attention weight map to generate an enhanced image feature set.

[0076] After obtaining the activated feature map, in order to highlight the importance of different spatial positions in the activated feature map, it is necessary to perform spatial attention weight assignment processing on it. The spatial attention weight assignment processing mainly includes three steps: dual-path pooling, channel concatenation, and convolution to generate a weight map.

[0077] First, perform a dual-path operation of max pooling and average pooling on the spatial dimension of the activated feature map. The max pooling operation divides the spatial dimension of the activated feature map into fixed-size regions, denoted as pooling windows W1. For each pooling window W1, the maximum value in it is taken as the pooling result of this region. The average pooling operation takes the average value of this region as the pooling result.

[0078] Assume that the size of the activation feature map F11 is H5×W5×D5×C8, the size of the pooling window W1 is p1×p1×p1, and the stride is s1. When performing the max pooling operation, the pooling window W1 slides on the activation feature map F11, sliding s1 units each time, and taking the maximum value of the elements in each pooling window to obtain the max pooling feature map F12. When performing the average pooling operation, similarly, the pooling window W1 slides on the activation feature map F11, sliding s1 units each time, and taking the average value of the elements in each pooling window to obtain the average pooling feature map F13.

[0079] Then, the max pooling feature map F12 and the average pooling feature map F13 are concatenated in the channel dimension to obtain the dual-path pooling feature map F14. Assume that the size of the max pooling feature map F12 is H6×W6×D6×C8, the size of the average pooling feature map F13 is H6×W6×D6×C8, and the size of the concatenated dual-path pooling feature map F14 is H6×W6×D6×(2*C8).

[0080] Next, the dual-path pooling feature map F14 is processed through a convolutional layer to generate the spatial attention weight map F15. The convolutional layer performs a convolution operation on the dual-path pooling feature map F14 to learn the importance information of different spatial positions. Assume that the convolution kernel of the convolutional layer is K7, and its size is h1×w1×d1×(2*C8)×1. Through the convolution operation, the dual-path pooling feature map F14 is converted into a spatial attention weight map F15 with a size of H6×W6×D6×1.

[0081] Finally, the activation feature map F11 is weighted in the spatial dimension based on the spatial attention weight map F15. The spatial attention weight map F15 and the activation feature map F11 are subjected to an element-wise multiplication operation to obtain the enhanced image feature set F16. For each element at each position in the activation feature map F11, it is multiplied by the element at the corresponding position in the spatial attention weight map F15 to obtain the element value at that position in the enhanced image feature set F16. Through this spatial attention weight distribution process, the feature information of important spatial positions in the activation feature map is enhanced, and the feature information of unimportant spatial positions is suppressed, making the enhanced image feature set better able to highlight the key local anatomical structure features.

[0082] Step S124: Perform global pooling processing on the enhanced image feature set to obtain the global spatial distribution features.

[0083] After obtaining the enhanced image feature set, in order to obtain the global information of the image data, it is necessary to perform global pooling processing on it. Global pooling processing can compress the enhanced image feature set in the spatial dimension to obtain a global feature that can represent the entire image data.

[0084] Step S1241: Perform channel - dimension average pooling on the local anatomical structure features in the enhanced image feature set to generate a channel - statistical feature vector.

[0085] First, perform channel - dimension average pooling on the local anatomical structure features in the enhanced image feature set. Assume the size of the enhanced image feature set F16 is H7×W7×D7×C9. Channel - dimension average pooling sums all the elements in each channel in the height, width, and depth directions, and then divides by the total number of elements to obtain the average value of that channel.

[0086] Specifically, for the j - th channel of the enhanced image feature set F16, the sum of all its elements is Sum(j), and the total number of elements is H7*W7*D7. Then the average value of this channel is Avg(j)=Sum(j) / (H7*W7*D7). Combining the average values of all channels forms a channel - statistical feature vector V1 with a length of C9. The channel - statistical feature vector V1 reflects the average feature information of each channel in the enhanced image feature set, providing a concise feature representation for subsequent processing.

[0087] Step S1242: Perform a fully - connected layer mapping on the channel - statistical feature vector to generate a dimensionality - reduced feature vector.

[0088] After obtaining the channel - statistical feature vector, in order to further reduce the dimensionality of the data, a fully - connected layer mapping needs to be performed on it. The fully - connected layer is a simple linear transformation layer, which can output a dimensionality - reduced feature vector according to the input of the channel - statistical feature vector.

[0089] Assume the length of the channel - statistical feature vector V1 is C9, the weight matrix of the fully - connected layer is W3 with a size of C9×C10 (C10 < C9), and the bias vector is b3. The mapping calculation of the fully - connected layer can be expressed as V2 = W3*V1 + b3, where V2 is the dimensionality - reduced feature vector. Through the fully - connected layer mapping process, the information in the channel - statistical feature vector is compressed and integrated to obtain a dimensionality - lower dimensionality - reduced feature vector V2. The dimensionality - reduced feature vector V2 retains the main information in the channel - statistical feature vector, while reducing data redundancy and improving computational efficiency.

[0090] Step S1243: Map the dimensionality - reduced feature vector through a fully - connected layer to be consistent with the channel dimension of the enhanced image feature set, and perform feature broadcasting processing to generate a spatial weight distribution map; the feature broadcasting processing includes copying and expanding the mapped feature vector to the same spatial size as the enhanced image feature set.

[0091] After obtaining the dimensionality-reduced feature vector, it is necessary to map it through a fully-connected layer to be consistent with the channel dimension of the enhanced image feature set. Assume that the length of the dimensionality-reduced feature vector V2 is C10, and the number of channels of the enhanced image feature set F16 is C9. The weight matrix of the fully-connected layer is W4, with its size being C10×C9, and the bias vector is b4. Through the mapping calculation of the fully-connected layer, the mapped feature vector V3 = W4 * V2 + b4 is obtained, and its length is C9.

[0092] Then, feature broadcasting processing is performed. Feature broadcasting processing is to copy and expand the mapped feature vector V3 to the same spatial size as the enhanced image feature set F16. Assume that the size of the enhanced image feature set F16 is H7×W7×D7×C9. Each element in the mapped feature vector V3 is copied H7*W7*D7 times and filled into a space with a size of H7×W7×D7×C9 to obtain the spatial weight distribution map F17. The spatial weight distribution map F17 assigns a weight value to each position in the enhanced image feature set F16, reflecting the importance of that position in the global space.

[0093] Step S1244: Based on the spatial weight distribution map, perform channel-dimensional weighting and spatial-dimensional aggregation processing on the enhanced image feature set to obtain the global spatial distribution feature.

[0094] After obtaining the spatial weight distribution map, perform channel-dimensional weighting and spatial-dimensional aggregation processing on the enhanced image feature set based on it. First, perform channel-dimensional weighting by performing an element-wise multiplication operation on the spatial weight distribution map F17 and the enhanced image feature set F16. For each element at each position in the enhanced image feature set F16, multiply it by the corresponding element in the spatial weight distribution map F17 to obtain the weighted enhanced image feature set F18.

[0095] Then, perform spatial-dimensional aggregation processing. Spatial-dimensional aggregation processing can be performed using methods such as average pooling or summation. Here, taking average pooling as an example, perform average pooling operations in the height, width, and depth directions of the weighted enhanced image feature set F18. Assume that the size of the weighted enhanced image feature set F18 is H7×W7×D7×C9. Sum all the elements in the height, width, and depth directions for each channel, and then divide by the total number of elements to obtain the average value of that channel. Combine the average values of all channels together to form the global spatial distribution feature V4. The global spatial distribution feature V4 is a vector with a length of C9, which synthesizes the information of the enhanced image feature set in the global space and provides an important global feature representation for subsequent feature fusion and lesion area recognition.

[0096] Step S125: Perform feature alignment processing on the local anatomical structure features and the global spatial distribution features to form a hierarchical image feature set; the feature alignment processing includes performing a spatial coordinate mapping operation on the local anatomical structure features to generate a spatial alignment feature vector corresponding to the global spatial distribution features; performing a feature broadcasting operation on the global spatial distribution features to expand its spatial dimension to be consistent with the spatial alignment feature vector and adjust the channel dimension to match; performing channel dimension splicing on the expanded global spatial distribution features and the spatial alignment feature vector to generate a hierarchical image feature set with a unified spatial dimension.

[0097] After obtaining the local anatomical structure features (i.e., the enhanced image feature set F16) and the global spatial distribution feature V4, in order to effectively fuse them together, feature alignment processing needs to be performed to form a hierarchical image feature set.

[0098] First, perform a spatial coordinate mapping operation on the local anatomical structure features. The spatial coordinate mapping operation is to map the spatial coordinates of each element in the local anatomical structure features to a unified coordinate system, so that its spatial coordinates are consistent with the spatial coordinates corresponding to the global spatial distribution features. Assume that the spatial coordinate system of the local anatomical structure feature F16 is coordinate system A, and the spatial coordinate system corresponding to the global spatial distribution feature V4 is coordinate system B. Through a coordinate transformation matrix T1, the spatial coordinates of each element in the local anatomical structure feature F16 are transformed from coordinate system A to coordinate system B to obtain a spatial alignment feature vector F19. The spatial coordinates of the spatial alignment feature vector F19 are consistent with the spatial coordinates corresponding to the global spatial distribution feature V4.

[0099] Then, perform a feature broadcasting operation on the global spatial distribution features. The feature broadcasting operation is to expand the spatial dimension of the global spatial distribution feature V4 to be consistent with the spatial alignment feature vector F19 and adjust the channel dimension to match. Assume that the global spatial distribution feature V4 is a vector of length C9, and the size of the spatial alignment feature vector F19 is H7×W7×D7×C9. Each element in the global spatial distribution feature V4 is copied H7*W7*D7 times and filled into a space of size H7×W7×D7×C9 to obtain the expanded global spatial distribution feature F20.

[0100] Finally, the extended global spatial distribution feature F20 and the spatially aligned feature vector F19 are concatenated in the channel dimension. Channel dimension concatenation means connecting the extended global spatial distribution feature F20 and the spatially aligned feature vector F19 in the channel dimension. Suppose the size of the extended global spatial distribution feature F20 is H7×W7×D7×C9, and the size of the spatially aligned feature vector F19 is H7×W7×D7×C9. The size of the concatenated hierarchical image feature set F21 is H7×W7×D7×(2*C9). Through this channel dimension concatenation, the local anatomical structure features and the global spatial distribution features are effectively fused together to form a hierarchical image feature set F21 with a unified spatial dimension. The hierarchical image feature set F21 synthesizes local and global feature information, providing a more comprehensive and richer feature representation for subsequent multi-scale feature fusion and lesion area recognition.

[0101] Step S130: Invoke a pre-trained multi-scale feature fusion network to perform multi-scale feature fusion on the hierarchical image feature set to generate a fused feature map.

[0102] After obtaining the hierarchical image feature set, in order to further mine and integrate the feature information therein, a pre-trained multi-scale feature fusion network is invoked to perform multi-scale feature fusion processing on it to generate a fused feature map. The multi-scale feature fusion network can analyze and fuse the hierarchical image feature set from different scales and angles, so that the fused feature map can more comprehensively and accurately reflect the structure and lesion information of the oral and maxillofacial region.

[0103] Step S131: Perform feature concatenation processing on the local anatomical structure features and the global spatial distribution features to generate an initial concatenated feature.

[0104] To make full use of the local anatomical structure features and the global spatial distribution features in the hierarchical image feature set, they are first subjected to feature concatenation processing. Suppose the local anatomical structure feature is F16 (with a size of H7×W7×D7×C9), and the global spatial distribution feature after feature broadcast operation is F20 (with a size of H7×W7×D7×C9). When performing feature concatenation, the local anatomical structure feature F16 and the global spatial distribution feature F20 are connected in the channel dimension to obtain the initial concatenated feature F22, whose size is H7×W7×D7×(2*C9). Through feature concatenation processing, local and global feature information is integrated together, providing a richer feature input for subsequent multi-scale feature fusion.

[0105] Step S132: Perform grouped axial self-attention mechanism processing on the initial splicing feature to generate a self-attention weight distribution map; the grouped axial self-attention mechanism processing includes decomposing the feature along the spatial dimension into axial sub-regions, respectively calculating the correlation matrix between spatial positions within each axial sub-region, and performing normalization processing on the correlation matrix to generate a local self-attention weight distribution map.

[0106] After obtaining the initial splicing feature, perform grouped axial self-attention mechanism processing on it to mine the correlation information between different spatial positions in the feature. The grouped axial self-attention mechanism processing mainly includes three steps: feature decomposition, correlation matrix calculation, and normalization processing.

[0107] First, decompose the initial splicing feature F22 along the spatial dimension into axial sub-regions. Assume that the size of the initial splicing feature F22 is H7×W7×D7×(2*C9). It can be divided in the height, width, and depth directions respectively to obtain multiple axial sub-regions. For example, divide it into m intervals in the height direction, n intervals in the width direction, and p intervals in the depth direction, thus obtaining m*n*p axial sub-regions. The size of each axial sub-region is (H7 / m)×(W7 / n)×(D7 / p)×(2*C9).

[0108] Then, calculate the correlation matrix between spatial positions within each axial sub-region respectively. For each axial sub-region, expand it into a two-dimensional matrix. The number of rows of the matrix is the number of spatial positions within the axial sub-region, and the number of columns is the number of channels of the feature. Then, calculate the similarity between different rows in this two-dimensional matrix to obtain the correlation matrix. The calculation of similarity can use methods such as dot product, cosine similarity, etc. Assume that the two-dimensional matrix after expansion of a certain axial sub-region is M, with the number of rows r and the number of columns (2*C9). The element R(i, j) of the correlation matrix R represents the similarity between the i-th row and the j-th row.

[0109] Finally, perform normalization processing on the correlation matrix to generate a local self-attention weight distribution map. The purpose of normalization processing is to map the elements in the correlation matrix to the interval [0, 1], so that each element represents the importance weight of this position within the axial sub-region. The Softmax function can be used to perform normalization processing on the correlation matrix. For each row of the correlation matrix R, calculate the normalized value of each element in this row through the Softmax function to obtain the local self-attention weight distribution map. Combine the local self-attention weight distribution maps of all axial sub-regions to obtain the self-attention weight distribution map F23 of the entire initial splicing feature. The size of the self-attention weight distribution map F23 is the same as that of the initial splicing feature F22, which reflects the importance correlation information between different spatial positions in the initial splicing feature.

[0110] Step S133: Dynamically weight the local anatomical structure features based on the self-attention weight distribution map to generate weighted local features; the dynamic weighting process includes performing an element-wise multiplication operation on the self-attention weight distribution map and the local anatomical structure features.

[0111] After obtaining the self-attention weight distribution map, perform dynamic weighting on the local anatomical structure features based on it. The local anatomical structure features have been obtained in the previous steps, denoted as F16, and the self-attention weight distribution map is F23. The dynamic weighting process is to perform an element-wise multiplication operation on the self-attention weight distribution map F23 and the local anatomical structure features F16.

[0112] For each element at each position in the local anatomical structure features F16, multiply it by the corresponding element in the self-attention weight distribution map F23 to obtain the weighted element value. For example, assume the element value at position (x, y, z, c) in the local anatomical structure features F16 is F16(x, y, z, c), and the element value at position (x, y, z, c) in the self-attention weight distribution map F23 is F23(x, y, z, c). Then the weighted element value is F16(x, y, z, c) * F23(x, y, z, c). Perform such processing on all elements at all positions in the local anatomical structure features F16 to obtain the weighted local features F24.

[0113] Through this dynamic weighting process, the feature information at important spatial positions in the local anatomical structure features is enhanced, and the feature information at unimportant spatial positions is suppressed. Since the self-attention weight distribution map reflects the relative importance of the associations between different spatial positions, the weighted local features can focus more on the key local anatomical structure features.

[0114] Step S134: Perform feature fusion on the weighted local features and the global spatial distribution features to generate intermediate fusion features.

[0115] After obtaining the weighted local features, perform feature fusion on them and the global spatial distribution features. The global spatial distribution features after feature broadcasting operation are F20, and the weighted local features are F24. The feature fusion process can be carried out in various ways. Here, a combination of concatenation and weighted addition is adopted.

[0116] First, concatenate the weighted local features F24 and the global spatial distribution features F20 along the channel dimension to obtain the concatenated features F25. Assume the size of the weighted local features F24 is H7×W7×D7×C9, and the size of the global spatial distribution features F20 is H7×W7×D7×C9. The size of the concatenated features F25 after concatenation is H7×W7×D7×(2*C9).

[0117] Then, perform a weighted summation operation on the spliced feature F25. To balance the contributions of the weighted local feature and the global spatial distribution feature, weights w1 and w2 are assigned to them respectively, and w1 + w2 = 1. For each element in each channel of the spliced feature F25, multiply it by the corresponding weight according to whether it belongs to the weighted local feature or the global spatial distribution feature part, and then sum the results.

[0118] For example, for the elements in the first C9 channels (corresponding to the weighted local feature F24) of the spliced feature F25, multiply by the weight w1; for the elements in the last C9 channels (corresponding to the global spatial distribution feature F20), multiply by the weight w2. Assume that the element value at position (x, y, z, c) in the spliced feature F25 is F25(x, y, z, c). If c < C9, the weighted element value is w1 * F25(x, y, z, c); if c >= C9, the weighted element value is w2 * F25(x, y, z, c). Process all elements in this way to obtain the intermediate fusion feature F26.

[0119] Through this feature fusion process, the information of the weighted local feature and the global spatial distribution feature is effectively integrated. The intermediate fusion feature F26 contains both the key local anatomical structure information and the global spatial distribution information, providing a more comprehensive feature representation for subsequent cross-channel feature interaction processing.

[0120] Step S135: Perform cross-channel feature interaction processing on the intermediate fusion feature to generate a fusion feature map; the cross-channel feature interaction processing includes decomposing the intermediate fusion feature into multiple channel groups, performing independent convolution operations on each channel group, and then rearranging the channel order to generate a fusion feature map with channel interaction characteristics.

[0121] After obtaining the intermediate fusion feature, to further promote information interaction between different channels, perform cross-channel feature interaction processing on it to generate a fusion feature map. The cross-channel feature interaction processing mainly includes three steps: channel grouping, independent convolution, and channel rearrangement.

[0122] Step S1351: Divide the intermediate fusion feature into multiple sub-feature groups along the channel dimension; where each sub-feature group contains a predetermined number of consecutive channels.

[0123] The size of the intermediate fusion feature F26 is H7×W7×D7×(2*C9). Divide it along the channel dimension, and each sub-feature group contains a predetermined number of consecutive channels. Assume that the channel dimension of the intermediate fusion feature F26 is evenly divided into k sub-feature groups, and each sub-feature group contains (2*C9) / k consecutive channels.

[0124] For example, the first sub - feature group contains channels 1 to (2*C9) / k, the second sub - feature group contains channels (2*C9) / k + 1 to 2*(2*C9) / k, and so on. In this way, k sub - feature groups are obtained, denoted as F26_1, F26_2,..., F26_k respectively. The size of each sub - feature group is H7×W7×D7×((2*C9) / k). Through channel splitting, the channel information of the intermediate fusion feature is grouped.

[0125] Step S1352: Perform grouped convolution processing on each of the sub - feature groups to generate grouped feature maps; the grouped convolution processing uses independent convolution kernels to act on each sub - feature group respectively.

[0126] After obtaining the sub - feature groups, perform grouped convolution processing on each sub - feature group. The grouped convolution processing uses independent convolution kernels to act on each sub - feature group respectively. Assume that for each sub - feature group F26_i (i = 1, 2,..., k), a separate convolution kernel K_i is used for convolution operation.

[0127] The size of the convolution kernel K_i is h2×w2×d2×((2*C9) / k)×C11, where h2, w2, and d2 respectively represent the sizes of the convolution kernel in the height, width, and depth directions, and C11 represents the number of output channels. Slide the convolution kernel K_i on the sub - feature group F26_i. Each time it slides, the convolution kernel K_i performs a convolution operation with the corresponding region of the sub - feature group F26_i to obtain a convolution result. As the convolution kernel K_i slides continuously, a series of convolution results are generated, and these convolution results combined together form the grouped feature map F27_i.

[0128] The size of the grouped feature map F27_i is H8×W8×D8×C11, where H8, W8, and D8 are determined according to parameters such as the stride and padding of the convolution operation. Through grouped convolution processing, each sub - feature group performs independent feature extraction, and the information between different sub - feature groups does not interact at this stage, but the channel information within each sub - feature group is further mined and integrated.

[0129] Step S1353: Perform channel shuffle processing on the grouped feature maps to generate shuffled feature maps; the channel shuffle processing rearranges the channels in each sub - feature group according to a predetermined rule to promote cross - group information interaction.

[0130] After obtaining the grouped feature maps, in order to promote information interaction between different sub - feature groups, perform channel shuffle processing on the grouped feature maps. The channel shuffle processing rearranges the channels in each sub - feature group according to a predetermined rule.

[0131] First, k grouped feature maps F27_1, F27_2, ..., F27_k are concatenated in the channel dimension to obtain the concatenated feature map F28 with a size of H8×W8×D8×(k*C11). Then, the channels of the concatenated feature map F28 are rearranged according to a predetermined rule.

[0132] For example, the channels are arranged by extracting and recombining every k channels. Assume the channel numbers of the concatenated feature map F28 range from 1 to k*C11. Channels 1, k + 1, 2*k + 1, ... are taken as a new group of channels, channels 2, k + 2, 2*k + 2, ... are taken as another group of channels, and so on. After such rearrangement, the shuffled feature map F29 is obtained.

[0133] Channel shuffling breaks the channel boundaries of the original sub - feature groups, enabling the channel information between different sub - feature groups to interact. In this way, the information between the originally independently processed sub - feature groups is fused.

[0134] Step S1354: Perform point - wise convolution processing on the shuffled feature map to generate cross - channel interaction features; the point - wise convolution processing uses a 1×1 convolution kernel to adjust the channel dimension.

[0135] After obtaining the shuffled feature map, perform point - wise convolution processing on it to generate cross - channel interaction features. The point - wise convolution processing is performed using a 1×1 convolution kernel. The role of the 1×1 convolution kernel is to adjust the channel dimension and promote information interaction between different channels.

[0136] Assume the size of the shuffled feature map F29 is H8×W8×D8×(k*C11). Use a 1×1 convolution kernel K8 with a size of 1×1×1×(k*C11)×C12 to perform point - wise convolution processing. Slide the convolution kernel K8 on the shuffled feature map F29. Each time it slides, the convolution kernel K8 performs a convolution operation with the corresponding region of the shuffled feature map F29.

[0137] Since the size of the convolution kernel K8 is 1×1×1, it only operates on the channel dimension of the shuffled feature map F29. Through the convolution operation, the channel dimension of the shuffled feature map F29 is adjusted from (k*C11) to C12, obtaining the cross - channel interaction feature F30 with a size of H8×W8×D8×C12.

[0138] Point - wise convolution processing enables the linear combination of information between different channels, further promoting cross - channel information interaction. By adjusting the channel dimension, feature redundancy is reduced while important information between different channels is retained.

[0139] Step S1355: Perform residual connection processing on the cross-channel interaction feature and the intermediate fusion feature to generate a fused feature map; the residual connection processing adds the cross-channel interaction feature and the original intermediate fusion feature element-wise.

[0140] After obtaining the cross-channel interaction feature, perform residual connection processing on it and the intermediate fusion feature to generate a fused feature map. The purpose of the residual connection processing is to integrate the information of the cross-channel interaction feature and the intermediate fusion feature, while avoiding losing the original intermediate fusion feature information during the feature fusion process.

[0141] First, the intermediate fusion feature F26 needs to go through a series of operations to make its size the same as that of the cross-channel interaction feature F30. The intermediate fusion feature F26 can be adjusted through operations such as convolution and pooling to obtain the adjusted intermediate fusion feature F31, whose size is H8×W8×D8×C12.

[0142] Then, perform an element-wise addition operation on the cross-channel interaction feature F30 and the adjusted intermediate fusion feature F31. For the element value F30(x, y, z, c) at position (x, y, z, c) in the cross-channel interaction feature F30 and the element value F31(x, y, z, c) at position (x, y, z, c) in the adjusted intermediate fusion feature F31, add them to obtain the element value F32(x, y, z, c) at position (x, y, z, c) in the fused feature map F32, where F32(x, y, z, c) = F30(x, y, z, c) + F31(x, y, z, c).

[0143] Through the residual connection processing, the fused feature map F32 contains both the new cross-channel information brought by the cross-channel interaction feature and retains the original information of the intermediate fusion feature. This integration of information enables the fused feature map to more comprehensively and accurately reflect the feature information of the oral and maxillofacial images, providing high-quality feature input for subsequent lesion area recognition processing.

[0144] Step S140: Perform lesion area recognition processing based on the fused feature map to determine the location information and morphological description information of the oral and maxillofacial abnormal area of the target patient.

[0145] After obtaining the fused feature map, perform lesion area recognition processing based on it to determine the location information and morphological description information of the oral and maxillofacial abnormal area of the target patient. Lesion area recognition processing is a key link in the entire image recognition and diagnosis method, which can find the areas where lesions may exist from the fused feature map and accurately describe the locations and morphologies of these areas.

[0146] Step S141: Perform candidate region generation processing on the fused feature map to obtain multiple candidate region bounding boxes; the candidate region generation processing uses a sliding window mechanism to generate initial bounding boxes on the fused feature map at different scales and aspect ratios.

[0147] To find the regions in the fused feature map that may contain lesions, first perform candidate region generation processing on it. The candidate region generation processing uses a sliding window mechanism. The sliding window is a rectangular box with a set size and shape, which slides on the fused feature map.

[0148] The size of the fused feature map F32 is H8×W8×D8×C12. In the sliding window mechanism, sliding windows with different scales and aspect ratios are used to slide on the fused feature map. For example, set different window sizes, such as small-sized windows, medium-sized windows, and large-sized windows, and each size of window has different aspect ratios, such as 1:1, 2:1, 1:2, etc.

[0149] For each sliding window, it slides on the fused feature map at a set step size. Each time it slides, record the position of the window on the fused feature map, and this position determines an initial bounding box. As the sliding window keeps sliding, multiple initial bounding boxes will be generated. These initial bounding boxes cover different regions of the fused feature map, and each bounding box represents a candidate region that may contain lesions. Combine all the generated initial bounding boxes together to obtain multiple candidate region bounding boxes.

[0150] Step S142: Perform feature cropping processing on each of the candidate region bounding boxes to generate candidate region feature maps; the feature cropping processing intercepts the feature data of the corresponding region from the fused feature map according to the bounding box coordinates.

[0151] After obtaining multiple candidate region bounding boxes, perform feature cropping processing on each candidate region bounding box. The purpose of the feature cropping processing is to extract the feature data corresponding to each candidate region from the fused feature map to generate candidate region feature maps.

[0152] For each candidate region bounding box, its coordinate information determines its position on the fused feature map. According to the coordinates of the bounding box, intercept the feature data of the corresponding region from the fused feature map F32. For example, assume that the upper left corner coordinates of a certain candidate region bounding box are (x1, y1, z1), and the lower right corner coordinates are (x2, y2, z2), then intercept the feature data within the range of x1 to x2, y1 to y2, and z1 to z2 from the fused feature map F32 to obtain a sub-feature map.

[0153] Use this sub - feature map as the candidate region feature map for the candidate region. Such feature clipping processing is performed on all candidate region bounding boxes to obtain multiple candidate region feature maps. Each candidate region feature map contains the feature information of the corresponding candidate region, providing specific feature inputs for subsequent abnormal probability prediction processing.

[0154] Step S143: Input the candidate region feature map into a pre - trained lesion classification network for abnormal probability prediction processing to obtain the abnormal confidence score for each candidate region; the lesion classification network includes multiple fully - connected layers for mapping the candidate region feature map to an abnormal probability value.

[0155] After obtaining the candidate region feature maps, input them into a pre - trained lesion classification network for abnormal probability prediction processing. The lesion classification network is a trained deep - learning model that can predict whether there is an abnormality in the region and the probability of the abnormality based on the input candidate region feature map.

[0156] The lesion classification network includes multiple fully - connected layers. A fully - connected layer is a simple linear transformation layer that maps the input feature vector to a new feature space. For each candidate region feature map, first expand it into a one - dimensional vector. Assume the size of the candidate region feature map is H9×W9×D9×C12, then expand it into a one - dimensional vector with a length of H9*W9*D9*C12.

[0157] Then, input this one - dimensional vector into the first fully - connected layer of the lesion classification network. The weight matrix of the first fully - connected layer performs a linear transformation on the input one - dimensional vector to obtain a new feature vector. Then, input this new feature vector into the next fully - connected layer, and so on, through the processing of multiple fully - connected layers.

[0158] The output of the last fully - connected layer is a vector with a length of 1. The value of this vector is mapped to the interval [0, 1] through the Sigmoid function to obtain the abnormal probability value of the candidate region. This abnormal probability value is the abnormal confidence score of the candidate region. Such abnormal probability prediction processing is performed on all candidate region feature maps to obtain the abnormal confidence score for each candidate region.

[0159] Step S144: Screen the candidate regions corresponding to the abnormal confidence scores according to a preset confidence threshold to generate a candidate set of abnormal regions.

[0160] After obtaining the abnormal confidence score for each candidate region, screen these candidate regions according to a preset confidence threshold to generate a candidate set of abnormal regions. The preset confidence threshold is a pre - set value used to determine whether a candidate region is likely to be an abnormal region.

[0161] For the abnormal confidence score of each candidate region, if the score is greater than the preset confidence threshold, it is considered that the candidate region may be an abnormal region, and it is added to the candidate set of abnormal regions; if the score is less than or equal to the preset confidence threshold, it is considered that the candidate region is unlikely to be an abnormal region and is excluded.

[0162] Through this screening operation, candidate regions with higher abnormal confidence scores are screened out from all candidate regions, forming a candidate set of abnormal regions. The candidate set of abnormal regions contains candidate regions where lesions may exist, reducing the workload of subsequent processing and improving the efficiency of lesion region identification.

[0163] Step S145: Sort each candidate region in the candidate set of abnormal regions in descending order according to the abnormal confidence score to generate a sorted candidate list.

[0164] After obtaining the candidate set of abnormal regions, in order to further screen out the regions most likely to be lesions, it is necessary to sort each candidate region in the candidate set of abnormal regions in descending order according to the abnormal confidence score to generate a sorted candidate list. The sorting process helps to give priority to processing candidate regions with high abnormal probabilities, improving the efficiency and accuracy of subsequent processing.

[0165] The sorting process is based on the abnormal confidence score. For each candidate region in the candidate set of abnormal regions, its abnormal confidence score is a quantified indicator representing the likelihood of the existence of an abnormality in that region. Sorting these candidate regions from high to low according to the abnormal confidence score gives the sorted candidate list.

[0166] In actual operation, common sorting algorithms such as quicksort and mergesort can be used. Taking quicksort as an example, its basic idea is to select a pivot element, divide the list into two parts such that the elements in the left part are all less than or equal to the pivot element, and the elements in the right part are all greater than or equal to the pivot element, and then recursively sort the left and right parts respectively.

[0167] Suppose the candidate set of abnormal regions is set A, and each candidate region in it has a corresponding abnormal confidence score. Select the first candidate region in set A as the pivot element, divide set A into two parts, one is the candidate region set B with an abnormal confidence score less than or equal to the pivot element, and the other is the candidate region set C with an abnormal confidence score greater than the pivot element. Then recursively sort set B and set C respectively, and finally obtain a sorted candidate list sorted in descending order according to the abnormal confidence score.

[0168] The first candidate region in the sorted candidate list has the highest anomaly confidence score, meaning it is most likely to be the lesion region, and the anomaly likelihood of subsequent candidate regions decreases in turn. This sorted candidate list provides an ordered set of candidate regions for subsequent deduplication and final anomaly region determination.

[0169] Step S146: Select the first candidate region from the sorted candidate list as the reference region, and calculate the spatial overlap degree between the remaining candidate regions and the reference region; the spatial overlap degree is calculated by the Intersection over Union (IoU) algorithm to calculate the proportion of the overlapping area of the bounding box of the reference region and the bounding boxes of other candidate regions.

[0170] After obtaining the sorted candidate list, select the first candidate region from the list as the reference region. Since this candidate region has the highest anomaly confidence score, it is preferentially used as a reference to determine whether other candidate regions overlap with it.

[0171] Next, calculate the spatial overlap degree between the remaining candidate regions and the reference region. The spatial overlap degree is an index to measure the degree of spatial overlap between two candidate regions, and is calculated by the Intersection over Union (IoU) algorithm.

[0172] The core of the Intersection over Union (IoU) algorithm is to calculate the intersection area and union area of the bounding box of the reference region and the bounding boxes of other candidate regions, and then divide the intersection area by the union area to obtain the proportion of the overlapping area.

[0173] Suppose the bounding box of the reference region is a rectangular box R1, its upper left corner coordinates are (x1, y1, z1), and its lower right corner coordinates are (x2, y2, z2); the bounding box of another candidate region is a rectangular box R2, its upper left corner coordinates are (x3, y3, z3), and its lower right corner coordinates are (x4, y4, z4).

[0174] First, calculate the intersection region of the two bounding boxes. The upper left corner coordinates of the intersection region are (max(x1, x3), max(y1, y3), max(z1, z3)), and the lower right corner coordinates are (min(x2, x4), min(y2, y4), min(z2, z4)). If the upper left corner coordinates of the intersection region are greater than the lower right corner coordinates, it means that the two bounding boxes have no intersection, and the intersection area is 0.

[0175] Then, calculate the volume V_intersection of the intersection region. Assume the length, width, and height of the intersection region are l, w, and h respectively, then V_intersection = l * w * h, where l = max(0, min(x2, x4) - max(x1, x3)), w = max(0, min(y2, y4) - max(y1, y3)), and h = max(0, min(z2, z4) - max(z1, z3)).

[0176] Next, calculate the union area of the two bounding boxes. The union area is equal to the sum of the volumes of the two bounding boxes minus the intersection area. The volume V1 of the reference region bounding box is (x2 - x1) * (y2 - y1) * (z2 - z1), and the volume V2 of the other candidate region bounding box is (x4 - x3) * (y4 - y3) * (z4 - z3). Then the union area V_union = V1 + V2 - V_intersection.

[0177] Finally, calculate the intersection over union IoU = V_intersection / V_union. The value of this intersection over union is the spatial overlap degree of the two candidate regions, which reflects the overlapping degree of the two candidate regions in space. The larger the value, the higher the overlapping degree.

[0178] For each candidate region in the sorted candidate list except the reference region, calculate its spatial overlap degree with the reference region according to the above method. These values of the spatial overlap degrees will be used for subsequent screening operations to remove those candidate regions that highly overlap with the reference region and avoid repeated identification of the same lesion region.

[0179] Step S147: Remove the candidate regions whose spatial overlap degrees meet the conditions according to the preset overlap degree threshold, and update the sorted candidate list.

[0180] After calculating the spatial overlap degrees of the remaining candidate regions with the reference region, screen these candidate regions according to the preset overlap degree threshold. The preset overlap degree threshold is a preset value used to determine whether the overlapping degree of two candidate regions is too high and needs to be removed.

[0181] If the spatial overlap degree of a certain candidate region with the reference region is greater than the preset overlap degree threshold, it means that these two candidate regions are very likely to be describing the same lesion region. To avoid repeated identification, this candidate region needs to be removed from the sorted candidate list.

[0182] During specific operations, each candidate region in the sorted candidate list except the reference region is traversed, and its spatial overlap with the reference region is compared with a preset overlap threshold. If the spatial overlap is greater than the preset overlap threshold, the candidate region is removed from the sorted candidate list; if the spatial overlap is less than or equal to the preset overlap threshold, the candidate region is retained.

[0183] For example, assume that the preset overlap threshold is a specific value T. For a candidate region R in the sorted candidate list, if its spatial overlap IoU(R) with the reference region is greater than T, the candidate region R is removed from the sorted candidate list; if IoU(R) <= T, the candidate region R remains in the sorted candidate list.

[0184] After such screening and removal processes, the sorted candidate list is updated. In the updated sorted candidate list, the overlap of each candidate region with the reference region is below the preset overlap threshold, reducing the possibility of repeated recognition and making the subsequent determined abnormal regions more accurate.

[0185] Step S148: Repeat the following operations until the sorted candidate list is an empty set: Select the candidate region with the highest anomaly confidence score in the current sorted candidate list as the reference region, and calculate the spatial overlap of the remaining candidate regions with the reference region; if the spatial overlap of a candidate region with the reference region exceeds the preset overlap threshold, remove the candidate region from the sorted candidate list; add the reference regions that have not been removed to the final abnormal region set, and record their position information and morphological description information; the morphological description information includes the maximum diameter, shape irregularity, and edge clarity index of the abnormal region.

[0186] After one screening and removal process, there may still be multiple candidate regions in the sorted candidate list. To ensure that all possible abnormal regions are accurately identified, a series of operations need to be repeated until the sorted candidate list is an empty set.

[0187] Each time the operation is repeated, first select the candidate region with the highest anomaly confidence score in the current sorted candidate list as the new reference region. Since the sorted candidate list is sorted in descending order of anomaly confidence score, the first candidate region is the one with the highest anomaly confidence score.

[0188] Then, calculate the spatial overlap of the remaining candidate regions with this new reference region. The calculation method is the same as before, and the ratio of the overlapping area of the reference region bounding box and the other candidate region bounding boxes is calculated through the intersection over union algorithm.

[0189] Next, these spatial overlap degrees are judged according to a preset overlap degree threshold. If the spatial overlap degree between a certain candidate region and the new reference region exceeds the preset overlap degree threshold, it indicates that these two candidate regions are very likely to describe the same lesion region, and this candidate region is removed from the sorted candidate list.

[0190] After removing the candidate regions with high overlap degrees, the new reference regions that have not been removed are added to the final abnormal region set. Meanwhile, the position information and morphological description information of this reference region are recorded. The position information can be represented by the coordinates of the bounding box, such as the upper left corner coordinates and the lower right corner coordinates. The morphological description information includes the maximum diameter of the abnormal region, the shape irregularity degree, and the edge clarity index.

[0191] The maximum diameter can be obtained by calculating the length of the longest line segment within the bounding box of the abnormal region. The shape irregularity degree can be measured by comparing the difference between the actual shape of the abnormal region and regular shapes (such as circles, rectangles). For example, the ratio of the perimeter of the abnormal region to the perimeter of a regular shape with the same area can be calculated. The edge clarity index can be determined by analyzing the gray level change gradient at the edge of the abnormal region. The larger the gray level change gradient, the clearer the edge.

[0192] Repeat the above operations. Each time, select the candidate region with the highest abnormal confidence score from the updated sorted candidate list as the reference region, perform spatial overlap degree calculation, screening, and removal operations until the sorted candidate list is an empty set. At this time, the final abnormal region set contains all the abnormal regions that have been screened and confirmed, and the position information and morphological description information of these regions will be used for generating subsequent diagnostic reports.

[0193] Step S150: Generate a diagnostic report according to the position information and morphological description information, and transmit the diagnostic report to a medical terminal device for display.

[0194] After determining the position information and morphological description information of the oral and maxillofacial abnormal regions of the target patient, it is necessary to generate a diagnostic report according to this information and transmit the report to a medical terminal device for display so that doctors can intuitively understand the patient's condition.

[0195] Step S151: Perform a matching process on the position information with a preset anatomical structure database to determine the anatomical part name corresponding to the abnormal region; the matching process determines the anatomical part name corresponding to the minimum distance by performing spatial registration on the center coordinates of the abnormal region and the standard coordinates of the anatomical structure, and combining the topological relationship constraints between anatomical structures.

[0196] To accurately describe the location of the abnormal region, it is necessary to match the location information with the preset anatomical structure database. The preset anatomical structure database contains the standard coordinate information of each anatomical part in the oral and maxillofacial region and the topological relationships between them.

[0197] First, calculate the center coordinates of the abnormal region. For the bounding box of each abnormal region, its center coordinates can be calculated from the upper-left corner coordinates and the lower-right corner coordinates of the bounding box. Suppose the upper-left corner coordinates of the bounding box are (x1, y1, z1) and the lower-right corner coordinates are (x2, y2, z2), then the center coordinates are ((x1 + x2) / 2, (y1 + y2) / 2, (z1 + z2) / 2).

[0198] Then, perform spatial registration on the center coordinates of the abnormal region and the standard coordinates of the anatomical structures in the preset anatomical structure database. The purpose of spatial registration is to map the center coordinates of the abnormal region into the coordinate system of the anatomical structure database for accurate matching.

[0199] Next, combine the topological relationship constraints between anatomical structures and calculate the distances between the center coordinates of the abnormal region and the standard coordinates of each anatomical structure. The topological relationship constraints can help exclude some matching results that do not conform to the anatomical logic. For example, an abnormal region cannot be located in a non-anatomical area between two anatomical parts.

[0200] Finally, determine the name of the anatomical part corresponding to the minimum distance. For the distances between the center coordinates of the abnormal region and the standard coordinates of each anatomical structure, select the anatomical part with the minimum distance as the anatomical part corresponding to the abnormal region. In this way, the location of the abnormal region in the oral and maxillofacial anatomical structure can be accurately determined.

[0201] Step S152: Compare the morphological description information with the preset lesion feature database to determine the predicted pathological type of the abnormal region; the comparison process uses the nearest neighbor algorithm to perform similarity matching between the morphological description feature vectors and the standard lesion features in the database.

[0202] After determining the name of the anatomical part corresponding to the abnormal region, it is necessary to compare the morphological description information with the preset lesion feature database to determine the predicted pathological type of the abnormal region. The preset lesion feature database contains the standard lesion features of various common pathological types, such as the eigenvalue characteristics of lesions of different pathological types in terms of maximum diameter, shape irregularity, edge clarity, etc.

[0203] First, convert the morphological description information into a morphological description feature vector. The morphological description feature vector is a multi-dimensional vector, and each dimension corresponds to a morphological description index, such as the maximum diameter, shape irregularity, edge clarity, etc. For example, the morphological description feature vector can be expressed as (d, s, e), where d represents the maximum diameter, s represents the shape irregularity, and e represents the edge clarity.

[0204] Then, use the nearest neighbor algorithm to perform similarity matching between the morphological description feature vector and the standard lesion features in the preset lesion feature database. The basic idea of the nearest neighbor algorithm is to calculate the similarity between the morphological description feature vector and each standard lesion feature vector in the database, and select the pathological type corresponding to the standard lesion feature vector with the highest similarity as the predicted result of the pathological type of the abnormal area.

[0205] The calculation of similarity can use methods such as Euclidean distance and cosine similarity.

[0206] Calculate the Euclidean distance between the morphological description feature vector and all standard lesion feature vectors in the database, and select the pathological type corresponding to the standard lesion feature vector with the smallest distance as the predicted result of the pathological type of the abnormal area.

[0207] Step S153: Generate a diagnostic advice text according to the anatomical part name and the predicted result of the pathological type; the diagnostic advice text is generated by semantically combining the anatomical part name and the pathological type through a predefined template.

[0208] After determining the anatomical part name and the predicted result of the pathological type corresponding to the abnormal area, generate a diagnostic advice text according to this information. The diagnostic advice text is an important reference for doctors to make diagnoses and treatments, and it is necessary to accurately and clearly express the situation and suggestions of the abnormal area.

[0209] The predefined template is a preset text format used to semantically combine the anatomical part name and the pathological type. For example, the predefined template can be "A [pathological type] lesion is found in the [anatomical part name], and it is recommended to further [specific examination or treatment suggestion]".

[0210] Substitute the anatomical part name and the predicted result of the pathological type into the predefined template to generate a diagnostic advice text. For example, if the anatomical part name is "maxilla" and the predicted result of the pathological type is "cyst", the diagnostic advice text can be "A cyst lesion is found in the maxilla, and it is recommended to further perform a contrast-enhanced CT scan to clarify the nature of the lesion".

[0211] Step S154: Perform a formatted combination process on the diagnostic advice text, the location information, and the morphological description information to generate a diagnostic report; the formatted combination process typesets and integrates the text information and the coordinate data according to the medical report standard format.

[0212] After generating the diagnostic advice text, it is combined and formatted with the location information and morphological description information to generate a diagnostic report. The purpose of the formatting and combination process is to typeset and integrate the text information and coordinate data according to the standard format of a medical report, making the diagnostic report standardized and readable.

[0213] First, present the location information in a clear manner, such as listing the bounding box coordinates or center coordinates of the abnormal area. Then, describe in detail the morphological description information such as the maximum diameter, shape irregularity, edge clarity, etc. Next, add the diagnostic advice text to the report so that it is correlated with the location information and morphological description information.

[0214] In terms of typesetting, set the formats of the title, body text, paragraphs, etc. according to the standard format of a medical report. For example, set the report title at the beginning of the report, such as "Oral and Maxillofacial Imaging Diagnostic Report", and then list the patient information, abnormal area location information, morphological description information, pathological type prediction results, diagnostic advice, etc. in sequence.

[0215] Finally, save the formatted and combined content as a diagnostic report file, such as in PDF format or Word format. Transmit the generated diagnostic report to a medical terminal device for display. Doctors can view the diagnostic report on the medical terminal device to understand the abnormalities of the patient's oral and maxillofacial region, providing a basis for subsequent diagnosis and treatment.

[0216] In the above embodiment, privacy-sensitive data such as the patient's oral and maxillofacial three-dimensional image data set is involved. To protect the patient's privacy and prevent data leakage, a series of privacy protection and anti-leakage technical means are adopted. For example, during the data collection process, the patient's identity information is encrypted. When collecting the oral and maxillofacial three-dimensional image data, sensitive identity information such as the patient's real name and ID number is not directly recorded, but these information are converted into encrypted codes using an encryption algorithm. For example, use a symmetric encryption algorithm, such as the AES algorithm, to encrypt the patient's identity information, and the encryption key is managed by a dedicated security system. At the same time, strict security management is carried out on the collection device. Ensure that the software and hardware of the collection device have passed security inspections to prevent data from being illegally obtained during the collection process. Strict access control is carried out on the collection device, and only authorized personnel can operate the collection device.

[0217] In this method, multiple pre-trained models are involved, such as three-dimensional convolutional neural network, lesion classification network, multi-scale feature fusion network, etc. The general steps of model training include data preparation, model construction, model training, and model evaluation. Specifically, the corresponding training process can be referred to according to the process of the above model application, and it will not be elaborated here.

[0218] Based on the above description, in another embodiment, the embodiment of the present invention further provides an oral and maxillofacial surgery image recognition and diagnosis system based on deep learning. Refer to Figure 2 , Figure 2 is a structural diagram of the oral and maxillofacial surgery image recognition and diagnosis system 100 based on deep learning provided by the embodiment of the present invention. The oral and maxillofacial surgery image recognition and diagnosis system 100 based on deep learning may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 112 (for example, one or more processors) and a memory 111. Among them, the memory 111 can be short-term storage or persistent storage. The program stored in the memory 111 may include one or more modules, and each module may include a series of instruction operations on the oral and maxillofacial surgery image recognition and diagnosis system 100 based on deep learning. Further, the central processing unit 112 may be set to communicate with the memory 111 and execute a series of instruction operations in the memory 111 on the oral and maxillofacial surgery image recognition and diagnosis system 100.

[0219] The oral and maxillofacial surgery image recognition and diagnosis system 100 based on deep learning may further include one or more power supplies, one or more communication units 113, one or more transfer to output interfaces, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0220] The steps performed by the oral and maxillofacial surgery image recognition and diagnosis system in the above embodiment can be combined with Figure 2 the structure of the oral and maxillofacial surgery image recognition and diagnosis system shown.

[0221] In addition, the embodiment of the present invention further provides a storage medium for storing a computer program for executing the method provided in the above embodiment.

[0222] The embodiment of the present invention further provides a computer program product including instructions, which when running on a computer, causes the computer to execute the method provided in the above embodiment.

[0223] Those of ordinary skill in the art will understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments. The foregoing storage medium can be at least one of the following media: read-only memory (ROM), RAM, magnetic disk, or optical disc, etc., which can store program codes.

[0224] It should be noted that the embodiments in this specification are all described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the partial description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0225] As described above, it is only a specific implementation step of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An image recognition and diagnosis method for oral and maxillofacial surgery based on deep learning, characterized in that, Including: Obtain a three-dimensional oral and maxillofacial image data set of a target patient, where the three-dimensional image data set includes multiple groups of continuous scanned layer image data; Perform image feature extraction processing on the three-dimensional image data set to obtain a hierarchical image feature set of the continuous scanned layer image data; the hierarchical image feature set includes local anatomical structure features and global spatial distribution features; Call a pre-trained multi-scale feature fusion network to perform multi-scale feature fusion on the hierarchical image feature set to generate a fused feature map; Based on the fused feature map, perform lesion area recognition processing to determine the location information and morphological description information of the abnormal area of the oral and maxillofacial region of the target patient; Generate a diagnostic report according to the location information and morphological description information, and transmit the diagnostic report to a medical terminal device for display.

2. The method for oral and maxillofacial surgical image recognition and diagnosis based on deep learning according to claim 1, characterized in that, The performing image feature extraction processing on the three-dimensional image data set to obtain a hierarchical image feature set of the continuous scanned layer image data includes: Perform spatial normalization processing on the continuous scanned layer image data to obtain a normalized three-dimensional image sequence; Call a three-dimensional convolutional neural network to perform shallow feature extraction processing on the normalized three-dimensional image sequence to obtain a primary image feature set; the primary image feature set includes edge contour features and texture distribution features; Perform cross-layer feature enhancement processing on the primary image feature set to obtain an enhanced image feature set; the enhanced image feature set contains local anatomical structure features with enhanced spatial resolution; Perform global pooling processing on the enhanced image feature set to obtain global spatial distribution features; Perform feature alignment processing on the local anatomical structure features and the global spatial distribution features to form a hierarchical image feature set; the feature alignment processing includes performing a spatial coordinate mapping operation on the local anatomical structure features to generate a spatially aligned feature vector corresponding to the global spatial distribution features; Perform a feature broadcast operation on the global spatial distribution features to expand its spatial dimension to be consistent with the spatially aligned feature vector and adjust the channel dimension to match; Perform channel dimension splicing on the expanded global spatial distribution features and the spatially aligned feature vector to generate a hierarchical image feature set with a unified spatial dimension.

3. The method for oral and maxillofacial surgical image recognition and diagnosis based on deep learning according to claim 2, characterized in that, The calling a three-dimensional convolutional neural network to perform shallow feature extraction processing on the normalized three-dimensional image sequence to obtain a primary image feature set includes: Perform convolutional kernel sliding processing on the normalized three-dimensional image sequence to generate an initial convolutional feature map, and perform max-pooling processing on the initial convolutional feature map to obtain a dimensionality-reduced feature map; Input the dimensionality-reduced feature map into a residual connection module for feature compensation processing to obtain a compensated feature map; the feature compensation processing includes performing an element-wise addition operation on the dimensionality-reduced feature map and the original feature map passed through the skip connection to generate a compensated feature map with the characteristic of retaining details; Perform channel attention weight assignment processing on the compensated feature map to generate a weighted feature map; the channel attention weight assignment processing includes performing a global average pooling operation on the channel dimension of the compensated feature map to generate a channel description vector; Input the channel description vector into a fully connected layer to generate channel attention weights, and recalibrate each channel of the compensated feature map based on the channel attention weights to generate a weighted feature map; Extract the edge contour features and texture distribution features from the weighted feature map to form a primary image feature set; the edge contour features are extracted by a learnable edge convolution kernel for gradient features of the weighted feature map, and the texture distribution features are extracted by a multi-scale convolution kernel group for adaptive texture pattern extraction of the weighted feature map.

4. The method for oral and maxillofacial surgical image recognition and diagnosis based on deep learning according to claim 2, wherein, Performing cross-layer feature enhancement processing on the primary image feature set to obtain an enhanced image feature set, including: Perform feature splicing processing on the edge contour features and texture distribution features in the primary image feature set to generate a spliced feature map; Perform dilated convolution processing on the spliced feature map to generate a multi-scale receptive field feature map; the dilated convolution processing uses convolution kernels with different dilation rates to act on the spliced feature map in parallel to generate feature sub-maps with different receptive field ranges; Merge the feature sub-maps in the channel dimension to generate a multi-scale receptive field feature map, and perform feature pyramid fusion processing on the multi-scale receptive field feature map to obtain a multi-scale fusion feature map; the feature pyramid fusion processing includes upsampling or downsampling the feature sub-maps at different scales to a unified resolution and performing an element-wise addition operation; Perform non-linear activation processing on the multi-scale fusion feature map to generate an activation feature map; the non-linear activation processing uses an activation function with a gating mechanism to dynamically activate each channel of the multi-scale fusion feature map; Perform spatial attention weight allocation processing on the activation feature map to obtain an enhanced image feature set; the spatial attention weight allocation processing includes performing a dual-path operation of max pooling and average pooling on the spatial dimension of the activation feature map to generate a dual-path pooled feature map; Perform channel splicing on the dual-path pooled feature map and generate a spatial attention weight map through a convolutional layer, and perform spatial dimension weighting on the activation feature map based on the spatial attention weight map to generate an enhanced image feature set.

5. The method for oral and maxillofacial surgical image recognition and diagnosis based on deep learning according to claim 2, characterized in that, Performing global pooling processing on the enhanced image feature set to obtain global spatial distribution features, including: Perform channel dimension average pooling processing on the local anatomical structure features in the enhanced image feature set to generate a channel statistical feature vector; Perform fully connected layer mapping processing on the channel statistical feature vector to generate a dimensionality-reduced feature vector; Map the dimensionality-reduced feature vector through a fully connected layer to be consistent with the channel dimension of the enhanced image feature set, and perform feature broadcasting processing to generate a spatial weight distribution map; the feature broadcasting processing includes copying and expanding the mapped feature vector to the same spatial size as the enhanced image feature set; Perform channel dimension weighting and spatial dimension aggregation processing on the enhanced image feature set based on the spatial weight distribution map to obtain global spatial distribution features.

6. The method for oral and maxillofacial surgical image recognition and diagnosis based on deep learning according to claim 1, wherein Call a pre-trained multi-scale feature fusion network to perform multi-scale feature fusion on the hierarchical image feature set to generate a fusion feature map, including: Perform feature splicing processing on the local anatomical structure features and the global spatial distribution features to generate initial splicing features; Perform grouped axial self-attention mechanism processing on the initial splicing features to generate a self-attention weight distribution map; the grouped axial self-attention mechanism processing includes decomposing the features into axial sub-regions along the spatial dimension, respectively calculating the correlation matrix between spatial positions within each axial sub-region, and normalizing the correlation matrix to generate a local self-attention weight distribution map; Perform dynamic weighting processing on the local anatomical structure features based on the self-attention weight distribution map to generate weighted local features; the dynamic weighting processing includes performing an element-wise multiplication operation on the self-attention weight distribution map and the local anatomical structure features; Perform feature fusion processing on the weighted local features and the global spatial distribution features to generate intermediate fusion features; Perform cross-channel feature interaction processing on the intermediate fusion features to generate a fusion feature map; the cross-channel feature interaction processing includes decomposing the intermediate fusion features into multiple channel groups, performing independent convolution operations on each channel group, and then rearranging the channel order to generate a fusion feature map with channel interaction characteristics.

7. The method for oral and maxillofacial surgical image recognition and diagnosis based on deep learning according to claim 6, wherein The performing cross-channel feature interaction processing on the intermediate fusion features to generate a fusion feature map includes: Partition the intermediate fusion features into multiple sub-feature groups along the channel dimension; wherein, each sub-feature group contains a predetermined number of consecutive channels; Perform grouped convolution processing on each of the sub-feature groups to generate grouped feature maps; the grouped convolution processing uses independent convolution kernels to act on each sub-feature group respectively; Perform channel shuffle processing on the grouped feature maps to generate shuffled feature maps; the channel shuffle processing rearranges the channels in each sub-feature group according to a predetermined rule to promote cross-group information interaction; Perform pointwise convolution processing on the shuffled feature maps to generate cross-channel interaction features; the pointwise convolution processing uses a 1×1 convolution kernel to adjust the channel dimension; Perform residual connection processing on the cross-channel interaction features and the intermediate fusion features to generate a fusion feature map; the residual connection processing performs element-wise addition on the cross-channel interaction features and the original intermediate fusion features.

8. The method for oral and maxillofacial surgical image recognition and diagnosis based on deep learning according to claim 1, characterized in that The performing lesion area recognition processing based on the fusion feature map to determine the position information and morphological description information of the oral and maxillofacial abnormal area of the target patient includes: Perform candidate region generation processing on the fusion feature map to obtain multiple candidate region bounding boxes; the candidate region generation processing uses a sliding window mechanism to generate initial bounding boxes on the fusion feature map at different scales and aspect ratios; Perform feature cropping processing on each of the candidate region bounding boxes to generate candidate region feature maps; the feature cropping processing intercepts the feature data of the corresponding region from the fusion feature map according to the bounding box coordinates; Input the candidate region feature maps into a pre-trained lesion classification network for abnormal probability prediction processing to obtain the abnormal confidence scores of each candidate region; the lesion classification network includes multiple fully connected layers for mapping the candidate region feature maps to abnormal probability values; Filter the candidate regions corresponding to the abnormal confidence scores according to a preset confidence threshold to generate a candidate set of abnormal regions; Sort each candidate region in the candidate set of abnormal regions in descending order according to the abnormal confidence score to generate a sorted candidate list; Select the first candidate region from the sorted candidate list as the reference region, and calculate the spatial overlap degree between the remaining candidate regions and the reference region; the spatial overlap degree calculates the proportion of the overlapping area of the reference region bounding box and the bounding boxes of other candidate regions through the intersection over union algorithm; Filter and remove the candidate regions whose spatial overlap degree meets the conditions according to a preset overlap threshold, and update the sorted candidate list; Repeat the following operations until the sorted candidate list is an empty set: select the candidate region with the highest abnormal confidence score in the current sorted candidate list as the reference region, and calculate the spatial overlap degree between the remaining candidate regions and the reference region; if the spatial overlap degree between the candidate region and the reference region exceeds the preset overlap threshold, remove the candidate region from the sorted candidate list; add the reference regions that have not been removed to the final abnormal region set, and record their position information and morphological description information; the morphological description information includes the maximum diameter, shape irregularity, and edge clarity index of the abnormal region.

9. The method for oral and maxillofacial surgical image recognition and diagnosis based on deep learning according to claim 1, wherein Generating a diagnostic report based on the position information and morphological description information, including: Perform a matching process on the position information with a preset anatomical structure database to determine the anatomical part name corresponding to the abnormal region; the matching process determines the anatomical part name corresponding to the minimum distance by performing spatial registration of the center coordinates of the abnormal region with the standard coordinates of the anatomical structure and combining the topological relationship constraints between the anatomical structures; Compare the morphological description information with a preset lesion feature database to determine the predicted pathological type result of the abnormal region; the comparison process uses the nearest neighbor algorithm to perform similarity matching between the morphological description feature vectors and the standard lesion features in the database; Generate a diagnostic suggestion text based on the anatomical part name and the predicted pathological type result; the diagnostic suggestion text is generated by semantically combining the anatomical part name and the pathological type through a predefined template; Perform a formatting combination process on the diagnostic suggestion text with the position information and morphological description information to generate a diagnostic report; the formatting combination process typesets and integrates the text information and coordinate data according to the medical report standard format.

10. An oral and maxillofacial surgery image recognition and diagnosis system based on deep learning, characterized in that, Including: A processor; A memory, in which a computer program is stored, and when the computer program is executed, it implements the deep learning-based oral and maxillofacial surgical image recognition and diagnosis method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Skull side surface image analysis method based on neural network and random forest, and system

    CN110246580A

  • Cerebral hemorrhage classification, positioning and prediction method based on three-dimensional deep learning model

    CN110503630A

  • Pancreatic cancer image segmentation system based on multi-view feature fusion network

    CN117994518A

  • Deep learning-based pulmonary embolism segmentation method

    CN118212411A

  • A deep learning texture analysis system for B-ultrasound images

    CN119784810A

Cited By

  • Oral periodontal image analysis method and system based on multi-scale feature fusion

    CN120953746A

  • Tooth health preliminary screening method and device based on image recognition, equipment and medium

    CN121504870A

  • Image recognition-based methods, devices, equipment, and media for initial screening of dental health.

    CN121504870B

  • Medical image diagnosis method and device and computing equipment

    CN122134724A

  • Medical imaging diagnostic methods and devices, computing equipment

    CN122134724B