Lower jawbone fracture detection system based on wavelet transform feature enhancement and multi-scale attention focusing
By introducing a wavelet convolution module and a multi-scale attention mechanism in the mandible fracture detection system, the problems of insufficient identification capabilities and prone to missed detection and missed detection in the prior art are solved, and higher detection accuracy and accuracy are achieved.
Patent Information
- Application Number
- CN202510223013.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The existing mandible fracture detection methods are insufficient in the recognition ability of complex fracture images, and are prone to missed and missed detection, and relatively insufficient research on children's population.
Using a detection system based on wavelet transformation feature enhancement and multi-scale attention focus, the accuracy and efficiency of the model in fracture line detection is significantly improved by introducing a wavelet convolution module and a multi-scale attention mechanism.
It significantly improves the accuracy and positioning accuracy of fracture line detection, reduces the probability of missed and missed detection, and shows strong recognition ability especially under complex interference and complex fracture morphology.
Smart Images

Figure CN120107218A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a mandibular fracture detection system based on wavelet transform feature enhancement and multi-scale attention focusing. Background Art
[0002] Although maxillofacial fractures in children account for only 1%-5% of all age groups, mandibular fractures account for about 40%-60% of mandibular fractures, showing its importance in clinical research. The mandible is an important independent bone structure in the lower third of the face, which not only determines the appearance of the face, but also plays a key role in supporting chewing function. As the only active bone protruding outward from the skull base, the mandible lacks adequate protection and is therefore very prone to fractures under high-impact external forces such as traffic accidents. Studies have shown that mandibular fractures are one of the common types of fractures in humans, and common fracture sites include the median symphysis, left and right mental foramina, mandibular angle, and condylar neck. When the external force is small, there is usually no obvious displacement in the fracture area due to muscle attachment; when the external force is large, the mandibular condyle breaks first to absorb the impact force, thereby protecting the intracranial organs from serious damage. Mandibular fractures are often accompanied by complex follow-up problems, including possible bone dislocation, bone compression, anteversion or retraction at the fracture site, which further increases the difficulty of imaging diagnosis. In CT images of mandibular fractures, the fracture area varies greatly in size and shape, and accompanying symptoms such as bleeding cause changes in the X-ray absorption rate locally, making it more difficult to identify the fracture line. These factors directly prolong the diagnosis time of radiologists and easily lead to misdiagnosis or missed diagnosis, affecting patients' timely treatment.
[0003] At the same time, compared with adults, the treatment of fractures in children is more complex and challenging because children are in the growth and development stage, with vigorous growth metabolism and strong tissue healing ability. The treatment plan needs to take into account the long-term effects of functional recovery and growth and development, which places higher demands on the accuracy of diagnosis and the timeliness of treatment. However, existing research mainly focuses on the imaging identification and diagnosis and treatment of mandibular fractures in adults, and research on the pediatric population is relatively insufficient.
[0004] With the rapid development of medical imaging technology, such as the continuous maturity of image fusion, target detection and image segmentation, its role in clinical diagnosis has become increasingly prominent. Computer-aided diagnosis systems based on high-performance deep learning have gradually been applied to the detection and analysis of mandibular fractures, hoping to improve the consistency and accuracy of diagnostic results. However, although deep learning combined with medical imaging has made significant progress in the field of fracture detection, existing methods still face many challenges:
[0005] 1. Difficulty in data collection: Due to the protection of patient privacy and imaging differences between different devices, it is difficult to obtain a large-scale and standardized mandibular fracture dataset, which limits the training effect of the deep learning model.
[0006] 2. Significant differences in fracture areas: In CT images, the size and shape of the fracture area vary greatly due to different anatomical positions (coronal, sagittal and axial), fracture sites (median symphysis, mental foramen, mandibular angle and condylar neck) and external force intensity, which makes the existing model insufficient in terms of regional recognition accuracy.
[0007] 3. Lack of multi-scale information: Existing algorithms pay little attention to multi-scale features, which limits the model's ability to recognize complex fracture areas, especially when dealing with small or atypical fractures.
[0008] 4. Identify interference factors: The fracture area is often accompanied by local bleeding, and the changes in X-ray absorption caused by bleeding make it more difficult to identify the fracture line in CT images. In addition, fractures in the same location may present different fracture sizes and shapes, which further increases the complexity of algorithm detection. Summary of the invention
[0009] In view of the technical difficulties existing in mandibular fracture line detection, such as insufficient ability to recognize small fracture lines in complex fracture images, easy missed detection and false detection under the influence of interference factors, and poor performance in small target detection, the present invention provides a mandibular fracture detection system based on wavelet transform feature enhancement and multi-scale attention focusing. By introducing wavelet convolution module and multi-scale attention mechanism, the accuracy and efficiency of the model in fracture line detection are significantly improved.
[0010] A mandibular fracture detection system based on wavelet transform feature enhancement and multi-scale attention focusing comprises a computer memory, a computer processor and a computer program stored in the computer memory and executable on the computer processor, wherein a trained mandibular fracture detection model is stored in the computer memory; the mandibular fracture detection model comprises a backbone network, a feature fusion network and a detection head part;
[0011] The backbone network is used to extract feature maps P3, P4 and P5 of different scales. The backbone network first uses two convolutional layers to downsample the input image, and then iteratively extracts the P3 feature map through two layers of C3k2_WT modules. Based on the P3 feature map, the P4 feature map is further extracted through the third layer of C3k2_WT module. Based on the P4 feature map, the fourth layer of C3k2_WT module, the spatial pyramid fast pooling module SPPF and the C2PSA_EMA module are sequentially passed to finally generate a low-resolution P5 feature map with high-level semantic representation capabilities.
[0012] The feature fusion network adopts a bidirectional fusion strategy, first performing top-down semantic enhancement, upsampling the P5 feature map by bilinear interpolation to align its size with the P4 feature map, and splicing it along the channel dimension to generate the fused feature map P4. new , and then P4 new After upsampling, it is concatenated with the P3 feature map to generate P3 new ; Then transfer the details from bottom to top, and make the spliced P3 new A C3k2_WT module is introduced to perform feature compression and nonlinear enhancement, complete downsampling, and compare it with the original P4 and P5 features. Figure 2 splicing to form a closed-loop fusion path and generate P4 new ', P5 new ';
[0013] The detection head part contains three independent detection heads, and preset multi-scale anchor frames for feature maps of different resolutions; the fused P3 new 、P4 new ', P5 new 'The feature map is passed to three independent detection heads, and the small-scale anchor boxes densely cover P3 new High-resolution grid, medium-scale anchor box adapted to P4 new 'Medium-resolution grid, large-scale anchor box adapted to P5 new '; each detection head outputs bounding box coordinates, category probability and confidence simultaneously, and the outputs of the three detection heads are fused to finally retain high-precision predictions;
[0014] When the computer processor executes the computer program, the following steps are implemented:
[0015] The image to be detected is input into the trained mandibular fracture detection model to obtain the prediction result of mandibular fracture.
[0016] Furthermore, the backbone network first uses two convolutional layers to downsample the input image, and the formula is as follows:
[0017] O 1 =Conv(I,64,3,2)
[0018] O 2 =Conv(O 1 ,128,3,2)
[0019] Among them, O 1 represents the output feature map after the first convolutional layer, O 2 Represents the output feature map after the second convolution layer; the input data I undergoes the first convolution operation, using 64 convolution kernels of size 3×3 and a stride of 2 to obtain the output feature map O 1; Perform the first convolution operation again, using 128 convolution kernels of size 3×3 and a stride of 2 to obtain the new output feature map O 2 .
[0020] Furthermore, the working process of the C3k2_WT module is as follows:
[0021] First, the feature map of the input C3k2_WT module is decomposed to obtain the feature map X 1 and X 2 ;
[0022] For feature map X 1 Perform lightweight convolution feature extraction to obtain X light ;
[0023] The wavelet transform transforms the feature map X through four filters 2 Decomposed into a low frequency component F LL and three high frequency components F LH 、F HL and F HH , and then use the inverse wavelet transform to fuse the low-frequency component and the three high-frequency components to reconstruct the same 2 Feature map X with consistent spatial resolution IWT ;
[0024] The obtained X light and X IWT Fusion is performed to obtain feature fusion results for subsequent operations.
[0025] Furthermore, for the feature map X 1 Perform lightweight convolution feature extraction, the formula is as follows:
[0026] X light =Conv light (X 1 )=Pointwise(Depthwise(X 1 , k = 2))
[0027] In the formula, Depthwise means using a 2×2 small kernel for depth convolution, Pointwise means using a 1×1 convolution for channel fusion, and k=2 means that the size of the convolution kernel is 2×2.
[0028] Furthermore, a low-frequency component F LL , reflecting the overall structural information of the input; the three high-frequency components are F LH 、F HL and F HH , where F LH Represents the high-frequency component in the horizontal direction, which is used to capture the edge details in the horizontal direction; F HLRepresents the high-frequency component in the vertical direction, which is used to capture the edge details in the vertical direction; F HH Represents the high-frequency component in the diagonal direction and is used to extract the texture features in the diagonal direction.
[0029] Furthermore, the working process of the C2PSA_EMA module is as follows:
[0030] After the feature map is input into the C2PSA_EMA module, it passes through a 1×1 convolutional layer to compress the number of channels and decompose it into retained feature A and feature B for further processing; feature A directly participates in the final fusion as a global feature; feature B enters the multi-scale attention path to extract spatial and contextual information;
[0031] Apply average pooling along the height and width directions to feature B to generate B respectively. h and B w ; Then B h and B w The attention weight matrix is generated by concatenation and 1×1 convolution, and the attention of the salient area is further strengthened by the sigmoid activation function σ to obtain the final attention weight matrix W EMA ; Finally, through weighted operation, the attention weight is applied to the feature map B to highlight the target area and obtain the attention weighted feature B EMA ;
[0032] To B EMA 3×3 convolution and 1×1 convolution are applied in parallel to extract local spatial correlation and global channel context dependency respectively: 3×3 convolution captures texture and edge details in the local neighborhood by expanding the spatial receptive field, and 1×1 convolution models the semantic dependency between global channels through cross-channel linear combination; then the convolution result B of the two is concatenated. FNN Through a two-layer feedforward network (including nonlinear activation function), nonlinear transformation is performed to further enhance the feature expression ability and compare the enhanced result with the original B EMA Add, retain the original information through residual connection, and finally obtain the enhanced feature B′;
[0033] The retained feature A and the enhanced feature B′ are concatenated in the channel dimension, and the number of output channels is restored through a 1×1 convolution to obtain the final feature map for subsequent operations.
[0034] Preferably, in the detection head part, the outputs of the three detection heads are fused across scales through non-maximum suppression (NMS).
[0035] Furthermore, the dataset for training the mandibular fracture detection model is constructed as follows:
[0036] Images were extracted from maxillofacial CT images in DICOM format and a dataset was constructed. The fracture area was annotated, including length, width, and center coordinates. The images were resized and formatted in a standardized manner to ensure the consistency and adaptability of the input data. Finally, the dataset was expanded through data enhancement.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] 1. Wavelet convolution: This paper introduces wavelet convolution (C3k2_WT) in the backbone network (Backbone) to enhance the feature extraction of fracture area details. During the detection process, the model can retain the structural information of the original input and capture local details in multiple scales and directions, reducing unnecessary losses in feature extraction, thereby improving the detection accuracy and positioning accuracy of fracture lines.
[0039] 2. Multi-scale attention mechanism: This invention innovatively introduces a multi-scale attention mechanism (C2PSA_EMA) in the feature extraction stage, and combines it with the rich detail features obtained by the above wavelet convolution to enable the model to focus on the specific areas in the image that are most relevant to detection, and can effectively focus on the key information of the mandibular area. Then, the cross-stage network (CSP) integrates low-level and high-level features to capture the slight changes in the fracture line at different scales, thereby improving the ability to identify the fracture line under complex interference and complex fracture morphology.
[0040] 3. Multi-scale feature fusion optimization: The present invention optimizes the feature fusion stage of the model by integrating the above-mentioned wavelet convolution (C3k2_WT) and multi-scale attention mechanism (C2PSA_EMA) in the feature fusion network (Neck). This module maps high-level semantic information to a finer pixel space by increasing the resolution to facilitate the detection of small objects. With the support of wavelet convolution and multi-scale attention mechanism, the key detail information of the image is accurately retained, and then the upsampled feature map is spliced with the adjacent high-resolution feature map in the channel dimension, effectively aggregating multi-scale feature information, thereby helping the model to better identify and locate complex fracture lines, especially when the fracture position changes little or is partially occluded, enhancing the ability to capture details.
[0041] 4. In specific experiments, the mandibular fracture line detection method proposed in the present invention is superior to the existing fracture detection method in various performance indicators. The detection accuracy, recall rate and F1 score are significantly improved. The detection accuracy rates on the sagittal plane, axial plane and coronal plane reached 87.97%, 89.31% and 92.66% respectively. This method provides a new idea for mandibular fracture detection and provides effective support for fracture diagnosis in clinical practice. It has the potential to be further promoted and applied to other medical imaging diagnoses. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is the overall architecture diagram of the mandibular fracture detection model in the present invention.
[0043] Figure 2 This is the core architecture diagram of the C3K2_WT module in the present invention.
[0044] Figure 3 This is the core architecture diagram of the C2PSA_EMA module in the present invention. DETAILED DESCRIPTION
[0045] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be pointed out that the embodiments described below are intended to facilitate the understanding of the present invention and do not have any limiting effect on the present invention.
[0046] A mandibular fracture detection system based on wavelet transform feature enhancement and multi-scale attention focusing includes a computer memory, a computer processor, and a computer program stored in the computer memory and executable on the computer processor. The computer memory stores a trained mandibular fracture detection model; the mandibular fracture detection model includes a backbone network, a feature fusion network, and a detection head part; the overall architecture of the mandibular fracture detection model is as follows Figure 1 shown.
[0047] The backbone network extracts multi-scale feature maps P3, P4, and P5 from the input image through layer-by-layer convolution operations, and uses step-by-step downsampling to reduce the spatial resolution and increase the channel depth to provide basic information for subsequent feature fusion. The backbone network first downsamples the input image through two convolutional layers: the first convolutional layer uses 64 3×3 convolution kernels (stride 2) to generate feature map O 1 ; The second convolutional layer uses 128 3×3 convolution kernels (stride 2) to generate feature map O 2 . Subsequently, O2 iteratively extracts the high-resolution feature map P3 through two layers of C3k2_WT modules in turn, where the output of each C3k2_WT module is adjusted by a 1×1 standard convolution to adjust the number of channels. After the P3 feature map is processed by the third layer of C3k2_WT module, it is downsampled through a 3×3 convolution layer (stride 2) to generate a medium-resolution feature map P4; the P4 feature map is further processed by the fourth layer of C3k2_WT module, the spatial pyramid fast pooling module (SPPF) and the C2PSA_EMA module to finally generate a low-resolution, high-level semantic feature map P5. The parameters of the convolutional layer following each layer of C3k2_WT module are dynamically adjusted according to the feature map hierarchy.
[0048] The feature fusion network adopts a closed-loop bidirectional strategy to achieve multi-scale feature complementarity. First, perform top-down semantic enhancement, perform bilinear interpolation upsampling on the P5 feature map to align its size with the P4 feature map, and then splice along the channel dimension to generate the fused feature map P4. new , and then P4 new After upsampling, it is concatenated with the P3 feature map to generate P3 new ; Then transfer the details from bottom to top, and make the spliced P3 new A C3k2_WT module is introduced to perform feature compression and nonlinear enhancement, complete downsampling, and compare it with the original P4 and P5 features. Figure 2 splicing to form a closed-loop fusion path and generate P4 new ', P5 new ';
[0049] The detection head part contains three independent detection heads, and preset multi-scale anchor frames for feature maps of different resolutions; the fused P3 new 、P4 new ', P5 new 'The feature map is passed to three independent detection heads, and the small-scale anchor boxes densely cover P3 new High-resolution grid, medium-scale anchor box adapted to P4 new 'Medium-resolution grid, large-scale anchor box adapted to P5 new '; each detection head synchronously outputs bounding box coordinates, category probability and confidence, and the outputs of the three detection heads are fused to finally retain high-precision predictions.
[0050] The present invention completes layer-by-layer convolution operations through a backbone network to extract basic features from multi-scale input images; a step-by-step downsampling method is used to reduce spatial resolution and increase feature map depth, thereby providing basic information for subsequent advanced feature extraction.
[0051] The backbone network first uses two convolutional layers to downsample the input image, the formula is as follows:
[0052] O 1 =Conv(I,64,3,2)
[0053] O 2 =Conv(O 1 ,128,3,2)
[0054] Among them, O 1 represents the output feature map after the first convolutional layer, O 2 Represents the output feature map after the second convolution layer; the input data I undergoes the first convolution operation, using 64 convolution kernels of size 3×3 and a stride of 2 to obtain the output feature map O 1; Perform the first convolution operation again, using 128 convolution kernels of size 3×3 and a stride of 2 to obtain the new output feature map O 2 .
[0055] The present invention is improved based on YOLOv11, and a C3K2_WT (wavelet convolution) module based on a cross-stage part (CSP) network and wavelet transform is introduced into the model to strengthen the capture of local details and reduce unnecessary losses in feature extraction. The module consists of lightweight convolution and wavelet convolution, and extends the one-dimensional wavelet transform (WT) to two dimensions, so that features of different scales and directions can be effectively modeled. Specifically, the module is based on a cross-stage part (CSP) network, and the input feature map is divided into two parts, one of which is extracted by lightweight convolution and directly fused with the output feature, and the other part is subjected to wavelet transform. By constructing four wavelet transform filters, it is used to decompose the input features into low-frequency and high-frequency components in three directions. Then the CSP structure splices and fuses the two parts of the features, so that the model not only retains the structural information of the original input, but also can capture local details in multiple scales and directions, which is helpful to identify fracture lines under complex backgrounds and improve the accuracy of the model.
[0056] like Figure 2 As shown, the working process of the C3K2_WT module is as follows:
[0057] 1.1 Input eigendecomposition:
[0058] X split = split(X) = {X 1 ,X 2}
[0059] The input feature map X is decomposed into two paths: Main path: X 1 , used for lightweight convolution; wavelet path: X 2 , used for wavelet transform expansion.
[0060] 1.2 Lightweight convolution feature extraction
[0061] X light =Conv light (X 1 )=Pointwise(Depthwise(X 1 ,k))
[0062] Process X using depthwise convolution and pointwise convolution 1 , to extract local features and efficiently fuse channel information. Depthwise means using a 2×2 small kernel for deep convolution, and Pointwise means using a 1×1 convolution for channel fusion.
[0063] 1.3 Two-dimensional wavelet transform extension:
[0064]
[0065] The formula shows that the wavelet transform decomposes the input feature map X into low-frequency and high-frequency components through four filters. LL Represents the low-frequency component, reflecting the overall structural information of the input; F LH Represents the high-frequency component in the horizontal direction, which is used to capture the edge details in the horizontal direction; F HL Represents the high-frequency component in the vertical direction, which is used to capture the edge details in the vertical direction; F HH Represents the high-frequency component in the diagonal direction and is used to extract the texture features in the diagonal direction.
[0066] 1.4 Wavelet reconstruction (inverse wavelet transform):
[0067] X IWT =Conv transposed ([F LL F LH F HL F HH ],[X LL X LH X HL X HH ])
[0068] The formula indicates that the feature map X is reconstructed back to the original space using inverse convolution (transposed convolution) and the high and low frequency components decomposed above.
[0069] 1.5 Output feature fusion:
[0070] X out =Concat(X light ,X IWT )
[0071] The formula represents the output X of the lightweight convolution path light and the output X of the wavelet transform path IWT To merge.
[0072] The specific execution process of the module is as follows:
[0073] 1. Feature Decomposition:
[0074] The input feature map X is decomposed into two parts: the first part X 1 , used for lightweight convolution processing and extracting local features; the second part X 2 , used for wavelet transform processing and decomposed into multi-scale high and low frequency features.
[0075] 2. Lightweight convolution feature extraction:
[0076] X 1 Apply depth convolution (channel-by-channel processing) and point-by-point convolution (fusion of channel information) in sequence to generate lightweight features X light , efficiently capture local details of the input.
[0077] 3. Wavelet transform feature expansion
[0078] X 2 Apply a two-dimensional wavelet transform through the filter F LL 、F LH 、F HL 、F HH The features are decomposed into a low-frequency component (global structure) and three high-frequency components (detail information in different directions).
[0079] 4. Wavelet reconstruction
[0080] Use inverse wavelet transform to transform the above components X LL , X LH , X HL , X HH Fusion to reconstruct feature X IWT , recovering multi-scale details and global features.
[0081] 5. Output feature fusion
[0082] The lightweight convolution output X light and wavelet reconstruction feature X IWT Concatenate channels to generate the final output feature X of the module out .
[0083] The present invention is improved based on YOLOv11. The model introduces a multi-scale attention mechanism and enhances the spatial attention of the feature map through the C2PSA_EMA module. This helps the model focus on the specific areas in the image that are most relevant to the detection, thereby improving the performance of small objects such as (mandibular fracture line). This module uses a cross-space learning method to handle short-term and long-term dependencies. Specifically, the model divides the input feature map into two parts, takes one part to strengthen the attention of the salient area through spatial attention generation (EMA module) and context enhancement (FNN module), and extracts local and global context information, and finally performs feature splicing with the other part, so that the model can retain more accurate location information, while weakening irrelevant information in the background area, and improving the accuracy of detection.
[0084] like Figure 3 As shown, the working process of the C2PSA_EMA module is as follows:
[0085] 2.1 Feature Decomposition
[0086] A,B=Split(Conv 1×1 X))
[0087] The formula indicates that through a 1×1 convolution layer, the number of channels is compressed and decomposed into two parts: the retained part A and the part B for further processing, where A directly participates in the final fusion as a global feature; B enters the multi-scale attention path to extract spatial and contextual information.
[0088] 2.2 Spatial Attention Generation (EMA Module)
[0089] B h =Pool h (B),B w =Pool w (B)
[0090] W EMA =σ(Conv 1×1 (Concat(B h ,B w )))
[0091] B EMA =W EMA ·B
[0092] Where σ represents the sigmoid activation function, EMA represents the exponential moving average technique, and the formula represents the average pooling applied to B along the height and width directions to generate B respectively. h and B w , extracting horizontal and vertical global context information; then B h and B w Generate the attention weight matrix W by concatenation and 1×1 convolution EMA , and further strengthen the attention of the salient area through the sigmoid activation function; finally, through weighted operation, the attention weight is applied to the feature map to highlight the target area.
[0093] 2.3 Context Enhancement (FNN Module)
[0094] B FNN =Conv 1×1 (Conv 3×3 (B EMA ))
[0095] B′=B EMA +B FNN
[0096] The formula represents the EMA3×3 convolution and 1×1 convolution are applied in parallel to extract local and global context information respectively; the convolution result is then transformed nonlinearly through a two-layer feedforward network to further enhance the feature expression capability.
[0097] 2.4 Feature Fusion
[0098] X out =Conv 1×1 (Concat(A,B'))
[0099] The formula indicates that the initial retained A and the enhanced B′ are concatenated in the channel dimension, and the number of output channels is restored through a 1×1 convolution to obtain the final feature map X out .
[0100] The specific execution process of the module is as follows:
[0101] 1. Feature Decomposition:
[0102] Apply 1×1 convolution to the input feature X for channel compression and decompose it into two parts:
[0103] A: Global features, directly used for the final splicing; B: Used for subsequent multi-scale attention calculation and enhancement.
[0104] 2. Spatial Attention Generation (EMA Module):
[0105] Average pooling of B: Generate B along the height direction h , generate B along the width direction w . Splice B h and B w , generate the attention weight W through 1×1 convolution EMA , and highlight the target area through sigmoid activation. EMA Applied to B, generating attention-weighted feature B EMA .
[0106] 3. Context Enhancement (FNN Module)
[0107] To B EMA Local (3×3 convolution) and global (1×1 convolution) feature extraction are performed. The 3×3 convolution captures the edge details and morphological changes of the fracture line through the local spatial convolution kernel, while the 1×1 convolution strengthens the feature response related to the fracture semantics through cross-channel weight adjustment. The parallel processing of the two realizes the synergistic enhancement of spatial details and channel context, and then nonlinear enhancement is performed through the feedforward network. The enhanced result is compared with the original B EMA Add together, retain the original information through residual connection, and finally obtain the enhanced feature B′.
[0108] 4. Feature Fusion
[0109] The retained feature A and the enhanced feature B′ are concatenated, and the number of output channels is restored through 1×1 convolution to obtain the final feature map X out .
[0110] Through the above calculations, the C2PSA_EMA module achieves efficient fusion of global and local features. The module uses the attention weight matrix generated by the spatial attention mechanism (EMA module) to highlight the salient areas, suppress background interference, and enhance the global expression ability of the feature map; through context enhancement (FNN module), it fuses local and global information to further improve the distinguishing ability and expression richness of the features; finally, through channel compression and splicing operations, it fully utilizes the global information of the initial features and the enhanced detail information, and outputs a high-quality feature map with both global and local features.
[0111] The present invention integrates a wavelet convolution module in the feature fusion network (Neck), adopts a multi-level feature extraction architecture and a bidirectional fusion strategy to aggregate feature maps from different resolutions and pass them to the corresponding detection head. Specifically, the above-mentioned C3k2_WT module is integrated in the neck network to improve the speed and performance of feature aggregation, and upsampling and splicing layers are applied to combine feature maps of different scales. After splicing, the C3k2_WT module is used to ensure efficient feature aggregation, which makes a very important contribution to improving accuracy in relatively blurred CT images.
[0112] At the same time, based on the above feature extraction and fusion, the model uses the detection head part (Head) of YOLOv11 to perform the positioning prediction task of the mandibular fracture line. The detection head part will output the bounding box, class probability and confidence score. Specifically, the model has three detection heads, called shallow detection heads, middle detection heads and deep detection heads, which make predictions on feature maps of different scales, and each detection head is responsible for targets of different sizes. During training, each detection head will participate in the loss calculation. During inference, the three detection heads will generate prediction results, and then merge the results through non-maximum suppression (NMS). The detection of small targets mainly relies on the shallow detection head, and the middle and deep detection heads will have a small amount of supplementary predictions for small targets (especially when small targets are in complex backgrounds).
[0113] The specific working process of the feature fusion network and the detection head is as follows:
[0114] Feature maps (P3, P4, P5) of multiple scales are extracted from the backbone network (Backbone). The P3 feature map is iteratively extracted by two layers of C3k2_WT modules to output a high-resolution feature map, which retains rich spatial detail information and is suitable for small target detection. The P4 feature map is further extracted by the third layer of C3k2_WT module on the basis of P3, with the resolution reduced to balance the details and semantic information. The P5 feature map is based on P4 and is sequentially compressed by the fourth layer of C3k2_WT module and the spatial pyramid fast pooling module (SPPF). The C2PSA_EMA module is introduced to enhance channel attention, and finally a low-resolution feature map is generated. The feature map has high-level semantic representation capabilities. The P4 and P5 feature maps will be used for secondary splicing later. To achieve cross-scale feature complementarity, this scheme adopts a bidirectional fusion strategy: first, perform top-down semantic enhancement, upsample low-resolution feature maps (such as P4, P5) through bilinear interpolation to match their size with the high-resolution feature map, thereby mapping high-level semantic information to a finer pixel space. Specifically, perform bilinear interpolation upsampling on P5 to align its size with P4, and splice along the channel dimension to generate the fused feature map P4. new , and then P4 new After upsampling, concatenate with P3 to generate P3 new ; Then transfer the details from bottom to top, and make the spliced P3 new The C3k2_WT module is introduced to perform feature compression and nonlinear enhancement, complete downsampling, and compare it with the original P4 and P5 features. Figure 2 splicing to form a closed-loop fusion path and generate P4 new ', P5 new ', ensuring the two-way interaction between shallow details and deep semantics, while improving the speed and effect of multi-scale feature aggregation. The fused P3 new 、P4 new ', P5 new 'The feature maps are passed to three independent detection heads respectively, and multi-scale anchor frames are preset for feature maps of different resolutions. The small-scale anchor frames densely cover P3 new High-resolution grid and large-scale anchor box adapted to P5 new ', each detection head synchronously outputs bounding box coordinates, category probability and confidence, and the results are fused across scales through non-maximum suppression (NMS), ultimately retaining high-precision predictions.
[0115] The mandibular fracture detection system proposed in the present invention significantly improves the accuracy and robustness of fracture line detection by combining lightweight feature extraction, spatial attention optimization and multi-scale feature fusion. The C3k2_WT module and C2PSA_EMA module are introduced in the model design to achieve efficient modeling of global and local features, and the feature expression ability of the model in complex fracture morphology and subtle changes is enhanced through the optimized feature fusion strategy. The optimized feature extraction and target detection strategy further improves the detection accuracy and positioning performance, providing an efficient and reliable solution for the automatic detection of mandibular fractures, which has important clinical application value and promotion potential.
[0116] Experimental results show that the mandibular fracture line detection system of the present invention outperforms existing mainstream methods in key performance indicators such as detection accuracy, recall rate, and F1 score. Through the multi-scale feature capture capability of wavelet transform and the significant regional enhancement of the spatial attention mechanism, the present invention can achieve richer feature expression in complex mandibular fracture detection tasks, especially showing superior performance in the detection of subtle fracture lines and small targets. The improved model shows strong robustness and adaptability when dealing with challenging fracture morphologies and complex background structures, and can effectively cope with CT image data of different patients, significantly reducing false detections and missed detections.
[0117] To verify the effect of the present invention, the present invention is tested below.
[0118] S1. Dataset
[0119] The original data used in the present invention comes from the CT images of the Children's Hospital Affiliated to Zhejiang University School of Medicine, including 284 fracture samples and 300 non-fracture (no fracture or only skull fracture) control samples. The present invention is mainly aimed at the automatic detection of mandibular fracture lines in minors. The age range of the 284 fracture samples ranges from 1 month to 15 years and 5 months, including 182 males and 102 females.
[0120] The data is large in scale, with 156,806 images in the training set and 39,202 images in the test set, with 24,008 and 5,989 annotations respectively, covering the three slice directions of coronal, sagittal and axial planes, which can effectively capture the multi-angle characteristics of the mandibular fracture line. Each image is accompanied by an annotation file, and medical experts annotate the fracture line in yolo format on Labelimg to generate a txt annotation file (the parameters are the label, the x, y coordinates of the center of the image, and the length and width of the image). A label of 0 indicates the presence of a fracture line.
[0121] S2. Baseline and Evaluation Metrics
[0122] In order to comprehensively evaluate the effectiveness of the mandibular fracture detection model (WE-YOLO) constructed by the present invention in the target detection task, the present invention sets up a series of baseline models and systematically compares and verifies their performance through multiple evaluation indicators.
[0123] Baseline model setting: In order to verify the improvement effect of the WE-YOLO model in fracture line detection, this study selected the following baseline models for comparative experiments:
[0124] (1) YOLOv5 model: We use the lightweight YOLOv5 model as a comparison baseline and focus on evaluating its performance in the fracture line detection task, especially in terms of the balance between inference speed and model accuracy.
[0125] (2) YOLOv8 model: The YOLOv8 standard model is selected as the strong baseline model, focusing on comparing the differences between the WE-YOLO model and the previous generation detection model in terms of accuracy, robustness, and small target detection capabilities.
[0126] (4) YOLOv11 model: Without adding the improved module, the YOLOv11 standard model was used alone to predict the fracture line as an important baseline for comparative experiments to evaluate the effect of the C3k2_WT module and the C2PSA_EMA module on improving the model performance.
[0127] (5) RT-DETR, EfficientDet, and faster-rcnn: Mainstream target detection models are introduced as baselines, including RT-DETR, EfficientDet, and Faster R-CNN, to comprehensively evaluate the advantages of the improved YOLOv11 model in detection efficiency and accuracy, especially its adaptability when processing complex medical images.
[0128] The above baseline models cover a variety of detection structures from lightweight models to large-scale models, which can systematically verify the innovation and effectiveness of the WE-YOLO model. Through comparative analysis, we can comprehensively evaluate the advantages of the present invention in improving small target detection capabilities, enhancing complex background adaptability, and optimizing multi-scale feature expression.
[0129] Evaluation indicators: The present invention evaluates the performance of the model through the following indicators:
[0130] (1) True Positive (TP): The number of instances correctly predicted by the model as positive. That is, the instances that are actually positive and correctly identified as positive by the model.
[0131] (2) False Positive (FP): The number of instances where the model incorrectly predicts the negative class as the positive class. This means that the instances actually belong to the negative class but are incorrectly labeled as the positive class.
[0132] (3) True Negative (TN): The number of instances correctly predicted by the model as negative. That is, the instances that are actually negative and correctly identified as negative by the model.
[0133] (4) False Negative (FN): The number of instances where the model incorrectly predicts the positive class as the negative class. That is, the instances that actually belong to the positive class but are mistakenly labeled as negative.
[0134] (5) Accuracy: It is the ratio of all correctly predicted samples (TP+TN) to the total number of samples. The calculation formula is: It measures the overall correctness of the model.
[0135] (6) Precision: The proportion of samples that are actually positive among all samples predicted by the model to be positive. The calculation formula is: This reflects how accurate the model is in predicting the positive class.
[0136] (7) Recall: Also known as sensitivity or detection rate, it is the proportion of samples that are actually positive and are correctly identified by the model. The calculation formula is: A higher Recall value indicates that the model has a low missed detection rate and can cover most of the actual fracture lines, which is crucial to reducing the risk of missed diagnosis in medical applications.
[0137] (8) F1 Score: It is the harmonic mean of precision and recall, and is used to comprehensively evaluate the performance of the model. It is used when both precision and recall are important. The calculation formula is: The higher the F1 score, the better the overall performance of the model.
[0138] (9) Average Precision (AP): For binary classification problems, AP is equivalent to Precision. This indicator reflects the model's ability to detect fracture lines at different confidence levels by calculating the area under the Precision-Recall curve. A higher AP50 value indicates that the model has better performance in locating and classifying the fracture line area and can accurately capture the fracture line features.
[0139] S3. Implementation details
[0140] The experiment of this invention is implemented on the PyTorch 12.4 framework, using the Python 3.10 programming language, and trained on the RTX3090 GPU. To ensure the convergence and stability of the model, the experiment sets a unified training strategy and hyperparameter configuration. The specific implementation details are as follows:
[0141] (1) Training Epochs: The training epoch is set to 200 to ensure that the model can fully learn the data features and capture the complex patterns of fracture lines, while avoiding overfitting due to too many training epochs. An appropriate number of training epochs helps the model gradually optimize its weights and improve detection accuracy.
[0142] (2) Batch Size: The batch size is set to 64. A moderate batch size can balance memory usage and computational efficiency while ensuring the stability of gradient calculation. When using a larger GPU, a larger batch size helps to increase the training speed and also enables the optimizer to have more data support each time the parameters are updated, thereby improving the stability of training.
[0143] (3) Optimizer: The Adam optimizer is used to quickly find the direction in the early iterations in high-dimensional space.
[0144] (4) Learning Rate: The initial learning rate is set to 0.01. This learning rate is used to control the convergence speed of the model to prevent overly fast or overly slow convergence.
[0145] The above implementation details ensure the stability and convergence of the WE-YOLO mandibular fracture detection model of the present invention during the training process, and provide a reliable experimental basis for the final fracture detection accuracy.
[0146] S4. Experimental Results
[0147] 4.4.1 Quantitative analysis
[0148] This paper systematically and quantitatively analyzes the prediction effect of the mandibular fracture detection model (WE-YOLO model) of the present invention by comparing the performance indicators of different models. The experimental results show that WE-YOLO significantly outperforms the YOLOv11 baseline model in all key indicators, and also surpasses other comparison models, verifying the effectiveness and robustness of the model. The following are the main analysis results:
[0149] (1) Axial surface:
[0150] The WE-YOLO model outperforms other models in AP, Precision, Recall, and F1 scores in the axial plane, with AP: 0.89309, Precision: 0.880, and F1 score: 0.85.
[0151] Compared with the YOLOv11 baseline model (AP: 0.81737, Precision: 0.770, F1: 0.76), AP is increased by about 7.6%, Precision is increased by 11.4%, and F1 score is increased by 9.2%, significantly enhancing the ability to capture axial surface features and detect small targets.
[0152] Yolov8+ContextAggregation performs second best in AP and F1 scores (AP: 0.883, F1: 0.831), but still lower than the improved YOLOv11.
[0153] The AP of YOLOv5, Faster R-CNN and EfficientDet are 0.733, 0.736 and 0.715 respectively. The Precision and Recall performances are significantly behind, indicating that they are difficult to adapt to complex fracture detection scenarios.
[0154] (2) Sagittal plane:
[0155] The WE-YOLO model is far superior to other models in terms of AP, Precision, Recall and F1 score in the sagittal plane, with AP: 0.8797, Precision: 0.830, and F1 score: 0.83.
[0156] Compared with YOLOv11 (AP: 0.88519, Precision: 0.830, F1: 0.84), the performance has slightly decreased, but its Precision and Recall remain at a high level, indicating that wavelet convolution and attention mechanisms still have strong adaptability in complex backgrounds.
[0157] Yolov8+ContextAggregation maintains high performance in the sagittal plane (AP: 0.882, F1: 0.839), but fails to surpass the improved models of the YOLOv11 series.
[0158] Compared with other models, the values of other models on the sagittal plane have dropped significantly, especially the DETR model, whose mAP50 on the sagittal plane has dropped by 0.2-0.3 compared with other planes, indicating that the adaptability of the model is limited. The WE-YOLO model still maintains a high mAP50, precision and recall value on the sagittal plane, which shows that it is well adapted to complex scenes.
[0159] (3) Coronal plane:
[0160] The WE-YOLO model outperforms other models in AP, Precision, Recall, and F1 scores on the coronal plane, with an AP of 0.92662, a Precision of 0.860, and an F1 score of 0.88. This shows that our model works best when the plane is the coronal plane.
[0161] Compared with the YOLOv11 baseline model (AP: 0.86335, Precision: 0.830, F1: 0.82), the AP is increased by 7.3%, the Precision is increased by 3.6%, and the F1 score is increased by 6.8%, showing the superiority of multi-scale feature fusion and spatial attention mechanism.
[0162] The AP of Yolov8+ContextAggregation is 0.913, but it fails to surpass WE-YOLO in terms of Precision and F1 scores. The quantitative analysis results of the present invention verify the superiority of the WE-YOLO mandibular fracture detection model in fracture feature extraction and accurate detection, and provide strong technical support for the application of the model in medical imaging.
[0163] The AP of RT-DETR in the coronal plane is close to WE-YOLO (0.926), but the F1 score is slightly inferior, which further verifies the strong adaptability of the improved model to small target detection.
[0164] The quantitative analysis results of the present invention verify the superiority of the WE-YOLO mandibular fracture detection model in fracture feature extraction and accurate detection, and provide strong data support for the application of the model in medical imaging.
[0165] 4.4.2 Qualitative analysis
[0166] This paper further verifies the effectiveness of the mandibular target detection method based on wavelet convolution and multi-scale attention improved YOLOv11 through qualitative analysis, focusing on the advantages of this method in feature extraction, attention optimization and multi-scale feature fusion. Through the visualization analysis of the detection results and feature attention areas of specific samples, it can be clearly seen how the model uses the improved YOLOv11 structure to improve the accuracy and robustness of mandibular fracture detection.
[0167] (1) Wavelet convolution enhances the effectiveness of local detail extraction:
[0168] In the mandibular fracture detection process of the WE-YOLO model, the C3k2_WT module captures multi-scale features through wavelet convolution, significantly enhancing the model's ability to focus on local details. Through the visualization analysis of the detection results, it can be seen that the model can effectively retain detailed information, laying the foundation for subsequent feature extraction, thereby significantly improving the accuracy of fracture line detection. This shows that wavelet convolution can effectively improve the feature expression ability in fracture detection tasks.
[0169] (2) Effectiveness of multi-scale attention focusing on key parts
[0170] In the mandibular fracture detection process of the WE-YOLO model, the C2PSA_EMA module focuses on key areas through the spatial attention mechanism, further enhancing the model's sensitivity to the mandibular fracture line. It receives detailed information from the C3k2_WT module, highlights the significance of the fracture area, and suppresses the interference of background noise, thereby significantly improving the detection robustness of the model in complex backgrounds.
[0171] (3) Complementarity of multi-scale feature fusion:
[0172] The WE-YOLO model combines feature maps of different scales and enhances the comprehensive judgment ability of fracture lines of different sizes through a multi-scale feature fusion mechanism. When dealing with complex fracture structures, the model can capture subtle boundary information from high-resolution features while obtaining global context from low-resolution features. Even when the fracture line is relatively small or the boundary is blurred, the model can still accurately locate the fracture area, significantly reducing the probability of missed detection and false detection. This complementarity of multi-scale features effectively enhances the adaptability and detection accuracy of the model.
[0173] In general, the automatic detection system for mandibular fractures of the present invention significantly improves the detection accuracy and robustness of the model through wavelet convolution optimization, multi-scale attention mechanism and multi-scale feature fusion, and has broad clinical application prospects.
[0174] Tables 1 to 3 show the experimental results.
[0175] Table 1
[0176]
[0177] Table 2
[0178]
[0179]
[0180] Table 3
[0181]
[0182] In summary, the qualitative analysis of the present invention further demonstrates the significant advantages of the present invention in feature extraction, multi-scale fusion and fracture line localization, proving the good performance and high accuracy of the model in the task of mandibular fracture detection.
[0183] The embodiments described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A mandibular fracture detection system based on wavelet transform feature enhancement and multi-scale attention focusing, characterized in that: The invention comprises a computer memory, a computer processor and a computer program stored in the computer memory and executable on the computer processor, wherein the computer memory stores a trained mandibular fracture detection model; the mandibular fracture detection model comprises a backbone network, a feature fusion network and a detection head part; The backbone network is used to extract feature maps P3, P4 and P5 of different scales. The backbone network first uses two convolutional layers to downsample the input image, and then iteratively extracts the P3 feature map through two layers of C3k2_WT modules. Based on the P3 feature map, the P4 feature map is further extracted through the third layer of C3k2_WT module. Based on the P4 feature map, the fourth layer of C3k2_WT module, the spatial pyramid fast pooling module SPPF and the C2PSA_EMA module are sequentially passed to finally generate a low-resolution P5 feature map with high-level semantic representation capabilities. The feature fusion network adopts a bidirectional fusion strategy, first performing top-down semantic enhancement, upsampling the P5 feature map by bilinear interpolation to align its size with the P4 feature map, and splicing it along the channel dimension to generate the fused feature map P4. new , and then P4 new After upsampling, it is concatenated with the P3 feature map to generate P3 new ; Then transfer the details from bottom to top, and make the spliced P3 new A C3k2_WT module is introduced to perform feature compression and nonlinear enhancement, complete downsampling, and re-join with the original P4 and P5 feature maps to form a closed-loop fusion path to generate P4 new ', P5 new '; The detection head part contains three independent detection heads, and preset multi-scale anchor frames for feature maps of different resolutions; the fused P3 new 、P4 new ', P5 new 'The feature map is passed to three independent detection heads, and the small-scale anchor boxes densely cover P3 new High-resolution grid, medium-scale anchor box adapted to P4 new 'Medium-resolution grid, large-scale anchor box adapted to P5 new '; Each detection head outputs bounding box coordinates, category probability and confidence simultaneously, and the outputs of the three detection heads are fused to finally retain high-precision predictions; When the computer processor executes the computer program, the following steps are implemented: The image to be detected is input into the trained mandibular fracture detection model to obtain the prediction result of mandibular fracture.
2. The mandibular fracture detection system based on wavelet transform feature enhancement and multi-scale attention focusing according to claim 1 is characterized in that: The backbone network first uses two convolutional layers to downsample the input image, the formula is as follows: O1=Conv(I,64,3,2) O2=Conv(O1,128,3,2) Among them, O1 represents the output feature map after the first convolution layer, and O2 represents the output feature map after the second convolution layer; the input data I undergoes the first convolution operation, using 64 convolution kernels of size 3×3 and a stride of 2 to obtain the output feature map O1; the first convolution operation is performed again, using 128 convolution kernels of size 3×3 and a stride of 2 to obtain the new output feature map O2.
3. The mandibular fracture detection system based on wavelet transform feature enhancement and multi-scale attention focusing according to claim 1 is characterized in that: The working process of the C3k2_WT module is as follows: First, the feature map of the input C3k2_WT module is decomposed to obtain feature maps X1 and X2; Perform lightweight convolution feature extraction on feature map X1 to obtain X light ; The wavelet transform decomposes the feature map X2 into a low-frequency component F through four filters. LL and three high frequency components F LH 、F HL and F HH , and then use the inverse wavelet transform to fuse the low-frequency component and the three high-frequency components to reconstruct the feature map X2 with the same spatial resolution. IWT ; The obtained X light and X IWT Fusion is performed to obtain feature fusion results for subsequent operations.
4. The mandibular fracture detection system based on wavelet transform feature enhancement and multi-scale attention focusing according to claim 3 is characterized in that: Perform lightweight convolution feature extraction on the feature map X1. The formula is as follows: X light =Conv light (X1)=Pointwise(Depthwise(X1,k=2)) In the formula, Depthwise means using a 2×2 small kernel for depth convolution, Pointwise means using a 1×1 convolution for channel fusion, and k=2 means that the size of the convolution kernel is 2×2.
5. The mandibular fracture detection system based on wavelet transform feature enhancement and multi-scale attention focusing according to claim 3 is characterized in that: A low frequency component is F LL , reflecting the overall structural information of the input; the three high-frequency components are F LH 、F HL and F HH ; F LH Represents the high-frequency component in the horizontal direction, which is used to capture the edge details in the horizontal direction; F HL Represents the high-frequency component in the vertical direction, which is used to capture the edge details in the vertical direction; F HH Represents the high-frequency component in the diagonal direction and is used to extract the texture features in the diagonal direction.
6. The mandibular fracture detection system based on wavelet transform feature enhancement and multi-scale attention focusing according to claim 1 is characterized in that: The working process of the C2PSA_EMA module is as follows: After the feature map is input into the C2PSA_EMA module, it passes through a 1×1 convolutional layer to compress the number of channels and decompose it into retained feature A and feature B for further processing; feature A directly participates in the final fusion as a global feature; feature B enters the multi-scale attention path to extract spatial and contextual information; Apply average pooling along the height and width directions to feature B to generate B respectively. h and B w ; Then B h and B w The attention weight matrix is generated by concatenation and 1×1 convolution, and the attention of the salient area is further strengthened by the sigmoid activation function σ to obtain the final attention weight matrix W EMA ; Finally, through weighted operation, the attention weight is applied to the feature map B to highlight the target area and obtain the attention weighted feature B EMA ; To B EMA 3×3 convolution and 1×1 convolution are applied in parallel to extract local spatial correlation and global channel context dependency respectively: 3×3 convolution captures texture and edge details in the local neighborhood by expanding the spatial receptive field, and 1×1 convolution models the semantic dependency between global channels through cross-channel linear combination; then the convolution result B of the two is concatenated. FNN Through a two-layer feedforward network (including nonlinear activation function), nonlinear transformation is performed to further enhance the feature expression ability and compare the enhanced result with the original B EMA Add, retain the original information through residual connection, and finally obtain the enhanced feature B′; The retained feature A and the enhanced feature B′ are concatenated in the channel dimension, and the number of output channels is restored through a 1×1 convolution to obtain the final feature map for subsequent operations.
7. The mandibular fracture detection system based on wavelet transform feature enhancement and multi-scale attention focusing according to claim 1 is characterized in that: In the detection head part, the outputs of the three detection heads are fused across scales through non-maximum suppression (NMS).
8. According to the mandibular fracture detection system based on wavelet transform feature enhancement and multi-scale attention focusing as claimed in claim 1, the data set for training the mandibular fracture detection model is constructed in the following process: Images were extracted from maxillofacial CT images in DICOM format and a dataset was constructed. The fracture area was annotated, including length, width, and center coordinates. The images were resized and formatted in a standardized manner to ensure the consistency and adaptability of the input data. Finally, the dataset was expanded through data enhancement.
Citation Information
Patent Citations
Infrared weak and small target detection method and system based on conflict filtering and feature focusing
CN116740371A
Infrared target detection method and model based on convolution attention
CN119273891A
X-ray image contraband detection method based on multi-scale feature fusion
CN119295886A
Contextual visual-based SAR target detection method and apparatus, and storage medium
US20230184927A1
Cited By
Microbial microscopic image target identification method based on wavelet enhanced convolutional neural network
CN120580689A