Wrist image fracture recognition and detection method and system based on improved YOLOv8 network
By improving the YOLOv8 network, adopting the fast network backbone structure and the bidirectional feature pyramid network structure, the problems of insufficient recognition ability and low computing efficiency of wrist fracture detection in complex backgrounds in the prior art are solved, and more efficient feature extraction and fusion are achieved, and detection accuracy and real-time performance are improved.
Patent Information
- Application Number
- CN202510474708.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing wrist fracture detection method based on convolutional neural networks has insufficient recognition ability to capture small target features and complex backgrounds, and is large in calculation and slow inference speed, making it difficult to meet the real-time requirements.
The improved YOLOv8 network is adopted to enhance the efficiency of feature extraction and fusion by introducing a bidirectional feature pyramid network structure of fast network backbone structure, global to local spatial aggregation module and dynamic cross-stage deformable convolutional network module.
It significantly improves the extraction accuracy and calculation efficiency of fracture characteristics, enhances the ability to capture fracture characteristics in complex backgrounds, and improves the accuracy and real-time recognition of wrist fracture images.
Smart Images

Figure CN119992603A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a wrist image fracture recognition and detection method and system based on an improved YOLOv8 network. Background Art
[0002] Medical imaging diagnostic methods based on artificial intelligence are gradually being widely used in the field of fracture diagnosis. More and more researchers are using convolutional neural networks (CNNs) for fracture detection and classification. At present, automated diagnosis based on convolutional neural networks mainly involves inputting wrist fracture X-ray images into the model, and then extracting local and global features from the images through multi-layer convolution and pooling operations. Finally, the fully connected layer or Softmax activation function is used to classify the features and output the prediction results. The relevant technology still has certain limitations. First, the fracture area in wrist fracture images is usually small and irregular in shape. Traditional convolution operations are limited by the receptive field and it is difficult to accurately capture the features of small targets. Secondly, the complex background and low contrast in X-ray images further increase the difficulty of detection. The existing feature fusion module is insufficient in the aggregation of detail information, which makes the model prone to missed detection or misdetection under complex imaging conditions. Finally, existing models usually use more complex feature extraction modules, which are computationally intensive and lack lightweight design. The reasoning speed is slow and it is difficult to meet real-time requirements. Summary of the invention
[0003] In order to solve the above technical problems, the purpose of the present invention is to provide a wrist image fracture recognition and detection method and system based on an improved YOLOv8 network, which can enhance the ability to capture wrist fracture features under complex backgrounds and the efficiency of feature extraction, thereby improving the accuracy of wrist fracture image recognition.
[0004] The first technical solution adopted by the present invention is: a wrist image fracture recognition and detection method based on an improved YOLOv8 network, comprising the following steps: Acquire a wrist fracture image data set to be processed and perform data preprocessing to obtain a wrist fracture image; Based on the fast network backbone structure, a bidirectional feature pyramid network structure integrating the global to local spatial aggregation module and the dynamic cross-stage deformable convolutional network module is introduced to build an improved YOLOv8 recognition network model. Based on the improved YOLOv8 recognition network model, image recognition processing is performed on wrist fracture images to obtain wrist image fracture recognition results.
[0005] Further, the step of obtaining a wrist fracture image dataset to be processed and performing data preprocessing to obtain a wrist fracture image specifically includes: Obtain a wrist fracture image dataset to be processed; Performing data cleaning and labeling on the wrist fracture image dataset to be processed to obtain a labeled wrist fracture image dataset; Data enhancement and segmentation are performed on the labeled wrist fracture image dataset to obtain wrist fracture images.
[0006] Furthermore, the improved YOLOv8 recognition network model includes a backbone network module, a neck network module and a detection head network module, and the backbone network module, the neck network module and the detection head network module are connected in sequence, wherein: The backbone network module adopts a fast network backbone structure. The backbone network module specifically includes an embedding layer, a fast network module, a backbone fusion block, a global pooling layer, a backbone convolution block, a fully connected layer and a spatial pyramid fast pooling module. The fast network module includes a depth-separable convolution layer and an attention mechanism. The neck network module adopts a bidirectional feature pyramid network structure. The neck network module specifically includes a global to local space aggregation module, an upsampling module, a neck fusion block and a dynamic cross-stage deformable convolutional network module. The global to local space aggregation module includes a first feature segmentation module, a first two-dimensional convolution module, a second two-dimensional convolution module, a global to local space aggregation convolution branch module, a context information aggregation module, a first feature splicing module and a first convolution block. The dynamic cross-stage deformable convolutional network module includes a second convolution block, a second feature segmentation module, several dynamic deformable convolution bottleneck modules and a second feature splicing module.
[0007] Furthermore, it also includes introducing a generalized intersection-and-union loss function based on the internal intersection-and-union loss function, and constructing an internal generalized intersection-and-union loss function as a loss function of the improved YOLOv8 recognition network model. The expression of the internal generalized intersection-and-union loss function is specifically as follows: ; In the above formula, represents the internal generalized intersection-over-union loss function, represents the internal intersection-over-union loss function, represents the generalized intersection and union ratio, represents the area of the minimum enclosing rectangle of two bounding boxes, Represents the area of the union.
[0008] Further, the step of performing image recognition processing on the wrist fracture image based on the improved YOLOv8 recognition network model to obtain the wrist image fracture recognition result specifically includes: Input the wrist fracture image into the improved YOLOv8 recognition network model; Based on the backbone network module of the improved YOLOv8 recognition network model, feature extraction is performed on the wrist fracture image to obtain the wrist fracture feature image; Based on the neck network module of the improved YOLOv8 recognition network model, feature fusion processing is performed on the wrist fracture feature image to obtain the fused wrist fracture feature image; Based on the detection head network module of the improved YOLOv8 recognition network model, target recognition detection is performed on the fused wrist fracture feature image to obtain the wrist image fracture recognition result.
[0009] Furthermore, the backbone network module based on the improved YOLOv8 recognition network model performs feature extraction processing on the wrist fracture image to obtain the wrist fracture feature image, which specifically includes: Input the wrist fracture image into the backbone network module of the improved YOLOv8 recognition network model; Based on the embedding layer of the backbone network module, data dimension reduction processing is performed on the wrist fracture image to obtain a wrist fracture image after dimension reduction; The fast network module based on the backbone network module extracts local detail features and global structural features from the wrist fracture image after dimensionality reduction to obtain a preliminary wrist fracture feature image. Based on the backbone fusion block of the backbone network module, multi-resolution feature fusion processing is performed on the preliminary wrist fracture feature image to obtain a fused wrist fracture feature image; Based on the global pooling layer of the backbone network module, the fused wrist fracture feature image is globally pooled and compressed to obtain a compressed wrist fracture feature image; Based on the backbone convolution block of the backbone network module, convolution processing is performed on the compressed wrist fracture feature image to obtain a convolved wrist fracture feature image; Based on the fully connected layer of the backbone network module, the convolutional wrist fracture feature image is mapped to obtain a mapped wrist fracture feature image; Based on the spatial pyramid fast pooling module of the backbone network module, multi-scale aggregation processing is performed on the mapped wrist fracture feature image to obtain the wrist fracture feature image.
[0010] Furthermore, the neck network module based on the improved YOLOv8 recognition network model performs feature fusion processing on the wrist fracture feature image to obtain a fused wrist fracture feature image, which specifically includes: Input the wrist fracture feature image into the neck network module of the improved YOLOv8 recognition network model; Based on the global to local space aggregation module of the neck network module, the wrist fracture feature image is subjected to global attention weighted calculation and feature aggregation processing to obtain the wrist fracture global feature image; An upsampling module based on the neck network module performs upsampling processing on the global feature image of wrist fracture to obtain an upsampled global feature image of wrist fracture; Based on the neck fusion block of the neck network module, multi-scale feature fusion processing is performed on the upsampled global feature image of wrist fracture to obtain a preliminary fused global feature image of wrist fracture; Based on the dynamic cross-stage deformable convolutional network module of the neck network module, cross-stage feature optimization processing is performed on the global feature image of wrist fracture after preliminary fusion to obtain the feature image of wrist fracture after fusion.
[0011] Furthermore, the global to local space aggregation module based on the neck network module performs global attention weighted calculation and feature aggregation processing on the wrist fracture feature image to obtain the global feature image of the wrist fracture, which specifically includes: The wrist fracture feature image is input into the global to local space aggregation module of the neck network module; Based on the first feature segmentation module of the global to local space aggregation module, feature segmentation processing is performed on the wrist fracture feature image to obtain a segmented wrist fracture feature image; Based on the first branch two-dimensional convolution module and the second branch two-dimensional convolution module of the global to local space aggregation module, convolution processing is performed on the segmented wrist fracture feature image to obtain a first convolution wrist fracture feature image and a second convolution wrist fracture feature image; A global-to-local spatial aggregation convolution branch module based on the global-to-local spatial aggregation module performs spatial feature aggregation processing on the first convolution wrist fracture feature image to obtain a first aggregated wrist fracture feature image; Based on the context information aggregation module of the global to local space aggregation module, the context information aggregation processing is performed on the second convolution wrist fracture feature image to obtain a second aggregated wrist fracture feature image; Based on the second feature splicing module of the global to local space aggregation module, the first aggregated wrist fracture feature image and the second aggregated wrist fracture feature image are spliced to obtain a spliced wrist fracture feature image; Based on the first convolution block of the global to local spatial aggregation module, the spliced wrist fracture feature image is convolved to obtain the global feature image of wrist fracture.
[0012] Furthermore, the dynamic cross-stage deformable convolutional network module based on the neck network module performs cross-stage feature optimization processing on the wrist fracture global feature image after preliminary fusion to obtain the fused wrist fracture feature image, which specifically includes: The global feature image of wrist fracture after preliminary fusion is input into the dynamic cross-stage deformable convolutional network module of the neck network module; Based on the second convolution block of the dynamic cross-stage deformable convolutional network module, the wrist fracture global feature image after preliminary fusion is convolved to obtain the convolved wrist fracture global feature image; The second feature segmentation module based on the dynamic cross-stage deformable convolutional network module performs feature segmentation processing on the convolved wrist fracture global feature image to obtain a segmented wrist fracture global feature image; A dynamic deformable convolution bottleneck module based on a dynamic cross-stage deformable convolution network module performs dynamic feature optimization processing on the segmented wrist fracture global feature image to obtain an optimized wrist fracture feature image; Based on the second feature splicing module of the dynamic cross-stage deformable convolutional network module, feature splicing and fusion processing is performed on the optimized wrist fracture feature image to obtain a fused wrist fracture feature image.
[0013] The second technical solution adopted by the present invention is: a wrist image fracture recognition and detection system based on an improved YOLOv8 network, comprising: The first module is used to obtain a wrist fracture image data set to be processed and perform data preprocessing to obtain a wrist fracture image; The second module is used to build an improved YOLOv8 recognition network model based on a fast network backbone structure by introducing a bidirectional feature pyramid network structure that integrates a global to local spatial aggregation module and a dynamic cross-stage deformable convolutional network module; The third module is used to perform image recognition processing on wrist fracture images based on the improved YOLOv8 recognition network model to obtain wrist image fracture recognition results.
[0014] The method and system of the present invention have the following beneficial effects: the present invention obtains a wrist fracture image data set to be processed and performs data preprocessing, and further introduces a bidirectional feature pyramid network structure that integrates a global to local space aggregation module and a dynamic cross-stage deformable convolutional network module based on a fast network backbone structure, optimizes the backbone structure of a YOLOv8 network through the fast network backbone structure, and enables the model to perform well in multi-scale feature extraction, and can significantly improve the extraction accuracy and computational efficiency of fracture features while reducing the number of parameters, and further through the global to local space aggregation module, can more meticulously aggregate local spatial information while maintaining the global receptive field, enhance the ability to capture fracture features under complex backgrounds, thereby improving the detection accuracy of the model, and further improves the accuracy of feature extraction and fusion through the dynamic cross-stage deformable convolutional network module, and finally performs image recognition processing on wrist fracture images based on an improved YOLOv8 recognition network model, thereby improving the accuracy of wrist fracture image recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a flowchart of the steps of the wrist image fracture recognition and detection method based on the improved YOLOv8 network of the present invention; Figure 2 It is a structural block diagram of the wrist image fracture recognition and detection system based on the improved YOLOv8 network of the present invention; Figure 3 It is a structural diagram of an improved YOLOv8 recognition network model provided by a specific embodiment of the present invention; Figure 4 is a schematic diagram of the structure of a backbone network provided by a specific embodiment of the present invention; Figure 5 It is a structural diagram of a fused bidirectional feature pyramid network provided by a specific embodiment of the present invention; Figure 6 is a schematic diagram of the structure of a dynamic deformable convolution module provided by a specific embodiment of the present invention; Figure 7 is a schematic diagram of the structure of a dynamic deformable convolution bottleneck module provided by a specific embodiment of the present invention; Figure 8 It is a structural schematic diagram of a dynamic cross-stage deformable convolutional network module provided by a specific embodiment of the present invention; Fig. 9 It is a structural schematic diagram of a global to local spatial aggregation convolution branch module provided in a specific embodiment of the present invention; Fig.10 It is a structural schematic diagram of a global to local space aggregation module provided in a specific embodiment of the present invention; Fig.11 is a schematic diagram of the changes in the heat map before and after the network improvement provided by a specific embodiment of the present invention; Fig.12 It is a schematic diagram of detecting changes before and after network improvement provided by a specific embodiment of the present invention. DETAILED DESCRIPTION
[0016] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. The step numbers in the following embodiments are only provided for the convenience of explanation and description, and the order between the steps is not limited in any way. The execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0017] Reference Figure 1 The present invention provides a wrist image fracture recognition and detection method based on an improved YOLOv8 network, the method comprising the following steps: S100, obtaining a wrist fracture image data set to be processed and performing data preprocessing to obtain a wrist fracture image; Specifically, a wrist fracture image dataset to be processed is obtained; data cleaning and labeling processing is performed on the wrist fracture image dataset to be processed to obtain a labeled wrist fracture image dataset; data enhancement and segmentation processing is performed on the labeled wrist fracture image dataset to obtain a wrist fracture image.
[0018] In this embodiment, based on the public dataset GRAZPEDWRI-DX and the wrist fracture image dataset provided by a traditional Chinese medicine hospital, a wrist fracture image dataset to be processed is constructed, and then data cleaning is performed, and the fracture site is annotated using labeling tools such as LabelImg to generate a YOLO format dataset containing the fracture location. The data enhancement library (Albumentations) is used to perform random rotation, scaling, translation and other data enhancement operations on the image data to enrich the diversity of the data set and improve the adaptability of the model to different imaging conditions. Finally, the annotated and enhanced data set is divided into a training set, a validation set and a test set in a ratio of 75%:15%:15% to ensure the effectiveness of model training and the accuracy of testing.
[0019] S200, based on the fast network backbone structure, introduces a bidirectional feature pyramid network structure that integrates a global to local spatial aggregation module and a dynamic cross-stage deformable convolutional network module to build an improved YOLOv8 recognition network model; Specifically, Figure 3 As shown, the improved YOLOv8 recognition network model includes a backbone network module, a neck network module and a detection head network module, and the backbone network module, the neck network module and the detection head network module are connected in sequence, wherein: Further, if Figure 4As shown, the backbone network module adopts a fast network backbone structure, and the backbone network module specifically includes an embedding layer, a fast network module, a backbone fusion block, a global pooling layer, a backbone convolution block, a fully connected layer and a spatial pyramid fast pooling module, and the fast network module includes a depth-separable convolution layer and an attention mechanism; In this embodiment, in the deep learning target detection model, the backbone network is responsible for feature extraction of the input image, converting the original image pixel information into a multi-dimensional feature representation, which is then passed to the subsequent neck network and detection head for deeper target detection and positioning. The feature extraction capability of the backbone network determines the detection accuracy, speed and generalization performance of the model, and is a key part of the entire detection process. Therefore, the optimization of the backbone network can significantly improve the overall performance of the model. In YOLOv8, the backbone network is mainly composed of multiple convolutional layers and pooling layers, which are used to extract image features at different levels. In order to improve the performance of the model in wrist fracture imaging detection, the embodiment of the present invention first refers to the design idea of FasterNet and combines the architecture of YOLOv8 for improvement. FasterNet is a lightweight and efficient backbone network architecture proposed in CVPR2023. FasterNet is then used to replace the backbone network of YOLOv8. At the same time, in order to retain the multi-scale feature extraction capability of YOLOv8, the spatial pyramid fast pooling module (SPPF) in the original backbone of YOLOv8 is retained in the backbone network. By combining the lightweight design of FasterNet and the multi-scale receptive field enhancement of the SPPF module, it can reduce the computational overhead while still maintaining a high feature expression capability.
[0020] Further, if Figure 5 As shown, the neck network module adopts a bidirectional feature pyramid network structure. The neck network module specifically includes a global to local space aggregation module, an upsampling module, a neck fusion block and a dynamic cross-stage deformable convolutional network module, wherein, Fig.10 As shown, the global to local space aggregation module includes a first feature segmentation module, a first two-dimensional convolution module, a second two-dimensional convolution module, a global to local space aggregation convolution branch module, a context information aggregation module, a first feature splicing module and a first convolution block. Further, as Figure 8 As shown, the dynamic cross-stage deformable convolutional network module includes a second convolution block, a second feature segmentation module, a plurality of dynamic deformable convolution bottleneck modules and a second feature splicing module, wherein, Figure 7 As shown, the dynamic deformable convolution bottleneck module includes a convolution block, a dynamic deformable convolution module and a jump connection block, wherein, Figure 6As shown in Figure 1, the dynamic deformable convolution module includes a two-dimensional convolution module, a multi-scale position-sensitive cross attention module, a feature partitioning module, a feature splicing module, an activation function module, a dynamic deformable convolution module, a two-dimensional batch normalization module, and a SiLU activation function module. Furthermore, as Fig. 9 As shown, the global to local space aggregation convolution branch module includes several convolution blocks and a sigmoid activation function module.
[0021] In this embodiment, in the deep learning target detection model, the neck network (Neck) is responsible for fusing features from different levels, thereby improving the model's ability to locate and classify targets, especially multi-scale feature fusion is crucial for accurately detecting targets of different sizes. In the task of wrist fracture image detection, the features of the fracture area are often complex and small, so the optimization of the neck network can significantly improve the detection accuracy and generalization performance of the model. In order to improve the performance of the algorithm in wrist fracture detection, the embodiment of the present invention first replaces the PAFPN in YOLOv8 with BiFPN in the neck network to improve the flexibility and efficiency of feature fusion. BiFPN allows features to flow bidirectionally between different resolutions, so that the model can better fuse multi-scale features and enhance the detection ability of small targets (such as fracture details). In order to further optimize BiFPN, the embodiment of the present invention replaces some convolution modules in BiFPN with "global to local spatial aggregation (GLSA)" modules. GLSA can aggregate local spatial information more meticulously while maintaining the global receptive field, enhance the ability to capture fracture features under complex backgrounds, and thus improve the detection accuracy of the model. Finally, the improved MPCA is used to enhance the deformable convolutional network (DCNv2) and replace the cross-stage partial network (C2f) module in BiFPN. Based on the cross-attention mechanism (CA) proposed by CVPR2021, the MPCA attention mechanism is designed to enhance the offset and mask generation capabilities in DCNv2. By combining MPCA with DCNv2 and embedding it into the C2f module, a dynamic cross-stage deformable convolutional network module (C2f_DCNv2_Dynamic) is formed, which further replaces the C2f module in BiFPN to further improve the accuracy of feature extraction and fusion, especially in complex wrist fracture imaging scenarios.
[0022] Furthermore, it should be noted that in deep learning target detection models, the loss function is used to evaluate the difference between the predicted results and the true annotations. Its design directly affects the training effect of the model and the final detection performance. , which takes into account the overlapping area between bounding boxes and the distance between their center points, and has a certain effect on improving the positioning accuracy. However, in the task of wrist fracture image detection, the complex fracture morphology and the smaller target area put forward higher requirements for positioning accuracy. In order to further improve the positioning accuracy of the model, the embodiment of the present invention improves the positioning loss function of YOLOv8. First, the internal intersection-over-union ratio ( )replace , The loss function focuses on the overlapping area inside the bounding box and is more sensitive to small objects and subtle boundaries, making it perform better in the wrist fracture detection task. ),exist Based on this, we introduce the generalized intersection-union ratio ( ) is improved so that the loss function can more accurately measure the overlap and difference between the predicted box and the true box, especially when the position and shape of the bounding box are complex. Can provide higher positioning accuracy.
[0023] The loss function of the embodiment of the present invention The specific definition is as follows: 1) Bounding box adjustment: According to the scale parameter ( , in this embodiment, it is defined ), reduce the size of each bounding box and calculate its reduced coordinates; For the first bounding box box1, its calculation expression is: ; For the second bounding box box2, its calculation expression is: ; 2) Intersection area calculation, the expression is: ; The above formula calculates the two bounding boxes in and The product of the overlapping distances in the direction is used to obtain the area of the overlapping region. If the two boxes do not overlap, the overlapping area is 0.
[0024] 3) Calculation of the union area, the expression is: ; The area of the union is the sum of the areas of the two boxes minus the area of their intersection, is a small value used to prevent the denominator from being zero.
[0025] 4) Calculate, the expression is: ; 5) Import: It is standard The minimum enclosing rectangle is introduced to evaluate the positional relationship between two bounding boxes.
[0026] The calculation formula is: ; in, is the area of the minimum enclosing rectangle of the two bounding boxes.
[0027] The specific expression is as follows: ; exist middle, Used to measure the matching degree between the reduced core areas, the second part ,Right now Used to penalize the relative position of two boxes. If the two boxes are far apart, this value will increase, thereby reducing the overall value.
[0028] In summary, the improved YOLOv8 recognition network model constructed by the embodiment of the present invention can more effectively cope with the characteristics of medical images and the subtle features of wrist fractures. By combining the lightweight design concept of the fast network (FasterNet), the backbone network of YOLOv8 is optimized to make feature extraction more efficient, especially more adaptable in multi-scale feature fusion. At the same time, the bidirectional feature pyramid network (BiFPN) is used in the neck network to replace the original path aggregation feature pyramid network (PAFPN), and the global to local space aggregation (GLSA) module and the improved multi-scale position sensitive cross attention mechanism (MPCA) are used to further enhance the aggregation ability of features. These improvements improve the detection accuracy and computational efficiency of the model, and can better support the auxiliary diagnosis needs of wrist fractures. In addition, by replacing the original complete intersection-over-union (CIOU) loss function with the internal generalized intersection-over-union (Inner_GIoU), the positioning ability in complex morphology and small target detection can be improved.
[0029] S300, performing image recognition processing on the wrist fracture image based on the improved YOLOv8 recognition network model to obtain a wrist fracture recognition result.
[0030] S310, inputting the wrist fracture image into the improved YOLOv8 recognition network model; S320, based on the backbone network module of the improved YOLOv8 recognition network model, performing feature extraction processing on the wrist fracture image to obtain a wrist fracture feature image; Specifically, the wrist fracture image is input into the backbone network module of the improved YOLOv8 recognition network model; based on the embedding layer of the backbone network module, the data dimension reduction processing is performed on the wrist fracture image to obtain the reduced dimension wrist fracture image; based on the fast network module of the backbone network module, the local detail feature and global structure feature extraction processing is performed on the reduced dimension wrist fracture image to obtain a preliminary wrist fracture feature image; based on the backbone fusion block of the backbone network module, the preliminary wrist fracture feature image is subjected to multi-resolution feature fusion processing to obtain a fused wrist fracture feature image; based on the global pooling layer of the backbone network module, the fused wrist fracture feature image is subjected to global pooling compression processing to obtain a compressed wrist fracture feature image; based on the backbone convolution block of the backbone network module, the compressed wrist fracture feature image is subjected to convolution processing to obtain a convolved wrist fracture feature image; based on the fully connected layer of the backbone network module, the convolved wrist fracture feature image is mapped to obtain a mapped wrist fracture feature image; based on the spatial pyramid fast pooling module of the backbone network module, the mapped wrist fracture feature image is subjected to multi-scale aggregation processing to obtain a wrist fracture feature image.
[0031] In this embodiment, the lightweight convolution module designed by FasterNet is used to replace some convolution modules in YOLOv8 to reduce the number of parameters and computational overhead. The input image is efficiently encoded through the Embedding layer, and the low-level texture information is converted into a high-dimensional feature representation, which reduces the data dimension and reduces the subsequent calculation burden; FasterNet Block combines deep separable convolution and attention mechanism to accurately extract local details and global structural features of the image, significantly enhancing the model's perception of the fracture area; the Merging layer integrates multi-resolution features to improve the model's adaptability to targets of different scales; the global pooling layer further integrates global features, compresses redundant information, and improves feature expression efficiency. These optimizations not only improve the detection accuracy and feature extraction efficiency of the model, but also enhance its adaptability to complex fracture morphology, while effectively reducing the computational complexity of fracture imaging diagnosis. The SPPF module in YOLOv8 is retained to enhance the multi-scale feature extraction capability, so that the model has a higher feature expression capability while maintaining lightweight. Specifically, the SPPF module can perform multi-scale aggregation on features of different scales and generate multi-layer feature representations by applying the maximum pooling operation to the feature map, thereby enhancing the detection capability of small targets (such as fracture details), making the model lightweight and having higher feature expression capabilities.
[0032] S330, based on the neck network module of the improved YOLOv8 recognition network model, performing feature fusion processing on the wrist fracture feature image to obtain a fused wrist fracture feature image; Specifically, the wrist fracture feature image is input into the neck network module of the improved YOLOv8 recognition network model; based on the global to local space aggregation module of the neck network module, the wrist fracture feature image is subjected to global attention weighted calculation and feature aggregation processing to obtain the global feature image of the wrist fracture; based on the upsampling module of the neck network module, the wrist fracture global feature image is subjected to upsampling processing to obtain the upsampled wrist fracture global feature image; based on the neck fusion block of the neck network module, the upsampled wrist fracture global feature image is subjected to multi-scale feature fusion processing to obtain the preliminary fused wrist fracture global feature image; based on the dynamic cross-stage deformable convolutional network module of the neck network module, the cross-stage feature optimization processing is performed on the preliminary fused wrist fracture global feature image to obtain the fused wrist fracture feature image.
[0033] Among them, for the global to local space aggregation module, the wrist fracture feature image is input into the global to local space aggregation module of the neck network module; based on the first feature segmentation module of the global to local space aggregation module, the wrist fracture feature image is subjected to feature segmentation processing to obtain the segmented wrist fracture feature image; based on the first branch two-dimensional convolution module and the second branch two-dimensional convolution module of the global to local space aggregation module, convolution processing is performed on the segmented wrist fracture feature image respectively to obtain the first convolution wrist fracture feature image and the second convolution wrist fracture feature image; based on the global to local space aggregation module, the first branch two-dimensional convolution module and the second branch two-dimensional convolution module are subjected to convolution processing on the segmented wrist fracture feature image respectively to obtain the first convolution wrist fracture feature image and the second convolution wrist fracture feature image; A convolutional wrist fracture feature image is subjected to spatial feature aggregation processing to obtain a first aggregated wrist fracture feature image; a context information aggregation module based on the global to local spatial aggregation module is subjected to context information aggregation processing on the second convolutional wrist fracture feature image to obtain a second aggregated wrist fracture feature image; a second feature splicing module based on the global to local spatial aggregation module is used to splice the first aggregated wrist fracture feature image with the second aggregated wrist fracture feature image to obtain a spliced wrist fracture feature image; a convolution block based on the global to local spatial aggregation module is used to convolve the spliced wrist fracture feature image to obtain a global feature image of wrist fracture.
[0034] Furthermore, for the dynamic cross-stage deformable convolutional network module, the global feature image of wrist fracture after preliminary fusion is input into the dynamic cross-stage deformable convolutional network module of the neck network module; based on the second convolution block of the dynamic cross-stage deformable convolutional network module, the global feature image of wrist fracture after preliminary fusion is convolved to obtain the convolved global feature image of wrist fracture; based on the second feature segmentation module of the dynamic cross-stage deformable convolutional network module, the global feature image of wrist fracture after convolution is subjected to feature segmentation processing to obtain the segmented global feature image of wrist fracture; based on the dynamic deformable convolution bottleneck module of the dynamic cross-stage deformable convolutional network module, the segmented global feature image of wrist fracture is subjected to dynamic feature optimization processing to obtain the optimized feature image of wrist fracture; based on the second feature splicing module of the dynamic cross-stage deformable convolutional network module, the optimized feature image of wrist fracture is subjected to feature splicing and fusion processing to obtain the fused feature image of wrist fracture.
[0035] In this embodiment, BiFPN is used to replace the original PAFPN. BiFPN allows features to flow bidirectionally between different resolutions. By cascading feature maps of different resolutions, the fusion capability of multi-scale features is enhanced. This improvement enables the model to effectively capture the detailed features of small targets in wrist fracture images, while reducing unnecessary computational overhead. The convolution module in BiFPN is replaced with the GLSA module to increase the fineness of feature aggregation to enhance the feature capture capability under complex backgrounds. Specifically, GLSA performs global attention weighted processing on feature maps from different resolutions, and then performs fine feature aggregation on local areas, thereby more accurately extracting features of the fracture area under complex backgrounds. The MPCA attention mechanism is used to enhance the offset and mask generation of DCNv2, improve the dynamics and pertinence of feature extraction, and embed it into the C2f module to form the C2f_DCNv2_Dynamic module, which then replaces the C2f module in BiFPN. The cross-stage part of the feature map can be optimized, and the key features of the fracture area can be effectively extracted, while suppressing the interference of complex backgrounds.
[0036] S340, based on the detection head network module of the improved YOLOv8 recognition network model, performing target recognition detection on the fused wrist fracture feature image to obtain a wrist image fracture recognition result.
[0037] Furthermore, the improved YOLOv8 network was trained with 300 training rounds, 40 early stops, 0.01 initial learning rate, 0.01 minimum learning rate, 64 batch size, SGD optimizer, and weight decay of 0.0005. The training process was iterated to the best effect, the optimal weight was saved, and the weight was used to test the test set to verify the model performance.
[0038] Finally, the model was evaluated, and the recall rate (R) and average precision (mAP50, mAP50-95) were used as evaluation indicators to evaluate the performance of the improved model. The experimental results show that the improved fracture detection algorithm has improved accuracy, computational efficiency and real-time performance. Compared with the original YOLOv8 algorithm, the recall rate and average precision of the model have been optimized, and the model has higher accuracy in wrist fracture positioning and detection. For wrist limb general category detection, the recall rate increased from 0.613 to 0.761, an increase of about 24%; MAP50 increased from 0.624 to 0.784, an increase of about 26%; MAP50-100 increased from 0.344 to 0.501, an increase of about 46%. For fracture category detection, the recall rate increased from 0.741 to 0.869, an increase of about 17%; MAP50 increased from 0.826 to 0.928, an increase of about 12%; MAP50-100 increased from 0.425 to 0.54, an increase of about 27%, as shown in Tables 1 and 2. The simulation experimental results are shown in Tables 1 and 2. Fig.11 as well as Fig.12 shown.
[0039] In order to verify the improvement of the optimized model in feature extraction, this embodiment compares the differences in feature extraction capabilities of the models through heat maps. Fig.11 As shown in the figure, by comparing the feature activation maps of the models on the same dataset, it can be seen that the optimized model shows stronger feature responses in the fracture and joint areas in the image, and can focus on key areas more accurately, thereby improving the accuracy and robustness of target detection.
[0040] Some test results of the validation set are as follows Fig.12 As shown in the figure, by comparison, the YOLOv8n model has false detection or missed detection in some samples. However, through improvement, the model effectively improves this limitation, can more accurately mark the fracture area, and improve the detection confidence.
[0041] Table 1. Current model training results data table
[0042] Table 2: Data table of model training results constructed by the present invention
[0043] In summary, the embodiment of the present invention optimizes the backbone network based on the YOLOv8 algorithm and combines the design concept of FasterNet, while retaining the original SPPF module of YOLOv8, so that the model performs well in multi-scale feature extraction. Compared with the original backbone network structure, this algorithm significantly improves the extraction accuracy and computational efficiency of fracture features while reducing the number of parameters.
[0044] In the neck network, BiFPN is used to replace PAFPN to enhance the flexibility of feature fusion, and the GLSA module is used to replace the convolution module in BiFPN to improve the ability to capture details in the fracture area. In addition, an improved MPCA attention mechanism is designed and embedded in the offset and mask generation modules in DCNv2, combined with the C2f module to form the C2f_DCNv2_Dynamic module to further enhance the feature fusion accuracy of BiFPN. This optimization scheme enables the model to focus more accurately on fracture features at different scales and complex backgrounds.
[0045] In terms of positioning loss, CIOU is replaced by Inner_GIoU, which enables the model to have higher positioning accuracy in small target and complex morphology boundary detection. Inner_GIoU can more effectively optimize the boundary matching of the fracture area during model training by reducing the bounding box size and combining the penalty mechanism of GIoU, thereby significantly improving positioning accuracy.
[0046] Reference Figure 2 , a wrist image fracture recognition and detection system based on an improved YOLOv8 network, including: The first module 201 is used to obtain a wrist fracture image data set to be processed and perform data preprocessing to obtain a wrist fracture image; The second module 202 is used to introduce a bidirectional feature pyramid network structure integrating a global to local spatial aggregation module and a dynamic cross-stage deformable convolutional network module based on a fast network backbone structure to build an improved YOLOv8 recognition network model; The third module 203 is used to perform image recognition processing on the wrist fracture image based on the improved YOLOv8 recognition network model to obtain a wrist image fracture recognition result.
[0047] The contents of the above method embodiments are all applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0048] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A wrist image fracture recognition and detection method based on an improved YOLOv8 network is characterized in that: The following steps are involved: Acquire a wrist fracture image data set to be processed and perform data preprocessing to obtain a wrist fracture image; Based on the fast network backbone structure, a bidirectional feature pyramid network structure integrating the global to local spatial aggregation module and the dynamic cross-stage deformable convolutional network module is introduced to build an improved YOLOv8 recognition network model. Based on the improved YOLOv8 recognition network model, image recognition processing is performed on wrist fracture images to obtain wrist image fracture recognition results.
2. According to the wrist image fracture recognition and detection method based on the improved YOLOv8 network according to claim 1, it is characterized in that: The step of obtaining a wrist fracture image dataset to be processed and performing data preprocessing to obtain a wrist fracture image specifically includes: Obtain a wrist fracture image dataset to be processed; Performing data cleaning and labeling on the wrist fracture image dataset to be processed to obtain a labeled wrist fracture image dataset; Data enhancement and segmentation are performed on the labeled wrist fracture image dataset to obtain wrist fracture images.
3. According to claim 2, the wrist image fracture recognition and detection method based on the improved YOLOv8 network is characterized in that: The improved YOLOv8 recognition network model includes a backbone network module, a neck network module and a detection head network module, and the backbone network module, the neck network module and the detection head network module are connected in sequence, wherein: The backbone network module adopts a fast network backbone structure. The backbone network module specifically includes an embedding layer, a fast network module, a backbone fusion block, a global pooling layer, a backbone convolution block, a fully connected layer and a spatial pyramid fast pooling module. The fast network module includes a depth-separable convolution layer and an attention mechanism. The neck network module adopts a bidirectional feature pyramid network structure. The neck network module specifically includes a global to local space aggregation module, an upsampling module, a neck fusion block and a dynamic cross-stage deformable convolutional network module. The global to local space aggregation module includes a first feature segmentation module, a first two-dimensional convolution module, a second two-dimensional convolution module, a global to local space aggregation convolution branch module, a context information aggregation module, a first feature splicing module and a first convolution block. The dynamic cross-stage deformable convolutional network module includes a second convolution block, a second feature segmentation module, several dynamic deformable convolution bottleneck modules and a second feature splicing module.
4. According to claim 3, the wrist image fracture recognition and detection method based on the improved YOLOv8 network is characterized in that: It also includes introducing a generalized intersection-and-union loss function based on the internal intersection-and-union loss function, and constructing an internal generalized intersection-and-union loss function as a loss function of the improved YOLOv8 recognition network model. The expression of the internal generalized intersection-and-union loss function is specifically as follows: ; In the above formula, represents the internal generalized intersection-over-union loss function, represents the internal intersection-over-union loss function, represents the generalized intersection and union ratio, represents the area of the minimum enclosing rectangle of two bounding boxes, Represents the area of the union.
5. The wrist image fracture recognition and detection method based on the improved YOLOv8 network according to claim 4 is characterized in that: The step of performing image recognition processing on the wrist fracture image based on the improved YOLOv8 recognition network model to obtain the wrist image fracture recognition result specifically includes: Input the wrist fracture image into the improved YOLOv8 recognition network model; Based on the backbone network module of the improved YOLOv8 recognition network model, feature extraction is performed on the wrist fracture image to obtain the wrist fracture feature image; Based on the neck network module of the improved YOLOv8 recognition network model, feature fusion processing is performed on the wrist fracture feature image to obtain the fused wrist fracture feature image; Based on the detection head network module of the improved YOLOv8 recognition network model, target recognition detection is performed on the fused wrist fracture feature image to obtain the wrist image fracture recognition result.
6. The wrist image fracture recognition and detection method based on the improved YOLOv8 network according to claim 5 is characterized in that: The backbone network module based on the improved YOLOv8 recognition network model performs feature extraction processing on the wrist fracture image to obtain the wrist fracture feature image, which specifically includes: Input the wrist fracture image into the backbone network module of the improved YOLOv8 recognition network model; Based on the embedding layer of the backbone network module, data dimension reduction processing is performed on the wrist fracture image to obtain a wrist fracture image after dimension reduction; The fast network module based on the backbone network module extracts local detail features and global structural features from the wrist fracture image after dimensionality reduction to obtain a preliminary wrist fracture feature image. Based on the backbone fusion block of the backbone network module, multi-resolution feature fusion processing is performed on the preliminary wrist fracture feature image to obtain a fused wrist fracture feature image; Based on the global pooling layer of the backbone network module, the fused wrist fracture feature image is globally pooled and compressed to obtain a compressed wrist fracture feature image; Based on the backbone convolution block of the backbone network module, convolution processing is performed on the compressed wrist fracture feature image to obtain a convolved wrist fracture feature image; Based on the fully connected layer of the backbone network module, the convolutional wrist fracture feature image is mapped to obtain a mapped wrist fracture feature image; Based on the spatial pyramid fast pooling module of the backbone network module, multi-scale aggregation processing is performed on the mapped wrist fracture feature image to obtain the wrist fracture feature image.
7. The wrist image fracture recognition and detection method based on the improved YOLOv8 network according to claim 6 is characterized in that: The step of performing feature fusion processing on the wrist fracture feature image based on the neck network module of the improved YOLOv8 recognition network model to obtain the fused wrist fracture feature image specifically includes: Input the wrist fracture feature image into the neck network module of the improved YOLOv8 recognition network model; Based on the global to local space aggregation module of the neck network module, the wrist fracture feature image is subjected to global attention weighted calculation and feature aggregation processing to obtain the wrist fracture global feature image; An upsampling module based on the neck network module performs upsampling processing on the global feature image of wrist fracture to obtain an upsampled global feature image of wrist fracture; Based on the neck fusion block of the neck network module, multi-scale feature fusion processing is performed on the upsampled global feature image of wrist fracture to obtain a preliminary fused global feature image of wrist fracture; A dynamic cross-stage deformable convolutional network module based on the neck network module performs cross-stage feature optimization processing on the global feature image of wrist fracture after preliminary fusion to obtain the feature image of wrist fracture after fusion.
8. The wrist image fracture recognition and detection method based on the improved YOLOv8 network according to claim 7 is characterized in that: The global to local space aggregation module based on the neck network module performs global attention weighted calculation and feature aggregation processing on the wrist fracture feature image to obtain the global feature image of the wrist fracture. This step specifically includes: The wrist fracture feature image is input into the global to local space aggregation module of the neck network module; Based on the first feature segmentation module of the global to local space aggregation module, feature segmentation processing is performed on the wrist fracture feature image to obtain a segmented wrist fracture feature image; Based on the first branch two-dimensional convolution module and the second branch two-dimensional convolution module of the global to local space aggregation module, convolution processing is performed on the segmented wrist fracture feature image to obtain a first convolution wrist fracture feature image and a second convolution wrist fracture feature image; A global-to-local spatial aggregation convolution branch module based on the global-to-local spatial aggregation module performs spatial feature aggregation processing on the first convolution wrist fracture feature image to obtain a first aggregated wrist fracture feature image; Based on the context information aggregation module of the global to local space aggregation module, the context information aggregation processing is performed on the second convolution wrist fracture feature image to obtain a second aggregated wrist fracture feature image; Based on the second feature splicing module of the global to local space aggregation module, the first aggregated wrist fracture feature image and the second aggregated wrist fracture feature image are spliced to obtain a spliced wrist fracture feature image; Based on the first convolution block of the global to local spatial aggregation module, the spliced wrist fracture feature image is convolved to obtain the global feature image of wrist fracture.
9. The wrist image fracture recognition and detection method based on the improved YOLOv8 network according to claim 8 is characterized in that: The step of performing cross-stage feature optimization processing on the wrist fracture global feature image after preliminary fusion by the dynamic cross-stage deformable convolutional network module based on the neck network module to obtain the fused wrist fracture feature image specifically includes: The global feature image of wrist fracture after preliminary fusion is input into the dynamic cross-stage deformable convolutional network module of the neck network module; Based on the second convolution block of the dynamic cross-stage deformable convolutional network module, convolution processing is performed on the preliminary fused global feature image of wrist fracture to obtain a convolved global feature image of wrist fracture; The second feature segmentation module based on the dynamic cross-stage deformable convolutional network module performs feature segmentation processing on the convolved wrist fracture global feature image to obtain a segmented wrist fracture global feature image; A dynamic deformable convolution bottleneck module based on a dynamic cross-stage deformable convolution network module performs dynamic feature optimization processing on the segmented wrist fracture global feature image to obtain an optimized wrist fracture feature image; Based on the second feature splicing module of the dynamic cross-stage deformable convolutional network module, feature splicing and fusion processing is performed on the optimized wrist fracture feature image to obtain a fused wrist fracture feature image.
10. Wrist image fracture recognition and detection system based on improved YOLOv8 network, characterized in that: Includes the following modules: The first module is used to obtain a wrist fracture image data set to be processed and perform data preprocessing to obtain a wrist fracture image; The second module is used to build an improved YOLOv8 recognition network model based on a fast network backbone structure, introducing a bidirectional feature pyramid network structure that integrates a global to local spatial aggregation module and a dynamic cross-stage deformable convolutional network module; The third module is used to perform image recognition processing on wrist fracture images based on the improved YOLOv8 recognition network model to obtain wrist image fracture recognition results.
Citation Information
Patent Citations
Lightweight underwater target real-time detection method
CN116977844A
Automobile target detection and fine grit recognition algorithm based on remote sensing image
CN118736412A
Hidden forbidden article detection method based on lightweight millimeter wave radar
CN118823311A
Center alignment positioning method and system based on YOLOv8-RBT
CN119131331A
Underwater target detection method based on improved YOLOv8
CN119229272A
Cited By
Conveying belt defect detection method based on GLSA-YOLOv11
CN120598872A
Litchi tree canopy segmentation method fusing FasterNet and BiFPN
CN121685957A