Child wrist fracture image detection method based on improved YOLOv8
By introducing the adaptive hybrid feature enhancement module AHF, receptive field attention convolution operation RFAConv and loss function weighted fusion in the YOLOv8 model, the small object detection difficulties and complex background interference problems of YOLOv8 in the detection of wrist fracture images of children are solved, and the detection accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202510234503.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-11
AI Technical Summary
The existing YOLOv8 model has problems such as small object detection difficulties, complex background interference, data imbalance, low contrast and noise interference in the detection of children's wrist fracture images, resulting in low detection accuracy.
The adaptive hybrid feature enhancement module AHF is introduced in YOLOv8's Neck network to enhance feature extraction capabilities; the receptive field attention convolution operation RFAConv is introduced in the Backbone network to improve multi-scale feature extraction; the weighted fusion is combined with Focal Loss and the binary classification cross-entropy loss function to increase the model's attention to difficult-to-classify samples.
It significantly improves the accuracy and efficiency of children's wrist fracture image detection, reduces the amount of calculation of the model and the probability of missed detection and missed detection.
Smart Images

Figure CN120298300A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting images of children's wrist fractures based on improved YOLOv8, belonging to the field of computer vision. Background Art
[0002] The technology for detecting images of children's wrist fractures mainly focuses on the automated analysis of medical images of children's wrist fractures, especially X-ray images. With the rapid development of medical imaging and artificial intelligence technologies, methods based on computer vision and deep learning have been widely applied in medical image analysis. Especially in the field of detecting images of children's wrist fractures, traditional manual analysis methods can no longer meet the requirements of high efficiency and accuracy. Therefore, automated image detection and analysis technologies are particularly important.
[0003] In images of children's wrist fractures, fractures in the wrist area usually present in relatively small regions and have a low contrast with the background structure, making them easily overlooked. Such images often have complex background information and low image quality (such as insufficient brightness, blurring, etc.), which greatly reduces the detection effect of traditional image processing methods. Traditional methods for detecting wrist fracture images mainly rely on manual feature extraction and rule-based image analysis technologies, which usually analyze using shallow features such as edges, textures, and shapes of the images. However, these traditional methods often cannot effectively identify fracture sites with more image noise or lower contrast, resulting in low detection accuracy.
[0004] In recent years, the rapid development of deep learning technologies has brought breakthrough progress to the field of image detection. Especially, the application of convolutional neural networks (CNNs) in object detection has achieved remarkable results. The YOLO series of algorithms, with their high detection speed and accuracy, have become important methods in the field of object detection, especially performing excellently in real-time object detection. As the most popular version in the YOLO series, YOLOv8 has stronger feature extraction capabilities and higher efficiency, and has been widely applied to various object detection tasks.
[0005] However, YOLOv8 still faces some challenges when dealing with images of children's wrist fractures, especially in the detection of complex backgrounds and small targets. Wrist fractures usually occur in relatively small regions and are adjacent to other bone structures, with complex and difficult-to-distinguish morphologies. Therefore, the traditional YOLOv8 model may not be able to fully capture these fine-grained features, resulting in low detection accuracy of the model for small targets. In addition, images of children's wrist fractures usually have low contrast and noise interference, which further increases the difficulty of the detection task. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for detecting children's wrist fracture images based on improved YOLOv8, aiming to solve the technical problems of difficult small target detection, complex background interference, data imbalance, and low contrast and noise interference in the existing technology for children's wrist fracture image detection tasks.
[0007] To achieve the above object, the technical solution adopted by the present invention is a method for detecting children's wrist fracture images based on improved YOLOv8, aiming to improve the recognition accuracy and detection efficiency of the fracture area, which specifically includes the following steps:
[0008] Step1: Select a children's wrist fracture image dataset, divide the dataset into a training set and a validation set, and perform data augmentation;
[0009] Step2: Based on the YOLOv8 network as the basic framework, introduce an Adaptive Hybrid Feature Enhancement module AHF in the Neck network to improve the network's feature extraction ability;
[0010] Step3: Introduce a Receptive Field Attention Convolution operation RFAConv in the Backbone network of YOLOv8 to enhance the multi-scale feature extraction ability;
[0011] Step4: Introduce the Focal Loss function, and combine it with the original binary cross-entropy loss function CE Loss in the YOLOv8 model, and perform weighted fusion on the two loss functions;
[0012] Step5: First, through the improvement of the original YOLOv8 model in Step2-Step4, obtain the Adaptive Multidimensional Enhancement network structure ARF_YOLO model, then use the training set to train the ARF_YOLO model to obtain the final training weights, and finally use the validation set to verify the model, and finally achieve the detection of children's wrist fracture images.
[0013] Specifically, Step2 is as follows:
[0014] YOLOv8 consists of three parts: Backbone, Neck, and Head. The Backbone part of the backbone network is composed of several CBS, C2f convolution modules, and a Spatial Pyramid Pooling Module (SPPF). The Neck part of the neck network is composed of several CBS, C2f, an Upsample module, and a Concat module. The Head part of the head network is composed of three Detect modules. Among them, the Backbone is the basic feature extraction network of YOLOv8, responsible for extracting multi-level features of the image. The CBS module is used for preliminary feature extraction and downsamples the feature map through convolution to reduce its size. The C2f module is used to improve the feature expression ability. The SPPF module extracts and fuses features at multiple scales to enhance context information. The Neck is used to fuse features at different levels and strengthen the multi-scale detection ability. The Upsample module upsamples the feature map through nearest neighbor interpolation to enlarge it. The Concat module is used for the splicing operation of feature maps. The Head is the final prediction module, directly outputting the category and bounding box of the target. The Detect module is used to finally detect the target from the feature map.
[0015] The proposed AHF module in the present invention is an adaptive hybrid feature enhancement module, which combines a channel attention module (SE, Squeeze-and-Excitation) and a spatial attention module (SA, Spatial Attention). It improves the feature extraction ability of the network through dynamic initialization (Dynamic Init) and feature enhancement mechanism. Among them, dynamic initialization is used to dynamically adjust the parameter settings of the SE and SA modules according to the dimension, size and channels of the input feature map, ensuring that the module adapts to input data of different sizes and characteristics; the SE module is mainly used for channel attention mechanism, which models the global information, assigns different weights to different channels, uses global average pooling (Global Average Pooling, GAP) to compress the spatial features of each channel into a single value, generates a global feature description of the channel, then uses a fully connected layer (FC) to model the weights of each channel, and is generated by the ReLU and Sigmoid activation functions. The Scale operation is to multiply the generated channel weights by the original input feature map to weight each channel; the SA module is mainly used for spatial attention mechanism, by capturing the spatial distribution information of the feature map, focusing on important regions, performing convolution operations on the input feature map, extracting the features of each spatial position, first using a convolutional layer to extract spatial features, BN (Batch Normalization) for normalization, ReLU to activate non-linear features, and then using a fully connected layer and Sigmoid activation function to generate the weights of each spatial position. Finally, the outputs of the SE and SA modules are fused by addition to generate enhanced features. This fusion method can capture rich channel information and spatial information and improve the overall feature expression ability.
[0016] Further, the input feature map X initializes the module parameters through the Dynamic Init, X ∈ R C×H×W , where C is the number of channels, and H and W are the height and width of the feature map respectively;
[0017] The SE module performs channel enhancement on the input feature map X, and uses global average pooling GAP to compress the input feature map X into channels:
[0018]
[0019] Generate channel weights W through a two-layer fully connected network and non-linear activation c :
[0020] W c = σ(W2 · ReLU(W1 · Z c ))
[0021] where, Z c , W c ∈ RC , Z c denotes the channel global feature vector obtained after global average pooling. X(c, i, j) represents the input feature map, c represents the channel index, (i, j) represents the pixel position, W1 and W2 are the weight matrices of the fully connected layer, σ is the Sigmoid activation function, and ReLU represents the activation function;
[0022] Using the channel weight W c to perform channel weighting on the input feature map:
[0023] X SE = W c ·X
[0024] The SA module performs spatial enhancement on the input feature map X and uses convolution operations to extract spatial distribution features:
[0025] F conv = ReLU(BN(Conv2d(X)))
[0026] Using F conv and the Sigmoid activation function to generate the spatial weight W SA :
[0027] W SA = σ(F conv )
[0028] Using the spatial weight W SA to perform spatial weighting on the input feature map:
[0029] X SA = W SA ·X
[0030] where F conv , W SA ∈ R C×H×W , X SE represents the output feature map of the SE module, X SA represents the output feature map of the SA module, F com represents the spatial features extracted from the convolution operation, BN represents batch normalization of the convolution result, and Conv2d(X) represents two-dimensional convolution of the input feature map X to extract spatial information;
[0031] Fuse the outputs of the SE module and the SA module by weighting, and the final output is:
[0032] X' = X + X SE + X SA
[0033] where X' is the final output.
[0034] The present invention adds the AHF module to the Neck network of YOLOv8, making feature fusion more flexible and adaptive, enabling efficient fusion between feature maps at different levels and scales, thereby enhancing the detection performance.
[0035] Specifically, Step 3 is as follows:
[0036] To solve the problem of parameter sharing in the convolutional kernel and better evaluate the importance of each feature in the receptive field, the receptive field attention convolution operation (RFAConv) is introduced. RFAConv is a lightweight and plug-and-play module, and the RFAConv operation can replace the standard convolution with very few additional parameters and computational costs. This significantly improves the network performance and effectively solves the parameter sharing problem existing in traditional convolution methods.
[0037] The structure of the RFAConv includes a 3×3 convolutional kernel, and uses fast grouped convolution to extract receptive field features, average pooling to aggregate global information, and a 1×1 grouped convolution operation with softmax to calculate the importance of each feature in the receptive field. Specifically:
[0038] The input feature map undergoes feature extraction through fast grouped convolution, which reduces the computational amount and improves the feature expression ability through channel grouping;
[0039] Use the AvgPool operation to aggregate global information and capture the context of the image;
[0040] Adjust the feature weights through 1×1 grouped convolution and normalize the importance of each feature using softmax;
[0041] Combine the adjusted features with the original features through a weighting mechanism to form a weighted feature map, which is passed to the next layer of the network to enhance the representation ability of the receptive field;
[0042] The calculation formula of the RFAConv is expressed as follows:
[0043] F = Softmax(g 1×1 (AvgPool(X))) × ReLU(Norm(g k×k (X))) = A rf ×F rf
[0044] where F represents the field space feature obtained by multiplying the attention map A rf with the transformed receptive field map F rf multiplied, g i ×iDenote grouped convolution of size i×i, k represents the size of the convolution kernel, Norm represents normalization, and X represents the input feature map; after shape adjustment, the height and width of the feature map increase by k times, and k×k convolution with a stride of k is required to extract information; this method significantly enhances the performance of the standard convolution kernel, especially in fine-grained feature extraction for complex tasks.
[0045] The present invention introduces the RFAConv module to replace the standard convolution module, reduces the computational amount by introducing grouped convolution, and uses 1×1 grouped convolution and softmax activation to dynamically adjust the feature weights at each position in the receptive field, enhancing the feature expression ability. It aggregates global information through average pooling, improves the understanding of image content, and enhances the representational ability of the receptive field, helping the model focus on important feature regions, especially performing outstandingly in complex tasks. By reducing parameters and computational costs while enhancing the feature extraction ability, the network performance is significantly improved.
[0046] The specific weighted fusion of the two loss functions is as follows:
[0047] The FocalLoss introduced in the present invention is a loss function used to handle the class imbalance problem, usually used in object detection tasks, especially in the case of severe imbalance between foreground and background samples. Focal Loss is proposed based on the traditional cross-entropy loss, by reducing the weights of easy-to-classify samples and focusing on hard-to-classify samples, thereby improving the model's learning ability for hard-to-classify samples.
[0048] In the detection task, the loss function of the YOLOv8 model includes regression loss and classification loss. The regression loss is mainly used to measure the difference between the predicted bounding box position and the true bounding box position of the model. The YOLOv8 model uses the CIoU loss function and the DEL loss function to calculate the overlap degree between the predicted box and the true box, and adjusts the position of the bounding box through the regression loss; the classification loss is used to measure the prediction accuracy of the target class within each bounding box. For each target, YOLOv8 calculates the probability that it belongs to each class and compares it with the actual class label. The commonly used classification loss function is binary cross-entropy loss (CELoss). The present invention introduces the FocalLoss loss function, combines it with the original CELoss loss function in the YOLOv8 model, and performs weighted fusion on the two loss functions, enabling the model to pay more attention to hard-to-classify samples while avoiding over-ignoring easy-to-classify samples, thus improving the overall detection performance.
[0049] The standard binary cross-entropy loss function is used to evaluate the difference between the predicted value and the true label, and is suitable for handling relatively balanced classes. Its calculation formula is:
[0050] CE(p, y) = -y·log(p) - (1 - y)·log(1 - p)
[0051] Among them, p is the predicted probability (obtained through the Sigmoid activation function output by the model); y is the true label (1 represents the target, 0 represents the background). This loss function is easily affected by background samples because the proportion of background samples in the dataset is much higher than that of foreground targets, resulting in the network being prone to predicting the background class and thus ignoring the target.
[0052] Focal Loss introduces a modulating factor (1 - p t ) Υ , and this factor can reduce the loss of samples that are easy to classify, thus making the network pay more attention to samples that are difficult to classify, especially the target objects. Its formula is:
[0053] FL(p t ) = -α·(1 - p t ) Υ ·log(p t )
[0054] Among them, p t is the predicted probability of the target class (when the target is the positive class, p t = p, when the target is the negative class, p t = 1 - p); α is the balancing factor, which adjusts the weights of positive and negative samples to prevent the excessive influence of negative samples; Υ is the modulating factor, which is used to control the weights of difficult samples. Focal Loss can reduce the loss impact of background samples and increase the attention to target objects, especially for small objects and target objects that are difficult to classify.
[0055] In order to combine Focal Loss and the traditional cross-entropy loss function, the advantages of both can be utilized simultaneously through weighted fusion. The weighted fusion loss function is expressed as:
[0056] Loss = λ·FL(p t ) + (1 - λ)·CE(p, y)
[0057] Among them, Loss is the total loss after weighted fusion of the two loss functions, FL(p t ) is the loss calculated through Focal Loss; CE(p, y) is the loss calculated through the traditional binary cross-entropy; λ is a hyperparameter used to balance the weights of the two losses; through the method of weighted fusion, the model can not only pay attention to target objects that are difficult to classify, but also utilize the stability and simplicity of the cross-entropy loss, avoiding excessive attention to the background area during training.
[0058] The beneficial effects of the present invention are as follows: By proposing a brand-new Adaptive Hybrid Feature Enhancement Module (AHF) and integrating it into the Neck network of YOLOv8, the present invention significantly improves the efficiency and performance of feature extraction; introducing the Receptive Field Attention Convolution operation (RFAConv) into the Backbone network of YOLOv8 and replacing the standard convolution module effectively enhances the multi-scale feature extraction ability of the model; introducing Focal Loss and combining it with the traditional binary cross-entropy loss, and by weighted fusion of the loss functions of both, the model pays more attention to difficult-to-classify samples and avoids over-neglecting easy-to-classify samples. This weighted fusion method effectively improves the detection ability of the model for fracture regions and enhances the detection performance. In summary, the model proposed by the present invention can effectively improve the detection performance of children's wrist fractures, increase the average precision of detection while reducing the computational amount of the model and the probability of target missed detection and false detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is the flowchart of the steps of the present invention;
[0060] Figure 2 is the structure diagram of the ARF_YOLO model proposed by the present invention;
[0061] Figure 3 is the structure diagram of the AHF module proposed by the present invention;
[0062] Figure 4 is the detection result diagram of children's wrist fractures of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0064] Example 1: A method for detecting children's wrist fracture images based on improved YOLOv8, the flowchart of the steps can be seen in Figure 1 , and the specific model structure diagram can be seen in Figure 2 . The technical solutions adopted mainly include the following steps:
[0065] Step1: Select a children's wrist fracture image dataset, divide the dataset into a training set and a validation set, and perform data augmentation. Specifically:
[0066] Prepare the experimental configurations required by the present invention: The operating system uses Ubuntu22.04.4, the NVDIARTX4090D graphics card, the selected deep learning computing framework is PyTorch, uses CUDA11.8 of the deep learning platform to manage the tool libraries required for running the code, and uses PyCharm software to run the model code and configure the virtual environment and tool libraries required by YOLOv8.
[0067] To train the improved model of the present invention and verify the effectiveness of the model subsequently, the present invention selects the GRAZPEDWRI-DX dataset to train and verify the model. The GRAZPEDWRI-DX dataset is a publicly available pediatric wrist trauma X-ray image dataset provided by the Medical University of Graz, containing 20,327 X-ray images. In addition, the dataset covers 6,091 pediatric patients, including 10,643 studies, and provides 74,459 image labels and 67,771 annotated objects, divided into 9 object categories. In addition, the dataset is divided into a training set, a validation set, and a test set according to a ratio of 7:2:1. Due to reasons such as low brightness and blurriness of the data images, the dataset is augmented, specifically by enhancing the contrast and brightness of the dataset images, and the augmented data is used for training, validation, and testing.
[0068] Step 2: Based on the YOLOv8 network as the basic framework, an Adaptive Hybrid Feature Enhancement module AHF is introduced in the Neck network to improve the network's feature extraction ability.
[0069] Specifically, the AHF combines the Channel Attention module SE and the Spatial Attention module SA, and improves the network's feature extraction ability through Dynamic Init and the feature enhancement mechanism. Specifically:
[0070] The input feature map X initializes the module parameters through the Dynamic Init, X ∈ R C×H×W , where C is the number of channels, and H and W are the height and width of the feature map respectively;
[0071] The SE module performs channel enhancement on the input feature map X, and uses global average pooling GAP to compress the input feature map X into channels:
[0072]
[0073] The channel weight W is generated through two fully connected networks and non-linear activation: c :
[0074] W c = σ(W2 · ReLU(W1 · Z c ))
[0075] where, Z c , W c ∈ R C ,Z c represents the channel global feature vector obtained after global average pooling, X(c, i, j) represents the input feature map, c represents the channel index, (i, j) represents the pixel position, W1 and W2 are the weight matrices of the fully connected layers, σ is the Sigmoid activation function, and ReLU represents the activation function;
[0076] Using channel weight W c Perform channel weighting on the input feature map:
[0077] X SE = W c · X
[0078] The SA module performs spatial enhancement on the input feature map X and uses convolutional operations to extract spatial distribution features:
[0079] F conv = ReLU(BN(Conv2d(X)))
[0080] Using F conv and the Sigmoid activation function to generate the spatial weight W SA :
[0081] W SA = σ(F conv )
[0082] Using the spatial weight W SA Perform spatial weighting on the input feature map:
[0083] X SA = W SA · X
[0084] where F conv , W SA ∈ R C×H×W , X SE represents the output feature map of the SE module, X SA represents the output feature map of the SA module, F com represents the spatial features extracted from the convolutional operation, BN represents batch normalization of the convolutional result, Conv2d(X) represents two-dimensional convolution of the input feature map X to extract spatial information;
[0085] Perform weighted fusion of the outputs of the SE module and the SA module, and the final output is:
[0086] X' = X + X SE + X SA
[0087] where X' is the final output.
[0088] Step3: Introduce the receptive field attention convolution operation RFAConv into the backbone network Backbone of YOLOv8 to enhance the multi-scale feature extraction ability.
[0089] Specifically, the structure of the RFAConv includes a 3×3 convolutional kernel, and uses fast grouped convolution to extract receptive field features, average pooling to aggregate global information, and 1×1 grouped convolution operation with softmax to calculate the importance of each feature in the receptive field, specifically as follows:
[0090] The input feature map undergoes feature extraction through fast grouped convolution;
[0091] The AvgPool operation is used to aggregate global information and capture the context of the image;
[0092] The feature weights are adjusted through 1×1 grouped convolution, and the importance of each feature is normalized using softmax;
[0093] The adjusted features are combined with the original features through a weighting mechanism to form a weighted feature map, which is passed to the next layer of the network;
[0094] The calculation formula of the RFAConv is expressed as follows:
[0095] F = Softmax(g 1×1 (AvgPool(X))) × ReLU(Norm(g k×k (X))) = A rf ×F rf
[0096] Among them, F represents the field space feature obtained by multiplying the attention map A rf with the transformed receptive field map F rf g i ×i represents grouped convolution of size i×i, k represents the size of the convolutional kernel, Norm represents normalization, and X represents the input feature map; after shape adjustment, the height and width of the feature map increase by k times, and k×k convolution with a stride of k is required to extract information.
[0097] Step4: Introduce the FocalLoss loss function, and combine it with the original binary cross-entropy loss function CELoss in the YOLOv8 model to perform weighted fusion on the two loss functions.
[0098] Specifically, the standard binary cross-entropy loss function is used to evaluate the difference between the predicted value and the true label, and is suitable for dealing with the case of relatively balanced categories. Its calculation formula is:
[0099] CE(p,y) = -y·log(p) - (1 - y)·log(1 - p)
[0100] Among them, p is the predicted probability (obtained through the Sigmoid activation function output by the model); y is the true label (1 represents the target, and 0 represents the background).
[0101] FocalLoss introduces a modulating factor (1 - p t ) Υ , which can reduce the loss of samples that are easy to classify, so that the network pays more attention to samples that are difficult to classify, especially the target objects. Its formula is:
[0102] FL(p t ) = -α·(1 - p t ) Υ ·log(p t )
[0103] Among them, p t is the predicted probability of the target class (when the target is the positive class, p t = p, when the target is the negative class, p t = 1 - p); α is the balance factor, which adjusts the weights of positive and negative samples to prevent the excessive influence of negative samples; Υ is the modulating factor, which is used to control the weights of difficult samples. In the present invention, Υ = 2 is set.
[0104] In order to combine Focal Loss and the traditional cross-entropy loss function, the advantages of both can be utilized simultaneously through weighted fusion. The weighted fusion loss function is expressed as:
[0105] Loss = λ·FL(p t )+(1 - λ)·CE(p,y)
[0106] Among them, Loss is the total loss after weighted fusion of the two loss functions, FL(p t ) is the loss calculated by FocalLoss; CE(p,y) is the loss calculated by the traditional binary cross-entropy; λ is a hyperparameter used to balance the weights of the two losses; through the method of weighted fusion, the model can not only pay attention to the target objects that are difficult to classify, but also utilize the stability and simplicity of the cross-entropy loss, and avoid excessive attention to the background area during training.
[0107] Step5: First, through the improvement of the original YOLOv8 model in Step2 - Step4, the ARF_YOLO model is obtained. Then, the training set is used to train the ARF_YOLO model to obtain the final training weights. Finally, the validation set is used to validate the model, and finally the detection of children's wrist fracture images is realized.
[0108] Specifically, after building the ARF_YOLO model, the GRAZPEDWRI-DX dataset is used to train and verify the superiority of the model. Under the conditions of the same experimental environment and the same network parameters, the ARF_YOLO model is compared with other object detection models to verify the effectiveness and superiority of the present invention.
[0109] YOLOv8 is used as the baseline model during training, and the SDG optimizer is used for optimization.
[0110] Table 1 Model training hyperparameter settings
[0111]
[0112] To verify the superiority of the proposed ARF_YOLO model of the present invention, under the conditions of the same experimental environment and the same network parameters, on the GRAZPEDWRI-DX dataset, it is compared with other object detection algorithms, and the results are shown in Table 2.
[0113] Table 2 Detection effects of different algorithms on the GRAZPEDWRI-DX dataset
[0114]
[0115] It can be seen from the experimental results that the detection effect of the ARF_YOLO model on the GRAZPEDWRI-DX dataset is better than that of other models. Especially, it reaches 65.9% in mAP50 (50% average precision), which is significantly higher than that of CAF-YOLO (62.6%) and YOLOv8-ResCBAM (62.3%). ARF_YOLO performs more balanced in terms of precision (P), recall (R), and computational cost (FLOPs), demonstrating that the improved ARF_YOLO model reduces the consumption of computing resources while maintaining high detection accuracy. This indicates that ARF_YOLO can effectively improve the detection performance in the detection task of children's wrist fracture images and also shows excellent performance in computational efficiency. Specific children's fracture detection diagrams are shown in Figure 4 . The green, blue, and pink labels in the figure represent fracture, metal, and periosteal reaction respectively. Multiple detection boxes match these target areas, indicating that the model can effectively identify pathological features. The confidence scores (such as 0.8, 0.9) are relatively high, especially in fracture detection, and most of the predicted values exceed 0.7, showing the advantages of the model in fracture recognition.
[0116] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, making the same replacement or change, shall be covered by the protection scope of the present invention.
Claims
1. A method for detecting children's wrist fracture images based on improved YOLOv8, characterized in that: Step1: Select a dataset of children's wrist fracture images, divide the dataset into a training set and a validation set, and perform data augmentation; Step2: Based on the YOLOv8 network as the basic framework, introduce an Adaptive Hybrid Feature Enhancement module AHF in the Neck network to improve the network's feature extraction ability; Step3: Introduce a Receptive Field Attention Convolution operation RFAConv in the Backbone of YOLOv8 to enhance the extraction ability of multi-scale features; Step4: Introduce the Focal Loss function, and combine it with the original binary cross-entropy loss function CE Loss in the YOLOv8 model, and perform weighted fusion on the two loss functions; Step5: First, through the improvement of the original YOLOv8 model in Step2 - Step4, obtain the ARF_YOLO model, then use the training set to train the ARF_YOLO model to obtain the final training weights, and finally use the validation set to verify the model, and finally realize the detection of children's wrist fracture images.
2. The method for detecting children's wrist fracture images based on improved YOLOv8 according to claim 1, wherein The specific content of Step2 is as follows: The AHF combines a channel attention module SE and a spatial attention module SA, and improves the network's feature extraction ability through DynamicInit and a feature enhancement mechanism, specifically: The input feature map X initializes the module parameters through the Dynamic Init, X ∈ R C×H×W , where C is the number of channels, and H and W are the height and width of the feature map respectively; The SE module performs channel enhancement on the input feature map X, and uses global average pooling GAP to compress the input feature map X into channels: Generate the channel weight W through two layers of fully connected networks and non-linear activation c : W c = σ(W2·ReLU(W1·Z c )) Among them, Z c , W c ∈R C , Z c represents the channel global feature vector obtained after global average pooling, X(c, i, j) represents the input feature map, c represents the channel index, (i, j) represents the pixel position, W1 and W2 are the weight matrices of the fully connected layers, σ is the Sigmoid activation function, and ReLU represents the activation function; Use the channel weight W c Perform channel weighting on the input feature map: X SE = W c · X The SA module performs spatial enhancement on the input feature map X, and uses convolution operations to extract spatial distribution features: F conv = ReLU(BN(Conv2d(X))) Use F conv and the Sigmoid activation function to generate the spatial weight W SA : W SA = σ(F conv ) Using the spatial weight W SA Perform spatial weighting on the input feature map: X SA = W SA ·X Among them, F conv , W SA ∈R C×H×W , X SE represents the output feature map of the SE module, and X SA represents the output feature map of the SA module. F com represents the spatial features extracted from the convolution operation. BN represents batch normalization of the result after convolution, and Conv2d(X) represents two-dimensional convolution of the input feature map X to extract spatial information; The outputs of the SE module and the SA module are weighted and fused, and the final output is: X' = X + X SE + X SA Among them, X' is the final output.
3. The method for detecting children's wrist fracture images based on improved YOLOv8 according to claim 1, wherein, The specific content of Step3 is as follows: The structure of the RFAConv includes a 3×3 convolution kernel, and uses fast grouped convolution to extract receptive field features, average pooling to aggregate global information, and a 1×1 grouped convolution operation with softmax to calculate the importance of each feature in the receptive field, specifically: The input feature map undergoes feature extraction through fast grouped convolution; Use the AvgPool operation to aggregate global information and capture the context of the image; Adjust the feature weights through 1×1 grouped convolution, and normalize the importance of each feature using softmax; Combine the adjusted features with the original features through a weighted mechanism to form a weighted feature map and pass it to the next layer of the network; The calculation formula of the RFAConv is expressed as follows: F = Softmax(g 1×1 (AvgPool(X))) × ReLU(Norm(g k×k (X))) = A rf × F rf Among them, F represents the field space feature obtained by multiplying the attention map A rf with the transformed receptive field map F rf to obtain the field space feature, g i×i represents grouped convolution of size i×i, k represents the size of the convolution kernel, Norm represents normalization, and X represents the input feature map; after shape adjustment, the height and width of the feature map are increased by k times, and k×k convolution with a stride of k is required to extract information.
4. A method for detecting children's wrist fracture images based on improved YOLOv8 according to claim 1, characterized in that, The specific method of performing weighted fusion on the two loss functions is as follows: Loss=λ·FL(p t )+(1-λ)·CE(p,y) Among them, Loss is the total loss after weighted fusion of two loss functions, and FL(p t ) is the loss calculated by Focal Loss; CE(p, y) is the loss calculated by the traditional binary cross-entropy; λ is a hyperparameter used to balance the weights of the two losses, p is the predicted probability obtained through the Sigmoid activation function output by the model; y is the true label, and p t is the predicted probability of the target class. When the target is the positive class, p t = p. When the target is the negative class, p t = 1 - p.
Citation Information
Cited By
Optimized YOLOv8-based anesthetic psychotropic drug identification model and training method
CN121259507A