Small target gesture area detection method based on improved YOLOv8 algorithm

By building an improved YOLOv8 algorithm network, the problem of insufficient accuracy in small-target gesture detection is solved, and the detection accuracy of small-target gestures is improved through adaptive downsampling and task alignment technology.

CN120236334AInactive Publication Date: 2025-07-01ANHUI UNIV

Patent Information

Application Number
CN202510382883.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing gesture recognition detection methods have insufficient accuracy in small-target gesture detection, especially due to the small proportion of gesture areas, low resolution and large positioning errors caused by multi-scale feature interference, resulting in missed or misdetection.

Method used

Build a small-target gesture dataset, build an improved YOLOv8 algorithm network, including backbone network, neck network and head network, and adopts an adaptive downsampling Adown module, AIFI module and a task-aligned one-stage task detection head T-head network to extract, fusion and detect small-target gestures through feature extraction, fusion and detection.

Benefits of technology

The accuracy of small-target gesture detection is improved, the problems of insufficient feature extraction of small-targets and multi-scale interference are solved, and higher detection accuracy is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236334A_ABST
    Figure CN120236334A_ABST
Patent Text Reader

Abstract

The invention discloses a small target gesture area detection method based on an improved YOLOv8 algorithm. The method comprises the following steps: constructing a small target gesture data set; establishing a small target gesture area detection network based on an improved YOLOv8 algorithm, wherein the small target gesture area detection network comprises a backbone network, a neck network and a head network; wherein the backbone network comprises a multi-layer convolution, a C2f module, an improved adaptive downsampling ADown module and an AIFI module, and the head network is an improved one-stage task detection head T-head network for task alignment; and performing feature extraction on an input image through the backbone network, fusing different convolutional feature maps through the neck network, and detecting a small hand target through the head network. According to the invention, the detection precision of the network on small target gestures is improved, and a technical reference can be provided for small-hand area detection in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of gesture recognition, and particularly relates to a small target gesture region detection method based on an improved YOLOv8 algorithm. Background Technique

[0002] With the rapid development of artificial intelligence technology, the human-computer interaction technology has been continuously innovated. As an important part of it, gesture recognition plays a crucial role in designing intelligent and efficient human-computer interaction due to its rich information. At present, the detection methods for gesture recognition mainly focus on optimizing gesture keypoint detection or introducing complex attention mechanisms. People solve the problem of inaccurate gesture estimation caused by different background conditions, occlusion of hands, etc. by increasing the network depth or introducing multi-scale detection modules, so as to improve the detection accuracy. In the field of object detection, the introduction of deep learning has significantly improved the detection accuracy. Supervised object detection algorithms are mainly divided into one-stage and two-stage object detection algorithms. They learn a large amount of labeled gesture sample information, enabling the network model to recognize small target gesture categories. Among them, YOLOv8, as an efficient one-stage object detection network, can achieve good detection results and meet the requirements for small target gesture detection in practical applications in terms of detection speed and accuracy.

[0003] However, there are still some deficiencies in the existing gesture recognition detection methods. Due to the high flexibility of gesture keypoints, self-occlusion and high coupling of fingers, and the small proportion of the gesture region in the image, etc., the detection accuracy of gesture recognition still needs to be further improved. Research has found that gesture data usually varies in size. Small hands account for a small proportion in the image and have low resolution, resulting in the network algorithm often failing to recognize or misrecognizing them, which has a great impact on the further improvement of the detection accuracy of gesture recognition. In addition, the original YOLOv8 backbone network directly fuses feature information of different scales, making the features of different scales influence each other, and the adaptability of the detection head to targets of different scales is poor, especially the positioning error of tiny targets is large. This leads to small target features being easily covered by large-scale features, resulting in missed detection or misdetection, and reducing the accuracy of the network. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a small target gesture region detection method based on an improved YOLOv8 algorithm to solve the problems existing in the above-mentioned prior art.

[0005] To achieve the above object, in the first aspect, the present invention provides a small target gesture region detection method based on an improved YOLOv8 algorithm, including:

[0006] Construct a small target gesture data set, and the small target gesture data set includes gesture images with a small proportion of the gesture region;

[0007] Build a small target gesture area detection network based on the improved YOLOv8 algorithm. The small target gesture area detection network includes a backbone network, a neck network, and a head network;

[0008] Among them, the backbone network includes: multiple layers of convolution, C2f module, improved adaptive downsampling Adown convolution, and AIFI module. The head network is an improved task-aligned one-stage task detection head T-head network. The input image is feature-extracted by the backbone network, the feature maps of different convolutions are fused by the neck network, and small hand targets are detected by the head network;

[0009] Train the small target gesture area detection network using the small target gesture dataset;

[0010] Input the gesture image to be detected into the trained small target gesture area detection network to obtain a result map containing small hand categories and small hand positions.

[0011] Preferably, the steps of constructing the small target gesture dataset include:

[0012] Collect gesture images with a gesture area ratio less than 30%. The gesture images cover various lighting conditions, background complexities, and hand occlusion scenarios;

[0013] Perform data augmentation operations on the collected gesture images to obtain an augmented dataset;

[0014] Label the augmented dataset and divide it into a training set, a validation set, and a test set according to a ratio.

[0015] Preferably, the adaptive downsampling Adown convolution uses a 2×2 average pooling window to integrate redundant features.

[0016] Preferably, the AIFI module performs interactive processing on features of the same scale through a self-attention mechanism to enhance the fine-grained feature expression ability of small targets.

[0017] Preferably, the neck network adopts a PAN-FPN network, and the PAN-FPN network uses upsampling and downsampling cross-layer fusion connections.

[0018] Preferably, the detection head T-ead network optimizes the consistency of classification and localization tasks through a dynamic task alignment strategy.

[0019] Preferably, the detection head T-head network includes a T-shaped detection head and a task alignment learning mechanism.

[0020] Preferably, the T-shaped detection head is used to generate multi-scale task interaction features from the features fused by the features through multiple layers of cascaded convolutional layers, and then through convolutional operations with different receptive fields, fuse multi-level semantic and localization information.

[0021] Preferably, the task alignment learning mechanism is used to select a fixed number of anchors with the largest predictions as positive sample information for each true label information, and the rest as negative samples.

[0022] In a second aspect, the present invention also discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0023] Compared with the prior art, the present invention has the following advantages and technical effects:

[0024] The present invention provides a small target gesture area detection method based on an improved YOLOv8 algorithm. First, a small target gesture data set is constructed, and the small target gesture data set includes gesture images with a small proportion of gesture areas; secondly, a small target gesture area detection network based on the improved YOLOv8 algorithm is built, and the small target gesture area detection network includes a backbone network, a neck network, and a head network; wherein, the backbone network includes: multiple layers of convolution, C2f module, improved adaptive downsampling Adown convolution, and AIFI module, and the head network is an improved task-aligned one-stage task detection head T-head network; the input image is subjected to feature extraction through the backbone network, the feature maps of different convolutions are fused through the neck network, and small hand targets are detected through the head network; then, the small target gesture area detection network is trained using the small target gesture data set; finally, the gesture image to be detected is input into the trained small target gesture area detection network to obtain a result map containing small hand categories and small hand positions.

[0025] Based on the deep learning object detection algorithm, for the detection of small target gesture areas, the present invention constructs a small target gesture area detection network based on the improved YOLOv8 algorithm; this network effectively solves the problems of insufficient small target feature extraction, multi-scale interference, and low small target localization accuracy existing in the original YOLOv8 network through the collaborative optimization of the adaptive downsampling Adown module, AIFI module, and T-head detection head, improves the detection accuracy of the network for small target gestures, and can provide a technical reference for the detection of small hand areas in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0027] Figure 1 Flow chart of the small target gesture region detection method based on the improved YOLOv8 algorithm in the embodiment of the present invention;

[0028] Figure 2 Structure diagram of the small target gesture region detection network based on the improved YOLOv8 algorithm in the embodiment of the present invention;

[0029] Figure 3 Structure diagram of the adaptive downsampling Adown in the embodiment of the present invention;

[0030] Figure 4 Structure diagram of the AIFI in the embodiment of the present invention;

[0031] Figure 5 Structure diagram of the T-head internal task alignment predictor TAP in the embodiment of the present invention. Detailed implementation manners

[0032] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0033] It should be noted that the steps shown in the flowchart of the accompanying drawings may be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0034] Embodiment 1

[0035] As Figure 1 shown, in this embodiment, a small target gesture region detection method based on the improved YOLOv8 algorithm is provided, including:

[0036] S1. Construct a small target gesture data set, where the small target gesture data set contains gesture images with the proportion of the gesture region less than 30% of the image;

[0037] Further, step S1 includes the following contents:

[0038] 1). Collect gesture images with the proportion of the gesture region less than 30%, and the gesture images cover various lighting conditions, background complexities and hand occlusion scenarios;

[0039] 2). Perform data augmentation operations on the collected gesture images, including operations such as flipping and cropping;

[0040] 3). Label the data set with the LabelImg software and divide it into a training set, a validation set and a test set according to the ratio of 8:1:1.

[0041] S2. Build a small target gesture area detection network based on the improved YOLOv8 algorithm. The small target gesture area detection network includes a backbone network, a neck network, and a head network;

[0042] Among them, the backbone network includes: multiple layers of convolution, C2f modules, improved adaptive downsampling Adown convolution, and AIFI modules. The head network is an improved task-aligned one-stage task detection head T-head network. The input image is subjected to feature extraction through the backbone network, the feature maps of different convolutions are fused through the neck network, and small hand targets are detected through the head network;

[0043] Specifically, the small target gesture area detection network in step S2 is based on the improved YOLOv8 algorithm, as Figure 2 shown.

[0044] First, the input image size is 640*640*3, and it is input into the feature extraction network.

[0045] Secondly, the feature extraction network will pass through two Conv2D_BN_SiLU layers and C2f layers to extract shallow feature information. The 3rd, 4th, and 5th layer convolution structures use the adaptive downsampling Adown structure as Figure 3 shown, followed by a C2f module, which uses a 2×2 average pooling window to integrate redundant features and reduce the loss of small target detail information; an AIFI module is added after the last C2f layer of the backbone network as Figure 4 shown. Through the self-attention mechanism, the features of the same scale are interactively processed to enhance the fine-grained feature expression ability of small targets, and to avoid the small target features being masked by the large-scale features during the fusion process of features at different scales, thereby enhancing the model's ability to extract and express small target features. The feature extraction network will extract three features Feature0, Feature1, and Feature2.

[0046] Finally, the three-layer features Feature0, Feature1, and Feature2 extracted by the feature extraction network will pass through the feature fusion network to extract a feature information, and then be input into the classification prediction network T-head by connecting two 3*3 convolutions. The T-head network can obtain the prediction box and the small target category within the prediction box.

[0047] The classification prediction network T-head mainly consists of a T-shaped detection head and a task alignment learning mechanism inside. Among them, the T-shaped detection head generates multi-scale task interaction features from the features fused by the features through multiple layers of cascaded convolution layers. These features fuse multi-level semantic and localization information through convolution operations with different receptive fields. Some of its expression formulas are as follows:

[0048]

[0049] Among them, δ is the activation function, and X fpn represents the feature map input to the feature fusion network.

[0050] In the T-shaped detection head, dynamic weights are calculated for the classification and localization tasks respectively. The task alignment predictor TAP (as Figure 5 shown) extracts task-related features from the interaction features and then performs alignment adjustment. Some of its related expression formulas are as follows:

[0051] w = δ(fc2(δ(fc1(x inter )))) (2)

[0052]

[0053] M = δ(conv2(δ(conv1(X inter )))) (4)

[0054] O = conv4(δ(conv3(X inter ))) (5)

[0055]

[0056] B align (i, j, c) = B(i + O(i, j, 2×c), j + O(i, j, 2×c + 1), c) (7)

[0057] Among them, the weight w is generated through a fully connected layer and the Sigmoid function; M is the spatial probability map, which is learned from the interaction features and is used to adjust the spatial distribution of the classification prediction; O is the spatial offset map, which learns the offset amount at each position and adjusts the predicted coordinates of the bounding box to make the localization more accurate.

[0058] The task alignment learning mechanism is to select a fixed number of anchors with the largest predictions as positive sample information for each real label information, and the rest as negative samples, so that the boxes with high classification and accurate localization are retained. The expression formula is as follows:

[0059] t = s α ×u β (8)

[0060] Among them, t is the alignment metric, s is the comprehensive classification score, u is the predicted box IOU, and α and β are hyperparameters that control the influence weights of the two.

[0061] The input features generate prediction results through the T-shaped detection head T-head. The task alignment learning mechanism TAL calculates the alignment metric and assigns samples, and backpropagation optimizes the classification and localization losses. Then, the adjusted classification score P is used. align and the bounding box B align , and the final detection result is output after NMS.

[0062] S3. Use the small target gesture dataset to train the small target gesture region detection network;

[0063] S4. Input the gesture image to be detected into the trained small target gesture region detection network to obtain a result map containing the small hand category and the small hand position.

[0064] Embodiment 2

[0065] This embodiment also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in Embodiment 1 are implemented.

[0066] The above is only the preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A small target gesture area detection method based on an improved YOLOv8 algorithm, characterized in that: The following steps are involved: Constructing a small target gesture dataset, wherein the small target gesture dataset includes gesture images in which the gesture area accounts for less than 30% of the image; Building a small target gesture area detection network based on the improved YOLOv8 algorithm, wherein the small target gesture area detection network includes a backbone network, a neck network, and a head network; The backbone network includes: multi-layer convolution, C2f module, improved adaptive downsampling Adown convolution and AIFI module, and the head network is an improved task-aligned one-stage task detection head T-head network; the backbone network is used to extract features of the input image, the neck network is used to fuse feature maps of different convolutions, and the head network is used to detect small hand targets; Using the small target gesture data set to train the small target gesture area detection network; The gesture image to be detected is input into the trained small target gesture area detection network to obtain a result image containing small hand categories and small hand positions.

2. The method according to claim 1, characterized in that The steps to construct a small target gesture dataset include: Collect gesture images with a small gesture area, which cover a variety of lighting conditions, background complexity, and hand occlusion scenarios; Perform data augmentation operations on the collected gesture images to obtain an augmented data set; The enhanced dataset is labeled and divided into a training set, a validation set, and a test set according to a certain ratio.

3. The method according to claim 1, characterized in that The adaptive downsampling Adown convolution adopts a 2×2 average pooling window to integrate redundant features.

4. The method according to claim 1, characterized in that: The AIFI module interactively processes features of the same scale through a self-attention mechanism to enhance the fine-grained feature expression capability of small targets.

5. The method according to claim 1, characterized in that The neck network adopts a PAN-FPN network, and the PAN-FPN network adopts up- and down-sampling cross-layer fusion connection.

6. The method according to claim 1, characterized in that The detection head T-head network optimizes the consistency of classification and positioning tasks through a dynamic task alignment strategy.

7. The method according to claim 1, characterized in that The detection head T-head network includes a T-shaped detection head and a task alignment learning mechanism.

8. The method according to claim 7, characterized in that The T-shaped detection head is used to generate multi-scale task interaction features through multiple layers of serially connected convolutional layers from the features fused, and then fuse multi-level semantic and positioning information through convolution operations with different receptive fields.

9. The method according to claim 7, characterized in that: The task alignment learning mechanism is used to select a fixed number of anchors with the maximum prediction as positive sample information for each real label information, and the rest are negative samples.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Gesture recognition method based on improved YOLOv5 network

    CN118116033A

  • Gesture recognition method under complex background

    CN118522036A

  • Improved target detection method based on YOLOv8s

    CN118982734A

  • Strip steel surface defect detection method based on RAC-YOLO algorithm

    CN119107282A

  • Target detection method based on YOLOv8 improved model

    CN119625494A

Cited By

  • Expressway foreign matter small target detection method, system and device and medium

    CN121191077A