Target identification method of intelligent table tennis picking robot based on improved YOLOv8
By introducing a dynamic detection head module and a lightweight cross-scale feature fusion module in the YOLOv8 model, the problem of intelligent picking table tennis robots identifying table tennis positions in complex backgrounds and high occlusion scenes is solved, and efficient, accurate and real-time target recognition effects are achieved.
Patent Information
- Application Number
- CN202411917804.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-27
AI Technical Summary
Existing intelligent pick-up table tennis robots are difficult to accurately identify table tennis positions in complex backgrounds and high occlusion scenarios, and the real-time requirements are high but the computing resources are limited, resulting in recognition delays and errors.
The improved YOLOv8 target recognition method is adopted, and a dynamic detection head module and a lightweight cross-scale feature fusion module are introduced. The feature processing method is dynamically adjusted to improve the target detection accuracy in complex backgrounds, and the multi-scale object detection capability is enhanced through cross-scale feature fusion to meet real-time requirements.
It significantly improves the accuracy and real-time target detection, can effectively identify and locate table tennis, reduce recognition delays and errors, and improves the overall performance of the intelligent pick-up table tennis robot.
Smart Images

Figure CN120047781A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of robot technology and target detection, and specifically to a target recognition method for an intelligent table tennis picking robot based on improved YOLOv8. Background Art
[0002] With the rapid development of artificial intelligence technology, intelligent table tennis training has become an important future direction. By applying advanced machines, such as intelligent ball throwers and intelligent ball pickers, exclusive technical training plans can be formulated for athletes, and their techniques can be tutored through professional teaching videos. In addition, through the intelligent table tennis training system, athletes can also complete various technical training tasks alone, thereby better mastering techniques and more easily achieving goals. However, a large amount of cumbersome ball picking work is a major problem in training. Therefore, by designing a robot with an automatic ball picking function, the intelligent table tennis training can be completed more effectively. It can quickly and accurately find the table tennis ball and pick it up, thus effectively reducing the input of manpower and greatly improving the quality and efficiency of training.
[0003] In an intelligent table tennis picking robot, target recognition technology is the key to achieving precise table tennis picking. Existing target recognition technologies for intelligent table tennis picking robots still pose challenges: in actual scenarios, the table tennis ball may be blocked by other objects or obstacles, resulting in the robot being unable to correctly identify the position of the table tennis ball; the intelligent table tennis picking robot needs to process sensor data in real time and make responses. Especially in high-speed movement and dynamic environments, the real-time requirement is very high. However, the computing resources of the robot platform are limited, and a too complex model may lead to system latency and recognition errors. Summary of the Invention
[0004] To address the above challenges, the present invention proposes a target recognition method for an intelligent table tennis picking robot based on improved YOLOv8, aiming to serve the intelligent table tennis training scenario. As an important sub-module of the training scenario, this system is expected to efficiently complete the recognition and positioning of table tennis balls, and assist the robot in intelligent ball picking operations through autonomous navigation functions and autonomous obstacle avoidance capabilities.
[0005] To meet the actual application requirements, the present invention provides an improved YOLOv8 target recognition method, aiming to solve the occlusion problem and real-time challenges existing in the prior art by introducing a dynamic detection head module and a lightweight cross-scale feature fusion module into the YOLOv8 model. Experimental results show that the method proposed in the present invention has achieved obvious effects in solving the problems of target detection accuracy and real-time performance.
[0006] The target recognition method for an intelligent table tennis picking robot based on improved YOLOv8 specifically includes:
[0007] (1) Dynamic Detection Head Module
[0008] The Dynamic Head structure automatically adjusts the processing method of head features according to the complexity of the input image and the target occlusion situation, thereby improving the target detection effect in complex backgrounds or occlusion scenarios. This structure significantly improves the representation ability of the target detection head by coherently combining multiple self-attention mechanisms between scale-aware feature levels, between spatially-aware spatial positions, and within task-aware output channels. In complex dynamic environments, especially in scenarios with target occlusion, it can effectively improve the recognition accuracy of small target objects by the model, thus enhancing the performance of the entire target detection model.
[0009] (2) Lightweight Cross-Scale Feature Fusion Module
[0010] The lightweight cross-scale feature fusion module is mainly used to improve the Neck structure of the traditional YOLOv8 model. By cross-scale feature fusion, it enhances the model's detection ability for multi-scale targets while maintaining a low computational cost. This module inserts N fusion blocks (N is default 3) composed of RepBlocks into the feature map fusion path. These fusion blocks fuse features of different scales in a certain order, which can not only fully extract the boundary information of the feature map but also meet the real-time requirements of the robot through lightweight design.
[0011] The specific process of the method includes the following steps:
[0012] S1 Input the picture into the backbone network of the model, extract features from it, and generate feature maps P1, P2, P3, P4, P5.
[0013] S2 Input the feature maps P3, P4, P5 extracted by the backbone network into the lightweight cross-scale feature fusion module. The fusion module fuses features of different scales, so that the fused feature map contains information from different scales.
[0014] S3 Input the fused feature map into the dynamic detection head module, automatically adjust the processing method of head features according to the complexity of the input image and the target occlusion situation, dynamically turn on or off feature channels to support different tasks, improve the classification accuracy and localization accuracy of the target, and finally output the target detection result, including the category and bounding box of the target. Description of the Drawings
[0015] Figure 1 It is the structure diagram of the intelligent table tennis picking robot target recognition method based on the improved YOLOv8
[0016] Figure 2 It is the schematic diagram of the dynamic detection head
[0017] Figure 3 The structural diagram of the lightweight cross-scale feature fusion module
[0018] Figure 4 The results of model pre-training Figure 5 The experimental results Specific implementation manners
[0019] The following further elaborates on the specific implementation manners of the present invention in conjunction with the accompanying drawings.
[0020] The present invention proposes an intelligent picking table tennis robot target recognition method based on improved YOLOv8, and the specific algorithm structural diagram is as shown in the appended Figure 1 figure.
[0021] (1) Dynamic detection head module
[0022] The detection head part of the YOLOv8 model is responsible for extracting the classification and localization information of the target from the feature map. However, the traditional detection head may have insufficient local feature extraction in the face of dynamic targets and high occlusion scenarios, thus affecting the accuracy and real-time performance of target detection.
[0023] To solve this problem, the present invention introduces a dynamic detection head module. The Dynamic Head structure is a kind of dynamic detection head, which is used to unify scale perception, spatial perception and task perception. The principle of the Dynamic Head structure is as shown in the appended Figure 2 figure, and it includes three different attention mechanisms: scale perception attention, spatial perception attention and task perception attention. First, an image is collected through a USB camera, and the collected picture is input into the model. After the model extracts features and fuses the features, an input feature composed of the features of L different levels in the feature pyramid can be obtained stitched together, where F i represents the feature tensor from the i-th layer (i represents the index of the layer). Through upsampling and downsampling operations, the features of consecutive levels are adjusted to the same scale as the intermediate layer features. The adjusted feature pyramid can be represented as a tensor F∈R L×S×C , where L represents the number of layers in the pyramid, S represents the height×width of the intermediate layer features, and C represents the number of channels of the intermediate layer features (in this paper, L is taken as 3, S is taken as 40×40, and C is taken as 256). The formula for using attention is as shown in formula (1):
[0024] W(F) = π(F)·F (1)
[0025] Among them, W(F) is the output feature tensor obtained after being processed by the attention mechanism, and π(·) is an attention function. Implementing this attention function directly through a fully connected layer is a simple solution. However, due to the high dimensionality of the tensor, the computational cost of directly learning the attention function for all dimensions is too high. We decompose the attention function into three consecutive attention mechanisms, each of which only focuses on one dimension, as shown in Equation (2):
[0026] W(F)
[0027] = π C (π S (π L (F)·F)·F)
[0028] ·F(2)
[0029] Among them, π L (·), π S (·), π C (·) are three different attention functions acting on dimensions L, S, and C respectively.
[0030] 1) Scale-aware attention module: First, introduce scale-aware attention to dynamically fuse features according to the semantic importance of different scales, as shown in Equation (3):
[0031]
[0032] Among them, f(·) is a linear function approximated by a 1×1 convolutional layer, is a hard-sigmoid function.
[0033] 2) Spatial-aware attention module: Adopt a spatial-aware attention module based on the fused features to focus on the discriminative regions of the consistent coexistence between spatial positions and feature hierarchies: First, use deformable convolution to make the attention learning sparse, and then aggregate features across layers at the same spatial position, as shown in Equation (4):
[0034]
[0035] where K is the number of sparse sampling positions, l represents the index of the feature layer, p k is the basic sampling position in the feature map, Δp k is the learned offset, representing the displacement adaptively learned from the feature map, used to focus on the discriminative region, Δm k is the weight learned by the network according to the input features at the sampling position p k +Δp k to adjust the contribution of this position to the final output feature. Δp k and Δm kIt is learned through convolution operations from the median layer of the input feature map F.
[0036] 3) Task-aware attention module: To achieve joint learning and generalization of different representations of objects, this module dynamically switches feature channels (ON means on, OFF means off), as shown in Equation (5):
[0037] π C (F)·F = max(α 1 (F)·F c +β 1 (F), α 2 (F)·F c +β 2 (F)) (5)
[0038] where F c is the feature slice of the feature map F in the c-th channel. [α 1 , α 2 , β 1 , β 2 T are dynamic parameters learned through the hyperfunction θ(·), and θ(·) is used to control the activation threshold of the feature channels, that is, the dynamic adjustment of each feature channel. The calculation process of α 1 , α 2 , β 1 , β 2 : First, global average pooling is performed on the L×S dimension of the input feature F (L represents the number of layers of the feature, and S represents the height×width of the feature). The pooled feature is processed through two fully connected layers, and then a normalization layer is used to further adjust the distribution of the feature. Finally, the Sigmoid function is used to normalize the output value to the range [-1, 1].
[0039] The detailed configuration of the dynamic detection head module is as shown in the appendix Figure 3 . This structure significantly improves the representation ability of the object detection head by coherently combining multiple self-attention mechanisms between scale-aware feature hierarchies, between space-aware spatial positions, and within task-aware output channels. In complex dynamic environments, especially in scenarios with object occlusion, it can effectively improve the recognition accuracy of small target objects by the model, thereby enhancing the performance of the entire object detection model.
[0040] (2) Lightweight cross-scale feature fusion module
[0041] Although the dynamic detection head module effectively improves the target recognition accuracy, the number of model parameters also increases accordingly. The computational complexity of the model is not suitable for the robot platform, and it must be lightweight improved to meet the high real-time requirements of the target recognition method for the intelligent table tennis picking robot. The present invention introduces a lightweight cross-scale feature fusion module based on CNN to improve the Neck structure of the traditional YOLOv8 model. The structure of the lightweight cross-scale feature fusion module based on CNN consists of: 6 convolutional layers, 2 upsampling layers, 4 C2f modules, and 4 feature splicing layers. The 1st, 3rd, 6th, 8th, 11th, and 14th layers are convolutional layers, and the size of the convolutional kernel is 1×1. After passing through the convolutional layers, all the feature channel numbers are adjusted to 256; the 2nd and 7th layers are upsampling layers, and the upsampling multiple is 2. The nearest interpolation method is used, and the size of the feature map is enlarged by 2 times after passing through the upsampling layers; the 5th, 10th, 13th, and 16th layers are C2f modules, and the number of channels in the module is 256; the 4th, 9th, 12th, and 15th layers are feature splicing layers, and the splicing dimension parameter is 1, and they are merged in the channel direction.
[0042] The specific process of the method includes the following steps:
[0043] S1 Input the picture into the backbone network of the model. The backbone network consists of 5 convolutional layers, 4 C2f modules, and 1 SPPF module. Each convolutional layer uses a 3×3 convolutional kernel to slide on the image for convolution operation to extract the local features of the image. The stride of the convolutional layer is 2, and each convolutional layer will halve the size of the input feature map. The number of channels of the convolutional layer gradually increases according to (64, 128, 256, 512, 1024). The C2f module first uses a 1×1 convolution to adjust the number of channels of the feature map, and then uses a 3×3 convolution to extract the local spatial information to enhance the expression ability of the features. It combines the low-level and high-level features through skip connections, and finally fuses the features of different scales through the Concat operation to improve the multi-scale target detection ability. SPPF is a fast spatial pooling pyramid layer, and the size of the pooling kernel is 5, which is used for pooling operations of different scales, and splices the feature maps of different scales together to improve the detection ability of targets of different scales. The output feature maps are P1, P2, P3, P4, and P5, and the output size of each layer gradually decreases, but the number of channels gradually increases.
[0044] S2 Input the feature maps P3, P4, and P5 layers extracted by the backbone network into the lightweight cross-scale feature fusion module. The fusion module fuses the features of different scales, so that the fused feature map contains information from different scales. First, after the feature map P5 passes through the 1st convolutional layer, the number of channels is adjusted to 256, and the feature map P 5-1 ; the feature map P 5-1 After passing through the 2nd upsampling layer, the size becomes 2 times, and the feature map P5-2 ; After the feature map P4 passes through the 3rd convolutional layer, the number of channels becomes 256, generating the feature map P 4-1 . After the feature map P3 passes through the 8th convolutional layer, the number of channels becomes 256, generating the feature map P 3-1 . The feature map P 5-2 and the feature map P 4-1 pass through the 4th feature concatenation layer to generate a new feature map P t1 ; The feature map P t1 undergoes 3 operations of the C2f module in the 5th layer, and then passes through the 6th convolutional layer, and the number of channels is adjusted to 256, generating the feature map P v1 ; The feature map P v1 after the 7th upsampling operation and the feature map P generated by the 8th layer 3-1 pass through the 9th feature concatenation layer to generate a new feature map P t2 ; The feature map P t2 undergoes 3 operations of the C2f module in the 10th layer to generate the feature map P f1 ; The feature map P f1 after passing through the 11th convolutional layer and the feature map P generated by the 6th layer v1 pass through the 12th feature concatenation layer to generate a new feature map P t3 ; The feature map P t3 undergoes 3 operations of the C2f module in the 13th layer to generate the feature map P f2 ; The feature map P f2 after passing through the 14th convolutional layer and the feature map P generated by the first layer 5-1 pass through the 15th feature concatenation layer to generate a new feature map P t4 ; The feature map P t4 undergoes 3 operations of the C2f module in the 16th layer to generate the feature map P f3 .
[0045] S3 inputs the fused feature maps P f1 , P f2 and P f3 into the dynamic detection head module. For each layer of feature map, a 3×3 convolutional kernel is used to generate the offset and mask, and the modulated deformable convolution is applied. According to the offset and mask, the feature map is dynamically convolved with a dynamic kernel size of 3×3 to obtain the enhanced feature map; for the feature maps P f1 , P f2 and P f3 , 1×1 convolution operations are performed to uniformly adjust the number of channels to 256; then for the feature maps P f1 , P f2 and P f3Point - by - point weighted fusion is performed through an attention mechanism (as shown in Equation 2), and the Softmax function is used as the activation function to obtain the fused feature map. The dynamic detection head predicts the class of the target: First, a 3×3 convolution operation is performed on the fused feature map, with the ReLU function as the activation function, then a 1×1 convolution is used to adjust the channel dimension to the number of classes \(n_c\) (in this paper, \(n_c = 47\)), and finally the Sigmoid activation function is used to output the confidence of each class. The dynamic detection head predicts the bounding box coordinates of the target: A 3×3 convolution operation is performed on the fused feature map, with the ReLU function as the activation function, and then a 1×1 convolution is used to adjust the channel dimension to the regression parameters (usually 4 values: \(x\), \(y\), \(w\), \(h\). Here, \(x\) and \(y\) represent the horizontal and vertical coordinates of the center point of the bounding box respectively; \(w\) and \(h\) represent the width and height of the bounding box respectively). Finally, the class and bounding box of the target are output.
[0046] The dataset used in the training of the present invention is composed of three scenarios in the large - scale indoor dynamic scene dataset (THUD) for mobile robots released by Tsinghua University, plus a self - made table tennis dataset, a total of four indoor scenarios. Among them, the training set contains about 6000 images, and the validation set contains about 1300 images. The dataset is collected through a real robot platform, which contains dynamic objects such as moving pedestrians, mobile robots, shopping carts, etc., with different degrees of dynamic complexity. In addition, in complex dynamic scenarios, this dataset contains different degrees of target occlusion phenomena.
[0047] During the training process of the present invention, an adaptive anchor box design is used to automatically analyze the target distribution of the dataset, generate the best anchor box set (Anchor Box Set) to minimize the matching error, reduce unnecessary parameter tuning work, and effectively cope with the diversity of target sizes and distributions in different scenarios. An anchor box is a set of bounding boxes with fixed shapes defined on the feature map, used as a reference starting point for predicting the target. The center position, width, and height of the predicted bounding box (Bounding Box) are adjusted based on the anchor box. Calculate the intersection - over - union (IoU) between the anchor box and the ground truth (GT) in the dataset. The anchor box with the largest IoU is regarded as matching this target. This method first performs statistical analysis on the target bounding boxes in the training data, extracts the width and height information of all target boxes, and uses the K - Means++ clustering algorithm to generate the anchor box sizes and ratios adapted to the dataset. By using the intersection - over - union (IoU) as the similarity metric, the matching relationship between the target box and the anchor box is automatically determined to ensure that the anchor box covers the main range of the target distribution. The clustering process is shown in Equation (14):
[0048]
[0049] Among them, B is the target box and A is the anchor box. The matching degree between each anchor box and the target box is calculated according to IoU, and the anchor boxes are allocated by optimizing IoU. After adjustment, the anchor boxes generated by clustering can adapt to the size changes of the targets in the multi-scale feature maps. In addition, during the training process, the anchor box design is dynamically optimized, and the anchor box sizes are readjusted according to the changes in the target distribution, enhancing the model's adaptability to dynamic and complex scenarios. This strategy reduces the complexity of manual parameter tuning, effectively improves the model's detection performance for small and complex targets, and ensures robustness and generalization in multiple scenarios, providing a solid foundation for efficient training.
[0050] In the target detection model of the intelligent table tennis picking robot based on the improved YOLOv8, the overall loss function is the weighted sum of the location loss (locatization loss, loc), the confidence loss (confidence loss, conf), and the classification loss (classification loss, cls), as shown in Equation (15):
[0051] L total = λ loc × L loc + λ conf × L conf + λ cls × L cls (15)
[0053] Among them, L loc is the location loss, which is used to measure the position deviation between the predicted box and the ground truth box; L conf is the confidence loss, which is used to measure the confidence level that the target box contains the target; L cls is the classification loss, which is used to measure the deviation between the predicted class and the ground truth class; λ loc , λ conf , and λ cls are the weights of each item.
[0054] The loss function used for the location loss (L loc ) is the CIOU Loss (Complete IoU Loss), and the formula is as shown in (16):
[0055]
[0056] Among them, ρ 2 (b, b gt ) is the square of the Euclidean distance between the center points of the predicted box and the ground truth box; c is the length of the diagonal of the smallest closed bounding box containing the predicted box and the ground truth box; and the calculation formulas of α are as shown in Formulas (17) and (18) respectively:
[0057]
[0058] where ω, h: the width and height of the prediction box; ω gt , h gt : the width and height of the ground truth box.
[0059] The loss function used for the confidence loss (L conf ) is the binary cross-entropy loss (BCE), and the formula is shown in (19):
[0060]
[0061] where N is the number of samples, y i is the target value (1 represents a positive sample, 0 represents a negative sample), and p i is the predicted confidence score.
[0062] The loss function used for the classification loss (L cls ) is the categorical cross-entropy loss, and the formula is shown in (20):
[0063]
[0064] where N is the number of samples, C is the number of classes, y i,c is the true label of the i-th sample. If the sample belongs to class c, then y i,c = 1, otherwise 0; p i,c is the probability that the i-th sample is predicted to be class c.
[0065] Subsequently, the loss function is minimized for training. The Stochastic Gradient Descent (SGD) method is used to minimize the above cost function. Predictions are calculated for the feature map results of multiple convolutional layers with different scales, and the prediction outputs of different layers are combined. Training requires normalizing all data sizes, so in the present invention, the original image is reset to 640×640 pixels for training. The learning rate is the most important parameter in the Stochastic Gradient Descent method, which determines the speed of weight update. The momentum parameter and the weight decay factor can improve the training adaptability. Through experimental observation, in the present invention, the initial value of the learning rate is set to 0.01, the momentum parameter is set to 0.937, and the weight decay factor is set to the default 0.0005. The Stochastic Gradient Descent learning process is accelerated by the NVIDIA GeForce RTX3090 device, with 16 images processed in each batch, and the number of training epochs is 500. During the training process, Non-Maximum Suppression (NMS) will be used for post-processing to suppress redundant boxes. All boxes are sorted according to the confidence level, and the box with the highest confidence level is selected. The IoU value between the currently selected box and the remaining boxes is calculated. If the IoU value between the currently selected box and other boxes is greater than the preset threshold of 0.7, these boxes are considered to be duplicates and are deleted. The above process is repeated continuously, and finally, the boxes with higher confidence levels are retained. The training result reaches the optimal value at the 333rd epoch. The experimental results are as shown in Figure 5 Figure . The object detection model of the intelligent table tennis ball picking robot based on the improved YOLOv8 can achieve an mAP (mean average precision) of 72.7%, and the GFLOPs (giga floating point operations per second) is 7.5. Therefore, this method is practical and has important application value for realizing efficient, accurate, and real-time object recognition of the intelligent table tennis ball picking robot.
Claims
1. The target recognition method of intelligent table tennis picking robot based on improved YOLOv8 is characterized by Includes the following modules: (1) Dynamic detection head module The dynamic detection head module is introduced; the Dynamic Head structure is a dynamic detection head that is used to unify scale perception, space perception, and task perception; the Dynamic Head structure contains three different attention mechanisms: scale-aware attention, space-aware attention, and task-aware attention; First, a USB camera is used to capture images, and the captured images are input into the model. The model extracts features and fuses the features to obtain a feature pyramid composed of L different levels of features. The concatenated input features, where F i Represents the feature tensor from the i-th layer, i represents the index of the layer, and the features of the consecutive layers are adjusted to the same scale as the intermediate layer features through upsampling and downsampling operations; the adjusted feature pyramid is represented as a tensor F∈R L×S×C , where L represents the number of layers in the pyramid, S represents the height × width of the middle layer feature, and C represents the number of channels of the middle layer feature. L is 3, S is 40 × 40, and C is 256. The formula for using attention is shown in formula (1): Where W(F) is the output feature tensor after attention mechanism processing, and π(·) is an attention function; The attention function is decomposed into three consecutive attention mechanisms, each of which focuses on only one dimension, as shown in formula (2): W(F)=π C (p S (p L (F)·F)·F)·F (2) Among them, π L (·), π S (·), π C (·) are three different attention functions acting on dimensions L, S, and C respectively; 1) Scale-aware attention module: First, we introduce scale-aware attention to dynamically fuse features according to the semantic importance of different scales, as shown in formula (3): where f(·) is a linear function approximated by a 1×1 convolutional layer, It is a hard-sigmoid function; 2) Spatial perception attention module: First, deformable convolution is used to make the attention learning sparse, and then the features are aggregated across layers at the same spatial location, as shown in formula (4): Where K is the number of sparse sampling positions, l represents the index of the feature layer, and p k is the base sampling position in the feature map, Δp k is the learned offset, which represents the displacement adaptively learned from the feature map; Δm k The network samples at position p according to the input features. k +Δp k The learned weights; 3) Task-aware attention module: This module dynamically switches feature channels to support different tasks, as shown in formula (5): p C (F)·F=max(α 1 (F)·F c +b 1 (F),a 2 (F)·F c +b 2 (F)) (5) Among them, F c is the feature slice of the cth channel of the feature map F, [α 1 ,α 2 ,β 1 ,β 2 ] T is a dynamic parameter learned by the hyperfunction θ(·), which is used to control the activation threshold of the feature channel, that is, the dynamic adjustment of each feature channel; α 1 ,α 2 ,β 1 ,β 2 The calculation process is as follows: First, the L×S dimension of the input feature F (L represents the number of layers of the feature, S represents the height×width of the feature) is globally averaged pooled. The pooled features are processed by two fully connected layers, and then the normalization layer is used to further adjust the distribution of the features. Finally, the Sigmoid function is used to normalize the output value to the range of [-1,1]; (2) Lightweight cross-scale feature fusion module A lightweight cross-scale feature fusion module based on CNN is introduced to improve the Neck structure of the traditional YOLOv8 model; the structure of the lightweight cross-scale feature fusion module based on CNN includes: 6 convolutional layers, 2 upsampling layers, 4 C2f modules, and 4 feature splicing layers. The 1st, 3rd, 6th, 8th, 11th, and 14th layers are convolutional layers with a convolution kernel size of 1×1. After the convolutional layers, the number of feature channels is adjusted to 256; the 2nd and 7th layers are upsampling layers with an upsampling multiple of 2. After the upsampling layers, the size of the feature map is enlarged by 2 times; the 5th, 10th, 13th, and 16th layers are C2f modules with a channel number of 256; the 4th, 9th, 12th, and 15th layers are feature splicing layers with a splicing dimension parameter of 1, which are merged in the channel direction; The specific process includes the following steps: S1 inputs the image into the backbone network of the model; the backbone network consists of 5 convolutional layers, 4 C2f modules and 1 SPPF module; each convolutional layer uses a 3×3 convolution kernel to slide on the image for convolution operation to extract local features of the image; the stride of the convolutional layer is 2, and each convolutional layer will halve the size of the input feature map; the convolutional layer gradually increases the number of channels according to (64, 128, 256, 512, 1024); the C2f module first uses 1×1 convolution to adjust the number of channels of the feature map, and then uses 3×3 convolution to extract local spatial information to enhance the expressiveness of the features, combines low-level and high-level features through jump connections, and finally fuses features of different scales through the Concat operation to improve the ability of multi-scale target detection; SPPF is a fast spatial pooling pyramid layer, the size of the pooling kernel is 5, and the output feature maps are P1, P2, P3, P4, P5; S2 inputs the feature maps P3, P4, and P5 extracted by the backbone network into the lightweight cross-scale feature fusion module; first, after the feature map P5 passes through the first convolution layer, the number of channels is adjusted to 256 to generate the feature map P 5-1 ; Feature map P 5-1 After the second upsampling layer, the size becomes twice as large, generating the feature map P 5-2 ; After the feature map P4 passes through the third convolution layer, the number of channels becomes 256, generating the feature map P 4-1 After the feature map P3 passes through the 8th convolutional layer, the number of channels becomes 256, generating the feature map P 3-1 ; Feature map P 5-2 and feature map P 4-1 After the 4th feature concatenation layer, a new feature map P is generated t1 ; Feature map P t1 After the fifth layer, the C2f module operation is performed three times, and then after the sixth convolution layer, the number of channels is adjusted to 256 to generate the feature map P v1 ; Feature map P v1 After the upsampling operation of the 7th layer and the feature map P generated by the 8th layer 3-1 After the 9th feature concatenation layer, a new feature map P is generated t2 ; Feature map P t2 After the 10th layer, the C2f module operation is performed three times to generate the feature map P f1 ; Feature map P f1 After the 11th convolutional layer and the feature map P generated by the 6th layer v1 After the 12th feature concatenation layer, a new feature map P is generated t3 ; Feature map P t3 After the 13th layer, the C2f module operation is performed three times to generate the feature map P f2 ; Feature map P f2 After the 14th convolutional layer and the feature map P generated by the first layer 5-1 After the 15th feature concatenation layer, a new feature map P is generated t4 ; Feature map P t4 After the 16th layer, the C2f module operation is performed three times to generate the feature map P f3 ; S3 takes the fused feature map P f1 , P f2 and P f3 The input is sent to the dynamic detection head module. For each layer of feature maps, a 3×3 convolution kernel is used to generate an offset Offset and a mask Mask. The feature map is dynamically convolved according to the offset and the mask. The dynamic kernel size is 3×3 to obtain the enhanced feature map. f1 , P f2 and P f3 Perform a 1×1 convolution operation to uniformly adjust the number of channels to 256; then perform a 1×1 convolution operation on the feature map P f1 , P f2 and P f3 The attention mechanism is used to perform point-by-point weighted fusion, and the activation function uses the Softmax function to obtain the fused feature map. The dynamic detection head predicts the category of the target: first, a 3×3 convolution operation is performed on the fused feature map, and the activation function is the ReLU function. Then, a 1×1 convolution is used to adjust the channel dimension to the number of categories ncnc = 47. Finally, the Sigmoid activation function is used to output the confidence of each category. The dynamic detection head predicts the bounding box coordinates of the target: a 3×3 convolution operation is performed on the fused feature map, the activation function is the ReLU function, and then a 1×1 convolution is used to adjust the channel dimension to the regression parameter, which is 4 values: x, y, w, h; x and y represent the horizontal and vertical coordinates of the center point of the bounding box, respectively; w and h represent the width and height of the bounding box, respectively, and finally output the category and bounding box of the target.
Citation Information
Cited By
Real-time target detection model and method fusing multi-scale feature enhancement and dynamic label distribution
CN121121160A