Object Detection Method and Device Based on YOLO-SD-Tiny

By introducing the Mish activation function and Self-DeConvolution module in the YOLO-SD-Tiny model, the YOLOV4-Tiny model is optimized, and the accuracy and real-time problems of object detection on low-performance devices are solved, achieving higher detection accuracy and speed.

CN114565959BActive Publication Date: 2025-07-11WUHAN DONGXIN TONGBANG INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210152654.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-18
Publication Date
2025-07-11
Estimated Expiration
2042-02-18

AI Technical Summary

Technical Problem

The existing object detection algorithm cannot be effectively deployed on low-performance devices, and there are problems of low accuracy and poor real-time performance.

Method used

Using the YOLO-SD-Tiny model, the YOLOV4-Tiny model is optimized to adapt to low-performance devices by introducing the Mish activation function in the backbone feature extraction network and the Self-DeConvolution module in the feature pyramid network.

Benefits of technology

The accuracy and detection speed of target detection have been improved. YOLO-SD-Tiny has improved the accuracy of 6.35% and the detection speed of 9.64% compared with YOLOv4-Tiny on the OccludeFace dataset.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114565959B_ABST
    Figure CN114565959B_ABST
Patent Text Reader

Abstract

The present invention discloses an object detection method and device based on YOLO-SD-Tiny, which relates to the field of object detection. The method includes: after replacing the activation functions adopted in the last CSP-Body and CBL in the YOLOV4-Tiny backbone feature extraction network with Mish activation functions, extracting information from the picture to be detected to obtain effective feature layers; performing Self-DeConvolution upsampling on the effective feature layers according to the Feature Pyramid Network FPN and outputting; using YOLO Head to predict the output values after upsampling. The present invention is applicable to general devices, especially low-performance devices with low computing power, and can improve the accuracy and detection speed of object detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of object detection, and particularly to an object detection method and device based on YOLO-SD-Tiny. Background Art

[0002] Face detection is a very important computer vision task and an important branch of object detection. Object detection algorithms based on deep learning are mainly divided into two categories, one based on region proposals and the other without region proposals.

[0003] Object detection algorithms based on region proposals mainly include R-CNN, Fast R-CNN, Faster R-CNN, etc. These object detection algorithms based on region proposals are divided into two steps. First, a series of candidate boxes are generated based on the object, and then classification and coordinate regression are performed by a convolutional neural network. Object detection algorithms based on region proposals have a high accuracy rate, but their model parameters are often too large and their real-time performance is also poor.

[0004] Object detection algorithms without region proposals mainly include the YOLO series (You Only Look Once) algorithms, which combine the candidate box generation process and classification regression, greatly reducing the computational complexity of the neural network. However, compared with the two-stage method based on region proposals, their accuracy rate is lower. The YOLO algorithm is also constantly evolving. Currently, for example, YOLOV4 is widely used, and its speed and accuracy have been greatly improved compared with the earlier algorithms in the series.

[0005] In recent years, some exquisitely designed network models and operators have been proposed, greatly reducing the amount of computation and the number of parameters. For example, YOLOV4-Tiny has been proposed based on YOLOV4. The YOLOV4-Tiny structure is a simplified version of YOLOV4 and belongs to a lightweight model. Its parameters are only 6 million, which is equivalent to one-tenth of the original, greatly improving the detection speed.

[0006] These mainstream object detection algorithms often require high-computing power devices, but mobile devices and embedded devices cannot support such complex models in terms of computing power. Currently, there is a lack of an object detector that can better meet the high accuracy rate and high real-time performance of different computing power devices. Summary of the Invention

[0007] Aiming at the defects existing in the prior art, the first aspect of the present invention provides an object detection method based on YOLO-SD-Tiny, which is applicable to general devices, especially low-performance devices with low computing power, and can improve the accuracy rate and detection speed of object detection.

[0008] To achieve the above purposes, the technical solution adopted by the present invention is:

[0009] A target detection method based on YOLO-SD-Tiny, the method comprising the following steps:

[0010] After replacing the activation functions used in the last CSP-Body and CBL in the YOLOV4-Tiny backbone feature extraction network with Mish activation functions, information extraction is performed on the image to be detected to obtain effective feature layers;

[0011] According to the Feature Pyramid Network FPN, perform Self-DeConvolution upsampling on the effective feature layers and output;

[0012] Use YOLO Head to predict the output values after upsampling.

[0013] In some embodiments, the performing Self-DeConvolution upsampling on the effective feature layers according to the Feature Pyramid Network FPN and outputting includes:

[0014] Compress the number of channels of the effective feature layer F with shape H×W×C to C through 1×1 convolution r ;

[0015] Set the upsampling rate to σ, and based on the compressed effective feature layer F, through convolution For the output feature layer A point l in t Predict an upsampling kernel related to the position information Where

[0016] The obtained kernel After reshaping through the weighted summation operator, obtain Where

[0017] Where For the output feature layer The point l in t The feature map of, θ is the weighted summation operator, k area Is a neighborhood of a point in the effective feature layer F, k encoder Is a neighborhood smaller than k area By one area;

[0018] Map the point l in the output feature layer Back to the corresponding point l in the effective feature layer F, and take out the k t Centered on l area ×k area Area, and the predicted upsampling kernel of this point Do a dot product to get the output value.

[0019] In some embodiments, after obtaining the output value, it further includes:

[0020] Through the softmax normalization kernel σH×σW×k area ×k area Make the sum of kernel weights equal to 1;

[0021] In some embodiments, when using the YOLO Head to predict the upsampled output value, the CIOU loss is used as the bounding box regression loss.

[0022] In some embodiments, when using the YOLO Head to predict the upsampled output value, the GHM loss is used as the classification loss.

[0023] The second aspect of the present invention provides an object detection device based on YOLO-SD-Tiny, which is applicable to general devices, especially low-performance devices with low computing power, and can improve the accuracy and detection speed of object detection.

[0024] To achieve the above objectives, the technical solution adopted by the present invention is:

[0025] An object detection device based on YOLO-SD-Tiny, comprising:

[0026] A backbone feature extraction network, which is formed by replacing the activation functions used in the last CSP-Body and CBL in the YOLOV4-Tiny backbone feature extraction network with Mish activation functions, and is used to extract information from the picture to be detected to obtain an effective feature layer;

[0027] A Feature Pyramid Network FPN, which is used to perform Self-DeConvolution upsampling on the effective feature layer and output;

[0028] YOLO Head, which is used to predict the upsampled output value.

[0029] In some embodiments, the Feature Pyramid Network FPN includes a Self-DeConvolution calculation unit, and the Self-DeConvolution calculation unit includes:

[0030] An upsampling kernel prediction module, which is used for:

[0031] Compress the number of channels of the effective feature layer F with the shape of H×W×C to C through a 1×1 convolution r ;

[0032] Set the upsampling rate to σ, and based on the compressed effective feature layer F, through convolution As the output feature layer a point l in t predict an upsampling kernel related to position information wherein

[0033] The obtained kernel is obtained after being reshaped by the weighted summation operator wherein

[0034] wherein is the output feature layer midpoint l t is the feature map of, θ is the weighted summation operator, k area is a neighborhood of a point in the effective feature layer F, k encoder is a neighborhood smaller than k area by one region;

[0035] Feature traversal module, which is used for: mapping the point l in the output feature layer back to the corresponding point l of the effective feature layer F, and taking out the k t centered on l area ×k area region, and the predicted upsampling kernel of this point to perform a dot product to obtain the output value.

[0036] In some embodiments, after obtaining the output value, the upsampling kernel prediction module is further used for:

[0037] Normalize the kernel weights to 1 through softmax for σH×σW×k area ×k area such that the sum of kernel weights is 1;

[0038] In some embodiments, when using YOLO Head to predict the upsampled output value, CIOU loss is used as the bounding box regression loss.

[0039] In some embodiments, when using YOLO Head to predict the upsampled output value, GHM loss is used as the classification loss.

[0040] Compared with the prior art, the advantages of the present invention are:

[0041] In the present invention, aiming at the problems that the target detection model is too large to be deployed on low-performance devices and the real-time performance is poor, the YOLO-SD-Tiny model is proposed. The MCSP-Body based on the Mish activation function is introduced in the backbone feature extraction network part, so that information can flow into the network better; the SD module is introduced in the feature pyramid network part to accelerate the speed of feature fusion and the receptive field. Through the analysis of experimental results, it can be seen that on the OccludeFace dataset, the YOLO-SD-Tiny proposed in the present invention has an AP improvement of 6.35% and a detection speed improvement of 9.64% compared with YOLOv4-Tiny, which solves the problems of detection speed and accuracy to a certain extent. Description of the Drawings

[0042] Figure 1 It is a flowchart of the target detection method based on YOLO-SD-Tiny in an embodiment of the present invention;

[0043] Figure 2 It is a schematic diagram of the overall network model of YOLO-SD-Tiny in an embodiment of the present invention;

[0044] Figure 3 It is a comparison diagram of the curves of the Mish activation function and the LeakyReLU activation function in an embodiment of the present invention;

[0045] Figure 4 It is a flowchart of the upsampling kernel prediction in an embodiment of the present invention;

[0046] Figure 5 It is a flowchart of the feature traversal in an embodiment of the present invention. Detailed Embodiment

[0047] For the purpose of making the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.

[0048] See Figure 1 As shown, an embodiment of the present invention provides a target detection method based on YOLO-SD-Tiny, and the method includes the following steps:

[0049] S1. After replacing the activation functions used in the last CSP-Body and CBL in the YOLOV4-Tiny backbone feature extraction network with the Mish activation function, information extraction is performed on the picture to be detected to obtain an effective feature layer.

[0050] The backbone feature extraction network of YOLOV4-Tiny consists of two CBLs, three CSP-Bodies, and one CBL in sequence.

[0051] It should be noted that CBL represents Convolution, Batch Normalization, and the LeakyReLU activation function. The CSP-Body consists of three CBL structures and one Maxpool. The CSP-Body divides the feature map passed from the previous layer into two parts and then combines them through a cross-stage hierarchical structure. The CSP-Body can enhance the learning ability of the neural network through the residual structure, reducing memory occupancy and computational complexity while ensuring no loss of the neural network's accuracy.

[0052] See Figure 2 As shown, in this embodiment, the LeakyReLU activation function used in the fifth CSP-Body and the sixth CBL is replaced with the Mish activation function. After replacement, they are denoted as MCSP-body and CBM respectively. During sampling, upsampling is performed based on Self-DeConvolution (abbreviated as SD), so the model in this embodiment of the present invention is called YOLO-SD-Tiny.

[0053] When the input is 416×416, the overall network model of YOLO-SD-Tiny is shown in Figure 2 , from Figure 2 it can be seen that YOLO-SD-Tiny is divided into three parts: the backbone feature extraction network, the feature pyramid network, and YOLO Head. The backbone feature extraction network consists of two CBLs, two CSP-Bodies, one MCSP-Body, and CMB. The MCSP-Body replaces the CBL structure in the CSP-Body with a CBM structure based on the Mish activation function, enabling information to flow better into the network.

[0054] The activation function can complete the non-linear transformation from the neuron input to the output, which is of great significance for the training of the neural network. Commonly used activation functions in neural networks include Sigmoid, Tanh, ReLU, LeakyReLU, etc., but they all have certain defects. Taking the ReLU activation function as an example, when the input is negative, the gradient will become zero, resulting in the vanishing gradient; when the LeakyReLU activation function accepts negative inputs, it allows a slight negative gradient, avoiding the influence of the vanishing gradient caused by negative inputs to a certain extent.

[0055] The calculation formula of the Mish activation function is as follows:

[0056] f(x) = xtanh(softplus(x)) = xtanh(ln(1 + e x ))

[0057] The value range of the Mish activation function is [≈ -0.31, +∞). The comparison between the curves of the Mish activation function and the LeakyReLU activation function is shown in Figure 3 . From Figure 3 it can be seen that the Mish activation function allows slight negative values in the value range and no maximum value set, which brings a better gradient flow. After the input of the neural network is mapped by the smooth activation function, the information can penetrate the network better, resulting in better accuracy.

[0058] S2. Upsample the effective feature layers according to the Feature Pyramid Network (FPN) using Self-DeConvolution and output.

[0059] FPN, that is, the Feature Pyramid Network, is a top-down feature fusion method. The Feature Pyramid Network adopts a simple top-down fusion, that is, upsampling the more abstract and semantically stronger high-level feature maps, and connecting the upsampled feature maps to the previous layer of feature maps through horizontal connections. After fusing the high-level features into the shallow features, it can help the shallow features detect targets better. Traditional upsampling is based on the interpolation method, which cannot utilize the semantic information of the feature maps and has a small receptive field.

[0060] Therefore, in this embodiment, upsampling is performed based on Self-DeConvolution (abbreviated as SD). SD involves two modules. The first module is the upsampling kernel prediction module, and the second module is the feature traversal module. For the effective feature layer F with the shape of H×W×C, given an integer upsampling rate of σ, an output feature layer with the shape of σH×σW×C will be obtained after SD For a certain point l in the output feature layer t =(x t , y t ), a corresponding point l=(x, y) can be found in the effective feature layer F, where Denote the neighborhood of l as N(F l ).

[0061] Step S2 includes an upsampling kernel prediction process and a feature traversal process. Specifically, it includes:

[0062] S21. Compress the number of channels of the effective feature layer F with the shape of H×W×C to C through a 1×1 convolution r ;

[0063] S22. Set the upsampling rate to σ, and based on the compressed effective feature layer F, perform convolution to obtain the output feature layer at a point l t to predict an upsampling kernel related to the position information where

[0064] S23. Reshape the obtained kernel through the weighted summation operator to obtain where

[0065] where is the feature map of the point l in the output feature layer, θ is the weighted summation operator, t k k area is a neighborhood of a point in the effective feature layer F, and k encoder is a neighborhood one region smaller than k area ;

[0066] See Figure 4 as shown. It can be understood that in the upsampling kernel prediction module, first compress the number of channels to C through a 1×1 convolution r , and then perform convolution to obtain the output feature layer at a point l t to predict an upsampling kernel related to the position information with parameters k encoder ×k encoder ×C r ×σ 2 ×k area 2 , where k encoder = k area - 1.

[0067] S24. Map the point l in the output feature layer t back to the corresponding point l in the effective feature layer F, and extract the k area ×k area region centered on l, and perform a dot product with the predicted upsampling kernel at this point to obtain the output value.

[0068] It can be understood that the feature traversal process is shown in Figure 5 . For the point l t in the output feature layer (upsampling kernel), map it back to the corresponding point l in the effective feature layer F, and extract the k area ×k areaThe area of ​​​​the point is dot-producted with the predicted upsampling kernel of the point to get the output value. Different channels at the same position share the same upsampling kernel, so that for a point l in the effective feature layer F, the neighborhood graph N(F l ,k area ), in k of l area ×k area Each pixel in the region outputs a feature layer Corresponding pixel l t The contribution is different, based on the content of the feature rather than the distance of the position, so that the semantics of the feature map reorganized by the feature can be stronger than the original feature map, because each pixel can focus on the information from the relevant points in the local area.

[0069] S3. Use YOLO Head to predict the upsampled output value.

[0070] It should be noted that the loss function of YOLO-SD-Tiny in this embodiment is divided into three parts, namely confidence loss, classification loss and bounding box regression loss.

[0071] In a preferred embodiment, the CIOU loss function is used as the bounding box regression loss. IOU refers to the ratio of the intersection and union of the predicted box and the true box. As a measure of the accuracy of bounding box regression, the IOU and CIOU calculation formulas are as follows:

[0072]

[0073] Among them, B is the prediction box, B gt For the real frame.

[0074]

[0075] Among them, b and b gt They represent the center points of the predicted box and the real box respectively, α is the weight function, ν is the parameter used to measure the aspect ratio of the bounding box, and c represents the diagonal distance of the minimum closure area that can contain both the predicted box and the real box.

[0076]

[0077] In terms of classification loss, GHM (gradient harmonizing mechanism) loss is introduced to solve the problem of imbalance between positive and negative samples and the problem of samples that are particularly difficult to distinguish (outliers). The gradient modulus d of outliers is much larger than that of general samples. If the model is forced to pay attention to these samples, it may reduce the accuracy of the model. In order to attenuate both easy-to-distinguish samples and particularly difficult-to-distinguish samples, the gradient density GD(g) is proposed, and the calculation formula is as follows:

[0078]

[0079] Among them, δ ε (g k , g) indicates the number of samples in samples 1 - N where the gradient magnitude distribution is within the range, and l ε (g) represents the length of the interval. Therefore, the physical meaning of the gradient density GD(g) is the total number of samples in the part where the unit gradient magnitude g is located. Next, multiplying the cross - entropy by the reciprocal of the gradient density of this sample can obtain the GHM loss, and the calculation formula is as follows:

[0080]

[0081]

[0082] Among them, N is the total number of samples, and L CE (p i , p i * ) is the binary cross - entropy loss, p ∈ [0, 1] is the probability predicted by the model, and p * ∈ {0, 1} is the true label of a certain class.

[0083] Thus, in the embodiment of the present invention, the loss function of YOLO - SD - Tiny uses the CIOU loss in the bounding box regression loss to accelerate the bounding box regression speed, and uses the GHM loss in the classification loss to solve the problem of unbalanced positive and negative samples and the problem of particularly difficult - to - distinguish samples (outliers).

[0084] The overall process of object detection of YOLO - SD - Tiny in the embodiment of the present invention is as follows:

[0085] First, it is necessary to divide the input image with S×S grids. Each grid in these S×S grids is only responsible for predicting the object whose center point falls within this grid, and calculate 3 prediction boxes. Each prediction box corresponds to 5 + C values; among them, C represents the total number of categories in the dataset, and 5 represents the center coordinates (x, y) of the predicted bounding box, the width and height dimensions (w, h) of the prediction box, and the confidence. Then, solve the class confidence predicted by the network, which is related to the probability P(n object ) that the object falls into the grid, the accuracy P(n class |n object ) of the grid predicting the i - th class object, and the intersection - over - union (IOU). The expression is as follows:

[0086]

[0087] If the object center falls into this grid, then P(n object ) = 1, otherwise it is 0; It is the intersection over union (IoU) between the predicted bounding box and the ground truth bounding box. Finally, DIOU NMS is used to filter out the predicted bounding box with the highest score as the object detection bounding box. The output feature maps are 26×26 and 13×13 respectively, thus realizing the localization and classification of the object. It is worth noting that NMS is an essential post-processing step in object detection, aiming to remove duplicate bounding boxes and leave the most accurate box. DIoU NMS suggests that two bounding boxes with far-apart central points may be located on different objects and should not be deleted (this is the biggest difference between DIoU NMS and NMS).

[0088] It is worth noting that there are various performance evaluation metrics in the current object detection field. For example, the most widely used precision and recall can be adopted to evaluate the model. The calculation formulas are as follows:

[0089]

[0090]

[0091] Among them, precision P is used to evaluate the prediction results. In the formula, TP (True Positive) represents the number of positive samples correctly predicted as positive samples by the model, and FP (False Positive) represents the number of negative samples predicted as positive samples. Recall is used to evaluate the samples, indicating how many positive samples in all samples are correctly predicted. FN (False Negative) represents that the model predicts the input that was originally a positive sample as a negative sample.

[0092] AP represents the area enclosed by the PR curve formed by precision and recall at different confidence thresholds for a single category and the coordinate axes, comprehensively considering precision and recall, and providing a more comprehensive evaluation of the recognition effect of single-class object detection. FPS is the number of images that the model can detect in one second. The larger the FPS value, the faster the detection speed of the model.

[0093] The YOLO-SD-Tiny algorithm in the embodiments of the present invention is compared with YOLOv4-Tiny on the OccludeFace dataset. At the same time, ablation experiments are conducted on YOLO-SD-Tiny (with MCSP-Body) and YOLO-SD-Tiny (with GHM&CIOU) to verify the influence of different modules on the model. The experimental results are shown in Table 1. As can be seen from Table 1, introducing MCSP-Body based on the Mish activation function improves the AP by 0.67% compared with YOLOv4-Tiny, indicating that the gradient does not disappear and the smooth activation function can make information penetrate deeper into the network, thereby improving the detection accuracy. YOLO-SD-Tiny, which introduces the GHM loss in the classification loss part and the CIOU loss in the bounding box regression part, improves the AP by 2.09% compared with the original YOLOv4-Tiny model, indicating that CIOU, which comprehensively considers the overlapping area, center point, and aspect ratio, and the GHM loss that solves the imbalance between positive and negative samples and difficult-to-separate samples can increase the detection accuracy of the model. YOLO-SD-Tiny improves the AP by 6.35% compared with YOLOv4-Tiny and improves the detection speed FPS by 9.64%. Through the comparison of various experimental data in Table 1, it can be verified that the improved method proposed in the present invention can effectively improve the target detection accuracy and detection speed.

[0094] Table 1

[0095]

[0096] In summary, for the problems of the target detection model being too large to be deployed on low-performance devices and poor real-time performance in the present invention, the YOLO-SD-Tiny model is proposed. The MCSP-Body based on the Mish activation function is introduced in the backbone feature extraction network part to allow information to flow better into the network; the SD module is introduced in the feature pyramid network part to accelerate the feature fusion speed and receptive field. Through the analysis of the experimental results, it can be seen that the YOLO-SD-Tiny proposed in the present invention improves the AP by 6.35% and the detection speed by 9.64% compared with YOLOv4-Tiny on the OccludeFace dataset, and solves the problems of detection speed and accuracy to a certain extent.

[0097] At the same time, the embodiments of the present invention also provide a target detection device based on YOLO-SD-Tiny, which includes a backbone feature extraction network, a feature pyramid network FPN, and a YOLO Head.

[0098] The backbone feature extraction network is formed by replacing the activation functions used in the last CSP-Body and CBL in the YOLOV4-Tiny backbone feature extraction network with the Mish activation function, and is used to extract information from the image to be detected to obtain effective feature layers.

[0099] The Feature Pyramid Network (FPN) is used to perform Self-DeConvolution upsampling on the effective feature layers and output. The YOLO Head is used to predict the output values after upsampling.

[0100] In some embodiments, the Feature Pyramid Network (FPN) includes a Self-DeConvolution calculation unit, and the Self-DeConvolution calculation unit includes an upsampling kernel prediction module and a feature traversal module.

[0101] The upsampling kernel prediction module is used for:

[0102] Compressing the number of channels of the effective feature layer F with a shape of H×W×C to C through a 1×1 convolution r ;

[0103] Setting the upsampling rate to σ, and predicting an upsampling kernel related to the position information for a point l in the output feature layer through convolution based on the compressed effective feature layer F where t the obtained kernel is reshaped through a weighted summation operator to obtain where

[0104] is the feature map of the point l in the output feature layer where θ is the weighted summation operator, where

[0105] where is the output feature layer at the midpoint l t k is a neighborhood of a point in the effective feature layer F, and k is a neighborhood smaller than k by one region; area encoder area t area area

[0106] The feature traversal module is used for: mapping the point l in the output feature layer back to the corresponding point l in the effective feature layer F, and taking out the k×k t region centered on l, and performing a dot product with the predicted upsampling kernel at this point to obtain the output value. area ×k area

[0107] ​In some embodiments, after obtaining the output value, the upsampling kernel prediction module is further configured to:

[0108] Through the softmax normalization kernel σH×σW×k area ×k area Make the sum of kernel weights equal to 1;

[0109] In some embodiments, when using YOLO Head to predict the upsampled output value, the CIOU loss is used as the bounding box regression loss.

[0110] In some embodiments, when using YOLO Head to predict the upsampled output value, the GHM loss is used as the classification loss.

[0111] In the face detection device based on YOLOV4-Tiny of the present invention, the MCSP-Body based on the Mish activation function is introduced in the backbone feature extraction network part, so that information can flow into the network better; the SD module is introduced in the feature pyramid network part to accelerate the speed of feature fusion and the receptive field. Through the analysis of experimental results, it can be seen that on the OccludeFace dataset, YOLO-SD-Tiny proposed by the present invention has an AP improvement of 6.35% compared with YOLOv4-Tiny, and the detection speed is increased by 9.64%, which solves the problems of detection speed and accuracy to a certain extent.

[0112] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including the element.

[0113] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments. The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A target detection method based on YOLO-SD-Tiny, characterized in that, The method includes the following steps: After replacing the activation functions used in the last CSP-Body and CBL in the YOLOV4-Tiny backbone feature extraction network with Mish activation functions, information extraction is performed on the image to be detected to obtain effective feature layers; Based on the Feature Pyramid Network (FPN), Self-DeConvolution upsampling is performed on the effective feature layers and output; YOLO Head is used to predict the output values after upsampling.

2. The object detection method based on YOLO-SD-Tiny according to claim 1, characterized in that, The step of performing Self-DeConvolution upsampling on the effective feature layers based on the Feature Pyramid Network (FPN) and outputting includes: Compress the number of channels of the effective feature layer F with shape H×W×C to C through 1×1 convolution r , where C r is the number of channels after compression of the effective feature layer F; Set the upsampling rate to σ, and through convolution based on the compressed effective feature layer F as the output feature layer at a point l in t predict an upsampling kernel related to the position information where N represents the neighborhood; The obtained core is obtained after reshaping by the weighted summation operator where Among them is the output feature layer The midpoint l t is the feature map, θ is the weighted summation operator, k area is a neighborhood of a point in the effective feature layer F, k encoder is a neighborhood of a region smaller than k area by one region; Map the point l in the output feature layer back to the point l corresponding to the valid feature layer F, and extract a k×k t region centered at l, and perform a dot product with the upsampling kernel of the predicted point l area ×k area to obtain the output value. t ​​ 3. The object detection method based on YOLO-SD-Tiny according to claim 2, wherein, After obtaining the output values, it further includes: Through the softmax normalization kernel σH×σW×k area ×k area Make the sum of kernel weights equal to 1.

4. The object detection method based on YOLO-SD-Tiny according to claim 1, characterized in that, When using YOLO Head to predict the output values after upsampling, CIOU loss is used as the bounding box regression loss.

5. The object detection method based on YOLO-SD-Tiny according to claim 1, wherein, When using YOLO Head to predict the output values after upsampling, GHM loss is used as the classification loss.

6. An object detection device based on YOLO-SD-Tiny, characterized in that, It includes: A backbone feature extraction network, which is formed by replacing the activation functions used in the last CSP-Body and CBL in the YOLOV4-Tiny backbone feature extraction network with Mish activation functions, and is used to perform information extraction on the image to be detected to obtain effective feature layers; The Feature Pyramid Network (FPN), which is used to perform Self-DeConvolution upsampling on the effective feature layers and output; YOLO Head, which is used to predict the output values after upsampling.

7. The object detection device based on YOLO-SD-Tiny according to claim 6, characterized in that The Feature Pyramid Network (FPN) includes a Self-DeConvolution calculation unit, and the Self-DeConvolution calculation unit includes: An upsampling kernel prediction module, which is used for: Compress the number of channels of the effective feature layer F with shape H×W×C to C through 1×1 convolution r , where C r is the number of channels after compression of the effective feature layer F; Set the upsampling rate to σ, and based on the compressed effective feature layer F, through convolution to obtain the output feature layer at a point l t predict an upsampling kernel related to the position information where N represents the neighborhood; The obtained core is obtained after reshaping by the weighted summation operator wherein Among them is the output feature layer midpoint l t 's feature map, θ is the weighted summation operator, k area is a neighborhood of a point in the effective feature layer F, k encoder is a neighborhood that is one region smaller than k area ; The feature traversal module is used to: output feature layer Point l t Map back to the point l corresponding to the effective feature layer F, and take out the k points centered on l area ×k area area, and the predicted point l t The upsampling kernel ω lt Do the dot product to get the output value.

8. The object detection device based on YOLO-SD-Tiny according to claim 7, wherein, After obtaining the output values, the upsampling kernel prediction module is further used for: Through the softmax normalization kernel σH×σW×k area ×k area The sum of the kernel weights is made to be 1.

9. The object detection device based on YOLO-SD-Tiny according to claim 6, wherein When using YOLO Head to predict the output values after upsampling, CIOU loss is used as the bounding box regression loss.

10. The object detection device based on YOLO-SD-Tiny according to claim 6, wherein, When using YOLO Head to predict the output values after upsampling, GHM loss is used as the classification loss.