Target detection method and device

By building a large-core attention network and a progressive feature pyramid network in the YOLO model and connecting the auxiliary training head network in parallel, the problem of low accuracy in detecting high-speed moving small targets in the prior art is solved, and higher detection accuracy and recognition capabilities are achieved.

CN120182572APending Publication Date: 2025-06-20SHENZHEN XUMI YUNTU SPACE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510188555.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the prior art, the accuracy of detecting small targets that move at high speed is low, making it difficult to effectively identify and locate these targets.

Method used

By building a large-nuclear attention network, combining deep separation convolution layer, deep expansion convolution layer and channel convolution layer, and adding a large-nuclear attention network to the backbone network of the YOLO model, a progressive feature pyramid network is used to replace the neck network of the YOLO model, and a parallel connection of the auxiliary training head network is carried out to form a new target detection model.

Benefits of technology

The accuracy of detecting small targets with high speed moving is improved, and the model's ability to identify and locate high speed moving targets is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182572A_ABST
    Figure CN120182572A_ABST
Patent Text Reader

Abstract

The invention provides a target detection method and device. The method comprises the following steps: adding a large kernel attention network in a fast spatial pyramid pooling network in a backbone network of a YOLO model, replacing a neck network of the YOLO model with a progressive feature pyramid network, connecting an auxiliary training head network in parallel at the progressive feature pyramid network in the YOLO model, and taking an obtained new model as a target detection model; detecting whether a target object exists in the training image by using a target detection model to obtain a first detection result output by an original head network of the YOLO model and a second detection result output by an auxiliary training head network; calculating a first loss between the first detection result and the label of the training image, and calculating a second loss between the second detection result and the label of the training image; and optimizing model parameters of the target detection model according to the first loss and the second loss. By adopting the technical means, the problem of low precision of detecting a small target moving at a high speed in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of object detection, and in particular, to an object detection method and device. Background Art

[0002] Due to the redundant characteristics of the convolutional network in traditional object detection algorithms, simply increasing the number of parameters cannot achieve good accuracy, especially for detecting small targets moving at high speeds. How to achieve high-precision detection of small targets moving at high speeds is a major technical pain point at present. Summary of the Invention

[0003] In view of this, embodiments of the present application provide an object detection method, device, electronic device, and computer-readable storage medium to solve the problem of low accuracy in detecting small targets moving at high speeds in the prior art.

[0004] In a first aspect of the embodiments of the present application, an object detection method is provided, including: constructing a large kernel attention network using a depthwise separable convolutional layer, a depthwise dilated convolutional layer, and a channel convolutional layer, and constructing an auxiliary training head network using a convolutional layer; adding the large kernel attention network to the fast spatial pyramid pooling network in the backbone network of the YOLO model, replacing the neck network of the YOLO model with a progressive feature pyramid network, and connecting the auxiliary training head network in parallel at the progressive feature pyramid network of the YOLO model, and using the obtained new model as the object detection model; obtaining a training image, using the object detection model to detect whether there is an object target in the training image, and obtaining a first detection result output by the original head network of the YOLO model and a second detection result output by the auxiliary training head network; calculating a first loss between the first detection result and the label of the training image, and calculating a second loss between the second detection result and the label of the training image; optimizing the model parameters of the object detection model according to the first loss and the second loss to complete the training of the object detection model.

[0005] In a second aspect of the embodiments of the present application, a target detection device is provided, including: a construction module configured to construct a large kernel attention network using a depthwise separable convolutional layer, a depthwise dilated convolutional layer, and a channel convolutional layer, and construct an auxiliary training head network using a convolutional layer; a model improvement module configured to add the large kernel attention network to a fast spatial pyramid pooling network in a backbone network of a YOLO model, replace the neck network of the YOLO model with a progressive feature pyramid network, connect the auxiliary training head network in parallel at the progressive feature pyramid network in the YOLO model, and use the obtained new model as a target detection model; an acquisition module configured to acquire a training image, and use the target detection model to detect whether there is a target object in the training image, obtaining a first detection result output by the original head network of the YOLO model and a second detection result output by the auxiliary training head network; a calculation module configured to calculate a first loss between the first detection result and the label of the training image, and calculate a second loss between the second detection result and the label of the training image; and an optimization module configured to optimize the model parameters of the target detection model according to the first loss and the second loss to complete the training of the target detection model.

[0006] In a third aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor implements the steps of the above method when executing the computer program.

[0007] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, where the computer-readable storage medium stores a computer program, and the computer program implements the steps of the above method when executed by a processor.

[0008] The beneficial effects of the embodiments of this application compared with the prior art are as follows: A large kernel attention network is constructed using depthwise separable convolutional layers, depthwise dilated convolutional layers, and channel convolutional layers, and an auxiliary training head network is constructed using convolutional layers; the large kernel attention network is added to the fast spatial pyramid pooling network in the backbone network of the YOLO model, the progressive feature pyramid network is used to replace the neck network of the YOLO model, the auxiliary training head network is connected in parallel at the progressive feature pyramid network in the YOLO model, and the new model obtained is used as the object detection model; training images are obtained, and the object detection model is used to detect whether there are target objects in the training images, obtaining the first detection result output by the original head network of the YOLO model and the second detection result output by the auxiliary training head network; the first loss between the first detection result and the label of the training image is calculated, and the second loss between the second detection result and the label of the training image is calculated; the model parameters of the object detection model are optimized based on the first loss and the second loss to complete the training of the object detection model. By adopting the above technical means, the problem of low accuracy in detecting small targets moving at high speed in the prior art can be solved, and thus the accuracy of detecting small targets moving at high speed can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of this application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0010] Figure 1 is a flowchart of an object detection method provided by an embodiment of this application;

[0011] Figure 2 is a flowchart of another object detection method provided by an embodiment of this application;

[0012] Figure 3 is a structural schematic diagram of an object detection device provided by an embodiment of this application;

[0013] Figure 4 is a structural schematic diagram of an electronic device provided by an embodiment of this application. DETAILED DESCRIPTION

[0014] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.

[0015] A target detection method and device according to an embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0016] Figure 1 It is a schematic flowchart of a target detection method provided by an embodiment of the present application. Figure 1 The target detection method can be executed by a computer or a server, or software on a computer or a server. As Figure 1 shown, the target detection method includes:

[0017] S101, constructing a large kernel attention network using a depthwise separable convolutional layer, a depthwise dilated convolutional layer, and a channel convolutional layer, and constructing an auxiliary training head network using a convolutional layer;

[0018] S102, adding a large kernel attention network to the fast spatial pyramid pooling network in the backbone network of the YOLO model, replacing the neck network of the YOLO model with a progressive feature pyramid network, connecting the auxiliary training head network in parallel at the input side of the progressive feature pyramid network in the YOLO model, and using the obtained new model as the target detection model;

[0019] S103, obtaining a training image, using the target detection model to detect whether there is a target object in the training image, and obtaining a first detection result output by the original head network of the YOLO model and a second detection result output by the auxiliary training head network;

[0020] S104, calculating a first loss between the first detection result and the label of the training image, and calculating a second loss between the second detection result and the label of the training image;

[0021] S105, optimizing the model parameters of the target detection model according to the first loss and the second loss to complete the training of the target detection model.

[0022] The YOLO (You Only Look Once) model is a series of neural network architectures for real-time target detection. Embodiments of the present application improve the YOLO model (such as the yolov8s version) to adapt to the target detection scenario, especially for small targets moving at high speed, and improve the accuracy of target detection. The YOLO model can be divided into three parts: a backbone network, a neck network, and a head network. The backbone network contains a fast spatial pyramid pooling network inside. In the embodiments of the present application, a large kernel attention network is added to the fast spatial pyramid pooling network, the neck network of the YOLO model is replaced with a progressive feature pyramid network, and an auxiliary training head network is connected in parallel at the input side of the progressive feature pyramid network in the YOLO model (the progressive feature pyramid network and the auxiliary training head network are in a parallel relationship), and the obtained new model is used as the target detection model.

[0023] The backbone network is generally a residual network, which contains multiple stage networks and outputs multiple stage features. The original neck network of the YOLO model processes the features of three stages, namely C3, C4, and C5, while the Progressive Feature Pyramid Network processes the features of four stages, namely C2, C3, C4, and C5.

[0024] Use the training images to train the object detection model. Calculate the first loss between the first detection result output by the original head network of the YOLO model and the label of the training image, and calculate the second loss between the second detection result output by the auxiliary training head network and the label of the training image. Optimize the model parameters of the object detection model according to the first loss and the second loss respectively.

[0025] According to the technical solution provided by the embodiments of the present application, a large kernel attention network is constructed using a depthwise separable convolutional layer, a depthwise dilated convolutional layer, and a channel convolutional layer, and an auxiliary training head network is constructed using a convolutional layer; a large kernel attention network is added to the Fast Spatial Pyramid Pooling Network in the backbone network of the YOLO model, and the Progressive Feature Pyramid Network is used to replace the neck network of the YOLO model. An auxiliary training head network is connected in parallel at the Progressive Feature Pyramid Network in the YOLO model, and the new model obtained is used as the object detection model; obtain the training images, use the object detection model to detect whether there are target objects in the training images, and obtain the first detection result output by the original head network of the YOLO model and the second detection result output by the auxiliary training head network; calculate the first loss between the first detection result and the label of the training image, and calculate the second loss between the second detection result and the label of the training image; optimize the model parameters of the object detection model according to the first loss and the second loss to complete the training of the object detection model. By adopting the above technical means, the problem of low accuracy in detecting small targets moving at high speed in the prior art can be solved, and thus the accuracy of detecting small targets moving at high speed can be improved.

[0026] Further, adding a large kernel attention network to the Fast Spatial Pyramid Pooling Network in the backbone network of the YOLO model includes: adding a large kernel attention network before the last convolutional layer in the Fast Spatial Pyramid Pooling Network.

[0027] After adding a large kernel attention network to the Fast Spatial Pyramid Pooling Network, the Fast Spatial Pyramid Pooling Network internally includes, in sequence: a convolutional layer, a normalization layer, an activation layer, three max pooling layers, a concatenation layer, a large kernel attention network, a convolutional layer, a normalization layer, and an activation layer.

[0028] The convolutional layer in this embodiment is a common convolutional layer, which can generally be a channel convolutional layer.

[0029] Furthermore, the target detection model is used to detect whether the target object exists in the training image, and a first detection result output by the original head network of the YOLO model and a second detection result output by the auxiliary training head network are obtained, including: inputting the training image into the target detection model, and within the target detection model: processing the training image through the backbone network to obtain image features; processing the image features through the progressive feature pyramid network to obtain progressive features; processing the progressive features through the original head network of the YOLO model to obtain the first detection result; processing the image features through the auxiliary training head network to obtain the second detection result.

[0030] The target detection model includes: backbone network, progressive feature pyramid network, original head network of YOLO model and auxiliary training head network. The training image is processed by backbone network and progressive feature pyramid network in turn to obtain progressive features. Finally, the original head network of YOLO model processes progressive features, and the auxiliary training head network processes image features to obtain the first detection result and the second detection result.

[0031] Furthermore, the input of the large core attention network is recorded as the first feature. Inside the large core attention network: the first feature is processed by a depthwise separable convolution layer to obtain the second feature; the second feature is processed by a depthwise dilated convolution layer to obtain the third feature; the third feature is processed by a channel convolution layer to obtain the fourth feature; the first feature and the fourth feature are subjected to matrix multiplication to obtain the fifth feature, wherein the fifth feature is the output of the large core attention network.

[0032] The auxiliary training head network includes a depthwise separable convolution layer, a depthwise dilated convolution layer, a channel convolution layer, and a network layer for matrix multiplication operations.

[0033] Depthwise Separable Convolution is an optimized convolution operation that is widely used in modern neural network architectures, especially lightweight models in mobile devices and embedded systems. It significantly reduces the amount of computation and the number of parameters while maintaining high performance by decomposing the standard convolution into two simpler steps - depthwise convolution and pointwise convolution.

[0034] Atrous Convolution (also called Dilated Convolution or Dilated Convolution) is a special convolution operation that expands the receptive field by inserting holes between the convolution kernel elements (i.e. skipping some input pixels) without increasing the number of parameters or the amount of calculation. This method is particularly suitable for tasks that need to capture a wider range of contextual information, such as semantic segmentation, target detection, etc.

[0035] The Channel-wise Convolution Layer performs convolution operations on each channel of the input data separately, rather than jointly on all channels. This type of convolution layer is typically used to preserve the unique features of each channel and allow the network to learn different representations between channels.

[0036] Furthermore, the input to the Fast Spatial Pyramid Pooling network is denoted as the sixth feature. Inside the Fast Spatial Pyramid Pooling network: the sixth feature is processed sequentially through a convolutional layer, a normalization layer, and an activation layer to obtain the seventh feature; the seventh feature is processed sequentially through three max-pooling layers to obtain the eighth feature, the ninth feature, and the tenth feature respectively; the eighth feature, the ninth feature, and the tenth feature are concatenated together through a concatenation layer to obtain the first feature; the first feature is processed through a large kernel attention network to obtain the fifth feature; the fifth feature is processed sequentially through a convolutional layer, a normalization layer, and an activation layer to obtain the eleventh feature, where the eleventh feature is the output of the Fast Spatial Pyramid Pooling network.

[0037] The sixth feature passes through a convolutional layer, a normalization layer, and an activation layer in sequence to obtain the seventh feature. Processing the seventh feature through three max-pooling layers in sequence includes: processing the seventh feature through the first max-pooling layer to obtain the eighth feature; processing the eighth feature through the second max-pooling layer to obtain the ninth feature; processing the ninth feature through the third max-pooling layer to obtain the tenth feature. The concatenation layer concatenates the eighth feature, the ninth feature, and the tenth feature together to obtain the first feature. The large kernel attention network processes the first feature to obtain the fifth feature. The fifth feature passes through a convolutional layer, a normalization layer, and an activation layer in sequence to obtain the eleventh feature.

[0038] The Fast Spatial Pyramid Pooling network SPPF (Spatial Pyramid Pooling–Fast) is an improved Spatial Pyramid Pooling (SPP) technique, initially proposed in Fast YOLO (a variant of YOLOv3). The SPPF module aims to enhance the model's detection ability for targets of different sizes by introducing multi-scale feature fusion, while maintaining high computational efficiency. It significantly improves the model's expressiveness without adding too many additional parameters.

[0039] Further, after optimizing the model parameters of the object detection model according to the first loss and the second loss to complete the training of the object detection model, the method further includes: pruning the auxiliary training head network in the object detection model; obtaining the image to be detected, and using the object detection model to detect whether there is an object of interest in the image to be detected, and obtaining the detection result output by the original head network of the YOLO model.

[0040] The auxiliary training head network is a lightweight network. Therefore, when using the auxiliary training head network to accelerate the training of the object detection model, after completing the training of the object detection model, the auxiliary training head network is pruned.

[0041] Figure 2 It is a schematic flowchart of another object detection method provided by an embodiment of the present application. As Figure 2 shown, the method includes:

[0042] S201, input the training image into the object detection model. Inside the object detection model:

[0043] S202, process the training image through the backbone network to obtain image features;

[0044] S203, process the image features through the progressive feature pyramid network to obtain progressive features;

[0045] S204, process the progressive features through the original head network of the YOLO model to obtain the first detection result;

[0046] S205, process the image features through the auxiliary training head network to obtain the second detection result.

[0047] All the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated here one by one.

[0048] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the present application. For the details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the present application.

[0049] Figure 3 It is a schematic diagram of an object detection device provided by an embodiment of the present application. As Figure 3 shown, the object detection device includes:

[0050] A construction module 301, configured to construct a large kernel attention network by using a depthwise separable convolutional layer, a depthwise dilated convolutional layer, and a channel convolutional layer, and construct an auxiliary training head network by using a convolutional layer;

[0051] A model improvement module 302, configured to add a large kernel attention network to the fast spatial pyramid pooling network in the backbone network of the YOLO model, use the progressive feature pyramid network to replace the neck network of the YOLO model, connect the auxiliary training head network in parallel at the progressive feature pyramid network of the YOLO model, and use the obtained new model as the object detection model;

[0052] An acquisition module 303 is configured to acquire training images, detect whether there are target objects in the training images by using a target detection model, and obtain a first detection result output by the original head network of the YOLO model and a second detection result output by the auxiliary training head network;

[0053] A calculation module 304 is configured to calculate a first loss between the first detection result and the label of the training image, and calculate a second loss between the second detection result and the label of the training image;

[0054] An optimization module 305 is configured to optimize the model parameters of the target detection model according to the first loss and the second loss to complete the training of the target detection model.

[0055] The YOLO (You Only Look Once) model is a series of neural network architectures for real-time target detection. The embodiments of this application improve the YOLO model to adapt to the target detection scenario, especially for small targets moving at high speed (such as motorcycles), and improve the accuracy of target detection. The YOLO model can be divided into three parts: a backbone network, a neck network, and a head network. The backbone network internally contains a fast spatial pyramid pooling network. In the embodiments of this application, a large kernel attention network is added to the fast spatial pyramid pooling network, and a progressive feature pyramid network is used to replace the neck network of the YOLO model. An auxiliary training head network is connected in parallel at the input side of the progressive feature pyramid network in the YOLO model (the progressive feature pyramid network and the auxiliary training head network are in a parallel relationship), and the obtained new model is used as the target detection model.

[0056] Use the training images to train the target detection model. For the first detection result output by the original head network of the YOLO model and the second detection result output by the auxiliary training head network, use the cross-entropy loss function to calculate the first loss between the first detection result and the label of the training image, and calculate the second loss between the second detection result and the label of the training image. Optimize the model parameters of the target detection model according to the first loss and the second loss respectively.

[0057] According to the technical solution provided by the embodiments of the present application, a large kernel attention network is constructed using a depthwise separable convolutional layer, a depthwise dilated convolutional layer, and a channel convolutional layer, and an auxiliary training head network is constructed using a convolutional layer; a large kernel attention network is added to the fast spatial pyramid pooling network in the backbone network of the YOLO model, and the progressive feature pyramid network is used to replace the neck network of the YOLO model. The auxiliary training head network is connected in parallel at the progressive feature pyramid network in the YOLO model, and the obtained new model is used as the object detection model; training images are obtained, and the object detection model is used to detect whether there are target objects in the training images, and the first detection result output by the original head network of the YOLO model and the second detection result output by the auxiliary training head network are obtained; the first loss between the first detection result and the label of the training image is calculated, and the second loss between the second detection result and the label of the training image is calculated; the model parameters of the object detection model are optimized according to the first loss and the second loss to complete the training of the object detection model. By adopting the above technical means, the problem of low accuracy in detecting small targets moving at high speed in the prior art can be solved, and thus the accuracy of detecting small targets moving at high speed can be improved.

[0058] In some embodiments, the model improvement module 302 is further configured to add a large kernel attention network before the last convolutional layer in the fast spatial pyramid pooling network.

[0059] After adding a large kernel attention network to the fast spatial pyramid pooling network, the fast spatial pyramid pooling network internally includes, in sequence: a convolutional layer, a normalization layer, an activation layer, three max pooling layers, a concatenation layer, a large kernel attention network, a convolutional layer, a normalization layer, and an activation layer.

[0060] The convolutional layer in this embodiment is a common convolutional layer, and generally can be a channel convolutional layer.

[0061] In some embodiments, the acquisition module 303 is further configured to input the training image into the object detection model. Inside the object detection model: the training image is processed by the backbone network to obtain image features; the image features are processed by the progressive feature pyramid network to obtain progressive features; the progressive features are processed by the original head network of the YOLO model to obtain the first detection result; the image features are processed by the auxiliary training head network to obtain the second detection result.

[0062] The object detection model internally includes: a backbone network, a progressive feature pyramid network, the original head network of the YOLO model, and an auxiliary training head network. The training image is processed by the backbone network and the progressive feature pyramid network in sequence to obtain progressive features. Finally, the original head network of the YOLO model processes the progressive features, and the auxiliary training head network processes the image features to obtain the first detection result and the second detection result.

[0063] In some embodiments, the acquisition module 303 is also configured to record the input of the large core attention network as the first feature, and inside the large core attention network: process the first feature through a depthwise separable convolution layer to obtain a second feature; process the second feature through a depthwise dilated convolution layer to obtain a third feature; process the third feature through a channel convolution layer to obtain a fourth feature; perform matrix multiplication on the first feature and the fourth feature to obtain a fifth feature, wherein the fifth feature is the output of the large core attention network.

[0064] The auxiliary training head network includes a depthwise separable convolution layer, a depthwise dilated convolution layer, a channel convolution layer, and a network layer for matrix multiplication operations.

[0065] Depthwise Separable Convolution is an optimized convolution operation that is widely used in modern neural network architectures, especially lightweight models in mobile devices and embedded systems. It significantly reduces the amount of computation and the number of parameters while maintaining high performance by decomposing the standard convolution into two simpler steps - depthwise convolution and pointwise convolution.

[0066] Atrous Convolution (also called Dilated Convolution or Dilated Convolution) is a special convolution operation that expands the receptive field by inserting holes between the convolution kernel elements (i.e. skipping some input pixels) without increasing the number of parameters or the amount of calculation. This method is particularly suitable for tasks that need to capture a wider range of contextual information, such as semantic segmentation, target detection, etc.

[0067] The channel-wise convolution layer performs convolution operations on each channel of the input data separately, rather than jointly operating on all channels. This type of convolution layer is often used to maintain the unique characteristics of each channel and allow the network to learn different representations between channels.

[0068] In some embodiments, the obtaining module 303 is further configured to record the input of the fast spatial pyramid pooling network as the sixth feature. Inside the fast spatial pyramid pooling network: the sixth feature is processed sequentially through a convolutional layer, a normalization layer, and an activation layer to obtain the seventh feature; the seventh feature is processed sequentially through three max-pooling layers to obtain the eighth feature, the ninth feature, and the tenth feature respectively; the eighth feature, the ninth feature, and the tenth feature are concatenated together through a concatenation layer to obtain the first feature; the first feature is processed through a large kernel attention network to obtain the fifth feature; the fifth feature is processed sequentially through a convolutional layer, a normalization layer, and an activation layer to obtain the eleventh feature, where the eleventh feature is the output of the fast spatial pyramid pooling network.

[0069] The sixth feature passes through a convolutional layer, a normalization layer, and an activation layer in sequence to obtain the seventh feature. Processing the seventh feature sequentially through three max-pooling layers includes: processing the seventh feature through the first max-pooling layer to obtain the eighth feature; processing the eighth feature through the second max-pooling layer to obtain the ninth feature; processing the ninth feature through the third max-pooling layer to obtain the tenth feature. The concatenation layer concatenates the eighth feature, the ninth feature, and the tenth feature together to obtain the first feature. The large kernel attention network processes the first feature to obtain the fifth feature. The fifth feature passes through a convolutional layer, a normalization layer, and an activation layer in sequence to obtain the eleventh feature.

[0070] The fast spatial pyramid pooling network SPPF is an improved spatial pyramid pooling (SPP) technique, which was initially proposed in Fast YOLO (a variant of YOLOv3). The SPPF module aims to enhance the model's detection ability for targets of different sizes by introducing multi-scale feature fusion while maintaining high computational efficiency. Without adding too many additional parameters, it significantly improves the model's expressiveness.

[0071] In some embodiments, the optimization module 305 is further configured to prune the auxiliary training head network in the object detection model; obtain the image to be detected, and use the object detection model to detect whether there is an object of interest in the image to be detected, so as to obtain the detection result output by the original head network of the YOLO model.

[0072] The auxiliary training head network is a lightweight network. Therefore, the auxiliary training head network is used to accelerate the training of the object detection model, and after the training of the object detection model is completed, the auxiliary training head network is pruned.

[0073] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not indicate the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0074] Figure 4 is a schematic diagram of the electronic device 4 provided by an embodiment of the present application. As Figure 4 shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of each module / unit in the above-mentioned device embodiments are implemented.

[0075] The electronic device 4 may be a desktop computer, a notebook, a palm computer, a cloud server, or other electronic devices. The electronic device 4 may include, but is not limited to, the processor 401 and the memory 402. Those skilled in the art can understand that Figure 4 merely examples of the electronic device 4, and do not constitute a limitation on the electronic device 4, and may include more or fewer components than those shown in the figure, or different components.

[0076] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0077] The memory 402 may be an internal storage unit of the electronic device 4, for example, the hard disk or memory of the electronic device 4. The memory 402 may also be an external storage device of the electronic device 4, for example, a plug-in hard disk equipped on the electronic device 4, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. The memory 402 may also include both the internal storage unit and the external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device.

[0078] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0079] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0080] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A target detection method, characterized in that: include: Use depthwise separable convolutional layers, depthwise dilated convolutional layers, and channel convolutional layers to build a large-core attention network, and use convolutional layers to build an auxiliary training head network; A large core attention network is added to the fast spatial pyramid pooling network in the backbone network of the YOLO model, a progressive feature pyramid network is used to replace the neck network of the YOLO model, an auxiliary training head network is connected in parallel to the progressive feature pyramid network in the YOLO model, and the obtained new model is used as the target detection model; Acquire a training image, use the target detection model to detect whether there is a target object in the training image, and obtain a first detection result output by the original head network of the YOLO model and a second detection result output by the auxiliary training head network; Calculating a first loss between the first detection result and the label of the training image, and calculating a second loss between the second detection result and the label of the training image; The model parameters of the target detection model are optimized according to the first loss and the second loss to complete the training of the target detection model.

2. The method according to claim 1, characterized in that Add a large core attention network to the fast spatial pyramid pooling network in the backbone network of the YOLO model, including: The large core attention network is added before the last convolutional layer in the fast spatial pyramid pooling network.

3. The method according to claim 1, characterized in that Using the target detection model to detect whether there is a target object in the training image, obtaining a first detection result output by the original head network of the YOLO model and a second detection result output by the auxiliary training head network, including: The training image is input into the object detection model, and inside the object detection model: Processing the training image through a backbone network to obtain image features; Processing the image features through a progressive feature pyramid network to obtain progressive features; Processing the progressive features through the original head network of the YOLO model to obtain a first detection result; The image features are processed by the auxiliary training head network to obtain a second detection result.

4. The method according to claim 1, characterized in that: The input of the large core attention network is recorded as the first feature. Inside the large core attention network: Processing the first feature through the depthwise separable convolutional layer to obtain a second feature; Processing the second feature through the deep dilated convolutional layer to obtain a third feature; Processing the third feature through the channel convolution layer to obtain a fourth feature; Perform a matrix multiplication operation on the first feature and the fourth feature to obtain a fifth feature, wherein the fifth feature is the output of the large core attention network.

5. The method according to claim 1, characterized in that After the large core attention network is added to the fast spatial pyramid pooling network, the fast spatial pyramid pooling network includes: a convolution layer, a normalization layer, an activation layer, three maximum pooling layers, a splicing layer, a large core attention network, a convolution layer, a normalization layer and an activation layer.

6. The method according to claim 5, characterized in that The input of the fast spatial pyramid pooling network is recorded as the sixth feature. Inside the fast spatial pyramid pooling network: The sixth feature is processed by the convolution layer, the normalization layer and the activation layer in sequence to obtain the seventh feature; Processing the seventh feature through three maximum pooling layers in sequence to obtain an eighth feature, a ninth feature, and a tenth feature respectively; The eighth feature, the ninth feature and the tenth feature are spliced ​​together through a splicing layer to obtain a first feature; Process the first feature through a large core attention network to obtain the fifth feature; The fifth feature is processed sequentially through a convolution layer, a normalization layer, and an activation layer to obtain an eleventh feature, wherein the eleventh feature is an output of the fast spatial pyramid pooling network.

7. The method according to claim 1, characterized in that After optimizing the model parameters of the target detection model according to the first loss and the second loss to complete the training of the target detection model, the method further includes: Pruning the auxiliary training head network in the target detection model; An image to be detected is obtained, and the target detection model is used to detect whether a target object exists in the image to be detected, and a detection result output by the original head network of the YOLO model is obtained.

8. A target detection device, characterized in that: include: A building module, configured to build a large-core attention network using depthwise separable convolutional layers, depthwise dilated convolutional layers, and channel-wise convolutional layers, and to build an auxiliary training head network using convolutional layers; A model improvement module is configured to add a large core attention network to a fast spatial pyramid pooling network in a backbone network of a YOLO model, replace a neck network of the YOLO model with a progressive feature pyramid network, connect an auxiliary training head network in parallel to the progressive feature pyramid network in the YOLO model, and use the obtained new model as a target detection model; An acquisition module is configured to acquire a training image, use the target detection model to detect whether a target object exists in the training image, and obtain a first detection result output by the original head network of the YOLO model and a second detection result output by the auxiliary training head network; a calculation module, configured to calculate a first loss between the first detection result and the label of the training image, and calculate a second loss between the second detection result and the label of the training image; An optimization module is configured to optimize the model parameters of the target detection model according to the first loss and the second loss to complete the training of the target detection model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.