Dynamic target detection method and system based on full convolutional neural network
By using a fully convolutional neural network in dynamic object detection, replacing the residual module of the backbone network as a Fire module, and combining feature pyramids and deformed convolutional networks, the problem of high computational complexity of deep learning models is solved, and efficient dynamic object detection is achieved, suitable for real-time applications.
Patent Information
- Application Number
- CN202510280803.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-13
AI Technical Summary
The high computational complexity of modern deep learning models in dynamic object detection leads to a decrease in processing speed, which is difficult to meet the needs of real-time applications, and existing model optimization and acceleration technologies usually sacrifice a certain detection accuracy.
The dynamic object detection method based on a fully convolutional neural network is adopted, and the Hourglass-104 network is used as the backbone network, and its residual module is replaced with the Fire module in CornerNet Lite, which reduces the amount of algorithm parameters and calculation complexity, and uses feature pyramid modules and deformed convolutional networks to improve detection accuracy.
While reducing the amount of algorithm parameters, it ensures the target detection performance and provides a reliable dynamic target detection solution, suitable for real-time application scenarios.
Smart Images

Figure CN120147661A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a dynamic target detection method and system based on a fully convolutional neural network. Background Art
[0002] Dynamic target detection is a key research direction in the field of computer vision, aiming to identify and track moving objects in the environment in real time. This technology is of great significance in many practical applications, especially in the fields of autonomous driving, intelligent monitoring, robot navigation, etc. With the rapid development of artificial intelligence and machine learning technologies, the application of dynamic target detection has become more and more extensive and important. In short, as a key technology, dynamic target detection plays an irreplaceable important role in improving the intelligence and safety of automated systems. With the continuous progress of technology, the application prospect of dynamic target detection will be broader, having a profound impact on the development of various industries. The basic task of dynamic target detection is to detect all existing targets from the input video frames and provide accurate bounding boxes and class labels for each target. Compared with static image detection, dynamic target detection needs to handle more complex scene changes, including the movement of targets, illumination changes, background interference, and the deformation of targets. In addition, real-time performance is also a major challenge for dynamic target detection, because many application scenarios, such as autonomous driving and monitoring systems, require the algorithm to quickly and accurately detect targets at high frame rates.
[0003] Although the application of modern deep learning models in dynamic target detection has improved the detection accuracy, their high computational complexity often leads to a decrease in processing speed, making it difficult to meet the requirements of real-time applications. For example, complex Convolutional Neural Networks (CNNs) and Transformer-based models have significant advantages in detection accuracy, but their application on embedded devices and mobile platforms still faces challenges. Therefore, how to reduce the computational complexity and energy consumption of the model while ensuring the detection accuracy is an urgent problem to be solved. To solve the real-time problem, researchers have proposed various model optimization and acceleration techniques. For example, model pruning and quantization techniques can accelerate the inference process by reducing the model parameters and computational amount. However, these methods usually sacrifice a certain amount of detection accuracy. In addition, the knowledge distillation method trains a lightweight model to imitate the behavior of a large and complex model to achieve accelerated inference, but the detection performance of the lightweight model usually cannot fully reach the level of the original large model. Therefore, in view of the above problems, more efficient dynamic target detection methods need to be proposed. Summary of the Invention
[0004] In view of the above problems, the purpose of the present invention is to provide a dynamic target detection method and system based on a fully convolutional neural network, which reduces the number of algorithm parameters while ensuring the target detection performance, providing a reliable basis for dynamic target detection.
[0005] The above object of the present invention is achieved by the following technical solutions: A dynamic target detection method based on a fully convolutional neural network, comprising the following steps: S1: Input the image to be subjected to dynamic target detection into the backbone network in the fully convolutional neural network for feature extraction; S2: Use the corner prediction branch to predict the upper left corner point and the lower right corner point of the target through the cascaded corner pooling layer; S3: Use the center point prediction branch to predict the center point of the target through the center point pooling layer, and screen the bounding box according to whether the center point is within the central region; S4: Use the flexible non-maximum suppression algorithm to remove the redundant bounding boxes, and obtain the final detection result through the loss function of the network model.
[0006] Further, in step S1, inputting the image to be subjected to dynamic target detection into the backbone network in the fully convolutional neural network for feature extraction specifically includes: Use the Hourglass-104 network as the backbone network, and at the same time replace the residual module in the Hourglass-104 with the Fire module in the CornerNet Lite; In the Hourglass-104, use 4×4 deconvolution instead of the upsampling operation, and remove the max pooling layer to reduce the checkerboard effect caused by the dilated convolution.
[0007] Further, replacing the residual module in the Hourglass-104 with the Fire module in the CornerNet Lite specifically includes: The Fire module includes a Squeeze layer and an Expand layer; The Squeeze layer uses a 1×1 convolutional kernel to reduce the number of channels of the input feature map, thereby reducing the computational amount and effectively compressing the information of the feature map; The Expand layer includes two types of convolutional kernels, 1×1 and 3×3. The convolutional kernels are used to expand the number of channels of the feature map. The 1×1 convolutional kernel is used for lightweight channel mixing, while the 3×3 convolutional kernel is used to capture more spatial information.
[0008] Further, in step S1, it further includes: Add a Feature Pyramid Network (FPN) module on the basis of the backbone network, and form the FPN module with feature maps of different resolution sizes generated at each stage. Change the upsampling operation in the original FPN module to a deformable convolutional network (DCN), and learn the offsets through a parallel network learning module to offset the sampling points of the convolutional kernel from the input feature map, focusing on the regions of interest.
[0009] Furthermore, in step S2, use the corner prediction branch to predict the upper left corner point and the lower right corner point of the target through the cascaded corner pooling layer, specifically: Adopt cascaded corner pooling to search along the boundary to find the boundary maximum value, and then search within the position of the boundary maximum value to determine the internal maximum value. Add the boundary maximum value and the internal maximum value together. The cascaded corner pooling enables the corner points to obtain both boundary information and the visual pattern of the object.
[0010] Furthermore, in step S3, use the center point prediction branch to predict the center point of the target through the center point pooling layer, and screen the bounding boxes based on whether the center point is within the central region, specifically: Adopt center pooling to capture the visual pattern. The detailed process of the center pooling is as follows: The backbone network outputs a feature map, find the maximum values in the horizontal and vertical directions and add these two values together to determine whether a certain pixel in the feature map is a central key point.
[0011] Furthermore, in step S4, use the soft non-maximum suppression algorithm to remove the redundant bounding boxes, and obtain the final detection result through the loss function of the network model, specifically: The loss function is specifically: Among them, and represent the focal loss, which is used to train the network to detect corner points and central key points respectively. is the pull loss of the corner points, which is used to minimize the distance between the embedding vectors belonging to the same object. is the push loss of the corner points, which is used to maximize the distance between the embedding vectors belonging to different objects. and are the l1 loss, which is used to train the network to predict the offsets of the corner points and the central key points. α, β, and γ respectively represent the weights of the corresponding losses.
[0012] A fully convolutional neural network-based dynamic object detection system for performing the above-mentioned fully convolutional neural network-based dynamic object detection method, including: The backbone network extraction module is used to input the image to be dynamically target-detected into the backbone network of the fully convolutional neural network for feature extraction; The corner pooling module is used to predict the upper left corner point and the lower right corner point of the target through the cascaded corner pooling layer by using the corner prediction branch; The center point pooling module is used to predict the center point of the target through the center point pooling layer by using the center point prediction branch, and screen the bounding box according to whether the center point is within the central area; The bounding box removal module is used to remove the redundant bounding boxes by using the flexible non-maximum suppression algorithm, and obtain the final detection result through the loss function of the network model.
[0013] A computer device includes a memory and one or more processors. Computer code is stored in the memory. When the computer code is executed by the one or more processors, the one or more processors execute the method as described above.
[0014] A computer-readable storage medium stores computer code. When the computer code is executed, the method as described above is executed.
[0015] Compared with the prior art, the present invention has at least one of the following beneficial effects: The dynamic target detection method based on the fully convolutional neural network provided by the present application first inputs the image into the backbone network of the fully convolutional neural network; then, uses the corner prediction branch to predict the upper left corner point and the lower right corner point of the target through the cascaded corner pooling layer, uses the center point prediction branch to predict the center point of the target through the center point pooling layer, and screens the bounding box according to whether the center point is within the central area; finally, uses the flexible non-maximum suppression algorithm to remove the redundant bounding boxes to obtain the final detection result. The present invention ensures the target detection performance while reducing the algorithm parameter quantity, and provides a reliable basis for dynamic target detection. Description of the Drawings
[0016] Figure 1 It is a flowchart of the dynamic target detection method based on the fully convolutional neural network provided by the embodiment of the present invention; Figure 2 It is a network architecture diagram of the dynamic target detection method based on the fully convolutional neural network provided by the embodiment of the present invention; Figure 3 It is a network architecture diagram of the Fire module provided by the embodiment of the present invention; Figure 4 It is a feature pyramid network architecture diagram provided by the embodiment of the present invention; Figure 5 It is a deformable convolutional network architecture diagram provided by the embodiment of the present invention; Figure 6 The structural diagram of central pooling and cascaded corner pooling provided by the embodiments of the present invention; Figure 7 The schematic diagram of the dynamic target detection result provided by the embodiments of the present invention. Detailed implementation manners
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0018] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups.
[0019] The following will describe some implementation manners of the present application in detail with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0020] 1. Network architecture The present invention proposes a dynamic target detection method based on a fully convolutional neural network, and its network architecture is as Figure 2 shown. This method uses the Hourglass-104 network as the backbone network, and at the same time replaces the residual modules in the Hourglass-104 with the Fire modules in the CornerNet Lite to reduce the number of algorithm parameters and improve the detection speed. In addition, in the Hourglass-104, 4×4 deconvolution is used instead of the upsampling operation, and the max pooling layer is removed to reduce the checkerboard effect caused by dilated convolution. Finally, more visual patterns are introduced for the central key points and corner points through central pooling and cascaded corner pooling to improve the detection accuracy.
[0021] The working process of this method is as follows: As Figure 1As shown in the figure, first, the image is input into the backbone network of the fully convolutional neural network; then, the corner prediction branch uses the cascaded corner pooling layer to predict the upper left corner point and the lower right corner point of the target, and the center point prediction branch uses the center point pooling layer to predict the center point of the target. The bounding box is screened based on whether the center point is within the central region; finally, the flexible non-maximum suppression algorithm is used to remove redundant bounding boxes to obtain the final detection result.
[0022] 2. Fire Module To ensure the target detection performance while reducing the number of algorithm parameters, the present invention uses the Fire module in CornerNet-Lite, and its network structure is as Figure 3 shown. The Fire module consists of two main parts: the Squeeze layer and the Expand layer. This design is inspired by the SqueezeNet architecture, which improves the efficiency of the network by reducing the amount of computation and the number of parameters. The Squeeze layer uses a 1×1 convolutional kernel to reduce the number of channels of the input feature map, thereby reducing the amount of computation. In this way, the information of the feature map is effectively compressed. The Expand layer includes 1×1 and 3×3 convolutional kernels, and these convolutional kernels are used to expand the number of channels of the feature map. The 1×1 convolutional kernel is mainly used for lightweight channel mixing, while the 3×3 convolutional kernel is used to capture more spatial information.
[0023] 3. FPN Module Since the existing backbone network is a single-mode network and does not combine high-resolution low-semantic feature maps and low-resolution high-semantic feature maps, the detection effect on small targets is poor. Therefore, the present invention adds a Feature Pyramid Network (FPN) on the basis of the backbone network. The FPN is formed by the feature maps of different resolution sizes generated at each stage. The FPN adopts the nearest neighbor upsampling algorithm, which has the advantages of small computational amount and fast speed. However, the nearest neighbor upsampling directly takes the nearest feature of the sampling point as the feature of the sampling point without considering the influence of other adjacent features. Therefore, the sampled features will be discontinuous. During the upsampling process, the features will be deformed and the relationship between the features will be affected. The Deformable Convolutional Network (DCN) uses a parallel network learning module to learn the offset, so that the convolutional kernel offsets the sampling points of the input feature map and focuses on the regions of interest. Therefore, the present invention changes the upsampling operation in the original FPN to DCN, correspondingly improving the accuracy. DCN is beneficial to fully detect complex scenes. The original FPN structure and the DCN structure are respectively as Figure 4 and Figure 5 shown.
[0024] 4. Center Pooling The geometric center of an object does not always convey a recognizable visual pattern (for example, a human head contains strong visual patterns, but the central key point is usually located in the middle of the human body). To solve this problem, the present invention proposes center pooling to capture richer and more recognizable visual patterns. The detailed process of center pooling is as follows: The backbone network outputs a feature map. To determine whether a certain pixel in the feature map is a central key point, it is necessary to find the maximum values in the horizontal and vertical directions and add these two values. Center pooling helps to improve the detection effect of central key points.
[0025] 5. Cascaded corner pooling Corners are usually located outside the object and lack local appearance features. Existing methods use corner pooling to solve this problem. Corner pooling aims to find the maximum value on the boundary to determine the corner. However, this method makes the corner sensitive to the edge. To solve this problem, the present invention proposes cascaded corner pooling, where the corner extracts features from the central region of the object. Cascaded corner pooling searches along the boundary to find the maximum value on the boundary, then searches within the position of the maximum boundary value to determine the maximum value inside, and then adds these two maximum values. Cascaded corner pooling enables the corner to obtain both boundary information and the visual pattern of the object.
[0026] Center pooling and cascaded corner pooling can be achieved by applying corner pooling in different directions. Figure 6 Shows the structure of the center pooling module. To determine the maximum value in a specific direction, such as the horizontal direction, simply connect the left and right pooling in sequence. Figure 6 Shows the structure of the cascaded corner pooling module, where the white rectangle represents a 3×3 convolution followed by batch normalization.
[0027] 6. Loss function Training the dynamic object detection network proposed by the present invention on an NVIDIA TITAN RTX (24GB) GPU, the loss function in this method draws on the loss function of CenterNet, which is specifically described as follows: Among them, and represent the focal loss, which is used to train the network to detect corners and central key points respectively. is the "pull" loss of the corner, which is used to minimize the distance between the embedding vectors belonging to the same object. is the "push" loss of the corner, which is used to maximize the distance between the embedding vectors belonging to different objects. and It is the L1 loss, which is used to train the network to predict the offsets of corner points and center key points. α, β, and γ respectively represent the weights of the corresponding losses, which are set to 0.1, 0.1, and 1. The batch size is 48, and the maximum number of training epochs is 100. In the first 88 epochs of training, the learning rate is set to 2.5×10-4, and in the remaining 12 epochs of training, the learning rate is set to 2.5×10-5. Finally, the dynamic object detection results proposed by the present invention are as Figure 7 shown.
[0028] In this embodiment, first, the image is input into the backbone network of the fully convolutional neural network; then, the corner prediction branch predicts the upper left corner point and the lower right corner point of the target through the cascaded corner pooling layer, and the center point prediction branch predicts the center point of the target through the center point pooling layer, and the bounding box is screened by whether the center point is within the central region; finally, the flexible non-maximum suppression algorithm is used to remove redundant bounding boxes to obtain the final detection result. The present invention ensures the object detection performance while reducing the algorithm parameter quantity, providing a reliable basis for dynamic object detection.
[0029] A computer-readable storage medium stores computer code, and when the computer code is executed, the above method is executed. Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0030] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
[0031] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0032] It should be noted that the above embodiments can be freely combined as needed. The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A dynamic target detection method based on a fully convolutional neural network, characterized in that: The following steps are involved: S101: Input the image that needs to be detected as a dynamic target into the backbone network of the fully convolutional neural network for feature extraction; S102: using the corner prediction branch to predict the upper left corner and lower right corner of the target through the cascaded corner pooling layer; S103: using the center point prediction branch to predict the center point of the target through the center point pooling layer, and filtering the bounding box according to whether the center point is in the center area; S104: Use a flexible non-maximum suppression algorithm to remove the redundant bounding boxes, and obtain the final detection result through the loss function of the network model.
2. The dynamic target detection method based on a fully convolutional neural network according to claim 1, characterized in that: In step S1, the image that needs to be detected for dynamic targets is input into the backbone network of the fully convolutional neural network for feature extraction, specifically: Use the Hourglass-104 network as the backbone network, and replace the residual module in Hourglass-104 with the Fire module in CornerNet Lite; In Hourglass-104, 4×4 deconvolution is used instead of upsampling operation, and the maximum pooling layer is removed to reduce the checkerboard effect caused by dilated convolution.
3. The dynamic target detection method based on a fully convolutional neural network according to claim 2 is characterized in that: Replace the residual module in Hourglass-104 with the Fire module in CornerNet Lite, specifically: The Fire module includes a Squeeze layer and an Expand layer; The Squeeze layer uses a 1×1 convolution kernel to reduce the number of channels of the input feature map, thereby reducing the amount of calculation and effectively compressing the information of the feature map; The Expand layer includes two types of convolution kernels, 1×1 and 3×3. The convolution kernels are used to expand the number of channels of the feature map. The 1×1 convolution kernel is used for lightweight channel mixing, while the 3×3 convolution kernel is used to capture more spatial information.
4. The dynamic target detection method based on a fully convolutional neural network according to claim 1, characterized in that: In step S1, it also includes: Adding a feature pyramid module FPN on the basis of the backbone network, and forming the feature pyramid module FPN through feature maps of different resolutions generated in each stage; The upsampling operation in the original feature pyramid module FPN is changed to a deformable convolutional network DCN, and a parallel network learning module is used to learn the offset so that the convolution kernel offsets the sampling points of the input feature map to focus on the area of interest.
5. The dynamic target detection method based on a fully convolutional neural network according to claim 1, characterized in that: In step S2, the corner prediction branch is used to predict the upper left corner and lower right corner of the target through the cascaded corner pooling layer, specifically: Cascaded corner pooling is used to search along the boundary to find the boundary maximum, and then search within the position of the boundary maximum to determine the internal maximum. The boundary maximum and the internal maximum are added together. Cascaded corner pooling enables the corner points to obtain boundary information and the visual pattern of the object at the same time.
6. The dynamic target detection method based on a fully convolutional neural network according to claim 1, characterized in that: In step S3, the center point prediction branch is used to predict the center point of the target through the center point pooling layer, and the bounding box is screened according to whether the center point is in the center area, specifically: Center pooling is used to capture visual patterns. The detailed process of center pooling is as follows: the backbone network outputs a feature map, finds the maximum value in the horizontal and vertical directions and adds the two values to determine whether a pixel in the feature map is a center key point.
7. The dynamic target detection method based on a fully convolutional neural network according to claim 1, characterized in that: In step S4, the redundant bounding boxes are removed using a flexible non-maximum suppression algorithm, and the final detection result is obtained through the loss function of the network model, specifically: The loss function is specifically: in, and represents the focal loss, which is used to train the network to detect corner points and center key points respectively. is the pull loss of corner points, which is used to minimize the distance between embedding vectors belonging to the same object. is the push loss of the corner points, which is used to maximize the distance between the embedding vectors belonging to different objects. and is the l1 loss, which is used to train the network to predict the offsets of corner points and center keypoints. α, β, and γ represent the weights of the corresponding losses.
8. A dynamic target detection system based on a fully convolutional neural network for executing the dynamic target detection method based on a fully convolutional neural network as described in any one of claims 1 to 7, characterized in that: include: The backbone network extraction module is used to input the image that needs to be detected by dynamic targets into the backbone network of the full convolutional neural network for feature extraction; Corner pooling module, used to predict the upper left corner and lower right corner of the target through cascaded corner pooling layers using the corner prediction branch; A center point pooling module, used to predict the center point of the target through the center point pooling layer using the center point prediction branch, and to filter the bounding box according to whether the center point is within the center area; The bounding box removal module is used to remove the redundant bounding boxes using a flexible non-maximum suppression algorithm, and obtain the final detection result through the loss function of the network model.
9. A computer device comprising a memory and one or more processors, wherein the memory stores computer codes, and when the computer codes are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 7. 10 . A computer-readable storage medium storing a computer code. When the computer code is executed, the method according to claim 1 is executed.