Lightweight panoramic target detection method for small target

By improving the YOLOv7 model, the lightweight convolution and attention mechanism were introduced, and the panoramic object detection network was optimized, which solved the calculation complexity and accuracy of small object detection in the panoramic monitoring system, and achieved efficient and accurate small object detection.

CN120374958APending Publication Date: 2025-07-25HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510790399.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, small object detection in panoramic monitoring systems has problems such as high computational complexity, large parameter scale, difficulty in deploying to edge computing platforms and low detection accuracy, especially in complex environments, small object detection results are not good.

Method used

Using the improved YOLOv7 model, by introducing lightweight convolution modules, attention mechanisms and feature fusion modules, we optimize the backbone and head networks, reduce the model parameter quantity and calculation complexity, and enhance the small object detection capability.

Benefits of technology

While maintaining detection accuracy, the detection speed and accuracy of small targets are significantly improved, and are suitable for resource-constrained equipment, improving the real-time and detection efficiency of the panoramic monitoring system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374958A_ABST
    Figure CN120374958A_ABST
Patent Text Reader

Abstract

The invention relates to a small-target-oriented lightweight panoramic target detection method. The method comprises the steps of collecting a to-be-detected panoramic image; the to-be-detected panoramic image is input into a real-time target detection model, a target detection result is output, the real-time target detection model is obtained by training an improved YOLOv7 model based on a target detection data set, the target detection data set comprises an image of a marked target, and the target detection result is obtained through the image of the marked target. The improved YOLOv7 model is obtained by improving a CBS unit of the YOLOv7 model and embedding an attention mechanism. According to the method, the bottleneck problems of large parameter scale and high calculation complexity of a traditional network model are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and particularly to a lightweight panoramic object detection method for small objects. Background Art

[0002] With the rapid explosive development of intelligent vision technology, panoramic monitoring systems are increasingly widely and deeply applied in many important fields such as smart cities, public security, and industrial automation. The significant advantage of panoramic images lies in their extremely wide field of view, which can achieve accurate and comprehensive overall perception of complex scenes, efficiently cover large areas, and reduce monitoring blind spots. However, this panoramic mode also brings huge challenges in small object detection. In high-resolution panoramic images, various objects such as pedestrians, vehicles, and components often appear in extremely small pixel proportions, even less than 0.1%. These objects are extremely small in size, extremely scattered in distribution, have no obvious pattern to follow, and are easily interfered by complex backgrounds, such as light changes, object occlusion, and complex textures of the surrounding environment. The traditional method relying on manual inspection is extremely inefficient, and each manual detection requires a large amount of time and effort. In addition, due to the limitations of the human eye vision itself, it is difficult to achieve real-time and accurate recognition of small objects in a large amount of panoramic data. This leads to problems such as untimely hidden danger investigation in the field of public security and a high missed detection rate in the industrial quality inspection link. Therefore, there is an urgent need to adopt intelligent small object detection technology to achieve automated management of the scene and improve the safety and operation efficiency of various fields.

[0003] The current mainstream object detection algorithms mainly include two-stage detection methods based on candidate region generation and single-stage detection methods based on end-to-end architectures. The two-stage algorithms first extract candidate regions and then perform fine-grained classification. Although they have high accuracy, they have problems such as high computational complexity and insufficient real-time performance; single-stage algorithms such as the YOLO series are known for their speed, but they still face difficulties in small object detection in complex environments and lightweight model deployment. In addition, as the panoramic monitoring scene continues to expand, small objects account for a very small proportion in the image, are sparsely distributed, and lack feature information, and are extremely vulnerable to factors such as light changes, object occlusion, and environmental textures. At the same time, the parameter scale of mainstream detection models is usually large, making it difficult to directly deploy them to edge computing platforms, and traditional model lightweight methods may lead to a decline in the network feature expression ability, thereby affecting the final detection effect.

[0004] In view of the above problems, the present invention provides a lightweight panoramic object detection method for small objects. Summary of the Invention

[0005] The object of the present invention is to provide a lightweight panoramic object detection method for small targets. Based on the widely used YOLOv7 framework with outstanding performance, in-depth optimization and improvement have been carried out around the features and difficulties of panoramic small target detection. The backbone network of YOLOv7 is designed to be lightweight, and lightweight convolutional modules are introduced, effectively reducing the number of model parameters and computational complexity, and improving the detection speed while maintaining the detection accuracy as much as possible.

[0006] To achieve the above object, the present invention provides the following solutions:

[0007] A lightweight panoramic object detection method for small targets, comprising:

[0008] Collect the panoramic image to be detected;

[0009] Input the panoramic image to be detected into a real-time object detection model, and output the object detection result. Among them, the real-time object detection model is obtained by training an improved YOLOv7 model based on an object detection data set. The object detection data set includes images with labeled targets. The improved YOLOv7 model is obtained by improving the CBS unit in the YOLOv7 model and embedding an attention mechanism.

[0010] Optionally, the improved YOLOv7 model includes: a backbone network and a head network;

[0011] The backbone network includes a first SE attention module, an L-ELAN module, an L-MP1 module, and a CBS unit. Among them, the L-ELAN module is obtained by lightweighting the CBS unit in the original ELAN module of the YOLOv7 model. The L-MP1 module is obtained by combining a max-pooling operation and a lightweight basic feature extraction module. The first SE attention module is connected to the L-ELAN module;

[0012] The head network includes a lightweight basic feature extraction module, a second SE attention module, an L-MP2 module, an L-ELAN-H module, an SPPF module, an L-RepConv module, and a small target detection layer. Among them, the SPPF module replaces the SPPCSPC module in the YOLOv7 model. The L-ELAN-H module is obtained by lightweighting the CBS unit in the original ELAN-H module of the YOLOv7 model. The L-RepConv module is obtained by lightweighting the CBS unit in the original RepConv module of the YOLOv7 model. The CBS unit is lightweighted by replacing the original CBS unit with a lightweight basic feature extraction module. The second SE attention module is connected to the L-ELAN-H module. The small target detection layer is used to obtain fine-grained small target information.

[0013] Optionally, the fine-grained small target information obtained by the small target detection layer includes: performing upsampling on the L-ELAN-H layer in the feature fusion module, and fusing it with the features output by the first L-ELAN module in the backbone network to generate a new high-resolution feature map.

[0014] Optionally, the SPPF module uses maximum pooling operations of different scales to replace the convolution operations with convolution kernels of different sizes, and generates a feature map with a fixed size.

[0015] Optionally, the lightweight basic feature extraction module is obtained by replacing the convolution module in the original CBS unit with a depthwise separable convolution module and replacing the SiLU activation function with an FReLU activation function. Optionally, obtaining the target detection dataset includes:

[0016] Collect airport surveillance video images to generate an image dataset;

[0017] Annotate the airplanes and vehicles in each image in the image dataset to obtain the target detection dataset.

[0018] Optionally, after training the improved YOLOv7 model, the detection performance is comprehensively verified using Precision, Recall, and mAP metrics.

[0019] The beneficial effects of the present invention are as follows: Based on the YOLOv7 object detection algorithm, the present invention designs a lightweight basic feature extraction module (LBF) to replace the original CBS structural unit; introduces the SPPF module in the head network to replace the computationally complex SPPCSPC module in the original network; at the same time, a detection layer dedicated to small targets is added to the head network, and the SE attention mechanism is integrated into the YOLOv7 backbone network to construct a lightweight panoramic small target detection method based on the improved YOLOv7 framework. It effectively solves the bottleneck problems of large parameter scale and high computational complexity of traditional network models; significantly improves the problems of insufficient small target feature expression ability and low detection accuracy, and realizes the coordinated improvement of detection accuracy and model lightweight level. Description of the Drawings

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a flowchart of a lightweight panoramic target detection method for small targets according to an embodiment of the present invention;

[0022] Figure 2 Schematic diagram of the LBF structure according to an embodiment of the present invention;

[0023] Figure 3 Structural diagram of the lightweight panoramic object detection network for small objects according to an embodiment of the present invention;

[0024] Figure 4 Schematic diagram of the basic convolutional structure of CBS according to an embodiment of the present invention;

[0025] Figure 5 Schematic diagram of the SE structure according to an embodiment of the present invention;

[0026] Figure 6 Schematic diagram of the L-ELAN structure according to an embodiment of the present invention;

[0027] Figure 7 Schematic diagram of the L-MP structure according to an embodiment of the present invention;

[0028] Figure 8 Schematic diagram of the L-ELAN-H structure according to an embodiment of the present invention;

[0029] Figure 9 Schematic diagram of the SPPF structure according to an embodiment of the present invention;

[0030] Figure 10 Schematic diagram of the L-RepConv structure according to an embodiment of the present invention. Detailed implementation manners

[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0032] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0033] This embodiment provides a lightweight panoramic object detection method for small objects, including:

[0034] Collect the panoramic image to be detected;

[0035] Input the panoramic image to be detected into the real-time object detection model, and output the object detection results. Among them, the real-time object detection model is obtained by training the improved YOLOv7 model based on the object detection dataset. The object detection dataset includes images with labeled objects. The improved YOLOv7 model is obtained by improving the CBS units in the YOLOv7 model and embedding the attention mechanism.

[0036] Further, the improved YOLOv7 model includes: a backbone network and a head network;

[0037] The backbone network includes a first SE attention module, an L-ELAN module, an L-MP1 module, and CBS units. Among them, the L-ELAN module is obtained by lightweighting the CBS units in the original ELAN module of the YOLOv7 model. The L-MP1 module is obtained by combining the max pooling operation with the lightweight basic feature extraction module. The first SE attention module is connected to the L-ELAN module;

[0038] The head network includes a lightweight basic feature extraction module, a second SE attention module, an L-MP2 module, an L-ELAN-H module, an SPPF module, an L-RepConv module, and a small object detection layer. Among them, the SPPF module replaces the SPPCSPC module in the YOLOv7 model. The L-ELAN-H module is obtained by lightweighting the CBS units in the original ELAN-H module of the YOLOv7 model. The L-RepConv module is obtained by lightweighting the CBS units in the original RepConv module of the YOLOv7 model. The CBS unit lightweighting is to replace the original CBS unit with the lightweight basic feature extraction module. The second SE attention module is connected to the L-ELAN-H module. The small object detection layer is used to obtain fine-grained small object information.

[0039] Specifically, from a structural perspective, the LBF module significantly reduces the overall parameter scale while ensuring the detection accuracy by reconstructing the convolution operation and reasonably optimizing the parameter allocation. Applying the SPPF module to the head network of YOLOv7 to replace the original SPPCSPC structure not only effectively reduces the model parameter quantity but also improves the diversity of feature extraction and the detection efficiency. The small object detection layer effectively improves the detection performance of small objects while controlling the network complexity. The SE module significantly enhances the network's perception ability of key features with extremely small parameter overhead.

[0040] Further, the small object detection layer obtaining fine-grained small object information includes: performing upsampling processing on the L-ELAN-H layer in the feature fusion module, and fusing it with the features output by the first L-ELAN module in the backbone network to generate a new high-resolution feature map.

[0041] Specifically, after further feature refinement by the ELAN module, a new high-resolution feature layer is formed as an additional small object detection head to improve the detection accuracy of small objects. At the same time, through downsampling operations, the detailed features extracted from the shallow layer are effectively introduced into the neck feature fusion structure to achieve efficient integration of features at different scales, thereby significantly enhancing the small object detection ability and overall robustness of the model in complex panoramic image environments such as airport scene monitoring.

[0042] Furthermore, SPPF uses max pooling operations at different scales to replace the convolutional operations with convolutional kernels of different sizes to generate feature maps of a fixed size.

[0043] Furthermore, the lightweight basic feature extraction module is obtained by replacing the convolutional module in the original CBS unit with a depthwise separable convolutional module and replacing the SiLU activation function with the FReLU activation function.

[0044] Specifically, as Figure 1 shown, replacing the traditional CBS unit in YOLOv7 with the LBF module can effectively relieve the computational pressure of the model, improve the running speed and deployment convenience, and provide excellent support for the application of the algorithm on resource-constrained devices.

[0045] Furthermore, obtaining the object detection dataset includes:

[0046] Collecting airport surveillance video images to generate an image dataset;

[0047] Annotating the airplanes and vehicles in each image in the image dataset to obtain the object detection dataset.

[0048] Furthermore, after training the improved YOLOv7 model, it includes: comprehensively verifying the detection performance using Precision, Recall, and mAP metrics.

[0049] The method of this embodiment will be further described below with reference to the accompanying drawings:

[0050] A lightweight panoramic object detection for small objects, the overall process is as Figure 1 shown and it includes the following steps:

[0051] Step 1: Collecting airport surveillance video images to generate an image dataset;

[0052] Step 2: Annotating the airplanes and vehicles in each image in the dataset, establishing a dedicated object detection dataset, and dividing it into a training set and a test set according to a ratio of 8:2;

[0053] Step 3: Based on the improved YOLOv7 framework, constructing a real-time object detection model for airport scenarios;

[0054] Step 4: Train the improved YOLOv7 model constructed in Step 3 to optimize the detection performance;

[0055] Step 5: Use the lightweight panoramic small target detection model trained in Step 4 to detect the test set and verify the performance through relevant evaluation indicators of object detection;

[0056] Step 6: Apply the trained weight parameters to the detect.py script, build a Python inference framework, run the improved model, and verify that its detection accuracy is higher than that of the original model in the airport scenario and can meet the real-time requirements.

[0057] Among them, the improved YOLOv7 framework innovatively designs a lightweight basic feature extraction module to replace the original CBS unit structure; introduces the SPPF module in the head network to improve the feature fusion efficiency, and at the same time adds a feature layer dedicated to small target detection; in addition, embeds the SE attention mechanism in the backbone network to further enhance the feature extraction ability.

[0058] 1. LBF basic module:

[0059] In the YOLOv7 object detection framework, the CBS basic module is a key building unit, which is composed of a standard convolutional layer, a normalization layer and a SiLU activation function stacked in sequence, and supports the entire network structure through different forms of combination. During the feature extraction stage, the standard convolution needs to traverse all channels of the input feature, perform convolution operations using multi-scale convolutional kernels, and fuse the results with the original feature. Although it helps to extract rich feature information, it also brings extremely high parameter quantities and computational overheads, posing great challenges to computing resources and processing speeds. Therefore, the present invention designs a lightweight basic feature module (LBF), whose design concept draws on the advantages of depthwise separable convolution in reducing parameter quantities and computational complexities, as Figure 2 shown. In specific implementation, the LBF module replaces the convolutional module (Conv) in the original CBS unit with a depthwise separable convolutional module (DConv), and at the same time replaces the SiLU activation function with the FReLU activation function, effectively improving the classification and detection performance while maintaining the lightweight characteristics. The original CBS unit structure is as Figure 4 shown. Different from the standard convolution in CBS where regions and channels are processed simultaneously, LBF transforms into a process of first performing depthwise convolution to extract regional features and then pointwise convolution to increase channels, realizing the separation of channels and regions. This method can reduce a large number of network parameters and maintain a certain accuracy.

[0060] As Figure 1As shown in the figure, replacing the traditional CBS unit in YOLOv7 with the LBF module can effectively relieve the computational pressure of the model, improve the running speed and deployment convenience, and provide excellent support for the application of the algorithm on resource-constrained devices.

[0061] 2. L-ELAN-H module:

[0062] In the L-ELAN-H module, the LBF is used to replace the traditional CBS convolution module. While ensuring the feature extraction ability, it reduces the computational cost and the number of model parameters. After the input feature map enters the L-ELAN-H module, the 1×1 convolution first adjusts the channel dimension to integrate the information between different channels, and then the 3×3 convolution further extracts features in the spatial dimension to mine the details such as the edges and textures of the objects in the image. Multiple LBFs are connected in series in the L-ELAN-H module, and the feature map passes through each LBF in turn to realize the gradual extraction and refinement of features. The output of the previous LBF is used as the input of the next LBF, continuously deepening the understanding and representation of features, gradually transitioning from the initial feature extraction to a more abstract and discriminative feature representation. Multiple LBFs are combined to extract features from different scales and angles, enriching the feature expression ability. The multi-branch structure (such as skip connection) constructed by LBF helps to optimize the gradient propagation path, relieve the problem of gradient disappearance in deep networks, promote the stable training of the network, and improve the model performance.

[0063] 3. L-RepConv module:

[0064] L-RepConv replaces the CBS convolution in the original RepConv module with the LBF lightweight convolution. The LBF reduces the amount of convolution operations, and the reparameterization technology simplifies the structure during inference, reducing the computational complexity. At the same time, a multi-branch structure is adopted, and different branches perform convolution and normalization operations respectively to extract and process features from different angles. The BN normalization layer normalizes the input data to make the data distribution stable, accelerates the model convergence, and prevents the gradient from disappearing or exploding. In L-RepConv, it is ensured that the data of each branch participates in the operation under a suitable distribution.

[0065] Two branches containing the LBF lightweight convolution extract different levels of spatial features, and the branch containing only the BN normalization provides the original feature information after normalization, enriching the feature diversity. The multi-branch structure extracts features from different aspects during training, and the fused features are more abundant and discriminative, improving the model's ability to extract and represent image features. Add the features of the two branches processed by LBF and BN to the features of the branch processed only by BN to achieve feature fusion and enrich the output feature representation.

[0066] 4. SPPF basic module:

[0067] During the process of deep learning research, to address the problem of the exponential growth of network parameters caused by convolutional operations during the training phase, the Spatial Pyramid Pooling (SPP) technique was innovatively introduced. Inspired by the Spatial Pyramid Matching (SPM) idea, the SPP-Net was developed, aiming to endow the network with the ability to process input images of different sizes and regularize them into a unified size, thereby optimizing the subsequent feature extraction and classification processes. SPP-Net solved the problem of the fixed size requirement of the input feature map in traditional networks by connecting a spatial pyramid pooling module after the convolutional layer. Due to the different sizes of convolutional kernels resulting in inconsistent feature map sizes, it was challenging to process them uniformly, while the SPP module effectively alleviated this contradiction. On this basis, the SPPF module further optimized the SPP structure, maintaining the same functionality but being more efficient in structure. SPPF uses max-pooling operations of different scales to replace the convolutional operations of convolutional kernels of different sizes, thereby generating feature maps of a fixed size. Specifically, the input first passes through a CBS unit for convolution to generate a preliminary feature map, and then through max-pooling operations at three scales of 5×5, 9×9, and 13×13 to extract multi-level features. The pooling results are then fused with the original feature map through a Concat operation. Finally, the fused feature map undergoes convolutional processing to further enrich the feature representation. Applying the SPPF module to the head network of YOLOv7 and replacing the original SPPCSPC structure not only effectively reduces the model parameters but also improves the diversity of feature extraction and detection efficiency.

[0068] 5. Add a small target detection layer:

[0069] In the YOLOv7 object detection framework, its head network covers the output of feature maps for objects of different scales, large, medium, and small. Through the systematic analysis of airport scene monitoring images and panoramic monitoring image data, it is found that the proportion of small objects in this type of dataset is extremely high. Since small objects are extremely small in size in the image, conventional object detection algorithms often suffer from insufficient feature extraction and loss of object information, resulting in frequent missed detections and severely limiting the improvement space of detection accuracy. To solve this technical problem, the present invention effectively improves the detection performance of small objects by adding a small object detection layer to the network while controlling the network complexity. In the design of the feature fusion network, upsampling, downsampling operations, and a dedicated detection layer for small objects are innovatively introduced to fully exploit the fine-grained small object information in the shallow features. The specific method is to perform upsampling on the L-ELAN-H layer in the feature fusion module and fuse it with the features output by the first L-ELAN module in the backbone network to generate a new high-resolution feature map. After further feature refinement by the L-ELAN module, a new high-resolution feature layer is formed as the newly added small object detection head to improve the detection accuracy of small objects. At the same time, through downsampling operations, the high-resolution features extracted from the shallow layer of the backbone network are introduced into the neck network to achieve efficient integration of features at different scales, thereby significantly enhancing the small object detection ability and overall robustness of the model in complex panoramic image environments such as airport scene monitoring.

[0070] 6. SE Attention Module:

[0071] The SE module (Squeeze-and-Excitation module) is a lightweight attention mechanism that focuses on modeling the dependencies between channels (as shown in the figure). It can adaptively recalibrate the channel feature weights, enabling the model to pay more attention to the key feature channels, suppressing irrelevant or redundant information, and thus effectively improving the model performance. The SE module first performs the Squeeze operation, that is, global average pooling is performed on the input feature map in the spatial dimension, compressing the spatial information of each channel into a single value to extract the global channel description. Then it enters the Excitation stage, and the pooled channel description vector is input into a bottleneck structure composed of two fully connected layers (FC layers). The first FC layer is used for dimensionality reduction, introducing non-linearity by combining with the ReLU activation function. The second FC layer is used for dimensionality restoration to the original number of channels, and a normalized channel weight coefficient is generated through the Sigmoid activation function. The generated channel weights reflect the importance of each channel. Finally, by element-wise multiplication (scaling) between channels, these weights are applied to the original feature map to complete the adaptive recalibration of the channel features, highlighting important features and suppressing irrelevant features. Overall, the SE module significantly enhances the network's ability to perceive key features with extremely small parameter overhead.

[0072] 7. Lightweight panoramic small target detection method:

[0073] Overall, the network structure proposed in this embodiment is consistent with YOLOv7 in terms of framework design and still consists of a backbone network and a head network, and its overall structure is as Figure 3 shown.

[0074] The backbone network part includes CBS units (as Figure 4 shown), the first SE attention module (as Figure 5 shown), a lightweight L-ELAN module (as Figure 6 shown), and an L-MP1 structure (as Figure 7 shown). Among them, the CBS unit consists of a lightweight convolutional layer, a normalization layer, and a SiLU activation function. The L-ELAN module is implemented by replacing the CBS unit in the original ELAN module of YOLOv7 with an LBF module, and the L-MP1 structure is composed of a max pooling operation and an LBF module, adopting a dual downsampling strategy to improve the feature extraction efficiency. The head network part integrates an LBF module, a second SE attention module, an L-MP2 structure, an L-ELAN-H module (as Figure 8 shown), an SPPF structure (as Figure 9 shown), an L-RepConv module (as Figure 10 shown), and a small target detection layer. The SPPF module first performs preliminary convolution processing through a CBS unit, then introduces max pooling operations of three different sizes (5×5, 9×9, 13×13) to construct a spatial pyramid pooling network to enrich the hierarchical expression of features. After that, the multi-scale pooling features are concatenated and fused with the original features in the channel dimension, and further refined through convolution operations, enhancing the diversity and expression ability of feature extraction. The output of the L-ELAN module in the backbone network is input to the LBF module.

[0075] Considering the characteristics that the targets in the airport scene are generally small-sized, this embodiment adds a set of small target detection feature maps on the basis of the original YOLOv7 framework. Specifically, the image undergoes downsampling operations of 4 times, 8 times, 16 times, and 32 times to generate feature maps suitable for detecting large, medium, and small-sized targets respectively, and additional feature layers are added for small target detection. These feature maps are uniformly sent to the detection head for inference. During the overall optimization process, the network design always takes reducing the number of parameters and computational overhead and improving the detection speed while maintaining the detection accuracy as the core goal, ensuring the efficiency and real-time performance of the model in practical applications.

[0076] 8. Improved YOLOv7 target detection algorithm file and parameter configuration:

[0077] In this embodiment, the airport scene dataset constructed above and the panoramic lightweight object detection method of YOLOv7 with a fusion attention mechanism are used to complete the verification of the real-time object detection algorithm for airports. The specific steps are as follows:

[0078] Step1: Configure the network training configuration file:

[0079] Select 100 images from the dataset as the training set and 10 images as the test set. Before training, the improved YOLOv7 data and model configuration files need to be modified. The number of object categories in the data file is changed to 3, and the object category names are modified in the category list names. According to the improved network structure, the feature scale and attention module are added. The SE model code is added to the common.py file, and the SE structure is introduced in the yolo.py file. Run the yolo.py file to check whether the network changes are correct, and use the K-means clustering algorithm to calculate the anchor box size of the model.

[0080] Step2: Configure the network training environment:

[0081] The experimental environment of this test is the Dawning cloud computing service system, which is equipped with a Hygon C8671 8532-core CPU and an NVIDIA RTX3080 GPU. The model is built on the deep learning framework PyTorch, and the Python version and Torch version are 3.9.1 and 1.12.1. The network test hardware environment configuration is shown in Table 1, the network test software part environment configuration is shown in Table 2, and the partial parameter settings of the improved YOLOv7 model are shown in Table 3. The specific configuration information is shown in the table.

[0082] Table 1

[0083]

[0084] Table 2

[0085]

[0086] Table 3

[0087]

[0088] 9. Verification based on object detection evaluation metrics:

[0089] Step1: Selection of object detection evaluation metrics:

[0090] All models are trained and tested based on the airport image dataset constructed in this embodiment. During the experiment, Recall, Precision, and mAP are selected as the main evaluation indicators for model performance, and the IOU threshold is uniformly set to 0.5. The mAP indicator represents the average value of the accuracy (Precision) of each category under different Recall conditions and is used to comprehensively measure the detection performance. Among them, the specific definitions of Recall, Precision, and mAP are as follows:

[0091]

[0092] Among them, TP represents the number of correctly identified positive samples, FP represents the number of negative samples misidentified as positive samples, FN represents the number of positive samples that fail to be identified, and m represents the total number of categories involved in the object detection task.

[0093] Step2: Analysis of experimental results:

[0094] To verify the effectiveness of the improved object detection algorithm proposed in the present invention, the experimental results are shown in Table 4 below:

[0095] Table 4

[0096]

[0097]

[0098] In the experiment, the resolution of all network input images is set to 640×640×3, and the number of training epochs is 300. SSD300, Faster-RCNN, CenterNet, YOLOv7, and the method proposed in the present invention are compared and tested. The experimental results show that the proposed real-time object detection method for airport scenarios has achieved the best performance in various evaluation indicators, with the mAP improved by 1.2% compared with the original YOLOv7 network. At the same time, the mAP is improved by 41.3% compared with SSD300, by 36.1% compared with Faster-RCNN, and by 13.0% in terms of the mAP0.5 index compared with CenterNet.

[0099] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention should fall within the protection scope determined by the claims of the present invention.

Claims

1. A lightweight panoramic object detection method for small targets, characterized in that, Including: Collecting the panoramic image to be detected; Inputting the panoramic image to be detected into a real-time object detection model to output an object detection result, where the real-time object detection model is obtained by training an improved YOLOv7 model based on an object detection dataset, the object detection dataset includes images with labeled objects, and the improved YOLOv7 model is obtained by improving the CBS unit in the YOLOv7 model and embedding an attention mechanism.

2. The lightweight panoramic object detection method for small objects according to claim 1, wherein The improved YOLOv7 model includes: a backbone network and a head network; The backbone network includes a first SE attention module, an L-ELAN module, an L-MP1 module, and a CBS unit, where the L-ELAN module is obtained by lightweighting the CBS unit in the original ELAN module of the YOLOv7 model, the L-MP1 module is obtained by combining a max pooling operation and a lightweight basic feature extraction module, and the first SE attention module is connected to the L-ELAN module; The head network includes a lightweight basic feature extraction module, a second SE attention module, an L-MP2 module, an L-ELAN-H module, an SPPF module, an L-RepConv module, and a small object detection layer, where the SPPF module replaces the SPPCSPC module in the YOLOv7 model, the L-ELAN-H module is obtained by lightweighting the CBS unit in the original ELAN-H module of the YOLOv7 model, the L-RepConv module is obtained by lightweighting the CBS unit in the original RepConv module of the YOLOv7 model, the CBS unit is lightweighted by replacing the original CBS unit with a lightweight basic feature extraction module, the second SE attention module is connected to the L-ELAN-H module, and the small object detection layer is used to obtain fine-grained small object information.

3. The lightweight panoramic object detection method for small targets according to claim 2, characterized in that, The small object detection layer obtaining fine-grained small object information includes: performing upsampling processing on the L-ELAN-H layer in the feature fusion module, and fusing it with the features output by the first L-ELAN module in the backbone network to generate a new high-resolution feature map.

4. The lightweight panoramic object detection method for small targets according to claim 2, wherein The SPPF module adopts max pooling operations of different scales to replace the convolutional operations of convolutional kernels of different sizes, and generates a feature map of a fixed size.

5. The lightweight panoramic object detection method for small objects according to claim 2, characterized in that The lightweight basic feature extraction module is obtained by replacing the convolutional module in the original CBS unit with a depthwise separable convolutional module and replacing the SiLU activation function with a FReLU activation function.

6. The lightweight panoramic object detection method for small objects according to claim 1, characterized in that Obtaining the object detection dataset includes: Collecting airport surveillance video images to generate an image dataset; Annotating the airplanes and vehicles in each image in the image dataset to obtain an object detection dataset.

7. The lightweight panoramic object detection method for small objects according to claim 1, characterized in that After training the improved YOLOv7 model, it includes: comprehensively verifying the detection performance using Precision, Recall, and mAP metrics.