A lightweight detection method for panoramic images

By building the panoramic image dataset and improving the YOLOv7 framework, adding the SE attention module and small object detection layer, and using the lightweight convolution module, the problems of inaccurate small object detection and excessive model parameters in panoramic images are solved, and efficient and lightweight object detection is achieved.

CN116681987BActive Publication Date: 2025-07-29HARBIN ENG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310674083.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-08
Publication Date
2025-07-29
Estimated Expiration
2043-06-08

AI Technical Summary

Technical Problem

The prior art lacks applicable data sets in panoramic images, resulting in low accuracy of small object detection and excessive amount of model parameters, making it difficult to realize real-time lightweight object detection on the device side.

Method used

The image data set based on real panoramic video data is constructed, and the YOLOv7 framework is improved by using SE attention module and small object detection layer. Combined with the lightweight convolution module, the network structure is optimized to adapt to the panoramic image characteristics and reduce the amount of model parameters.

Benefits of technology

It improves the detection accuracy and recall of small targets in panoramic images, reduces the amount of model parameters, and realizes real-time lightweight target detection on the device side.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116681987B_ABST
    Figure CN116681987B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight detection method for panoramic images, including: acquiring real panoramic video data, and obtaining an image data set based on the real panoramic video data; performing image feature annotation on the image data set to obtain a target detection data set; constructing a lightweight target detection model, training the lightweight target detection model based on the target detection data set to obtain a lightweight target detection model for panoramic images; and performing target detection based on the lightweight target detection model for panoramic images. The present invention solves the problems of lack of data in the research on panoramic image features, inaccurate detection of too small targets in panoramic images, and excessive model parameter quantities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image recognition, and particularly relates to a lightweight detection method for panoramic images. Background Art

[0002] Nowadays, panoramic images are widely used in fields such as intelligent ships, intelligent driving, and airport security. The panoramic images collected by panoramic cameras are the cornerstones of technologies such as information perception, communication navigation, route planning, and early warning. The targets in panoramic images are extremely small compared to the overall image. Existing target detection technologies cannot correctly detect small target objects, and the network models of small target detection technologies are more massive. In their work, this requires a relatively large information storage space and program running space to achieve target detection quickly and efficiently. Therefore, it is very meaningful to develop and design a lightweight target detection algorithm for panoramic images using computer vision technology to achieve end-to-end real-time detection capabilities on the device side.

[0003] Currently, common target detection technologies can be mainly divided into two categories: two-stage detection and one-stage detection. Among them, two-stage detection methods first generate a set of candidate detection boxes, and then predict the position and category of each box, such as algorithms like R-CNN, Fast R-CNN, and Faster R-CNN. While one-stage detection methods directly predict the position and category in the detection network, such as target detection algorithms like SDD and YOLO.

[0004] The above aspects have the following problems in constructing lightweight target detection for panoramic images: First, the current target detection methods only improve and optimize common datasets, such as PASCAL VOC and MS COCO, etc., and lack datasets suitable for panoramic images. Second, the accuracy of small target detection in the existing YOLO target detection series is relatively low, and most of the targets in panoramic images are small target objects, which affects the accuracy of lightweight target detection for panoramic images. Third, the models of existing small target detection technologies are relatively large. Summary of the Invention

[0005] The purpose of the present invention is to provide a lightweight detection method for panoramic images to solve the problems existing in the above-mentioned prior art.

[0006] To achieve the above purpose, the present invention provides a lightweight detection method for panoramic images, including:

[0007] Obtain real panoramic video data, and obtain an image dataset based on the real panoramic video data;

[0008] Perform image feature annotation on the image dataset to obtain a target detection dataset;

[0009] Build a lightweight object detection model, train the lightweight object detection model based on the object detection dataset, and obtain a panoramic image lightweight object detection model;

[0010] Perform object detection based on the panoramic image lightweight object detection model.

[0011] Preferably, the process of obtaining the image dataset includes:

[0012] Obtain the real panoramic video data based on a panoramic camera;

[0013] Process the real panoramic video data to obtain the image dataset.

[0014] Preferably, the process of obtaining the object detection dataset includes:

[0015] Establish object position labels based on the characteristics of panoramic image objects;

[0016] Perform image feature annotation on the image dataset based on the object position labels to obtain the object detection dataset.

[0017] Preferably, the process of building the lightweight object detection model includes:

[0018] Build a convolutional neural network model;

[0019] Build an attention module and a small object detection layer, introduce the attention module and the small object detection layer into the convolutional neural network model for optimization, and obtain an optimized model;

[0020] Build a lightweight convolutional module, add the lightweight convolutional module to the optimized model to obtain a detection model;

[0021] Optimize the network structure in the detection model to obtain the lightweight object detection model.

[0022] Preferably, the formula for calculating the number of parameters in the lightweight convolutional module is:

[0023]

[0024] where N m is the total number of lightweight convolutional operation parameters, K is the convolutional kernel size, C1 is the number of channels of the input feature map, and C2 is the number of channels of the output feature map.

[0025] Preferably, the network structure in the detection model includes a backbone network and a head network;

[0026] The backbone network includes an SE structure, an ECBS structure, an EELAN structure, and an EMP1 structure;

[0027] The head network includes an SE structure, an FPN structure, a PAN structure, an EMP2 structure, an EELAH-H structure, an ESPPCCSPC structure, and a small target detection layer of an EREP structure.

[0028] Preferably, the process of obtaining the lightweight object detection model for panoramic images includes:

[0029] Select a training set and a test set from the object detection dataset, and configure the network training environment at the same time;

[0030] Based on the network training environment, input the training set into the lightweight object detection model for training to obtain a trained model;

[0031] Input the test set into the trained model for testing to obtain the lightweight object detection model for panoramic images.

[0032] The technical effects of the present invention are:

[0033] The present invention constructs a dataset that conforms to the characteristics of panoramic images based on the real collected images of airport panoramic cameras. The YOLOv7 object detection framework includes a backbone network and a head network. The present invention adds an SE attention module to the backbone network and the head network, adds a small target detection layer to the head network, and uses a lightweight convolutional network to replace the original convolutional network to lightweight the overall network architecture, thereby constructing a lightweight object detection method for panoramic images. Its advantages are: (1) Solve the problem of lacking data in the research on the characteristics of panoramic images. (2) Solve the problem that the targets in panoramic images are too small and the detection is inaccurate. (3) Solve the problem of too large model parameter quantity. Description of the Drawings

[0034] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0035] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0036] Figure 1 It is a flowchart of the lightweight detection method for panoramic images in the embodiment of the present invention;

[0037] Figure 2 It is the network structure of the SE module in the embodiment of the present invention;

[0038] Figure 3 It is the EConv structure in the embodiment of the present invention;

[0039] Figure 4 This is the panoramic image lightweight object detection network structure in the embodiments of the present invention;

[0040] Figure 5 This is the schematic diagram of the ECBS basic convolution structure in the embodiments of the present invention;

[0041] Figure 6 This is the schematic diagram of the EELAN structure in the embodiments of the present invention;

[0042] Figure 7 This is the schematic diagram of the EMP structure in the embodiments of the present invention;

[0043] Figure 8 This is the schematic diagram of the EELAN-H structure in the embodiments of the present invention;

[0044] Figure 9 This is the schematic diagram of the ESPPCCSPC structure in the embodiments of the present invention;

[0045] Figure 10 This is the schematic diagram of the EREP structure in the embodiments of the present invention. Detailed implementation manners

[0046] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0047] It should be noted that the steps shown in the flowchart of the drawings may be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0048] Embodiment 1

[0049] As Figure 1 shown, this embodiment provides a lightweight detection method for panoramic images, including:

[0050] Step 1: Collect real panoramic video data from the airport panoramic camera to obtain an image dataset;

[0051] Step 2: Label the image features, label the airplanes and vehicles in each picture in the dataset, construct an object detection dataset, and divide the training set and the test set according to 7:3;

[0052] Step 3: Construct a panoramic image lightweight object detection model that improves YOLOv7;

[0053] Step 4: Train the panoramic image lightweight object detection model constructed in Step 3;

[0054] Step 5: Use the trained lightweight object detection model for panoramic images of the improved YOLOv7 to perform object detection on the test set and verify through object detection evaluation metrics;

[0055] Step 6: Write the obtained weight parameters during training into detect.py, use Python to build a framework to run the improved model, and verify that the lightweight object detection model for panoramic images of the improved YOLOv7 is more accurate and meets the requirements of lightweight object detection for panoramic images.

[0056] In this embodiment, by improving the YOLOv7 framework, a lightweight object detection method for panoramic images is constructed. An SE attention module is added to the backbone network and the head network, a small object detection layer is added to the head network, and the original convolutional network is replaced with a lightweight convolutional network to lightweight the overall network architecture, thus constructing a lightweight object detection method for panoramic images. The established airport panoramic camera is used to collect a real panoramic video data dataset to verify the training network model, and verification is carried out through object detection evaluation metrics.

[0057] For the further optimization scheme, first, a panoramic image dataset is established. The experimental data comes from the real data collected by the airport panoramic camera. In this embodiment, common objects in the airport panoramic images are used, mainly including airplanes and vehicles, etc. An object label dataset is established according to the characteristics of the panoramic image objects. The labeled categories of the object position labels are: two categories of airplanes and cars, and the position label format is the YOLO format label of the center point plus height and width. The position information labels are used to label the dataset with the annotation program Labelimg. A total of 100 pictures are labeled, with a total of 8660 labeled objects. According to the definition of small objects, when the target occupies less than 0.1 of the image size, the target is considered a small object. In this dataset, there are 8648 small objects, accounting for 99.8% of the dataset.

[0058] For the further optimization scheme, such as Figure 2As shown, an SE attention module is constructed. The attention network mechanism enables the model to be more interested in a certain type of target, reduces useless information to enhance the detection ability, and thus improves the network's detection ability for small target objects in panoramic images. The SE (Squeeze and Excitation) attention module is a method proposed in recent years to further improve the performance of convolutional neural networks (CNNs). The SE attention module provides a method to adaptively adjust the channel weights of feature maps, enabling the information in specific channels to receive more attention during the feature extraction process, thereby effectively improving the performance of convolutional neural networks. The SE attention module is implemented by using two fully connected layers, one called the Squeeze operation and the other called the Excitation operation. The Squeeze operation resizes the feature map to a size of 1×1×C channels and then performs a non-linear transformation through the activation function ReLU. This generates a channel descriptor that condenses all the feature map activations of a given channel. The Excitation operation feeds the channel descriptor into two fully connected layers and uses the Sigmoid function and the ReLU activation function to generate a channel weight vector. This vector is multiplied by the feature map to control the activation degree of each channel in the feature map, that is, each element of the vector is used to control the weight of the corresponding channel. By introducing the SE attention module, convolutional neural networks can better focus on the information in different channels, reduce the weights of useless features, and perform feature modeling more accurately.

[0059] Next, a small target detection layer is added. Before the original image is sent to the feature detection network in the Yolov7 framework, it undergoes downsampling by 8 times, 16 times, and 32 times, resulting in a large target detection feature map of 20*20, a medium target feature detection map of 40*40, and a small target feature detection map of 80*80. For small targets in the YOLOv7 object detection algorithm, due to their small size, they are often ignored or missed by the algorithm, affecting the detection accuracy and recall rate. All the detected targets in the panoramic image are small targets, and their pixel features are similar to the background, so the object detection algorithm may ignore or miss them. Based on the original algorithm, we add a new scale feature map. The original image undergoes downsampling by 4 times, 8 times, 16 times, and 32 times, resulting in a large target detection feature map of 20*20, a medium target feature detection map of 40*40, a small target feature detection map of 80*80, and a small target feature detection map of 160*160, as follows: Upsample the feature fusion EELAN-H layer and then fuse it with the first EELAN layer in the backbone network to generate a new feature map. The generated feature map passes through EELAN to obtain a larger feature layer. This feature layer serves as the new head detection layer for small target detection and also undergoes downsampling operations to integrate shallow detail features into the neck feature fusion network structure. The four detection layers improve the detection accuracy and recall rate of small targets in the panoramic image, making the algorithm more applicable in panoramic images.

[0060] Then, a lightweight convolution module EConv is constructed. The idea of the lightweight convolution module is an operation applied to a convolutional network with lightweight parameters and computational complexity. It transforms the conventional convolution operation into a process of stacking multiple lightweight convolutions to achieve lightweight and high-performance convolution processing. When an image is input into the network in matrix form, the conventional convolution operation can be regarded as a stacking of filtering operations. The calculation of the number of parameters in the convolution operation is shown in Equation 1. The lightweight convolution module used in this paper consists of three convolution operations with a convolution kernel size of 1×1 to generate a convolution channel of 3×C2 / 4 and one convolution operation with a convolution kernel of K to generate a convolution channel of C2 / 4, and then uses the Concat operation to output a feature map with a channel number of C2. The structure is as Figure 3 shown. The calculation formula for its number of parameters is shown in Equation 2. The reduction in the number of lightweight convolution parameters is shown in Equation 3

[0061] N c = K×K×C1×C2 + C2(1)

[0062]

[0063]

[0064] In the formula, N c is the total number of parameters for the conventional convolution operation, N mLet \(K\) be the size of the convolutional kernel, \(C1\) be the number of channels of the input feature map, and \(C2\) be the number of channels of the output feature map. The total number of parameters for lightweight convolution operations is

[0065] Overall, the improved YOLOv7 network structure is the same as the YOLOv7 network framework. The network consists of a backbone network and a head network. The overall network structure is as Figure 4 shown.

[0066] The backbone network includes an SE structure, an ECBS structure (as Figure 5 shown), an EELAN structure (as Figure 6 shown), and an EMP1 structure (as Figure 7 shown). Among them, the ECBS structure consists of a lightweight convolution module, a normalization module, and a SiLU activation function. The EELAN structure is based on the original ELAN structure in YOLOv7 and uses lightweight network modules. The EMP1 structure is composed of a max-pooling module and a convolution with a stride of 2, and is reconstructed by two downsampling methods to enhance the network's learning ability without destroying the structure.

[0067] The head network includes an SE structure, an FPN structure, a PAN structure, an EMP2 structure, an EELAH-H structure (as Figure 8 shown), an ESPPCCSPC structure (as Figure 9 shown), and an EREP structure (as Figure 10 shown) for small object detection layers. Among them, the FPN structure is a top-down feature pyramid that uses upsampling to improve the ability to detect small objects. The PAN is a bottom-up feature pyramid that uses the information of the lower layer to be passed to the upper layer to improve the ability to detect occluded objects. The input feature map of the ESPPCCSPC structure passes through three ECBS operations, and the three max-pooling operations are combined using the concat operation, and then passed through two ECBS operations and combined with the input using the concat operation. According to the characteristics that most objects in panoramic images are small objects, a new scale feature map is added on the basis of the original algorithm. The original image is downsampled by 4 times, 8 times, 16 times, and 32 times to obtain a 20×20 large object detection feature map, a 40×40 medium object feature detection map, an 80×80 small object feature detection map, and a 160×160 small object feature detection map, and then sent into the detection network. The improvement of the network structure for the target features of panoramic images in this paper is beneficial to improving the comprehensive performance of the lightweight object detection method for panoramic images.

[0068] The present invention verifies the lightweight object detection algorithm for panoramic images for the above-mentioned constructed panoramic image dataset and improved YOLOv7 object detection algorithm. The specific steps are as follows:

[0069] Step1: Configure the network training configuration file

[0070] Select 70 images from the dataset as the training set and 30 images as the test set. Before training, it is necessary to modify the improved YOLOv7 data and model configuration files. Change the number of object categories to 3 in the data file, and modify the object category names in the category list names. Increase the feature scale and attention module according to the improved network structure. Add the SE model code to the common.py file and introduce the SE structure in the yolo.py file. Run the yolo.py file to check whether the network changes are correct. Use the K-means clustering algorithm to calculate the anchor box size of the model.

[0071] Step2: Configure the network training environment

[0072] The experimental environment of this test is the Dawning cloud computing service system. Each node is configured with 1 x86 processor with 32 cores and a main frequency of 2.5GHz and 1 NVIDIA Tesla V100 accelerator card. Each node is configured with 2 16GB DDR4 2666 ECC REG memories, and two sets of Dawning Parastor300S parallel storage systems are configured to provide large-capacity data storage. In terms of network communication, the cluster adopts a full-line speed and non-blocking 200Gb HDR Infiniband dedicated computing network and uses the pytorch1.9.0 deep learning framework. The specific configuration information is shown in Table 1-3.

[0073] Table 1

[0074]

[0075] Table 2

[0076]

[0077] Table 3

[0078]

[0079] Then, it is verified based on the object detection evaluation metrics:

[0080] Step1: Selection of object detection evaluation metrics

[0081] All models are trained and tested using the panoramic image data established in this embodiment. In the experiment, Recall, Precision, F1, mAP, and the number of model parameters (Parameters) are used as the evaluation metrics of the model, and the threshold of IOU is set to 0.5. F1 is the harmonic mean of Recall and Precision, which can give a more accurate response to the model. mAP is the average value of the average accuracy of multiple targets under different Recall conditions. The definitions of Recall, Precision, F1, and mAP are as follows:

[0082]

[0083]

[0084]

[0085]

[0086] Among them, TP represents the number of correctly recognized positive samples, FP represents the number of misrecognized positive samples, FN represents the number of missed recognized positive samples, and m represents the number of recognition categories.

[0087] Step2: Analysis of experimental results

[0088] To verify the effectiveness of the improved object detection algorithm proposed by the present invention, the experimental results are shown in Table 4 below:

[0089] Table 4

[0090]

[0091]

[0092] The resolution of the network input image is 640×640×3, and it is trained for 5000 epochs. The experimental comparison results of the detection accuracies of YOLOv3, YOLOv4, YOLOv5l, YOLOv7 and the lightweight detection method of the present invention are shown. The experimental results show that the lightweight object detection method for panoramic images proposed by the present invention reduces the number of network parameters by 48.5% and has the best performance in the evaluation indexes of the present invention. The mAP0.5 of the present invention is improved by 20.7% compared with YOLOv3, by 12.7% compared with YOLOv4, by 10.1% compared with YOLOv5l, and by 16.7% compared with YOLOv7. The detection results of panoramic images using YOLOv3, YOLO, YOLOv5l, YOLOv7 and the object detection algorithm of the present invention show that the YOLO algorithm fails to detect the target, while the improved YOLOv7 can detect the target with higher accuracy.

[0093] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A lightweight detection method for panoramic images, characterized in that, It includes the following steps: Obtain real panoramic video data, and obtain an image dataset based on the real panoramic video data; Perform image feature annotation on the image dataset to obtain a target detection dataset; Construct a lightweight target detection model, and train the lightweight target detection model based on the target detection dataset to obtain a panoramic image lightweight target detection model; Perform target detection based on the panoramic image lightweight target detection model; The process of constructing the lightweight target detection model includes: Construct a convolutional neural network model; Construct an attention module and a small target detection layer, and introduce the attention module and the small target detection layer into the convolutional neural network model for optimization to obtain an optimized model; Construct a lightweight convolutional module, and add the lightweight convolutional module to the optimized model to obtain a detection model; Optimize the network structure in the detection model to obtain the lightweight target detection model; The parameter calculation formula in the construction of the lightweight convolutional module is: ; Among them, is the total amount of lightweight convolution operation parameters, K is the convolution kernel size, C1 is the number of channels of the input feature map, and C2 is the number of channels of the output feature map; The network structure in the detection model includes a backbone network and a head network; The backbone network includes an SE structure, an ECBS structure, an EELAN structure, and an EMP1 structure; The head network includes an SE structure, an FPN structure, a PAN structure, an EMP2 structure, an EELAH-H structure, an ESPPCCSPC structure, and a small target detection layer of an EREP structure; The FPN structure uses a top-down feature pyramid to improve the small target detection ability by the top-down method. The PAN uses a bottom-up feature pyramid to transfer the lower-layer information to the upper layer to improve the detection ability of occluded targets. The input feature map of the ESPPCCSPC structure undergoes three ECBS operations, and the three maximum pooling operations are combined using the concat operation, and then undergo two ECBS operations and are combined with the input using the concat operation.

2. The lightweight detection method for panoramic images according to claim 1, wherein The process of obtaining the image dataset includes: Obtain the real panoramic video data based on a panoramic camera; Process the real panoramic video data to obtain the image dataset.

3. The lightweight detection method for panoramic images according to claim 1, wherein The process of obtaining the target detection dataset includes: Establish a target position label based on the target features of the panoramic image; Perform image feature annotation on the image dataset based on the target position label to obtain the target detection dataset.

4. The lightweight detection method for panoramic images according to claim 1, wherein The process of obtaining the panoramic image lightweight target detection model includes: Select a training set and a test set from the target detection dataset, and configure the network training environment at the same time; Input the training set into the lightweight target detection model for training based on the network training environment to obtain a training model; Input the test set into the training model for testing to obtain the panoramic image lightweight target detection model.

Citation Information

Patent Citations

  • Real-time identification method of live panoramic traffic signs

    CN109325438A

  • Lightweight deep convolutional neural network model based on expansion convolution

    CN110490298A

  • Rapid pest detection method based on improved YOLO V4

    CN114220035A

  • Deep learning-based complex road lane recognition method and chip

    CN115294545A