A method for fish-eye image object detection based on improved YOLOv5

By improving the YOLOv5 network model, combining the initial correction and interpolation technology of fisheye images, a hybrid attention mechanism and CIoU loss function are introduced, which solves the problems of low accuracy and long-term detection of fisheye images, and achieves efficient and accurate small object detection.

CN115187768BActive Publication Date: 2025-07-01XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210610350.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-07-01
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

In the prior art, when processing fisheye images, there are problems such as low detection accuracy, long time consumption and difficulty in detecting small objects, especially in high resolution correction images.

Method used

Using the fisheye image object detection method based on improved YOLOv5, an improved YOLOv5 network model is constructed by initial correction and interpolation of the original fisheye image, and a hybrid attention mechanism is introduced into the backbone network, using CIoU as the loss function of border regression.

Benefits of technology

The detection accuracy and efficiency of fisheye images are improved, the blur of the corrected image is reduced, the detection ability of small objects under high-resolution fisheye images is enhanced, and the accuracy and robustness of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115187768B_ABST
    Figure CN115187768B_ABST
Patent Text Reader

Abstract

The present invention discloses a fish-eye image target detection method based on improved YOLOv5, which includes: initially correcting the original fish-eye image to obtain the longitude and latitude corresponding to each pixel point of the corrected image; performing interpolation based on the longitude and latitude corresponding to each pixel point of the corrected image to obtain the pixel value of each pixel point in the corrected image; constructing an improved YOLOv5 network model and training the network model; inputting the corrected image of the original fish-eye image into the trained improved YOLOv5 network model to obtain the target information in the picture. The fish-eye image target detection method based on improved YOLOv5 according to the present invention combines the longitude and latitude of each pixel point, introduces a hybrid attention mechanism into the backbone network, and simultaneously uses CIoU as the bounding box regression loss function, which can effectively improve the phenomenon of difficult small target detection under high-resolution corrected images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image target detection, and particularly relates to a fish-eye image target detection method based on improved YOLOv5. Background Art

[0002] Due to the booming development of computer technology and the continuous replacement of hardware resources, the field of deep learning has witnessed unprecedented development. The target detection technology driven by deep learning is also advancing rapidly. The target detection technology based on neural networks has shown amazing performance compared with the traditional target detection technology based on features. Although the target detection technology based on neural networks is relatively mature, most of its detection objects are normal images without distortion. For detection objects with serious distortion such as fish-eye images, the research is not yet perfect.

[0003] Target detection has always been one of the most cutting-edge directions in the field of computer vision research. Its purpose is to detect specific target objects and their specific position coordinates in a real scene or an input image, and at the same time assign corresponding class labels to each detected target object. The target detection technology has a very wide range of applications and has attracted great research enthusiasm in both the industrial and scientific research fields. It is mainly divided into two methods, namely the traditional image feature-based method and the deep learning-based method.

[0004] The traditional method based on image features is mainly implemented through image features, sliding windows, and support vector machines (SVM). The implementation process is as follows: First, an image feature is designed for subsequent steps. The sliding window method is used on the input image to achieve the effect of feature extraction, and then classification and regression operations are performed. Commonly used image features in traditional target detection methods include Histogram of Oriented Gradients (HOG), Scale Invariant Feature Transform (SIFT), and Harr-like features. Due to the large number of redundant operations in the detection process of traditional target detection algorithms based on image features and the poor robustness of manually designed image features, they are gradually being replaced by deep learning-based target detection methods.

[0005] Object detection methods based on deep learning are the mainstream of current research, which are usually divided into two categories. One is the two-stage object detection algorithm, and the other is the one-stage object detection algorithm. Among them, the two-stage algorithm is also called the object detection algorithm based on candidate regions, which is completed in two steps. First, a large number of candidate regions are selected in the input image, and then they are sent to the network to determine the category through a classifier. Its representative algorithms are mainly the R-CNN series. The core idea of the one-stage detection algorithm is that object detection can be achieved only by extracting features once. Compared with the R-CNN series algorithms, the detection accuracy of this type of algorithm has decreased, but it has a higher detection speed. Its representative algorithms are mainly SSD and YOLO series. Since the resolution of the image after fisheye image distortion correction is relatively high, the object detection task for a single image is relatively more time-consuming, and it is difficult to detect small objects in the high-resolution corrected image, and the accuracy is relatively low. Summary of the Invention

[0006] To solve the above problems existing in the prior art, the present invention provides a fisheye image object detection method based on improved YOLOv5. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0007] One aspect of the present invention provides a fisheye image object detection method based on improved YOLOv5, including:

[0008] S1: Perform initial correction on the original fisheye image to obtain the longitude and latitude corresponding to each pixel point of the corrected image;

[0009] S2: Perform interpolation based on the longitude and latitude corresponding to each pixel point of the corrected image to obtain the pixel value of each pixel point in the corrected image;

[0010] S3: Construct an improved YOLOv5 network model and train the network model. Among them, the improved YOLOv5 network model includes an input unit, a backbone network, a Neck unit, and a Prediction unit. The backbone network includes a Focus module, a convolutional module, a hybrid attention module, a convolutional module, a hybrid attention module, a convolutional module, a hybrid attention module, a convolutional module, a spatial pyramid pooling module, and a hybrid attention module connected in sequence. Among them, the hybrid attention module is used to obtain the hybrid attention feature of the channel domain feature and the spatial domain feature of the input feature vector;

[0011] S4: Input the corrected image of the original fisheye image into the trained improved YOLOv5 network model to obtain the target information in the picture.

[0012] In an embodiment of the present invention, the S1 includes:

[0013] S1a: Extract the valid region from the original fish-eye image to obtain the valid region image;

[0014] S1b: Perform preliminary correction on the valid region image according to the longitude-latitude mapping model to obtain the corrected rectangular image, and obtain the longitude α and latitude β of each pixel point in the corrected rectangular image.

[0015] In an embodiment of the present invention, the S1b includes:

[0016] Obtain the relationship between the coordinates (x, y) of each pixel point in the corrected rectangular image and its longitude α and latitude β:

[0017]

[0018] where x and y respectively represent the horizontal and vertical coordinates of the corresponding point on the corrected rectangular image, α and β are the longitude and latitude of the point (x, y), a is the length of the corrected rectangular image, and b is the width of the corrected rectangular image;

[0019] According to the corresponding relationship of the coordinates of the same pixel point before and after image correction, use the coordinates of the corrected rectangular image and its corresponding longitude α and latitude β to obtain the floating-point coordinate data corresponding to the mapped uncorrected fish-eye image:

[0020]

[0021] where x' and y' represent the floating-point coordinates corresponding to the point on the uncorrected valid region image, and r0 is the radius of the uncorrected valid region image.

[0022] In an embodiment of the present invention, the S2 includes:

[0023] S2a: Introduce a direction vector according to the longitude and latitude (α, β) of the floating point P obtained by mapping The direction vector represents the direction of the floating point P corresponding to the point on the corrected rectangular image after mapping, and is expressed as:

[0024]

[0025] where α and β are the longitude and latitude of the floating point P, and y represents the slope of the direction vector of;

[0026] S2b: Obtain the four nearest integer points P11, P12, P21, and P22 around the floating point P, and use the pixel values of the integer points P11, P12, P21, and P22 to obtain the pixel value of the floating point P by using the bilinear interpolation algorithm based on longitude and latitude;

[0027] S2c: According to the pixel value of each floating point and the corresponding relationship between the corrected rectangular image coordinate points and the floating point, obtain the pixel value of each point in the corrected image, so as to form the corrected rectangular image.

[0028] In one embodiment of the present invention, the hybrid attention module includes a channel attention sub-module and a spatial attention sub-module, wherein,

[0029] The channel attention sub-module is used to obtain the channel domain feature M of the input feature vector c , and the spatial attention sub-module is used to obtain the spatial domain feature M of the input feature vector S , and the hybrid attention module is used to fuse the channel domain feature M c and the spatial domain feature M S to obtain the hybrid attention feature F":

[0030]

[0031]

[0032] wherein, F ∈ R C*H*W represents the feature map input to the hybrid attention module, C represents the number of channels of the input feature map, H represents the height of the input feature map, W represents the width of the input feature map, M c ∈ R C*1*1 , M S ∈ R 1*H*W .

[0033] In one embodiment of the present invention, training the improved YOLOv5 network model includes:

[0034] Construct a training data set: Obtain a large number of original fish-eye images for initial correction, obtain the corrected rectangular image corresponding to each original fish-eye image, and label the true box of the target for each image in the training data set to form a training data set;

[0035] Input the pictures in the training data set into the improved YOLOv5 network model in turn, and obtain the predicted box position and the category to which the target object belongs through forward propagation;

[0036] Use the CIoU function as the loss function for bounding box regression, and update the weight matrix and bias in the improved YOLOv5 network model through gradient descent iteration to reduce the difference between the predicted box and the true box;

[0037] Use the training data set to perform multiple iterative trainings on the network model, and finally obtain the optimal network model when the loss function takes the minimum value.

[0038] In one embodiment of the present invention, as the CIoU loss function for bounding box regression, it includes:

[0039] Calculate the IOU value between the predicted bounding box and the ground truth bounding box:

[0040]

[0041] where P represents the area of the predicted bounding box, and G represents the area of the ground truth bounding box;

[0042] If the IOU value is less than the set threshold, then the CIoU loss function is obtained as:

[0043]

[0044] where c represents the diagonal distance of the minimum enclosing box of the predicted bounding box and the ground truth bounding box, and d represents the distance between the center points of the predicted bounding box P and the ground truth bounding box G;

[0045] If the IOU value is greater than or equal to the set value, then the CIoU loss function is obtained as:

[0046]

[0047] where,

[0048]

[0049]

[0050] where G w and G h respectively represent the width and height of the ground truth bounding box, and P w and P h respectively represent the width and height of the predicted bounding box.

[0051] Another aspect of the present invention provides a storage medium, in which a computer program is stored, and the computer program is used to execute the steps of the fish-eye image target detection method based on the improved YOLOv5 described in any one of the above embodiments.

[0052] Another aspect of the present invention provides an electronic device, which includes a memory and a processor. When the processor calls the computer program stored in the memory, it implements the steps of the fish-eye image target detection method based on the improved YOLOv5 described in any one of the above embodiments.

[0053] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0054] 1. The object detection method for fisheye images based on improved YOLOv5 of the present invention performs interpolation work after rough correction of longitude and latitude for each pixel point, which can better consider the direction factors of each pixel during the fisheye image correction process, thereby reducing the blurriness of the corrected image, achieving high-quality generation of the corrected image, and the higher the quality of the finally generated image, the more conducive to further work on this basis.

[0055] 2. The present invention introduces a hybrid attention mechanism into the backbone network Backbone, enabling the feature map tensor to obtain attention features in both the spatial domain and the channel domain, improving the focusing ability on small target objects on the horizontal line under high-resolution fisheye images.

[0056] 3. The present invention uses CIOU as the bounding box regression algorithm, achieving a balance between the center point position and the aspect ratio matching situation, enabling the loss function to maintain a good convergence speed at any position, effectively improving the problem of the penalty term degradation of the original loss function, and enhancing the accuracy and robustness of the model.

[0057] The following will further elaborate on the present invention in conjunction with the drawings and embodiments. Description of the Drawings

[0058] Figure 1 is the overall flowchart of an object detection method for fisheye images based on improved YOLOv5 provided by an embodiment of the present invention;

[0059] Figure 2 is the flowchart of a fisheye correction algorithm provided by an embodiment of the present invention;

[0060] Figure 3 is a schematic diagram of the effective area extraction process provided by an embodiment of the present invention;

[0061] Figure 4 is a schematic diagram of an original fisheye image and a corrected rectangular image provided by an embodiment of the present invention.

[0062] Figure 5 is a schematic diagram of a bilinear interpolation algorithm based on longitude and latitude provided by an embodiment of the present invention;

[0063] Figure 6 is a schematic diagram of the improvement of the Backbone part of the YOLOv5s network provided by an embodiment of the present invention;

[0064] Figure 7 is an overall view of a hybrid attention module provided by an embodiment of the present invention;

[0065] Figure 8 is a flowchart of the processing process of a hybrid attention module provided by an embodiment of the present invention;

[0066] Figure 9 It is a schematic structural diagram of a channel attention sub-module provided by an embodiment of the present invention;

[0067] Figure 10 It is a schematic structural diagram of a spatial attention sub-module provided by an embodiment of the present invention;

[0068] Figure 11 It is a schematic diagram of a CIOU bounding box regression algorithm provided by an embodiment of the present invention;

[0069] Figure 12 It is a fisheye correction image dataset collected in an embodiment of the present invention;

[0070] Figure 13 It is a comparison graph of PR curves of a traditional algorithm and the object detection algorithm of the present invention under the use of a pre-trained model;

[0071] Figure 14 It is a comparison graph of PR curves of a traditional algorithm and the object detection algorithm of the present invention without using a pre-trained model;

[0072] Figure 15 It is a comparison graph of experimental results of a traditional algorithm and the object detection algorithm of the present invention. Detailed implementation manners

[0073] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following provides a detailed description of a fisheye image object detection method based on improved YOLOv5 proposed according to the present invention in combination with the accompanying drawings and specific implementation manners.

[0074] The foregoing and other technical contents, features, and effects of the present invention can be clearly presented in the following detailed description in conjunction with the accompanying drawings. Through the description of the specific implementation manners, a more in-depth and specific understanding of the technical means and effects adopted by the present invention to achieve the predetermined purpose can be obtained. However, the accompanying drawings are only for reference and illustration, and are not used to limit the technical solutions of the present invention.

[0075] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, so that an article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of another identical element in the article or device including the said element.

[0076] Please refer to Figure 1 , Figure 1 which is the overall flowchart of a fisheye image target detection method based on improved YOLOv5 provided by an embodiment of the present invention. The fisheye image target detection method of this embodiment includes:

[0077] S1: Perform initial correction on the original fisheye image and obtain the longitude and latitude corresponding to each pixel point of the corrected image.

[0078] Please refer to Figure 2 and Figure 3 , the high-quality correction process of the fisheye image in this embodiment specifically includes the following steps:

[0079] S1a: Extract the effective region in the original fisheye image to obtain an effective region image.

[0080] Specifically, first adjust the original fisheye image to a grayscale image, set a grayscale threshold T, and perform grayscale detection in the four directions of up, down, left, and right. Specifically, start from the four sides of the original fisheye image and use four approximation lines parallel to the corresponding sides to approach the middle of the image. At the same time, detect the grayscale values of the pixel points on each approximation line. When there is a pixel point greater than the above grayscale threshold T on a certain approximation line, the approximation in the current direction ends, and use this approximation line as a boundary of the effective region. In this way, determine the four-sided boundaries of the effective region in the original fisheye image, so as to obtain the effective region image of the fisheye image after extraction. The effective region image obtained in this embodiment is a rectangular image. In the actual processing process, the value of the grayscale threshold T can be set according to the situation of the collected fisheye image.

[0081] Furthermore, obtain the radius r0 of this effective region image.

[0082] S1b: Perform preliminary correction on the effective region image according to the longitude-latitude mapping model to obtain a corrected rectangular image, and obtain the longitude α and latitude β where each pixel point of the corrected rectangular image is located.

[0083] Specifically, first obtain the relationship between the coordinates (x, y) of each pixel point of the corrected rectangular image and its longitude α and latitude β as follows:

[0084]

[0085] where x and y respectively represent the horizontal and vertical coordinates of the corresponding point on the corrected rectangular image, α and β are the longitude and latitude where the point (x, y) is located, a is the length of the corrected rectangular image, and b is the width of the corrected rectangular image. It should be noted that in practice, the length and width of the corrected rectangular image can be set according to needs.

[0086] Further, according to the corresponding relationship of the coordinates of the same pixel point before and after image correction, using the coordinates of the corrected rectangular image and its corresponding longitude α and latitude β, the floating-point coordinate data corresponding to the mapped uncorrected fisheye image is obtained:

[0087]

[0088] Among them, x′ and y′ represent the floating-point coordinates corresponding to the points on the uncorrected valid region image, and r0 is the radius of the uncorrected valid region image. Please refer to Figure 4 , Figure 4 which is a schematic diagram of an original fisheye image and a corrected rectangular image provided by an embodiment of the present invention.

[0089] S2: Interpolation is performed based on the longitude and latitude corresponding to each pixel point of the corrected image to obtain the pixel value of each pixel point in the corrected image.

[0090] It should be noted that the reverse mapping method is used in the correction process of step S1. However, in the actual mapping process, the integer pixel points in the corrected rectangular image will appear as floating-point numbers after mapping, and cannot exactly correspond to the integer pixel points of the fisheye image before correction. This floating-point number reflects the weight of the influence on the target pixel near the target pixel. This problem can be solved by the bilinear interpolation method.

[0091] Please refer to Figure 3 , Figure 3 which is a schematic diagram of a bilinear interpolation algorithm based on longitude and latitude provided by an embodiment of the present invention. The specific steps of interpolation based on longitude and latitude in this embodiment are as follows:

[0092] S2a: Interpolation is performed based on the longitude and latitude information of each pixel point. According to the longitude and latitude (α, β) of the target point P, the direction vector represents the direction of the floating-point P corresponding to the point on the corrected rectangular image after mapping. The equation of the direction vector of the target point P can be expressed as:

[0093]

[0094] where α and β are the longitude and latitude of the target point P.

[0095] S2b: Obtain the four nearest integer points P11, P12, P21, and P22 around the floating-point P, and use the pixel values of the four integer points P11, P12, P21, and P22 to obtain the pixel value of the floating-point P by using the bilinear interpolation algorithm based on longitude and latitude.

[0096] Specifically, please refer to Figure 5 , Figure 5It is a schematic diagram of a bilinear interpolation algorithm based on longitude and latitude provided by an embodiment of the present invention. The four nearest points P11, P12, P21, and P22 around the floating-point P are obtained. It should be noted that the target point P in the image after being corrected in step 1 is a floating-point coordinate, denoted as P(i + u, j + v), where i and j respectively represent the integer parts of the abscissa and ordinate of the target point P, and u and v respectively represent the decimal parts of the integer parts of the abscissa and ordinate of the target point P. The four nearest points around the target point P are respectively: P 11 (i, j), P 12 (i, j + 1), P 21 (i + 1, j) and P 22 (i + 1, j + 1). For example, assuming the coordinate of the target point P is (1.3, 1.4), the four nearest points around the target point P are respectively P11(1, 1), P12(1, 2), P21(2, 1), and P22(2, 2).

[0097] Subsequently, it is set that the two intersection points of the rectangle formed by the four points P 11 (i, j), P 12 (i, j + 1), P 21 (i + 1, j) and P 22 (i + 1, j + 1) and the vector are respectively A and B. The distances between point A and P11(i, j) as d A1 , the distance between point A and P 12 (i, j + 1) as d A2 , the distance between point B and P 21 (i + 1, j) as d B1 , the distance between point B and P 22 (i + 1, j + 1) as d B2 , the distance between point A and point P as d AP , and the distance between point B and point P as d BP .

[0098] Among them, the calculation formula for the distance d between two points is:

[0099]

[0100] where x1, y1 are the abscissa and ordinate of the first point, and x2, y2 are the abscissa and ordinate of the second point.

[0101] Interpolating from the pixel values of point P 11 and point P 12 , the pixel value of point A is obtained:

[0102]

[0103] Among them, f(A) represents the pixel value of point A, and P in the formula 11 represents the pixel value of point P 11 , P 12 represents the pixel value of point P 12 , and d A1 is the distance between point A and point P 11 ; d A2 is the distance between point A and point P 12 . It should be noted that since the horizontal and vertical coordinates of point P 11 and point P 12 are both integers, and their corresponding pixel values are all known, therefore, the pixel value of point A can be obtained by interpolation using the above formula.

[0104] Interpolate based on the pixel values of point P 21 and point P 22 to obtain the pixel value of point B:

[0105]

[0106] Among them, f(B) represents the pixel value of point B, and P in the formula 21 represents the pixel value of point P 21 , P 22 represents the pixel value of point P 22 , and d B1 is the distance between point B and point P 21 ; d B2 is the distance between point B and point P 22 . It should be noted that since the horizontal and vertical coordinates of point P 21 and point P 22 are both integers, and their corresponding pixel values are all known, therefore, the pixel value of point B can be obtained by interpolation using the above formula.

[0107] Interpolate based on the pixel values of point A and point B to obtain the pixel value of point P:

[0108]

[0109] Among them, d AP is the distance between point A and point P; d BP is the distance between point B and point P.

[0110] According to the above steps, the pixel value of each point in the corrected image can be obtained.

[0111] S2c: According to the pixel value of each floating point and the correspondence between the corrected rectangular image coordinate points and the floating points, obtain the pixel value of each point in the corrected image, thereby constructing the corrected rectangular image.

[0112] S3: Construct an improved YOLOv5 network model and train the improved YOLOv5 network model.

[0113] The improved YOLOv5 network model of this embodiment includes an input unit, a backbone network, a Neck unit, and a Prediction unit. Among them, the input unit adopts the Mosaic data augmentation method to enrich the data set. The Backbone unit is mainly used to extract features from the input image data. The role of the Neck unit is to further enhance the diversity and robustness of the features. The Prediction unit is mainly used to predict the image features and generate bounding boxes and target categories. This embodiment of the present invention mainly improves the network structure in the Backbone part and improves its bounding box regression loss function.

[0114] Please refer to Figure 6 , Figure 6 which is a schematic diagram of the improvement of the Backbone part of a YOLOv5s network provided by this embodiment of the present invention. In this embodiment, the Convolutional Block Attention Module (CBAM) is introduced into the backbone network. This hybrid attention module can enable the feature map tensor to implement the attention mechanism in both the spatial domain and the channel domain. The backbone network of this embodiment includes a Focus module, a convolutional module, a hybrid attention module, a convolutional module, a hybrid attention module, a convolutional module, a hybrid attention module, a convolutional module, a Spatial Pyramid Pooling module, and a hybrid attention module connected in sequence. Among them, the hybrid attention module is used to obtain the hybrid attention features of the channel domain features and the spatial domain features of the input feature vector.

[0115] Furthermore, please refer to Figure 7 , Figure 7 which is the overall structure diagram of a hybrid attention module provided by this embodiment of the present invention. This hybrid attention module mainly includes two sub-modules, namely the channel attention sub-module and the spatial attention sub-module, and these two sub-modules are serially connected. For an input feature map F ∈ R C*H*W (where C represents the number of channels of the input feature map, H represents the height of the input feature map, and W represents the width of the input feature map), this hybrid attention module combines the channel domain feature M c ∈ R C*1*1 and the spatial domain feature M S ∈ R 1*H*W to obtain the final hybrid attention feature F", and the processing process is as follows:

[0116]

[0117]

[0118] Among them, F represents the feature map input to this hybrid attention module.

[0119] Next, referring to Figures 8 to 10 together, after the input image is convolved, the number of channels is the same as the number of convolution kernels. Different channels have different effects on achieving the task, and even some channels will have negative effects. Therefore, the channel attention sub-module of this embodiment mainly realizes highlighting the features of the target by learning the weight relationship between different channels.

[0120] Specifically, the channel attention sub-module of this embodiment first compresses the input feature map, and then performs global max pooling and global average pooling on the channels respectively, encoding the entire spatial feature into a global feature. The role of global max pooling is to screen out the features with prominent targets, and the role of global average pooling is to screen out the global background information of the target. Then, the channel information is screened through the MLP (Multi-Layer Perceptron) shared network unit, the two outputs are added element-wise, and then activated using an activation function to obtain the features in the channel domain. The entire processing process of the channel attention sub-module is shown in the formula:

[0121]

[0122] Among them, F represents the feature map input to this channel attention sub-module, W1 ∈ R C×C / r and W0 ∈ R C / r×C represent the weight parameters in the MLP (Multi-Layer Perceptron), σ represents the Sigmoid activation operation, and r represents the reduction rate.

[0123] It should be noted that the channel attention sub-module directly performs global average pooling on the information within one channel, ignoring the local information within each channel. In fact, the importance of pixels at different coordinate positions within the same channel is not the same, and the spatial attention sub-module can solve this problem.

[0124] Specifically, please refer to Figure 8 and Figure 10 , the spatial attention sub-module of this embodiment performs pooling operations on each channel of the input feature using the max pooling layer and the average pooling layer to generate two features representing different information, and then concatenates them according to the channels. After passing through a 7×7 convolutional layer, it is activated to generate the spatial attention feature map M S . The entire processing process of the spatial attention sub-module is shown in the formula:

[0125]

[0126] Among them, F' represents the feature map input to this spatial attention sub-module, σ represents the Sigmoid activation operation, and 7*7 represents the size of the convolutional kernel.

[0127] Subsequently, the obtained channel attention feature and the spatial domain attention feature are concatenated and combined to obtain the final hybrid attention feature F". The backbone network of this embodiment mainly completes the work of feature extraction of the input picture, and then inputs the extracted features into the Neck unit to complete the feature fusion work, and finally inputs them into the Prediction unit to complete the prediction work and generate the prediction box for object detection.

[0128] In this embodiment, training the improved YOLOv5 network model includes:

[0129] Construct a training data set, obtain a large number of original fisheye images, perform initial correction according to steps S1 and S2 to obtain the corrected rectangular image corresponding to each original fisheye image, form a training data set Fisheye_detect, and then use the labeimg tool to perform object annotation on each picture in the training data set, that is, mark the true box of the target object.

[0130] Input the pictures in the training data set into the improved YOLOv5 network model in turn, and obtain the predicted box position and the category to which the target object belongs through forward propagation;

[0131] Adopt the CIoU function as the loss function for bounding box regression, and iteratively update the weight matrix and bias in the improved YOLOv5 network model through gradient descent to reduce the difference between the predicted box and the true box.

[0132] Specifically, please refer to Figure 11 , Figure 11 is a schematic diagram of a CIOU bounding box regression algorithm provided by an embodiment of the present invention. The dashed rectangular box is the minimum circumscribed box between the predicted box and the true box. The CIoU loss function is further improved on the basis of DIoU, considering the matching of the center point position and the aspect ratio.

[0133] Specifically, calculate the IOU value between the predicted box and the true box, as shown in the formula:

[0134]

[0135] Among them, P represents the area of the predicted box, and G represents the area of the true box.

[0136] If the IOU value obtained in step 5a is less than the set threshold, which is set to 0.5 in this embodiment, the formula of the CIoU loss function is as follows:

[0137]

[0138] Among them, c represents the diagonal distance of the dashed rectangular box, and d represents the distance between the center points of the predicted box P and the ground truth box G.

[0139] If the IOU value obtained in step 5a is greater than or equal to the set threshold, which is set to 0.5 in this embodiment, the formula of the CIoU loss function is as follows:

[0140]

[0141] Among them, γ is used to calculate the matching situation of the aspect ratio, and the definitions of γ and α are shown in the following formulas respectively:

[0142]

[0143]

[0144] Among them, G w and G h represent the width and height of the ground truth box respectively, and P w and P h represent the width and height of the predicted box respectively.

[0145] During the training process, the network model is iteratively trained multiple times using the training dataset, and finally the optimal network model when the loss function takes the minimum value can be obtained.

[0146] S4: Input the corrected image of the original fisheye image into the trained improved YOLOv5 network model to obtain the target information in the picture.

[0147] Specifically, input the rectangular image obtained by preliminarily correcting the original fisheye image to be detected through steps S1 and S2 into the optimal network model obtained in the above steps, and the information of the target to be detected in this picture can be obtained, that is, the position of the predicted box of the target object and the category to which it belongs.

[0148] The detection effect of the fisheye image target detection method based on the improved YOLOv5 of the present invention can be illustrated by the following detection experiments.

[0149] Please refer to Figure 12 , Figure 12 which is the fisheye correction image dataset collected in the embodiments of the present invention. The experiments of the embodiments of the present invention are all completed based on this dataset. This dataset contains 3588 pictures, including 5676 vehicle targets and 4745 person targets.

[0150] In the embodiments of the present invention, transfer training for 100 epochs was carried out using the traditional algorithm (the original YOLOv5s network) and the method of the embodiments of the present invention with and without using a pre-trained model. Please refer to Figure 13 and Figure 14 , where Figure 13 (a) and Figure 13 (b) respectively represent the PR curves (Precision-Recall curves) of the traditional algorithm and the object detection method of the embodiments of the present invention under the use of a pre-trained model. Among them, the abscissa is the recall rate R, and the ordinate is the precision rate P. Figure 13 In (a), the value of person@.5 of the traditional algorithm is 86.9%, and the value of car@.5 is 93.8%. The mAP value of the two is 90.4%. Figure 13 In (b), the value of person@.5 of the method of the embodiments of the present invention is 87.2%, and the value of car@.5 is 94.0%. The mAP value of the two is 90.6%. Figure 14 (a) and Figure 14 (b) respectively represent the PR curves of the traditional algorithm and the method of the embodiments of the present invention without using a pre-trained model. Figure 14 In (a), the value of person@.5 of the traditional algorithm is 74.1%, and the value of car@.5 is 90.6%. The mAP value of the two is 82.3%. Figure 14 In (b), the value of person@.5 of the algorithm of the present invention is 76.1%, and the value of car@.5 is 90.7%. The mAP value of the two is 83.4%. The initial learning rate set in this experiment is 0.001, and the decay coefficient is 0.0005. The results show that the method of the present invention can achieve a better performance level and has better performance than the traditional algorithm. Without using a pre-trained model, the mAP of the method of the present invention is increased by 1.1% compared with the traditional algorithm.

[0151] Please refer to Figure 15 , Figure 15 is the result comparison of two models trained by the traditional algorithm and the method of the present invention without using a pre-trained model in the detection task. Among them, Figure 15 The upper and lower pictures in (a) are the object detection results of the traditional algorithm model. Figure 15 The upper and lower pictures in (b) are the object detection results of the algorithm model of the present invention. It can be seen that the present invention can effectively detect dense targets in high-resolution fisheye images, and at the same time, the present invention can accurately identify small targets in high-resolution fisheye images.

[0152] In summary, the object detection method for fisheye images based on improved YOLOv5 in the embodiments of the present invention performs interpolation work after rough longitude and latitude correction by combining the longitude and latitude of each pixel point, which can better consider the direction factors of each pixel during the fisheye image correction process, thereby reducing the blurriness of the corrected image, achieving high-quality generation of the corrected image, and the higher the quality of the finally generated image, the more conducive to further work on this basis; introducing the hybrid attention mechanism into the backbone network Backbone enables the feature map tensor to obtain attention features in both the spatial domain and the channel domain, improving the focusing ability on small target objects on the horizontal line under high-resolution fisheye images.

[0153] In addition, the embodiments of the present invention adopt CIOU as the bounding box regression algorithm, achieving a balance between the center point position and the aspect ratio matching situation, enabling the loss function to maintain a good convergence speed at any position, effectively improving the problem of the penalty term degradation of the original loss function, and enhancing the accuracy and robustness of the model.

[0154] In several embodiments provided by the present invention, it should be understood that the devices and methods disclosed in the present invention can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0155] In addition, each functional module in the various embodiments of the present invention can be integrated in one processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0156] Another embodiment of the present invention provides a storage medium, in which a computer program is stored, and the computer program is used to execute the steps of the method for detecting fish-eye image targets based on the improved YOLOv5 in the above embodiment. Another aspect of the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the method for detecting fish-eye image targets based on the improved YOLOv5 as described in the above embodiment are implemented. Specifically, the integrated modules implemented in the form of software function modules can be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium and include several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0157] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for fish-eye image target detection based on improved YOLOv5, characterized in that, Including: S1: Perform initial correction on the original fisheye image to obtain the longitude and latitude corresponding to each pixel point of the corrected image; S2: Perform interpolation based on the longitude and latitude corresponding to each pixel point of the corrected image to obtain the pixel value of each pixel point in the corrected image; S3: Construct an improved YOLOv5 network model and train the network model. Among them, the improved YOLOv5 network model includes an input unit, a backbone network, a Neck unit, and a Prediction unit. The backbone network includes a Focus module, a convolutional module, a hybrid attention module, a convolutional module, a hybrid attention module, a convolutional module, a hybrid attention module, a convolutional module, a spatial pyramid pooling module, and a hybrid attention module connected in sequence. Among them, the hybrid attention module is used to obtain the hybrid attention feature of the channel domain feature and the spatial domain feature of the input feature vector; S4: Input the corrected image of the original fisheye image into the trained improved YOLOv5 network model to obtain the target information in the picture.

2. The fish-eye image target detection method based on improved YOLOv5 according to claim 1, characterized in that The S1 includes: S1a: Extract the effective region in the original fisheye image to obtain the effective region image; S1b: Perform preliminary correction on the effective region image according to the longitude and latitude mapping model to obtain the corrected rectangular image, and obtain the longitude α and latitude β where each pixel point of the corrected rectangular image is located.

3. The method for fish-eye image target detection based on improved YOLOv5 according to claim 2, wherein, The S1b includes: Obtain the relationship between the coordinates (x, y) of each pixel point of the corrected rectangular image and its longitude α and latitude β: Among them, x and y respectively represent the horizontal and vertical coordinates of the corresponding point on the corrected rectangular image, α and β are the longitude and latitude where the point (x, y) is located, a is the length of the corrected rectangular image, and b is the width of the corrected rectangular image; According to the corresponding relationship of the coordinates of the same pixel point before and after image correction, use the coordinates of the corrected rectangular image and its corresponding longitude α and latitude β to obtain the floating-point coordinate data corresponding to the mapped uncorrected fisheye image: Among them, x′ and y′ represent the floating-point coordinates corresponding to the points on the uncorrected effective region image, and r0 is the radius of the uncorrected effective region image.

4. The method for fisheye image object detection based on improved YOLOv5 according to claim 3, wherein The S2 includes: S2a: Introduce a direction vector with the longitude and latitude (α, β) of the floating-point P obtained according to the mapping The said direction vector represents the direction of the floating-point P corresponding to the point on the corrected rectangular image after mapping, expressed as: Among them, α and β are the longitude and latitude of floating point P, and y represents the slope of the direction vector; S2b: Obtain the four nearest integer points P11, P12, P21, and P22 around the floating point P, and use the pixel values of the integer points P11, P12, P21, and P22 to obtain the pixel value of the floating point P by using the bilinear interpolation algorithm based on longitude and latitude; S2c: According to the pixel value of each floating point and the corresponding relationship between the coordinate points of the corrected rectangular image and the floating point, obtain the pixel value of each point in the corrected image, thereby constructing the corrected rectangular image.

5. The method for fish-eye image target detection based on improved YOLOv5 according to claim 1, characterized in that, The hybrid attention module includes a channel attention sub-module and a spatial attention sub-module, where The channel attention sub-module is used to obtain the channel-domain feature M of the input feature vector c , and the spatial attention sub-module is used to obtain the spatial-domain feature M of the input feature vector S , and the hybrid attention module is used to fuse the channel-domain feature M c and the spatial-domain feature M S to obtain the hybrid attention feature F": where \(F\in\mathbb{R}\) C*H*W represents the feature map input to the hybrid attention module, \(C\) represents the number of channels of the input feature map, \(H\) represents the height of the input feature map, \(W\) represents the width of the input feature map, \(M\) c \(\in\mathbb{R}\) C*1*1 , \(M\) S \(\in\mathbb{R}\) 1*H*W .

6. The method for fish-eye image object detection based on improved YOLOv5 according to claim 1, characterized in that, Training the improved YOLOv5 network model includes: Construct a training data set: Obtain a large number of original fisheye images for initial correction, obtain the corrected rectangular image corresponding to each original fisheye image, and mark the true box of the target for each image in the training data set to form a training data set; Input the pictures in the training data set into the improved YOLOv5 network model in sequence, and obtain the predicted box position and category of the target object through forward propagation; The CIoU function is used as the loss function of the bounding box regression, and the weight matrix and the bias in the improved YOLOv5 network model are iteratively updated by gradient descent to reduce the difference between the predicted box and the real box; The network model is trained iteratively multiple times using the training data set, and finally an optimal network model is obtained when the loss function takes a minimum value.

7. The method for fish-eye image object detection based on improved YOLOv5 according to claim 6, wherein The CIoU loss function for bounding box regression includes: Calculate the IOU value between the predicted box and the true box: Among them, P represents the area of ​​the predicted box, and G represents the area of ​​the real box; If the IOU value is less than the set threshold, the CIoU loss function is obtained as: Where c represents the diagonal distance between the predicted box and the minimum bounding box of the real box, and d represents the distance between the predicted box P and the center point of the real box G; If the IOU value is greater than or equal to the setting, the CIoU loss function is obtained as: in, Among them, G w and G h represent the width and height of the ground truth box respectively, and P w and P h represent the width and height of the predicted box respectively.

8. A storage medium, characterized in that, The storage medium stores a computer program, which is used to execute the steps of the fisheye image target detection method based on improved YOLOv5 as described in any one of claims 1 to 6.

9. An electronic device, characterized in that, The invention comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the fisheye image target detection method based on the improved YOLOv5 are implemented as claimed in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Fisheye image target identification method and system

    CN111260539A

  • Small target detection algorithm based on improved YOLOv5

    CN114241548A