A Nighttime Target Detection Method Based on Feature Aggregation Network
Through the design of feature aggregation network, the accuracy and robustness of night object detection are improved, the problems of insufficient light and high noise in night environments are solved, and more efficient feature extraction and detection effects are achieved.
Patent Information
- Application Number
- CN202510570235.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The existing night object detection methods are insufficient in the case of insufficient light and noise in the night environment, and the detection accuracy and robustness are insufficient, and the data set labeling is difficult, making it difficult to meet the real-time requirements.
Using a feature aggregation network-based method, general feature extraction is performed by designing the backbone of feature aggregation, combining the aggregation feature extraction module, the two-layer feature aggregation module and the pyramid aggregation attention mechanism to further integrate features, and using the Neck part of the pyramid feature fusion to improve feature extraction efficiency and accuracy. At the same time, model training is guided through data augmentation technology and comprehensive loss function.
It improves the accuracy and robustness of object detection in night environments, can better adapt to different lighting conditions and noise environments, enhances the model's detection ability of targets at different scales, and improves the accuracy and real-timeness of detection results.
Smart Images

Figure CN120107564B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing and detection, and in particular, to a method for night target detection based on a feature aggregation network. Background Art
[0002] With the acceleration of the urbanization process and the continuous development of society, night activities and night security have become increasingly important. Night target detection has important application values in traffic management, public safety, military surveillance, etc. However, due to insufficient illumination in the night environment, the image quality is poor and the noise increases, resulting in a significant reduction in the performance of traditional target detection methods in the night environment. Therefore, how to improve the accuracy and robustness of night target detection has become an important research topic in the fields of computer vision and image processing.
[0003] In recent years, night target detection methods have been widely used in computer vision and image processing, and many scholars at home and abroad have studied this topic and achieved corresponding progress. With the rapid development of deep learning technology, deep learning-based target detection methods have shown great potential in night target detection. By constructing a deep neural network, the deep learning method can automatically extract high-level features of images, and has strong robustness and generalization ability. This method directly inputs night images and outputs target detection results, which can provide additional scene information for night scenes and improve the accuracy and robustness of target detection.
[0004] Although deep learning-based methods have shown excellent performance in night target detection, they still face some challenges. First, the illumination changes violently and there is more noise in the night environment, which poses higher requirements for the training and inference of deep learning models. Second, the annotation and collection of night target detection datasets are relatively difficult, and insufficient data volume may lead to overfitting of the model and poor generalization ability. In addition, application scenarios with high real-time requirements pose strict requirements on the computational efficiency of the model, and a balance needs to be found between model complexity and detection accuracy.
[0005] In summary, the development of the night target detection field requires the comprehensive application of various technical means to improve the accuracy and robustness of night target detection to meet the needs of practical applications. Summary of the Invention
[0006] According to the above-mentioned technical problems, a nighttime object detection method based on a feature aggregation network is provided. The present invention performs general feature extraction based on the backbone designed by feature aggregation, improves the efficiency and accuracy of feature extraction through the focus module, and uses convolutional feature extraction, double-layer feature aggregation module and pyramid aggregation attention mechanism for feature extraction. And a Neck based on pyramid feature fusion is designed to further fuse the initial features extracted from the backbone. The method of the present invention can effectively improve the effect and robustness of object detection in nighttime environments.
[0007] The technical means adopted by the present invention are as follows:
[0008] A nighttime object detection method based on a feature aggregation network, comprising:
[0009] S1. Obtain a nighttime object detection dataset, and randomly divide the dataset into a training set and a test set according to a ratio of 7:3;
[0010] S2. Use image preprocessing operations to augment the nighttime object detection dataset obtained in step S1;
[0011] S3. Select the images after the preprocessing operation in step S2, and input the selected images into the backbone designed by feature aggregation for general feature extraction, and extract high-resolution initial features, medium-resolution initial features and low-resolution initial features;
[0012] S4. Input the high-resolution initial features, medium-resolution initial features and low-resolution initial features obtained in step S3 into the Neck part based on pyramid feature fusion to further extract high-resolution intermediate features, medium-resolution intermediate features and low-resolution intermediate features with diversity and robustness;
[0013] S5. Input the high-resolution intermediate features, medium-resolution intermediate features and low-resolution intermediate features extracted in step S4 into the Head part designed based on convolutional layers to complete object detection prediction.
[0014] Furthermore, the method further includes:
[0015] S6. Design a comprehensive loss function using classification loss, bounding box regression loss, and confidence loss to constrain the detection process based on the feature aggregation network.
[0016] Furthermore, step S2 specifically includes:
[0017] S21. Adopt the Mosaic technique to randomly splice four images into a new image, and the formula is as follows:
[0018] ;
[0019] Among them, denotes randomly selecting four night images from the dataset, denotes arranging together according to a random layout, denotes the resulting image of the mosaic data augmentation method;
[0020] S22. Apply the RandomAffine technique to perform rotation, translation, scaling, and shearing operations. The formula is as follows:
[0021] ;
[0022] Among them, denotes the horizontal and vertical coordinates in the input original image, denotes the horizontal and vertical coordinates after the random affine transformation, denotes the parameters of the affine transformation matrix for controlling rotation, scaling, and shearing, denotes the translation parameter;
[0023] S23. Apply the MixUp technique to linearly mix two images and their labels in a certain proportion. The formula is as follows:
[0024] ;
[0025] ;
[0026] Among them, and denote randomly selecting two original images from the night object detection dataset, and denote and corresponding labels, denotes the mixing ratio randomly sampled from the uniform distribution between [0, 1];
[0027] S24. Apply the image blurring technique and select Gaussian blurring to obtain the blurred image. The formula is as follows:
[0028] ;
[0029] Among them, denotes the blurred image, denotes the normalization factor, denotes the image blurring radius, denotes the weight of the Gaussian kernel, denotes the convolution operation, Represents the original image randomly selected from the night dataset;
[0030] S25. Adopt the HSV color space enhancement technology to simulate different lighting conditions by changing the hue, saturation, and brightness of the image. The formula is as follows:
[0031] ;
[0032] ;
[0033] ;
[0034] Among them, Represents the hue, saturation, and brightness components of the original image; Represents the hue, saturation, and brightness components of the enhanced image; Represents the random adjustment parameter;
[0035] S26. Adopt the random horizontal flipping technology to increase the robustness of the model to mirror changes. The formula is as follows:
[0036] ;
[0037] Among them, Represents the image after the random horizontal flipping operation, Represents the width of the image.
[0038] Furthermore, the Backbone based on feature aggregation design in step S3 includes an aggregation (Focus) feature extraction module, a convolutional feature extraction module, a double-layer feature aggregation module, and a pyramid aggregation attention mechanism. Step S3 specifically includes:
[0039] S31. The Focus feature extraction module extracts efficient features from the input image, reorganizes the spatial information of the input image, reduces the width dimension and height dimension of the image by half through slicing operations and splicing operations, and at the same time increases the number of channels, so as to perform subsequent convolutional operations. Among them:
[0040] The formula for the slicing operation is as follows:
[0041] ;
[0042] ;
[0043] ;
[0044] ;
[0045] Among them, Represents the original image, and the size of the original image is (C , H , W ), where C is the number of channels, H is the height, W is the width, represents the 4 parts obtained by slicing the original image according to odd and even indices, The sizes of C , H / 2, W / 2);
[0046] The formula for the splicing operation is as follows:
[0047] ;
[0048] Among them, represents the splicing operation, represents the feature image after splicing, The sizes of C , H / 2, W / 2);
[0049] Perform a convolution operation on the feature image after the splicing operation. The formula is as follows:
[0050] ;
[0051] Among them, represents the output feature image after convolution, represents the convolution operation;
[0052] S32. Design the function of the convolutional feature extraction module. The formula is as follows:
[0053] ;
[0054] Among them, represents the function of the convolutional feature extraction module, represents the input variable of the function of the convolutional feature extraction module, represents the LeakyReLU activation function, represents the batch normalization operation;
[0055] S33. Design the function of the double-layer feature aggregation module. The formula is as follows:
[0056] ;
[0057] Among them, represents the function of the double-layer feature aggregation module, represents the input variable of the function of the double-layer feature aggregation module, represents the channel direction splitting operation;
[0058] S34. Design a function for the pyramid aggregation attention mechanism, and the formula is as follows:
[0059]
[0060] Among them, represents the pyramid aggregation attention mechanism function, represents the input variable of the pyramid aggregation attention mechanism function, represents the max pooling operation;
[0061] S35. Input the output feature image of step S31 into the convolutional feature extraction module function and the double-layer feature aggregation module function in sequence to extract the high-resolution initial features, and the formula is as follows:
[0062]
[0063] Among them, F F11 represents the extracted high-resolution initial features;
[0064] S36. Input the high-resolution initial features F11 extracted in step S35 into the convolutional feature extraction module function and the double-layer feature aggregation module function in sequence to extract the medium-resolution initial features, and the formula is as follows:
[0065]
[0066] Among them, F12 represents the extracted medium-resolution initial features;
[0067] S37. Input the medium-resolution initial features extracted in step S36 into the convolutional feature extraction module function and the pyramid aggregation attention mechanism function in sequence to extract the low-resolution initial features, and the formula is as follows:
[0068]
[0069] Among them, F13 represents the extracted low-resolution initial features.
[0070] Furthermore, the Neck based on pyramid feature fusion in step S4 includes a convolutional feature extraction module and a double-layer feature aggregation module. Step S4 specifically includes:
[0071] S41. Design a function for the convolutional feature extraction module, and the formula is as follows:
[0072] ;
[0073] S42. Design a function for the double-layer feature aggregation module, with the formula as follows:
[0074] ;
[0075] S43. Input the low-resolution initial features extracted in step S37 into the double-layer feature aggregation module function and the convolutional feature extraction module function in sequence to extract low-resolution residual features, with the formula as follows:
[0076]
[0077] where, represents the extracted low-resolution residual features;
[0078] S44. Input the medium-resolution initial features extracted in step S36 and the low-resolution residual features extracted in step S43 into the double-layer feature aggregation module function and the convolutional feature extraction module function to extract residual medium-resolution features, with the formula as follows:
[0079]
[0080] where, represents the extracted residual medium-resolution features, UP () represents the upsampling operation;
[0081] S45. Input the high-resolution initial features extracted in step S35 F 11 and the residual medium-resolution features extracted in step S44 into the double-layer feature aggregation module function to extract high-resolution intermediate features, with the formula as follows:
[0082]
[0083] where, represents the extracted high-resolution intermediate features;
[0084] S46. Input the residual medium-resolution features extracted in step S44 and the high-resolution intermediate features extracted in step S45 into the double-layer feature aggregation module function and the convolutional feature extraction module function to extract medium-resolution intermediate features, with the formula as follows:
[0085]
[0086] where, represents the extracted medium-resolution intermediate features;
[0087] S47. Input the low-resolution residual features extracted in step S43 and the medium-resolution intermediate features extracted in step S46 into the double-layer feature aggregation module function and the convolutional feature extraction module function to extract low-resolution intermediate features. The formula is as follows:
[0088]
[0089] where represents the extracted low-resolution intermediate features.
[0090] Furthermore, step S5 specifically includes:
[0091] S51. Input the high-resolution intermediate feature F21, the medium-resolution intermediate feature F22, and the low-resolution intermediate feature F23 extracted in step S4 into the Head designed based on the convolutional layer to complete object detection prediction, and extract the high-resolution prediction feature map, the medium-resolution prediction feature map, and the low-resolution prediction feature map. The formula is as follows:
[0092] ;
[0093] ;
[0094] ;
[0095] where represents the high-resolution prediction feature map, represents the medium-resolution prediction feature map, represents the low-resolution prediction feature map;
[0096] S52. Based on the extracted high-resolution prediction feature map, medium-resolution prediction feature map, and low-resolution prediction feature map, obtain the feature map , and the formula is as follows:
[0097] ;
[0098] where represents the coordinate position of the center point of the bounding box of the feature map , represents the width and height of the bounding box of the feature map , represents the category of the bounding box of the feature map , and represents the confidence of the bounding box of the feature map
[0099] S53. To merge between different resolutions, convert the bounding box coordinates of different resolutions to the same scale. The normalization formula is as follows:
[0100] ;
[0101] ;
[0102] ;
[0103] ;
[0104] Among them, represents the coordinates of the normalized bounding box, represents the height and width of the input image, and represent the height and width of the feature map ;
[0105] S54. Calculate the upper-left and lower-right coordinates of the bounding box based on the normalized bounding box coordinates. The formula is as follows:
[0106] ;
[0107] ;
[0108] ;
[0109] ;
[0110] Among them, represents the upper-left coordinate of the bounding box, represents the lower-right coordinate of the bounding box;
[0111] S55. Filter out the candidate bounding boxes with confidence greater than the threshold. The formula is as follows:
[0112] ;
[0113] Among them, represents the confidence threshold;
[0114] S56. Filter the bounding boxes according to the specified category. The formula is as follows:
[0115] ;
[0116] Among them, represents the specified category, represents the category, ;
[0117] S57. Apply non-maximum suppression to the remaining bounding boxes. The formula is as follows:
[0118]
[0119] Among them, represents two different remaining bounding boxes after the above operations, represents the bounding box and the bounding box of the intersection over union (IoU), represents the bounding box and the bounding box of the intersection area, represents the bounding box and the bounding box of the union area;
[0120] S58. Sort all the bounding boxes, select the bounding box with the highest confidence, and remove other bounding boxes whose intersection over union (IoU) with it is greater than the set threshold. Repeat this process until no bounding box meets the condition, and return the remaining bounding boxes as the final detection result. The formula is as follows:
[0121]
[0122] Among them, represents the final detection result, and NMS represents the non-maximum suppression process.
[0123] Furthermore, in step S6, the designed comprehensive loss function is as follows:
[0124]
[0125] Among them, represents the weight hyperparameter of the classification loss, represents the weight hyperparameter of the bounding box regression loss, represents the weight hyperparameter of the confidence loss, represents the classification loss, which is used to measure the performance of the model in the classification task, represents the bounding box regression loss, which is used to measure the performance of the model in localizing the bounding box, represents the low-confidence loss, which is used to measure the confidence prediction of the model for the existence of the target, represents the comprehensive loss, which is the weighted sum of the classification loss, the bounding box regression loss, and the confidence loss.
[0126] Compared with the prior art, the present invention has the following advantages:
[0127] 1. A nighttime target detection method based on a feature aggregation network provided by the present invention. The Backbone designed based on feature aggregation can effectively extract features of different resolutions. By designing a double-layer feature aggregation module and a pyramid aggregation attention mechanism, more detailed information can be captured, thereby improving the detection accuracy.
[0128] 2. A night-time target detection method based on a feature aggregation network provided by the present invention. The Neck based on pyramid feature fusion can further fuse features at different levels, improving the model's detection ability for targets of different scales. This Neck includes a convolutional feature extraction and a double-layer feature aggregation module, which can better maintain the diversity of features and improve the robustness of the model.
[0129] 3. A night-time target detection method based on a feature aggregation network provided by the present invention. By combining the spatial pyramid pooling technique with an attention mechanism, it can effectively extract information at different scales, and strengthen important features through the attention mechanism, further enhancing the detection performance of the model.
[0130] 4. A night-time target detection method based on a feature aggregation network provided by the present invention. The comprehensive loss function designed using classification loss, bounding box regression loss, and confidence loss can more comprehensively guide the training of the model and improve the accuracy of the detection results.
[0131] For the above reasons, the present invention can be widely promoted in fields such as target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0132] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0133] Figure 1 It is a flowchart of the method of the present invention.
[0134] Figure 2 It is a comparison chart of the precision-recall curves of the method of the present invention and other algorithms.
[0135] Figure 3 It is the target detection results of the method of the present invention and other algorithms in a night-time indoor scene.
[0136] Figure 4 It is the target detection results of the method of the present invention and other algorithms in a night-time outdoor scene. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0137] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.
[0138] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0139] As Figure 1 shown, the present invention provides a nighttime target detection method based on a feature aggregation network, including:
[0140] S1. Obtain a nighttime target detection dataset, and randomly divide the dataset into a training set and a test set according to a ratio of 7:3;
[0141] S2. Use image preprocessing operations to augment the nighttime target detection dataset obtained in step S1;
[0142] S3. Select the images after the preprocessing operation in step S2, and input the selected images into the Backbone part designed based on feature aggregation for general feature extraction to extract high-resolution initial features, medium-resolution initial features, and low-resolution initial features;
[0143] S4. Input the high-resolution initial features, medium-resolution initial features, and low-resolution initial features obtained in step S3 into the Neck part based on pyramid feature fusion to further extract high-resolution intermediate features, medium-resolution intermediate features, and low-resolution intermediate features with diversity and robustness;
[0144] S5. Input the high-resolution intermediate features, medium-resolution intermediate features, and low-resolution intermediate features extracted in step S4 into the Head part designed based on convolutional layers to complete target detection prediction.
[0145] S6. Design a comprehensive loss function using classification loss, bounding box regression loss, and confidence loss to constrain the detection process based on the feature aggregation network.
[0146] Specifically, as a preferred implementation manner of the present invention, step S2 specifically includes:
[0147] S21. Adopt the Mosaic technique to randomly splice four images into a new image, thereby simulating a more complex scene to increase the diversity of the dataset and enhance the model's adaptability to different backgrounds. The formula is as follows:
[0148] ;
[0149] Wherein, represents randomly selecting four night images from the dataset, represents splicing together according to a random layout, represents the resulting image of the Mosaic data augmentation method;
[0150] S22. Adopt the RandomAffine random affine transformation technique to perform rotation, translation, scaling, and shearing operations. The formula is as follows:
[0151] ;
[0152] Wherein, represents the horizontal and vertical coordinates in the input original image, represents the horizontal and vertical coordinates after the RandomAffine random affine transformation, represents the parameters of the affine transformation matrix for controlling rotation, scaling, and shearing, represents the translation parameter;
[0153] S23. Adopt the MixUp technique to linearly mix two images and their labels in a certain proportion. The formula is as follows:
[0154] ;
[0155] ;
[0156] Wherein, and represent randomly selecting two original images from the night target detection dataset, and represent and corresponding labels, represents the mixing ratio randomly sampled from the uniform distribution between [0, 1];
[0157] S24. Employ the image blurring technique, and select Gaussian blurring to obtain the blurred image. The formula is as follows:
[0158] ;
[0159] Wherein, represents the blurred image; represents the normalization factor; represents the image blurring radius; represents the weight of the Gaussian kernel; represents the convolution operation; represents the original image randomly selected from the night dataset;
[0160] S25. Employ the HSV color space enhancement technique to simulate different lighting conditions by changing the hue, saturation, and brightness of the image. The formula is as follows:
[0161] ;
[0162] ;
[0163] ;
[0164] Wherein, represents the hue, saturation, and brightness components of the original image; represents the hue, saturation, and brightness components of the enhanced image; represents the randomly adjusted parameter;
[0165] S26. Employ the random horizontal flipping technique to increase the robustness of the model to mirror changes. The formula is as follows:
[0166] ;
[0167] Wherein, represents the image after the random horizontal flipping operation; represents the width of the image.
[0168] Specifically in implementation, as a preferred implementation manner of the present invention, the Backbone designed based on feature aggregation in step S3 includes a Focus feature extraction module, a convolutional feature extraction module, a double-layer feature aggregation module, and a pyramid aggregation attention mechanism. Step S3 specifically includes:
[0169] S31. The Focus feature extraction module extracts efficient features from the input image. Its main function is to reorganize the spatial information of the input image, reduce the width dimension and height dimension of the image by half through slicing and splicing operations, and at the same time increase the number of channels, so as to perform subsequent convolution operations. Among them:
[0170] The formula for the slicing operation is as follows:
[0171] ;
[0172] ;
[0173] ;
[0174] ;
[0175] Among them, represents the original image, and the size of the original image is ([[]] C , H , W ), where C is the number of channels, H is the height, W is the width, represents the 4 parts into which the original image is sliced according to odd and even indices, The sizes of all are ([[]] C , H / 2, W / 2);
[0176] The formula for the splicing operation is as follows:
[0177] ;
[0178] Among them, represents the splicing operation, represents the feature image after splicing, The sizes of all are (4[[[]] C , H / 2, W / 2);
[0179] For further feature extraction, a convolution operation is performed on the feature image after the splicing operation, and the formula is as follows:
[0180] ;
[0181] Among them, represents the output feature image after convolution, represents the convolution operation;
[0182] S32. For further feature extraction, the function of the convolution feature extraction module is designed, and the formula is as follows:
[0183] ;
[0184] Among them, represents the function of the convolution feature extraction module, Represents the input variable of the convolutional feature extraction module function, Represents the LeakyReLU activation function, Represents the batch normalization operation;
[0185] S33. Design the function of the double-layer feature aggregation module, and the formula is as follows:
[0186] ;
[0187] Among them, Represents the double-layer feature aggregation module function, Represents the input variable of the double-layer feature aggregation module function, Represents the channel direction splitting operation;
[0188] S34. Design the function of the pyramid aggregation attention mechanism, and the formula is as follows:
[0189]
[0190] Among them, Represents the pyramid aggregation attention mechanism function, Represents the input variable of the pyramid aggregation attention mechanism function, Represents the max pooling operation;
[0191] S35. Input the output feature image of step S31 into the convolutional feature extraction module function and the double-layer feature aggregation module function in sequence to extract the high-resolution initial feature, and the formula is as follows:
[0192]
[0193] Among them, F 11 represents the extracted high-resolution initial feature;
[0194] S36. Input the high-resolution initial feature F11 extracted in step S35 into the convolutional feature extraction module function and the double-layer feature aggregation module function in sequence to extract the medium-resolution initial feature, and the formula is as follows:
[0195]
[0196] Among them, represents the extracted medium-resolution initial feature;
[0197] S37. Input the medium-resolution initial feature extracted in step S36 into the convolutional feature extraction module function and the pyramid aggregation attention mechanism function in sequence to extract the low-resolution initial feature, and the formula is as follows:
[0198]
[0199] Among them, represents the extracted low-resolution initial features.
[0200] In specific implementation, as a preferred implementation manner of the present invention, the Neck based on pyramid feature fusion in step S4 includes a convolutional feature extraction module and a double-layer feature aggregation module. Step S4 specifically includes:
[0201] S41. Design the function of the convolutional feature extraction module, and the formula is as follows:
[0202] ;
[0203] S42. Design the function of the double-layer feature aggregation module, and the formula is as follows:
[0204] ;
[0205] S43. Input the low-resolution initial features extracted in step S37 into the functions of the double-layer feature aggregation module and the convolutional feature extraction module in sequence to extract low-resolution residual features. The formula is as follows:
[0206]
[0207] Among them, represents the extracted low-resolution residual features;
[0208] S44. Input the medium-resolution initial features extracted in step S36 and the low-resolution residual features extracted in step S43 into the functions of the double-layer feature aggregation module and the convolutional feature extraction module to extract residual medium-resolution features. The formula is as follows:
[0209]
[0210] Among them, represents the extracted residual medium-resolution features, UP () represents the upsampling operation;
[0211] S45. Input the high-resolution initial features F extracted in step S35 and the residual medium-resolution features extracted in step S44 into the function of the double-layer feature aggregation module to extract high-resolution intermediate features. The formula is as follows:
[0212]
[0213] Among them, represents the extracted high-resolution intermediate features;
[0214] S46. Input the medium-resolution feature of the residuals extracted in step S44 and the high-resolution intermediate feature extracted in step S45 into the double-layer feature aggregation module function and the convolutional feature extraction module function to extract the medium-resolution intermediate feature. The formula is as follows:
[0215]
[0216] where represents the extracted medium-resolution intermediate feature;
[0217] S47. Input the low-resolution residual feature extracted in step S43 and the medium-resolution intermediate feature extracted in step S46 into the double-layer feature aggregation module function and the convolutional feature extraction module function to extract the low-resolution intermediate feature. The formula is as follows:
[0218]
[0219] where represents the extracted low-resolution intermediate feature.
[0220] In specific implementation, as a preferred implementation manner of the present invention, step S5 specifically includes:
[0221] S51. Input the high-resolution intermediate feature F21, the medium-resolution intermediate feature F22, and the low-resolution intermediate feature F23 extracted in step S4 into the Head designed based on the convolutional layer to complete the target detection prediction, and extract the high-resolution prediction feature map, the medium-resolution prediction feature map, and the low-resolution prediction feature map. The formula is as follows:
[0222] ;
[0223] ;
[0224] ;
[0225] where represents the high-resolution prediction feature map, represents the medium-resolution prediction feature map, represents the low-resolution prediction feature map;
[0226] S52. Based on the extracted high-resolution prediction feature map, medium-resolution prediction feature map, and low-resolution prediction feature map, obtain the feature map , and the formula is as follows:
[0227] ;
[0228] Among them, represents the center point coordinate position of the bounding box of the feature map ; represents the width and height of the bounding box of the feature map ; represents the class of the bounding box of the feature map ; is the confidence of the
[0229] S53. To merge between different resolutions, the bounding box coordinates of different resolutions are converted to the same scale, and the normalization formula is as follows:
[0230] ;
[0231] ;
[0232] ;
[0233] ;
[0234] Among them, represents the coordinates of the normalized bounding box, represents the height and width of the input image, and represents the height and width of the feature map ;
[0235] S54. According to the coordinates of the normalized bounding box, calculate the upper left corner coordinates and lower right corner coordinates of the bounding box, and the formula is as follows:
[0236] ;
[0237] ;
[0238] ;
[0239] ;
[0240] Among them, represents the upper left corner coordinates of the bounding box, represents the lower right corner coordinates of the bounding box;
[0241] S55. Filter out the candidate bounding boxes with a confidence greater than the threshold, and the formula is as follows:
[0242] ;
[0243] Among them, represents the confidence threshold;
[0244] S56. Filter the bounding boxes according to the specified category, and the formula is as follows:
[0245] ;
[0246] wherein, represents the specified category, represents the category, ;
[0247] S57. Apply non-maximum suppression to the remaining bounding boxes, and the formula is as follows:
[0248]
[0249] wherein, represents two different remaining bounding boxes after the above operations, represents the bounding box and the bounding box of the intersection over union, represents the bounding box and the bounding box of the intersection area, represents the bounding box and the bounding box of the union area; in this embodiment, the intersection over union (IoU) is used to measure the overlap degree between the predicted bounding box and the ground truth bounding box. The IoU value ranges from 0 to 1, and the higher the value, the higher the positioning accuracy. The IoU value of 1.0 indicates complete alignment. Usually, the threshold of IoU is 0.50, which is used to define the true positive in metrics such as mAP. A lower IoU value indicates that the model has difficulty in accurately positioning the object, and it can be improved by improving the bounding box regression or increasing the annotation accuracy.
[0250] S58. Sort all the bounding boxes, select the bounding box with the highest confidence, and remove other bounding boxes whose intersection over union IoU with it is greater than the set threshold. Repeat this process until no bounding box meets the condition, and return the remaining bounding boxes as the final detection result. The formula is as follows:
[0251]
[0252] wherein, represents the final detection result, and NMS represents the non-maximum suppression process.
[0253] Specifically, in the implementation, as a preferred implementation manner of the present invention, in step S6, the designed comprehensive loss function is as follows:
[0254]
[0255] Among them, represents the weight hyperparameter of the classification loss, represents the weight hyperparameter of the bounding box regression loss, represents the weight hyperparameter of the confidence loss, represents the classification loss, which is used to measure the performance of the model in the classification task, represents the bounding box regression loss, which is used to measure the performance of the model in localizing the bounding box, represents the low-confidence loss, which is used to measure the confidence prediction of the model for the existence of the target, represents the combined loss, which is the weighted sum of the classification loss, the bounding box regression loss, and the confidence loss.
[0256] Embodiment
[0257] As Figure 2 shown, it shows the Precision-Recall Curve of YOLOv5s, YOLOv6s, and the method of the present invention. Figure 2 Among them, (a) represents the Precision-Recall curve graph using the YOLOv5s method; (b) represents the Precision-Recall curve graph using the YOLOv6s method; (c) represents the Precision-Recall curve graph using the method of the present invention; Precision and Recall are important indicators for evaluating the performance of the classification model. The Precision-Recall curve shows the trade-off relationship between precision and recall at different thresholds and is used for the evaluation of imbalanced datasets or when focusing on the detection performance of the positive class. Figure 2 It can be seen from that the mAP@0.5 (mAP50) index of the method of the present invention in all classes is 0.637, higher than the mAP@0.5 (mAP50) index of the YOLOv5s method, which is 0.618, and higher than the mAP@0.5 (mAP50) index of the YOLOv6s method, which is 0.625. The method of the present invention performs excellently in both precision and recall, indicating better performance in the night target detection task.
[0258] As Figure 3 shown, it shows the target detection results of the present invention and other algorithms in the night indoor scene. Figure 3 Among them, (a) represents the detection result of using the YOLOv5s method for the target; (b) represents the detection result of using the YOLOv6s method for the target; (c) represents the detection result of using the method of the present invention for the target; (d) represents the real detection result. From Figure 3It can be seen that the mAP@0.5 (mAP50) index of the YOLOv5s method and the YOLOv6s method in all classes is the same, both being 0.3. However, both the YOLOv5s method and the YOLOv6s method wrongly detect the cat in the night image as a dog. The mAP@0.5 (mAP50) index of the method of the present invention in all classes is 0.8, and the method of the present invention successfully detects the cat in the image.
[0259] As Figure 4 shown, the target detection results of the present invention and other algorithms in the night outdoor scene are presented. Figure 4 Among them, (a) represents the result of detecting the target using the YOLOv5s method; (b) represents the result of detecting the target using the YOLOv6s method; (c) represents the result of detecting the target using the method of the present invention; (d) represents the real detection result. According to Figure 4 it can be seen that the mAP@0.5 (mAP50) index of the method of the present invention, the YOLOv5s method and the YOLOv6s method in all classes is the same, all being 0.4. Although the YOLOv5s method, the YOLOv6s method and the method of the present invention all successfully detect the hull, the detection box of the YOLOv5s method only marks half of the hull, while the YOLOv6s method generates three detection boxes for the same hull, and two of the detection boxes have incorrect detection pixel ranges. To sum up, the method of the present invention can robustly detect objects in the night scene and at the same time provide a more accurate target detection range.
[0260] In this embodiment, different algorithms are compared in terms of accuracy, recall rate, mAP50, and mAP50-95 objective indicators; accuracy quantifies the proportion of true positives in all positive predictions and evaluates the ability of the model to avoid false positives. The recall rate calculates the proportion of true positive predictions among all actual positive predictions and measures the ability of the model to detect all instances of a certain class. mAP50 is the average precision calculated according to the intersection over union (IoU) threshold of 0.50, which is a measure of the precision of the model considering only "easy" detections. mAP50-95 is the average of the average precisions calculated at different IoU thresholds between 0.50 and 0.95, which comprehensively reflects the performance of the model under different detection difficulties. As shown in the following table:
[0261] Table 1 Comparison of the precision of the processing results of the model of the present invention and other advanced algorithms
[0262]
[0263] Table 2 Comparison of the recall rate of the processing results of the model of the present invention and other advanced algorithms
[0264]
[0265] Table 3 Comparison of mAP50 of the model of the present invention and the processing results of other advanced algorithms
[0266]
[0267] Table 4 Comparison of mAP50-95 of the model of the present invention and the processing results of other advanced algorithms
[0268]
[0269] In summary, the model of the present invention is superior to YOLOv5s and YOLOv6s in terms of accuracy, recall rate, mAP50, and mAP50-95. Specifically, the algorithm of the present invention has higher precision, recall rate, and average precision on the night target detection dataset. In particular, the improvement of mAP50-95 shows the consistent advantage of the model under different IoU thresholds.
[0270] Therefore, the present invention is more suitable for application in night target detection tasks, can detect and locate targets more effectively, while reducing false alarms and missed detections, and is particularly suitable for deployment and use in night monitoring scenarios.
[0271] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A nighttime target detection method based on a feature aggregation network, characterized in that, Including: S1. Obtain a night target detection dataset, and randomly divide the dataset into a training set and a test set according to a ratio of 7:
3. S2. Use image preprocessing operations to augment the night target detection dataset obtained in step S1. S3. Select the images after the preprocessing operation in step S2, and input the selected images into the backbone part designed based on feature aggregation for general feature extraction, extracting high-resolution initial features, medium-resolution initial features, and low-resolution initial features. The backbone based on feature aggregation design includes an aggregated feature extraction module, a convolutional feature extraction module, a two-layer feature aggregation module, and a pyramid aggregation attention mechanism, specifically including: S31. The aggregated feature extraction module extracts efficient features from the input image, reorganizes the spatial information of the input image, reduces the width dimension and height dimension of the image by half through slicing operations and splicing operations, and simultaneously increases the number of channels, where: The formula for the slicing operation is as follows: ; ; ; ; Among them, represents the original image, and the size of the original image is ([[]] C , H , W ), where C is the number of channels, H is the height, W is the width, represents 4 parts obtained by slicing the original image according to odd and even indices, and the size of is all ([[]] C , H / 2, W / 2); The formula for the splicing operation is as follows: ; Among them, represents a splicing operation, represents the feature image after splicing, The sizes of all are (4 C , H / 2, W / 2); Perform a convolution operation on the feature image after the splicing operation, and the formula is as follows: ; Among them, represents the output feature image after convolution, represents the convolution operation; S32. Design the function of the convolutional feature extraction module, and the formula is as follows: ; Among them, represents the convolutional feature extraction module function, represents the input variable of the convolutional feature extraction module function, represents the LeakyReLU activation function, represents the batch normalization operation; S33. Design the function of the two-layer feature aggregation module, and the formula is as follows: ; Among them, represents the double-layer feature aggregation module function, represents the input variable of the double-layer feature aggregation module function, represents the channel-wise splitting operation; S34. Design the function of the pyramid aggregation attention mechanism, and the formula is as follows: Among them, represents the pyramid pooling attention mechanism function, represents the input variable of the pyramid pooling attention mechanism function, represents the max pooling operation; S35. Input the output feature image of step S31 into the convolutional feature extraction module function and the double-layer feature aggregation module function in sequence to extract high-resolution initial features. The formula is as follows: Among them, F 11 represents the extracted high-resolution initial features; S36. Input the high-resolution initial feature F11 extracted in step S35 into the convolutional feature extraction module function and the two-layer feature aggregation module function in sequence to extract medium-resolution initial features, and the formula is as follows: Among them, represents the extracted initial medium-resolution features; S37. Input the initial medium-resolution features extracted in step S36 into the convolutional feature extraction module function and the pyramid aggregation attention mechanism function in sequence to extract the initial low-resolution features. The formula is as follows: Among them, represents the extracted low-resolution initial feature; S4. Input the high-resolution initial features, medium-resolution initial features, and low-resolution initial features obtained in step S3 into the neck part based on pyramid feature fusion to further extract high-resolution intermediate features, medium-resolution intermediate features, and low-resolution intermediate features with diversity and robustness; the neck based on pyramid feature fusion includes a convolutional feature extraction module and a two-layer feature aggregation module, specifically including: S41. Design the function of the convolutional feature extraction module, and the formula is as follows: ; S42. Design the function of the two-layer feature aggregation module, and the formula is as follows: ; S43. Input the low-resolution initial features extracted in step S37 into the double-layer feature aggregation module function and the convolutional feature extraction module function in sequence to extract low-resolution residual features. The formula is as follows: Among them, represents the extracted low-resolution residual features; S44. Input the initial medium-resolution features extracted in step S36 and the low-resolution residual features extracted in step S43 into the double-layer feature aggregation module function and the convolutional feature extraction module function to extract residual medium-resolution features. The formula is as follows: Among them, represents the resolution feature in the extracted residual, UP () represents the upsampling operation; S45. Input the high-resolution initial features extracted in step S35 F 11 and the residual medium-resolution features extracted in step S44 into the double-layer feature aggregation module function to extract high-resolution intermediate features. The formula is as follows: Among them, represents the extracted high-resolution intermediate features; S46. Input the medium-resolution features of the residuals extracted in step S44 and the high-resolution intermediate features extracted in step S45 into the double-layer feature aggregation module function and the convolutional feature extraction module function to extract medium-resolution intermediate features. The formula is as follows: Among them, represents the extracted medium-resolution intermediate features; S47. Input the low-resolution residual features extracted in step S43 and the mid-resolution intermediate features extracted in step S46 into the double-layer feature aggregation module function and the convolutional feature extraction module function to extract low-resolution intermediate features. The formula is as follows: Among them, represents the extracted low-resolution intermediate features; S5. Input the high-resolution intermediate features, medium-resolution intermediate features, and low-resolution intermediate features extracted in step S4 into the prediction head part designed based on convolutional layers to complete target detection prediction.
2. The night target detection method based on a feature aggregation network according to claim 1, characterized in that The method further includes: S6. Use classification loss, bounding box regression loss, and confidence loss to design a comprehensive loss function to constrain the detection process based on the feature aggregation network.
3. The method for night-time target detection based on a feature aggregation network according to claim 1, wherein Step S2 specifically includes: S21. Adopt the mosaic technique to randomly splice four images into a new image, and the formula is as follows: ; Among them, represents randomly selecting four night-time images from the dataset, represents arranging together according to a random layout, represents the resulting image of the mosaic data augmentation method; S22. Adopt the random affine transformation technique to perform rotation, translation, scaling, and shearing operations, and the formula is as follows: ; Among them, represents the horizontal and vertical coordinates in the original input image, represents the horizontal and vertical coordinates after random affine transformation, represents the parameters of the affine transformation matrix, which are used to control rotation, scaling, and shearing, represents the translation parameter; S23. Adopt the linear mixing technique to linearly mix two pictures and their labels, and the formula is as follows: ; ; Among them, and represent two original images randomly selected from the night target detection dataset, and represent and the corresponding labels, represents the mixing ratio randomly sampled from the uniform distribution between [0, 1]; S24. Adopt the image blurring technique to obtain a blurred image by selecting Gaussian blurring, and the formula is as follows: ; Among them, represents the blurred image, represents the normalization factor, represents the image blur radius, represents the weight of the Gaussian kernel, represents the convolution operation, represents the original image randomly selected from the night dataset; S25. Adopt the HSV color space enhancement technique to simulate different lighting conditions by changing the hue, saturation, and brightness of the image, and the formula is as follows: ; ; ; Among them, represents the hue, saturation, and brightness components of the original image; represents the hue, saturation, and brightness components of the enhanced image; represents the random adjustment parameter; S26. Adopt the random horizontal flipping technique to increase the robustness of the model to mirror changes. The formula is as follows: ; Among them, represents the image after the random horizontal flipping operation, represents the width of the image.
4. A method for night-time target detection based on a feature aggregation network according to claim 1, wherein, Step S5 specifically includes: S51. Input the high-resolution intermediate feature F21, medium-resolution intermediate feature F22, and low-resolution intermediate feature F23 extracted in step S4 into the prediction head designed based on the convolutional layer to complete object detection prediction, and extract the high-resolution prediction feature map, medium-resolution prediction feature map, and low-resolution prediction feature map. The formula is as follows: ; ; ; Among them, represents the high-resolution prediction feature map, represents the medium-resolution prediction feature map, represents the low-resolution prediction feature map; S52. Based on the extracted high-resolution prediction feature map, medium-resolution prediction feature map, and low-resolution prediction feature map, obtain a feature map , and the formula is as follows: ; Among them, represents the center point coordinate position of the bounding box of the feature map , represents the width and height of the bounding box of the feature map , represents the class of the bounding box of the feature map , and the confidence of S53. To merge between different resolutions, convert the bounding box coordinates of different resolutions to the same scale. The normalization formula is as follows: ; ; ; ; Among them, represents the coordinates of the normalized bounding box, represents the height and width of the input image, and represents the feature map height and width; S54. Calculate the upper-left coordinate and lower-right coordinate of the bounding box according to the normalized bounding box coordinates. The formula is as follows: ; ; ; ; Among them, represents the upper left corner coordinates of the bounding box, represents the lower right corner coordinates of the bounding box; S55. Screen out the candidate bounding boxes with a confidence greater than the threshold. The formula is as follows: ; wherein, represents a confidence threshold; S56. Filter the bounding boxes according to the specified category. The formula is as follows: ; Among them, represents a specified category, represents a category, ; S57. Apply non-maximum suppression to the remaining bounding boxes. The formula is as follows: Among them, represents two different remaining bounding boxes after the above operations, represents the bounding box and the bounding box 's intersection over union (IoU), represents the intersection area of the bounding box and the bounding box ; represents the union area of the bounding box and the bounding box . S58. Sort all the bounding boxes, select the bounding box with the highest confidence, and remove other bounding boxes with an intersection over union (IoU) greater than the set threshold with it. Repeat this process until no bounding box meets the condition, and return the remaining bounding boxes as the final detection result. The formula is as follows: Among them, represents the final detection result, and NMS represents the non-maximum suppression process.
5. A night target detection method based on a feature aggregation network according to claim 2, characterized in that, In step S6, the designed comprehensive loss function is as follows: Among them, represents the weight hyperparameter of the classification loss, represents the weight hyperparameter of the bounding box regression loss, represents the weight hyperparameter of the confidence loss, represents the classification loss, which is used to measure the performance of the model in the classification task, represents the bounding box regression loss, which is used to measure the performance of the model in localizing the bounding box, represents the low confidence loss, which is used to measure the confidence prediction of the model for the existence of the target, represents the comprehensive loss, which is the weighted sum of the classification loss, the bounding box regression loss, and the confidence loss.
Citation Information
Patent Citations
Night vehicle detection method and system based on improved YOLOv5 convolutional neural network
CN116524319A
Cloth defect detection method and system based on multi-scale feature fusion and diffusion pyramid network
CN119515861A