Night target detection method based on feature aggregation network

By adopting a feature aggregation network-based method in night object detection, combining the two-layer feature aggregation module and the pyramid aggregation attention mechanism, the performance reduction problem of night object detection in environments with insufficient light and noise increase is solved, and higher detection accuracy and robustness are achieved.

CN120107564AActive Publication Date: 2025-06-06ZHOUSHAN YONGXIANG SHIPPING CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510570235.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-06
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

Night object detection has greatly reduced performance in environments with insufficient light and increased noise, and data annotation and collection are difficult, resulting in low model overfitting and computational efficiency.

Method used

Using a feature aggregation network-based method, general feature extraction is performed by designing the main part of feature aggregation, combining the two-layer feature aggregation module and the pyramid aggregation attention mechanism, further fusing features at different levels, and using data augmentation technology and comprehensive loss function for model training.

Benefits of technology

It improves the accuracy and robustness of night object detection, enhances the model's detection ability of targets at different scales, reduces the risk of overfitting, and improves the computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107564A_ABST
    Figure CN120107564A_ABST
Patent Text Reader

Abstract

The invention provides a night target detection method based on a feature aggregation network, belongs to the technical field of image data processing and detection, effectively extracts image information through feature aggregation operation, and aims to improve the accuracy and robustness of night target detection. The method comprises the following steps: firstly, dividing a night target detection data set into a training set and a test set, and performing data augmentation on a night image data set by using image preprocessing operation; then, general feature extraction is carried out by using a trunk part based on feature aggregation design, the efficiency and precision of feature extraction are improved through an aggregation feature extraction architecture, and feature extraction is carried out by using convolutional feature extraction, a double-layer feature aggregation module and a pyramid aggregation attention mechanism; further fusing features through the neck based on pyramid feature fusion; and finally, a target detection task is completed through the prediction head. The whole detection process is constrained through a comprehensive loss function, including classification loss, bounding box regression loss and confidence loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image data processing and detection technology, and in particular to a nighttime target detection method based on a feature aggregation network. Background Art

[0002] With the acceleration of urbanization and the continuous development of society, nighttime activities and nighttime safety are becoming more and more important. Nighttime object detection has important application value in traffic management, public safety, military monitoring, etc. However, due to insufficient illumination, poor image quality and increased noise in the nighttime environment, the performance of traditional object detection methods is greatly reduced in the nighttime environment. Therefore, how to improve the accuracy and robustness of nighttime object detection has become an important research topic in the field of computer vision and image processing.

[0003] In recent years, nighttime target detection methods have been widely used in computer vision and image processing. Many scholars at home and abroad have studied this topic and made corresponding progress. With the rapid development of deep learning technology, target detection methods based on deep learning have shown great potential in nighttime target detection. Deep learning methods can automatically extract high-level features of images by constructing deep neural networks, and have strong robustness and generalization capabilities. This method directly inputs from nighttime images to target detection results, which can provide additional scene information for nighttime scenes and improve the accuracy and robustness of target detection.

[0004] Although deep learning-based methods have shown excellent performance in nighttime object detection, they still face some challenges. First, the illumination changes dramatically and there is a lot of noise in the nighttime environment, which puts higher requirements on the training and inference of deep learning models. Secondly, it is difficult to label and collect nighttime object detection datasets, and insufficient data may lead to overfitting and poor generalization of the model. In addition, application scenarios with high real-time requirements place strict requirements on the computational efficiency of the model, and it is necessary to find a balance between model complexity and detection accuracy.

[0005] In summary, the development of the field of night target detection requires the comprehensive application of multiple technical means to improve the accuracy and robustness of night target detection to meet the needs of practical applications. Summary of the invention

[0006] According to the technical problems raised above, a method for nighttime target detection based on a feature aggregation network is provided. The present invention performs general feature extraction based on the backbone part (Backbone) designed based on feature aggregation, improves the efficiency and accuracy of feature extraction by aggregating feature extraction modules (Focus), and uses convolutional feature extraction, a two-layer feature aggregation module and a pyramid aggregation attention mechanism for feature extraction. And a Neck based on pyramid feature fusion is designed to further fuse the initial features extracted from the backbone part (Backbone). The method of the present invention can effectively improve the effect and robustness of target detection in nighttime environments.

[0007] The technical means adopted by the present invention are as follows: A nighttime target detection method based on a feature aggregation network, comprising: S1. Obtain a nighttime target detection dataset and randomly divide the dataset into a training set and a test set in a ratio of 7:3; S2, using image preprocessing operations to perform data set augmentation on the nighttime target detection data set obtained in step S1; S3, selecting the image after the preprocessing operation in step S2, and inputting the selected image into the backbone part (Backbone) designed based on feature aggregation to extract general features, and extract high-resolution initial features, medium-resolution initial features and low-resolution initial features; S4, inputting the high-resolution initial features, medium-resolution initial features and low-resolution initial features obtained in step S3 into the neck part based on pyramid feature fusion, and further extracting high-resolution intermediate features, medium-resolution intermediate features and low-resolution intermediate features with diversity and robustness; S5. Input the high-resolution intermediate features, medium-resolution intermediate features, and low-resolution intermediate features extracted in step S4 into the prediction head (Head) part designed based on the convolution layer to complete the target detection prediction.

[0008] Furthermore, the method further comprises: S6. Use classification loss, bounding box regression loss, and confidence loss to design a comprehensive loss function to constrain the detection process based on the feature aggregation network.

[0009] Furthermore, step S2 specifically includes: S21, Mosaic technology is used to randomly splice four images into a new image. The formula is as follows: ; in, means randomly selecting four night images from the dataset, Indicates that according to the random layout Stitched together, Represents the result image of the mosaic data augmentation method; S22, using linear mixing technology (MixUp), the two images and their labels are linearly mixed in a certain proportion, the formula is as follows: ; in, Represents the horizontal and vertical coordinates of the input original image, represents the horizontal and vertical coordinates after random affine transformation, Represents the parameters of the affine transformation matrix, which is used to control rotation, scaling, and shearing. represents the translation parameter; S23, using random affine transformation technology (RandomAffine) to perform rotation, translation, scaling and shearing operations, the formula is as follows: ; ; in, and Indicates that two original images are randomly selected from the nighttime target detection dataset. and express and The corresponding label, represents the mixing ratio randomly sampled from a uniform distribution between [0, 1]; S24, using image blur technology, select Gaussian blur to obtain a blurred image, the formula is as follows: ; in, represents the image after image blur, represents the normalization factor, represents the image blur radius, represents the weight of the Gaussian kernel, represents the convolution operation, represents a random raw image from the nighttime dataset; S25, using HSV color space enhancement technology, by changing the hue, saturation and brightness of the image to simulate different lighting conditions, the formula is as follows: ; ; ; in, Represents the hue, saturation, and brightness components of the original image; Indicates the hue, saturation, and brightness components of the enhanced image; represents the random adjustment parameter; S26, using random horizontal flipping technology to increase the model's robustness to mirror changes, the formula is as follows: ; in, represents the image after random horizontal flipping operation, Indicates the width of the image.

[0010] Furthermore, the Backbone designed based on feature aggregation in step S3 includes a focus feature extraction module, a convolution feature extraction module, a double-layer feature aggregation module and a pyramid aggregation attention mechanism. Step S3 specifically includes: S31, Focus feature extraction module extracts efficient features from the input image, reorganizes the spatial information of the input image, reduces the width and height dimensions of the image by half through slicing and splicing operations, and increases the number of channels, so as to perform subsequent convolution operations. Among them: The formula for the slice operation is as follows: ; ; ; ; in, Represents the original image, the size of the original image is ( C , H , W ),in C is the number of channels, H is the height, W is the width, Indicates that the original image is cut into 4 parts according to odd and even indexes. The dimensions are ( C , H / 2, W / 2); The formula for the splicing operation is as follows: ; in, Represents a splicing operation, represents the concatenated feature image, The dimensions are (4 C , H / 2, W / 2); The convolution operation is performed on the feature image after the splicing operation. The formula is as follows: ; in, represents the output feature image after convolution, Represents the convolution operation; S32. Design the function of the convolution feature extraction module. The formula is as follows: ; in, Represents the convolution feature extraction module function, Represents the input variable of the convolution feature extraction module function, represents the LeakyReLU activation function, Represents a batch normalization operation; S33. Design the function of the double-layer feature aggregation module. The formula is as follows: ; in, Represents a two-layer feature aggregation module function, Represents the input variable of the double-layer feature aggregation module function, Indicates the channel direction split operation; S34. Design the function of the pyramid aggregation attention mechanism. The formula is as follows:

[0011] in, represents the pyramid aggregation attention mechanism function, represents the input variable of the pyramid aggregation attention mechanism function, Represents the maximum pooling operation; S35: Output feature image of step S31 Input into the convolution feature extraction module function and the double-layer feature aggregation module function in turn to extract high-resolution initial features. The formula is as follows:

[0012] in, F 11 represents the extracted high-resolution initial features; S36, input the high-resolution initial feature F11 extracted in step S35 into the convolution feature extraction module function and the double-layer feature aggregation module function in sequence to extract the medium-resolution initial feature, the formula is as follows:

[0013] in, represents the extracted medium-resolution initial features; S37, the medium resolution initial feature extracted in step S36 Input into the convolution feature extraction module function and the pyramid aggregation attention mechanism function in turn to extract low-resolution initial features. The formula is as follows:

[0014] in, Represents the extracted low-resolution initial features.

[0015] Furthermore, the pyramid feature fusion-based Neck in step S4 includes a convolution feature extraction module and a double-layer feature aggregation module. Step S4 specifically includes: S41. Design the function of the convolution feature extraction module. The formula is as follows: ; S42. Design a function for the double-layer feature aggregation module. The formula is as follows: ; S43, the low-resolution initial features extracted in step S37 Input into the double-layer feature aggregation module function and the convolution feature extraction module function in turn to extract low-resolution residual features. The formula is as follows:

[0016] in, Represents the extracted low-resolution residual features; S44, the medium resolution initial features extracted in step S36 and the low-resolution residual features extracted in step S43 Input into the double-layer feature aggregation module function and the convolution feature extraction module function to extract the residual mid-resolution features. The formula is as follows:

[0017] in, represents the extracted residual mid-resolution feature, UP () indicates upsampling operation; S45, the high-resolution initial features extracted in step S35 F 11 and the residual medium-resolution feature extracted in step S44 Input into the double-layer feature aggregation module function to extract high-resolution intermediate features. The formula is as follows:

[0018] in, Represents the extracted high-resolution mid-level features; S46, the residual mid-resolution feature extracted in step S44 and the high-resolution intermediate features extracted in step S45 Input into the double-layer feature aggregation module function and the convolution feature extraction module function to extract medium-resolution intermediate features. The formula is as follows:

[0019] in, Represents the extracted medium-resolution mid-level features; S47, the low-resolution residual features extracted in step S43 and the medium-resolution intermediate features extracted in step S46 Input into the double-layer feature aggregation module function and the convolution feature extraction module function to extract low-resolution intermediate features. The formula is as follows:

[0020] in, Represents the extracted low-resolution mid-level features.

[0021] Furthermore, step S5 specifically includes: S51, input the high-resolution intermediate features F21, medium-resolution intermediate features F22, and low-resolution intermediate features F23 extracted in step S4 into the Head designed based on the convolutional layer to complete the target detection prediction, and extract the high-resolution prediction feature map, the medium-resolution prediction feature map, and the low-resolution prediction feature map. The formula is as follows: ; ; ; in, represents the high-resolution prediction feature map, represents the medium-resolution prediction feature map, Represents a low-resolution prediction feature map; S52, based on the extracted high-resolution prediction feature map, medium-resolution prediction feature map and low-resolution prediction feature map, obtain a feature map , the formula is as follows: ; in, Representation feature map The coordinate position of the center point of the bounding box, Representation feature map The width and height of the bounding box, Representation feature map The category of the bounding box Confidence level; S53. In order to merge between different resolutions, the bounding box coordinates of different resolutions are converted to the same scale. The normalization formula is as follows: ; ; ; ; in, represents the coordinates of the normalized bounding box, represents the height and width of the input image, and Representation feature map height and width; S54. Calculate the upper left corner coordinates and the lower right corner coordinates of the bounding box according to the standardized bounding box coordinates. The formula is as follows: ; ; ; ; in, Represents the coordinates of the upper left corner of the bounding box, Represents the coordinates of the lower right corner of the bounding box; S55. Filter out candidate bounding boxes whose confidence is greater than the threshold. The formula is as follows: ; in, represents the confidence threshold; S56. Filter the bounding box according to the specified category. The formula is as follows: ; in, Indicates the specified category. Indicates category, ; S57. Apply non-maximum suppression to the remaining bounding boxes. The formula is as follows:

[0022] in, represents the two different remaining bounding boxes after the above operations, Represents a bounding box and bounding box The intersection ratio of Represents a bounding box and bounding box The intersection area of Represents a bounding box and bounding box The area of ​​the union of S58. Sort all bounding boxes, select the bounding box with the highest confidence, and remove other bounding boxes whose intersection over union (IoU) is greater than the set threshold. Repeat this process until no bounding box meets the conditions, and return the remaining bounding boxes as the final detection result. The formula is as follows:

[0023] in, Represents the final detection result, and NMS represents the non-maximum suppression process.

[0024] Furthermore, in step S6, the designed comprehensive loss function is as follows:

[0025] in, represents the weight hyperparameter of the classification loss, represents the weight hyperparameter of the bounding box regression loss, represents the weight hyperparameter of the confidence loss, Represents classification loss, which is used to measure the performance of the model on the classification task. represents the bounding box regression loss, which is used to measure the performance of the model in locating the bounding box. Represents low confidence loss, which is used to measure the model's confidence prediction of the target's existence. represents the comprehensive loss, which is the weighted sum of classification loss, bounding box regression loss and confidence loss.

[0026] Compared with the prior art, the present invention has the following advantages: 1. The present invention provides a night target detection method based on a feature aggregation network. The Backbone designed based on feature aggregation can effectively extract features of different resolutions. By designing a double-layer feature aggregation module and a pyramid aggregation attention mechanism, more detailed information can be captured, thereby improving detection accuracy.

[0027] 2. The present invention provides a nighttime target detection method based on a feature aggregation network. The Neck based on pyramid feature fusion can further fuse features at different levels and improve the model's detection ability for targets of different scales. The Neck includes a convolutional feature extraction and a double-layer feature aggregation module, which can better maintain the diversity of features and improve the robustness of the model.

[0028] 3. The present invention provides a night target detection method based on a feature aggregation network, which can effectively extract information of different scales through the spatial pyramid pooling technology combined with the attention mechanism, and strengthen important features through the attention mechanism to further improve the detection performance of the model.

[0029] 4. The present invention provides a night target detection method based on a feature aggregation network. The comprehensive loss function designed using classification loss, bounding box regression loss, and confidence loss can more comprehensively guide model training and improve the accuracy of detection results.

[0030] Based on the above reasons, the present invention can be widely promoted in the fields of target detection and the like. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0032] Figure 1 The figure is a flow chart of the method of the present invention.

[0033] Figure 2 The figure is a comparison chart of the precision-recall curves of the method of the present invention and other algorithms.

[0034] Figure 3 The results of target detection in indoor scenes at night using the method of the present invention and other algorithms.

[0035] Figure 4 The results of target detection in outdoor scenes at night using the method of the present invention and other algorithms. DETAILED DESCRIPTION

[0036] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0038] like Figure 1 As shown, the present invention provides a nighttime target detection method based on a feature aggregation network, comprising: S1. Obtain a nighttime target detection dataset and randomly divide the dataset into a training set and a test set in a ratio of 7:3; S2, using image preprocessing operations to perform data set augmentation on the nighttime target detection data set obtained in step S1; S3, selecting the image after the preprocessing operation in step S2, and inputting the selected image into the Backbone part based on feature aggregation design to extract general features, and extracting high-resolution initial features, medium-resolution initial features and low-resolution initial features; S4, inputting the high-resolution initial features, medium-resolution initial features and low-resolution initial features obtained in step S3 into the Neck part based on pyramid feature fusion, and further extracting high-resolution medium-level features, medium-resolution medium-level features and low-resolution medium-level features with diversity and robustness; S5. Input the high-resolution intermediate features, medium-resolution intermediate features, and low-resolution intermediate features extracted in step S4 into the Head part designed based on the convolutional layer to complete the target detection prediction.

[0039] S6. Use classification loss, bounding box regression loss, and confidence loss to design a comprehensive loss function to constrain the detection process based on the feature aggregation network.

[0040] In specific implementation, as a preferred embodiment of the present invention, step S2 specifically includes: S21, Mosaic mosaic technology is used to randomly splice four images into a new image to simulate more complex scenes, increase the diversity of the data set, and enhance the model's adaptability to different backgrounds. The formula is as follows: ; in, means randomly selecting four night images from the dataset, Indicates that according to the random layout Stitched together, Represents the result image of the Mosaic data augmentation method; S22, using MixUp technology, linearly mix the two images and their labels in a certain proportion. The formula is as follows: ; in, Represents the horizontal and vertical coordinates of the input original image, Represents the horizontal and vertical coordinates after RandomAffine random affine transformation, Represents the parameters of the affine transformation matrix, which is used to control rotation, scaling, and shearing. represents the translation parameter; S23, using RandomAffine random affine transformation technology to perform rotation, translation, scaling and shearing operations, the formula is as follows: ; ; in, and Indicates that two original images are randomly selected from the nighttime target detection dataset. and express and The corresponding label, represents the mixing ratio randomly sampled from a uniform distribution between [0, 1]; S24, using image blur technology, select Gaussian blur to obtain a blurred image, the formula is as follows: ; in, represents the image after image blur, represents the normalization factor, represents the image blur radius, represents the weight of the Gaussian kernel, represents the convolution operation, represents a random raw image from the nighttime dataset; S25, using HSV color space enhancement technology, by changing the hue, saturation and brightness of the image to simulate different lighting conditions, the formula is as follows: ; ; ; in, Represents the hue, saturation, and brightness components of the original image; Indicates the hue, saturation, and brightness components of the enhanced image; represents the random adjustment parameter; S26, using random horizontal flipping technology to increase the model's robustness to mirror changes, the formula is as follows: ; in, represents the image after random horizontal flipping operation, Indicates the width of the image.

[0041] In specific implementation, as a preferred embodiment of the present invention, the Backbone based on feature aggregation design in step S3 includes a Focus feature extraction module, a convolution feature extraction module, a double-layer feature aggregation module and a pyramid aggregation attention mechanism. Step S3 specifically includes: S31, Focus feature extraction module extracts efficient features from the input image. Its main function is to reorganize the spatial information of the input image, reduce the width and height dimensions of the image by half through slicing and splicing operations, and increase the number of channels, so as to perform subsequent convolution operations. Among them: The formula for the slice operation is as follows: ; ; ; ; in, Represents the original image, the size of the original image is ( C , H , W ),in C is the number of channels, H is the height, W is the width, Indicates that the original image is cut into 4 parts according to odd and even indexes. The dimensions are ( C , H / 2, W / 2); The formula for the splicing operation is as follows: ; in, Represents a splicing operation, represents the concatenated feature image, The dimensions are (4 C , H / 2, W / 2); In order to further extract features, a convolution operation is performed on the feature image after the splicing operation. The formula is as follows: ; in, represents the output feature image after convolution, Represents the convolution operation; S32. To further extract features, a function of a convolution feature extraction module is designed. The formula is as follows: ; in, Represents the convolution feature extraction module function, Represents the input variable of the convolution feature extraction module function, represents the LeakyReLU activation function, Represents a batch normalization operation; S33. Design the function of the double-layer feature aggregation module. The formula is as follows: ; in, Represents a two-layer feature aggregation module function, Represents the input variable of the double-layer feature aggregation module function, Indicates the channel direction split operation; S34. Design the function of the pyramid aggregation attention mechanism. The formula is as follows:

[0042] in, represents the pyramid aggregation attention mechanism function, represents the input variable of the pyramid aggregation attention mechanism function, Represents the maximum pooling operation; S35: Output feature image of step S31 Input into the convolution feature extraction module function and the double-layer feature aggregation module function in turn to extract high-resolution initial features. The formula is as follows:

[0043] in, F 11 represents the extracted high-resolution initial features; S36, input the high-resolution initial feature F11 extracted in step S35 into the convolution feature extraction module function and the double-layer feature aggregation module function in sequence to extract the medium-resolution initial feature, the formula is as follows:

[0044] in, represents the extracted medium-resolution initial features; S37, the medium resolution initial feature extracted in step S36 Input into the convolution feature extraction module function and the pyramid aggregation attention mechanism function in turn to extract low-resolution initial features. The formula is as follows:

[0045] in, Represents the extracted low-resolution initial features.

[0046] In specific implementation, as a preferred embodiment of the present invention, the Neck based on pyramid feature fusion in step S4 includes a convolution feature extraction module and a double-layer feature aggregation module, and step S4 specifically includes: S41. Design the function of the convolution feature extraction module. The formula is as follows: ; S42. Design a function for the double-layer feature aggregation module. The formula is as follows: ; S43, the low-resolution initial features extracted in step S37 Input into the double-layer feature aggregation module function and the convolution feature extraction module function in turn to extract low-resolution residual features. The formula is as follows:

[0047] in, Represents the extracted low-resolution residual features; S44, the medium resolution initial features extracted in step S36 and the low-resolution residual features extracted in step S43 Input into the double-layer feature aggregation module function and the convolution feature extraction module function to extract the residual mid-resolution features. The formula is as follows:

[0048] in, represents the extracted residual mid-resolution feature, UP () indicates upsampling operation; S45, the high-resolution initial features extracted in step S35 F 11 and the residual medium-resolution feature extracted in step S44 Input into the double-layer feature aggregation module function to extract high-resolution intermediate features. The formula is as follows:

[0049] in, Represents the extracted high-resolution mid-level features; S46, the residual mid-resolution feature extracted in step S44 and the high-resolution intermediate features extracted in step S45 Input into the double-layer feature aggregation module function and the convolution feature extraction module function to extract medium-resolution intermediate features. The formula is as follows:

[0050] in, Represents the extracted medium-resolution mid-level features; S47, the low-resolution residual features extracted in step S43 and the medium-resolution intermediate features extracted in step S46 Input into the double-layer feature aggregation module function and the convolution feature extraction module function to extract low-resolution intermediate features. The formula is as follows:

[0051] in, Represents the extracted low-resolution mid-level features.

[0052] In specific implementation, as a preferred embodiment of the present invention, step S5 specifically includes: S51, input the high-resolution intermediate features F21, medium-resolution intermediate features F22, and low-resolution intermediate features F23 extracted in step S4 into the Head designed based on the convolutional layer to complete the target detection prediction, and extract the high-resolution prediction feature map, the medium-resolution prediction feature map, and the low-resolution prediction feature map. The formula is as follows: ; ; ; in, represents the high-resolution prediction feature map, represents the medium-resolution prediction feature map, Represents a low-resolution prediction feature map; S52, based on the extracted high-resolution prediction feature map, medium-resolution prediction feature map and low-resolution prediction feature map, obtain a feature map , the formula is as follows: ; in, Representation feature map The coordinate position of the center point of the bounding box, Representation feature map The width and height of the bounding box, Representation feature map The category of the bounding box Confidence level; S53. In order to merge between different resolutions, the bounding box coordinates of different resolutions are converted to the same scale. The normalization formula is as follows: ; ; ; ; in, represents the coordinates of the normalized bounding box, represents the height and width of the input image, and Representation feature map height and width; S54. Calculate the upper left corner coordinates and the lower right corner coordinates of the bounding box according to the standardized bounding box coordinates. The formula is as follows: ; ; ; ; in, Represents the coordinates of the upper left corner of the bounding box, Represents the coordinates of the lower right corner of the bounding box; S55. Filter out candidate bounding boxes whose confidence is greater than the threshold. The formula is as follows: ; in, represents the confidence threshold; S56. Filter the bounding box according to the specified category. The formula is as follows: ; in, Indicates the specified category. Indicates category, ; S57. Apply non-maximum suppression to the remaining bounding boxes. The formula is as follows:

[0053] in, represents the two different remaining bounding boxes after the above operations, Represents a bounding box and bounding box The intersection ratio of Represents a bounding box and bounding box The intersection area of Represents a bounding box and bounding box In this embodiment, the intersection over union (IoU) is used to measure the overlap between the predicted bounding box and the ground truth bounding box. IoU values ​​range from 0 to 1, with higher values ​​indicating higher positioning accuracy. An IoU value of 1.0 indicates perfect alignment. Typically, a threshold of 0.50 is used for IoU, which is used to define true positives in metrics such as mAP. Lower IoU values ​​indicate that the model has difficulty accurately locating objects, which can be improved by improving bounding box regression or increasing annotation accuracy.

[0054] S58. Sort all bounding boxes, select the bounding box with the highest confidence, and remove other bounding boxes whose intersection over union (IoU) is greater than the set threshold. Repeat this process until no bounding box meets the conditions, and return the remaining bounding boxes as the final detection result. The formula is as follows:

[0055] in, Represents the final detection result, and NMS represents the non-maximum suppression process.

[0056] In specific implementation, as a preferred embodiment of the present invention, in step S6, the designed comprehensive loss function is as follows:

[0057] in, represents the weight hyperparameter of the classification loss, represents the weight hyperparameter of the bounding box regression loss, represents the weight hyperparameter of the confidence loss, Represents classification loss, which is used to measure the performance of the model on the classification task. represents the bounding box regression loss, which is used to measure the performance of the model in locating the bounding box. Represents low confidence loss, which is used to measure the model's confidence prediction of the target's existence. represents the comprehensive loss, which is the weighted sum of classification loss, bounding box regression loss and confidence loss.

[0058] Example like Figure 2 As shown, the Precision-Recall Curve of YOLOv5s, YOLOv6s and the method of the present invention is shown. Figure 2In the figure, (a) represents the Precision-Recall curve using the YOLOv5s method; (b) represents the Precision-Recall curve using the YOLOv6s method; (c) represents the Precision-Recall curve using the method of the present invention; Precision and recall are important indicators for evaluating the performance of classification models. The Precision-Recall curve shows the trade-off between the precision and recall of the model at different thresholds, and is used for evaluation of unbalanced data sets or when focusing on positive class detection performance. Figure 2 It can be seen that the mAP@0.5 (mAP50) index of the proposed method in all classes is 0.637, which is higher than the mAP@0.5 (mAP50) index of the YOLOv5s method of 0.618 and higher than the mAP@0.5 (mAP50) index of the YOLOv6s method of 0.625. The proposed method performs well in both precision and recall, indicating that it has better performance in night target detection tasks.

[0059] like Figure 3 As shown, the target detection results of the present invention and other algorithms in indoor scenes at night are demonstrated. Figure 3 In the figure, (a) represents the result of using the YOLOv5s method to detect the target; (b) represents the result of using the YOLOv6s method to detect the target; (c) represents the result of using the method of the present invention to detect the target; (d) represents the real detection result. Figure 3 It can be seen that the mAP@0.5 (mAP50) index of the YOLOv5s method and the YOLOv6s method in all classes is the same, which is 0.3. However, the YOLOv5s method and the YOLOv6s method both mistakenly detect the cat in the night image as a dog. The mAP@0.5 (mAP50) index of the method of the present invention in all classes is 0.8, and the method of the present invention successfully detects the cat in the image.

[0060] like Figure 4 As shown, the target detection results of the present invention and other algorithms in outdoor scenes at night are demonstrated. Figure 4 In the figure, (a) represents the result of using the YOLOv5s method to detect the target; (b) represents the result of using the YOLOv6s method to detect the target; (c) represents the result of using the method of the present invention to detect the target; (d) represents the actual detection result. Figure 4It can be seen that the mAP@0.5 (mAP50) index of the method of the present invention is the same as that of the YOLOv5s method and the YOLOv6s method in all classes, which is 0.4. Although the YOLOv5s method, the YOLOv6s method and the method of the present invention all successfully detect the hull, the detection frame of the YOLOv5s method only marks half of the hull, and the YOLOv6s method generates three detection frames for the same hull, two of which have incorrect detection pixel ranges. In summary, the method of the present invention can robustly detect objects in night scenes and provide a more accurate target detection range.

[0061] This embodiment compares different algorithms from the objective indicators of accuracy, recall, mAP50, and mAP50-95; the accuracy quantifies the proportion of true positives in all positive predictions, and evaluates the model's ability to avoid false positives. The recall rate calculates the proportion of true positive predictions in all actual positive predictions, and measures the model's ability to detect all instances of a certain class. mAP50 is the average precision calculated at the intersection greater than the union (IoU) threshold of 0.50, and is a measure of the accuracy of the model that only considers "easy" detection. mAP50-95 is the average of the average precisions calculated at different IoU thresholds between 0.50 and 0.95, which fully reflects the performance of the model under different detection difficulties. As shown in the following table: Table 1 Comparison of the accuracy of the model of the present invention and other advanced algorithms

[0062] Table 2 Comparison of recall rates between the proposed model and other advanced algorithms

[0063] Table 3 Comparison of mAP50 between the proposed model and other advanced algorithms

[0064] Table 4 Comparison of mAP50-95 of the processing results of the proposed model and other advanced algorithms

[0065] In summary, the model of the present invention is superior to YOLOv5s and YOLOv6s in terms of accuracy, recall, mAP50, and mAP50-95. Specifically, the algorithm of the present invention has higher accuracy, recall, and average precision on the night object detection dataset, especially the improvement of mAP50-95 shows the consistency advantage of the model under different IoU thresholds.

[0066] Therefore, the present invention is more suitable for application in nighttime target detection tasks, and can more effectively detect and locate targets while reducing false positives and false negatives, and is particularly suitable for deployment and use in nighttime monitoring scenarios.

[0067] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A nighttime target detection method based on feature aggregation network, characterized in that: include: S1. Obtain a nighttime target detection dataset and randomly divide the dataset into a training set and a test set in a ratio of 7:3; S2, using image preprocessing operations to perform data set augmentation on the nighttime target detection data set obtained in step S1; S3, selecting the image after the preprocessing operation in step S2, and inputting the selected image into the trunk part based on feature aggregation design to perform general feature extraction, and extracting high-resolution initial features, medium-resolution initial features and low-resolution initial features; S4, inputting the high-resolution initial features, medium-resolution initial features and low-resolution initial features obtained in step S3 into the neck part based on pyramid feature fusion, and further extracting high-resolution medium-level features, medium-resolution medium-level features and low-resolution medium-level features with diversity and robustness; S5. Input the high-resolution intermediate features, medium-resolution intermediate features, and low-resolution intermediate features extracted in step S4 into the prediction head part designed based on the convolution layer to complete the target detection prediction.

2. A nighttime target detection method based on feature aggregation network according to claim 1, characterized in that: The method further comprises: S6. Use classification loss, bounding box regression loss, and confidence loss to design a comprehensive loss function to constrain the detection process based on the feature aggregation network.

3. The nighttime target detection method based on feature aggregation network according to claim 1 is characterized in that: Step S2 specifically includes: S21, using mosaic technology, randomly splice the four images into a new image, the formula is as follows: ; in, means randomly selecting four night images from the dataset, Indicates that according to the random layout Stitched together, Represents the result image of the mosaic data augmentation method; S22, using linear blending technology, the two images and their labels are linearly blended in a certain proportion, the formula is as follows: ; in, Represents the horizontal and vertical coordinates of the input original image, represents the horizontal and vertical coordinates after random affine transformation, Represents the parameters of the affine transformation matrix, which is used to control rotation, scaling, and shearing. represents the translation parameter; S23, using random affine transformation technology to perform rotation, translation, scaling and shearing operations, the formula is as follows: ; ; in, and Indicates that two original images are randomly selected from the nighttime target detection dataset. and express and The corresponding label, represents the mixing ratio randomly sampled from a uniform distribution between [0, 1]; S24, using image blur technology, select Gaussian blur to obtain a blurred image, the formula is as follows: ; in, represents the image after image blur, represents the normalization factor, represents the image blur radius, represents the weight of the Gaussian kernel, represents the convolution operation, represents a random raw image from the nighttime dataset; S25, using HSV color space enhancement technology, by changing the hue, saturation and brightness of the image to simulate different lighting conditions, the formula is as follows: ; ; ; in, Represents the hue, saturation, and brightness components of the original image; Indicates the hue, saturation, and brightness components of the enhanced image; represents the random adjustment parameter; S26, using random horizontal flipping technology to increase the model's robustness to mirror changes, the formula is as follows: ; in, represents the image after random horizontal flipping operation, Indicates the width of the image.

4. The nighttime target detection method based on feature aggregation network according to claim 1 is characterized in that: The backbone based on feature aggregation design in step S3 includes an aggregate feature extraction module, a convolutional feature extraction module, a two-layer feature aggregation module and a pyramid aggregation attention mechanism. Step S3 specifically includes: S31, the aggregation feature extraction module extracts efficient features from the input image, reorganizes the spatial information of the input image, reduces the width and height dimensions of the image by half through slicing and splicing operations, and increases the number of channels, where: The formula for the slice operation is as follows: ; ; ; ; in, Represents the original image, the size of the original image is ( C , H , W ),in C is the number of channels, H is the height, W is the width, Indicates that the original image is cut into 4 parts according to odd and even indexes. The dimensions are ( C , H / 2, W / 2); The formula for the splicing operation is as follows: ; in, Represents a splicing operation, represents the concatenated feature image, The dimensions are (4 C , H / 2, W / 2); The convolution operation is performed on the feature image after the splicing operation. The formula is as follows: ; in, represents the output feature image after convolution, Represents the convolution operation; S32. Design the function of the convolution feature extraction module. The formula is as follows: ; in, Represents the convolution feature extraction module function, Represents the input variable of the convolution feature extraction module function, represents the LeakyReLU activation function, Represents a batch normalization operation; S33. Design the function of the double-layer feature aggregation module. The formula is as follows: ; in, Represents a two-layer feature aggregation module function, Represents the input variable of the double-layer feature aggregation module function, Indicates the channel direction split operation; S34. Design the function of the pyramid aggregation attention mechanism. The formula is as follows: in, represents the pyramid aggregation attention mechanism function, represents the input variable of the pyramid aggregation attention mechanism function, Represents the maximum pooling operation; S35: Output feature image of step S31 Input into the convolution feature extraction module function and the double-layer feature aggregation module function in turn to extract high-resolution initial features. The formula is as follows: in, F 11 represents the extracted high-resolution initial features; S36, input the high-resolution initial feature F11 extracted in step S35 into the convolution feature extraction module function and the double-layer feature aggregation module function in sequence to extract the medium-resolution initial feature, the formula is as follows: in, represents the extracted medium-resolution initial features; S37, the medium resolution initial feature extracted in step S36 Input into the convolution feature extraction module function and the pyramid aggregation attention mechanism function in turn to extract low-resolution initial features. The formula is as follows: in, Represents the extracted low-resolution initial features.

5. The nighttime target detection method based on feature aggregation network according to claim 1 is characterized in that: The neck based on pyramid feature fusion in step S4 includes a convolution feature extraction module and a double-layer feature aggregation module. Step S4 specifically includes: S41. Design the function of the convolution feature extraction module. The formula is as follows: ; S42. Design a function for the double-layer feature aggregation module. The formula is as follows: ; S43, the low-resolution initial features extracted in step S37 Input into the double-layer feature aggregation module function and the convolution feature extraction module function in turn to extract low-resolution residual features. The formula is as follows: in, Represents the extracted low-resolution residual features; S44, the medium resolution initial features extracted in step S36 and the low-resolution residual features extracted in step S43 Input into the double-layer feature aggregation module function and the convolution feature extraction module function to extract the residual mid-resolution features. The formula is as follows: in, represents the extracted residual mid-resolution feature, UP () indicates upsampling operation; S45, the high-resolution initial features extracted in step S35 F 11 and the residual medium-resolution feature extracted in step S44 Input into the double-layer feature aggregation module function to extract high-resolution intermediate features. The formula is as follows: in, Represents the extracted high-resolution mid-level features; S46, the residual mid-resolution feature extracted in step S44 and the high-resolution intermediate features extracted in step S45 Input into the double-layer feature aggregation module function and the convolution feature extraction module function to extract medium-resolution intermediate features. The formula is as follows: in, Represents the extracted medium-resolution mid-level features; S47, the low-resolution residual features extracted in step S43 and the medium-resolution intermediate features extracted in step S46 Input into the double-layer feature aggregation module function and the convolution feature extraction module function to extract low-resolution intermediate features. The formula is as follows: in, Represents the extracted low-resolution mid-level features.

6. The nighttime target detection method based on feature aggregation network according to claim 1 is characterized in that: Step S5 specifically includes: S51, input the high-resolution intermediate features F21, medium-resolution intermediate features F22, and low-resolution intermediate features F23 extracted in step S4 into the prediction head designed based on the convolutional layer to complete the target detection prediction, and extract the high-resolution prediction feature map, the medium-resolution prediction feature map, and the low-resolution prediction feature map. The formula is as follows: ; ; ; in, represents the high-resolution prediction feature map, represents the medium-resolution prediction feature map, Represents the low-resolution prediction feature map; S52, based on the extracted high-resolution prediction feature map, medium-resolution prediction feature map and low-resolution prediction feature map, obtain a feature map , the formula is as follows: ; in, Representation feature map The coordinate position of the center point of the bounding box, Representation feature map The width and height of the bounding box, Representation feature map The category of the bounding box confidence level; S53. In order to merge between different resolutions, the bounding box coordinates of different resolutions are converted to the same scale. The normalization formula is as follows: ; ; ; ; in, represents the coordinates of the normalized bounding box, represents the height and width of the input image, and Representation feature map height and width; S54. Calculate the upper left corner coordinates and the lower right corner coordinates of the bounding box according to the standardized bounding box coordinates. The formula is as follows: ; ; ; ; in, represents the coordinates of the upper left corner of the bounding box, Represents the coordinates of the lower right corner of the bounding box; S55. Filter out candidate bounding boxes whose confidence is greater than the threshold. The formula is as follows: ; in, represents the confidence threshold; S56. Filter the bounding box according to the specified category. The formula is as follows: ; in, Indicates the specified category. Indicates category, ; S57. Apply non-maximum suppression to the remaining bounding boxes. The formula is as follows: in, represents the two different remaining bounding boxes after the above operations, Represents a bounding box and bounding box The intersection ratio of Represents a bounding box and bounding box The intersection area of Represents a bounding box and bounding box The area of ​​the union of S58. Sort all bounding boxes, select the bounding box with the highest confidence, and remove other bounding boxes whose intersection over union (IoU) is greater than the set threshold. Repeat this process until no bounding box meets the conditions, and return the remaining bounding boxes as the final detection result. The formula is as follows: in, Represents the final detection result, and NMS represents the non-maximum suppression process.

7. The nighttime target detection method based on feature aggregation network according to claim 2 is characterized in that: In step S6, the designed comprehensive loss function is as follows: in, represents the weight hyperparameter of the classification loss, represents the weight hyperparameter of the bounding box regression loss, represents the weight hyperparameter of the confidence loss, Represents classification loss, which is used to measure the performance of the model on the classification task. represents the bounding box regression loss, which is used to measure the performance of the model in locating the bounding box. Represents low confidence loss, which is used to measure the model's confidence prediction of the target's existence. represents the comprehensive loss, which is the weighted sum of classification loss, bounding box regression loss and confidence loss.

Citation Information

Patent Citations

  • Night vehicle detection method and system based on improved YOLOv5 convolutional neural network

    CN116524319A

  • Cloth defect detection method and system based on multi-scale feature fusion and diffusion pyramid network

    CN119515861A

  • Method and apparatus for obstacle detection under complex weather

    US20240005626A1