A video object detection method in extreme dark light environment based on feature supervision

By combining feature supervision and anti-noise branches, the noise interference problem of video target detection in extreme dark light environments is solved, efficient video target detection effect is achieved, and detection accuracy is improved.

CN116469023BActive Publication Date: 2025-09-19NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210018825.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-09
Publication Date
2025-09-19
Estimated Expiration
2042-01-09

AI Technical Summary

Technical Problem

In extremely dark environments, existing technologies find it difficult to effectively detect video targets. The reason is that the sensor receives few photons, resulting in a low signal-to-noise ratio. The noise seriously affects the image quality. Existing denoising methods have large parameters and lose detailed texture information, resulting in reduced detection accuracy.

Method used

A feature-based supervision method is adopted. By training video data without adding synthetic dark light noise in the video target detection algorithm model, the backbone network parameters are fixed, clean features are obtained and the noisy backbone network is supervised. Anti-noise branches are added to enhance the anti-noise performance, and the spatiotemporal aggregation module is combined to improve the detection accuracy.

Benefits of technology

It effectively resists extreme dark light noise interference, reduces the number of parameters and calculations, makes full use of video redundant information, improves detection accuracy, and increases detection accuracy by 16.3%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116469023B_ABST
    Figure CN116469023B_ABST
Patent Text Reader

Abstract

This invention discloses a method for video object detection in extreme low-light environments based on feature supervision. This method divides the training of the feature supervision strategy into two stages: in the first stage, a set of weight parameters is trained on the SELSA model of the video object detection algorithm using video data without synthetic low-light noise. In the second stage, the video without synthetic low-light noise is fed into the backbone network of the SELSA model trained in the first stage, and its parameters are fixed without backpropagation optimization. Clean features are then obtained at different depths of the backbone network. The video with synthetic low-light noise is then fed into a new, noisy backbone network to be trained to obtain noise features at different depths. Finally, the clean features are used to supervise the noise features at the corresponding depths to improve the noise resistance of the noisy backbone network. This method can significantly reduce the number of network parameters, computational complexity, and inference time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer vision, and in particular to a video target detection method in an extreme dark light environment based on feature supervision. Background Art

[0002] Object detection in extreme low-light environments is a challenging problem in computer vision, widely applicable to nighttime surveillance and navigation equipment. However, existing research on this topic is extremely scarce. This is primarily due to the low signal-to-noise ratio (SNR) generated by sensors in extremely low-light conditions, which severely impacts image quality and significantly reduces detection accuracy.

[0003] Simply considering denoising extreme low-light videos before performing object detection presents two problems: 1. Neural network-based video denoising methods require a significant number of parameters and computational complexity, making them unsuitable for real-time detection tasks. 2. While denoising first improves the visual experience, it also loses a significant amount of useful texture detail, resulting in blurry and distorted images and a reduction in the accuracy of the subsequent detection network. Summary of the Invention

[0004] In response to the challenges brought about by solving the above-mentioned problem of video target detection in extreme dark light environments, the purpose of the present invention is to propose a video target detection method in extreme dark light environments based on feature supervision.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A method for detecting video targets in extreme dark environments based on feature supervision includes the following steps:

[0007] Step 1: Use video data without synthetic dark light noise to train the video object detection algorithm model and save the weight parameters of the model network;

[0008] Step 2: Obtain a clean video without synthetic dark light noise and a video with synthetic dark light noise. Input the clean video without synthetic dark light noise into the backbone network of the video object detection algorithm model trained in step 1, fix its parameters and do not perform backpropagation optimization. Then, obtain clean features at different depths of the backbone network. Input the video with synthetic dark light noise into a new noisy backbone network to be trained to obtain noise features at different depths. Then, use the clean features to supervise the noise features at the corresponding depths to improve the noise resistance performance of the noisy backbone network.

[0009] Step 3: After completing the network training in steps 1 and 2, test inference is performed. That is, the noisy video frame in an extreme dark light environment and its adjacent reference frames are input into the noisy backbone network trained in step 2 to obtain high-semantic features.

[0010] Step 4: The high-semantic features output by the noisy backbone network are fed into the proposal generation network and proposal alignment module. The proposal generation network generates some location box information where objects may be located based on the high-semantic features. These location box information is fed into the proposal alignment module, which extracts the corresponding location features from the original feature map and then pools them into a uniform size to output as proposal features.

[0011] Step 5: Send the proposed features output by the proposal alignment module to the spatiotemporal aggregation module of the video object detection algorithm to aggregate the relevance of the proposed features in the spatiotemporal dimension, thereby improving the accuracy of video detection;

[0012] In step 6, based on the features output by the spatiotemporal aggregation module, the object position box of the proposed features is classified and regressed, and the category and position information of the proposed box are output.

[0013] Furthermore, in step 2, in order to further improve the anti-noise performance of the noisy backbone network, an enhanced anti-noise branch is constructed in the noisy backbone network. The enhanced anti-noise branch is composed of multiple enhanced anti-noise sub-modules. Each enhanced anti-noise sub-module collects the noise characteristics at the current backbone network depth and the output characteristics of the previous level enhanced anti-noise sub-module, and after fusing the information of the two and performing spatial domain anti-noise and time domain anti-noise, the enhanced characteristics are output.

[0014] Furthermore, in step 2, the specific process of processing is:

[0015] In step 2-1, the input features of the enhanced anti-noise submodule can be expressed as:

[0016]

[0017] Among them, s∈{0,1,2,3,4} represents the step index of the backbone network, t∈[-T,T] represents the continuous video frames in the time dimension, and the total length is 2T+1. represents the input features of the enhanced anti-noise submodule, represents the noisy characteristics of the current step of the backbone network, represents the noisy characteristics of the previous step of the backbone network, represents the output feature after the previous stage enhanced anti-noise submodule enhances the anti-noise, Concat(· represents the feature channel dimension merging operation;

[0018] In step 2-2, the spatial dimension of the input features of the enhanced anti-noise submodule is enhanced and denoised using dense residual connections. For the anti-noise operation in the temporal dimension, the temporal self-attention mechanism is used. The temporal self-attention mechanism can be expressed by the following formula:

[0019]

[0020] Among them, i,j∈[-T,T] is the length of the current video frame, i represents the target frame that is about to be subjected to anti-noise, and j represents the support frame for the fusion of time domain information. represents the characteristics after anti-noise, is the feature before anti-noise, is the weight mapping function between i and j, and ⊙ is the multiplication operation of the corresponding feature pixel value;

[0021] Weight mapping function It can be further expressed as:

[0022]

[0023]

[0024] in, is the pre-weight coefficient, offset(· is the feature and The features after convolution operation after merging are used to obtain the bias information required for deformable convolution. DConv(·) is the deformable convolution used for alignment and Features, Emb(· consists of convolutional layers to map features;

[0025] Step 2-3, for the features of the enhanced anti-noise submodule output The mean square error loss is used for constraint, which can be expressed as follows:

[0026]

[0027] in Indicates the supervised features at the corresponding depth of the clean backbone network.

[0028] Based on the video object detection algorithm (SELSA), the present invention can effectively resist noise interference in extreme dark light noise environments by using constrained features and adding anti-noise branches during the backbone network (Backbone) training process. It not only effectively reduces the overhead in terms of parameter quantity and computational complexity, but also can fully utilize the redundant information of the original dark light noise video to improve the accuracy of the detection task. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 2 is a neural network structure diagram of the extreme dark light detection method according to an embodiment of the present invention;

[0030] Figure 2 This is the backbone training diagram of the extreme dark light detection network in an embodiment of the present invention. DETAILED DESCRIPTION

[0031] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0032] Reference Figure 1 2. The video target detection method in an extreme dark light environment based on feature supervision of this embodiment has the following specific steps:

[0033] Step 1: Use video data without adding synthetic dark light noise to train the video object detection algorithm (SELSA) model and save the weight parameters of the model network.

[0034] The video object detection algorithm SELSA is an improvement on the two-stage FasterRCNN algorithm for image object detection. The two stages of the FasterRCNN algorithm are as follows: In the first stage, a backbone network (typically ResNet) is used to extract high-semantic features from the image. In the second stage, the high-semantic features extracted by the backbone network are passed through a proposal generation network (Region Proposal Network, RPN), a proposal alignment module (Region of Interest Align, ROI Align), and classification and regression modules to obtain the detected object category and location information. Unlike image object detection, the SELSA algorithm targets video object detection, taking into account the efficient use of information in both temporal and spatial dimensions. Following the proposal alignment module in the second stage of the FasterRCNN algorithm, a spatiotemporal aggregation module is added. This performs spatiotemporal aggregation operations on each proposed feature, improving the accuracy of video object detection.

[0035] ResNet is used as the backbone network. The ResNet network structure can be divided into five stages. As the network depth increases, the output feature resolution of each stage becomes smaller, the number of channels increases, and the semantics becomes higher. This embodiment uses ResNet50 as the backbone network of the algorithm model.

[0036] Step 2: Obtain a video without synthetic dark light noise and a video with synthetic dark light noise. Input the clean video without synthetic dark light noise into the backbone network of the video target detection algorithm model trained in step 1 and fix its parameters without backpropagation optimization. Then obtain clean features at different depths of the backbone network. Input the video with synthetic dark light noise into the new noisy backbone network to be trained to obtain noise features at different depths.

[0037] The added synthetic dark noise is given by the following noise model:

[0038]

[0039] Where i represents the position index of the i-th pixel on the image, y i Indicates the value of the pixel after adding dark light noise, c is the RGB three-channel index of the image, represents the photon shooting noise, which obeys the expectation Poisson distribution, Represents dark current noise, which has a mean value of N d Poisson distribution, represents the sensor readout noise, which obeys a Gaussian distribution with a mean of 0 and a variance of σ. Represents stripe noise, which has a mean of 1 and a variance of σ beta Gaussian distribution, K c is a system global constant, Indicates truncation operation, truncating negative numbers to zero.

[0040] An enhanced anti-noise branch is added to the above noisy backbone network. This branch consists of enhanced anti-noise submodules (EnhancedDenoisingBlock, EDB). The input of any enhanced anti-noise submodule is:

[0041]

[0042] Among them, s∈{0,1,2,3,4} represents the stage index of the backbone network, t∈[-T,T] represents the continuous video frames in the time dimension, with a total length of 2T+1. represents the input features of the enhanced anti-noise submodule, represents the noisy characteristics of the current step of the backbone network, represents the noisy characteristics of the previous step of the backbone network, It represents the output feature after the anti-noise enhancement of the previous anti-noise submodule, and Concat(·) represents the feature channel dimension merging operation.

[0043] Each enhanced anti-noise submodule contains spatial anti-noise and temporal anti-noise. Spatial anti-noise consists of dense residual connection convolutional layers, and temporal anti-noise consists of temporal self-attention mechanism. The temporal self-attention mechanism can be expressed as follows:

[0044]

[0045] Among them, i,j∈[-T,T] is the length of the current video frame, i represents the target frame that is about to be subjected to anti-noise, and j represents the support frame for the fusion of time domain information. represents the characteristics after anti-noise, is the feature before anti-noise, is the weight mapping function between i and j, and ⊙ is the multiplication operation of the corresponding feature pixel value.

[0046] The weight mapping function can be further expressed as:

[0047]

[0048]

[0049] The previous formula is the pre-weight coefficient in the time domain Perform softmax operation to obtain weights The latter is the specific pre-weight coefficient Acquisition method, where offset(·) is a feature The features after convolution operation after merging are used to obtain the bias information required for deformable convolution. DConv(·) is the deformable convolution used for alignment Features, Emb(·) consists of convolutional layers to map features.

[0050] For the feature supervision strategy during training, the clean features output by the first backbone network are used to supervise the features output by the enhanced anti-noise submodule. The mean square error loss (MSELoss) is used for constraint, which can be expressed as follows:

[0051]

[0052] in Indicates the supervised features at the corresponding depth of the clean backbone network.

[0053] Finally, the output features of the noisy backbone can be expressed as:

[0054]

[0055] in, is the output of the noisy backbone, They are the output features of the last step and the output features of the last enhanced denoising module respectively.

[0056] Step 3: After completing the training of the extreme low-light video object detection network in steps 1 and 2, enter the test inference phase. The inference phase does not require the pre-trained backbone network and its clean features in step 1 as supervision. Instead, it only needs to input the noisy video frame in the extreme low-light environment and its adjacent reference frames into the noisy backbone network trained in step 2 to obtain high-semantic features.

[0057] In step 4, the high-semantic features output by the noisy backbone network are fed into the proposal generation network and proposal alignment module. The proposal generation network generates some location box information where objects may be located based on the high-semantic features. This location information is fed into the proposal alignment module to extract the corresponding location features from the original feature map. The features are then pooled to a uniform size and output as proposal features.

[0058] Step 5: The proposed features output by the proposal alignment module are fed into the spatiotemporal aggregation module in the SELSA algorithm. The correlation of the proposed features in the spatiotemporal dimension is expressed as follows:

[0059]

[0060]

[0061] Among them, k,l∈[-T,T] represents the two frame indices in the time domain, where k is the target frame, l is the reference frame, and i,j are the two proposed feature vector indices on the kth and lth frames respectively. is the i-th proposed feature vector on the target frame k, is the j-th proposed feature vector on the reference frame l, φ(·), The vector mapping function is composed of a fully connected layer, The weight coefficient obtained by taking the cosine similarity of the two proposed feature vectors after mapping, and finally It represents the aggregation result of the i-th proposed feature on the target frame k and all the proposed features in the time domain.

[0062] In step 6, based on the features output by the spatiotemporal aggregation module, the object position box classification (Classification) and regression (Regression) of the proposed features are performed, and the category and location information of the proposed box are output.

[0063] This method can directly suppress the interference of low-light environment noise in the backbone network of the detection algorithm model. Compared with the network model that denoises first and then detects (FastDVDNet denoising algorithm + SELSA detection algorithm), it directly omits the redundant denoising network, which can greatly reduce video memory usage, computational complexity and inference speed. It also improves detection accuracy by 16.3%. The average precision (AP) indicators of these two methods on the same test data are 64.3 and 74.8, respectively. The following table shows the specific performance and indicator comparison results of the two methods.

[0064] Table 1 Comparison results of different methods

[0065]

Claims

1. A method for video target detection in extreme dark environments based on feature supervision, characterized in that: The steps include: Step 1: Use video data without synthetic dark light noise to train the video object detection algorithm model and save the weight parameters of the model network; Step 2: obtain a clean video without synthetic dark light noise and a video with synthetic dark light noise added, input the clean video without synthetic dark light noise added into the backbone network of the video target detection algorithm model trained in step 1 and fix its parameters without back propagation optimization, and then obtain clean features at different depths of the backbone network; input the video with synthetic dark light noise added into the new noisy backbone network to be trained to obtain noise features at different depths, and then use the clean features to supervise the noise features at the corresponding depths to improve the anti-noise performance of the noisy backbone network; in order to further improve the anti-noise performance of the noisy backbone network, construct an enhanced anti-noise branch in the noisy backbone network, which is composed of multiple enhanced anti-noise sub-modules, each enhanced anti-noise sub-module collects the noise features at the current backbone network depth and the output features of the previous level enhanced anti-noise sub-module, fuses the information of the two and outputs the enhanced features after spatial domain anti-noise and temporal domain anti-noise; Step 3: After completing the network training in steps 1 and 2, test inference is performed. That is, the noisy video frame in an extreme dark light environment and its adjacent reference frames are input into the noisy backbone network trained in step 2 to obtain high-semantic features. Step 4: The high-semantic features output by the noisy backbone network are fed into the proposal generation network and proposal alignment module. The proposal generation network generates some location box information where objects may be located based on the high-semantic features. These location box information is fed into the proposal alignment module, which extracts the corresponding location features from the original feature map and then pools them into a uniform size to output as proposal features. Step 5: Send the proposed features output by the proposal alignment module to the spatiotemporal aggregation module of the video object detection algorithm to aggregate the relevance of the proposed features in the spatiotemporal dimension, thereby improving the accuracy of video detection; In step 6, based on the features output by the spatiotemporal aggregation module, the object position box of the proposed features is classified and regressed, and the category and position information of the proposed box are output.

2. The method for video target detection in extreme dark light environments based on feature supervision according to claim 1, characterized in that: In step 2, the specific process of processing is: In step 2-1, the input features of the enhanced anti-noise submodule can be expressed as: Among them, s∈{0,1,2,3,4} represents the step index of the backbone network, t∈[-T,T] represents the continuous video frames in the time dimension, and the total length is 2T+1. represents the input features of the enhanced anti-noise submodule, represents the noisy characteristics of the current step of the backbone network, represents the noisy characteristics of the previous step of the backbone network, represents the output feature after the previous stage enhanced anti-noise submodule enhances the anti-noise, and Concat(·) represents the feature channel dimension merging operation; In step 2-2, the spatial dimension of the input features of the enhanced anti-noise submodule is enhanced and denoised using dense residual connections. For the anti-noise operation in the temporal dimension, the temporal self-attention mechanism is used. The temporal self-attention mechanism can be expressed by the following formula: Among them, i,j∈[-T,T] is the length of the current video frame, i represents the target frame that is about to be subjected to anti-noise, and j represents the support frame for the fusion of time domain information. represents the characteristics after anti-noise, is the feature before anti-noise, is the weight mapping function between i and j, and ⊙ is the multiplication operation of the corresponding feature pixel value; Weight mapping function It can be further expressed as: in, is the pre-weight coefficient, offset(·) is the feature and The features after convolution operation after merging are used to obtain the bias information required for deformable convolution. DConv(·) is the deformable convolution used for alignment and Features,Emb(·) consists of convolutional layers to map features; Step 2-3, for the features of the enhanced anti-noise submodule output The mean square error loss is used for constraint, which can be expressed as follows: in Indicates the supervised features at the corresponding depth of the clean backbone network.