A Real-time Adaptive Clarity Method for Highway Video Surveillance

The real-time adaptive image enhancement method addresses the limitations of existing methods by aligning target features and maintaining scene integrity, improving target detection performance in highway video surveillance systems under adverse conditions.

CN117274084BActive Publication Date: 2025-07-15HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311203073.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-18
Publication Date
2025-07-15
Estimated Expiration
2043-09-18

AI Technical Summary

Technical Problem

The existing image clarification methods cannot effectively improve the performance of the target detection task when processing degraded images, especially in dense fog and low-illumination environments, which lack adaptability and target feature enhancement effects.

Method used

By training feature enhancement networks, the target feature difference between degraded images and clear images is sensed, unsupervised target feature alignment loss and overall feature maintenance constraints of scenes are introduced, and adaptive clarity processing of degraded images is realized.

Benefits of technology

The performance of target detection tasks in degraded images has been improved, the ability to identify target features has been enhanced, and the application effect of intelligent monitoring systems has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117274084B_ABST
    Figure CN117274084B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time adaptive clarity method for highway video surveillance, including: training a feature enhancement model. On the one hand, a reconstruction loss is constructed to constrain the input original image and the output enhanced image to be consistent in scene structure. On the other hand, the response features of the target in the clear image are used as constraints, and the features of the target in the high-level response of the detector are aligned through an unsupervised training method to improve the generalization performance of the model. At the same time, in order to determine whether the target in a certain scene needs to be enhanced, the confidence score output by the detector and the feature information entropy of the response are used as indicators to measure the scene target difference and achieve adaptive enhancement. The real-time adaptive clarity device for highway video surveillance implementing this method includes a scene target difference perception module and a scene target feature enhancement module. This device can be used as a preprocessing plug-in for the video surveillance system to provide clear image input for the target detection algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a real-time adaptive image enhancement method for highway video surveillance. Background Art

[0002] With the wide application of highway video surveillance, a large amount of video data is generated every day. To achieve an intelligent highway video surveillance system, advanced computer vision tasks such as object detection have become important auxiliary means. However, these intelligent algorithms usually require high-quality input images. Unfortunately, the imaging quality of highway video surveillance systems is easily affected by various degradation factors. At night, due to the low environmental illumination, the reflected light signal of the scene becomes weak, and the imaging contrast decreases. In thick fog weather, the scattering of atmospheric particles attenuates the scene signal and mixes in the atmospheric light, also resulting in a decrease in imaging contrast. The existence of these degradation factors limits the effective application scenarios of intelligent algorithms in highway video surveillance systems. To solve this problem, people need to focus on developing image enhancement technologies that can process degraded images to enhance or restore degraded images so that their quality meets the requirements of intelligent algorithms. By modeling the degradation factors and using physical model-based or data-driven methods, the quality of degraded images can be effectively improved, thereby enhancing the application effect of intelligent algorithms in highway video surveillance.

[0003] Currently, people are striving to enhance or restore degraded images through image enhancement methods to provide more effective information for the surveillance system. The main image enhancement methods are divided into two categories: physical model-based and data-driven methods.

[0004] Physical model-based methods need to physically model specific degradation factors and then eliminate the specific degradation factors by solving the characteristic components in the physical model. For example, in a night scene, the illumination component can be solved according to the retina cortex theory and then the brightness can be remapped; in the case of thick fog, the illumination component and the transmittance component can be solved and inversely deduced according to the atmospheric scattering model. These methods usually produce good visual effects because they have a deep understanding of the degradation process. However, there are still some degradation factors that cannot be modeled, resulting in limited application.

[0005] Data-driven methods mainly rely on generative models and are data-driven in a supervised or semi-supervised manner. These methods implicitly model the degradation process by designing a loss function to achieve the removal of degraded images. The advantage of this method is that it can handle some complex degradation situations and does not require in-depth knowledge of the physical principles of the degradation process.

[0006] However, the imaging process of an image may be affected by various factors, such as color bias noise caused by the sensor and changes in ambient light. These effects are difficult to accurately quantify, and ignoring these effects in the degradation model will result in obvious visual differences between the restored image and the real clear image. In addition, in the feature domain, there will also be more differences between the restored image and the clear image. Currently, existing degradation restoration methods mainly focus on enhancing the apparent features of images. However, in practical applications, the target analysis task of the monitoring system pays more attention to target features. Due to the difference between the pixel domain and the target feature domain, the performance of the target analysis algorithm is limited. Existing target feature enhancement methods lack attention to scene target features in the alignment of scene features, which may affect the generalization ability of the feature enhancement model. In addition, existing target feature enhancement methods do not perceive the differences in scene targets and cannot determine whether the targets in this scene need to be feature-enhanced, lacking environmental adaptability. Summary of the Invention

[0007] In this context, the inventor, aiming at the target detection task and considering the improvement of the performance of machine analysis tasks by target feature enhancement, further enhances the degraded image from the aspects of adaptive perception of scene target differences and alignment of target semantic features on the basis of enhancing the apparent features of the degraded image. Furthermore, a real-time adaptive clarification method and device for highway video surveillance are proposed to realize image clarification in different harsh environments such as thick fog and low illumination, improve the performance of target detection tasks in these scenarios, and are of great significance for intelligent monitoring systems.

[0008] The technical problem to be solved by the present invention is to improve the performance of the target detection task through image clarification processing. A real-time adaptive clarification method and device for highway video surveillance are proposed. First, the differences between the target features of the degraded image affected by factors such as thick fog and low illumination and the target features of the clear image are perceived to make a decision on whether enhancement is needed. Then, by introducing an unsupervised target feature alignment loss and at the same time introducing a constraint on maintaining the overall scene features, the enhancement model pays more attention to the targets in the scene, realizes the enhancement of target features in the degraded scene, and improves the performance of subsequent target detection tasks on the clarified processing results.

[0009] According to one aspect of the present invention, there is provided a real-time adaptive clarification method for highway video surveillance, which is characterized by comprising:

[0010] A) Training a feature enhancement network to obtain a feature enhancement model G for the degraded scene A. This step includes:

[0011] A1) Selecting sample images a and b from the degraded scene A and the clear scene B respectively;

[0012] A2) Input the sample image a into the feature enhancement model G to obtain the enhanced image a1. Then input the enhanced image a1 and the sample image a into the scene preservation module, which will impose consistency constraints on a1 and a in terms of scene pixel grayscale, scene structure perception, and color.

[0013] A3) Input the sample image b with annotation information into the pre-trained detector to obtain the target feature vector corresponding to the detector's response according to the label information. Input the sample image a1 without annotation information into the detector. Based on the high-confidence target boxes detected by the detector, correct the candidate boxes corresponding to the detector's response. Then, according to the corrected candidate boxes, obtain the target feature vector corresponding to the detector's response. After that, align the targets based on the feature vectors of the target responses in the sample image b and the sample image a1.

[0014] B) Scene target difference measurement. Perceive the scene target difference based on the detection confidence and response feature information entropy of the degraded scene target and the clear scene target as statistics.

[0015] C) During the test phase, sample the image data of the test scene. After target difference measurement, determine whether enhancement is needed. If enhancement is needed, use the scene data as the input to the feature enhancement model G to obtain the clarity processing result J of the present invention. Description of the Drawings

[0016] Figure 1 It is a schematic flowchart of a real-time adaptive clarity method for highway video surveillance according to the present invention.

[0017] Figure 2 It is a schematic diagram of the module structure of the real-time adaptive clarity method for highway video surveillance according to the present invention.

[0018] Figures 3(a) - 3(c) are schematic diagrams showing the enhancement effects of the image clarity method according to an embodiment of the present invention. Among them, Figure 3(a) is the detection result of the original low-light image, Figure 3(b) is the detection result after the apparent feature enhancement of the original low-light image, and Figure 3(c) is the detection result after the feature enhancement of the present invention. Detailed Embodiment

[0019] As Figure 1 shown, a real-time adaptive clarity method for highway video surveillance includes:

[0020] A) Train the feature enhancement network to obtain the feature enhancement model G for the degraded scene A. This step specifically includes:

[0021] A1) Select the sample images a and b from the degraded scene A and the clear scene B respectively;

[0022] A2) Input the sample image a into the feature enhancement model G to obtain the enhanced image a1. Then input the enhanced image a1 and the sample image a into the scene preservation module, which will impose consistency constraints on a1 and a in terms of scene pixel grayscale, scene structure perception, and color.

[0023] A3) Input the sample image b with annotation information into the pre-trained detector to obtain the target feature vector corresponding to the detector's response according to the label information. Input the sample image a1 without annotation information into the detector. According to the high-confidence target boxes detected by the detector, correct the candidate boxes corresponding to the detector's response. Then, according to the corrected candidate boxes, obtain the target feature vector corresponding to the detector's response. After that, align the targets based on the feature vectors of the target responses in the sample image b and the sample image a1.

[0024] B) Scene target difference measurement, including perceiving the scene target difference based on the detection confidence and response feature information entropy of the degraded scene target and the clear scene target as statistics.

[0025] C) During the test phase, sample the image data of the test scene. After target difference measurement, determine whether enhancement is needed. If enhancement is needed, input the scene data into the feature enhancement model G to obtain the clarification processing result J of the present invention.

[0026] According to a further embodiment of the present invention, in the above step A), the feature enhancement model G is composed of two convolutional layers that change the number of feature map channels and 4 residual blocks that do not change the size of the feature map. The activation function is ReLU, and the convolution kernel size is 3*3. It is a very lightweight network model. First, there is a convolutional layer that changes the number of feature map channels, which contains 64 convolutional kernels and is used to increase the number of channels of the feature map to 64. Subsequently, there are 4 residual blocks, each of which consists of two convolutional layers with 64 convolutional kernels. Then, there is a convolutional layer that changes the number of feature map channels, which contains 3 convolutional kernels and is used to reduce the number of channels of the feature map to 3. Its output serves as the final output of the model.,,

[0027] According to a further embodiment of the present invention, in the above step A1), the two input image domains are unpaired degraded images and clear images.

[0028] According to a further embodiment of the present invention, in the above step A2), the scene preservation module mainly realizes enhanced image reconstruction to ensure that the enhanced image and the original image are consistent in terms of scene structure. Among them, the constraint on pixel grayscale is

[0029]

[0030] Among them, I(x) represents the original image, and G(I(x)) represents the enhanced image. In this example, I(x) corresponds to image a, and G(I(x)) corresponds to image a1. The scene structure perception constraint uses a pre-trained VGG19 model, taking the feature maps of each activation layer of the restored image and the enhanced image generated by the model in the VGG19 network as constraints, so that the generated image is consistent with the original restored image in the overall structure:

[0031]

[0032] Among them, K represents the number of activation layers, k represents the k-th activation layer of the VGG19 model, and VGG k represents the response feature of the k-th activation layer. The color constraint for the scene is

[0033]

[0034] Among them, are the pixel means of the generated image a1 in the three color channels respectively. According to the gray world assumption, the three-channel means of a white-balanced image are approximately equal, and the difference between the maximum and minimum values of the three-channel means is used to constrain the overall color cast problem of the scene.

[0035] According to a further embodiment of the present invention, in the above step A3), for the Faster RCNN detector, the input sample image b is passed through the backbone network to obtain the global feature basefeat, and the candidate boxes output by the detector RPN module are obtained according to the label information label of the sample image b:

[0036] Bbox candidate =F rpn (basefeat,label)

[0037] After that, the features corresponding to the target responses are found on the global feature according to the candidate boxes, and the target feature vector f linear is obtained through a fully connected layer F b :

[0038] f b =F linear (basefeat,Bbox candidate )

[0039] For the input sample image a1, since there is no annotation information, the candidate boxes predicted by the detector RPN module are corrected according to the detection results with high confidence of the pre-trained detector model on the image a1, so that the module outputs more accurate candidate boxes, and then the target feature vector f a1 of the image a1 is obtained in the same way.

[0040] After obtaining the target feature vectors of the sample image b and the sample image a1, the target feature alignment loss is introduced:

[0041]

[0042] where represents the mean of the target feature vectors of the cls-th class in the response of image b in the detector, represents the mean of the target feature vectors of the cls-th class in the response of image a1 in the detector. This loss, together with the three losses in step A2), constitutes the optimization function of the model. The model parameter optimizer is Adam, the learning rate is 2e-4, and the model converges after about 50 rounds of iterative training. The clear image samples in the training data come from the VOC2007 dataset, the heavily fog-degraded image samples come from the RTTS dataset, and the low-illumination degraded image samples come from the exdark dataset.

[0043] According to a further embodiment of the present invention, in step B), for an input image, the average confidence of its object detection result can be expressed as

[0044]

[0045] N cand represents the number of candidate boxes, F score (Bbox i ) is the confidence score of the i-th candidate box, T is the confidence threshold, and candidate boxes with too small scores are excluded, set to 0.02. For an input image, the information entropy of the feature map where the object in it responds in the detector can reflect the amount of information expressed by the object in the detector:

[0046]

[0047] p i represents the probability that the i-th value appears in the feature. According to the definitions of the above two statistics, offline statistics are performed on the restored degraded image data and clear image data, and the statistic thresholds are obtained according to the 3σ rule to determine the decision conditions:

[0048] D = δ S<0.55orE<8.85

[0049] That is, when the condition δ S<0.55orE<8.85 is satisfied, the data of this scenario needs to be feature enhanced. In this step, the degraded image data comes from the heavily fogged images of the RTTS dataset, and the clear image data comes from the VOC2007 dataset.

[0050] According to a further embodiment of the present invention, in step C), during the test phase, the image data of the test scenario is first sampled, and then the confidence level and the feature information entropy of the image data are statistically analyzed online to determine whether the scenario needs to be enhanced. If enhancement is required, the degraded image is used as the input to the feature enhancement model G to obtain the sharpened processing result J of the present invention.

[0051] According to another aspect of the present invention, there is provided a real-time adaptive sharpening device for highway video surveillance, as Figure 2 shown, the device includes:

[0052] A scene target difference perception module and a scene target feature enhancement module. The scene target difference perception module is mainly used to determine whether the current scene needs to be enhanced. The scene target feature enhancement module realizes the enhancement of the scene target.

[0053] The scene target difference perception module mainly receives the video image data of the highway scene and determines whether the scene needs to be enhanced based on the scene target difference measurement method in the above step B).

[0054] The scene target feature enhancement module uses the degraded image as the input and realizes the enhancement of the scene target based on the feature enhancement method in the above step A).

[0055] Figures 3(a) - 3(c) show the enhancement effect of the image sharpening method according to an embodiment of the present invention. (In practice, 3(a) - 3(c) are color pictures.) By comparing the detection results in Figures 3(a) and 3(b), it can be seen that after the apparent feature enhancement is restored, although the visual quality improves, the improvement in detection performance is limited. The vehicle on the right still cannot be detected after the apparent feature enhancement, while in Figure 3(c), after the target feature enhancement of the present invention, the vehicle on the right can be detected.

[0056] It should be noted that the above-disclosed are only specific implementation examples of the present invention. According to the technical idea provided by the present invention, the changes that can be conceived by those of ordinary skill in the art should fall within the protection scope of the present invention.

Claims

1. A real-time adaptive clarity method for highway video surveillance, characterized in that Including: A) Training a feature enhancement network to obtain a feature enhancement model G for the degraded scene A, specifically including: A1) Selecting sample images a and b from the degraded scene A and the clear scene B respectively; A2) Inputting the sample image a into the feature enhancement model G to obtain an enhanced image a1, and inputting the enhanced image a1 and the sample image a into the scene preservation module, which will perform consistency constraints on the scene pixel gray level, scene structure perception, and color of a1 and a; A3) Inputting the sample image b with annotation information into the pre-trained detector, and obtaining the target feature vector corresponding to the detector response according to the label information; inputting the enhanced image a1 without annotation information into the detector, correcting the candidate box corresponding to the detector response according to the high-confidence target box detected by the detector, and obtaining the target feature vector corresponding to the detector response according to the corrected candidate box; then, aligning the targets according to the feature vectors of the target responses in the sample image b and the sample image a1; B) Measuring the scene target difference, and perceiving the scene target difference according to the detection confidence and response feature information entropy of the degraded scene target and the clear scene target as statistics; C) Sampling the image data of the test scene in the test stage, and judging whether enhancement is needed through the target difference measurement. If enhancement is needed, the scene data is used as the input of the feature enhancement model G to obtain the clarification processing result J of the present invention. Wherein: In the above step A1), the input sample images a and b are unpaired degraded images and clear images; The scene preservation module performs enhanced image reconstruction to ensure that the enhanced image and the original image are consistent in scene structure. Among them, the constraint on the pixel gray level is: Where I(x) represents the original image, that is, the corresponding sample image a, and G(I(x)) represents the enhanced image, that is, the corresponding image a1; The scene structure perception constraint adopts a pre-trained VGG19 model, and uses the feature maps of each activation layer of the VGG19 network of the restored image and the enhanced image generated by the model as constraints to make the generated image consistent with the original restored image in the overall structure: Among them, K represents the number of activation layers, k represents the k-th activation layer of the VGG19 model, and VGG k represents the response feature of the k-th activation layer. The color constraint on the scene is: Among them, are the pixel means of the generated image a1 in three color channels respectively. Assuming that the three-channel means of the white-balanced image are approximately equal, the difference between the maximum and minimum values of the three-channel means is used to constrain the overall color cast problem of the scene. Step A3) includes: for the FasterRCNN detector, inputting the sample image b, obtaining the global feature basefeat after passing through the backbone network, and obtaining the candidate box output by the detector RPN module according to the label information label of the sample image b; Bbox candidate = F rpn (basefeat, label) After that, the feature corresponding to the target response is found on the global feature according to the candidate box, and passes through a fully connected layer F linear Obtain the target feature vector f b : f b = F linear (basefeat, Bbox candidate ) Among them, F linear is a fully connected layer. For the input enhanced image a1, since there is no annotation information, the candidate boxes predicted by the detector RPN module are corrected according to the detection results with high confidence of the pre-trained detector model on the image a1, so that the module outputs more accurate candidate boxes. Subsequently, in the same way, the target feature vector f of the image a1 is obtained a1 , After obtaining the target feature vectors of the sample image b and the enhanced image a1, introducing the target feature alignment loss: The lower bound of the summation is 0, and there is no upper bound; Among them represents the mean of the target feature vectors of the cls-th class in the response of image b in the detector, represents the mean of the target feature vectors of the cls-th class in the response of image a1 in the detector. This loss, together with the pixel gray-scale loss Loss pix , the perceptual constraint loss Loss percep and the color constraint loss Loss color These three losses together constitute the optimization function of the model. The model parameter optimizer is Adam, the learning rate is 2e-4, and the model converges after about 50 rounds of iterative training. The clear image samples in the training data come from the VOC2007 dataset, the heavily fog-degraded image samples come from the RTTS dataset, and the low-illumination degraded image samples come from the exdark dataset. In step B), for an image collected in a degraded scene or a clear scene, the average confidence of its target detection result can be expressed as: The lower limit of summation is 0, the upper limit is 1, and N cand represents the number of candidate boxes, F score (Bbox i ) is the confidence score of the i-th candidate box, T is the confidence threshold, candidate boxes with too small scores are excluded, set to 0.

02. For an input image, the information entropy of the feature map where the target responds in the detector, which can reflect the amount of information expressed by the target in the detector: p i Indicates the probability that the i-th value appears in the feature, According to the definition of the average confidence S of the target detection result and the information entropy E of the feature map corresponding to the response in the detector, offline statistics are performed on the restored degraded image data and clear image data, and the statistical threshold is obtained according to the 3-sigma rule to determine the decision condition. D = δ S<0.55orE<8.85 , That is, when the condition δ S<0.55orE<8.85 is satisfied, the data of this scenario needs to be feature enhanced. Among them, the degraded image data comes from the thick fog images of the RTTS dataset, and the clear image data comes from the VOC2007 dataset. In step C), in the testing phase, the image data of the test scenario is first sampled, and then the confidence level and the feature information entropy of the image data are statistically analyzed online to determine whether the scenario needs to be enhanced. When enhancement is required, the degraded image is used as the input of the feature enhancement model G to obtain the sharpened processing result J.

2. The image sharpening processing method according to claim 1, characterized in that: The feature enhancement model G is composed of 4 residual blocks and is a very lightweight network model.

Citation Information

Patent Citations

  • Single-image defogging method based on ResNet neural network

    CN108230264A

  • A video monitoring-oriented moving object active sensing method and system

    CN109887040A