A method for enhancing images of road monitoring at night
By combining multimodal data fusion and GAN models with super-resolution and spatiotemporal continuity modeling, the problems of insufficient image detail, distortion of dynamic objects, and spatiotemporal discontinuity in nighttime image processing are solved, achieving high-quality image restoration and enhanced spatiotemporal consistency.
Patent Information
- Application Number
- CN202510186681.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-02-20
AI Technical Summary
Existing technologies cannot effectively improve image details in nighttime image processing, especially the details of dynamic objects. The color and shape of objects are severely distorted, and multimodal data cannot be fully integrated, which limits the accuracy and quality of the image restoration process. Furthermore, it cannot guarantee the spatiotemporal consistency of the image restoration process, resulting in unsmooth image conversion or jumps.
By deeply fusing multimodal data, inferring the shape and color of dynamic objects, and modeling super-resolution and spatiotemporal continuity, a static daytime scene model is generated using visible light, infrared, and depth images. Image restoration and detail recovery are performed by combining a GAN model. A super-resolution and spatiotemporal continuity enhancement method is designed to ensure image quality and spatiotemporal consistency.
It effectively improves the quality of nighttime images, accurately restores the appearance of dynamic objects, ensures the spatiotemporal consistency of image sequences, solves the problems of precision, realism and continuity in image processing in existing technologies, and improves the resolution and detail restoration of images.
Smart Images

Figure CN120047334B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image enhancement technology, and in particular relates to a method for enhancing images of road monitoring at night. Background Technology
[0002] With the rapid development of autonomous driving technology, intelligent traffic monitoring systems, and urban infrastructure, nighttime road monitoring and image processing have become crucial components of modern traffic management. Especially in the perception systems of autonomous vehicles, the identification and understanding of the nighttime road environment requires extremely high accuracy to ensure safe navigation and real-time response to traffic changes. However, traditional road monitoring systems face numerous challenges due to limitations in nighttime lighting conditions. In darkness, image quality deteriorates significantly, particularly visible light images, which often become blurry, have low contrast, and lose detail. Dynamic objects (such as moving vehicles and pedestrians) and static backgrounds (such as road markings and traffic signs) are difficult to clearly identify, severely impacting the realization of functions such as target detection, path planning, and traffic prediction.
[0003] Current technologies primarily rely on low-light image enhancement, image denoising, and object detection algorithms to improve nighttime image quality and object recognition capabilities. Low-light enhancement algorithms attempt to improve visibility in dark scenes by increasing image brightness and contrast, but these methods often struggle to preserve image details and frequently exhibit distortion in handling dynamic objects. Image denoising techniques are effective in removing noise and improving image clarity, but they often ignore the dynamic changes of different objects, failing to accurately preserve the shape, motion trajectory, or color of target objects during the restoration process, resulting in severe distortion and unnatural appearance in the restored image. Object detection algorithms, such as detection models based on convolutional neural networks (CNNs), also have limited performance in dark environments because they cannot accurately extract effective features from low-quality images and lack modeling for spatiotemporal information, failing to guarantee image consistency between consecutive frames.
[0004] Furthermore, while traditional night-to-day image conversion methods (such as altering image brightness and hue through image enhancement or color correction) can improve image quality to some extent, they do not fully utilize the relationship between the background and dynamic objects in a scene, especially failing to consider the matching and restoration of dynamic objects at night with daytime scenes. Specifically, the color, shape, and trajectory of targets such as vehicles and pedestrians moving at night may deviate significantly due to changes in lighting. Existing technologies typically rely on a single image enhancement strategy, which cannot accurately restore the true appearance of these dynamic objects. In addition, the lack of comprehensive utilization of multimodal data (such as infrared images and depth images) also limits existing technologies in target detection, dynamic object inference, and image restoration.
[0005] Therefore, existing nighttime image processing techniques generally suffer from the following problems:
[0006] It cannot effectively improve image details, especially the details of moving objects in the dark;
[0007] During the restoration of nighttime images, the color and shape of objects are severely distorted;
[0008] Image enhancement techniques cannot fully integrate multimodal data, which limits the accuracy and quality of the image restoration process;
[0009] The inability to guarantee spatiotemporal consistency during image restoration results in unsmooth image transitions or jumps in the video sequence. Summary of the Invention
[0010] The purpose of this invention is to propose a method for enhancing images from nighttime road monitoring. Through innovative methods such as deep fusion of multimodal data, shape and color inference of dynamic objects, and super-resolution and spatiotemporal continuity modeling, it overcomes the limitations of existing technologies in nighttime image restoration. This method not only improves image quality but also accurately restores the appearance of dynamic objects and ensures the spatiotemporal consistency of the image sequence. These innovations effectively solve the problems of accuracy, realism, and continuity in image processing under nighttime conditions, demonstrating significant technical advantages.
[0011] To achieve the above objectives, the present invention provides a method for enhancing road monitoring images at night, the method comprising:
[0012] Road monitoring images from different sensors during the daytime are collected and preprocessed. Multimodal feature fusion is then performed on the preprocessed road monitoring images to generate fused features. The fused features are obtained by weighted fusion of the multimodal features. The road monitoring images include visible light images, infrared images, and depth images.
[0013] In multiple road monitoring images acquired by different sensors, a daytime static scene model is constructed based on multiple preprocessed road monitoring images. Based on the daytime static scene model, dynamic objects are inferred using infrared and depth images in the dark. The results of dynamic object inference are matched with the output of the daytime static scene model to recover dynamic objects in the dark scene and generate a dynamic object inference image based on the dark scene.
[0014] Based on the dynamic object inference image of the night scene, a GAN model is designed for training and optimization. The model is used to restore the image and recover the details of the dynamic object inference image of the night scene and output a low-resolution image. The low-resolution image can approximate the real daytime scene in terms of details and object shape.
[0015] Based on low-resolution images of nighttime scenes, super-resolution enhancement and spatiotemporal continuity enhancement are performed on the low-resolution images of nighttime scenes, generating super-resolution loss and spatiotemporal consistency loss respectively. Based on the determined super-resolution loss and spatiotemporal consistency loss, joint optimization is performed to ensure the compatibility of the two objectives, generating images based on super-resolution and spatiotemporal continuity enhancement of nighttime scenes.
[0016] In generating super-resolution and spatiotemporal continuity enhanced images based on nighttime scenes, high-frequency texture and edge information are restored from the input image to enhance local details. Simultaneously, texture restoration and subtle edge enhancement are performed in low-contrast areas to generate a detail-enhanced image. Finally, the detail-enhanced image is globally optimized to generate the final enhanced nighttime road monitoring image.
[0017] Furthermore, the road monitoring images undergo preprocessing, including:
[0018] First, perform sensor data alignment between different sensors;
[0019] Secondly, the sensor data for each mode is standardized and denoised.
[0020] Finally, feature extraction is performed using the standardized and denoised data.
[0021] Furthermore, feature extraction is performed on different road monitoring images, including:
[0022] Edge detection and color histogram extraction methods are used to extract features from visible light images, extracting edge information and color distribution:
[0023] E vis =Canny(I vis )
[0024] H vis =Histogram(I vis )
[0025] Among them, E vis H represents the edge information of a visible light image. vis For color histograms;
[0026] Features are extracted from infrared images using temperature gradient and thermal zone identification methods:
[0027]
[0028] in, This represents the temperature gradient in an infrared image, reflecting the location and intensity of temperature changes. This method helps identify heat source regions in infrared images;
[0029] Deep edge detection and spatial connectivity analysis methods are used to extract features from depth images.
[0030]
[0031] in, It represents the spatial gradient of a depth image, reflecting areas in the image where depth changes significantly.
[0032] 4. The method for enhancing images for nighttime road monitoring according to claim 1, characterized in that, based on visible light images and infrared images, geometric information of static objects is extracted, including:
[0033] Edge detection is used to extract object boundaries in visible light images, while temperature gradients in infrared images are used to distinguish heat source areas.
[0034] The spatial distance information of an object is extracted from a depth image to obtain the object's three-dimensional shape;
[0035] A daytime static scene model based on the geometric information of static objects, including the position information, shape information and relative positional relationships between each object.
[0036] Based on a static daytime scene model, the position, shape, and trajectory of dynamic objects are inferred using infrared and depth images from nighttime. The resulting dynamic object inferences include:
[0037] For inferring the shape of dynamic objects:
[0038] Assuming that dynamic objects appear as regions with large temperature gradients in infrared images, and that this gradient change is closely related to the shape of the object, the outline of the dynamic object can be identified by analyzing the temperature changes in the infrared image. In this process, the geometric information of the static field model during the day is combined to confirm the position and shape of the object.
[0039] Color estimation for dynamic objects:
[0040] The color of a moving object during the day is inferred by extracting heat source distribution from visible light and infrared images. Assuming that the color change of the moving object is finite between day and night, a color mapping function is used to infer the color of the moving object.
[0041] C dvnamic =f(I vis , I IR S scene )
[0042] Among them, C dynamicThe color of a dynamic object is represented by f, which is a color mapping function extracted from daytime visible light and infrared images. The color of the dynamic object is calculated using the daytime static field model as a reference.
[0043] Motion prediction for dynamic objects:
[0044] By inferring the shape of dynamic objects from depth images in a night scene, the trajectory of the objects can be calculated. Depth information provides information about the changes in the position of the objects in space. Assuming that the dynamic objects in the night scene maintain a certain speed and trajectory, the displacement of the objects can be calculated from continuous depth images, and their future direction and speed can be inferred.
[0045] Furthermore, the result inferred based on dynamic objects is matched with the output of the daytime static scene model, including:
[0046] First, spatial alignment and color matching are performed using a daytime static scene model and dynamic object inference:
[0047] In nighttime scenes, the color difference between dynamic objects and those in daytime scenes needs to be corrected through image transformation. This invention designs a mapping function h(·) to map the colors of dynamic objects to the color range of daytime scenes:
[0048] ΔC match =h(C dynamic C day )
[0049] Where, ΔC match It represents the difference between the colors of dynamic objects and daytime scenes, and h(·) is the color mapping function;
[0050] By fusing the inferred dynamic objects with the original night scene image, a reconstructed image of the night scene is obtained:
[0051] By matching the colors and shapes of dynamic objects to daytime scenes, a similar effect to daytime can be reconstructed in nighttime images:
[0052]
[0053] Among them, I night_to_day For the restored image, Indicates image fusion operation, O dynamic It contains information about inferred dynamic objects.
[0054] Furthermore, the GAN model is constructed as follows:
[0055] The generator G not only receives the input night scene image I night_to_day It also receives prior information about daytime scenes. The process of generating a daytime scene template, which is estimated from a static daytime scene model, is as follows:
[0056]
[0057] Among them, I fake Generate an image for generator G, I night_to_day The restored image is used as the input image. θ represents prior information for daytime scenes generated based on historical data or environmental models. G These are the parameters of the generator;
[0058] The discriminator D aims to determine whether an image is a true daytime image, and is constructed as follows:
[0059] The first part is a standard binary classification network used to determine the authenticity of an image;
[0060] The second part is a local evaluation network based on dynamic objects, which is used to evaluate the restoration quality of dynamic objects in the image. The local evaluation network performs detailed analysis of moving objects in the image by learning the spatial distribution and motion trajectory of objects in the image.
[0061] The generator and discriminator of the GAN model are optimized through an adversarial approach, specifically including:
[0062] The optimization objectives of the generator include standard adversarial loss and dynamic object reconstruction loss, specifically including:
[0063] Standard adversarial loss: Optimizes the generator by minimizing the discriminative difference between the generated image and the real image.
[0064] Dynamic object reconstruction loss: Additional regularization is applied to the dynamic object regions in the generated image, using the obtained dynamic object regions to constrain the shape and motion trajectory of these objects in the generated image.
[0065] The optimization objectives of the discriminator include standard adversarial loss and dynamic object evaluation loss, specifically including:
[0066] Standard adversarial loss: used to optimize the discriminator's ability to distinguish between generated and real images.
[0067] Dynamic object evaluation loss: Based on the dynamic object annotation, the discriminator evaluates the restoration quality of the dynamic object and adds a dynamic object error term to the loss.
[0068] Furthermore, after adversarial training and optimization of the GAN model, peak signal-to-noise ratio and structural similarity index were used as the main quality assessment metrics. Simultaneously, for the reconstruction quality of dynamic objects, the Dynamic Object Quality Assessment (DQM) metric was introduced.
[0069]
[0070] in, It is the restored object in the i-th image. It is the actual label of the dynamic object inferred in the previous step.
[0071] Furthermore, the super-resolution enhancement uses a convolutional neural network architecture that incorporates a spatial attention mechanism to improve image detail recovery and resolution, specifically including:
[0072] Convolutional networks are used to extract features from low-resolution images, and spatial attention mechanisms are introduced to enhance attention to dynamic object regions, and then high-resolution images are generated.
[0073] The spatiotemporal continuity enhancement ensures the consistency of motion of dynamic objects between consecutive frames, avoiding unnatural motion or breaks, specifically including:
[0074] Design a loss function to measure image differences between consecutive frames, while constraining the spatiotemporal consistency of dynamic object regions. The loss function L temporal Measures that include both spatial and temporal differences:
[0075]
[0076] Among them, I HR and These are the high-resolution images of the current frame and the previous frame, respectively; ΔI HR and This represents the change in the region of dynamic objects between the current frame and the previous frame; α and β are weighting coefficients that control the influence of the loss term.
[0077] Joint optimizations are performed on image super-resolution restoration and spatiotemporal consistency enhancement, specifically including:
[0078] L final =L SR +λ·L temporal
[0079] Among them, L SR L represents super-resolution loss, measuring the difference between the high-resolution image and the target image; temporal λ represents the spatiotemporal consistency loss, which measures the difference between consecutive frames; λ is a hyperparameter that adjusts the trade-off between super-resolution and spatiotemporal consistency.
[0080] PSNR and SSIM are used to measure image sharpness and structural similarity to quantitatively evaluate the quality of the generated images. At the same time, the temporal-spatial consistency evaluation index TQI is introduced to measure the motion coherence between consecutive frames.
[0081]
[0082] in, Let N represent the high-resolution image of the i-th frame, where N is the total number of frames.
[0083] Furthermore, by introducing an adaptive enhancement function f enhance For input image I HR Enhance the details of the high-frequency components:
[0084] I detail =f enhance (I HR ,θ enhance )
[0085] Where, θ enhance To enhance the parameters of the function, I detail It is an image with enhanced details;
[0086] Redesign the adaptive factor A adapt This factor dynamically adjusts the intensity of detail enhancement based on the characteristics of local regions in the image. The specific calculation method is as follows:
[0087]
[0088] in, Image I HR At the gradient at position (x,y), ||·||2 is the L2 norm.
[0089] Furthermore, a deconvolutional layer based on a convolutional neural network is designed to recover image details. At the same time, an edge enhancement module is designed using the gradient information and local texture features of the image to enhance the edge details of the image. The edge enhancement module significantly improves the edge sharpness in the image by enhancing high-frequency information in the low-frequency part.
[0090] After restoring image details and enhancing edge details, global optimization is performed to obtain the final enhanced nighttime road monitoring image.
[0091] The beneficial technical effects of the present invention are at least as follows:
[0092] (1) This invention utilizes multimodal data (including visible light, infrared, and depth images) and combines high-resolution images of daytime scenes as a reference to train a generative adversarial network (GAN) to generate clear-sky mode images of nighttime scenes. By comprehensively utilizing multimodal inputs, the generator can recover the lost illumination, color, and detail information in nighttime scenes, especially the shape and color of dynamic objects (such as vehicles and pedestrians). Infrared images are used to infer the shape of dynamic objects, while the color and illumination information of daytime scenes are used to accurately restore the color and details of objects, overcoming the shortcomings of traditional image enhancement methods that cannot retain accurate features of dynamic objects.
[0093] (2) Traditional methods often fail to accurately reconstruct the color and shape of dynamic objects in nighttime scenes. This invention, however, infers the thermal distribution shape of objects using infrared images and combines this with illumination and color information from daytime images to infer the color of dynamic objects, achieving accurate reconstruction of dynamic objects in the dark. This technology solves the problem of existing methods failing to accurately restore the appearance of dynamic objects, thereby improving the realism and accuracy of image reconstruction.
[0094] (3) To address the issues of insufficient image detail and spatiotemporal discontinuities in video sequences, this invention combines super-resolution technology and spatiotemporal continuity modeling. Super-resolution technology effectively improves the resolution of nighttime images, making the details in the restored image clearer, especially those of roads, vehicles, and pedestrians. Spatiotemporal continuity modeling ensures smooth transitions between multiple time points by modeling the spatiotemporal relationships in the video frame sequence, avoiding unnatural flickering or jumps during the restoration process. This technology overcomes the problem that existing image enhancement techniques cannot guarantee temporal consistency.
[0095] (4) This invention overcomes the limitations of existing technologies in nighttime image restoration through innovative methods such as deep fusion of multimodal data, shape and color inference of dynamic objects, and super-resolution and spatiotemporal continuity modeling. It not only improves image quality but also accurately restores the appearance of dynamic objects and ensures the spatiotemporal consistency of image sequences. These innovations effectively solve the problems of accuracy, realism, and continuity in image processing in nighttime scenes, and have significant technical advantages. Attached Figure Description
[0096] The present invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained based on the following drawings without creative effort.
[0097] Figure 1 This is a flowchart of a nighttime road monitoring image enhancement method according to the present invention. Detailed Implementation
[0098] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0099] In one or more embodiments, such as Figure 1 As shown, a method for enhancing road monitoring images at night is disclosed, including:
[0100] S1. Multimodal feature fusion is performed on the monitored image to generate fused features, wherein the fused features are obtained by weighted fusion of the multimodal features; wherein the road monitoring image includes visible light image, infrared image and depth image.
[0101] Specifically, in this step, the present invention first needs to collect multimodal data, including visible light images (I... vis ), infrared image (I IR ) and depth images (I depth These data come from different sensors, and it is necessary to ensure that they can be fused into a unified coordinate system.
[0102] Furthermore, firstly, sensor data alignment is performed. Assume this invention has a sensor array, where images captured by each sensor may exist at different resolutions and timestamps. Spatiotemporal consistency of the image data is ensured through geometric correction and time synchronization. Specifically, a time window is set to guarantee that images captured by different sensors are aligned at the same time:
[0103] I aligned (t)={I vis (t), I IR (t), I depth (t)},t∈T
[0104] Here, T is a set of image timestamps, ensuring the spatiotemporal consistency of sensor data.
[0105] Furthermore, for the data of each modality, the present invention requires standardization and denoising processing to eliminate noise interference and inconsistencies between different sensors.
[0106] Standardization: For data of each modality, this invention first performs normalization processing, mapping the value of each pixel to the range [0, 1]. Image I is set to be normalized as follows:
[0107]
[0108] Here, min(I) and max(I) are the minimum and maximum values of pixels in modality I, respectively, ensuring that the data of each modality are within the same standard range.
[0109] Noise reduction processing: To eliminate the effects of low temperature environment or sensor noise in infrared images, this invention uses Gaussian filtering:
[0110] I IR.denoise =G σ *I IR
[0111] Among them, G σ It is a Gaussian kernel, and * indicates a convolution operation. This operation helps remove noise from images and enhances image stability.
[0112] Furthermore, using the standardized and denoised data, this invention performs feature extraction and further fuses multimodal features. Different physical methods are designed to extract features for different modalities of data.
[0113] Visible light image feature extraction: The features of visible light images mainly consist of color, texture, and edge information. Therefore, this invention employs edge detection (e.g., the Canny algorithm) and color histogram extraction methods to extract the edge information and color distribution of the image.
[0114] E vis =Canny(I vis )
[0115] H vis =Histogram(I vis )
[0116] Among them, E vis H represents the edge information of a visible light image. vis This is a color histogram. The features extracted using these two methods provide structural information about the image for subsequent steps.
[0117] Infrared Image Feature Extraction: Infrared images primarily provide information on heat distribution. This invention uses temperature gradient (or change detection) and thermal zone identification methods to extract features from infrared images. Assuming that the pixel values of an infrared image represent temperature information, the temperature gradient can be calculated as follows:
[0118]
[0119] in, This represents the temperature gradient in an infrared image, reflecting the location and intensity of temperature changes. This method helps identify heat source regions in infrared images.
[0120] Depth Image Feature Extraction: Depth images primarily provide spatial structure information. This invention uses depth edge detection (e.g., gradient-based edge detection) and spatial connectivity analysis methods to extract features from depth images. The edges of a depth image can be calculated using the following formula:
[0121]
[0122] in, This represents the spatial gradient of a depth image, reflecting regions with significant depth variations. This method helps to capture object edges in an image.
[0123] Furthermore, features extracted from different modalities (visible light, infrared, and depth) are fused. Based on the contribution of each modality to image restoration, features from each modality are combined using a weighted average method.
[0124] Feature-weighted fusion: To integrate the advantageous features of each modality, a weighted method is used to fuse them. Weighting coefficients are set as α, β, and γ, and the features of each modality are summed using the weights.
[0125]
[0126] Among them, F fused These are the fused features, where α, β, and γ are learned weights, representing the feature contributions of the visible light, infrared, and depth modes, respectively. This approach leverages the data advantages of each modality to form a comprehensive multimodal feature representation.
[0127] Furthermore, the dimensionality of the fused features is reduced to decrease computational complexity while retaining key information.
[0128] Dimensionality reduction: Principal component analysis (PCA) is used to reduce the dimensionality of the fused features, resulting in the dimensionality-reduced feature representation F. final Let the dimension of the feature after dimensionality reduction be d:
[0129] F final =PCA(F fused ,d)
[0130] Here, PCA is the principal component analysis operation, and d is the dimension of the reduced feature space. The reduced features will serve as input for subsequent steps in the generative model.
[0131] S2. In multiple road monitoring images acquired by different sensors, a daytime static scene model is constructed based on multiple preprocessed road monitoring images. Based on the daytime static scene model, dynamic objects are inferred using infrared and depth images in the dark. The results of dynamic object inference are matched with the output of the daytime static scene model to recover dynamic objects in the dark scene and generate a dynamic object inference image based on the dark scene.
[0132] Specifically, in this step, the present invention uses the modal data (visible light image I) processed in the previous stage. vis Infrared image I IR and depth image I depth This is used to construct a static model of the daytime scene. The goal is to obtain a reference model for comparison with the nighttime scene.
[0133] Furthermore, based on visible light image I vis and infrared image I IR This invention extracts the geometric information of static objects. It extracts I through edge detection. vis The object boundary in the middle, while using I IR Temperature gradient in To distinguish heat source regions. This invention assumes that fixed objects (such as buildings or roads) have relatively stable shapes and positions in different modes, and therefore uses these modes to generate an object contour and depth map to model the object. Through depth image I depth By extracting the spatial distance information of the object, the three-dimensional shape of the object can be obtained.
[0134] The final output static scene model S scene It consists of objects in the scene that do not change over time and their spatial information. The location information for each object is... i Shape information i The relative positions of objects will be accurately modeled.
[0135] Furthermore, in this step, the goal is to combine the previously generated daytime static scene model S scene Infrared and depth images from nighttime are used to infer the position, shape, and trajectory of moving objects.
[0136] Shape Inference for Dynamic Objects: Dynamic objects in nighttime scenes typically exhibit different temperature distributions in infrared images compared to static objects. This invention hypothesizes that dynamic objects appear as regions with large temperature gradients in infrared images, and that these gradient changes are closely related to the object's shape. This is achieved by analyzing temperature variations in infrared images. This invention is capable of recognizing the contours of dynamic objects. In this process, the invention combines the previous stage's static scene model S... sceneGeometric information is used to determine the position and shape of an object.
[0137] Color estimation of dynamic objects: Due to the lack of visible light information in the dark, the color of dynamic objects needs to be estimated based on information from the daytime scene. This is achieved using daytime visible light images I... vis and infrared image I IR Based on the extracted heat source distribution, this invention can infer the color of a dynamic object during the day. This invention assumes that the color change of a dynamic object between day and night is finite; therefore, it can use a color mapping function to infer the color of the dynamic object.
[0138] C dynamic =f(I vis ,I IR ,S scene )
[0139] Among them, C dynamic The color of a dynamic object is represented by f, a color mapping function extracted from daytime visible light and infrared images, using a daytime scene model S. scene Used as a reference to calculate the color of dynamic objects.
[0140] Motion estimation of dynamic objects: using depth images in a night scene depth The shape S of the dynamic object predicted in the previous step dynamic This invention can calculate the trajectory of an object. Depth information provides information about the object's positional changes in space. Assuming a dynamic object in a nighttime scene maintains a certain speed and trajectory, this invention calculates the object's displacement using continuous depth images and infers its future direction and speed of motion.
[0141] Furthermore, daytime and nighttime scenes are matched:
[0142] In this step, the present invention will output the results of the first two steps (daytime scene model S). scene And dynamic object speculation O dynamic The system performs matching, using color, shape, and motion information to blend dynamic objects with daytime scenes in order to restore dynamic objects in nighttime scenes.
[0143] Scene matching method: First, this invention uses a daytime scene model S scene And dynamic object speculation O dynamic Spatial alignment and color matching are performed. In nighttime scenes, the color difference between dynamic objects and those in daytime scenes needs to be corrected through image transformation. This invention designs a mapping function h(·) to map the colors of dynamic objects to the color range of daytime scenes:
[0144] ΔC match =h(C dynamicC day )
[0145] Where, ΔC match It represents the difference between the colors of dynamic objects and daytime scenes, and h(·) is the color mapping function.
[0146] Scene restoration output: Finally, this invention compares the inferred dynamic objects with the original night scene image I. night The images are fused to obtain a restored image of the night scene. night_to_day By matching the colors and shapes of dynamic objects to daytime scenes, this invention is able to reconstruct daytime-like effects in nighttime images.
[0147]
[0148] Among them, I night_to_day For the restored image, Indicates image fusion operation, O dynamic It contains information about inferred dynamic objects.
[0149] Through the above steps, this invention can effectively transform nighttime scenes into images with daytime-like effects, thereby improving the accuracy of road monitoring and object detection, especially in low-light environments. Each step, by combining physical models and image processing techniques, ensures that the recovery of the shape, color, and motion of dynamic objects in nighttime scenes is consistent with that of daytime scenes, thus achieving a complete nighttime to daytime image conversion.
[0150] S3. Based on the dynamic object inference image of the night scene, design a GAN model for training and optimization, perform image restoration and detail recovery on the dynamic object inference image of the night scene, and output a low-resolution image. The low-resolution image can approximate the real daytime scene in terms of detail and object shape.
[0151] Specifically, the core objective of this step is to use a Generative Adversarial Network (GAN) to infer the dynamic object image I generated in the previous step. night_to_day Image restoration and detail recovery are performed. By specifically designing a GAN model, the inference and detail reconstruction of dynamic objects in complex nighttime scenes are enhanced. This step is one of the core technologies of this patent and solves the problem of dynamic object recovery in low-light environments.
[0152] Furthermore, in order to better meet the needs of dynamic object reconstruction in low-light environments, this invention designs an innovative generative adversarial network architecture that combines the basic framework of traditional GANs and improves its performance in dynamic object inference and daytime scene reconstruction through some patented innovative elements.
[0153] Generator network design:
[0154] The generator G is responsible for processing the input image I night_to_day Through a series of convolutional and deconvolutional layers, image details are gradually recovered to generate images that closely resemble real daytime scenes. To enhance the generator's generalization ability, a special conditional generator G is designed, which not only receives the input nighttime scene image I... night_to_day It also receives prior information about daytime scenes. This refers to the daytime scene template predicted by the model. Specifically, the generator's generation process is as follows:
[0155]
[0156] Among them, I night_to_day For the input image, θ represents prior information for daytime scenes generated based on historical data or environmental models. G These are the parameters of the generator.
[0157] Discriminator network design:
[0158] The discriminator D determines whether an image is a genuine daytime image. Unlike traditional GAN designs, this patent employs a multi-level discriminator network that not only judges the realism of an image but also scores the reconstruction quality of dynamic objects. The discriminator network consists of two parts:
[0159] The first part is a standard binary classification network used to determine the authenticity of an image.
[0160] The second part is a local evaluation network based on dynamic objects, used to evaluate the quality of the restored dynamic objects in an image. This local evaluation module performs a detailed analysis of moving objects in the image by learning the spatial distribution and motion trajectory of objects in the image.
[0161] The goal of the discriminator is to determine whether an image is a true daytime image.
[0162]
[0163] Furthermore, during training, the generator and discriminator are optimized through an adversarial process. This invention designs a comprehensive optimization objective to enhance the accuracy of image restoration, particularly in the recovery of details from dynamic objects. This optimization objective not only considers the differences between the generated and real images but also specifically introduces a regularization term related to dynamic object inference.
[0164] Furthermore, the optimization objective of the generator is:
[0165] The generator's loss function L G It consists of two parts:
[0166] Standard adversarial loss: Optimizes the generator by minimizing the discriminative difference between the generated image and the real image.
[0167] Dynamic object reconstruction loss: Additional regularization is applied to the dynamic object regions in the generated image, utilizing the dynamic object information obtained from the previous step. This constrains the shape and trajectory of objects in the generated image. The specific loss function is:
[0168]
[0169] in, λ represents the region of dynamic objects inferred from the previous step. dynamic It is a weighting coefficient used to adjust the degree of influence of dynamic object reconstruction loss.
[0170] Furthermore, the optimization objective of the discriminator is:
[0171] The loss function L of the discriminator D It consists of two parts:
[0172] Standard adversarial loss: used to optimize the discriminator's ability to distinguish between generated and real images.
[0173] Dynamic object evaluation loss: Based on the dynamic object annotations obtained in the previous step, the discriminator evaluates the reconstruction quality of the dynamic objects and adds a specific dynamic object error term to the loss:
[0174]
[0175] in, λ represents the actual evaluation information of a dynamic object. dynamic_eval It is an adjustment term used to enhance the impact of dynamic object evaluation loss on generator optimization.
[0176] Furthermore, after adversarial training and optimization, the generator outputs image I reconstructed It can closely approximate realistic daytime scenes in terms of detail and object shape. This restoration process not only reproduces changes in ambient lighting but also accurately restores the details of dynamic objects, solving the problem of restoring dynamic objects in nighttime scenes that is difficult to handle with traditional methods.
[0177] Furthermore, in evaluating the restoration effect, this invention uses Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM) as the main quality assessment indicators. Simultaneously, for the restoration quality of dynamic objects, this invention introduces the Dynamic Object Quality Assessment (DQM) indicator:
[0178]
[0179] in, It is the restored object in the i-th image. It is the actual label of the dynamic object inferred in the previous step.
[0180] Through the above steps, this invention not only effectively utilizes GANs to restore the details of dynamic objects, but also improves the accuracy of image restoration by introducing a regularization loss for dynamic object reconstruction, especially in low-light environments. Furthermore, the innovation of this solution lies in its optimization of dynamic object restoration capabilities through the design of a multi-level discriminator and a dynamic object-specific evaluation module, further enhancing the quality and usability of the generated images. This solution provides strong technical support for the practical application of this invention in dynamic object inference and daytime scene restoration, and is particularly suitable for applications in low-light environments such as traffic monitoring and security monitoring.
[0181] S4. Based on the low-resolution image of the night scene, perform super-resolution enhancement and spatiotemporal continuity enhancement on the low-resolution image of the night scene, generate super-resolution loss and spatiotemporal consistency loss respectively, and perform joint optimization based on the determined super-resolution loss and spatiotemporal consistency loss to ensure the compatibility of the two objectives, and generate an image based on the super-resolution and spatiotemporal continuity enhancement of the night scene.
[0182] Specifically, from low-resolution image I reconstructed This invention restores more details, especially those of moving objects. It employs a convolutional neural network (CNN) architecture that incorporates a spatial attention mechanism to improve image detail recovery and resolution. Steps:
[0183] Convolutional networks are used to extract features from low-resolution images, and a spatial attention mechanism is introduced to enhance attention to regions of dynamic objects. The role of the spatial attention mechanism is to automatically learn which regions are most important for detail recovery, thereby improving the effect of detail recovery.
[0184] Generate high-resolution image I HR :
[0185] I HR =f SR (I reconstructed ;θ SR )
[0186] Among them, f SR Let θ represent the super-resolution function. SR These are its parameters.
[0187] Furthermore, based on the generated high-resolution images, the motion consistency of dynamic objects between consecutive frames is ensured to avoid unnatural motion or breaks. To this end, a spatiotemporal consistency loss function is introduced to optimize the similarity between consecutive frames. Steps:
[0188] Spatiotemporal consistency loss function:
[0189] To enhance spatiotemporal continuity, this invention designs a loss function that measures image differences between consecutive frames while constraining the spatiotemporal consistency of dynamic object regions. This loss function incorporates measures of both spatial and temporal differences:
[0190]
[0191] Among them, I HR and These are the high-resolution images of the current frame and the previous frame, respectively; ΔI HR and This represents the change in the region of the moving object in the current frame and the previous frame; α and β are weighting coefficients that control the influence of the loss term.
[0192] Super-resolution and spatiotemporal consistency:
[0193] Target:
[0194] Joint optimization is performed on image super-resolution restoration and spatiotemporal consistency enhancement to ensure compatibility between the two objectives, thereby generating higher quality images.
[0195] The ultimate optimization goal is to combine the super-resolution loss with the spatiotemporal consistency loss, and obtain a total loss function through weighted summation:
[0196] L final =L SR +λ·L temporal
[0197] Among them, L SR L represents super-resolution loss, measuring the difference between the high-resolution image and the target image; temporal λ represents the spatiotemporal consistency loss, which measures the difference between consecutive frames; λ is a hyperparameter that adjusts the trade-off between super-resolution and spatiotemporal consistency.
[0198] Furthermore, the generated image sequence should maintain high resolution while ensuring that the motion of dynamic objects is coherent and consistent.
[0199] step:
[0200] The quality of the generated images is evaluated using PSNR and SSIM to measure image sharpness and structural similarity.
[0201] Introducing the temporal-spatial consistency index (TQI) to measure motion coherence between consecutive frames:
[0202]
[0203] This formula measures the difference between consecutive frames. Let N represent the high-resolution image of the i-th frame, where N is the total number of frames.
[0204] This step enhances image detail recovery and spatiotemporal consistency across consecutive frames by combining super-resolution and spatiotemporal consistency enhancement. An innovative spatiotemporal consistency loss function ensures coherent motion of dynamic objects, while a spatial attention mechanism enhances detail recovery of dynamic objects. Through joint optimization, this invention can generate high-quality, high-resolution image sequences, solving the spatiotemporal consistency problem in low-resolution image and dynamic scene restoration.
[0205] S5. In generating super-resolution and spatiotemporal continuity enhanced images based on nighttime scenes, high-frequency texture and edge information are restored from the input image to enhance local details. At the same time, texture restoration and subtle edge enhancement are performed in low-contrast areas to generate images with enhanced details. Finally, global optimization is performed on the enhanced images to generate the final enhanced nighttime road monitoring image.
[0206] Specifically, the restored image may have visual problems, such as poor contrast and blurred details, which require post-processing optimization.
[0207] Furthermore, while maintaining the overall naturalness of the image, high-frequency texture and edge information are restored through detail enhancement methods to enhance local details and make the image more delicate and realistic. Steps:
[0208] Detail enhancement functions:
[0209] By introducing an adaptive enhancement function f enhance For input image I HR The function enhances the details of high-frequency components in the image. It combines edge detection and texture restoration to preserve the naturalness of the original image as much as possible while enhancing its appearance.
[0210] I detail =f enhance (I HR ,θ enhance )
[0211] Where, θ enhance To enhance the parameters of the function, I detail This is the image after detail enhancement. This function adjusts the enhancement intensity based on the local features and contrast of the image, automatically adjusting the degree of detail restoration in each region.
[0212] Adaptive enhancement factor:
[0213] To further optimize the effect of detail enhancement, an adaptive factor A was designed. adaptThis factor dynamically adjusts the intensity of detail enhancement based on the characteristics of local image regions (such as edges, texture density, etc.). The specific calculation method is as follows:
[0214]
[0215] in, Image I HR At the gradient at position (x,y), ||·||² is the L2 norm. This factor enhances the details of image edges and high-contrast regions while suppressing unnecessary enhancement in flat areas.
[0216] Texture restoration and edge enhancement:
[0217] Target:
[0218] This process restores texture and subtle edges in the image, especially in low-contrast areas. By incorporating local gradient information and combining it with global optimization, detail is improved without over-enhancing. Steps:
[0219] Texture restoration function:
[0220] This invention designs a deconvolutional layer based on a convolutional neural network (CNN) to recover image details. This layer focuses on high-frequency texture recovery, especially in small areas of the image, such as skin texture and object surfaces. The recovery process is represented by the following function:
[0221] I texture =f deconv (I detail ,θ deconv )
[0222] Among them, f deconv It is a deconvolution function, θ deconv These are the weights obtained during network training. Deconvolutional layers recover subtle textures from the image through inverse convolution operations, thereby restoring high-frequency details.
[0223] Edge enhancement:
[0224] This invention designs an edge enhancement module f by utilizing the gradient information and local texture features of an image. edge This module enhances edge details in images. It significantly improves edge sharpness by amplifying high-frequency information in the low-frequency range.
[0225] E enhanced =f edge (I texture ,θ edge )
[0226] Where, θ edgeThese are the parameters of the edge enhancement module. This module uses gradient information to enhance the edge features of an image, making the image details clearer.
[0227] Furthermore, global detail optimization:
[0228] Target:
[0229] Perform global optimization on the enhanced image to ensure that the enhanced details are consistent with the overall image style, while avoiding over-sharpening or unnatural effects. Steps:
[0230] Global consistency optimization:
[0231] To ensure that detail enhancement does not affect the global consistency of the image, this invention introduces a globally optimized loss function L. global This function combines detail enhancement, edge restoration, and the naturalness of global structure:
[0232]
[0233] in, It is the L2 norm loss, which measures the difference between the restored image details and the original image. λ1 is the L1 norm of the image gradient, used to preserve edge features; λ2 and λ1 are parameters that adjust the weights of different loss terms. This optimization ensures that details are enhanced while maintaining the naturalness of the image.
[0234] Furthermore, after detail enhancement and global optimization, the final image I obtained by this invention is... final Enhanced details are preserved while avoiding over-enhancement and distortion, ensuring a natural and clear image.
[0235] Furthermore, output and evaluation:
[0236] Objective: The final output image should possess sharp details, enhanced edges, and natural textures, while maintaining overall structural consistency. Steps:
[0237] Evaluate image quality: Use metrics such as Structural Similarity Index (SSIM) and Peak Signal-to-Noise Ratio (PSNR) to evaluate the detail enhancement effect and ensure that the image quality improvement meets expectations.
[0238] Perceptual loss function: The perceptual loss function is used to further optimize the visual naturalness of the image, ensuring that the image with enhanced details can achieve the best effect in human visual perception.
[0239] This step designs an innovative image detail enhancement and texture restoration method that combines adaptive enhancement, deconvolutional networks, and edge enhancement techniques. While ensuring detail restoration, it avoids the visual unnaturalness caused by over-enhancement. By introducing a global optimization loss function, the consistency and naturalness of the image during the detail enhancement process are further ensured, thereby generating a high-quality restored image that meets the high-precision image restoration requirements of the patent.
[0240] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0241] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0242] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
[0243] The various illustrative logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein can be implemented or performed using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor may be a microprocessor, but in alternatives, it may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.
[0244] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor such that the processor can read and write information to / from the storage medium. In an alternative, the storage medium may be integrated into the processor. The processor and storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor and storage medium may reside as discrete components in the user terminal.
[0245] In one or more exemplary embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software as a computer program product, the functionality may be stored or transmitted as one or more instructions or code on or through a computer-readable medium. A computer-readable medium includes both computer storage media and communication media, encompassing any medium that facilitates the transfer of a computer program from one location to another. A storage medium may be any available medium accessible to a computer. By way of example and not limitation, such a computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and is accessible to a computer. Any connection is also legitimately referred to as a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of a medium. As used in this article, disk and disc include compact discs (CDs), laser discs, optical discs, digital multi-purpose discs (DVDs), floppy disks, and Blu-ray discs. Disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of these should also be included within the scope of computer-readable media.
[0246] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but should be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for enhancing images used in nighttime road monitoring, characterized in that, The method includes: Road monitoring images from different sensors during the daytime are collected and preprocessed. Multimodal feature fusion is performed on the preprocessed road monitoring images to generate fused features, which are then input into a GAN model. The fused features are obtained by weighted fusion of the multimodal features. The road monitoring images include visible light images, infrared images, and depth images. In multiple road monitoring images acquired by different sensors, a daytime static scene model is constructed based on multiple preprocessed road monitoring images. Based on the daytime static scene model, dynamic objects are inferred using infrared and depth images in the dark. The results of dynamic object inference are matched with the output of the daytime static scene model to recover dynamic objects in the dark scene and generate a dynamic object inference image based on the dark scene. Based on the dynamic object inference image of the night scene, a GAN model is designed for training and optimization. The model is used to restore the image and recover the details of the dynamic object inference image of the night scene and output a low-resolution image. The low-resolution image can approximate the real daytime scene in terms of details and object shape. Based on low-resolution images of nighttime scenes, super-resolution enhancement and spatiotemporal continuity enhancement are performed on the low-resolution images of nighttime scenes, generating super-resolution loss and spatiotemporal consistency loss respectively. Based on the determined super-resolution loss and spatiotemporal consistency loss, joint optimization is performed to ensure the compatibility of the two objectives, generating images based on super-resolution and spatiotemporal continuity enhancement of nighttime scenes. In generating super-resolution and spatiotemporal continuity enhanced images based on nighttime scenes, high-frequency texture and edge information are restored from the input image to enhance local details. Simultaneously, texture restoration and subtle edge enhancement are performed in low-contrast areas to generate a detail-enhanced image. Finally, the detail-enhanced image is globally optimized to generate the final enhanced nighttime road monitoring image.
2. The method for enhancing images of road monitoring at night according to claim 1, characterized in that, Preprocessing of road monitoring images includes: First, perform sensor data alignment between different sensors; Secondly, the sensor data for each mode is standardized and denoised. Finally, feature extraction is performed using the standardized and denoised data.
3. The method for enhancing images of road monitoring at night according to claim 2, characterized in that, Feature extraction is performed on different road monitoring images, including: Edge detection and color histogram extraction methods are used to extract features from visible light images, extracting edge information and color distribution: ; ;in, Represents the edge information of a visible light image. For color histograms; Features are extracted from infrared images using temperature gradient and thermal zone identification methods: ; in, The temperature gradient in the infrared image reflects the location and intensity of temperature changes; Depth edge detection and spatial connectivity analysis methods are used to extract features from depth images: ; in, It represents the spatial gradient of a depth image, reflecting regions in the image where depth changes significantly.
4. The method for enhancing images of road monitoring at night according to claim 1, characterized in that, Based on visible light and infrared images, geometric information of static objects is extracted, including: Edge detection is used to extract object boundaries in visible light images, while temperature gradients in infrared images are used to distinguish heat source areas. The spatial distance information of an object is extracted from a depth image to obtain the object's three-dimensional shape; A daytime static scene model based on the geometric information of static objects, including the position information, shape information and relative positional relationships between each object; Based on a static daytime scene model, the position, shape, and trajectory of dynamic objects are inferred using infrared and depth images from nighttime. The resulting dynamic object inferences include: For inferring the shape of dynamic objects: Assuming that dynamic objects appear as regions with large temperature gradients in infrared images, and that this gradient change is closely related to the shape of the object, the outline of the dynamic object can be identified by analyzing the temperature changes in the infrared image. In this process, the geometric information of the static field model during the day is combined to confirm the position and shape of the object. Color estimation for dynamic objects: The color of a moving object during the day is inferred by extracting heat source distribution from visible light and infrared images. Assuming that the color change of the moving object is finite between day and night, a color mapping function is used to infer the color of the moving object. ; in, Indicates the color of a moving object. It is a color mapping function extracted from daytime visible light and infrared images, and uses the daytime static field model as a reference to calculate the color of dynamic objects; Motion prediction for dynamic objects: By inferring the shape of dynamic objects from depth images in a night scene, the trajectory of the objects can be calculated. Depth information provides information about the changes in the position of the objects in space. Assuming that the dynamic objects in the night scene maintain a certain speed and trajectory, the displacement of the objects can be calculated from continuous depth images, and their future direction and speed can be inferred.
5. The method for enhancing images of road monitoring at night according to claim 4, characterized in that, The matching of the results inferred from dynamic objects with the output of the daytime static scene model includes: First, spatial alignment and color matching are performed using a daytime static scene model and dynamic object inference: In nighttime scenes, the color difference between dynamic objects and those in daytime scenes needs to be corrected through image transformation; a mapping function needs to be designed. Map the colors of dynamic objects to the color range of a daytime scene: ; in, It is the difference between the colors of moving objects and those of a daytime scene. It is a color mapping function; By fusing the inferred dynamic objects with the original night scene image, a reconstructed image of the night scene is obtained: By matching the colors and shapes of dynamic objects to daytime scenes, a similar effect to daytime can be reconstructed in nighttime images: ; in, For the restored image, This indicates an image fusion operation. It contains information about inferred dynamic objects.
6. The method for enhancing images of road monitoring at night according to claim 5, characterized in that, The GAN model is constructed as follows: generator Not only receiving input night scene images It also receives prior information about daytime scenes. That is, the daytime scene template is estimated through a daytime static scene model, and the generation process is as follows: ; in, For generator Generate an image. The restored image is used as the input image. This refers to prior information about daytime scenes generated based on historical data or environmental models. These are the parameters of the generator; Discriminator The goal is to determine whether an image is a true daytime image, constructed as follows: The first part is a standard binary classification network used to determine the authenticity of an image; The second part is a local evaluation network based on dynamic objects, which is used to evaluate the restoration quality of dynamic objects in the image. The local evaluation network performs detailed analysis of moving objects in the image by learning the spatial distribution and motion trajectory of objects in the image. The generator and discriminator of the GAN model are optimized through an adversarial approach, specifically including: The optimization objectives of the generator include standard adversarial loss and dynamic object reconstruction loss, specifically including: Standard adversarial loss: Optimizes the generator by minimizing the discriminative difference between the generated image and the real image; Dynamic object reconstruction loss: Additional regularization is applied to the dynamic object regions in the generated image, using the obtained dynamic object regions to constrain the shape and motion trajectory of these objects in the generated image. The optimization objectives of the discriminator include standard adversarial loss and dynamic object evaluation loss, specifically including: Standard adversarial loss: used to optimize the discriminator's ability to distinguish between generated and real images; Dynamic object evaluation loss: Based on the dynamic object annotation, the discriminator evaluates the restoration quality of the dynamic object and adds a dynamic object error term to the loss.
7. The method for enhancing images of road monitoring at night according to claim 6, characterized in that, After adversarial training and optimization of the GAN model, peak signal-to-noise ratio (PSNR) and structural similarity index were used as the main quality assessment metrics. Additionally, for the reconstruction quality of dynamic objects, the Dynamic Object Quality Metric (DQM) was introduced. ; in, It is the first The restored object in the image, It is the actual label of the dynamic object inferred in the previous step.
8. The method for enhancing images of road monitoring at night according to claim 6, characterized in that, The super-resolution enhancement uses a convolutional neural network architecture that incorporates a spatial attention mechanism to improve image detail recovery and resolution, specifically including: Convolutional networks are used to extract features from low-resolution images, and spatial attention mechanisms are introduced to enhance attention to dynamic object regions, and then high-resolution images are generated. The spatiotemporal continuity enhancement ensures the consistency of motion of dynamic objects between consecutive frames, avoiding unnatural motion or breaks, specifically including: Design a loss function to measure image differences between consecutive frames, while constraining the spatiotemporal consistency of dynamic object regions. Measures that include both spatial and temporal differences: ; in, and These are high-resolution images of the current frame and the previous frame, respectively. and This represents the changes in the region of dynamic objects between the current frame and the previous frame. and These are weighting coefficients to control the impact of the loss term; Joint optimizations are performed on image super-resolution restoration and spatiotemporal consistency enhancement, specifically including: ; in, This represents super-resolution loss, which measures the difference between the high-resolution image and the target image. It represents the spatiotemporal consistency loss and measures the difference between consecutive frames; Hyperparameters used to adjust the trade-off between super-resolution and spatiotemporal consistency; PSNR and SSIM are used to measure image sharpness and structural similarity, quantitatively evaluating the quality of the generated images, while also introducing a spatiotemporal consistency evaluation index. To measure motion coherence between consecutive frames: ; in, Indicates the first High-resolution images of frames, This represents the total number of frames.
9. The method for enhancing images of road monitoring at night according to claim 8, characterized in that, By introducing an adaptive enhancement function For the input image Enhance the details of the high-frequency components: ; in, To enhance the parameters of the function, It is an image with enhanced details; Redesign Adaptive Factors This factor dynamically adjusts the intensity of detail enhancement based on the characteristics of local regions in the image. The specific calculation method is as follows: ; in, Representing an image In position gradient, It is an L2 norm.
10. The method for enhancing images of road monitoring at night according to claim 9, characterized in that, The design incorporates a deconvolutional layer based on a convolutional neural network to recover image details. Simultaneously, it utilizes the gradient information and local texture features of the image to design an edge enhancement module to enhance the edge details of the image. This edge enhancement module significantly improves the edge sharpness in the image by enhancing high-frequency information in the low-frequency region. After restoring image details and enhancing edge details, global optimization is performed to obtain the final enhanced nighttime road monitoring image.
Citation Information
Patent Citations
Multi-sensor fusion low-illumination video image enhancement method
CN105809640A
Generative adversarial network-based power transmission line image enhancement method under low illuminance
CN115601644A