Night road monitoring image enhancement method

Through multimodal data fusion and dynamic object estimation, combined with super-resolution and spatiotemporal continuity modeling, GAN and CNN architectures are used to solve the problems of insufficient details, color morphology distortion and spatiotemporal discontinuity in night image processing, and high-quality image restoration and spatiotemporal consistency are achieved.

CN120047334AActive Publication Date: 2025-05-27SHENZHEN ZHONGTING TECH CO LTD +1

Patent Information

Application Number
CN202510186681.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-27
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The prior art cannot effectively improve image details in night image processing, especially the details of dynamic objects. In the process of image restoration, the color and morphology of objects are severely distorted, and the multimodal data cannot be fully combined, resulting in limited accuracy and quality of image restoration and the time and space consistency cannot be guaranteed.

Method used

Using deep fusion of multimodal data, estimation of shape and color of dynamic objects, super-resolution and spatiotemporal continuity modeling, the resolution and spatiotemporal consistency of images are improved by generating convolutional neural network (CNN) architectures with adversarial network (GAN) and spatial attention mechanisms.

Benefits of technology

It significantly improves the quality of night images, accurately restores the appearance of dynamic objects, ensures the time and space consistency of image sequences, and solves the accuracy, authenticity and continuity of image processing in dark scenes in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047334A_ABST
    Figure CN120047334A_ABST
Patent Text Reader

Abstract

The invention provides a night road monitoring image enhancement method. The method comprises the steps that road monitoring images from different sensors in a daytime scene are collected and preprocessed; speculating an image based on a dynamic object of a night scene; and designing a GAN model for training and optimization. Generating a super-resolution and space-time continuity enhanced image based on the night scene; according to the method, high-frequency texture and edge information of an input image are recovered, local details of the image are enhanced, texture recovery and fine edge enhancement are carried out in a low-contrast area, an image after detail enhancement is generated, finally, global optimization is carried out on the image after detail enhancement, and a night road monitoring image after final enhancement is generated. According to the invention, many problems in night image processing in the prior art are solved. A visible light image, an infrared image, a depth image and other sensor data are fused, and a dynamic object at night is speculated by using daytime scene modeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image enhancement, and in particular relates to a method for enhancing images for nighttime road monitoring. Background Art

[0002] With the rapid development of autonomous driving technology, intelligent traffic monitoring systems, and urban infrastructure, nighttime road monitoring and image processing have become a vital part of modern traffic management. Especially in the perception system of autonomous vehicles, the recognition and understanding of the nighttime road environment requires extremely high precision to ensure that the vehicle can navigate safely and respond to traffic changes in real time. However, due to the limitations of nighttime lighting conditions, traditional road monitoring systems face many challenges. In a dark environment, image quality is greatly reduced, especially visible light images often become blurred, low in contrast, and lost in details. Dynamic objects (such as moving vehicles, pedestrians, etc.) and static backgrounds (such as road markings, traffic signs) are difficult to clearly identify, which seriously affects the realization of functions such as target detection, path planning, and traffic prediction.

[0003] Existing technologies mainly rely on low-light image enhancement, image denoising and target detection algorithms to improve night image quality and target recognition capabilities. Low-light enhancement algorithms attempt to improve the visibility of night scenes by increasing image brightness and contrast, but these methods often have difficulty retaining image details and often have distortions in the processing of dynamic objects. Image denoising techniques have been effective in removing noise and improving image clarity, but they often ignore the dynamic changes of different objects and cannot accurately retain the shape, motion trajectory or color of the target object during the restoration process, resulting in serious distortion and unnatural phenomena in the restored image. Target detection algorithms, such as detection models based on convolutional neural networks (CNNs), also have limited performance in dark environments because they cannot accurately extract effective features from low-quality images, lack modeling for spatiotemporal information, and cannot ensure the consistency of images between consecutive frames.

[0004] In addition, although traditional night-to-day image conversion methods (such as changing the brightness and hue of the image through image enhancement or color correction) can improve the quality of the image to a certain extent, they do not fully utilize the relationship between the background and dynamic objects in the scene, especially without considering the matching and restoration of nighttime dynamic objects with daytime scenes. Specifically, the color, shape and motion trajectory of targets such as vehicles and pedestrians traveling at night may deviate significantly due to changes in lighting. Existing technologies usually rely on a single image enhancement strategy and cannot accurately restore the true appearance of these dynamic objects. In addition, the lack of comprehensive use of multimodal data (such as infrared images, depth images, etc.) also limits existing technologies in target detection, dynamic object inference and image restoration.

[0005] Therefore, the existing night image processing technology generally has the following problems:

[0006] It cannot effectively improve image details, especially the details of dynamic objects in the dark;

[0007] In the process of restoring night images, the color and shape of objects are seriously distorted;

[0008] Image enhancement technology cannot fully combine multimodal data, resulting in limited accuracy and quality of the image restoration process;

[0009] The temporal and spatial consistency of the image restoration process cannot be guaranteed, resulting in uneven image transitions or jumps in video sequences. Summary of the invention

[0010] The purpose of this invention is to propose a method for enhancing images of road monitoring at night. Through innovative methods such as deep fusion of multimodal data, shape and color inference of dynamic objects, super-resolution and spatiotemporal continuity modeling, the limitations of existing technologies in night image restoration are overcome, which not only improves the quality of images, but also accurately restores the appearance of dynamic objects and ensures the spatiotemporal consistency of image sequences. These innovations effectively solve the problems of accuracy, authenticity and continuity of image processing in night scenes in existing technologies, and have significant technical advantages.

[0011] In order to achieve the above object, the present invention provides a method for enhancing a nighttime road monitoring image, the method comprising:

[0012] collecting and preprocessing road monitoring images from different sensors in daytime scenes, and performing multimodal feature fusion based on the preprocessed road monitoring images to generate fused features, wherein the fused features are obtained by weighted fusion of the multimodal features; wherein the road monitoring images include visible light images, infrared images, and depth images;

[0013] In a plurality of road monitoring images acquired by different sensors, a daytime static scene model is constructed based on a plurality of preprocessed road monitoring images, dynamic objects are inferred based on the daytime static scene model using infrared images and depth images at night, and the result of the dynamic object inference is matched with the output of the daytime static scene model to restore the dynamic objects in the night scene, and generate a dynamic object inference image based on the night scene;

[0014] Based on the inferred images of dynamic objects in night scenes, a GAN model is designed for training and optimization, image restoration and detail recovery are performed on the inferred images of dynamic objects in night scenes, and low-resolution images are output. The low-resolution images can be close to real daytime scenes in terms of details and object morphology;

[0015] Based on a low-resolution image of a night scene, super-resolution enhancement and spatiotemporal continuity enhancement are performed on the low-resolution image based on the night scene, a super-resolution loss and a spatiotemporal consistency loss are generated respectively, and joint optimization is performed based on the determined super-resolution loss and spatiotemporal consistency loss to ensure the compatibility of the two objectives, thereby generating an image based on the night scene with super-resolution and spatiotemporal continuity enhancement;

[0016] In generating super-resolution and spatiotemporal continuity enhanced images based on night scenes, high-frequency texture and edge information of the input image are restored to enhance local details of the image. At the same time, texture restoration and subtle edge enhancement are performed in low-contrast areas to generate detail-enhanced images. Finally, the detail-enhanced images are globally optimized to generate the final enhanced night road monitoring images.

[0017] Furthermore, the road monitoring image is preprocessed, including:

[0018] First, sensor data alignment of different sensors is performed;

[0019] Secondly, the sensor data of each modality is normalized and denoised;

[0020] Finally, feature extraction is performed on the standardized and denoised data.

[0021] Furthermore, feature extraction is performed on different road monitoring images, including:

[0022] Edge detection and color histogram extraction methods are used to extract features of visible light images, and to extract edge information and color distribution of images:

[0023] E vis =Canny(I vis )

[0024] H vis =Histogram(I vis )

[0025] Among them, E vis Represents the edge information of the visible light image, H vis is the color histogram;

[0026] Extract features of infrared images using temperature gradient and hotspot identification methods:

[0027]

[0028] in, Represents the temperature gradient of infrared images, reflecting the location and intensity of temperature changes. This method helps to identify heat source areas in infrared images;

[0029] Deep edge detection and spatial connectivity analysis methods extract features of depth images:

[0030]

[0031] in, Represents the spatial gradient of the depth image, reflecting the areas with large depth changes in the image.

[0032] 4. The method for enhancing images for nighttime road monitoring according to claim 1 is characterized in that the geometric information of static objects is extracted based on visible light images and infrared images, including:

[0033] The object boundaries in the visible light image are extracted by edge detection, and the temperature gradient in the infrared image is used to distinguish the heat source area;

[0034] The spatial distance information of the object is extracted through the depth image to obtain the three-dimensional shape of the object;

[0035] A daytime static scene model output based on the geometric information of static objects, including the position information and shape information of each object and the relative position relationship between objects;

[0036] Based on the daytime static scene model, the position, shape and motion trajectory of dynamic objects are inferred using infrared images and depth images at night to obtain dynamic object inference, including:

[0037] Shape inference for dynamic objects:

[0038] Assuming that dynamic objects appear as areas with large temperature gradients in infrared images, and that this gradient change is closely related to the shape of the object, the contours of dynamic objects can be identified by analyzing the temperature changes in infrared images. In this process, the geometric information of the daytime static field model is combined to confirm the position and shape of the object.

[0039] Color inference for dynamic objects:

[0040] The color of dynamic objects during the day is inferred through the heat source distribution extracted from the daytime visible light image and the infrared image. Assuming that the color change of dynamic objects during the day and night is limited, the color mapping function is used to infer the color of dynamic objects:

[0041] C dvnamic =f(I vis , I IR , S scene )

[0042] Among them, C dynamicrepresents the color of dynamic objects, f is the color mapping function extracted from daytime visible light images and infrared images, and the color of dynamic objects is estimated using the daytime static field model as a reference;

[0043] For the motion estimation of dynamic objects:

[0044] The shape of dynamic objects is inferred through depth images in night scenes, and the motion trajectory of the objects is calculated. The depth information provides the position change of the objects in space. Assuming that the dynamic objects in the night scenes maintain a certain speed and motion trajectory, the displacement of the objects is calculated through continuous depth images, and their future motion direction and speed are inferred.

[0045] Furthermore, the result based on the dynamic object inference is matched with the output of the daytime static scene model, including:

[0046] First, use the daytime static scene model and dynamic object inference for spatial alignment and color matching:

[0047] In a night scene, the color difference between the dynamic object and the day scene needs to be corrected by image transformation. The present invention designs a mapping function h(·) to map the color of the dynamic object to the color range of the day scene:

[0048] ΔC match =h(C dynamic , C day )

[0049] Where, ΔC match is the color difference between the dynamic object and the daytime scene, h(·) is the color mapping function;

[0050] The inferred dynamic object is fused with the original night scene image to obtain the restored image of the night scene:

[0051] By matching the color and shape of dynamic objects to the daytime scene, a daytime-like effect can be recreated in night images:

[0052]

[0053] Among them, I night_to_day is the restored image, represents the image fusion operation, O dynamic Contains information about inferred dynamic objects.

[0054] Furthermore, the GAN model is constructed as follows:

[0055] The generator G not only receives the input night scene image I night_to_day , and also receives prior information about daytime scenes That is, the daytime scene template estimated by the daytime static scene model is generated in the following process:

[0056]

[0057] Among them, I fake Generate images for the generator G, I night_to_day is the restored image, as the input image, is the daytime scene prior information generated based on historical data or environmental models, θ G are the parameters of the generator;

[0058] The discriminator D aims to determine whether the image is a real daytime image, which is constructed as follows:

[0059] The first part is a standard two-classification network, which is used to judge the authenticity of the image;

[0060] The second part is a local evaluation network based on dynamic objects, which is used to evaluate the restoration quality of dynamic objects in the image. The local evaluation network conducts a detailed analysis of the moving objects in the image by learning the spatial distribution and motion trajectory of the objects in the image.

[0061] The generator and discriminator of the GAN model are optimized through adversarial means, specifically including:

[0062] The optimization objectives of the generator include standard adversarial loss and dynamic object reconstruction loss, specifically:

[0063] Standard adversarial loss: optimize the generator by minimizing the discriminative difference between generated images and real images.

[0064] Dynamic object reconstruction loss: additional regularization is performed on the dynamic object regions in the generated image, and the obtained dynamic object regions are used to constrain the shape and motion trajectory of these objects in the generated image;

[0065] The optimization objectives of the discriminator include standard adversarial loss and dynamic object evaluation loss, specifically:

[0066] Standard adversarial loss: used to optimize the discriminator's ability to distinguish between generated images and real images.

[0067] Dynamic object evaluation loss: Based on dynamic object annotations, the discriminator evaluates the restoration quality of dynamic objects and adds the dynamic object error term to the loss.

[0068] Furthermore, after adversarial training and optimization of the GAN model, peak signal-to-noise ratio and structural similarity index were used as the main quality evaluation indicators. At the same time, the dynamic object quality evaluation indicator DQM was introduced for the restoration quality of dynamic objects:

[0069]

[0070] in, is the restored object in the i-th image, is the true label of the dynamic object inferred in the previous step.

[0071] Furthermore, the super-resolution enhancement is to use a convolutional neural network architecture that integrates a spatial attention mechanism to improve image detail recovery and resolution, specifically including:

[0072] Use convolutional networks to extract features from low-resolution images, and introduce a spatial attention mechanism to strengthen attention to dynamic object areas, and then generate high-resolution images;

[0073] The spatiotemporal continuity enhancement ensures the consistency of motion of dynamic objects between consecutive frames to avoid unnatural motion or breaks, specifically including:

[0074] Design a loss function to measure the image difference between consecutive frames and constrain the spatiotemporal consistency of dynamic object areas. The loss function L temporal Contains measures of spatial and temporal differences:

[0075]

[0076] Among them, I HR and are high-resolution images of the current frame and the previous frame respectively; ΔI HR and is the change of the dynamic object area between the current frame and the previous frame; α and β are weight coefficients that control the impact of the loss term;

[0077] Joint optimization is performed in image super-resolution restoration and spatiotemporal consistency enhancement, including:

[0078] L final =L SR +λ·L temporal

[0079] Among them, L SR represents the super-resolution loss, which measures the difference between the high-resolution image and the target image; L temporal represents the spatiotemporal consistency loss, measuring the difference between consecutive frames; λ is a hyperparameter that adjusts the trade-off between super-resolution and spatiotemporal consistency;

[0080] PSNR and SSIM are used to measure the clarity and structural similarity of the image, and to quantitatively evaluate the quality of the generated image. At the same time, the temporal and spatial consistency evaluation index TQI is introduced to measure the motion coherence between consecutive frames:

[0081]

[0082] in, represents the high-resolution image of the i-th frame, and N is the total number of frames.

[0083] Furthermore, by introducing the adaptive enhancement function f enhance , for the input image I HR Enhance the details of the high-frequency components in:

[0084] I detail =f enhance (I HR ,θ enhance )

[0085] Among them, θ enhance is the parameter of the enhancement function, I detail It is the image after detail enhancement;

[0086] Redesign adaptive factor A adapt , this factor dynamically adjusts the intensity of detail enhancement according to the characteristics of the local area of ​​the image. The specific calculation method is as follows:

[0087]

[0088] in, Represents image I HR The gradient at position (x, y), ||·|| 2 is the L2 norm.

[0089] Furthermore, a deconvolution layer based on a convolutional neural network is designed to restore image details. At the same time, an edge enhancement module is designed by utilizing the gradient information and local texture features of the image to enhance the edge details of the image. The edge enhancement module significantly improves the edge sharpness in the image by enhancing the high-frequency information in the low-frequency part.

[0090] After restoring image details and enhancing image edge details, global optimization is performed to obtain the final enhanced night road monitoring image.

[0091] The beneficial technical effects of the present invention are at least as follows:

[0092] (1) The present invention utilizes multimodal data (including visible light, infrared, depth images, etc.) and combines high-resolution images of daytime scenes as references to train a generative adversarial network (GAN) to generate sunny mode images of nighttime scenes. Through the comprehensive use of multimodal inputs, the generator is able to restore the lost illumination, color, and detail information in nighttime scenes, especially the shape and color of dynamic objects (such as vehicles, pedestrians, etc.). Infrared images are used to infer the shape of dynamic objects, while the color and illumination information of daytime scenes are used to accurately restore the color and details of objects, overcoming the shortcomings of traditional image enhancement methods that cannot retain the accurate features of dynamic objects.

[0093] (2) Traditional methods often cannot accurately restore the color and shape of dynamic objects in night scenes. However, the present invention uses infrared images to infer the thermal distribution shape of objects, and combines the illumination and color information of daytime images to infer the color of dynamic objects, thus achieving accurate restoration of dynamic objects in the dark. This technology solves the problem that existing methods cannot accurately restore the appearance of dynamic objects, thereby improving the authenticity and accuracy of image restoration.

[0094] (3) In order to solve the problems of insufficient image details and spatiotemporal discontinuity in video sequences, the present invention combines super-resolution technology and spatiotemporal continuity modeling. Super-resolution technology effectively improves the resolution of night images, making the details in the restored images clearer, especially the details of roads, vehicles and pedestrians. Spatiotemporal continuity modeling ensures smooth transition of images between multiple time points by modeling the spatiotemporal relationship in the video frame sequence, avoiding unnatural flickering or jumping during the restoration process. This technology breaks through the problem that existing image enhancement technology cannot ensure temporal consistency.

[0095] (4) The present invention overcomes the limitations of existing technologies in night image restoration through innovative methods such as deep fusion of multimodal data, shape and color inference of dynamic objects, super-resolution and spatiotemporal continuity modeling, which not only improves the quality of images, but also accurately restores the appearance of dynamic objects and ensures the spatiotemporal consistency of image sequences. These innovations effectively solve the problems of accuracy, authenticity and continuity of image processing in night scenes in existing technologies, and have significant technical advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] The present invention is further described using the accompanying drawings, but the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative work.

[0097] Figure 1 The present invention is a flowchart of a method for enhancing images of nighttime road monitoring. DETAILED DESCRIPTION

[0098] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0099] In one or more embodiments, Figure 1 As shown, a method for enhancing a nighttime road monitoring image is disclosed, comprising:

[0100] S1. Perform multimodal feature fusion on a monitoring image to generate fused features, wherein the fused features are obtained by weighted fusion of the multimodal features; wherein the road monitoring image includes a visible light image, an infrared image, and a depth image.

[0101] Specifically, in this step, the present invention first needs to collect multimodal data, including visible light images (I vis ), infrared image (I IR ) and depth image (I depth ). These data come from different sensors and need to be fused into a unified coordinate system.

[0102] Furthermore, first, sensor data alignment is performed. Assume that the present invention has a sensor array, where the images captured by each sensor may exist at different resolutions and timestamps. Through geometric correction and time synchronization, the temporal and spatial consistency of the image data is ensured. Specifically, a time window is set to ensure that the images captured by different sensors are aligned at the same time:

[0103] I aligned (t) = {I vis (t), I IR (t), I depth (t)}, t∈T

[0104] Where T is the set of image timestamps, ensuring the spatiotemporal consistency of sensor data.

[0105] Furthermore, for the data of each modality, the present invention needs to perform standardization and denoising processing to eliminate noise interference and inconsistency between different sensors.

[0106] Normalization: For each modality of data, the present invention first performs normalization processing to map the value of each pixel point to the range of [0, 1]. Set image I to be normalized:

[0107]

[0108] Among them, min(I) and max(I) are the minimum and maximum values ​​of the pixels in modality I, respectively, ensuring that the data of each modality is within the same standard range.

[0109] De-noising: In order to eliminate the influence of low temperature environment or sensor noise in infrared images, the present invention uses Gaussian filtering:

[0110] I IR.denoise =G σ *I IR

[0111] Among them, G σ is a Gaussian kernel and * denotes a convolution operation. This operation helps remove noise from the image and enhances the stability of the image.

[0112] Furthermore, the present invention extracts features from the standardized and denoised data, and further fuses multi-modal features. Different physical methods are designed to extract features for different modal data.

[0113] Visible light image feature extraction: The features of visible light images are mainly composed of color, texture and edge information. Therefore, the present invention uses edge detection (such as Canny algorithm) and color histogram extraction method to extract the edge information and color distribution of the image:

[0114] E vis =Canny(I vis )

[0115] H vis =Histogram(I vis )

[0116] Among them, E vis Represents the edge information of the visible light image, H vis is the color histogram. The features extracted by these two methods provide the structural information of the image for the subsequent steps.

[0117] Infrared image feature extraction: Infrared images mainly provide heat distribution information. This invention uses temperature gradient (or change detection) and hot zone recognition methods to extract the features of infrared images. Assuming that the pixel values ​​of infrared images represent temperature information, the temperature gradient can be calculated in the following way:

[0118]

[0119] in, Represents the temperature gradient of an infrared image, reflecting the location and intensity of temperature changes. This method helps to identify heat source areas in infrared images.

[0120] Depth image feature extraction: Depth images mainly provide spatial structure information. The present invention uses depth edge detection (such as gradient-based edge detection) and spatial connectivity analysis methods to extract the features of depth images. The edge of the depth image can be calculated by the following formula:

[0121]

[0122] in, Represents the spatial gradient of the depth image, reflecting the areas in the image where the depth changes greatly. This method helps capture the edges of objects in the image.

[0123] Furthermore, the features extracted from different modalities (visible light, infrared and depth) are fused. According to the contribution of different modalities to image restoration, the features of each modality are combined by weighted averaging.

[0124] Weighted feature fusion: In order to combine the advantages of each modality, a weighted method is used to fuse them. The weighting coefficients are set to α, β, γ, and the features of each modality are weighted and summed by weight:

[0125]

[0126] Among them, F fused is the fused feature, α, β, γ are the learned weights, representing the feature contribution of visible light, infrared, and depth modalities, respectively. In this way, the data advantages of each modality can be utilized to form a comprehensive multimodal feature representation.

[0127] Furthermore, the fused features are subjected to dimensionality reduction to reduce computational complexity while retaining key information.

[0128] Dimensionality reduction: Use principal component analysis (PCA) to reduce the dimensionality of the fusion features and obtain the feature representation F after dimensionality reduction. final . Set the dimension of the feature after dimensionality reduction to d:

[0129] F final = PCA(F fused ,d)

[0130] Among them, PCA is the principal component analysis operation, and d is the dimension of the feature space after dimensionality reduction. The reduced features will be used as input in the subsequent steps for the generation model.

[0131] S2. Among multiple road monitoring images acquired by different sensors, a daytime static scene model is constructed based on multiple preprocessed road monitoring images. Dynamic objects are inferred based on the daytime static scene model using infrared images and depth images at night. The result of dynamic object inference is matched with the output of the daytime static scene model to restore dynamic objects in the night scene, and generate a dynamic object inference image based on the night scene.

[0132] Specifically, in this step, the present invention uses the modal data (visible light image I vis 、Infrared image I IR and depth image I depth ) to construct a static model of the daytime scene. The goal is to obtain a reference model for comparison with the nighttime scene.

[0133] Furthermore, based on the visible light image I vis and infrared image I IR , the present invention extracts the geometric information of static objects. Extracting I vis The object boundaries in I IR The temperature gradient in To distinguish the heat source area. The present invention assumes that fixed objects (such as buildings or roads) have relatively stable shapes and positions in different modes, so these modes are used to generate an object contour and depth map to model the object. depth Extract the spatial distance information of the object and obtain the three-dimensional shape of the object.

[0134] The final output static scene model S scene It consists of objects in the scene that do not change over time and their spatial information. i , Shape information i And the relative position relationships between objects will be accurately modeled.

[0135] Furthermore, in this step, the goal is to combine the previously generated daytime static scene model S scene , using infrared images and depth images in the dark to infer the position, shape and motion trajectory of dynamic objects.

[0136] Shape estimation of dynamic objects: Dynamic objects in night scenes usually have different temperature distributions from static objects in infrared images. This paper assumes that dynamic objects appear as areas with large temperature gradients in infrared images, and that this gradient change is closely related to the shape of the object. The present invention can identify the contours of dynamic objects. In this process, the present invention combines the static scene model S sceneThe geometric information of the object is used to determine its position and shape.

[0137] Color estimation of dynamic objects: Due to the lack of visible light information at night, the color of dynamic objects needs to be estimated based on the information of the daytime scene. vis and infrared image I IR The extracted heat source distribution can be used to infer the color of dynamic objects during the day. The present invention assumes that the color changes of dynamic objects during the day and night are limited, so the present invention can use a color mapping function to infer the color of dynamic objects:

[0138] C dynamic =f(I vis ,I IR ,S scene )

[0139] Among them, C dynamic represents the color of dynamic objects, f is the color mapping function extracted from the daytime visible light image and infrared image, and the daytime scene model S scene Used as a reference to estimate the color of dynamic objects.

[0140] Motion estimation of dynamic objects: Depth images in dark scenes I depth and the shape S of the dynamic object inferred in the previous step dynamic , the present invention can calculate the motion trajectory of an object. Depth information provides the position change of an object in space. Assuming that a dynamic object in a night scene maintains a certain speed and motion trajectory, the present invention calculates the displacement of the object through continuous depth images and infers its future motion direction and speed.

[0141] Furthermore, day and night scenes are matched:

[0142] In this step, the present invention converts the output of the first two steps (daytime scene model S scene and dynamic object inference O dynamic ) are matched and dynamic objects are fused with daytime scenes through color, shape and motion information to restore dynamic objects in night scenes.

[0143] Scene matching method: First, the present invention uses the daytime scene model S scene and dynamic object inference O dynamic Perform spatial alignment and color matching. In a night scene, the color difference between the dynamic object and the day scene needs to be corrected by image transformation. The present invention designs a mapping function h(·) to map the color of the dynamic object to the color range of the day scene:

[0144] ΔC match =h(C dynamic,C day )

[0145] Where, ΔC match is the color difference between the dynamic object and the daytime scene, and h(·) is the color mapping function.

[0146] Scene restoration output: Finally, the present invention combines the inferred dynamic objects with the original night scene image I night Fusion is performed to obtain the restored image I of the night scene night_to_day By matching the color and shape of dynamic objects with daytime scenes, the present invention can reconstruct a daytime-like effect in night images:

[0147]

[0148] Among them, I night_to_day is the restored image, represents the image fusion operation, O dynamic Contains information about inferred dynamic objects.

[0149] Through the above steps, the present invention can effectively transform a night scene into an image with a daytime effect, thereby improving the accuracy of road monitoring and object detection, especially in low-light environments. Each step combines physical models and image processing technology to ensure that the shape, color and motion of dynamic objects in the night scene are restored to be consistent with the daytime scene, thereby achieving a complete night-to-sunny image conversion.

[0150] S3. Based on the inferred images of dynamic objects in night scenes, a GAN model is designed for training and optimization, and image restoration and detail recovery are performed on the inferred images of dynamic objects in night scenes, and a low-resolution image is output. The low-resolution image can be close to the real daytime scene in terms of details and object morphology.

[0151] Specifically, the core goal of this step is to use the generative adversarial network (GAN) to infer the image I of the dynamic object generated in the previous step. night_to_day Image restoration and detail recovery. By specifically designing the GAN model, the inference and detail reconstruction of dynamic objects in complex night scenes are enhanced. This step is one of the core technologies of this patent and solves the problem of dynamic object recovery in low-light environments.

[0152] Furthermore, in order to better meet the needs of dynamic object restoration in low-light environments, the present invention designs an innovative generative adversarial network architecture, which combines the basic framework of traditional GAN ​​and uses some patented innovative elements to improve its effect in dynamic object inference and daytime scene restoration.

[0153] Generator Network Design:

[0154] The generator G is responsible for transforming the input image I night_to_day Through a series of convolutional and deconvolutional layers, the image details are gradually restored to generate images close to real daytime scenes. In order to enhance the generalization ability of the generator, a special conditional generator G is designed, which not only receives the input night scene image I night_to_day , and also receives prior information about daytime scenes That is, the daytime scene template estimated by the model. Specifically, the generation process of the generator is:

[0155]

[0156] Among them, I night_to_day is the input image, is the daytime scene prior information generated based on historical data or environmental models, θ G are the parameters of the generator.

[0157] Discriminator network design:

[0158] The function of the discriminator D is to determine whether the image is a real daytime image. Different from the design of traditional GAN, this patent designs a multi-level discriminator network, which not only judges the authenticity of the image, but also scores the reconstruction quality of dynamic objects. The discriminator network consists of two parts:

[0159] The first part is a standard two-classification network, which is used to judge the authenticity of the image.

[0160] The second part is a local evaluation network based on dynamic objects, which is used to evaluate the restoration quality of dynamic objects in the image. This local evaluation module conducts a detailed analysis of the moving objects in the image by learning the spatial distribution and motion trajectory of the objects in the image.

[0161] The goal of the discriminator is to determine whether the image is a real daytime image:

[0162]

[0163] Furthermore, during the training process, the generator and the discriminator are optimized in an adversarial manner. The present invention designs a comprehensive optimization objective to enhance the accuracy of image restoration, especially in the restoration of details of dynamic objects. This optimization objective not only considers the difference between the generated image and the real image, but also specifically introduces regularization terms related to dynamic object inference.

[0164] Furthermore, the optimization goal of the generator is:

[0165] The loss function L of the generator G It consists of two parts:

[0166] Standard adversarial loss: optimize the generator by minimizing the discriminative difference between generated images and real images.

[0167] Dynamic object reconstruction loss: additional regularization is performed on the dynamic object regions in the generated image, using the dynamic object information obtained from the previous step To constrain the shape and motion trajectory of these objects in the generated image. The specific loss function is:

[0168]

[0169] in, represents the dynamic object region inferred from the previous step, λ dynamic It is a weight coefficient used to adjust the influence of dynamic object reconstruction loss.

[0170] Furthermore, the optimization goal of the discriminator is:

[0171] The loss function L of the discriminator D It consists of two parts:

[0172] Standard adversarial loss: used to optimize the discriminator's ability to distinguish between generated images and real images.

[0173] Dynamic object evaluation loss: Based on the dynamic object annotations obtained in the previous step, the discriminator evaluates the restoration quality of dynamic objects and adds a specific dynamic object error term to the loss:

[0174]

[0175] in, represents the actual evaluation information of the dynamic object, λ dynamic_eval It is a regulation term used to strengthen the promotion of dynamic object evaluation loss on generator optimization.

[0176] Furthermore, after adversarial training and optimization, the image I output by the generator is reconstructed It can approximate the real daytime scene in terms of details and object morphology. This restoration process not only restores the changes in ambient lighting, but also accurately restores the details of dynamic objects, solving the problem of restoring dynamic objects in night scenes that is difficult to handle with traditional methods.

[0177] Furthermore, when evaluating the restoration effect, the present invention uses the peak signal-to-noise ratio (PSNR) and the structural similarity index (SSIM) as the main quality evaluation indicators. At the same time, for the restoration quality of dynamic objects, the present invention introduces the dynamic object quality evaluation indicator (DQM):

[0178]

[0179] in, is the restored object in the i-th image, is the true label of the dynamic object inferred in the previous step.

[0180] Through the above steps, the present invention not only effectively utilizes GAN to restore the details of dynamic objects, but also makes the image restoration process more accurate, especially in low-light environments, by introducing the regularization loss of dynamic object reconstruction. In addition, the innovation of this solution is that by designing a multi-level discriminator and a dynamic object-specific evaluation module, the restoration ability of dynamic objects is specially optimized, further improving the quality and usability of the generated images. This solution provides strong technical support for the practical application of the present invention in dynamic object inference and daytime scene restoration, and is particularly suitable for application scenarios in low-light environments such as traffic monitoring and security monitoring.

[0181] S4. Based on a low-resolution image of a night scene, super-resolution enhancement and spatiotemporal continuity enhancement are performed on the low-resolution image based on the night scene, and a super-resolution loss and a spatiotemporal consistency loss are generated respectively. Joint optimization is performed based on the determined super-resolution loss and spatiotemporal consistency loss to ensure the compatibility of the two goals, and an image based on the night scene with super-resolution and spatiotemporal continuity enhancement is generated.

[0182] Specifically, from the low-resolution image I reconstructed More details are restored, especially those of dynamic objects. The present invention adopts a convolutional neural network (CNN) architecture that integrates a spatial attention mechanism to improve image detail restoration and resolution. Steps:

[0183] A convolutional network is used to extract features from low-resolution images, and a spatial attention mechanism is introduced to strengthen attention to dynamic object areas. The role of the spatial attention mechanism is to automatically learn which areas are most important for recovering details, thereby improving the effect of detail recovery.

[0184] Generate high-resolution image I HR :

[0185] I HR =f SR (I reconstructed θ SR )

[0186] Among them, f SR represents the super-resolution function, θ SR is its parameter.

[0187] Furthermore, on the basis of generating high-resolution images, the motion consistency of dynamic objects between consecutive frames is ensured to avoid unnatural motion or breakage. To this end, a spatiotemporal consistency loss function is introduced to optimize the similarity between consecutive frames. Steps:

[0188] Spatiotemporal consistency loss function:

[0189] In order to enhance the spatiotemporal continuity, the present invention designs a loss function to measure the image difference between consecutive frames and constrain the spatiotemporal consistency of dynamic object regions. This loss function includes the measurement of spatial difference and temporal difference:

[0190]

[0191] Among them, I HR and are high-resolution images of the current frame and the previous frame respectively; ΔI HR and is the change of the dynamic object area between the current frame and the previous frame; α and β are weight coefficients that control the impact of the loss term.

[0192] Super-resolution and spatiotemporal consistency:

[0193] Target:

[0194] Joint optimization is performed on image super-resolution restoration and spatiotemporal consistency enhancement to ensure the compatibility of these two goals, thereby generating higher quality images.

[0195] The final optimization goal is to combine the super-resolution loss with the spatiotemporal consistency loss and obtain an overall loss function through weighted summation:

[0196] L final =L SR +λ·L temporal

[0197] Among them, L SR represents the super-resolution loss, which measures the difference between the high-resolution image and the target image; L temporal represents the spatiotemporal consistency loss, measuring the difference between consecutive frames; λ is a hyperparameter that adjusts the trade-off between super-resolution and spatiotemporal consistency.

[0198] Furthermore, the generated image sequences should maintain high resolution while the motion of dynamic objects should be coherent and consistent.

[0199] step:

[0200] To evaluate the quality of the generated images, PSNR and SSIM are used to measure the image clarity and structural similarity.

[0201] The temporal and spatial consistency evaluation index TQI is introduced to measure the motion coherence between consecutive frames:

[0202]

[0203] This formula measures the difference between consecutive frames. represents the high-resolution image of the i-th frame, and N is the total number of frames.

[0204] This step improves the image detail recovery capability and the spatiotemporal consistency between consecutive frames by combining super-resolution and spatiotemporal consistency enhancement. The innovative spatiotemporal consistency loss function ensures the coherent motion of dynamic objects, while the spatial attention mechanism enhances the detail recovery of dynamic objects. Through joint optimization, the present invention can generate high-quality high-resolution image sequences and solve the spatiotemporal consistency problem in low-resolution images and dynamic scene restoration.

[0205] S5. In generating super-resolution and spatiotemporal continuity enhanced images based on night scenes, high-frequency texture and edge information of the input image are restored to enhance local details of the image. Texture restoration and subtle edge enhancement are performed in low-contrast areas to generate detail-enhanced images. Finally, the detail-enhanced images are globally optimized to generate the final enhanced night road monitoring images.

[0206] Specifically, the restored image may have visual problems, such as poor contrast and blurred details, which need to be optimized through post-processing.

[0207] Furthermore, under the premise of ensuring the overall naturalness of the image, the high-frequency texture and edge information are restored through the detail enhancement method, and the local details of the image are enhanced to make the image more delicate and realistic. Steps:

[0208] Detail enhancement function:

[0209] By introducing an adaptive enhancement function f enhance , for the input image I HR The high-frequency components in the image are enhanced for details. This function combines edge detection and texture restoration to preserve the natural feel of the original image as much as possible while enhancing it:

[0210] I detail =f enhance (I HR ,θ enhance )

[0211] Among them, θ enhance is the parameter of the enhancement function, I detail is the image after detail enhancement. This function adjusts the enhancement strength according to the local features and contrast of the image, and automatically adjusts the degree of detail restoration in each area.

[0212] Adaptive enhancement factor:

[0213] In order to further optimize the effect of detail enhancement, an adaptive factor A is designed. adaptThis factor dynamically adjusts the intensity of detail enhancement according to the characteristics of the local area of ​​the image (such as edges, texture density, etc.). The specific calculation method is as follows:

[0214]

[0215] in, Represents image I HR The gradient at position (x, y), ||·|| 2 is the L2 norm. This factor enhances the details of image edges and high-contrast areas while suppressing unnecessary enhancement in flat areas.

[0216] Texture restoration and edge enhancement:

[0217] Target:

[0218] Restore textures and subtle edges in images, especially in low-contrast areas. By introducing local gradient information of the image and combining it with global optimization, the details can be improved without over-enhancement. Steps:

[0219] Texture restoration function:

[0220] The present invention designs a deconvolution layer based on a convolutional neural network (CNN) to restore image details. This layer focuses on restoring high-frequency textures of the image, especially in small areas of the image, such as skin texture, object surface, etc. The restoration process is represented by the following function:

[0221] I texture =f deconv (I detail ,θ deconv )

[0222] Among them, f deconv is the deconvolution function, θ deconv is the weight obtained by network training. The deconvolution layer restores the tiny texture of the image through the reverse convolution operation, thereby restoring the high-frequency details.

[0223] Edge Enhancement:

[0224] Using the gradient information and local texture features of the image, the present invention designs an edge enhancement module f edge , which is used to enhance the edge details of the image. This module significantly improves the edge sharpness in the image by enhancing the high-frequency information in the low-frequency part:

[0225] E enhanced =f edge (I texture ,θ edge )

[0226] Among them, θ edgeis the parameter of the edge enhancement module. This module uses gradient information to enhance the edge features of the image and make the details of the image clearer.

[0227] Further, global detail optimization:

[0228] Target:

[0229] Perform global optimization on the detail-enhanced image to ensure that the enhanced details are consistent with the overall image style while avoiding over-sharpening or unnatural effects. Steps:

[0230] Global consistency optimization:

[0231] In order to ensure that detail enhancement does not affect the global consistency of the image, the present invention introduces a global optimization loss function L global , which combines the naturalness of detail enhancement, edge restoration, and global structure:

[0232]

[0233] in, It is the L2 norm loss, which measures the difference between the restored image details and the original image; is the L1 norm of the image gradient, which is used to preserve edge features; 1 and λ 2 is a parameter that adjusts the weights of different loss terms. This optimization ensures that the details are enhanced while maintaining the naturalness of the image.

[0234] Furthermore, after detail enhancement and global optimization, the final image I obtained by the present invention is final Enhanced details are retained while over-enhancement and distortion are avoided, ensuring the naturalness and clarity of the image.

[0235] Further, output and evaluation:

[0236] Goal: The final output image should have clear details, enhanced edges and natural textures while maintaining the consistency of the overall structure. Steps:

[0237] Evaluate image quality: Use indicators such as the structural similarity index (SSIM) and peak signal-to-noise ratio (PSNR) to evaluate detail enhancement effects to ensure that the image quality improvement meets expectations.

[0238] Perceptual loss function: The perceptual loss function is used to further optimize the visual naturalness of the image and ensure that the image with enhanced details can achieve the best effect in human eye perception.

[0239] In this step, an innovative method for image detail enhancement and texture restoration is designed, which combines adaptive enhancement, deconvolution network and edge enhancement technology to ensure detail restoration while avoiding visual unnaturalness caused by over-enhancement. By introducing a global optimization loss function, the consistency and naturalness of the image in the detail enhancement process are further ensured, thereby generating high-quality restored images that meet the high-precision image restoration requirements in the patent.

[0240] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0241] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0242] Although embodiments of the present invention have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

[0243] The various illustrative logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein may be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in cooperation with a DSP core, or any other such configuration.

[0244] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor so that the processor can read and write information from / to the storage medium. In an alternative, a storage medium may be integrated into a processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor and the storage medium may reside in a user terminal as discrete components.

[0245] In one or more exemplary embodiments, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented as a computer program product in software, each function may be stored on or transmitted by a computer-readable medium as one or more instructions or codes. Computer-readable media include both computer storage media and communication media, including any medium that facilitates the transfer of a computer program from one place to another. Storage media may be any available medium that can be accessed by a computer. As an example and not limitation, such a computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer. Any connection is also properly referred to as a computer-readable medium. For example, if the software is transmitted from a website, a server, or other remote source using a coaxial cable, a fiber optic cable, a twisted pair, a digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwaves, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of the medium. Disk and disc as used herein include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, wherein disk often reproduces data magnetically, while disc reproduces data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0246] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein, but should be granted the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for enhancing images of road monitoring at night, characterized in that: The method comprises: collecting and preprocessing road monitoring images from different sensors in daytime scenes, and performing multimodal feature fusion based on the preprocessed road monitoring images to generate fused features, wherein the fused features are obtained by weighted fusion of the multimodal features; wherein the road monitoring images include visible light images, infrared images, and depth images; In a plurality of road monitoring images acquired by different sensors, a daytime static scene model is constructed based on a plurality of preprocessed road monitoring images, dynamic objects are inferred based on the daytime static scene model using infrared images and depth images at night, and the result of the dynamic object inference is matched with the output of the daytime static scene model to restore the dynamic objects in the night scene, and generate a dynamic object inference image based on the night scene; Based on the inferred images of dynamic objects in night scenes, a GAN model is designed for training and optimization, image restoration and detail recovery are performed on the inferred images of dynamic objects in night scenes, and low-resolution images are output. The low-resolution images can be close to real daytime scenes in terms of details and object morphology; Based on a low-resolution image of a night scene, super-resolution enhancement and spatiotemporal continuity enhancement are performed on the low-resolution image based on the night scene, a super-resolution loss and a spatiotemporal consistency loss are generated respectively, and joint optimization is performed based on the determined super-resolution loss and spatiotemporal consistency loss to ensure the compatibility of the two objectives, thereby generating an image based on the night scene with super-resolution and spatiotemporal continuity enhancement; In generating super-resolution and spatiotemporal continuity enhanced images based on night scenes, high-frequency texture and edge information of the input image are restored to enhance local details of the image. At the same time, texture restoration and subtle edge enhancement are performed in low-contrast areas to generate detail-enhanced images. Finally, the detail-enhanced images are globally optimized to generate the final enhanced night road monitoring images.

2. The method for enhancing a nighttime road monitoring image according to claim 1, characterized in that: Preprocess the road monitoring images, including: First, sensor data alignment of different sensors is performed; Secondly, the sensor data of each modality is normalized and denoised; Finally, feature extraction is performed on the standardized and denoised data.

3. The method for enhancing nighttime road monitoring images according to claim 2, characterized in that: Feature extraction for different road monitoring images, including: Edge detection and color histogram extraction methods are used to extract features of visible light images, and to extract edge information and color distribution of images: E vis =Canny(I vis ) H vis =Histogram(I vis ) Among them, E vis Represents the edge information of the visible light image, H vis is the color histogram; Extract features of infrared images using temperature gradient and hotspot identification methods: in, Represents the temperature gradient of infrared images, reflecting the location and intensity of temperature changes. This method helps to identify heat source areas in infrared images; Deep edge detection and spatial connectivity analysis methods extract features of depth images: in, Represents the spatial gradient of the depth image, reflecting the areas with large depth changes in the image.

4. The method for enhancing images of nighttime road monitoring according to claim 1, characterized in that: Extract geometric information of static objects based on visible light images and infrared images, including: The object boundaries in the visible light image are extracted by edge detection, and the temperature gradient in the infrared image is used to distinguish the heat source area; The spatial distance information of the object is extracted through the depth image to obtain the three-dimensional shape of the object; A daytime static scene model output based on the geometric information of static objects, including the position information and shape information of each object and the relative position relationship between objects; Based on the daytime static scene model, the position, shape and motion trajectory of dynamic objects are inferred using infrared images and depth images at night to obtain dynamic object inference, including: Shape inference for dynamic objects: Assuming that dynamic objects appear as areas with large temperature gradients in infrared images, and that this gradient change is closely related to the shape of the object, the contours of dynamic objects can be identified by analyzing the temperature changes in infrared images. In this process, the geometric information of the daytime static field model is combined to confirm the position and shape of the object. Color inference for dynamic objects: The color of dynamic objects during the day is inferred through the heat source distribution extracted from the daytime visible light image and the infrared image. Assuming that the color change of dynamic objects during the day and night is limited, the color mapping function is used to infer the color of dynamic objects: C dynamic =f(I vis ,I IR ,S scene ) Among them, C dynamic represents the color of dynamic objects, f is the color mapping function extracted from daytime visible light images and infrared images, and the color of dynamic objects is estimated using the daytime static field model as a reference; For the motion estimation of dynamic objects: The shape of dynamic objects is inferred through depth images in night scenes, and the motion trajectory of the objects is calculated. The depth information provides the position change of the objects in space. Assuming that the dynamic objects in the night scenes maintain a certain speed and motion trajectory, the displacement of the objects is calculated through continuous depth images, and their future motion direction and speed are inferred.

5. The method for enhancing images of nighttime road monitoring according to claim 4, characterized in that: The result based on the dynamic object inference is matched with the output of the daytime static scene model, including: First, use the daytime static scene model and dynamic object inference for spatial alignment and color matching: In a night scene, the color difference between the dynamic object and the day scene needs to be corrected by image transformation. The present invention designs a mapping function h(·) to map the color of the dynamic object to the color range of the day scene: ΔC match =h(C dynamic ,C day ) Where, ΔC match is the color difference between the dynamic object and the daytime scene, h(·) is the color mapping function; The inferred dynamic object is fused with the original night scene image to obtain the restored image of the night scene: By matching the color and shape of dynamic objects to the daytime scene, a daytime-like effect can be recreated in night images: Among them, I night_to_day is the restored image, represents the image fusion operation, O dynamic Contains information about inferred dynamic objects.

6. The method for enhancing images for nighttime road monitoring according to claim 5, characterized in that: The GAN model is constructed as follows: The generator G not only receives the input night scene image I night_to_day , and also receives prior information about daytime scenes That is, the daytime scene template estimated by the daytime static scene model is generated in the following process: Among them, I fake Generate images for the generator G, I night_to_day is the restored image, as the input image, is the daytime scene prior information generated based on historical data or environmental models, θ G are the parameters of the generator; The discriminator D aims to determine whether the image is a real daytime image, which is constructed as follows: The first part is a standard two-classification network, which is used to judge the authenticity of the image; The second part is a local evaluation network based on dynamic objects, which is used to evaluate the restoration quality of dynamic objects in the image. The local evaluation network conducts a detailed analysis of the moving objects in the image by learning the spatial distribution and motion trajectory of the objects in the image. The generator and discriminator of the GAN model are optimized through adversarial means, specifically including: The optimization objectives of the generator include standard adversarial loss and dynamic object reconstruction loss, specifically: Standard adversarial loss: optimize the generator by minimizing the discriminative difference between generated images and real images. Dynamic object reconstruction loss: additional regularization is performed on the dynamic object regions in the generated image, and the obtained dynamic object regions are used to constrain the shape and motion trajectory of these objects in the generated image; The optimization objectives of the discriminator include standard adversarial loss and dynamic object evaluation loss, specifically: Standard adversarial loss: used to optimize the discriminator's ability to distinguish between generated images and real images. Dynamic object evaluation loss: Based on dynamic object annotations, the discriminator evaluates the restoration quality of dynamic objects and adds the dynamic object error term to the loss.

7. The method for enhancing images for nighttime road monitoring according to claim 6, characterized in that: After adversarial training and optimization of the GAN model, peak signal-to-noise ratio and structural similarity index were used as the main quality evaluation indicators. At the same time, the dynamic object quality evaluation indicator DQM was introduced for the restoration quality of dynamic objects: in, is the restored object in the i-th image, is the true label of the dynamic object inferred in the previous step.

8. The method for enhancing images of nighttime road monitoring according to claim 6, characterized in that: The super-resolution enhancement uses a convolutional neural network architecture that integrates a spatial attention mechanism to improve image detail recovery and resolution, specifically including: Use convolutional networks to extract features from low-resolution images, and introduce a spatial attention mechanism to strengthen attention to dynamic object areas, and then generate high-resolution images; The spatiotemporal continuity enhancement ensures the consistency of motion of dynamic objects between consecutive frames to avoid unnatural motion or breaks, specifically including: Design a loss function to measure the image difference between consecutive frames and constrain the spatiotemporal consistency of dynamic object areas. The loss function L temporal Contains measures of spatial and temporal differences: Among them, I HR and are high-resolution images of the current frame and the previous frame respectively; ΔI HR and is the change of the dynamic object area between the current frame and the previous frame; α and β are weight coefficients that control the impact of the loss term; Joint optimization is performed in image super-resolution restoration and spatiotemporal consistency enhancement, including: L final =L SR +λ·L temporal Among them, L SR represents the super-resolution loss, which measures the difference between the high-resolution image and the target image; L temporal represents the spatiotemporal consistency loss, measuring the difference between consecutive frames; λ is a hyperparameter that adjusts the trade-off between super-resolution and spatiotemporal consistency; PSNR and SSIM are used to measure the clarity and structural similarity of the image, and to quantitatively evaluate the quality of the generated image. At the same time, the temporal and spatial consistency evaluation index TQI is introduced to measure the motion coherence between consecutive frames: in, represents the high-resolution image of the i-th frame, and N is the total number of frames.

9. The method for enhancing images for nighttime road monitoring according to claim 8, characterized in that: By introducing the adaptive enhancement function f enhance , for the input image I HR Enhance the details of the high-frequency components in: I detail =f enhance (I HR ,θ enhance ) Among them, θ enhance is the parameter of the enhancement function, I detail It is the image after detail enhancement; Redesign adaptive factor A adapt , this factor dynamically adjusts the intensity of detail enhancement according to the characteristics of the local area of ​​the image. The specific calculation method is as follows: in, Represents image I HR The gradient at position (x,y), ||·||2 is the L2 norm.

10. The method for enhancing images of nighttime road monitoring according to claim 9, characterized in that: A deconvolution layer based on a convolutional neural network is designed to restore image details. At the same time, an edge enhancement module is designed by utilizing the gradient information and local texture features of the image to enhance the edge details of the image. The edge enhancement module significantly improves the edge sharpness in the image by enhancing high-frequency information in the low-frequency part. After restoring image details and enhancing image edge details, global optimization is performed to obtain the final enhanced night road monitoring image.

Citation Information

Patent Citations

  • Nighttime monitoring video enhancing method

    CN103020930A

  • Multi-sensor fusion low-illumination video image enhancement method

    CN105809640A

  • Night light image optimization method based on spatial information fusion multi-attention network

    CN114926349A

  • Generative adversarial network-based power transmission line image enhancement method under low illuminance

    CN115601644A

  • End-to-end color and detail enhancement method, device and equipment in low-illumination scene

    CN117274107A

Cited By

  • Super-resolution method and system based on edge enhancement and frequency domain optimization, and medium

    CN121481847A

  • Multi-mode perception data fusion enhancement method and system for coal mine unmanned vehicle

    CN122023144A

  • Multi-modal perception data fusion enhancement method and system for coal mine unmanned vehicle

    CN122023144B