A park night low-illumination image enhancement method

CN122736900APending Publication Date: 2026-09-11SHAANXI CONSTR ENG HLDG GRP FUTURE CITY INNOVATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610899462.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0011]为了克服以上现有技术存在的缺陷,本发明的目的在于提供一种园区夜晚低照度图像增强方法,基于深度学习技术,结合频域分析与空域注意力机制的卷积神经网络,在单一模型中实现去噪、去模糊和光照增强的联合优化,解决现有技术在混合降质场景下效果不佳及计算效率低的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122736900A_ABST
    Figure CN122736900A_ABST
Patent Text Reader

Abstract

A park night low-illumination image enhancement method, the method comprises the following steps: step S1: extracting a key frame as an initial monitoring image; step S2: performing format conversion, Bayer array removal operation and normalization processing on the initial monitoring image to obtain a to-be-processed input tensor; step S3: inputting the to-be-processed input tensor into a low-illumination image enhancement neural network model which is pre-constructed and trained; step S4: outputting a deep low-resolution feature map; step S5: the deep low-resolution feature map output by the encoder is used to generate a low-resolution illumination guide map, and the features containing illumination information are transmitted to the corresponding level of the decoder through the jump connection; step S6: outputting a high-resolution feature; step S7: outputting a final high-quality clear monitoring image. The present application is based on deep learning technology, and combines a convolutional neural network of frequency domain analysis and spatial attention mechanism to realize the joint optimization of denoising, deblurring and illumination enhancement in a single model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of nighttime image enhancement technology for parks, and specifically to a method for enhancing low-light nighttime images of parks. Background Technology

[0002] Video surveillance systems are core infrastructure for maintaining park security, preventing crime, tracing incidents, and managing production. Unlike the well-lit environment of daytime, nighttime surveillance faces more severe imaging challenges, mainly in the following aspects: The park area is vast, and lighting facilities often cannot cover all corners (such as the edges of walls, gaps in storage areas, and deep within woods). In environments with extremely low illumination (<5 Lux), the number of photons received by the sensors (CMOS / CCD) of surveillance cameras is scarce. To obtain a visible image, the ISP (Image Signal Processor) automatically increases the gain (ISO). However, high ISO directly introduces severe shot noise and readout noise. This noise not only manifests as dense specks but also destroys high-frequency details in the image, causing a "snow screen" phenomenon that obscures crucial security features (such as the intruder's facial features and clothing texture).

[0003] To increase the amount of light entering the camera, nighttime surveillance cameras typically extend the exposure time (e.g., from 1 / 50 second to 1 / 25 second or longer). However, this results in a severe side effect: motion blur.

[0004] When vehicles are moving or people are running within the park, the objects shift during the exposure time, resulting in motion blur and ghosting in the image. Furthermore, cameras mounted on poles may experience global blur due to wind or vibration. This blur is an irreversible degradation, rendering license plate recognition (LPR) and facial recognition algorithms virtually ineffective at night.

[0005] The park experiences extremely uneven lighting at night. Areas under streetlights may be overexposed, while areas in shadow are completely black. While existing wide dynamic range (WDR) technology can alleviate this to some extent, its effectiveness is limited in extremely low light conditions.

[0006] In low-light images, noise, blur, and insufficient brightness are tightly coupled. Traditional image processing pipelines typically address these issues step-by-step: denoising first and then brightening, or brightening first and then denoising. However, traditional denoising algorithms (such as mean filtering and BM3D) tend to smooth out textures, leading to increased blur; while traditional deblurring algorithms (such as Wiener filtering and RL deconvolution) are extremely sensitive to noise, easily amplifying it and producing ringing artifacts.

[0007] To address the above problems, academia and industry have proposed various solutions, but limitations still exist: Methods such as histogram equalization (HE), gamma correction, and the Retinex Theoretical Algorithm (MSRCR) often amplify noise while enhancing brightness, and are prone to color shift, making them unable to handle motion blur.

[0008] Single-task deep learning methods: (1) Low Light Enhancement (LLIE) Networks: such as RetinexNet and Zero-DCE. These networks focus on adjusting the illumination curve or decomposing the reflectance / illuminance map. They usually assume that the input image is sharp. When the input is a blurry nighttime image, the model will enhance the blur as texture, resulting in an output image that is brighter but still blurry, or even produces artifacts.

[0009] (2) Deblurring networks: such as DeblurGAN-v2 and MIMO-UNet. These networks are mostly trained on synthetic datasets (such as GoPro) and assume that the input image is well lit and has low noise. When directly applied to low-light images at night, the kernel estimation or feature extraction of the model will fail due to severe noise interference, resulting in a significant drop in deblurring performance.

[0010] Therefore, there is an urgent need for an end-to-end image enhancement method that can simultaneously handle low illumination, strong noise, and non-uniform blur, while also being computationally efficient, having a small number of parameters, and exhibiting strong robustness, in order to meet the actual needs of all-weather security monitoring in smart parks. Summary of the Invention

[0011] To overcome the shortcomings of the existing technologies, the present invention aims to provide a method for enhancing low-light images of parks at night. Based on deep learning technology, it combines a convolutional neural network with frequency domain analysis and spatial attention mechanism to achieve joint optimization of denoising, deblurring and illumination enhancement in a single model, thus solving the problems of poor performance and low computational efficiency of existing technologies in mixed degraded scenarios.

[0012] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for enhancing low-light images of a park at night, the method comprising the following steps: Step S1: Data Acquisition The initial monitoring video stream under low light conditions at night is obtained by image acquisition equipment deployed at various key monitoring points in the park, and key frames are extracted as the initial monitoring images. Step S2: Preprocessing steps: The initial monitoring image is subjected to format conversion, Bayer array removal, and normalization to map pixel values ​​to intervals, thus obtaining the input tensor to be processed. ; Step S3: Deep network inference steps: The input tensor to be processed The input is into a pre-built and trained low-light image enhancement neural network model (DARNet); the low-light image enhancement neural network model adopts an asymmetric encoder-decoder architecture; it aims to jointly solve the tasks of illumination restoration and deblurring; Step S4: Encoding and Illumination Enhancement Steps: In the encoder, a hierarchical encoding module (EBlock) is used to process the input tensor. Perform downsampling and feature extraction to output a deep, low-resolution feature map; Step S5: Illumination Guidance and Bottleneck Transfer Steps: The deep, low-resolution feature map output by the encoder is used to generate a low-resolution illumination guide map. And through Skip Connections, features containing lighting information are passed to the corresponding layer of the decoder; Step S6: Decoding and Deblurring Steps: In the decoder, a hierarchical decoding module (DBlock) is used to upsample and reconstruct details of features containing illumination information, outputting high-resolution features; Step S7: Post-processing and output steps: The high-resolution features output by the decoder are used to generate a residual map or directly generate an enhanced image through the output convolutional layer. ,right After performing inverse normalization and color space correction, the final high-quality and clear surveillance image is output.

[0013] In step S1, the initial monitoring images are collected by deploying high-definition network cameras in key areas such as the perimeter fence, main entrances and exits, cargo yards and parking lots of the park.

[0014] The initial surveillance images are characterized by low signal-to-noise ratio, low brightness contrast, and non-uniform blur caused by relative motion of objects or camera shake.

[0015] The preprocessing step S2 specifically involves: The initial monitoring image is subjected to format conversion and deemosaicing operations to convert the RAW format data of the initial monitoring image into an RGB three-channel image; Subsequently, in order to eliminate the differences in numerical range caused by different devices and lighting conditions, and to adapt to the input requirements of the neural network, the RGB three-channel image was normalized. The specific formula is as follows: Let the pixel values ​​of the RGB three-channel image be... ,in For spatial coordinates, For color channels; first convert integer pixel values ​​to floating-point numbers, then map them to color channels using a linear transformation. Interval: in, The range of values ​​is After normalization, the input tensor to be processed is obtained. Its numerical range is .

[0016] In step S3, the design of the low-light image enhancement neural network model (DARNet) is based on the Metaformer paradigm. The specific calculation process of the frequency domain enhancement module (EBlock) in the encoder is as follows: Let the input tensor to be processed be... RGB three-channel image pixel values After layer normalization (LayerNorm), we get ; First, the Spatial Attention (SAM) submodule is used, and its calculation formula is as follows: in, Represents the normalized features of the input. This indicates the spatial attention submodule. This represents the features output by the spatial attention module; The system then proceeds to the frequency domain feedforward network submodule (FGEM), which includes: in, Fast Fourier Transform; : Amplitude component of frequency domain signal; Phase components of frequency domain signals; Multilayer sensing for processing frequency domain amplitude; : A complex exponential term constructed from phase information; Inverse Fast Fourier Transform; Characteristics of the output of a frequency domain feedforward network; The final output feature of the entire module.

[0017] This indicates a multilayer perceptron that operates only on the amplitude spectrum, used to learn the global gain of illumination enhancement. The neural network model (DARNet) constructed in this invention adopts an asymmetric encoder-decoder architecture. The term "asymmetric" means that the encoder and decoder have different focuses in terms of functional design and structural complexity, and are not simply mirror-symmetric.

[0018] Its core design philosophy is to achieve task decoupling, allowing different parts of the network to focus on solving different degradation problems: The encoder focuses on "illuminance restoration." Its role is to progressively downsample and extract deep semantics and global illumination information from the image. Degradation in low-light images is primarily manifested in the attenuation of the frequency domain amplitude spectrum. Therefore, a frequency-guided enhancement module (FGEM) is designed into the encoder. Utilizing the global properties of the Fourier transform, it efficiently adjusts the illumination distribution in the low-resolution feature space to suppress noise. The encoder's structure is relatively deeper, emphasizing the extraction of global information and illumination correction.

[0019] The decoder focuses on deblurring and detail reconstruction. Its role is to progressively upsample the low-resolution features extracted by the encoder to restore the image's details and spatial structure. Blur, especially non-uniform motion blur, manifests as the loss of local information or ghosting in the spatial domain, requiring a large receptive field to cover it. Therefore, the decoder incorporates a multi-scale receptive field attention module (MSRFA), utilizing multi-branch dilated convolutions to expand the receptive field without significantly increasing the number of parameters, thus deconstructing complex blur kernels. The decoder's structure prioritizes the fusion of multi-scale features and the reconstruction of fine structures.

[0020] This asymmetric design introduces low-light guiding loss ( This further strengthens task separation. The loss function supervises the deepest layer (bottleneck layer) of the encoder, forcing the encoder to generate a feature map with correct brightness but low resolution, thus ensuring that the illumination restoration task is basically completed in the encoder stage. The decoder, on this basis, focuses on using the original structural information passed from skip connections and its own large receptive field to complete the tasks of deblurring and detail sharpening. This design avoids mutual interference between the two tasks, making the entire model more efficient and robust in solving the mixed degradation problem of "dark + noisy + blurry". In step S3 (deep network inference), the overall architecture of DARNet consists of an encoder and a decoder. Step S4 (encoding and illumination enhancement) specifically implements the EBlock processing in the encoder, including the SAM and FGEM modules, for illumination restoration. Step S5 (illumination guidance and bottleneck transport) uses the features of the deepest layer of the encoder to generate a low-resolution illumination guidance map and passes it to the decoder through skip connections. Step S6 (decoding and deblurring step) specifically implements the DBlock processing in the decoder, including the MSRFA and AGFFN modules, for deblurring and detail reconstruction. Step S7 (Post-processing and Output) reconstructs the final enhanced image from the decoder output. These steps are sequentially linked to complete the end-to-end enhancement.

[0021] In step S4, the EBlock includes a spatial attention submodule (SAM) and a frequency domain feedforward network submodule (FGEM). The FGEM analyzes feature maps. Perform a two-dimensional real fast Fourier transform (RFFT2) to separate the amplitude spectrum and phase spectrum in the frequency domain. Use a learnable global gain mask to perform nonlinear mapping enhancement on the amplitude spectrum, keep the phase spectrum unchanged or make fine adjustments, and then restore the spatial domain features through inverse transform (IRFFT2), thereby recovering the global illumination distribution in the low-resolution feature space.

[0022] In step S6, the DBlock includes a Multi-Scale Receptive Field Attention (MSRFA) submodule and an Adaptive Gated Feedforward Network (AGFFN) submodule. The MSRFA utilizes multi-path parallel dilated convolution branches to capture multi-scale spatial context, thereby resolving the non-uniform blur kernel. The AGFFN employs a simple gating mechanism (SimpleGate) for feature selection and nonlinear transformation. The specific structure of the Multi-Scale Receptive Field Attention (MSRFA) submodule in the decoder includes: Input features go through After dimensionality reduction by convolution, the stream splits into three parallel branches of depthwise convolution. The input feature X is an intermediate feature output by the encoder or the previous module of the DARNet model. This feature is a high-dimensional feature representation of the original low-light image after encoding and illumination enhancement. The first branch uses the expansion rate of Depthwise convolution is used to capture local detailed features; The second branch uses the expansion rate of Depth convolution is used to capture contextual information at a medium scale. The third branch uses the expansion rate or larger Deep convolution is used to capture long-range spatial dependencies and large-scale motion blur features; The output features of the three branches are concatenated or summed element-wise along the channel dimension, and then passed through a... Convolution is used for feature fusion, and adaptive recalibration of channel weights is performed through a Simplified Channel Attention (SCA) module.

[0023] The Simplified Channel Attention (SCA) module removes the Sigmoid activation function and complex fully connected layers from the traditional channel attention mechanism. Its calculation method is as follows: global average pooling is performed on the input features, channel weights are calculated through a linear layer, and channel-wise multiplication is performed directly with the original features to preserve the dynamic range of feature magnitudes and adapt to the numerical sensitivity requirements of image restoration tasks.

[0024] The specific calculation formula for the SimpleGate mechanism used in the Adaptive Gated Feedforward Network (AGFFN) submodule is as follows: Input features pass Convolution doubles the number of channels, resulting in a feature map. ; Will Divided into two parts along the channel dimension ; Output features ,in This represents element-wise multiplication; The SimpleGate mechanism does not include explicit activation functions such as ReLU, GELU, or Sigmoid; instead, it utilizes the mutual gating of the data itself to achieve nonlinear transformation.

[0025] The training process of the neural network model employs a composite loss function. Perform end-to-end supervised optimization, the composite loss function It consists of the following four weighted parts: in: , is the pixel-domain L1 loss, used to constrain the enhancement image. With real and clear images Pixel-level consistency between pixels ensures accuracy in color and brightness; To detect loss, multi-layer feature maps are extracted using a pre-trained VGG19 network, and the loss is calculated. and The Euclidean distance in the feature space is used to improve the visual perception quality of the image and restore texture details that conform to human vision. For edge loss, the L2 norm of the feed difference in the horizontal and vertical directions of the image is calculated to force the network to sharpen the image edges and suppress the ringing effect generated during the deblurring process. To address the low-light guiding loss, the feature map output by the encoder is mapped to an RGB image downsampled by 8 times. and compared with the downsampled real image Calculate the L1 distance to supervise the encoder's focus on the illumination recovery task.

[0026] The preferred value of the weighting coefficient is: .

[0027] A low-light image enhancement system for a park at night, the system architecture includes a front-end perception layer, an edge computing layer, a data transmission layer and a central management platform, which can fully support the implementation of the entire process of low-light image enhancement methods for parks at night, and realize end-to-end operation from image acquisition, preprocessing, deep network inference to enhanced output and intelligent analysis; Front-end perception layer: It consists of several low-light high-definition cameras distributed around the perimeter of the park, entrances and exits, parking lots and main roads, used to collect video data 24 hours a day. Edge computing layer: Edge computing boxes or NVR devices deployed at the park's aggregation nodes, with built-in AI acceleration chips, are used to run a lightweight version of the DARNet model and perform frame-by-frame or frame-by-frame enhancement processing on the real-time video stream. Data transmission layer: Based on fiber optic LAN or 5G private network, responsible for transmitting enhanced video streams or key frame alarm data to the central server; Central Management Platform: Deployed in the park's monitoring center, it includes a video storage server, a large-screen display system, and an intelligent analysis server; the intelligent analysis server receives enhanced, clear images and further performs tasks such as face recognition, license plate recognition, or intrusion detection.

[0028] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor implements the method when executing the program, and the processor may be a GPU, NPU, FPGA, or dedicated ASIC chip.

[0029] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method.

[0030] The park includes, but is not limited to, industrial production parks, logistics and warehousing parks, scientific research and development parks, and urban public parks; the method is designed for data augmentation training targeting complex light source interference (such as strong street light and vehicle headlight glare) and extremely low illumination shadow areas (illuminance below 0.1 Lux) unique to the park at night. The training dataset includes synthetic samples simulating different ISO noise levels and different motion speed blur kernels, as well as real collected samples.

[0031] The beneficial effects of this invention are: (1) All-in-One: The DARNet model proposed in this invention achieves task decoupling and joint optimization through an asymmetric encoder-decoder architecture. The frequency domain guided enhancement module (FGEM) in the encoder uses Fourier transform to adjust the amplitude spectrum in the frequency domain, thereby enhancing illumination and suppressing noise. The multi-scale receptive field attention module (MSRFA) in the decoder captures blur features at different scales through multi-branch dilated convolution, achieving spatial deblurring. This design enables a single model to solve the three major problems of low illumination, strong noise, and non-uniform blur simultaneously end-to-end, avoiding the error accumulation and computational resource waste caused by traditional serial connection of multiple models.

[0032] (2) Efficiency: The number of parameters is reduced by about 55% compared to LEDNet and by about 88% compared to Restormer. By combining FFT and dilated convolution, MACs (multiply-accumulate operations) are significantly reduced while maintaining state-of-the-art performance, making it possible to deploy on edge devices.

[0033] (3) Scene Adaptability (Robustness): This method is specifically optimized for complex nighttime scenes in the park. During the training phase, data augmentation with different ISO noise levels and motion blur kernels is used to enable the model to learn to handle non-Gaussian noise in the real world (such as shot noise and readout noise). The FGEM module performs nonlinear mapping of the amplitude spectrum in the frequency domain, which can adaptively suppress various types of noise; the multi-scale dilated convolution of the MSRFA module can cover various blur kernel sizes from slight jitter to fast motion, and uses simplified channel attention (SCA) to adaptively weight the blurred regions, thereby effectively handling non-uniform blur. The augmented images significantly improve the accuracy of subsequent AI analysis (such as face recognition and license plate recognition). Attached Figure Description

[0034] Figure 1 This is a hardware deployment topology diagram of the system of the present invention, including front-end cameras, edge computing nodes and central servers.

[0035] Figure 2 This is a schematic diagram of the overall architecture of the deep neural network model (DARNet) in this invention, showing the topology of the encoder, decoder and intermediate feature transmission.

[0036] Figure 3 This is a detailed structural diagram of the encoder module (EBlock) in this invention, showing the connection relationship between Spatial Adaptive Modulation (SAM) and Frequency Domain Guided Enhancement (FGEM).

[0037] Figure 4 This is a detailed structural diagram of the decoder module (DBlock) in this invention, showing the connection relationship between the multi-scale receptive field attention (MSRFA) and the adaptive gated feedforward network (AGFFN). Detailed Implementation

[0038] The present invention will now be described in further detail with reference to the accompanying drawings.

[0039] This invention discloses a method for enhancing low-light images of a park at night, which specifically includes the following parts: Part 1: Hardware and Deployment of the Park's Nighttime Surveillance Image Enhancement System like Figure 1 As shown in the figure, this embodiment constructs a complete intelligent nighttime monitoring system for industrial parks, aiming to solve the problem of poor visibility at night in industrial parks.

[0040] 1. Front-end data acquisition unit: (1) Deploy high-definition network cameras (IPCs) in key areas such as the perimeter fence, main entrance and exit, cargo yard and parking lot of the park.

[0041] (2) Camera configuration: Resolution of 1920x1080 or higher, supporting H.265 encoding. In night mode, the aperture is wide open (e.g., F1.4), and the ISO auto gain limit is set to 3200-6400 to ensure brightness, but this will introduce a lot of noise. The shutter speed is set to 1 / 25s to accommodate live video, but this will cause blurring of moving vehicles.

[0042] (3) Data source characteristics: The acquired images have low illumination (<5 Lux), high noise (Gaussian + Poisson mixed distribution) and motion blur characteristics.

[0043] 2. Edge Computing Node: (1) Deploy edge computing boxes (such as those based on NVIDIA Jetson Xavier or Orin series, or domestic AI computing cards) in the camera aggregation box or NVR backend.

[0044] (2) Function: It is responsible for decoding the video stream, extracting frames (such as processing 5-15 key frames per second), and running the DARNet enhancement model of this invention.

[0045] (3) Advantages: Enhancement processing at the edge can reduce transmission bandwidth pressure and ensure localized processing of privacy data, uploading only the enhanced clear images or alarm information to the center.

[0046] 3. Enhance processing servers and central platform: (1) Deployed in the monitoring center and equipped with a high-performance GPU server.

[0047] (2) Function: Receive data returned from the edge end, or perform real-time full frame rate enhancement on key channels.

[0048] (3) Application: The enhanced video stream is sent to the large-screen display system for security personnel to monitor, and simultaneously sent to the AI ​​analysis server for secondary analysis (such as face recognition and license plate recognition). A closed-loop system of "end-edge-cloud" collaboration is formed between the front-end acquisition unit, edge computing nodes and the central platform: the front-end camera acquires the original video stream in real time and transmits it to the edge computing node through the park network; the edge node, as the core processing unit, runs a lightweight DARNet model to enhance the video stream in real time and uploads the enhanced key frames or alarm information to the central platform; the central platform is responsible for global storage, large-screen display and deep intelligent analysis, and can remotely configure and update the algorithms of the edge nodes. The data transmission layer ensures efficient and reliable communication between the layers, realizing full-process collaboration from data acquisition, real-time processing to deep applications.

[0049] Network architecture design: The neural network model constructed in this invention adopts an asymmetric encoder-decoder structure and follows the Metaformer architecture paradigm (i.e., the structure of Token Mixer + Channel Mixer).

[0050] Encoder: Focused on illumination recovery; Based on the physical characteristic that illumination information is mainly concentrated in the frequency domain amplitude spectrum, the encoder introduces a frequency domain guided enhancement module (FGEM). Decoder: Focuses on deblurring and detail recovery; based on the characteristics of blurring which usually has non-uniform spatial distribution and long-distance dependence, the decoder introduces a multi-scale receptive field attention module (MSRFA).

[0051] In low-light images, the degradation of the frequency domain guided enhancement module (FGEM) is mainly manifested as attenuation and distortion of the amplitude spectrum, while the phase spectrum mainly encodes the structural information (edges, shape) of the image.

[0052] Implementation: This module first performs a 2D Fast Fourier Transform (FFT) on the feature map. In the frequency domain, a lightweight multilayer perceptron (MLP) or convolutional network is applied to the amplitude component for nonlinear mapping to enhance low-frequency illumination components and suppress high-frequency noise components. The phase component remains unchanged or is only fine-tuned to preserve the original structural information to the greatest extent possible. Finally, the spatial domain is recovered through an inverse Fourier transform (IFFT). This design leverages the global receptive field characteristics of the frequency domain to achieve global illumination adjustment with extremely low computational cost.

[0053] The Multi-Scale Receptive Field Attention Module (MSRFA) deblurring task requires a sufficiently large receptive field to cover the size of the blur kernel. Traditional methods increase the kernel size (e.g., ...). , This can lead to a surge in the number of parameters.

[0054] Implementation: MSRFA employs a multi-path parallel dilated convolution strategy. Branch 1: Dilation Rate (Ordinary convolution), focusing on local texture; Branch 2: dilation rate Focus on mesoscale object structures; Branch 3: Expansion rate (or larger), focusing on large-scale motion trajectories and long-distance contexts.

[0055] In the MSRFA module, three parallel dilated convolutional branches capture local details, medium-scale context, and large-scale motion features, respectively, achieving multi-scale perception. Subsequently, the outputs of each branch are concatenated or added to fuse features, aggregating multi-scale information into a unified feature representation. Finally, the fused features are channel-weighted by the Simplified Channel Attention (SCA) module's gating mechanism, achieving adaptive recalibration. Multi-scale perception is the foundation of feature extraction, feature fusion summarizes the information, and the gating mechanism selectively enhances the summarized features. These three elements work in a progressive and synergistic manner to constitute MSRFA's complete contextual modeling capability. This design significantly expands the effective receptive field without increasing the number of parameters, effectively addressing the blurring caused by the rapid motion of vehicles and people in the park.

[0056] SimpleGate and Simplified Channel Attention (SCA): To achieve lightweight design and avoid limit vanishing, the front bound network (FFN) in the network removes traditional activation functions such as ReLU / GELU and uses SimpleGate instead. (Split the feature map channel in half and multiply them together).

[0057] Channel attention employs SCA: it removes the Sigmoid and complex dimensionality reduction operations, directly uses linear layers to calculate channel scaling factors, preserves the amplitude information of features, and is more suitable for image restoration and regression tasks.

[0058] Composite loss function optimization: This invention employs a combination of loss functions tailored for hybrid degradation tasks during the training phase: Basic pixel-level reconstruction error ensures the fidelity of image content.

[0059] Feature-space perceptual distance based on VGG network. Compared to traditional MSE / PSNR, LPIPS can better optimize the subjective visual quality of images, avoiding overly smooth or "plastic" images.

[0060] Constraints enhance the consistency between the image and the real image within the limit domain, forcing the network to recover sharp edges and combat blur.

[0061] This is an architecture-guided loss. A supervisory signal is added to the deepest layer (low-resolution layer) of the encoder, forcing the encoder to output the correct low-resolution lighting map. This achieves implicit task decoupling: the encoder is responsible for lighting correction, and the decoder is responsible for detail correction.

[0062] This invention is particularly applicable to solving the problem of decreased video surveillance image quality in large parks such as industrial parks, logistics and warehousing areas, and urban parks under conditions of insufficient nighttime lighting, strong noise interference, and motion blur, aiming to improve the effectiveness and intelligence level of nighttime security monitoring.

[0063] Part Two: Core Process of Image Enhancement Methods like Figure 2 As shown, the method flow of the present invention includes the following steps: Step S1: Input Acquisition and Preprocessing S1.1: By calling the camera SDK or pulling the RTSP video stream, the video stream is decoded, and frames are extracted from the decoded image sequence at fixed time intervals as the original image frames. This operation is performed by the edge computing node.

[0064] S1.2 Normalization: The original image frame obtained in S1.1 Typically in 8-bit integer format, pixel values ​​range from 0 to 255. S1.2 converts it to a floating-point number and normalizes it: Mapped to From the interval, we obtain the tensor. Map pixel values ​​from floating-point space.

[0065] S1.3 Size Adjustment: To adapt to network inference efficiency, the image size can be adjusted. Scale or crop to a fixed size (e.g.) or Alternatively, a patch-based inference strategy can be used to process 4K large images.

[0066] Step S2: DARNet network inference: The preprocessed input tensor is then processed... Input the DARNet model.

[0067] The model is divided into four-scale encoder stages and four-scale decoder stages.

[0068] In step S3, DARNet employs a four-level encoder and a four-level decoder: input tensor After shallow convolution, the image sequentially enters four encoder stages (E-Stage 1–4), where the feature map resolution is halved and the number of channels is doubled in each stage. After the bottleneck layer, the features sequentially enter four decoder stages (D-Stage 4–1), where the feature map resolution is doubled and the number of channels is halved in each stage, ultimately reconstructing an enhanced image with the same resolution as the input. .

[0069] Step S3: Encoder Illumination Recovery (Encoder Stage) S3.1 Shallow Feature Extraction: First, through a... Convolution converts images The mapping is to the feature space.

[0070] This operation serves as the starting point for DARNet inference, and will The feature map is mapped to a shallow layer using a 3×3 convolution. , This then serves as the input for the first encoder stage (E-Stage1), upon which all subsequent encoding operations are performed.

[0071] S3.2 Level Processing: The encoder consists of 4 levels (Level 1-4). Each level is composed of several stacked EBlocks.

[0072] EBlock Explained (see) Figure 3 ): (1) Input features First, it goes through LayerNorm.

[0073] (2) SAM (Spatial Adaptive Modulation): It uses depthwise separable convolution to extract spatial features and adjusts the channel weights through SCA to make the network focus on the main region in the image.

[0074] (3) FGEM (Detailed Logic): For input features Perform RFFT2 transform to obtain the complex spectrum. .

[0075] Calculate the amplitude spectrum and phase spectrum .

[0076] Amplitude enhancement: Amplitude spectrum Input to a lightweight MLP (consisting of two (Composed of convolution and ReLU), learn a global gain mask. .

[0077] This step utilizes the global nature of frequency domain amplitude, adjusting the amplitude of low-frequency components to enhance overall brightness and suppressing the amplitude of high-frequency components to remove noise.

[0078] Inverse transform: utilizing the enhanced amplitude and the original phase The complex spectrum is reconstructed and converted back to spatial features using IRFFT2.

[0079] (4) Residual connection: The output features are fused with the enhanced illumination information in the frequency domain.

[0080] S3.3 Downsampling: S3.2 (i.e., EBlock processing) describes feature extraction within each encoder stage; between layers, downsampling is performed using convolutions with a stride of 2, gradually reducing the feature map resolution. The number of channels gradually increases. ).

[0081] Step S4: Low-light guidance (Bottleneck Guidance): Step S4 is a sub-step within step S3 (deep network inference), located between the encoder output and the decoder input. The feature maps from the deepest layer of the encoder are used to generate low-resolution illumination guidance maps. and through Loss supervision ensures that the illumination correction task is completed during the encoder stage. Only then is the feature map (containing the corrected illumination information) fed into the decoder to begin deblurring and reconstruction.

[0082] S4.1: In the deepest layer of the encoder (Level 4), the feature map resolution is... The network here is accessed through a... Convolution outputs a 3-channel low-resolution image. .

[0083] S4.2: During training, compute low-resolution images. Downsampled version of the real, clear image L1 loss between This forces the encoder to complete the illumination correction task at this stage, ensuring that the features fed into the decoder are "normally bright but blurry".

[0084] Step S5: Deblurring and Detail Reconstruction (Decoder Stage): First, the features from the previous level are upsampled, and then fused with the skip connection features from the corresponding level of the encoder in step S4. The fused features are used as the input to the first DBlock in the current decoder stage.

[0085] S5.1: Upsampling: Upsampling is performed using transposed convolution or bilinear interpolation + convolution.

[0086] S5.2: Skip connection: Concatenate or add the features of the corresponding level of the encoder with the features of the decoder to supplement the lost spatial details.

[0087] DBlock Explained (see...) Figure 4 ): (1)MSRFA (Multi-Scale Receptive Field Attention Module): The input features are processed by LayerNorm.

[0088] Multi-scale perception: Input tensor to be processed It is fed into three parallel depthwise convolution branches. Branch A: Conv, Dilation=1. Receptive field Branch B: Conv, Dilation=4; Sensitive field Branch C: Conv, Dilation = 9 (or greater). Receptive field .

[0089] Fusion: The outputs of the three branches are fused to cover blur kernels of various sizes (from tiny jitters to fast motion blurs).

[0090] Attention weighting: The fused features are recalibrated by calculating weights through the SCA module.

[0091] (2) AGFFN (Adaptive Gated Feed-Forward Network): Input features ,pass Convolution increases the dimension of channels to Divided into , Output The recalibration is applied to the enhanced features of the final output. This mechanism achieves efficient non-linear activation without using Sigmoid / ReLU, preserving more limit flow and aiding in the training of deep networks.

[0092] Step S6: Image Reconstruction and Output: Step S5 (decoding and deblurring) gradually restores image details through multiple DBlocks, ultimately obtaining a high-quality feature map with the same resolution as the input. Step S6 performs output convolution on this feature map (generating a residual map or directly predicting the image), followed by post-processing such as inverse normalization, to output the final enhanced image. .

[0093] S6.1: Output residual map of the last layer of the decoder .

[0094] S6.2: Final Enhanced Image (or direct prediction) ).

[0095] S6.3: For Perform inverse normalization, truncate to the desired value, and output the video stream.

[0096] Three: Model training strategy: To obtain robust model parameters, this invention employs a sophisticated training strategy.

[0097] 1. Dataset Construction: (1) Basic data: The public dataset LOLBlur was used, which contains 10,200 pairs of synthesized low-light blurred / sharp images.

[0098] (2) Real-world data adaptation: To adapt to the specific park scenario, this invention collected real nighttime videos of the park. Since a perfect "ground truth" of "clear illumination" cannot be obtained, a semi-supervised learning or transfer learning strategy was adopted. Fine-tuning was performed using the Real-LOLBlur dataset.

[0099] (3) Data augmentation: Gaussian noise, Poisson noise, random occlusion, and random rotation and flipping are randomly added during training to simulate the harsh environment of the park.

[0100] 2. Detailed Explanation of the Loss Function: The total loss function defined in this invention is: Pixel consistency ( Using L1 Norm, the formula is: Compared to MSE (L2), L1 is better at preserving the edge sharpness of the image.

[0101] Perceived quality ( The algorithm employs LPIPS (Learned Perceptual Image PatchSimilarity). The predicted and ground truth images are input into a pre-trained VGG19 network, and feature maps are extracted from layers conv1_2, conv2_2, conv3_2, conv4_2, and conv5_2. The mean squared error of these feature maps is calculated and then summed using weighted averages. This loss function makes the enhanced image more consistent with human perception in terms of texture and structure, eliminating the "smearing" effect.

[0102] Edge Fidelity The Sobel operator is applied to the image to obtain a limit map. The L2 loss between the limit maps is calculated. This directly constrains the network to eliminate edge diffusion caused by motion blur.

[0103] Low-light guidance ( Specifically designed for encoder output. Feature maps are used for supervision. This ensures that the FGEM module truly learns the lighting enhancements, rather than deferring the lighting task to the decoder, thus achieving decoupling between structure and lighting.

[0104] 3. Hyperparameter settings: Input image size cropped to Batch Size set to 32; Optimizer AdamW, Weight decay Initial learning rate The cosine annealing strategy was used to decay the value to... Training cycles are 200 epochs.

[0105] IV. Performance Comparison and Experimental Results: To verify the effectiveness of the method of the present invention, comparative experiments were conducted on a standard test set and in a real campus scenario.

[0106] 1. Comparison Experiment with Standard Test Set Table 1: Comparison of Objective Indicators (LOLBlur Dataset) method PSNR (dB)↑ SSIM ↑ LPIPS ↓ Number of parameters (M) Computational complexity (G-MACs) RetinexNet 17.68 0.542 0.510 - - DeblurGAN-v2 22.30 0.745 0.356 60.9 - NAFNet 25.36 0.882 0.158 12.05 30+ LEDNet 25.74 0.850 0.224 7.4 33.74 DARNet (This invention) 27.00 0.883 0.162 3.31 7.25 analyze: (1) Image quality: The present invention (DARNet) surpasses the previous best method LEDNet by about 1.26 dB in PSNR, which is a significant improvement in the field of image restoration. The SSIM and LPIPS indicators also reach the best or second best, indicating that the image structure is well preserved and the visual perception is clear.

[0107] (2) Lightweight Advantage: The number of parameters in this invention is only 3.31M, less than half that of LEDNet (7.4M), and only one-twentieth that of DeblurGAN-v2. The computational cost (MACs) is further reduced to 7.25G, which means that under the same hardware, the inference speed of this invention is more than 4 times that of LEDNet. This feature makes it a perfect fit for the real-time and edge deployment requirements of campus monitoring.

[0108] 2. Real-world scenario testing In a nighttime parking lot test at a technology park (light intensity < 1 Lux), the original surveillance footage was completely black and full of noise. After enabling the algorithm of this invention, the license plate recognition rate increased from 15% to over 85%, and the face detection rate within 20 meters of the camera increased from 0% to 70%. For vehicles traveling at 30 km / h, motion blur was effectively removed, and the vehicle outlines were clear.

Claims

1. A park night low-illumination image enhancement method, characterized in that, The method includes the following steps: Step S1: Acquire the initial monitoring video stream under low light conditions at night by using image acquisition devices deployed at various key monitoring points in the park, and extract key frames as the initial monitoring images; Step S2: Perform format conversion, Bayer array removal, and normalization on the initial monitoring image to map pixel values ​​to intervals, thereby obtaining the input tensor to be processed. ; Step S3: Convert the input tensor to be processed The input is a pre-built and trained low-light image enhancement neural network model; the low-light image enhancement neural network model adopts an asymmetric encoder-decoder architecture; Step S4: In the encoder, the hierarchical encoding module EBlock is used to process the input tensor. Perform downsampling and feature extraction to output a deep, low-resolution feature map; Step S5: The deep low-resolution feature map output by the encoder is used to generate a low-resolution illumination guide map. And through skip connections, the features containing lighting information are passed to the corresponding layer of the decoder; Step S6: In the decoder, a hierarchical decoding module is used to upsample and reconstruct details of features containing illumination information, and output high-resolution features; Step S7: The high-resolution features output by the decoder are processed by the output convolutional layer to generate a residual map or directly to generate an enhanced image. ,right After performing inverse normalization and color space correction, the final high-quality and clear surveillance image is output.

2. The method for enhancing low-light images in a park at night according to claim 1, characterized in that, In step S1, the initial monitoring images are collected by deploying high-definition network cameras in key areas of the park, including the perimeter fence, main entrances and exits, cargo yards, and parking lots.

3. The method for enhancing low-light images in a park at night according to claim 1, characterized in that, The preprocessing step S2 specifically involves: The initial monitoring image is format converted and Bayer array removed to convert the RAW format data of the initial monitoring image into an RGB three-channel image; Normalize the RGB three-channel image; The normalization process is as follows: Let the pixel values ​​of the RGB three-channel image be... ,in For spatial coordinates, For color channels; first convert integer pixel values ​​to floating-point numbers, then map them to color channels using a linear transformation. Interval: in, The range of values ​​is After normalization, the input tensor to be processed is obtained. Its numerical range is .

4. The method for enhancing low-light images in a park at night according to claim 3, characterized in that, In step S3, the specific calculation process of the frequency domain enhancement module in the encoder is as follows: Let the input tensor to be processed be... RGB three-channel image pixel values After layer normalization, we get ; First, through the spatial attention submodule, the calculation formula is as follows: in, Represents the normalized features of the input. This indicates the spatial attention submodule. This represents the features output by the spatial attention module; The program then proceeds to the frequency domain feedforward network submodule, which includes: in, Fast Fourier Transform; : Amplitude component of frequency domain signal; Phase component of a frequency domain signal; Multilayer sensing for processing frequency domain amplitude; : A complex exponential term constructed from phase information; Inverse Fast Fourier Transform; Characteristics of the output of a frequency domain feedforward network; The final output feature of the entire module.

5. The method for enhancing low-light images in a park at night according to claim 4, characterized in that, In step S4, the EBlock includes a spatial attention submodule and a frequency domain feedforward network submodule; FGEM analyzes feature maps A two-dimensional real-number fast Fourier transform is performed to separate the amplitude spectrum and phase spectrum in the frequency domain. The amplitude spectrum is enhanced by nonlinear mapping using a learnable global gain mask, while the phase spectrum remains unchanged or is finely adjusted. Then, the spatial domain features are restored through inverse transform, thereby restoring the global illumination distribution in the low-resolution feature space. The enhanced features output by FGEM, as the features of the corresponding layer of the encoder, are passed to the decoder stage in step S5 through skip connections. They are then fused with the upsampled features of the decoder to form the input features of the DBlock module in the decoder. .

6. The method for enhancing low-light images in a park at night according to claim 5, characterized in that, In step S6, the DBlock includes a multi-scale receptive field attention submodule MSRFA and an adaptive gated feedforward network submodule AGFFN. MSRFA utilizes multi-path parallel dilated convolution branches to capture multi-scale spatial context in order to deconstruct non-uniform blur kernels; AGFFN utilizes a simple gating mechanism for feature selection and nonlinear transformation; The specific structure of the multi-scale receptive field attention submodule in the decoder includes: Input features go through After dimensionality reduction by convolution, the stream is split into three parallel depthwise separable convolution branches; The first branch uses the expansion rate of Depthwise convolution is used to capture local detailed features; The second branch uses the expansion rate of Depth convolution is used to capture contextual information at a medium scale. The third branch uses the expansion rate or larger Deep convolution is used to capture long-range spatial dependencies and large-scale motion blur features; The output features of the three branches are concatenated or element-wise added along the channel dimension, and then passed through a... Convolution is used for feature fusion, and adaptive recalibration of channel weights is performed by simplifying the channel attention module.

7. The method for enhancing low-light images in a park at night according to claim 6, characterized in that, In step S6, the adaptive gated feedforward network submodule adopts a simplified channel attention module, and its calculation method is as follows: For input features Global average pooling is performed, and channel weights are calculated through a linear layer and directly compared with the original features. Perform channel-by-channel multiplication.

8. The method for enhancing low-light images in a park at night according to claim 7, characterized in that, In step S6, the specific calculation formula for the simple gating mechanism is as follows: Input features pass Convolution doubles the number of channels, resulting in a feature map. ; feature map Divided into two parts along the channel dimension ; Output features ,in This represents element-wise multiplication; The SimpleGate mechanism does not contain explicit activation functions such as ReLU, GELU, or Sigmoid, but uses the mutual gating of the data itself to achieve nonlinear transformation. The training process of the neural network model employs a composite loss function. Perform end-to-end supervised optimization, the composite loss function; Weighted by the following four parts composition: in: , is the pixel-domain L1 loss, used to constrain the enhancement image. With real and clear images Pixel-level consistency between them; To detect loss, multi-layer feature maps are extracted using a pre-trained VGG19 network to calculate the true sharp image. With constrained image enhancement The Euclidean distance in the feature space is used to improve the visual perception quality of the image and restore texture details that conform to human vision. For edge loss, the L2 norm of the feed difference in the horizontal and vertical directions of the image is calculated to force the network to sharpen the image edges and suppress the ringing effect generated during the deblurring process. To address the low-light guiding loss, the feature map output by the encoder is mapped to an RGB image downsampled by 8 times. and compared with the RGB image after downsampling from the real image. Calculate the L1 distance to supervise the encoder's focus on the illumination recovery task.

9. A low-light image enhancement system for a park at night for implementing the method according to any one of claims 1-8, characterized in that, The system architecture includes a front-end perception layer, an edge computing layer, a data transmission layer, and a central management platform; The front-end perception layer consists of several low-light high-definition cameras distributed around the perimeter of the park, entrances and exits, parking lots and main roads. It is used to continuously collect the initial monitoring video stream in low-light environments at night for 24 hours a day, extract key frames as the initial monitoring images, and provide the raw input for subsequent enhancement processing. The edge computing layer is deployed in edge computing boxes or NVR devices at the aggregation nodes of the park. It has a built-in AI acceleration chip to run a lightweight version of the DARNet model. It sequentially completes the format conversion of the initial monitoring image, Bayer array removal, normalization preprocessing, as well as network inference for encoding illumination enhancement, illumination-guided transmission, decoding deblurring and detail reconstruction, and performs frame-by-frame or frame-by-frame enhancement processing on the real-time video stream. The data transmission layer is based on a fiber optic LAN or a 5G private network and is responsible for transmitting high-quality enhanced video streams or key frame alarm data that have been denormalized and color space corrected to the central server. The central management platform is deployed in the park's monitoring center and includes a video storage server, a large-screen display system, and an intelligent analysis server. The intelligent analysis server receives clear enhanced images output from the edge and performs tasks such as face recognition, license plate recognition, or intrusion detection to support the park's security monitoring and intelligent management.