Night monitoring video enhancement method based on space-time joint and photometric adaptive mapping

By employing a spatiotemporal joint and photometric adaptive mapping method, the problems of noise avalanche and temporal flicker in nighttime surveillance video enhancement are solved, achieving high-quality video enhancement effects. This method is suitable for embedded security monitoring terminals and improves the performance of security systems.

CN122312424APending Publication Date: 2026-06-30CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNIV OF GEOSCIENCES (WUHAN)
Filing Date
2026-03-21
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing methods for enhancing nighttime surveillance videos are prone to noise avalanche when brightening dark areas, and the lack of consistency in inter-frame processing leads to temporal flickering. At the same time, it is difficult to achieve a balance between suppressing noise and preserving the sharpness of motion edges.

Method used

Based on the spatiotemporal joint and photometric adaptive mapping method, a temporally correlated monitoring video sequence is constructed. The video sequence is decoupled into low-frequency illumination features and high-frequency reflection features by using a spatiotemporal joint feature decoupling network. Differentiated processing strategies are used to process these two types of features respectively. Finally, a bright video frame with high dynamic range, high signal-to-noise ratio and temporal smoothness is reconstructed by the Retinex imaging model.

Benefits of technology

It completely solves the timing flicker problem of traditional algorithms, avoids the noise avalanche effect, and achieves a balance between brightening dark areas, preventing overexposure in bright areas, sharpening moving edges, and real-time processing. It is compatible with embedded security monitoring terminals and improves the early warning capability and evidence collection effectiveness of security monitoring systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122312424A_ABST
    Figure CN122312424A_ABST
Patent Text Reader

Abstract

This application provides a nighttime surveillance video enhancement method based on spatiotemporal joint and photometric adaptive mapping, belonging to the field of video image processing. The method includes: firstly, constructing a temporally correlated video frame sequence; then, decoupling the video sequence into a low-frequency illuminance feature sequence and a high-frequency reflectance feature sequence through a dual-branch decoder network; next, applying an iterative photometric adaptive mapping function to adjust the brightness of the low-frequency illuminance features, and applying deformable convolution and three-dimensional dynamic filtering to perform spatiotemporal alignment and denoising processing on the high-frequency reflectance features; finally, fusing the processed illuminance and reflectance features to reconstruct the enhanced video frame. This invention solves the problem of inter-frame temporal flicker and noise avalanche caused by brute-force brightening in traditional algorithms by completely decoupling and differentially parallelizing illuminance and reflectance components, improving details in dark areas while preserving edge sharpness, and achieving high-quality real-time enhancement of nighttime surveillance videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video image processing, and in particular to a method for enhancing nighttime surveillance video based on spatiotemporal joint and photometric adaptive mapping. Background Technology

[0002] With the comprehensive construction of safe cities and intelligent transportation systems, video surveillance systems have achieved full coverage of public areas and key locations. Among them, nighttime monitoring, as a core link in public security prevention and control, accident investigation, and evidence collection, directly determines the overall effectiveness of the security system through its imaging quality. When public security incidents, traffic violations, and safety accidents occur in low-light environments at night, the raw videos captured by surveillance cameras generally suffer from defects such as dark images, loss of detail, extremely low signal-to-noise ratio, and poor target recognition. These defects severely affect the recognition accuracy of backend intelligent analysis algorithms, leading to problems such as missed target detection, misjudgment, and tracking failure, and failing to meet the industrial-grade application requirements of security monitoring.

[0003] Nighttime surveillance scenarios generally suffer from inherent problems such as low ambient light, uneven light distribution, coexistence of localized strong light and large dark areas, and severe motion blur of moving targets. In addition, the CMOS / CCD image sensors of surveillance cameras have a significant decrease in photon collection efficiency in low-light environments. As a result, the proportion of thermal noise, shot noise, and fixed-pattern noise in the signal increases sharply. The signal-to-noise ratio (SNR) of the original video is usually below 10dB, which is in the extremely low SNR range, further exacerbating the degradation of video image quality.

[0004] To address the aforementioned issues, existing technologies have primarily focused on low-light image / video enhancement. However, these technologies suffer from insurmountable core pain points in both industrial implementation and technical applications, specifically: 1. Current mainstream low-light enhancement solutions often treat single-frame images as independent processing units, lacking customized designs for the temporal context of video sequences. Because the illumination estimation and brightness mapping for each frame are solved independently, minute illumination differences between adjacent frames are nonlinearly amplified, leading to severe inter-frame temporal flicker. This flicker causes backend intelligent analysis algorithms to generate numerous false features, resulting in target misses, frequent tracking switching, and misjudgments, making the enhanced video unable to meet the industrial-grade application requirements of security monitoring. Existing inter-frame smoothing filtering post-processing methods can only superficially suppress flicker and are prone to secondary defects such as blurred edges of moving targets. 2. Regarding brightening mechanisms, existing global brightness stretching schemes do not distinguish between brightness and detail components. While brightening dark areas, this leads to an exponential noise avalanche effect in high-frequency noise. Combined with the long exposure mode of surveillance cameras, aggressive brightening further amplifies edge diffusion and ghosting of moving targets, resulting in the complete loss of key texture details such as faces and license plates. 3. Existing sequential processing workflows, such as brightening first and then denoising or denoising first and then brightening, cannot effectively suppress noise while preserving weak texture details. In addition, existing denoising solutions mostly use fixed-kernel 3D convolution for spatiotemporal filtering, which cannot adapt to the dynamic motion offset of targets in monitoring and easily aggravates motion trajectory ghosting. Summary of the Invention

[0005] The purpose of this invention is to address the problems of existing nighttime surveillance video enhancement methods, such as the noise avalanche effect when brightening dark areas, the lack of consistency in inter-frame processing leading to temporal flicker, and the difficulty in achieving a balance between noise suppression and preservation of motion edge sharpness. This invention provides a nighttime surveillance video enhancement method based on spatiotemporal joint and photometric adaptive mapping.

[0006] The above-mentioned objective of this application is achieved through the following technical solution: S1: Based on the acquired low-light raw video stream, construct a time-series correlated monitoring video sequence; S2: Input the surveillance video sequence into a spatiotemporal joint feature decoupling network based on a physical optics model to obtain a low-frequency illumination feature sequence and a high-frequency reflection feature sequence. S3: The decoupled low-frequency illuminance feature sequence and high-frequency reflection feature sequence are processed in parallel using differentiated processing strategies based on their physical characteristics to obtain target illuminance features and target reflection features. S4: Element-wise multiplication is used to fuse the target illumination features and target reflection features, and the bright-state video frames with high dynamic range, high signal-to-noise ratio and smooth temporal sequence are reconstructed based on the Retinex imaging model to enhance the nighttime surveillance video.

[0007] A nighttime surveillance video enhancement system based on spatiotemporal joint and photometric adaptive mapping, deployed in security monitoring terminals, edge computing nodes, or network video recorder equipment, includes the following functional modules: Video sequence construction module: Composed of video acquisition unit, preprocessing unit and sequence buffering unit, it is used to acquire the low-light raw video stream collected by night security equipment, and construct a monitoring video sequence with temporal context after preprocessing; Spatiotemporal joint feature decoupling module: Built-in pre-trained weight-shared Siamese dual-branch spatiotemporal joint feature decoupling network, which decouples the surveillance video sequence into low-frequency illumination feature sequence and high-frequency reflection feature sequence based on the Retinex physical optics model; The differentiated parallel processing module includes a photometric adaptive mapping unit and a spatiotemporal alignment denoising unit, which work in parallel. The photometric adaptive mapping unit uses an iterative exposure control equation to adjust the brightness of the low-frequency illuminance feature sequence and outputs the target illuminance feature. The spatiotemporal alignment denoising unit uses deformable convolution and three-dimensional dynamic filtering to perform spatiotemporal alignment and denoising on the high-frequency reflection feature sequence and outputs the target reflection feature. The reconstruction output module consists of a feature fusion unit, a post-processing unit, and an output unit. It performs element-wise multiplication fusion of target illumination features and target reflection features to reconstruct bright-state video frames, which are then output to display devices, storage devices, or back-end intelligent analysis platforms after post-processing. The video acquisition unit acquires video streams via ONVIF, RTSP, or GB / T 28181 national standard protocol; the buffer capacity of the sequence buffer unit is flexibly adjusted according to the number of historical frames N; the output unit outputs to the display device in real time via HDMI interface, writes to the storage device via SATA interface, and transmits to the back-end intelligent analysis platform via network interface.

[0008] An electronic device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to enable the electronic device to perform a nighttime surveillance video enhancement method based on spatiotemporal joint and photometric adaptive mapping.

[0009] A computer-readable storage medium storing instructions that, when executed, perform a nighttime surveillance video enhancement method based on spatiotemporal joint and photometric adaptive mapping.

[0010] The beneficial effects of the technical solution provided in this application are: 1. This invention fundamentally solves the temporal flicker problem of traditional algorithms, achieving highly smooth and consistent brightness between frames. Traditional low-light enhancement algorithms treat each frame as an independent processing unit, resulting in uncorrelated illumination estimations between frames. Minor differences are nonlinearly amplified, causing severe flicker and leading to misjudgments and missed detections in backend intelligent analysis. This invention introduces a temporal context frame sequence from the input source, abandoning the independent processing mode of each frame and fundamentally avoiding the problem of inconsistent illumination estimation between frames. Through a weight-sharing twin encoder structure, it retains the independent feature dimensions of each frame, preventing premature fusion of temporal information. It applies specific temporal constraints to the illumination features that determine the brightness of the image, and simultaneously constrains the consistency of illumination between adjacent frames from the bottom layer through temporal smoothing loss during the training phase. This forms a triple guarantee of input temporal correlation, processing temporal constraints, and training temporal optimization, completely eliminating the temporal flicker problem caused by brightness jumps between frames in traditional algorithms, and providing stable video input for backend intelligent analysis.

[0011] 2. This invention resolves the core contradiction between brightening and denoising in traditional algorithms, completely avoiding the noise avalanche effect caused by brute-force brightening. Traditional algorithms employ a global brightness stretching scheme, failing to distinguish between brightness and detail components. Brightening dark areas leads to an exponential amplification of high-frequency noise. Furthermore, the sequential process of brightening before denoising or vice versa fails to preserve weak texture details while suppressing noise. This invention, based on the Retinex physical optics model, completely decouples the illuminance component, which determines brightness, from the reflection component, which determines detail. It employs a differentiated parallel processing strategy for the physical characteristics of these two components: nonlinear adaptive brightening of the low-frequency illuminance component adjusts only the image brightness without affecting high-frequency details, thus avoiding noise amplification; and time-aligned spatiotemporal joint denoising of the high-frequency reflection component is performed before brightening, fundamentally avoiding the noise avalanche problem caused by brute-force brightening in traditional algorithms.

[0012] 3. Customized optimization for security monitoring scenarios, addressing the poor adaptability and high computational cost of traditional algorithms, and possessing strong engineering applicability. Traditional denoising and inter-frame smoothing algorithms mostly use fixed-kernel 3D convolution for spatiotemporal filtering, which cannot adapt to the dynamic offset of moving targets in monitoring scenarios, easily exacerbating motion blur and incurring high computational costs, making it difficult to meet the real-time requirements of embedded devices. This invention is customized for the characteristics of security monitoring, such as numerous moving targets, high real-time requirements, and limited embedded computing power: it uses deformable convolution to achieve precise feature alignment of moving targets. Compared with traditional fixed-kernel 3D convolution, it can significantly reduce the computational cost while adapting to the dynamic offset of targets, avoiding motion blur, preserving edge sharpness, and conforming to the computing power limitations of security hardware, meeting the requirements of industrial-grade real-time processing.

[0013] 4. Improves scene adaptability and robustness, solving the problems of weak generalization ability and the need for manual parameter tuning in traditional algorithms, covering all types of nighttime monitoring scenarios. Traditional low-light enhancement algorithms have fixed brightness mapping parameters, requiring manual parameter tuning for different nighttime monitoring scenarios, resulting in poor generalization ability and problems such as overexposure and insufficient brightening of dark areas in complex lighting scenarios. The photometric adaptive mapping function of this invention adopts a learnable parameterized iterative equation, which can automatically adjust the brightness mapping curve according to the light distribution of different scenarios, eliminating the need for manual parameter tuning; at the same time, the network is trained with a large amount of nighttime monitoring video data with different lighting and scenarios, possessing extremely strong generalization ability, adapting to video streams of different resolutions such as 1080P and 4K, and compatible with various security hardware devices such as IPC cameras, network video recorders, and edge computing nodes. It can maintain stable enhancement effects in various complex nighttime monitoring scenarios such as no light, low light, backlight, local strong light, and vehicle headlight flicker, and can be widely used in all types of nighttime low-light monitoring scenarios such as urban security, park monitoring, traffic checkpoints, border control, and warehouse guarding.

[0014] 5. Enhanced hardware adaptability and reduced deployment costs address the issues of traditional algorithms' high dependence on high-performance hardware and difficulty in integrating with existing security systems. Traditional high-end low-light enhancement algorithms are complex and computationally intensive, requiring deployment only on high-performance servers. They are incompatible with existing embedded security monitoring terminals, exhibit poor protocol compatibility, and face significant challenges in integrating with existing security systems and high upgrade costs. The algorithm of this invention employs a lightweight network architecture design, allowing direct deployment on existing embedded security monitoring terminals, network video recorders, and edge computing nodes. No new hardware is required; functional upgrades to existing monitoring systems are achieved solely through software upgrades. Furthermore, the algorithm fully supports mainstream video transmission protocols such as ONVIF, RTSP, and GB / T 28181, enabling seamless integration with existing security monitoring systems. This flexible and convenient deployment significantly reduces the upgrade and maintenance costs of security systems. Attached Figure Description

[0015] The present application will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a step diagram of an embodiment of this application; Figure 2 This is a schematic diagram of the structure in an embodiment of this application; Figure 3 This is a schematic diagram of curve comparison in the embodiments of this application; Figure 4 This is a schematic diagram of the processing flow in the embodiments of this application; Figure 5 These are comparison images of enhanced image effects in the embodiments of this application; Figure 6 This is a schematic diagram of the electronic device structure in the embodiments of this application. Detailed Implementation

[0016] To provide a clearer understanding of the technical features, objectives, and effects of this application, the specific embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0017] The embodiments of this application provide a method for enhancing nighttime surveillance video based on spatiotemporal joint and photometric adaptive mapping.

[0018] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating the steps of a nighttime surveillance video enhancement method based on spatiotemporal joint and photometric adaptive mapping in an embodiment of this application, including: S1: Based on the acquired low-light raw video stream, construct a time-series correlated monitoring video sequence; S2: Input the surveillance video sequence into a spatiotemporal joint feature decoupling network based on a physical optics model to obtain a low-frequency illumination feature sequence and a high-frequency reflection feature sequence. S3: The decoupled low-frequency illuminance feature sequence and high-frequency reflection feature sequence are processed in parallel using differentiated processing strategies based on their physical characteristics to obtain target illuminance features and target reflection features. As one example, for the decoupled low-frequency illuminance feature sequence and high-frequency reflection feature sequence, a completely differentiated processing strategy is adopted to process them in parallel according to their physical characteristics and optimization objectives, so as to solve the core problems of brightness adjustment and noise suppression respectively, while ensuring the real-time performance of the processing.

[0019] S4: Element-wise multiplication is used to fuse the target illumination features and target reflection features, and the bright-state video frames with high dynamic range, high signal-to-noise ratio and smooth temporal sequence are reconstructed based on the Retinex imaging model to enhance the nighttime surveillance video.

[0020] This application addresses two core pain points—temporal flickering caused by single-frame processing and noise avalanche caused by aggressive brightening—by adopting the aforementioned technical solution. Simultaneously, it achieves a balance between multiple objectives: brightening of dark areas, prevention of overexposure in bright areas, sharpening of moving edges, and real-time processing. It is compatible with various security hardware devices such as embedded security monitoring terminals, network video recorders, and edge computing nodes, fully meeting the industrial-grade application requirements of nighttime security monitoring, and providing high-quality video input for backend intelligent analysis, target detection, and security evidence collection.

[0021] As one embodiment, this invention can be widely applied to real-time enhancement processing of surveillance videos in various low-light nighttime scenarios, such as urban security, park monitoring, traffic checkpoints, border control, and warehouse guarding. Specifically, it is compatible with hardware devices such as embedded security monitoring terminals, network video recorders, and edge computing nodes, providing high-quality, high-temporal-consistency, and high-edge-sharp video input for backend intelligent analysis, target detection and tracking, behavioral event recognition, and security evidence collection. It can effectively solve core problems such as dark surveillance images, loss of detail, temporal flicker, and excessive noise in low-light nighttime scenarios, significantly improving the early warning capabilities and evidence collection effectiveness of security monitoring systems.

[0022] As one embodiment, firstly, the brightened target illuminance feature L* and the purified target reflectance feature R* are fused using element-wise multiplication. Based on the Retinex imaging model, a bright video frame I* with high dynamic range, high signal-to-noise ratio, and temporal smoothness is reconstructed. The reconstructed bright video frame I* is then post-processed. The post-processing operations are as follows: pixel values ​​are denormalized from the [0,1] interval to the pixel bit depth range of the original video ([0,255] for 8-bit and [0,65535] for 16-bit); simultaneously, contrast fine-tuning and color correction are performed to ensure realistic color reproduction that conforms to human visual perception. The post-processed bright video frames are output through multiple channels: real-time output to the display device of the security monitoring terminal for real-time image display and observation by security personnel; written to storage devices such as solid-state drives (SSDs) via SATA interface for subsequent security incident evidence collection, violation investigation and intelligent analysis; and transmitted to the back-end intelligent analysis platform via network interface to provide high-quality input data for algorithms such as target detection, multi-target tracking and behavior event recognition.

[0023] Step S1 includes: Acquire raw low-light video streams collected by nighttime security monitoring equipment; Based on the current processing time point, extract the target video frames to be enhanced from the original low-light video stream. Simultaneously extract N historical reference frames that are temporally adjacent to the target video frame. Construct a video frame sequence with temporal context association. , where N is the preset number of historical frames, and the value is a positive integer; The video frame sequence is preprocessed to obtain the monitoring video sequence; The preprocessing operations include, in sequence: linearly normalizing pixel values ​​to the [0,1] interval, using Gaussian filtering to remove salt-and-pepper noise from the original video, and performing light deblurring on frames with slight motion blur.

[0024] This application provides an embodiment as follows, where N specifically ranges from 3 to 6, and can be flexibly adjusted according to the hardware computing power and video frame rate of the security monitoring terminal. N = 4 is preferred, balancing the integrity of temporal context information with real-time processing requirements. Finally, the constructed video frame sequence is preprocessed. The preprocessing operations are as follows: pixel values ​​are linearly normalized to the [0, 1] interval to eliminate the impact of pixel value magnitude differences on subsequent feature extraction; Gaussian filtering is used to remove a small amount of salt-and-pepper noise from the original video to avoid noise interference with feature decoupling; frames with slight motion blur are lightly deblurred to provide high-quality input data for subsequent spatiotemporal joint feature decoupling.

[0025] Step S2 includes: Pre-construct and train a two-branch spatiotemporal joint feature decoupling network; Based on the Retinex physical optics imaging model I = L ⊙ R, a dual-branch spatiotemporal joint feature decoupling network decouples the surveillance video sequence into two independent feature streams, including: a low-frequency illuminance feature sequence L and a high-frequency reflection feature sequence R; where I is the input video frame, L is the illuminance component, R is the reflection component, and ⊙ is the element-wise multiplication operator. The dual-branch spatiotemporal joint feature decoupling network specifically includes: The encoder adopts a weight-sharing Siamese network structure, consisting of 5 depthwise separable convolutional layers, 2 batch normalization layers, and 1 pooling layer. All weights are completely shared among the frames of the surveillance video sequence. Each video frame in the surveillance video sequence is independently forward-propagated to extract the corresponding spatiotemporal basic feature map. Decoder: It is divided into two independent parallel branches, namely the illumination branch decoder and the reflection branch decoder; The illumination branch decoder consists of 4 transposed convolutional layers and 2 batch normalization (BN) layers. It uses a 5×5 large convolutional kernel to capture the low-frequency distribution of global illumination. The output channel number is 1, corresponding to the low-frequency illumination features of a single frame. Finally, it outputs the low-frequency illumination feature sequence L of the monitoring video sequence. The reflection branch decoder consists of 4 transposed convolutional layers and 2 batch normalization (BN) layers. It uses 3×3 small convolutional kernels to capture high-frequency features of detailed textures. The number of output channels is the same as the number of color channels of the input video frame. Finally, it outputs the high-frequency reflection feature sequence R of the monitoring video sequence.

[0026] As one embodiment, the preprocessed video frame sequence with temporal context is input into a pre-built and trained dual-branch spatiotemporal joint feature decoupling network. This dual-branch spatiotemporal joint feature decoupling network is deployed on the edge computing node or embedded processor of the security monitoring terminal, adopting a lightweight network architecture to adapt to low-computing-power hardware environments. Then, based on the Retinex physical optics imaging model, the features of the input video frame sequence are decoupled into two independent feature streams through the dual-branch network.

[0027] As one embodiment, each element of the low-frequency illumination feature sequence L corresponds to the low-frequency illumination feature of a single frame, representing the spatial distribution and temporal variation of global ambient illumination in the video frame, and has the characteristics of low rank, low frequency, and temporal smoothness; each element of the high-frequency reflection feature sequence R corresponds to the high-frequency reflection feature of a single frame, representing the texture details, edge information, and dark random noise of the target in the monitoring scene, and has the characteristics of high frequency, high rank, and temporal correlation. The two are independent of each other and can be processed in a differentiated parallel manner.

[0028] Step S3 includes: Step S31: Photometric adaptive mapping processing of the low-frequency illuminance feature sequence to obtain the brightened target illuminance features, as detailed below: The luminance of the dark region is nonlinearly stretched using an iterative mapping equation. The network first predicts a set of learnable curve adjustment parameters. Then, adaptive brightness adjustment is achieved through higher-order iterative mapping; Step S32: Spatiotemporal alignment and dynamic denoising of the high-frequency reflection feature sequence to obtain pure target reflection features, specifically as follows: using the motion offset information between the historical reference frames of the monitoring video sequence and the target video frames, spatiotemporal feature alignment and three-dimensional dynamic filtering are performed to output pure target reflection features.

[0029] Step S31 includes: The iterative mapping equation for photometric adaptive mapping is as follows:

[0030] in It serves as the spatial index of a pixel in the feature map, representing the two-dimensional coordinate position of a pixel in the feature map; The number of iterations. Where N is the preset total number of iterations, which is a positive integer; For the first After the iteration, the illuminance value of pixel x remains within a certain range. The interval has no boundary violations; For the first Illuminance value after the next iteration, initial iteration value , that is, the original illumination feature value of the target frame; The curve adjustment parameter for the nth iteration is predicted and generated by the illumination branch network through a fully connected layer based on the context information of the current and historical frames. , The larger the value, the greater the brightening effect on dark areas. The smaller the value, the stronger the effect of suppressing excessive brightening in bright areas; The iterative mapping equation enables regional brightness adjustment: when When the curve exhibits non-linear stretching, it significantly enhances the brightness of dark areas; when At this time, the curve shows a compression trend to avoid overexposure in bright areas; when At that time, the curve adjusts linearly.

[0031] As one example, the equation is an S-shaped growth curve, which enables precise regional brightness adjustment: when (In dark areas), the curve exhibits non-linear stretching, significantly increasing the brightness of dark areas, with the stretching reaching 2 to 3 times; when (In bright areas), the curve tends to compress to avoid overexposure in bright areas, with a compression ratio of 0.5 to 0.8 times; when (In the mid-brightness region), the curve adjusts linearly to ensure a natural transition in image brightness. Simultaneously, the network adjusts the prediction parameters... At the same time, the mapping parameters of adjacent historical frames will be referenced, and the consistency of the mapping curves of adjacent frames will be ensured through timing constraints, thereby suppressing timing flickering from the root.

[0032] Step S32 includes: Motion offset prediction: Using a 3×3 deformable convolution operator, based on the high-frequency reflection feature sequence, motion offset prediction is performed on the high-frequency reflection feature sequence of each historical reference frame, and two offset feature maps corresponding to each reference frame are output to obtain the deformation field and motion offset matrix between adjacent frames. Temporal feature alignment: Based on the motion offset matrix, the high-frequency reflection features of all historical reference frames are aligned pixel by pixel to the coordinate space of the target video frame using a bilinear interpolation algorithm to construct a temporally aligned high-frequency reflection feature sequence; 3D dynamic filtering and noise reduction: The aligned high-frequency reflection feature sequence is jointly filtered by a 3×3×3 spatial-temporal three-dimensional Gaussian filter kernel. The spatial kernel size is 3×3 to suppress random noise within a single frame, and the temporal kernel size is 3 to suppress noise fluctuations between frames. Set edge-aware weights, reduce the filtering intensity for moving edge regions, and increase the filtering intensity for flat dark areas to output target reflection features.

[0033] As one example, for high-frequency reflection feature sequences, spatiotemporal feature alignment and three-dimensional dynamic filtering are performed using motion offset information between historical reference frames and target video frames. This suppresses dark random noise while maintaining the edge sharpness of the moving target, outputting clean and noise-free target reflection features. These target reflection features retain the texture details and edge information of the target, providing a high-quality detail foundation for subsequent video frame reconstruction. The specific processing flow is as follows: 1. Motion offset prediction: First, a 3×3 deformable convolution (DCN) operator is used. Based on the reflection feature Rt of the target video frame, motion offset prediction is performed on the reflection features of each historical reference frame. Two offset feature maps corresponding to each reference frame are output. The offset value range is [-3,3]. The deformation field and motion offset matrix between adjacent frames are obtained. This motion offset matrix records the pixel offset relationship between each reference frame and the target frame, which can accurately adapt to the dynamic offset of the moving target. 2. Temporal feature alignment: Based on the motion offset matrix, the reflection features of all historical reference frames are aligned pixel by pixel to the coordinate space of the target video frame using a bilinear interpolation algorithm, eliminating the inter-frame feature misalignment caused by the movement of the moving target and constructing a temporally aligned reflection feature sequence.

[0034] Step S4 includes: The spatiotemporal joint feature decoupling network employs a multi-constraint joint loss function for end-to-end training during the training phase. This multi-constraint joint loss function includes a temporal smoothing loss term, an image reconstruction loss term, and a noise suppression loss term. The total loss function is:

[0035] in, The total loss value is the result of multiple constraints, and the optimization objective of network training is the result of training. The smaller the value, the better the network training effect. The image reconstruction loss is defined as follows: ,in For enhanced video frames, These serve as high-quality reference frames to ensure the image quality reproduction of the enhanced video frames. This is the time-series smoothing loss; To mitigate noise loss, random noise is suppressed by constraining the spatial smoothness of the reflection features, while allowing edges to retain sharpness. These are the balancing hyperparameters for each loss, all of which are non-negative real numbers;

[0036] in The temporal smoothing loss term is used to quantify the degree of difference in illumination features between adjacent video frames. The smaller the value, the better the consistency of illumination between adjacent frames and the weaker the temporal flicker phenomenon; t is the frame index in the video sequence, which is a positive integer. The summation symbol indicates that the loss is calculated and accumulated for all adjacent frame pairs in the current training batch; The target illumination feature map for frame t is output by the illumination branch decoder of the dual-branch spatiotemporal joint feature decoupling network, with dimension . H is the height, W is the width, and each pixel value... This characterizes the ambient light intensity of the pixel. For the first The target illumination feature map of the frame, with dimensions and Completely identical; It is an L1 norm.

[0037] A nighttime surveillance video enhancement system based on spatiotemporal joint and photometric adaptive mapping, deployed in security monitoring terminals, edge computing nodes, or network video recorder equipment, includes the following functional modules: Video sequence construction module: Composed of video acquisition unit, preprocessing unit and sequence buffering unit, it is used to acquire the low-light raw video stream collected by night security equipment, and construct a monitoring video sequence with temporal context after preprocessing; Spatiotemporal joint feature decoupling module: Built-in pre-trained weight-shared Siamese dual-branch spatiotemporal joint feature decoupling network, which decouples the surveillance video sequence into low-frequency illumination feature sequence and high-frequency reflection feature sequence based on the Retinex physical optics model; The differentiated parallel processing module includes a photometric adaptive mapping unit and a spatiotemporal alignment denoising unit, which work in parallel. The photometric adaptive mapping unit uses an iterative exposure control equation to adjust the brightness of the low-frequency illuminance feature sequence and outputs the target illuminance feature. The spatiotemporal alignment denoising unit uses deformable convolution and three-dimensional dynamic filtering to perform spatiotemporal alignment and denoising on the high-frequency reflection feature sequence and outputs the target reflection feature. The reconstruction output module consists of a feature fusion unit, a post-processing unit, and an output unit. It performs element-wise multiplication fusion of target illumination features and target reflection features to reconstruct bright-state video frames, which are then output to display devices, storage devices, or back-end intelligent analysis platforms after post-processing. The video acquisition unit acquires video streams via ONVIF, RTSP, or GB / T 28181 national standard protocol; the buffer capacity of the sequence buffer unit is flexibly adjusted according to the number of historical frames N; the output unit outputs to the display device in real time via HDMI interface, writes to the storage device via SATA interface, and transmits to the back-end intelligent analysis platform via network interface.

[0038] Based on the application scenario of urban park security monitoring systems, this embodiment provides a complete implementation description of the technical solution of the present invention. The enhancement method and system are deployed in the network video recorder terminal and edge computing nodes of the park security monitoring system. Real-time enhancement processing is performed on the monitoring video stream in ultra-low light scenarios at night in the park. The core service is provided to the backend target detection and tracking, event recognition, and intelligent analysis system, enabling clear identification of pedestrians and vehicles within the park at night and facilitating security evidence collection. The specific implementation steps, equipment deployment, network training, and effect verification are as follows: Step 1: Constructing a time-series correlated surveillance video sequence: First, select and deploy video acquisition equipment: Hikvision DS-2CD3T25-I3 IPC cameras are used. These cameras support 2-megapixel, 1920×1080 resolution acquisition, support ONVIF protocol transmission, and have infrared fill light function (infrared fill light is turned off in this embodiment to simulate a pure low-light scene). Eight cameras are deployed in key areas such as entrances and exits, main roads, and parking lots in urban parks to form a full-coverage monitoring network.

[0039] Then, the video stream transmission configuration is completed: the raw low-light video stream captured by the IPC camera is transmitted to the network video recorder terminal of the park's security monitoring center via gigabit Ethernet. The video stream parameters are: resolution 1920×1080, frame rate 25fps, pixel depth 8bit, color format RGB three-channel, and transmission bit rate 4Mbps. Next, video preprocessing is performed: the built-in video sequence construction module of the network video recorder terminal preprocesses the raw video stream. First, the pixel values ​​are linearly normalized from [0,255] to the [0,1] range to eliminate differences in pixel value magnitude; a 5×5 Gaussian filter with a standard deviation of 1.0 is used to remove a small amount of salt-and-pepper noise in the raw video; for frames with slight motion blur, a Gaussian difference deblurring algorithm is used for light deblurring to improve frame quality. Finally, a time-series correlated video frame sequence is constructed: based on the current processing time node, the target video frame It to be enhanced is extracted, and four historical reference frames that are temporally adjacent to the target video frame are extracted simultaneously. That is, the preset number of historical frames N=4, and the final construction of a video frame sequence containing 5 frames with temporal context association. The sequence construction takes 5ms and is cached in the DDR4 8G memory of the network video recorder terminal to ensure real-time processing requirements. The reason for choosing N=4 is that this sequence length can cover the normal movement trajectories of pedestrians walking at a speed of 1-2m / s and vehicles traveling at a speed of 10-20km / h in security monitoring, ensuring sufficient temporal context information and effectively capturing inter-frame motion correlations; at the same time, the moderate sequence length can avoid the increase in computation and decrease in real-time performance caused by excessively long sequences, and can achieve the optimal balance between computing power and effect on embedded terminals.

[0040] Step 2: Decoupling of spatiotemporal joint features based on the physical optics model: First, the constructed video frame sequence is input into the pre-trained dual-branch spatiotemporal joint feature decoupling network, such as... Figure 2As shown, the network is deployed at the edge computing node (Huawei Atlas 200I DKA2, equipped with Ascend 310B processor, computing power 8 TOPS, memory 8G, storage 64GSSD) in the park's security monitoring center. Before input, the video frames are normalized, and the pixel values ​​are linearly mapped from [0,255] to the [0,1] range. Then, feature decoupling is completed through the weight-shared twin encoder-dual-branch decoder architecture of the dual-branch spatiotemporal joint feature decoupling network. The specific structure and forward propagation logic of the network in this embodiment are as follows: (1) Weight-shared twin encoder: It is composed of 5 layers of depthwise separable convolution with a stride of 2. The feature map resolution is reduced layer by layer from 1920×1080 to 60×34. All weights of the encoder are completely shared among the 5 video frames to avoid parameter redundancy, reduce the amount of computation, and adapt to the computing power of edge computing nodes. The 5 video frames in the sequence are propagated independently one by one. The corresponding spatiotemporal basic feature map is extracted for each frame. The output feature map size is 60×34×128, ensuring that the features of each frame are completely independent and there is no pre-fusion of time dimension. Specific parameters for each layer of the encoder: Layer 1 is a depthwise separable convolution with a kernel size of 7×7, stride of 2, 16 output channels, LeakyReLU activation function, and normalized BN layer; Layer 2 is a depthwise separable convolution with a kernel size of 5×5, stride of 2, 32 output channels, LeakyReLU activation function, and normalized BN layer; Layer 3 is a depthwise separable convolution with a kernel size of 3×3, stride of 2, 64 output channels, LeakyReLU activation function, and normalized BN layer; Layer 4 is a depthwise separable convolution with a kernel size of 3×3, stride of 1, 128 output channels, LeakyReLU activation function, and normalized BN layer; Layer 5 is a depthwise separable convolution with a kernel size of 3×3, stride of 1, 128 output channels, LeakyReLU activation function, and a pooling layer. (2) Illuminance branch decoder: It is composed of 4 layers of transposed convolution, and gradually upsamples to restore the feature map resolution to the input size of 1920×1080; the convolution kernel size is 5×5, which is used to capture the low frequency distribution of global illumination. The number of output channels is 64, 32, 16 and 1 respectively. The basic feature map of each frame is decoded independently, and finally the single-channel low frequency illumination feature sequence L of 5 frames is output. The activation function is Sigmoid to ensure that the output value is in the range of [0,1].The specific parameters of each layer of the decoder are as follows: Layer 1 is a transposed convolution with a kernel of 5×5, stride of 2, 64 output channels, LeakyReLU activation function, and normalized BN layer; Layer 2 is a transposed convolution with a kernel of 5×5, stride of 2, 32 output channels, LeakyReLU activation function, and normalized BN layer; Layer 3 is a transposed convolution with a kernel of 5×5, stride of 2, 16 output channels, LeakyReLU activation function, and normalized BN layer; Layer 4 is a transposed convolution with a kernel of 5×5, stride of 2, 1 output channel, and Sigmoid activation function. (3) Reflection branch decoder: It is composed of 4 layers of transposed convolution, and gradually upsamples to restore the feature map resolution to the input size of 1920×1080; the convolution kernel size is 3×3, which is used to capture the high-frequency features of the detailed texture. The number of output channels is 64, 32, 16 and 3 respectively. The basic feature map of each frame is decoded independently, and finally the three-channel high-frequency reflection feature sequence R of 5 frames is output. The activation function is Sigmoid. The specific parameters of each layer of the decoder are as follows: Layer 1 is a transposed convolution with a 3×3 kernel, a stride of 2, 64 output channels, and the activation function LeakyReLU, with normalized BN layers; Layer 2 is a transposed convolution with a 3×3 kernel, a stride of 2, 32 output channels, and the activation function LeakyReLU, with normalized BN layers; Layer 3 is a transposed convolution with a 3×3 kernel, a stride of 2, 16 output channels, and the activation function LeakyReLU, with normalized BN layers; Layer 4 is a transposed convolution with a 3×3 kernel, a stride of 2, 3 output channels, and the activation function Sigmoid.

[0041] Finally, feature decoupling was completed based on the Retinex physical optics imaging model. The input video frame achieved accurate separation of the illuminance component and the reflection component. The illuminance feature sequence only contains the global illumination information of the scene and has no detailed texture. The reflection feature sequence contains the inherent texture details and noise of the target and is not affected by changes in illumination, laying the foundation for subsequent differential processing.

[0042] Step 3: Parallel processing of differential illuminance and reflectance characteristic flows: For the decoupled low-frequency illuminance feature sequence and high-frequency reflection feature sequence, two independent processing units are used for parallel processing. Two independent processor cores are deployed on the edge computing node to ensure real-time processing. The process is divided into two sub-steps: Step S31: Photometric Adaptive Mapping Processing of Low-Frequency Illuminance Feature Flow like Figure 3 As shown, firstly, for the original illumination feature Lt(x) of the decoupled target frame, an iterative photometric adaptive mapping function is selected for brightness adjustment. In this embodiment, the number of iterations n=3, that is, the brightness mapping is completed through 3 iterations. The specific iteration process is as follows:

[0043] Where x is the pixel spatial location index, representing the two-dimensional coordinate position of the pixel in the feature map; This is the initial illuminance value. , which is the original illumination feature value of the target frame, output by the illumination branch decoder of the dual-branch spatiotemporal joint feature decoupling network, and its value range is . , representing the original ambient light intensity at pixel x; These are the illumination feature values ​​of pixel x after the 1st, 2nd, and 3rd iterations, respectively. The output value of each iteration is automatically maintained. Within the range; The optimal number of iterations can be adjusted to 2 to 4 times based on the actual brightness requirements of the scene. The parameters are adjusted for the curve predicted by the illuminance branch network through the fully connected layer, each Its physical meaning is to simulate the gain control in the exposure adjustment process of a camera. The larger the value, the stronger the brightening effect on dark areas in the current iteration. The smaller the value, the stronger the effect of suppressing overexposure in bright areas, and the better it can prevent overexposure in bright areas.

[0044] Step S32: Spatiotemporal alignment and dynamic denoising of high-frequency reflection feature stream: like Figure 4 As shown, the high-frequency reflection feature sequence obtained by decoupling In this embodiment, deformable convolution is used for spatiotemporal feature alignment, followed by 3D dynamic filtering for noise reduction. The specific process is as follows: (1) Motion offset prediction: First, a 3×3 deformable convolution operator is used, with the reflection feature Rt of the target frame It as the reference, and the reflection features of the four historical reference frames are respectively analyzed. ~ Motion offset prediction is performed, and two offset feature maps corresponding to each reference frame are output. The offset value range is [-3,3]. The deformation field and motion offset matrix between adjacent frames are obtained. This matrix records the pixel offset relationship between each reference frame and the target frame, which can accurately adapt to the dynamic movement of pedestrians and vehicles in the park. (2) Temporal feature alignment: Based on the predicted motion offset matrix, the reflection features of all historical reference frames are aligned pixel by pixel to the coordinate space of the target frame through bilinear interpolation algorithm to eliminate the inter-frame feature misalignment caused by the movement of the moving target and construct the temporally aligned reflection feature sequence:

[0045] in: The time-aligned reflection feature sequence is a three-channel feature sequence of 5 frames, with dimensions completely consistent with the original reflection feature sequence. These are historical reference frames. After motion offset compensation, the reflection features are aligned to the target frame coordinate space; The original reflection features of the target frame do not require alignment processing and can be directly included in the alignment sequence.

[0046] In this embodiment, the alignment error is controlled within 1 pixel to ensure that the spatial position of the same target is highly consistent between frames, providing accurate temporal correlation data for subsequent spatiotemporal joint denoising.

[0047] (3) Three-dimensional dynamic filtering and denoising: for the aligned reflection feature sequence A 3×3×3 spatial-temporal three-dimensional Gaussian filter kernel is used for joint filtering. The spatial kernel, with a size of 3×3 and a standard deviation of 0.8, is used to suppress random shot noise and thermal noise within a single frame. The temporal kernel, with a size of 3 and a standard deviation of 0.5, is used to smooth noise fluctuations between frames while preserving true temporal variations. To further protect the edge sharpness of moving targets, an edge-aware weighting mechanism is introduced: the Canny edge detection operator identifies moving edge regions in real time. Canny edge detection uses a high-to-low threshold ratio of 2:1, with the high threshold set to 80 and the low threshold set to 40. The edge detection results are used as a mask, with a filtering weight of 0.2 for edge regions within the mask and a filtering weight of 0.8 for flat dark areas outside the mask. The final weighted output filter result enhances noise suppression capabilities. The final output shows clean target reflection features. It effectively removes noise while fully preserving texture details and motion edges.

[0048] Step 4: Video Frame Reconstruction and System Output: First, the enhanced target illumination features With pure target reflection characteristics Element-wise multiplication fusion is performed to reconstruct the enhanced bright-state video frames. Then, post-processing is performed on the reconstructed video frames, changing the pixel values ​​from... Inverse normalization of intervals to The system uses the 8-bit standard video format and performs contrast fine-tuning. Finally, it completes multi-channel system output: 1. Real-time output via HDMI 2.0 interface to the display device in the security monitoring center (Hikvision DS-D5043U 43-inch monitor, resolution...). ,brightness 1. Provide security personnel with real-time monitoring of the park's nighttime situation; 2. Write the enhanced video stream to the Western Digital Purple disk of the network video recorder terminal via the SATA3.0 interface, with a storage format of H.265 and a storage duration of 30 days, for subsequent security incident evidence collection and intelligent analysis processing; 3. Transmit the enhanced video stream to the backend Hikvision IVMS-8700 intelligent analysis platform via Gigabit Ethernet to provide high-quality input data for algorithms such as target detection, multi-target tracking, and behavioral event recognition.

[0049] In this embodiment, the processing time for a single frame from video acquisition to output is 38ms, which is less than 40ms. It can achieve full real-time processing at 25fps, fully meeting the real-time requirements of security monitoring, without any stuttering or delay.

[0050] Training a two-branch spatiotemporal joint feature decoupling network: The dual-branch spatiotemporal joint feature decoupling network in this embodiment adopts an end-to-end supervised training method. The training platform is deployed on a laboratory server. The specific training configuration, process, and loss function are as follows: 1. Training platform configuration: The server uses an Intel Xeon E5-2690v4 processor, 64GB of memory, an NVIDIA RTX 4090 graphics card (24GB of video memory), a 2TB SSD hard drive, an Ubuntu 20.04 LTS operating system, PyTorch 1.13.1 deep learning framework, CUDA version 11.7, and CUDNN version 8.5 to ensure training efficiency and stability.

[0051] 2. Training Dataset and Data Augmentation: First, the training datasets were selected, using the publicly available low-light video datasets ExDark and LOL-v2, as well as a self-collected urban park nighttime security surveillance video dataset. Then, data augmentation was performed on the datasets, including random cropping, random flipping, random rotation, and random brightness perturbation, to improve the network's generalization ability.

[0052] 3. Core training parameters: All parameters have been adjusted multiple times to ensure the optimal balance between training effect and network performance. The specific parameter settings are shown in Table 1.

[0053] Table 1: Core Parameter Settings for Network Training

[0054] 4. Loss Function: End-to-end training is performed using a multi-constraint joint loss function. The total loss formula is as follows:

[0055] in , where is the total loss value of multiple constraints during the training phase, and is the optimization objective of network training. The smaller the value, the better the network training effect. The balancing hyperparameter for each loss is fixed at in this embodiment. ; Image reconstruction loss is used to measure the pixel-level difference between the enhanced video frame and the corresponding high-quality reference frame, ensuring image quality restoration. Temporal smoothing loss is used to constrain the consistency of illumination features between adjacent frames, thus solving the temporal flicker problem from the training layer. This is a noise suppression loss used to suppress random noise in reflection features while preserving target edge details.

[0056] The specific definitions and calculation formulas for each type of loss are as follows: (1) Reconstruction Loss: Measures the pixel-level difference between the enhanced video frame and the corresponding high-quality reference frame to ensure image quality restoration. The calculation formula is:

[0057] Where H, W, and C are the height, width, and number of color channels of the video frame, respectively; To enhance the pixel position of the video frame The pixel value of the c-th channel; For the corresponding standard normal lighting frame (tag frame) at the pixel position The pixel value of the c-th channel; This is for absolute value operations; The normalization coefficient is used to eliminate the influence of video frame size on the loss value. This loss enhances the difference between the result and the true value through L1 norm constraint, which can effectively avoid gradient explosion and improve training stability.

[0058] (2) Temporal smoothing loss: Constrains the consistency of illumination features between adjacent frames, thoroughly solving the temporal flicker problem from the training layer. The calculation formula is as follows:

[0059] Where T is the total number of frames in the training video sequence, which is a positive integer; t is the frame index in the video sequence, with a value ranging from 1 to T. This is the illumination feature map for frame t. For the first The illumination feature maps of the frames, both with dimensions of 1. ; for The norm is used to calculate the sum of the absolute differences of all corresponding pixels between two illuminance feature maps; the summation is used to accumulate the loss values ​​of all adjacent frame pairs in the training sequence. By minimizing this loss, the illuminance components of adjacent frames are forced to maintain a smooth transition, thus suppressing the brightness jump between frames from the root.

[0060] (3) Noise Suppression Loss: Suppresses random noise in reflection features while preserving target edge details. The calculation formula is as follows:

[0061] in The reflection feature map of the target frame, with dimension . ; For the two-dimensional coordinate index of the pixel, , This formula represents the total variation (TV) loss of the reflection feature. By calculating and summing the gradient magnitudes between adjacent pixels, it constrains the spatial smoothness of the reflection feature, forces noise reduction in flat areas, and allows the original sharpness to be preserved in edge areas, thus avoiding image blurring caused by noise reduction.

[0062] 5. Training Process: First, network initialization is performed using a He normal distribution to initialize network weights. The initial learning rate is set to 1e-4, the batch size to 8, and the total training epochs to 100. Then, forward propagation is performed, inputting the video frame sequences from the training dataset into the network to sequentially complete feature decoupling, differential parallel processing, and video frame reconstruction, outputting the enhanced frames. Next, loss calculation is performed, calculating the total loss between the enhanced frames and the label frames based on the multi-constraint joint loss function, and updating the network weights through backpropagation. Subsequently, the learning rate is adjusted, decreasing by 0.5 times every 20 training epochs to ensure fine-tuning of network parameters in the later stages of training and achieve convergence. Simultaneously, model saving is performed, saving the model weights every 10 epochs. After training, the model with the highest PSNR and lowest TBD on the validation set is selected as the final training model and saved in .onnx format for easy deployment to embedded edge computing nodes. Finally, convergence is judged. When the total loss on the validation set decreases by less than 1e-5 for 10 consecutive epochs, the network training is considered converged, and training is stopped. In this embodiment, the final convergence loss value of the trained network is 0.02.

[0063] 6. Model Deployment Optimization: The trained network model is quantized and optimized using TensorRT to reduce model size, improve inference speed, and adapt to the low computing power environment of edge computing nodes; at the same time, the model is pruned to remove redundant convolutional layers, further reducing the amount of computation.

[0064] To verify the technical effects of the present invention, comparative and ablation experiments were conducted in this embodiment. All experiments were carried out under the same dataset and hardware environment to ensure the fairness and reproducibility of the results. At the same time, scenario adaptability tests were added to verify the enhancement effect of the present invention in different nighttime monitoring scenarios, further highlighting the generalization ability of the present invention.

[0065] The experiment uses 5 core evaluation indicators to comprehensively cover image quality, temporal stability, noise suppression capability, and real-time implementation of engineering. The specific definitions and evaluation criteria of each indicator are as follows: (1) Peak signal-to-noise ratio (PSNR): measures the pixel-level similarity between the enhanced image and the standard normal illumination image, in dB. The higher the value, the better the image quality restoration. (2) Structural similarity (SSIM): measures the degree of preservation of structural information and texture details of the image, with a value range of The higher the value, the more complete the structure and details are preserved; (3) Timing brightness fluctuation standard deviation (TBD): measures the degree of brightness fluctuation between adjacent frames, the unit is pixel gray value. The lower the value, the better the timing stability and the weaker the flicker phenomenon. TBD1.0 can completely eliminate the flicker that can be perceived by the human eye; (4) Noise suppression ratio (NRR): measures the algorithm's ability to suppress dark random noise. The higher the value, the better the denoising effect and avoid noise interference with target recognition; (5) Single frame inference time: measures the algorithm's real-time processing capability, the unit is ms. The lower the value, the stronger the real-time performance and the more suitable it is for embedded security terminal deployment.

[0066] The comparative experiment selected the current mainstream low-light enhancement algorithms, including three widely used single-frame image enhancement algorithms and three mainstream video enhancement algorithms, and compared them with the method of this invention. All algorithms were deployed on the same edge computing node and input the same low-light video stream. The experimental results are shown in Table 2.

[0067] Table 2: Performance Comparison Experiment Results of Different Algorithms

[0068] Note: ↑ indicates that the higher the indicator, the better; ↓ indicates that the lower the indicator, the better. Peak signal-to-noise ratio, For structural similarity, The standard deviation of time-series brightness fluctuation, This represents the noise suppression ratio.

[0069] The comparative experimental results show that the method of this invention is significantly better than the existing mainstream algorithms in all evaluation indicators.

[0070] To verify the independent technical contributions of each core innovative module of the present invention, an ablation experiment was conducted in this embodiment. Based on the traditional single-frame Retinex enhancement settings, the core modules of the present invention were added one by one. The experimental conditions were the same as those in the comparative experiment, and the experimental results are shown in Table 3.

[0071] Table 3: Results of Ablation Experiments on Core Modules

[0072] Note: √ indicates that the module is included, × indicates that it is not included; PSNR is peak signal-to-noise ratio (in dB), TBD is the standard deviation of timing brightness fluctuation; ↑ indicates that the higher the indicator, the better, ↓ indicates that the lower the indicator, the better.

[0073] The results of the ablation experiment show that each core module of the present invention brings significant positive gains to image quality improvement and temporal stability optimization.

[0074] To verify the scenario adaptability of the present invention, four typical nighttime monitoring scenarios were selected and field tests were conducted. The test equipment was the same as that in this embodiment, and the test results are shown in Table 4.

[0075] Table 4: Results of Scene Adaptability Tests in Different Scenarios

[0076] Test results show that the method of the present invention can be adapted to various low-light monitoring scenarios at night, and can maintain excellent enhancement effects under different lighting conditions and different scene complexities.

[0077] To further and intuitively verify the actual technical effectiveness and backend business empowerment value of the algorithm of this invention in high-risk industrial security scenarios, a comparative test was conducted using a real nighttime construction scenario. Specific results are as follows: Figure 5 As shown.

[0078] from Figure 5 The comparison results clearly show that the enhanced image effect of the invention embodiment and the existing mainstream low-light enhancement algorithm in real night monitoring scenarios is shown in the comparison chart from left to right: original video frame at low light at night, enhancement effect of LLFlow algorithm, enhancement effect of Zero-DCE algorithm, and enhancement effect of the spatiotemporal joint decoupling enhancement algorithm of the present invention. This is used to intuitively demonstrate the technical advantages of the present invention in brightening dark areas, preventing overexposure of highlights, preserving motion edges, and restoring details and textures.

[0079] In one embodiment, Figure 2 This is a schematic diagram of the weight-shared twin dual-branch spatiotemporal joint feature decoupling network described in this invention, used to illustrate the encoder and dual-branch decoder architecture of the network; Figure 3 This is a schematic diagram comparing the curves of the photometric adaptive mapping function described in this invention, used to demonstrate the brightness mapping effect under different iteration numbers; Figure 4 This is a schematic diagram of the spatiotemporal feature alignment and three-dimensional dynamic filtering based on deformable convolution described in this invention, used to illustrate the denoising processing logic of the reflection feature stream; Figure 5 This is a comparison chart showing the enhanced image effect of the embodiment of the present invention and existing low-light enhancement algorithms in real-world nighttime surveillance scenarios.

[0080] This application also discloses an electronic device. (See reference...) Figure 6 , Figure 6 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. The electronic device 500 may include: at least one processor 501, at least one network interface 504, a user interface 503, a memory 505, and at least one communication bus 502.

[0081] The communication bus 502 is used to enable communication between these components.

[0082] The user interface 503 may include a display screen, and optionally, the user interface 503 may also include a standard wired interface or a wireless interface.

[0083] The network interface 504 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0084] This application also discloses a computer-readable storage medium storing multiple instructions adapted for loading by a processor to execute the above-described method for enhancing nighttime surveillance video based on spatiotemporal joint and photometric adaptive mapping.

[0085] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure.

[0086] This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

Claims

1. A night monitoring video enhancement method based on space-time joint and photometric adaptive mapping, characterized in that, The method includes the following steps: S1: Based on the acquired low-light raw video stream, construct a time-series correlated monitoring video sequence; S2: Input the surveillance video sequence into a spatiotemporal joint feature decoupling network based on a physical optics model to obtain a low-frequency illumination feature sequence and a high-frequency reflection feature sequence. S3: The decoupled low-frequency illuminance feature sequence and high-frequency reflection feature sequence are processed in parallel using differentiated processing strategies based on their physical characteristics to obtain target illuminance features and target reflection features. S4: Element-wise multiplication is used to fuse the target illumination features and target reflection features, and the bright-state video frames with high dynamic range, high signal-to-noise ratio and smooth temporal sequence are reconstructed based on the Retinex imaging model to enhance the nighttime surveillance video.

2. The night monitoring video enhancement method based on space-time joint and photometric adaptive mapping according to claim 1, characterized in that, Step S1 includes: Acquire raw low-light video streams collected by nighttime security monitoring equipment; Based on the current processing time point, extract the target video frames to be enhanced from the original low-light video stream. Simultaneously extract N historical reference frames that are temporally adjacent to the target video frame. Construct a video frame sequence with temporal context association. , where N is the preset number of historical frames, and the value is a positive integer; The video frame sequence is preprocessed to obtain the monitoring video sequence; The preprocessing operations include, in sequence: linearly normalizing pixel values ​​to the [0,1] interval, using Gaussian filtering to remove salt-and-pepper noise from the original video, and performing light deblurring on frames with slight motion blur.

3. The nighttime surveillance video enhancement method based on spatiotemporal joint and photometric adaptive mapping as described in claim 1, characterized in that, Step S2 includes: Pre-construct and train a two-branch spatiotemporal joint feature decoupling network; Based on the Retinex physical optics imaging model I = L ⊙ R, a dual-branch spatiotemporal joint feature decoupling network decouples the surveillance video sequence into two independent feature streams, including: a low-frequency illuminance feature sequence L and a high-frequency reflection feature sequence R; where I is the input video frame, L is the illuminance component, R is the reflection component, and ⊙ is the element-wise multiplication operator. The dual-branch spatiotemporal joint feature decoupling network specifically includes: The encoder adopts a weight-sharing Siamese network structure, consisting of 5 depthwise separable convolutional layers, 2 batch normalization layers, and 1 pooling layer. All weights are completely shared among the frames of the surveillance video sequence. Each video frame in the surveillance video sequence is independently forward-propagated to extract the corresponding spatiotemporal basic feature map. Decoder: It is divided into two independent parallel branches, namely the illumination branch decoder and the reflection branch decoder; The illumination branch decoder consists of 4 transposed convolutional layers and 2 batch normalization (BN) layers. It uses a 5×5 large convolutional kernel to capture the low-frequency distribution of global illumination. The output channel number is 1, corresponding to the low-frequency illumination features of a single frame. Finally, it outputs the low-frequency illumination feature sequence L of the monitoring video sequence. The reflection branch decoder consists of 4 transposed convolutional layers and 2 batch normalization (BN) layers. It uses 3×3 small convolutional kernels to capture high-frequency features of detailed textures. The number of output channels is the same as the number of color channels of the input video frame. Finally, it outputs the high-frequency reflection feature sequence R of the monitoring video sequence.

4. The nighttime surveillance video enhancement method based on spatiotemporal joint and photometric adaptive mapping as described in claim 1, characterized in that, Step S3 includes: Step S31: Photometric adaptive mapping processing of the low-frequency illuminance feature sequence to obtain the brightened target illuminance features, as detailed below: The luminance of the dark region is nonlinearly stretched using an iterative mapping equation. The network first predicts a set of learnable curve adjustment parameters. Then, adaptive brightness adjustment is achieved through higher-order iterative mapping; Step S32: Spatiotemporal alignment and dynamic denoising of the high-frequency reflection feature sequence to obtain pure target reflection features, specifically as follows: using the motion offset information between the historical reference frames of the monitoring video sequence and the target video frames, spatiotemporal feature alignment and three-dimensional dynamic filtering are performed to output pure target reflection features.

5. The nighttime surveillance video enhancement method based on spatiotemporal joint and photometric adaptive mapping as described in claim 4, characterized in that, Step S31 includes: The iterative mapping equation for photometric adaptive mapping is as follows: in It serves as the spatial index of a pixel in the feature map, representing the two-dimensional coordinate position of a pixel in the feature map; The number of iterations. Where N is the preset total number of iterations, which is a positive integer; For the first After the iteration, the illuminance value of pixel x remains within a certain range. The interval has no boundary violations; For the first Illuminance value after the next iteration, initial iteration value , that is, the original illumination feature value of the target frame; The curve adjustment parameter for the nth iteration is predicted and generated by the illumination branch network through a fully connected layer based on the context information of the current and historical frames. , The larger the value, the greater the brightening effect on dark areas. The smaller the value, the stronger the effect of suppressing excessive brightening in bright areas; The iterative mapping equation enables regional brightness adjustment: when When the curve exhibits non-linear stretching, it significantly enhances the brightness of dark areas; when At this time, the curve shows a compression trend to avoid overexposure in bright areas; when At that time, the curve adjusts linearly.

6. The nighttime surveillance video enhancement method based on spatiotemporal joint and photometric adaptive mapping as described in claim 4, characterized in that, Step S32 includes: Motion offset prediction: Using a 3×3 deformable convolution operator, based on the high-frequency reflection feature sequence, motion offset prediction is performed on the high-frequency reflection feature sequence of each historical reference frame, and two offset feature maps corresponding to each reference frame are output to obtain the deformation field and motion offset matrix between adjacent frames. Temporal feature alignment: Based on the motion offset matrix, the high-frequency reflection features of all historical reference frames are aligned pixel by pixel to the coordinate space of the target video frame using a bilinear interpolation algorithm to construct a temporally aligned high-frequency reflection feature sequence; 3D dynamic filtering and noise reduction: The aligned high-frequency reflection feature sequence is jointly filtered by a 3×3×3 spatial-temporal three-dimensional Gaussian filter kernel. The spatial kernel size is 3×3 to suppress random noise within a single frame, and the temporal kernel size is 3 to suppress noise fluctuations between frames. Set edge-aware weights, reduce the filtering intensity for moving edge regions, and increase the filtering intensity for flat dark areas to output target reflection features.

7. The nighttime surveillance video enhancement method based on spatiotemporal joint and photometric adaptive mapping as described in claim 1, characterized in that, Step S4 includes: The spatiotemporal joint feature decoupling network employs a multi-constraint joint loss function for end-to-end training during the training phase. This multi-constraint joint loss function includes a temporal smoothing loss term, an image reconstruction loss term, and a noise suppression loss term. The total loss function is: in, The total loss value is the result of multiple constraints, and the optimization objective of network training is the result of training. The smaller the value, the better the network training effect. The image reconstruction loss is defined as follows: ,in For enhanced video frames, These serve as high-quality reference frames to ensure the image quality reproduction of the enhanced video frames. This is the time-series smoothing loss; To mitigate noise loss, random noise is suppressed by constraining the spatial smoothness of the reflection features, while allowing edges to retain sharpness. These are the balancing hyperparameters for each loss, all of which are non-negative real numbers; in The temporal smoothing loss term is used to quantify the degree of difference in illumination features between adjacent video frames. The smaller the value, the better the consistency of illumination between adjacent frames and the weaker the temporal flicker phenomenon; t is the frame index in the video sequence, which is a positive integer. The summation symbol indicates that the loss is calculated and accumulated for all adjacent frame pairs in the current training batch; The target illumination feature map for frame t is output by the illumination branch decoder of the dual-branch spatiotemporal joint feature decoupling network, with dimension . H is the height, W is the width, and each pixel value... This characterizes the ambient light intensity of the pixel. For the first The target illumination feature map of the frame, with dimensions and Completely identical; It is an L1 norm.

8. A nighttime surveillance video enhancement system based on spatiotemporal joint and photometric adaptive mapping, used to implement the nighttime surveillance video enhancement method based on spatiotemporal joint and photometric adaptive mapping as described in any one of claims 1-7, characterized in that, Deployed in security monitoring terminals, edge computing nodes, or network video recorder equipment, it includes the following functional modules: Video sequence construction module: Composed of video acquisition unit, preprocessing unit and sequence buffering unit, it is used to acquire the low-light raw video stream collected by night security equipment, and construct a monitoring video sequence with temporal context after preprocessing; Spatiotemporal joint feature decoupling module: Built-in pre-trained weight-shared Siamese dual-branch spatiotemporal joint feature decoupling network, which decouples the surveillance video sequence into low-frequency illumination feature sequence and high-frequency reflection feature sequence based on the Retinex physical optics model; The differentiated parallel processing module includes a photometric adaptive mapping unit and a spatiotemporal alignment denoising unit, which work in parallel. The photometric adaptive mapping unit uses an iterative exposure control equation to adjust the brightness of the low-frequency illuminance feature sequence and outputs the target illuminance feature. The spatiotemporal alignment denoising unit uses deformable convolution and three-dimensional dynamic filtering to perform spatiotemporal alignment and denoising on the high-frequency reflection feature sequence and outputs the target reflection feature. The reconstruction output module consists of a feature fusion unit, a post-processing unit, and an output unit. It performs element-wise multiplication fusion of target illumination features and target reflection features to reconstruct bright-state video frames, which are then output to display devices, storage devices, or back-end intelligent analysis platforms after post-processing. The video acquisition unit acquires video streams via ONVIF, RTSP, or GB / T 28181 national standard protocol; the buffer capacity of the sequence buffer unit is flexibly adjusted according to the number of historical frames N; the output unit outputs to the display device in real time via HDMI interface, writes to the storage device via SATA interface, and transmits to the back-end intelligent analysis platform via network interface.

9. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to enable the electronic device to perform the nighttime surveillance video enhancement method based on spatiotemporal joint and photometric adaptive mapping as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed by a computer, perform the nighttime surveillance video enhancement method based on spatiotemporal joint and photometric adaptive mapping as described in any one of claims 1-7.