Security and protection monitoring method, system and equipment and storage medium
By constructing a three-dimensional illumination feature vector and a parameterized regression subnetwork, the problem of unstable image quality of the security monitoring system in complex lighting environments is solved, and the continuity of image processing and the improvement of key target recognition under extreme lighting conditions are achieved.
Patent Information
- Application Number
- CN202510877158.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-21
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing security monitoring systems have difficulty maintaining image quality stability and clear identifiability of key targets when faced with complex and changing lighting environments, especially under extreme lighting conditions such as dim light, backlight, and strong light.
By constructing a three-dimensional illumination feature vector and combining it with a parameterized regression subnetwork for nonlinear mapping calculation and image processing, an intelligent switching system for low-light enhancement, backlight compensation, strong light suppression and standard processing modes is established. Combined with boundary smoothing transition and historical memory mechanism, dynamic learning and optimization of ISP parameters are performed to ensure the continuity and stability of image processing.
It achieves continuity and stability in image processing under different lighting conditions, improves image quality and recognition of key targets, and meets the strict quality requirements of security monitoring.
Smart Images

Figure CN120823554A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of security monitoring technology, and in particular to a security monitoring method, system, equipment and storage medium. Background Art
[0002] Existing ISP processing methods are usually based on preset parameter configurations and lack the ability to perceive and adaptively adjust changes in lighting conditions in real time. This leads to poor processing effects under extreme lighting conditions such as dim light, backlight, and strong light. In particular, although certain effects can be achieved when processing single lighting scenes, they are often unable to cope with complex environments with frequently changing lighting conditions.
[0003] As a crucial security measure, security surveillance systems have extremely stringent requirements for image quality. They must not only maintain stable imaging in a variety of complex lighting environments, but also ensure the clear and identifiable identification of key objects, such as faces and license plates. However, security surveillance scenarios often face the challenges of complex and changing lighting conditions, including differences in indoor and outdoor lighting, day-night shifts, backlighting, and interference from various artificial light sources. These complex lighting environments place higher demands on traditional ISP processing technology. Furthermore, security surveillance systems typically process large amounts of high-resolution image data, requiring processing algorithms to not only guarantee image quality but also exhibit excellent real-time performance and stability. Summary of the Invention
[0004] The present invention provides a security monitoring method, system, device and storage medium. Compared with the traditional single lighting parameter detection method, the present invention can comprehensively and accurately describe the multi-dimensional characteristics of complex lighting environments, ensuring that the output image continues to meet the strict requirements of security monitoring quality standards.
[0005] In a first aspect, the present invention provides a security monitoring method, the security monitoring method comprising: Perform illumination detection on the original security monitoring image to obtain a three-dimensional illumination feature vector; According to the three-dimensional illumination feature vector, a matching determination is performed on preset low-light enhancement mode determination conditions, backlight compensation mode determination conditions, strong light suppression mode determination conditions, and standard processing mode determination conditions to obtain the ISP processing mode type corresponding to the current frame; Inputting the three-dimensional illumination feature vector and the ISP processing mode type into a parameterized regression subnetwork for nonlinear mapping calculation and image processing to obtain a first enhanced image; performing penalty calculation on the first enhanced image to obtain a total penalty value, and performing gradient reverse adjustment based on the total penalty value to obtain a second enhanced image; Performing quality inspection on the second enhanced image to obtain a target enhanced image that meets security monitoring quality standards.
[0006] In combination with the first aspect, in a first implementation of the first aspect of the present invention, performing illumination detection on the original security monitoring image to obtain a three-dimensional illumination feature vector includes: The original security monitoring image is segmented into pixel blocks to obtain multiple pixel blocks, and the mean and variance of the grayscale values in each pixel block are calculated to obtain the original light intensity value; Extracting white area pixels according to the original security monitoring image, and performing ratio calculation and standard color temperature curve interpolation operation on the red, green and blue three-channel values of the white area pixels to obtain an original color temperature deviation value; Performing differential gradient amplitude calculation and edge strength statistical analysis based on the original security monitoring image to obtain differential gradient amplitude distribution data, and performing a maximum value to average value ratio operation on the differential gradient amplitude distribution data to obtain an original contrast value; The original illumination intensity value, the original color temperature deviation value, and the original contrast value are normalized and vector-combined to obtain a three-dimensional illumination feature vector including the normalized illumination intensity value, the normalized color temperature deviation value, and the normalized contrast value.
[0007] In combination with the first aspect, in a second implementation of the first aspect of the present invention, matching and determining preset low-light enhancement mode determination conditions, backlight compensation mode determination conditions, strong light suppression mode determination conditions, and standard processing mode determination conditions based on the three-dimensional illumination feature vector to obtain the ISP processing mode type corresponding to the current frame includes: Comparing the normalized light intensity value in the three-dimensional light feature vector with a first preset threshold to obtain a light intensity interval classification result, and comparing the normalized contrast value with a second preset threshold to obtain a contrast level classification result; According to the illumination intensity interval classification result and the contrast level classification result, the preset low light enhancement mode determination condition, backlight compensation mode determination condition, strong light suppression mode determination condition and standard processing mode determination condition are matched and calculated one by one to obtain the matching degree value of each mode and the initial mode determination result; Performing boundary condition detection based on the initial mode determination result, and when it is detected that the three-dimensional illumination feature vector is in the mode boundary area, performing weighted fusion calculation on the matching degree values of two adjacent modes and generating a smooth transition weight coefficient to obtain a boundary fusion processing result; The boundary fusion processing result of the current frame is compared with the historical mode selection records of the previous N frames to obtain the difference degree. When the difference degree exceeds the third preset threshold, the mode switching amplitude is adjusted to obtain the ISP processing mode type corresponding to the current frame.
[0008] In combination with the first aspect, in a third implementation of the first aspect of the present invention, inputting the three-dimensional illumination feature vector and the ISP processing mode type into a parameterized regression subnetwork for nonlinear mapping calculation and image processing to obtain a first enhanced image includes: Matching the corresponding over-parameterized regression subnetwork based on the ISP processing mode type; Inputting the three-dimensional illumination feature vector into the regional importance analysis module of the parameterized regression subnetwork to perform multi-scale convolution feature extraction to obtain regional adaptive feature data; Inputting the regional adaptive feature data into the time series memory module of the parameterized regression subnetwork to perform inter-frame illumination change trend analysis to generate a time series enhanced feature vector; Inputting the temporal enhancement feature vector into the multi-branch parameter generation network of the parameterized regression subnetwork for parallel nonlinear mapping to obtain an ISP parameter set; Image processing is performed on the security monitoring original image based on the ISP parameter set to obtain a first enhanced image.
[0009] In combination with the first aspect, in a fourth implementation of the first aspect of the present invention, performing image processing on the security monitoring original image based on the ISP parameter set to obtain a first enhanced image includes: Perform bad pixel detection and neighborhood median filtering correction on the security monitoring original image according to the detection threshold parameter in the ISP parameter set to obtain RGB three-channel basic image data; Performing a linear transformation of pixel brightness on the RGB three-channel basic image data according to the exposure time and gain value in the ISP parameter set to obtain intermediate image data; Apply the color correction matrix in the ISP parameter set to perform matrix multiplication transformation on the intermediate image data to obtain pre-enhanced image data, and perform USM sharpening based on the pre-enhanced image data to obtain a first enhanced image.
[0010] In combination with the first aspect, in a fifth implementation of the first aspect of the present invention, performing penalty calculation on the first enhanced image to obtain a total penalty value, and performing gradient reverse adjustment based on the total penalty value to obtain the second enhanced image includes: Performing brightness histogram statistical analysis and bit-by-bit absolute value accumulation of the difference between the first enhanced image and the reference standard image to obtain a brightness consistency penalty item value; performing color saturation over-enhancement detection and hue shift angle calculation based on the RGB three-channel data of the first enhanced image to obtain a color naturalness penalty item value; and performing gradient amplitude change statistics and detail loss degree quantitative analysis on the first enhanced image to obtain a detail preservation penalty item value; Performing a weighted summation of the brightness consistency penalty item value, the color naturalness penalty item value, and the detail preservation penalty item value to obtain a total penalty value, and comparing the total penalty value with a preset penalty threshold value. When the total penalty value exceeds the preset penalty threshold value, a parameter callback processing mechanism is triggered to obtain a parameter adjustment trigger signal; Based on the parameter adjustment trigger signal, a gradient descent algorithm is started to perform reverse optimization calculation on the ISP parameter set to obtain an optimized ISP parameter set; The optimized ISP parameter set is reapplied to the security monitoring original image to perform ISP pipeline reprocessing calculation to obtain a second enhanced image.
[0011] In combination with the first aspect, in a sixth implementation of the first aspect of the present invention, performing quality detection on the second enhanced image to obtain a target enhanced image that meets security monitoring quality standards includes: Performing peak signal-to-noise ratio calculation and maximum-to-minimum brightness ratio calculation on the second enhanced image to obtain a dynamic range evaluation value. Simultaneously, based on Lab color space conversion and color difference formula calculation, a color fidelity evaluation value is obtained. Furthermore, a detail clarity evaluation value is obtained through modulation transfer function calculation and edge sharpness index statistical analysis. Furthermore, a visual quality evaluation value is obtained by using a structural similarity index and a multi-scale SSIM algorithm. Comparing the dynamic range evaluation value with a fourth preset threshold, comparing the color fidelity evaluation value with a fifth preset threshold, comparing the detail clarity evaluation value with a sixth preset threshold, and comparing the visual quality evaluation value with a seventh preset threshold, to obtain a four-dimensional quality compliance status determination result; Based on the four-dimensional quality compliance status determination result, the second enhanced image is format converted and encoded for output to obtain a target enhanced image that meets the security monitoring quality standard.
[0012] In a second aspect, the present invention provides a security monitoring system, comprising: Light detection module, used to perform light detection on the original security monitoring image and obtain a three-dimensional light feature vector; a matching determination module, configured to perform matching determination on preset low-light enhancement mode determination conditions, backlight compensation mode determination conditions, strong light suppression mode determination conditions, and standard processing mode determination conditions according to the three-dimensional illumination feature vector, to obtain the ISP processing mode type corresponding to the current frame; An image processing module, configured to input the three-dimensional illumination feature vector and the ISP processing mode type into a parameterized regression subnetwork for performing nonlinear mapping calculation and image processing to obtain a first enhanced image; a penalty calculation module, configured to perform penalty calculation on the first enhanced image to obtain a total penalty value, and perform gradient reverse adjustment based on the total penalty value to obtain a second enhanced image; The quality detection module is used to perform quality detection on the second enhanced image to obtain a target enhanced image that meets the security monitoring quality standard.
[0013] A third aspect of the present invention provides a security monitoring device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the security monitoring device executes the above-mentioned security monitoring method.
[0014] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, enable the computer to execute the above-mentioned security monitoring method.
[0015] The technical solution provided by the present invention constructs a three-dimensional illumination feature vector by simultaneously detecting light intensity, color temperature, and contrast. Compared with traditional single-light parameter detection methods, this method can comprehensively and accurately describe the multidimensional characteristics of complex lighting environments. It also establishes an intelligent switching system for four modes: low-light enhancement, backlight compensation, strong light suppression, and standard processing. Combined with boundary smoothing and historical memory mechanisms, it can automatically select the optimal processing strategy based on real-time lighting conditions, avoiding the limitations of traditional fixed-parameter processing methods and ensuring processing continuity and stability under different lighting conditions. A regional importance analysis module performs differentiated processing on facial regions, license plate regions, background regions, and boundary regions. Combined with a time series memory module and a multi-branch parameter generation network, it can be specifically optimized for the core task requirements of security monitoring. An over-parameterized regression subnetwork is used for dynamic learning of ISP parameters. By constructing a complex nonlinear mapping relationship, it can automatically generate the optimal ISP parameter set based on lighting conditions and scene characteristics, breaking through the limitations of fixed parameter settings in traditional ISP processing and achieving true parameter adaptive adjustment. By establishing triple penalty constraints for brightness consistency, color naturalness, and detail preservation, and optimizing ISP parameters in real time through a gradient inverse adjustment mechanism, the system effectively prevents over-enhancement and color shift during image processing, ensuring that the resulting image quality is enhanced while maintaining a natural and realistic visual experience. Within the complete ISP processing flow—defective pixel correction, demosaicing, exposure control, white balance correction, color correction, gamma correction, and sharpening enhancement—each step dynamically fine-tunes parameters based on the current frame's processing results, achieving more refined and adaptive image enhancement than traditional fixed pipeline processing. A four-dimensional evaluation system encompassing dynamic range, color fidelity, detail clarity, and visual quality, combined with quality anomaly detection and parameter relearning mechanisms, comprehensively monitors and automatically optimizes processing results, ensuring that output images consistently meet the stringent quality standards of security surveillance. To address the unique needs of security surveillance scenarios, the sharpening enhancement process implements differentiated processing for facial and license plate areas. Quality inspection involves format conversion and encoded output according to security surveillance system requirements, better serving the needs of real-world security surveillance applications.
[0016] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the present invention. The purposes and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0017] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A schematic diagram of an embodiment of a security monitoring method according to an embodiment of the present invention; Figure 2 A schematic diagram of a security monitoring system according to an embodiment of the present invention; Figure 3 Schematic diagram of a security monitoring device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0020] The terms "including," "having," and any variations thereof, as used in the embodiments of the present invention are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device comprising a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or device.
[0021] To facilitate understanding of this embodiment, a security monitoring method disclosed in an embodiment of the present invention is first described in detail. Figure 1 As shown, this method includes the following steps: 101. Perform illumination detection on the original security monitoring image to obtain a three-dimensional illumination feature vector; It is understandable that the execution subject of the present invention may be a security monitoring system, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.
[0022] Specifically, representative statistical information is extracted from the raw input image data to construct a multidimensional description of the lighting environment. A pixel-level block processing operation is performed on the input security surveillance raw image. This process divides the entire image into several small pixel blocks, such as 8×8 or 16×16 rectangular regions. The mean and variance of the grayscale values of all pixels within each pixel block are then calculated to generate local brightness distribution data corresponding to each region. The raw light intensity value is obtained by averaging the means of all pixel blocks, reflecting the global brightness level of the current image. To extract information reflecting changes in ambient color temperature, highlighted pixel regions with brightness above a set threshold (e.g., 200) are selected in the raw image as white region candidates. The raw RGB channel data of these regions is then extracted. The R / G and B / G ratios are modeled and mapped onto a predefined standard color temperature curve for interpolation. The original color temperature deviation value of the current image is obtained. This deviation value can represent the degree of deviation between cool and warm colors in the image. To capture the richness of edge detail and spatial structure contrast in an image, a differential gradient magnitude calculation is performed on the entire image. The Sobel or Prewitt operator is used to extract the image's gradient changes in all directions. The distribution of gradient magnitudes across the entire image is then statistically analyzed. The ratio of the maximum to average value in this distribution is then calculated as the raw contrast value, reflecting the image's local texture clarity and overall contrast level. The illumination intensity, color temperature deviation, and contrast values are input into a normalization module. Linear interval scaling is used to map these values to the [0, 1] normalized interval. These three normalized indicators are then assembled into a three-dimensional illumination feature vector [L, T, C], where L is the normalized illumination intensity, T is the normalized color temperature deviation, and C is the normalized contrast.
[0023] 102. Match and determine preset low-light enhancement mode determination conditions, backlight compensation mode determination conditions, strong light suppression mode determination conditions, and standard processing mode determination conditions based on the three-dimensional illumination feature vector to obtain the ISP processing mode type corresponding to the current frame; Specifically, the three-dimensional illumination feature vector [L, T, C] is analyzed at the parameter level, where L is the normalized illumination intensity value, T is the normalized color temperature deviation value, and C is the normalized contrast value. To perform pattern classification, multiple judgment thresholds are preset. The first preset threshold is used to divide the L value into intervals, for example, setting L less than 0.3 as a low-light area, and when L is greater than 0.7 as a high-light area, while the intermediate interval represents a normal-light area. Meanwhile, the second preset threshold is used to classify the C value into levels, for example, setting 0.4 as the boundary value, to determine whether the image is in a low-contrast state. These thresholds are used to perform parallel judgments on the two dimensions of illumination intensity and contrast, respectively obtaining the illumination interval classification results and contrast level classification results for the current frame image. The classification results are then applied to a set of predefined mode determination criteria, and the degree of match between the current frame and each mode is calculated. For example, in low-light enhancement mode, the dual conditions of L < 0.3 and C < 0.4 must be met simultaneously. Backlight compensation mode requires L > 0.7 and a local brightness variance exceeding a certain threshold. High-light suppression mode requires L > 0.8 and minimal T deviation. Standard mode covers scenarios with normal lighting and good color balance. By performing conditional matching logic operations on the determination criteria of each mode, a matching degree value is calculated for each mode, which reflects the degree of closeness between the current 3D lighting characteristics and the corresponding mode. The highest matching degree constitutes the initial mode determination result. To avoid sudden image processing changes at the mode switching boundary, boundary condition detection is performed. When the illumination feature vector is detected as falling within the boundary of multiple modes—for example, when the L value is close to 0.3 or 0.7, and the T and C values also fall within the intersection of multiple conditions—the system does not directly select a single mode. Instead, it performs a weighted fusion based on the matching degree between two adjacent modes. The fusion coefficient is dynamically calculated based on the inverse distance or relative proportion of the matching degree, resulting in a smooth fusion result, effectively alleviating the image style inconsistency caused by mode switching. Based on the boundary fusion results, a historical analysis mechanism is introduced to compare the fusion result of the current frame with the mode selection history of the previous N frames (e.g., the previous five frames). A difference index is calculated using vector differences or the frequency of mode label changes. When this difference exceeds a third preset threshold, indicating a drastic mode change, the mode switching amplitude of the current frame is limited based on the degree of difference. In other words, a full mode switch is not allowed within a single frame. Instead, a gradual mechanism is used to incrementally apply the new mode parameters frame by frame, updating only a certain proportion of the total parameters, such as 20%, each frame to ensure the continuity of the mode transition and the visual consistency of the output image. The ISP processing mode type that best matches the real-time lighting environment and has stable processing for the current frame is obtained.
[0024] 103. Input the three-dimensional illumination feature vector and the ISP processing mode type into the parameterized regression subnetwork for nonlinear mapping calculation and image processing to obtain a first enhanced image; Specifically, based on the ISP processing mode type (such as low-light enhancement, backlight compensation, strong light suppression, or standard mode), a corresponding set of over-parameterized regression sub-network models are selected. Each sub-network structure adopts a unified four-layer structure, but its parameter initialization, weight training process, and branch structure are discretely optimized for specific lighting scenarios, ensuring that the image processing parameters in each mode accurately fit the scene characteristics. The three-dimensional illumination feature vector is input into the regional importance analysis module in the sub-network. This module uses a multi-scale convolution kernel structure, such as a combination of 3×3, 5×5, and 7×7 convolution kernels, combined with a feature pyramid mechanism. It can extract the changes in illumination distribution, regional contrast differences, and detail intensity differences in the scene at different resolutions. Through the channel attention mechanism, the feature maps at each scale are dynamically weighted, and the output is adaptive feature data with region recognition capabilities to accurately capture the characteristics of key areas in the image that are susceptible to illumination. The region-adaptive feature data is fed into the subnetwork's time series memory module, which is implemented using a bidirectional gated recurrent unit (Bi-GRU) or a lightweight temporal convolutional network. Its core task is to model and analyze the changing trends of regional features in previous frames, thereby extracting a time-enhanced feature vector that reflects dynamic illumination changes. This vector not only encodes the illumination structure of the current frame but also embeds the historical trajectory of brightness, contrast, and color changes in similar regions in previous frames, enabling the system to retain contextual memory for inter-frame continuity. The time-enhanced feature vector is then fed into the subnetwork's multi-branch parameter generation network, where parallel nonlinear mapping is performed. This network consists of multiple parallel fully connected layers, each responsible for independently predicting sub-items within the ISP parameter set for exposure control, white balance adjustment, color correction, and focus control. Each branch of the network incorporates a nonlinear activation function (such as LeakyReLU or Swish), batch normalization, and dropout mechanisms to ensure that the mapping result retains nonlinear representation capabilities while being resistant to overfitting. The ISP parameter set control parameters are input into the image processing pipeline, and image enhancement operations including bad pixel correction, demosaicing interpolation, automatic exposure adjustment, red and blue gain white balance correction, 3×3 matrix color correction, mode-specific gamma mapping and edge sharpening are performed in sequence. Each processing step configures the corresponding operation intensity and adjustment factor in real time according to the network output parameters, thereby achieving high-precision, high-dynamic response enhancement processing of the original image, and while maintaining the naturalness of color and the realism of details, improving the recognizability and stability of the picture in extreme lighting environments, and finally outputting the first enhanced image as the current frame processing result.
[0025] The ISP parameter set contains the optimal parameter configuration for the current lighting environment and image characteristics, including the bad pixel detection threshold, exposure control parameters such as exposure time and gain, and the 3×3 matrix coefficients required for color correction. In practice, bad pixel identification is performed on the original image based on the detection threshold parameters. This process scans each pixel in the image and compares it with the set threshold. If a pixel value significantly deviates from the distribution of its surrounding pixels or exhibits saturation anomalies, it is marked as a bad pixel. A 3×3 neighborhood window is then extracted with the bad pixel as the center, and a median filter is performed on the valid pixels within the window. The median value is used as the replacement value for the bad pixel to complete the bad pixel repair operation. This process removes isolated outliers generated by the sensor or transmission process, while outputting basic RGB three-channel image data with good structural continuity. Based on the exposure time and gain values in the ISP parameter set, a linear brightness transformation is performed on the RGB base image data. This linearly amplifies the channel values of each pixel to achieve an overall image brightness boost or suppress. In the linear transformation formula, the exposure time determines the basis for the pixel's integrated energy, while the gain value serves as an amplification factor to adjust the output intensity. The resulting intermediate image data exhibits lighting that better matches the actual environment, dynamically optimizing the image's visibility in dark areas and detail in bright areas. The color correction matrix parameters are then used to perform a matrix multiplication on the intermediate image data, multiplying each pixel's RGB vector by the 3×3 correction matrix to complete the affine transformation of the color gamut. This operation effectively corrects color casts caused by inconsistent sensor responses across different color channels, while accurately mapping the sensor's color gamut to the target display's color gamut (e.g., sRGB), generating pre-enhanced image data. USM sharpening is performed based on the pre-enhanced image data. The image is Gaussian blurred to obtain a low-frequency image. The blurred image is then subtracted from the original image to obtain the edge enhancement component, and the component is superimposed back onto the original image through the set sharpening coefficient, thereby enhancing the clarity of the image edge structure and the contrast of texture details. The parameter adjustment of USM sharpening is controlled by the ISP parameter set. For example, the sharpening radius and intensity factor can be adapted and adjusted according to the current processing mode, so that the final output first enhanced image achieves the best performance in terms of detail level, structural integrity and visual transparency.
[0026] 104. Perform penalty calculation on the first enhanced image to obtain a total penalty value, and perform gradient reverse adjustment based on the total penalty value to obtain a second enhanced image; Specifically, the brightness histograms of the first enhanced image and the reference standard image are statistically analyzed to construct brightness probability distribution curves for the two images. Then, the absolute value operation of the bit-by-bit difference is performed on each grayscale position of the histogram, and all the differences are accumulated to form a brightness consistency penalty value. This value reflects the overall degree of deviation of the enhanced image from the reference image in terms of brightness distribution structure. At the same time, color naturalness evaluation is performed based on the RGB three-channel data of the first enhanced image. By analyzing the increase in color saturation and the hue offset angle of each pixel (the change in Hue distribution can be calculated by converting RGB to HSV space), it is detected whether there is color oversaturation or color cast, and based on this, a color naturalness penalty value is generated to measure whether the color enhancement is distorted or deviates from the natural lighting state. The image's spatial detail preservation capability is modeled and analyzed. Gradient amplitude extraction is performed on both the original and enhanced images, using the Sobel or Scharr operator to capture changes in edge strength. The degree of gradient loss at the edge structure of the image before and after enhancement is then statistically analyzed. The detail preservation penalty term is then constructed, combining the energy change of the local texture. If the enhanced image exhibits gradient drop or texture blurring at the original edge location, detail loss is considered. After calculating the penalty terms for these three dimensions, the brightness consistency penalty, color naturalness penalty, and detail preservation penalty are weighted and summed, with weight coefficients preset based on visual perception sensitivity, to calculate the total penalty value for the current frame. This total penalty value is then compared to a pre-set penalty threshold (e.g., 0.5). If the total penalty value exceeds this threshold, the current ISP processing parameters are deemed to pose a risk of over-enhancement or distortion, automatically triggering the parameter callback mechanism and outputting a parameter adjustment trigger signal. Based on this signal, the gradient descent optimization module retained in the deep network is called to perform a reverse optimization process based on the ISP parameter set. This process uses the current ISP parameters as the initial state. Under the guidance of the loss function as the penalty value, it uses a small step size (such as 0.01) to iteratively optimize for up to 5 rounds, gradually converging to a new parameter solution with a low penalty state, thus forming the optimized ISP parameter set. Using the updated ISP parameter set, the original security surveillance image is re-processed through the full ISP pipeline, repeating operations such as bad pixel correction, exposure adjustment, color correction, and sharpening enhancement, and outputting a second enhanced image.
[0027] 105. Perform quality inspection on the second enhanced image to obtain a target enhanced image that meets security monitoring quality standards.
[0028] Specifically, image signal features are extracted from the second enhanced image, and the peak signal-to-noise ratio is calculated based on this. By comparing the pixel-level error information between the enhanced image and the original image or the standard reference image, the degree of noise introduced by the image enhancement is quantified. Subsequently, a ratio operation is performed on the maximum and minimum brightness values of the image to obtain a dynamic range evaluation value, which reflects the image's performance in terms of light and dark contrast and brightness level transitions. If the dynamic range is insufficient, there will be darkened or overexposed areas, affecting detail perception and target recognition capabilities. At the same time, the image is converted from RGB space to Lab color space, and the standard color difference formula is called to calculate the color deviation between the current image and the reference image in the L (brightness), a (red and green channels), and b (yellow and blue channels) dimensions to obtain a color fidelity evaluation value. This value is used to determine whether the image has obvious color cast or color distortion, and is suitable for color correction result testing in low illumination and strong backlight scenes. To evaluate the image's texture detail and edge preservation, a modulation transfer function (MTF) is used to calculate the image's edge response strength in the frequency domain. This is combined with edge sharpness metrics (such as the ratio of maximum gradient to average gradient) to statistically analyze local edge variations, yielding a detail clarity assessment value. This ensures that the image enhancement process avoids blurring, oversmoothing, or artifacts. A structural similarity index and its multi-scale expansion algorithm are introduced to assess overall visual quality at the image structural level. This method comprehensively considers the matching degree of brightness, contrast, and structural information, simulating the human eye's subjective perception of image quality, and outputs a global visual quality score as the final perceptual metric. After completing the quantitative evaluation of these four dimensions, the dynamic range assessment value is compared with the fourth preset threshold, the color fidelity assessment value with the fifth preset threshold, the detail clarity assessment value with the sixth preset threshold, and the visual quality assessment value with the seventh preset threshold. When all dimensions exceed their respective thresholds, the second enhanced image is deemed to meet the technical standards for security surveillance image quality. At this point, format conversion and encoding processing are performed to map the processed linear RGB data to the standard sRGB color space, and generate JPEG images or H.264 encoded video frames according to the requirements of the target platform. The final output is a target enhanced image with rich brightness levels, true and natural colors, sharp and clear details, and good visual subjective perception.
[0029] In a specific embodiment, the process of executing step 101 may specifically include the following steps: The original security monitoring image is segmented into pixel blocks to obtain multiple pixel blocks, and the mean and variance of the grayscale values in each pixel block are calculated to obtain the original light intensity value; Extract white area pixels from the original security monitoring image, perform ratio calculation on the red, green and blue channel values of the white area pixels and perform standard color temperature curve interpolation to obtain the original color temperature deviation value; Based on the original security monitoring image, differential gradient amplitude calculation and edge intensity statistical analysis are performed to obtain differential gradient amplitude distribution data, and the ratio of the maximum value to the average value of the differential gradient amplitude distribution data is calculated to obtain the original contrast value; The original light intensity value, the original color temperature deviation value, and the original contrast value are normalized and vector-combined to obtain a three-dimensional light feature vector including the normalized light intensity value, the normalized color temperature deviation value, and the normalized contrast value.
[0030] Specifically, in the calculation path of light intensity, pixel block-level segmentation is performed on the input original image. This operation evenly divides the image size into several fixed-size regional units, typically 8×8 pixel blocks, to construct a distributed set of pixel blocks covering the entire image. Subsequently, the grayscale values in each pixel block are statistically processed block by block, including calculating the grayscale mean and variance of all pixels in each block. The grayscale mean is used to measure the local brightness level, while the variance characterizes the local dispersion of brightness. By averaging the grayscale means of all pixel blocks, the original light intensity value at the full image level is obtained. At the same time, the color deviation of the current image is extracted from the color temperature dimension. To this end, the white areas in the image are located and analyzed. This operation filters high-brightness pixels by setting a brightness threshold, for example, extracting areas with grayscale values above 200, and further restricting the color saturation variation range of these areas to avoid false white misjudgments caused by saturated highlights. For the filtered white area pixels, the values are extracted from their RGB channels, and the red / green channel ratio (R / G) and the blue / green channel ratio (B / G) of each pixel are calculated. The obtained ratios are used to characterize the color shift direction and intensity of the current white pixel. By consulting a standard color temperature curve (such as the CIE standard light source D65 curve or the Plankian locus color temperature trajectory), an interpolation algorithm is used to map the combination of R / G and B / G to a color temperature positioning curve, and the closest point on the curve is solved to obtain the original color temperature deviation value of the current image. This value reflects the degree of deviation between the image white balance and the ideal lighting state. On this basis, to obtain the original contrast index of the image, analysis and processing are performed along the structural gradient path. Differential gradient operations are performed on the original image. Gradient calculation methods include Sobel, Scharr, or Roberts operators. The brightness changes in the horizontal and vertical directions are calculated at each pixel point, and then the gradient amplitude of the pixel is synthesized. The gradient amplitude of all pixels is statistically analyzed, and the maximum and average values are extracted. The ratio of the maximum to the average value is then calculated as the original contrast value. This ratio has stronger edge perception ability and can simultaneously characterize the degree of contrast separation between sharp edges and background details in the image. The original light intensity value, original color temperature deviation value, and original contrast value are normalized to unify the dimension and scaling interval, eliminating the differences in the value range of the original features. The normalization method uses linear scaling to perform a minimum-maximum normalization mapping of each value according to its physically defined range (for example, brightness ranges from 0 to 255, color temperature deviation ranges from 0 to 1000, and contrast ratio ranges from 1 to 10), so that all normalized values fall into the closed interval [0, 1]. The three normalized indicators are sequentially combined to form a three-dimensional illumination feature vector [L, T, C], where L represents the normalized light intensity, T represents the normalized color temperature deviation, and C represents the normalized contrast value.
[0031] In a specific embodiment, the process of executing step 102 may specifically include the following steps: Comparing the normalized light intensity value in the three-dimensional light feature vector with a first preset threshold to obtain a light intensity interval classification result, and comparing the normalized contrast value with a second preset threshold to obtain a contrast level classification result; Based on the illumination intensity interval classification results and the contrast level classification results, the preset low-light enhancement mode determination conditions, backlight compensation mode determination conditions, strong light suppression mode determination conditions, and standard processing mode determination conditions are matched and calculated one by one to obtain the matching degree value of each mode and the initial mode determination result; Based on the initial mode determination result, boundary condition detection is performed. When the three-dimensional illumination feature vector is detected in the mode boundary area, the matching degree values of the two adjacent modes are weighted fused and a smooth transition weight coefficient is generated to obtain the boundary fusion processing result. The boundary fusion processing result of the current frame is compared with the historical mode selection records of the previous N frames to obtain the difference degree. When the difference degree exceeds the third preset threshold, the mode switching amplitude is adjusted to obtain the ISP processing mode type corresponding to the current frame.
[0032] Specifically, threshold comparison operations are performed on the normalized light intensity value and the normalized contrast value in the vector, respectively, so as to achieve a preliminary interval judgment on the type of lighting scene in which the current frame is located. Among them, the system sets a first preset threshold value for segmenting the normalized light intensity value L, for example, L<0.3 is defined as a low light interval, L>0.7 is defined as a strong light interval, and 0.3≤L≤0.7 is classified as a medium light interval; the second preset threshold value is used to classify the normalized contrast value C, for example, setting C<0.4 as low contrast and C≥0.4 as high contrast. Through these two comparison operations, the category labels of the current frame in the two key dimensions of brightness and contrast are obtained respectively, that is, the light intensity interval classification result and the contrast level classification result. The above two classification results are used as the basis for judgment, and are matched one by one with the four pre-defined processing mode conditions in the system. These four modes are low light enhancement, backlight compensation, strong light suppression and standard processing. Each mode has a clear combination of judgment conditions. For example, the low light enhancement mode requires L < 0.3 and C < 0.4, which is suitable for night or dark indoor images; the backlight compensation mode requires L > 0.7 and the local brightness variance exceeds the set threshold σ 2=0.6, primarily for scenes with strong backlight windows. The strong light suppression mode requires L > 0.8 and a small color temperature deviation, T, to address overexposed areas caused by strong sunlight or strong lighting. The standard processing mode requires 0.3 ≤ L ≤ 0.7, C ≥ 0.4, and a T deviation less than 0.3, for normal indoor and outdoor lighting conditions. The system substitutes the actual L, T, and C values of the current frame into each mode decision function and calculates the distance or offset between the decision threshold and the threshold to generate a matching score for each mode. A higher value indicates a greater degree of fit for the current frame's lighting conditions. After determining the matching scores for all modes, the mode with the highest matching score is selected based on the principle of maximum matching as the initial processing mode decision for the current frame. Given the frequent fluctuations in ambient lighting in real applications or the critical state between multiple modes, using hard switching logic to directly switch to a new mode can easily lead to abrupt changes in image style or reduced processing stability. Therefore, a boundary condition detection mechanism is designed to identify whether illumination features are within mode boundary regions. This mechanism determines this by analyzing the Euclidean distance or matching difference between the current frame's feature vector and the center points of each mode. When the matching degrees of two modes are simultaneously high and the difference is below a preset boundary threshold, the current frame is considered to be within the mode switching transition zone. In this case, the system does not immediately switch to the new target mode. Instead, it performs a weighted fusion of the matching values of the two adjacent modes to generate smooth transition weight coefficients. The parameters or processing styles of the two modes are proportionally blended according to these weights, resulting in a boundary fusion result. This ensures a natural and continuous transition between multiple modes, without artifacts such as sudden color changes or exposure shifts. To complete the boundary fusion system, an inter-frame historical mode buffering mechanism is introduced to improve processing continuity and system response stability. The system continuously maintains a record of mode selections for the previous N frames and performs a difference analysis between the current frame's boundary fusion result and the mode labels or blending weight vectors of the previous N frames. This difference is calculated based on the average vector distance or mode switching frequency within a temporal sliding window. When the calculation results indicate that the mode change between the current frame and the historical frame exceeds a third preset threshold, the system determines that there is a potential risk of mode jump. At this time, it will automatically adjust the mode switching amplitude of the current frame, that is, limit the parameter adjustment step from the historical mode to the target mode, so that the activation process of the new mode shows a gradual evolution trend. For example, a limit strategy of no more than 20% per frame is adopted, and processing parameters such as exposure, white balance, and sharpening intensity are smoothly adjusted frame by frame until the parameter settings corresponding to the new mode are fully transitioned. The ISP processing mode type corresponding to the current frame is obtained.
[0033] In a specific embodiment, the process of executing step 103 may specifically include the following steps: Based on the ISP processing mode type matching corresponding over-parameterized regression sub-network; The 3D illumination feature vector is input into the regional importance analysis module of the parameterized regression sub-network for multi-scale convolution feature extraction to obtain regional adaptive feature data. The regional adaptive feature data is input into the time series memory module of the parameterized regression sub-network to analyze the inter-frame illumination change trend and generate a time series enhanced feature vector; The time series enhancement feature vector is input into the multi-branch parameter generation network of the parameterized regression sub-network for parallel nonlinear mapping to obtain the ISP parameter set; The security monitoring original image is processed based on the ISP parameter set to obtain a first enhanced image.
[0034] Specifically, the system identifies the lighting environment of the current frame image and matches and classifies the 3D illumination feature vector [L, T, C] based on four a priori defined modes: low-light enhancement, backlight compensation, high-light suppression, and standard processing. After classification, the system selects an over-parameterized regression subnetwork model corresponding to the mode type. While each subnetwork maintains a consistent architecture, its detailed configuration, such as training weights, input channels, activation sensitivity, and dropout strategy, is tailored to the lighting scenario. This ensures that the network can extract highly adaptable and high-resolution feature representations under diverse lighting conditions and outputs a refined control scheme for the corresponding ISP parameters. After mode matching, the 3D illumination feature vector of the current frame is fed into the first-layer module of the selected subnetwork, the Region Importance Analysis module. This module utilizes a multi-scale convolutional kernel structure and incorporates convolutional operators with different receptive fields, such as 3×3, 5×5, and 7×7, to fully extract the response characteristics of each region in the image to illumination changes at different scales. To enhance the network's ability to focus on key regions, a channel attention mechanism (such as SE or CBAM) is introduced between convolutional channels. Dynamic weights are assigned to each convolutional channel based on its contribution to the final output, resulting in region-adaptive feature data with both regional recognition and scale responsiveness. This region-adaptive feature data is fed into the time series memory module of the parameterized regression subnetwork. This module models inter-frame illumination trends and employs bidirectional gated recurrent units or a one-dimensional temporal convolutional network to capture the temporal illumination path of the image. By sequentially encoding regional features across multiple frames, the system effectively extracts dynamic patterns such as illumination intensity fluctuations, gradual changes in color temperature, and contrast fluctuations. A temporal enhancement feature vector is constructed. This vector not only incorporates the enhancement strategy required for the current frame but also implicitly incorporates illumination behavior characteristics from previous frames. This ensures forward-looking and inter-frame consistency in ISP parameter generation, preventing parameter abrupt changes caused by short-term illumination perturbations. The temporal enhancement feature vector is then fed into a multi-branch parameter generation network to perform parallelized nonlinear mapping computations for different ISP parameter types. The network's overall structure consists of a collection of independent sub-networks, each branch corresponding to an ISP functional module, such as auto-exposure (outputting exposure time and gain values), auto-white balance (outputting red and blue channel gain and color temperature compensation factors), color correction (outputting a 3×3 color conversion matrix), and autofocus (outputting focal length and depth of field parameters). Each branch contains a two-layer fully connected structure, using LeakyReLU or Swish as activation functions to preserve the output's high sensitivity to subtle changes in input features. Dropout and BatchNorm mechanisms are also introduced to ensure output stability across multiple frames. By generating multiple ISP sub-parameters in parallel through this branching structure, the system ensures overall coordination and joint optimization of the entire ISP parameter set while maintaining mutual decoupling.After completing the ISP parameter set inference, this structured parameter data is fed into the image processing pipeline, initiating a parameter-controlled image enhancement process. This process follows the signal processing path of the hardware-level ISP chip, identifying pixel-level outliers and performing median filtering correction based on a bad pixel detection threshold to restore the coherence and stability of the underlying RGB channels. It then linearly amplifies the pixel brightness of the three RGB channels using exposure time and gain parameters to adjust the overall image exposure level. A 3×3 color correction matrix is then applied to each pixel's RGB vector for color space mapping and color gamut compression, ensuring that the image color meets the standards of the monitoring display system. The red and blue channel gain parameters are then used for white balance correction to level the image color temperature deviation. Finally, gamma modulation and sharpening are performed, controlling the nonlinear image brightness mapping curve using the corresponding gamma value. The unsharp mask (USM) edge enhancement algorithm is then applied based on the target sharpness coefficient, ultimately generating a first enhanced image with coherent structure, natural colors, balanced brightness, and clear details.
[0035] In a specific embodiment, the step of performing image processing on the original security monitoring image based on the ISP parameter set to obtain the first enhanced image may specifically include the following steps: According to the detection threshold parameters in the ISP parameter set, the security monitoring original image is subjected to bad pixel detection and neighborhood median filtering correction to obtain RGB three-channel basic image data; According to the exposure time and gain value in the ISP parameter set, the pixel brightness linear transformation is performed on the RGB three-channel basic image data to obtain the intermediate image data; The color correction matrix in the ISP parameter set is applied to perform matrix multiplication transformation on the intermediate image data to obtain pre-enhanced image data, and USM sharpening is performed based on the pre-enhanced image data to obtain a first enhanced image.
[0036] Specifically, based on the detection threshold parameters specified in the ISP parameter set, the original image data is traversed pixel by pixel. The absolute difference between each pixel value and the grayscale mean of its 3×3 or 5×5 neighboring pixels is calculated. This difference is then compared with the bad pixel detection threshold. If the value exceeds the threshold, the pixel is identified as an outlier. All pixels marked as bad pixels are corrected using a neighborhood median filter, replacing the current pixel with the median grayscale value of all pixels in its neighborhood. This method effectively eliminates isolated noise points while preserving local texture structure. After bad pixel correction, the RGB three-channel basic image data is restored. Based on this, the overall image brightness is compensated to adapt to changing scene lighting conditions. The exposure time and gain parameters included in the ISP parameter set are used to perform a pixel-level linear brightness transformation. A set of linear transformation functions is applied to the pixel values of each RGB channel in the image, multiplying each channel's pixel value by a gain factor. This factor also takes into account the proportional effect of exposure time on overall brightness, thereby achieving a balanced improvement in dark or overly bright areas of the original image. The transformed image forms the intermediate image data. Image color accuracy is corrected to ensure realistic and natural visuals on various display devices. The system extracts a pre-trained 3×3 color correction matrix from the ISP parameter set and applies it to the RGB vector of each pixel in the intermediate image data, performing a matrix multiplication operation. This operation treats each RGB triplet as a one-dimensional column vector and multiplies it by the 3×3 correction matrix on the left to complete a linear mapping of the color space. For example, when the original sensor response color gamut does not conform to the standard sRGB space, the color correction matrix accurately maps the image from the sensor color gamut to the color gamut supported by standard display devices, eliminating color distortion caused by sensor manufacturing variations, lens spectral transmittance variations, or illumination color temperature deviations. The resulting pre-enhanced image data is more realistic in terms of hue, saturation, and contrast. To enhance image clarity, the pre-enhanced image is spatially sharpened using Unsharp Mask (USM), a representative nonlinear high-pass enhancement technique. The basic process includes first performing Gaussian blur on the pre-enhanced image to generate a low-frequency smooth version, then subtracting the blurred image from the original image to obtain the high-frequency information component of the image, which represents the edge and texture details in the image. The high-frequency component is then linearly amplified according to the sharpening intensity coefficient, and finally superimposed with the original image to achieve edge enhancement and contour enhancement without destroying the overall brightness and color structure of the image. The intensity factor of this process is dynamically adjusted by the sharpening control parameters provided by the ISP parameter set. Different sharpening strategies can be automatically selected in different lighting modes. For example, the sharpening coefficient is slightly lower in low-light mode to suppress noise amplification, while it can be appropriately enhanced in backlight and strong light modes to highlight structural boundaries. Through the above steps, the first enhanced image is finally obtained.
[0037] In a specific embodiment, the process of executing step 104 may specifically include the following steps: Perform brightness histogram statistical analysis and bit-by-bit absolute value accumulation of the difference between the first enhanced image and the reference standard image to obtain the brightness consistency penalty term value. Simultaneously, perform color saturation over-enhancement detection and hue offset angle calculation based on the RGB three-channel data of the first enhanced image to obtain the color naturalness penalty term value. Furthermore, perform gradient amplitude change statistics and detail loss degree quantitative analysis on the first enhanced image to obtain the detail preservation penalty term value. The brightness consistency penalty item value, the color naturalness penalty item value, and the detail preservation penalty item value are weighted and summed to obtain a total penalty value, and the total penalty value is compared with a preset penalty threshold. When the total penalty value exceeds the preset penalty threshold, the parameter callback processing mechanism is triggered to obtain a parameter adjustment trigger signal; Based on the parameter adjustment trigger signal, the gradient descent algorithm is started to perform reverse optimization calculation on the ISP parameter set to obtain the optimized ISP parameter set; The optimized ISP parameter set is reapplied to the original security monitoring image for ISP pipeline reprocessing calculation to obtain a second enhanced image.
[0038] Specifically, a multi-dimensional comparative evaluation is conducted between the first enhanced image and the reference standard image, establishing a quantitative index system for brightness structure, color representation, and detail preservation. For brightness consistency, a grayscale histogram statistical analysis is performed on the two images to extract the frequency of occurrence of each grayscale level. By comparing the differences between the first enhanced image and the reference standard image across all grayscale levels, the cumulative deviation between the overall brightness distribution is calculated, resulting in a brightness consistency penalty value. A large value indicates that the current enhancement result has issues with overall exposure control or local brightness adjustment, such as overexposure, brightness compression, or abnormal brightness contrast in local areas. For color naturalness, the RGB data of the first enhanced image is converted to the HSV color space, and the saturation (S) channel and hue (H) channel information are extracted. The system detects whether there are large areas of high saturation in the image, such as a significantly higher proportion of pixels with a saturation close to 1 than in normal images, indicating over-enhancement. Furthermore, the average hue angle shift between the enhanced image and the reference image is statistically analyzed to quantify hue shift during the enhancement process. If an abnormal increase in saturation and a large hue shift are detected, it indicates that the image has lost naturalness during the color adjustment process. A color naturalness penalty is assigned, with higher values indicating greater color deviation. The enhanced image also analyzes the preservation of structural detail. Image gradients are calculated for both the original and first enhanced images, extracting edge and texture information. The gradient amplitude change between the two images is then statistically analyzed. The system assesses the extent of image detail loss during the enhancement process by analyzing whether edges become blurred and whether textures appear compressed or blurred. This is then quantified as a detail preservation penalty. This penalty is significantly increased if the enhanced image exhibits edge blurring or loss of detail in areas with significant original structure. After these three penalty values are generated, they are weighted and summed, with weights set to 0.4 (brightness consistency), 0.3 (color naturalness), and 0.3 (detail preservation), resulting in a unified total penalty value. This total value is used to comprehensively assess image quality deviation and is compared against a preset penalty threshold set at 0.5. When the total penalty value exceeds this threshold, it indicates that at least one of the current image's brightness, color, or detail fails to meet quality standards. Therefore, the system triggers the parameter callback mechanism and outputs a parameter adjustment trigger signal. Based on this trigger signal, the system calls the built-in gradient descent optimization module to perform a small, localized reverse optimization operation on the key parameters in the current ISP parameter set. The optimization goal is to reduce the sum of the three penalty terms to an acceptable range. This process controls the learning rate to 0.01 and limits the maximum number of iterations to 5 to ensure the convergence of the parameter callback and the stability of the image style.After optimization, a new set of ISP parameters is obtained. This set is more robust in image quality than the previous version and effectively avoids typical issues such as over-enhancement, color cast, and detail loss. These optimized ISP parameters are then reapplied to the original surveillance image, restarting the ISP image processing pipeline. Operations such as bad pixel repair, exposure adjustment, white balance compensation, color correction, gamma shift, and edge sharpening are performed in sequence to generate a second enhanced image.
[0039] In a specific embodiment, the process of executing step 105 may specifically include the following steps: The peak signal-to-noise ratio and maximum-to-minimum brightness ratio of the second enhanced image are calculated to obtain a dynamic range evaluation value. The color fidelity evaluation value is obtained based on Lab color space conversion and color difference formula calculation. The detail clarity evaluation value is obtained through modulation transfer function calculation and edge sharpness index statistical analysis. The visual quality evaluation value is obtained by using the structural similarity index and multi-scale SSIM algorithm. Comparing the dynamic range evaluation value with a fourth preset threshold, comparing the color fidelity evaluation value with a fifth preset threshold, comparing the detail clarity evaluation value with a sixth preset threshold, and comparing the visual quality evaluation value with a seventh preset threshold, to obtain a four-dimensional quality compliance status determination result; Based on the four-dimensional quality compliance status judgment result, the second enhanced image is format converted and encoded and output to obtain a target enhanced image that meets the security monitoring quality standard.
[0040] Specifically, a dynamic range analysis is performed on the second enhanced image, calculating two core brightness metrics. The first is the peak signal-to-noise ratio (PSNR), calculated based on image brightness data. This metric measures whether excessive signal noise was introduced during brightness adjustment. The second is the ratio of the maximum to minimum brightness values. This value reflects the image's brightness coverage and is particularly useful in high-contrast scenes or bright light environments. These two calculations yield a dynamic range assessment value, which is used to determine the integrity and reasonableness of the image's brightness structure. The second enhanced image is converted to Lab color space, a color model that exhibits stronger perceptual isometry than RGB and is therefore suitable for color deviation calculation. After the conversion, the L channel (luminance), a channel (red-green axis), and b channel (yellow-blue axis) are extracted from the enhanced image and the reference image, respectively. The color differences between the two images along these three dimensions are then calculated, resulting in a color fidelity assessment value that represents the overall degree of color difference. The closer this metric is to the reference image, the more natural the original colors are preserved during the enhancement process, and the less likely oversaturation or hue shifts are caused by parameter adjustments. The system then calculates detail clarity, using modulation transfer functions (MTFs) and edge sharpness metrics. The system extracts multiple edge regions from the image and constructs a modulation transfer function (MTF) by analyzing the brightness gradients in these regions. This function characterizes the image's responsiveness at different spatial frequencies, thereby measuring the sharpness of edge transitions and the integrity of edge information. Statistical analysis is also performed on the overall edge strength of the image, calculating the ratio of the average gradient change to the maximum gradient difference to form an edge sharpness metric. This metric quantifies the perceptibility and visual sharpness of structural edges within the image. These two metrics are combined to produce a numerical value for detail clarity. After completing the structural, color, and brightness layer assessments, the subjective visual quality analysis phase begins. The second enhanced image is compared to the reference image using a structural similarity index and a multi-scale structural similarity algorithm. The system extracts and compares multi-level features based on brightness similarity, contrast consistency, and structural information preservation. In the multi-scale algorithm, similarity calculations are repeated at multiple image resolutions to avoid misjudgments caused by image scaling or local deformation. From these statistics, a visual quality assessment value based on structure perception is derived. A higher value indicates that the image is closer to the original visual perception and more suitable for storage or transmission to a surveillance terminal. After all four indicators are evaluated, they are compared against four predefined quality thresholds.The dynamic range assessment value must be greater than a fourth preset threshold, such as a brightness ratio greater than 20; the color fidelity assessment value must be less than a fifth preset threshold, such as a color difference value less than 5; the detail clarity assessment value must be greater than a sixth preset threshold, such as an edge sharpness index of at least 1.5; and the visual quality assessment value must be greater than a seventh preset threshold, requiring a structural similarity index of at least 0.8. The four judgment results are combined for a logical operation. If all results meet their respective thresholds, a four-dimensional quality compliance status determination result is generated, which serves as the basis for output control. If the image quality is determined to meet the standard, the second enhanced image enters the format conversion and standard encoding stage. Based on the image data type and terminal device compatibility requirements, the image is mapped from the internal linear RGB space to the standard sRGB space to ensure color consistency across different devices. The appropriate output encoding method is then selected based on the security system configuration requirements, such as using JPEG for still image storage or the H.264 encoding standard to compress the image into frame data for video encoding. The conversion process performs image compression ratio calculation, bitrate control, and data integrity verification to ensure that the output image meets both visual quality standards and network transmission and storage efficiency constraints. The final output image is the target enhanced image.
[0041] The above describes the security monitoring method in the embodiment of the present invention. The following describes the security monitoring system in the embodiment of the present invention. Figure 2 In one embodiment of the present invention, a security monitoring system includes: Light detection module 201, used to perform light detection on the original security monitoring image to obtain a three-dimensional light feature vector; Matching determination module 202, configured to match preset low-light enhancement mode determination conditions, backlight compensation mode determination conditions, high-light suppression mode determination conditions, and standard processing mode determination conditions based on the three-dimensional illumination feature vector to obtain the ISP processing mode type corresponding to the current frame; An image processing module 203 is configured to input the three-dimensional illumination feature vector and the ISP processing mode type into a parameterized regression subnetwork for performing nonlinear mapping calculation and image processing to obtain a first enhanced image; a penalty calculation module 204 configured to perform penalty calculation on the first enhanced image to obtain a total penalty value, and perform gradient reverse adjustment based on the total penalty value to obtain a second enhanced image; The quality detection module 205 is used to perform quality detection on the second enhanced image to obtain a target enhanced image that meets the security monitoring quality standard.
[0042] Through the collaborative efforts of these components, a three-dimensional illumination feature vector is constructed by simultaneously detecting light intensity, color temperature, and contrast. Compared to traditional single-parameter illumination detection methods, this algorithm can comprehensively and accurately describe the multi-dimensional characteristics of complex lighting environments. It also establishes an intelligent switching system among four modes: low-light enhancement, backlight compensation, bright light suppression, and standard processing. Combined with boundary smoothing and a historical memory mechanism, it automatically selects the optimal processing strategy based on real-time lighting conditions, avoiding the limitations of traditional fixed-parameter processing and ensuring processing continuity and stability under varying lighting conditions. A region importance analysis module performs differentiated processing for face, license plate, background, and boundary regions. Combined with a time series memory module and a multi-branch parameter generation network, it enables specialized optimization for the core mission requirements of security monitoring. An over-parameterized regression subnetwork is used for dynamic learning of ISP parameters. By constructing complex nonlinear mapping relationships, it automatically generates the optimal ISP parameter set based on lighting conditions and scene characteristics, breaking through the limitations of fixed parameter settings in traditional ISP processing and achieving true parameter adaptive adjustment. By establishing triple penalty constraints for brightness consistency, color naturalness, and detail preservation, and optimizing ISP parameters in real time through a gradient inverse adjustment mechanism, the system effectively prevents over-enhancement and color shift during image processing, ensuring that the resulting image quality is enhanced while maintaining a natural and realistic visual experience. Within the complete ISP processing flow—defective pixel correction, demosaicing, exposure control, white balance correction, color correction, gamma correction, and sharpening enhancement—each step dynamically fine-tunes parameters based on the current frame's processing results, achieving more refined and adaptive image enhancement than traditional fixed pipeline processing. A four-dimensional evaluation system encompassing dynamic range, color fidelity, detail clarity, and visual quality, combined with quality anomaly detection and parameter relearning mechanisms, comprehensively monitors and automatically optimizes processing results, ensuring that output images consistently meet the stringent quality standards of security surveillance. To address the unique needs of security surveillance scenarios, the sharpening enhancement process implements differentiated processing for facial and license plate areas. Quality inspection involves format conversion and encoded output according to security surveillance system requirements, better serving the needs of real-world security surveillance applications.
[0043] above Figure 2 The security monitoring system in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The security monitoring device in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0044] Figure 3The diagram below is a schematic diagram of the structure of a security monitoring device provided by an embodiment of the present invention. The security monitoring device 300 may vary significantly depending on its configuration or performance. It may include one or more processors (central processing units, CPUs) 310 (e.g., one or more processors), memory 320, and one or more storage media 330 (e.g., one or more mass storage devices) storing application programs 333 or data 332. The memory 320 and storage media 330 may be either transient or persistent storage. The program stored in the storage medium 330 may include one or more modules (not shown), each of which may include a series of instructions operating on the security monitoring device 300. Furthermore, the processor 310 may be configured to communicate with the storage medium 330, executing the series of instructions stored in the storage medium 330 on the security monitoring device 300 to implement the steps of the security monitoring method described above.
[0045] The security monitoring device 300 may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be understood by those skilled in the art that Figure 3 The illustrated security monitoring device structure does not constitute a limitation on the security monitoring device provided by the present invention, and may include more or fewer components than illustrated, or a combination of certain components, or a different arrangement of components.
[0046] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, cause the computer to execute the steps of the security monitoring method.
[0047] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0048] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0049] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A security monitoring method, characterized in that: include: Perform illumination detection on the original security monitoring image to obtain a three-dimensional illumination feature vector; According to the three-dimensional illumination feature vector, a matching determination is performed on preset low-light enhancement mode determination conditions, backlight compensation mode determination conditions, strong light suppression mode determination conditions, and standard processing mode determination conditions to obtain the ISP processing mode type corresponding to the current frame; Inputting the three-dimensional illumination feature vector and the ISP processing mode type into a parameterized regression subnetwork for nonlinear mapping calculation and image processing to obtain a first enhanced image; performing penalty calculation on the first enhanced image to obtain a total penalty value, and performing gradient reverse adjustment based on the total penalty value to obtain a second enhanced image; Performing quality inspection on the second enhanced image to obtain a target enhanced image that meets security monitoring quality standards.
2. The security monitoring method according to claim 1, characterized in that: The illumination detection is performed on the original security monitoring image to obtain a three-dimensional illumination feature vector, including: The original security monitoring image is segmented into pixel blocks to obtain multiple pixel blocks, and the mean and variance of the grayscale values in each pixel block are calculated to obtain the original light intensity value; Extracting white area pixels according to the original security monitoring image, and performing ratio calculation and standard color temperature curve interpolation operation on the red, green and blue three-channel values of the white area pixels to obtain an original color temperature deviation value; Performing differential gradient amplitude calculation and edge strength statistical analysis based on the original security monitoring image to obtain differential gradient amplitude distribution data, and performing a maximum value to average value ratio operation on the differential gradient amplitude distribution data to obtain an original contrast value; The original illumination intensity value, the original color temperature deviation value, and the original contrast value are normalized and vector-combined to obtain a three-dimensional illumination feature vector including the normalized illumination intensity value, the normalized color temperature deviation value, and the normalized contrast value.
3. The security monitoring method according to claim 2, characterized in that: The matching and determination of preset low-light enhancement mode determination conditions, backlight compensation mode determination conditions, strong light suppression mode determination conditions, and standard processing mode determination conditions are performed according to the three-dimensional illumination feature vector to obtain the ISP processing mode type corresponding to the current frame, including: Comparing the normalized light intensity value in the three-dimensional light feature vector with a first preset threshold to obtain a light intensity interval classification result, and comparing the normalized contrast value with a second preset threshold to obtain a contrast level classification result; According to the illumination intensity interval classification result and the contrast level classification result, the preset low light enhancement mode determination condition, backlight compensation mode determination condition, strong light suppression mode determination condition and standard processing mode determination condition are matched and calculated one by one to obtain the matching degree value of each mode and the initial mode determination result; Performing boundary condition detection based on the initial mode determination result, and when it is detected that the three-dimensional illumination feature vector is in the mode boundary area, performing weighted fusion calculation on the matching degree values of two adjacent modes and generating a smooth transition weight coefficient to obtain a boundary fusion processing result; The boundary fusion processing result of the current frame is compared with the historical mode selection records of the previous N frames to obtain the difference degree. When the difference degree exceeds the third preset threshold, the mode switching amplitude is adjusted to obtain the ISP processing mode type corresponding to the current frame.
4. The security monitoring method according to claim 1, characterized in that: The step of inputting the three-dimensional illumination feature vector and the ISP processing mode type into a parameterized regression subnetwork for performing nonlinear mapping calculation and image processing to obtain a first enhanced image includes: Matching the corresponding over-parameterized regression subnetwork based on the ISP processing mode type; Inputting the three-dimensional illumination feature vector into the regional importance analysis module of the parameterized regression subnetwork to perform multi-scale convolution feature extraction to obtain regional adaptive feature data; Inputting the regional adaptive feature data into the time series memory module of the parameterized regression subnetwork to perform inter-frame illumination change trend analysis to generate a time series enhanced feature vector; Inputting the temporal enhancement feature vector into the multi-branch parameter generation network of the parameterized regression subnetwork for parallel nonlinear mapping to obtain an ISP parameter set; Image processing is performed on the security monitoring original image based on the ISP parameter set to obtain a first enhanced image.
5. The security monitoring method according to claim 4, characterized in that: The performing image processing on the security monitoring original image based on the ISP parameter set to obtain a first enhanced image includes: Perform bad pixel detection and neighborhood median filtering correction on the security monitoring original image according to the detection threshold parameter in the ISP parameter set to obtain RGB three-channel basic image data; Performing a linear transformation of pixel brightness on the RGB three-channel basic image data according to the exposure time and gain value in the ISP parameter set to obtain intermediate image data; Apply the color correction matrix in the ISP parameter set to perform matrix multiplication transformation on the intermediate image data to obtain pre-enhanced image data, and perform USM sharpening based on the pre-enhanced image data to obtain a first enhanced image.
6. The security monitoring method according to claim 5, characterized in that: The step of performing penalty calculation on the first enhanced image to obtain a total penalty value, and performing gradient reverse adjustment based on the total penalty value to obtain a second enhanced image includes: Performing brightness histogram statistical analysis and bit-by-bit absolute value accumulation of the difference between the first enhanced image and the reference standard image to obtain a brightness consistency penalty item value; performing color saturation over-enhancement detection and hue shift angle calculation based on the RGB three-channel data of the first enhanced image to obtain a color naturalness penalty item value; and performing gradient amplitude change statistics and detail loss degree quantitative analysis on the first enhanced image to obtain a detail preservation penalty item value; Performing a weighted summation of the brightness consistency penalty item value, the color naturalness penalty item value, and the detail preservation penalty item value to obtain a total penalty value, and comparing the total penalty value with a preset penalty threshold value. When the total penalty value exceeds the preset penalty threshold value, a parameter callback processing mechanism is triggered to obtain a parameter adjustment trigger signal; Based on the parameter adjustment trigger signal, a gradient descent algorithm is started to perform reverse optimization calculation on the ISP parameter set to obtain an optimized ISP parameter set; The optimized ISP parameter set is reapplied to the security monitoring original image to perform ISP pipeline reprocessing calculation to obtain a second enhanced image.
7. The security monitoring method according to claim 1, characterized in that: The performing quality detection on the second enhanced image to obtain a target enhanced image that meets the security monitoring quality standard includes: Performing peak signal-to-noise ratio calculation and maximum-to-minimum brightness ratio calculation on the second enhanced image to obtain a dynamic range evaluation value. Simultaneously, based on Lab color space conversion and color difference formula calculation, a color fidelity evaluation value is obtained. Furthermore, a detail clarity evaluation value is obtained through modulation transfer function calculation and edge sharpness index statistical analysis. Furthermore, a visual quality evaluation value is obtained by using a structural similarity index and a multi-scale SSIM algorithm. Comparing the dynamic range evaluation value with a fourth preset threshold, comparing the color fidelity evaluation value with a fifth preset threshold, comparing the detail clarity evaluation value with a sixth preset threshold, and comparing the visual quality evaluation value with a seventh preset threshold, to obtain a four-dimensional quality compliance status determination result; Based on the four-dimensional quality compliance status determination result, the second enhanced image is format converted and encoded for output to obtain a target enhanced image that meets the security monitoring quality standard.
8. A security monitoring system, characterized in that: Used to perform the security monitoring method according to any one of claims 1 to 7, the security monitoring system comprises: Light detection module, used to perform light detection on the original security monitoring image and obtain a three-dimensional light feature vector; a matching determination module, configured to perform matching determination on preset low-light enhancement mode determination conditions, backlight compensation mode determination conditions, strong light suppression mode determination conditions, and standard processing mode determination conditions according to the three-dimensional illumination feature vector, to obtain the ISP processing mode type corresponding to the current frame; An image processing module, configured to input the three-dimensional illumination feature vector and the ISP processing mode type into a parameterized regression subnetwork for performing nonlinear mapping calculation and image processing to obtain a first enhanced image; a penalty calculation module, configured to perform penalty calculation on the first enhanced image to obtain a total penalty value, and perform gradient reverse adjustment based on the total penalty value to obtain a second enhanced image; The quality detection module is used to perform quality detection on the second enhanced image to obtain a target enhanced image that meets the security monitoring quality standard.
9. A security monitoring device, characterized in that: The security monitoring device includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor calls the instructions in the memory to enable the security monitoring device to execute the security monitoring method according to any one of claims 1 to 7.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the security monitoring method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Anti-intrusion monitoring and early warning method and system
CN121747041A