A low-light video quality measurement system based on deep learning

Through the deep learning-based low-light video quality measurement system, the problem that traditional methods are difficult to accurately evaluate video quality in low-light environments is solved. Accurate measurement of video clarity, noise and color is achieved, and the accuracy and adaptability of low-light video quality assessment are improved.

CN119919780BActive Publication Date: 2025-09-12HUNAN UNIV OF ARTS & SCI
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510322505.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-09-12
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

Traditional video quality measurement methods have difficulty in fully capturing complex visual features in low-light environments and cannot adapt to the diversity of different scenes, resulting in low accuracy in low-light video quality measurement.

Method used

A low-light video quality measurement system based on deep learning is adopted, including a low-light video frame processing module, a video quality dimension measurement and annotation module, a video quality model prediction module and a prediction error loss optimization module. A convolutional neural network is used to construct a video quality measurement model, and video frame quality evaluation is performed through image processing and deep learning.

Benefits of technology

It achieves accurate capture of video clarity, noise, and color reproduction in low-light environments, improves the accuracy and reliability of video quality measurement, provides high-quality annotation data for subsequent analysis, adapts to changes in video quality in different environments, and provides high-precision overall quality assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919780B_ABST
    Figure CN119919780B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of quality measurement technology, and in particular to a low-light video quality measurement system based on deep learning. The system includes a low-light video frame processing module, a video quality dimension measurement and annotation module, a video quality model prediction module, and a prediction error loss optimization module. The system can use a video acquisition device to collect corresponding low-light video data in real time and perform video frame processing, while performing video index measurement and quality dimension measurement annotation to obtain a low-light video quality annotation score; use a convolutional neural network to construct a video quality measurement model architecture and perform model training optimization and quality weighted summation to obtain a low-light video quality prediction score; based on the low-light video quality annotation score and the low-light video quality prediction score, error loss reinforcement learning and video quality assessment measurement are performed to output a corresponding low-light video quality comprehensive measurement score. The present invention can achieve accurate assessment and measurement of low-light video quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of quality measurement technology, and in particular to a low-illumination video quality measurement system based on deep learning. Background Art

[0002] During video capture, low-light environments are common and challenging scenarios. Low light levels can cause video quality issues such as increased noise, reduced contrast, color distortion, and loss of detail, seriously affecting the viewing experience and subsequent analysis and utilization of the video.

[0003] Traditional video quality measurement methods are mainly based on manual features and statistical models, such as pixel-based quality assessment indicators (such as mean square error, peak signal-to-noise ratio) and structure-based assessment indicators (such as structural similarity index). However, these methods have obvious shortcomings when processing low-light videos. Manual features are difficult to fully capture the complex visual characteristics of low-light videos, and statistical models are unable to adapt to the diversity of low-light videos in different scenarios. Moreover, traditional methods often only focus on a single quality variable and cannot achieve a comprehensive and accurate assessment of the overall quality of low-light videos, thereby reducing the accuracy of low-light video quality measurement. Summary of the Invention

[0004] Based on this, it is necessary for the present invention to provide a low-illumination video quality measurement system based on deep learning to solve at least one of the above technical problems.

[0005] To achieve the above objectives, a low-light video quality measurement system based on deep learning is proposed, which includes the following modules:

[0006] A low-light video frame processing module is used to collect corresponding low-light video data in real time by using a video acquisition device in different low-light scenes, and perform video frame processing on the low-light video data to obtain a low-light video standard frame image set;

[0007] A video quality dimension measurement and annotation module is used to measure the video indicators of each frame of video image in the low-light video standard frame image set to obtain the video clarity, video noise level and video color reproduction corresponding to each low-light video frame; based on the video clarity, video noise level and video color reproduction corresponding to each low-light video frame, the corresponding low-light video is measured and annotated in terms of quality dimensions to obtain a low-light video quality annotation score;

[0008] A video quality model prediction module is used to construct a video quality measurement model architecture using a convolutional neural network, and to perform model training and optimization on the video quality measurement model architecture using a set of standard low-light video frame images to generate a low-light video measurement model based on deep learning. The module outputs a prediction score corresponding to each video frame quality dimension, and performs a quality-weighted summation on the corresponding low-light videos based on the prediction score corresponding to each video frame quality dimension to obtain a low-light video quality prediction score.

[0009] The prediction error loss optimization module is used to perform error loss reinforcement learning on the deep learning-based low-light video measurement model based on the low-light video quality annotation score and the low-light video quality prediction score to generate a low-light video measurement comprehensive model; the low-light video standard frame image set is re-input into the low-light video measurement comprehensive model for video quality assessment measurement to output the corresponding low-light video quality comprehensive measurement score.

[0010] Furthermore, the low-light video frame processing module includes the following functions:

[0011] By using video acquisition equipment to collect corresponding low-light video data in real time under different low-light scenes;

[0012] The low-light video data is extracted into single frames at a fixed frame rate of 24 fps to obtain a low-light video single-frame image set;

[0013] Measure the frame image blurriness of a low-illumination video single-frame image set to obtain the low-illumination video frame image blurriness;

[0014] Performing video frame image filtering on each low-illumination video frame in a low-illumination video single frame image set based on the low-illumination video frame image blurriness to obtain a low-illumination video frame filtered image set;

[0015] The pixel values ​​corresponding to each low-illumination video image frame in the low-illumination video frame filtered image set are normalized to obtain a low-illumination video standard frame image set.

[0016] Furthermore, the different low-light scenes are specifically outdoor scenes at night, underground parking lots, and indoor low-light environment scenes.

[0017] Furthermore, the video quality dimension measurement and annotation module includes the following functions:

[0018] Perform multi-channel decomposition on each frame of video image in the low-illumination video standard frame image set to obtain R, G, and B channel images corresponding to each frame of video image; obtain the grayscale mean and grayscale variance corresponding to each channel through the R, G, and B channel images corresponding to each frame of video image, and measure the channel energy ratio of each frame of video image in the low-illumination video standard frame image set based on the grayscale mean and grayscale variance corresponding to each channel to obtain the channel image energy ratio corresponding to each frame of video image; perform video clarity analysis on each corresponding frame of video image according to the channel image energy ratio corresponding to each frame of video image to obtain the video clarity corresponding to each low-illumination video frame;

[0019] Measuring the noise level of each frame of the low-illumination video standard frame image set to obtain the video noise level corresponding to each low-illumination video frame;

[0020] Performing color restoration measurement on each frame of a low-light video standard frame image set to obtain a video color restoration degree corresponding to each low-light video frame;

[0021] Based on the video clarity, video noise level, and video color restoration corresponding to each low-light video frame, a video quality dimension measurement calculation formula is used to measure and calculate the quality dimension of the corresponding low-light video to obtain the low-light video quality dimension;

[0022] Based on the low-light video quality dimension, the corresponding low-light video is labeled with a quality dimension system, with a quality dimension of 0%-20% labeled as 1 point, 21%-40% labeled as 2 points, 41%-60% labeled as 3 points, 61%-80% labeled as 4 points, and 81%-100% labeled as 5 points, so as to obtain the low-light video quality labeling score.

[0023] Furthermore, measuring the noise level of each frame of video image in the low-illumination video standard frame image set includes:

[0024] Performing image local area division on each frame of video image in the low-illumination video standard frame image set to obtain a video local sub-region corresponding to each frame of video image;

[0025] Calculating the texture fractal dimension of the local sub-region of the video corresponding to each frame of the video image to obtain the fractal dimension of the local region corresponding to each frame of the video image;

[0026] Calculating the local pixel variance of the local sub-region of the video corresponding to each frame of the video image to obtain the local region pixel variance corresponding to each frame of the video image;

[0027] The noise level of each frame of video image in the low-illumination video standard frame image set is measured based on the local region fractal dimension and local region pixel variance corresponding to each frame of video image to obtain the video noise level corresponding to each low-illumination video frame.

[0028] Furthermore, the color restoration measurement of each frame of the low-illumination video standard frame image set includes:

[0029] Calculating the chromaticity coordinates of different color regions corresponding to each frame of the video image in the low-light video standard frame image set to obtain the chromaticity coordinates corresponding to the different color regions in each frame of the video image;

[0030] Based on the chromaticity coordinates corresponding to the different color areas in each frame of the video image, the color saturation of the corresponding different color areas in each frame of the video image in the low-light video standard frame image set is measured to obtain the color saturation corresponding to the different color areas in each frame of the video image;

[0031] Performing hue distribution statistics on different color regions corresponding to each frame of the video image in the low-light video standard frame image set to obtain the hue distribution corresponding to the different color regions in each frame of the video image;

[0032] The color restoration degree of each video frame in the low-light video standard frame image set is measured based on the color saturation and hue distribution corresponding to different color areas in each video frame to obtain the video color restoration degree corresponding to each low-light video frame.

[0033] Furthermore, the video quality dimension measurement calculation formula is specifically as follows:

[0034] ;

[0035] Where, is the low-light video quality dimension, is the total number of low-light video frames, For the Low-light video frames, For the The video clarity corresponding to the low-light video frame, For the The video noise level corresponding to a low-light video frame, For the The video color restoration degree corresponding to the low-light video frame, is the video frame time interval, specifically from arrive timeframe, For in time The corresponding low-light video frame is as follows: is the time frame parameter, is the reference video frame, is the noise penalty factor.

[0036] Furthermore, the video quality model prediction module includes the following functions:

[0037] A video quality measurement model architecture is constructed using a convolutional neural network, which includes a sharp feature extraction branch, a noise feature extraction branch, and a fully connected neural network layer.

[0038] The low-light video standard frame image set is divided into a training set, a validation set, and a test set according to the division ratio of 7:2:1;

[0039] The training set is fed into the video quality measurement model architecture for model training. Stochastic gradient descent is used as an optimization algorithm in conjunction with the validation set to continuously adjust the model's corresponding hyperparameters, including the learning rate, batch size, and number of network layers. The test set is also used to optimize model performance. This generates a deep learning-based low-light video measurement model, which outputs a predicted score for each video frame quality dimension at each fully connected neural network layer.

[0040] Based on the prediction score corresponding to each video frame quality dimension, the quality-weighted sum of the corresponding low-light video is performed to obtain the low-light video quality prediction score.

[0041] Furthermore, the video quality measurement model architecture is specifically composed of a clear feature extraction branch, a noise feature extraction branch and multiple fully connected neural network layers, wherein the clear feature extraction branch is composed of 5 5x5 convolutional layers and 5 1x1 pooling layers, and the noise feature extraction branch is composed of 5 3x3 convolutional layers and 5 1x1 pooling layers. The fully connected neural network layer is determined by the total number corresponding to the quality dimension, so as to weightedly fuse the video frame features extracted by each branch into a comprehensive feature vector, and input the comprehensive feature vector into each fully connected neural network layer for nonlinear transformation to output the prediction score corresponding to each quality dimension.

[0042] Furthermore, the prediction error loss optimization module includes the following functions:

[0043] The quality mean square error loss is calculated between the low-light video quality annotation score and the low-light video quality prediction score to obtain the video quality error loss between the model prediction score and the annotation score;

[0044] Based on the video quality error loss between the model prediction score and the annotated score, the deep learning-based low-light video measurement model is subjected to error loss reinforcement learning. The corresponding hyperparameters of the low-light video measurement model are constrained from overfitting by introducing an L2 regularization term to generate a comprehensive low-light video measurement model.

[0045] The low-light video standard frame image set is re-input into the low-light video measurement comprehensive model for video quality assessment measurement to output the corresponding low-light video quality comprehensive measurement score.

[0046] Beneficial effects of the present invention:

[0047] The low-light video quality measurement system based on deep learning proposed in the present invention is composed of a low-light video frame processing module, a video quality dimension measurement and annotation module, a video quality model prediction module and a prediction error loss optimization module. Compared with the existing technology, the beneficial effect of the present application is that by using a video acquisition device to collect low-light video data in different low-light scenes, the clarity, noise and color reproduction of the video data can be accurately captured. In a low-light environment, the video acquisition device needs to have higher sensitivity and performance to cope with the noise and detail loss problems caused by insufficient light. The video frame processing technology further optimizes the quality of each frame image. By denoising, enhancing contrast, adjusting exposure, etc., the quality of low-light video can be made more stable and controllable. This processing process not only makes the video image clearer, but also improves the reliability and accuracy of the video data in subsequent analysis, thereby ensuring the accuracy and effectiveness of subsequent measurement and analysis. Secondly, by measuring the video indicators of each frame of low-light video image to obtain video clarity, noise level and color reproduction, this process is an important part of evaluating video quality. By performing in-depth image analysis on each frame of low-light video, the performance of the video image in different quality dimensions can be quantified. The measurement of video clarity helps to judge the degree of detail and blur in the image; the noise level measurement evaluates the degree of noise interference in the image. Excessive noise will affect the viewing experience of the video; the color reproduction measurement judges the accuracy and reproduction of the color in the video image. Color distortion will greatly affect the realism and visual effects of the video. Through these measurement indicators, a quality labeling score can be given to each video frame. The key to this step is that through comprehensive indicator analysis, it can not only accurately reflect the quality characteristics of each frame, but also provide real and high-quality labeled data for subsequent deep learning model training, effectively supporting the precise learning and optimization of the model. Then, by constructing a video quality measurement model based on convolutional neural network (CNN) and using a set of standard low-light video frame images for model training and optimization, the deep learning model can automatically extract complex feature information in the video frames, such as image clarity, noise, color and other quality indicators through self-learning and optimization. The convolutional neural network has powerful image processing capabilities and can automatically identify and analyze subtle differences and details in videos under low-light environments. It is particularly suitable for processing noise and blur problems in low-light videos. Through multi-layer network structure and a large amount of low-light video data training, the deep learning model can learn the complex relationship between image quality and video frames, and give an accurate quality prediction score. By weighted prediction of the quality dimension of each video frame, the low-light video quality prediction score output by the model can provide an accurate basis for the final video quality evaluation, thereby achieving a comprehensive and accurate evaluation of the overall quality of the low-light video.Finally, by performing error loss reinforcement learning based on the quality annotation scores of low-light videos and the model prediction scores, the prediction accuracy and robustness of the model can be significantly improved through the reinforcement learning method. The error loss function plays an important role in deep learning and can guide the model to adjust according to the difference between the predicted results and the true annotations. Through reinforcement learning, the model can better adapt to changes in different environments when facing low-light videos, reduce overfitting, and improve the model's adaptability to complex scenes. The comprehensive model after reinforcement learning can not only provide accurate quality prediction scores for each video frame, but also combine multiple quality dimensions for comprehensive evaluation, thereby obtaining the overall quality score of the low-light video, thereby ensuring that the final model can provide high-precision and reliable video quality measurement results in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments thereof made with reference to the following drawings:

[0049] Figure 1 Schematic diagram of the modules of the low-light video quality measurement system based on deep learning of the present invention;

[0050] Figure 2 for Figure 1 Schematic diagram of the functional flow of the low-light video frame processing module;

[0051] Figure 3 for Figure 1 Schematic diagram of the functional flow of the video quality dimension measurement and annotation module. DETAILED DESCRIPTION

[0052] The following is a clear and complete description of the technical system of the present invention in conjunction with the accompanying drawings. It is obvious that the embodiments described are part of the embodiments of the present invention, but not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of the present invention.

[0053] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor systems and / or microcontroller systems.

[0054] It should be understood that although the terms "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. The term "and / or" as used herein includes any and all combinations of one or more of the listed associated items.

[0055] To achieve this, please refer to Figures 1 to 3 The present invention provides a low-light video quality measurement system based on deep learning, which includes the following modules:

[0056] A low-light video frame processing module is used to collect corresponding low-light video data in real time by using a video acquisition device in different low-light scenes, and perform video frame processing on the low-light video data to obtain a low-light video standard frame image set;

[0057] A video quality dimension measurement and annotation module is used to measure the video indicators of each frame of video image in the low-light video standard frame image set to obtain the video clarity, video noise level and video color reproduction corresponding to each low-light video frame; based on the video clarity, video noise level and video color reproduction corresponding to each low-light video frame, the corresponding low-light video is measured and annotated in terms of quality dimensions to obtain a low-light video quality annotation score;

[0058] A video quality model prediction module is used to construct a video quality measurement model architecture using a convolutional neural network, and to perform model training and optimization on the video quality measurement model architecture using a set of standard low-light video frame images to generate a low-light video measurement model based on deep learning. The module outputs a prediction score corresponding to each video frame quality dimension, and performs a quality-weighted summation on the corresponding low-light videos based on the prediction score corresponding to each video frame quality dimension to obtain a low-light video quality prediction score.

[0059] The prediction error loss optimization module is used to perform error loss reinforcement learning on the deep learning-based low-light video measurement model based on the low-light video quality annotation score and the low-light video quality prediction score to generate a low-light video measurement comprehensive model; the low-light video standard frame image set is re-input into the low-light video measurement comprehensive model for video quality assessment measurement to output the corresponding low-light video quality comprehensive measurement score.

[0060] In the embodiment of the present invention, please refer to Figure 1FIG. 1 is a schematic diagram of a module of a low-light video quality measurement system based on deep learning according to the present invention. In this example, the low-light video quality measurement system based on deep learning includes the following modules:

[0061] S1: Low-illumination video frame processing module, used to collect corresponding low-illumination video data in real time by using video acquisition equipment in different low-illumination scenes, and perform video frame processing on the low-illumination video data to obtain a low-illumination video standard frame image set;

[0062] In an embodiment of the present invention, when performing video capture in a low-light environment, first select a high-sensitivity video capture device suitable for working under low-light conditions, such as a high-performance CMOS sensor or an image sensor based on back-illuminated technology. The capture device should be able to perform stable video recording in a low-light environment to avoid video blur or motion blur caused by long exposure time. In different low-light scenes, the capture device continues to record video to ensure that sufficient low-light video data is captured. The ambient light conditions during the video capture process should be fully considered to ensure data diversity, especially at night or in a dark indoor environment. The captured data should cover different low-light levels. Next, use image processing software to perform frame processing on the recorded video, extract each frame of image data, and retain representative standard frame images by removing redundant or damaged frames to ultimately form a low-light video standard frame image set.

[0063] S2: Video quality dimension measurement and annotation module, used to measure video indicators of each frame of video image in the low-light video standard frame image set to obtain the video clarity, video noise level and video color reproduction corresponding to each low-light video frame; based on the video clarity, video noise level and video color reproduction corresponding to each low-light video frame, perform quality dimension measurement and annotation on the corresponding low-light video to obtain a low-light video quality annotation score;

[0064] In an embodiment of the present invention, after acquiring a set of standard low-light video frames, video metrics are measured for each frame. First, a clarity measurement algorithm is used to analyze each frame. Common clarity evaluation metrics include edge strength, contrast, and texture features. By calculating the values ​​of these features, a clarity score for each frame can be obtained. Next, a noise detection method is used to assess the noise level of each frame. For example, image gradient analysis or wavelet transform methods are used to extract image noise features. The noise level measurement result can be reflected by calculating the signal-to-noise ratio (SNR) value of the image. In addition, the color reproduction of the video image can be measured by comparing it with a standard color chart. A color difference measurement method, such as CIEDE2000 or the Delta E formula, is used to compare the deviation between the video frame and the standard color model to obtain a color reproduction score. By combining the clarity, noise level, and color reproduction of each frame for a comprehensive analysis, a corresponding low-light video quality annotation score is generated for each frame, ultimately obtaining a low-light video quality annotation score.

[0065] S3: Video quality model prediction module, used to construct a video quality measurement model architecture using a convolutional neural network, and optimize the video quality measurement model architecture using a low-light video standard frame image set to generate a deep learning-based low-light video measurement model and output a prediction score corresponding to each video frame quality dimension. Based on the prediction score corresponding to each video frame quality dimension, a quality-weighted summation is performed on the corresponding low-light video to obtain a low-light video quality prediction score.

[0066] In an embodiment of the present invention, a framework for a low-light video quality measurement model is constructed based on deep learning technology. A convolutional neural network (CNN) model is first designed. This model can extract spatial features from video images through multiple convolutional layers and predict quality indicators through fully connected layers. To measure the quality of low-light videos, the CNN input is a set of standard low-light video frames. Each frame undergoes preprocessing, such as grayscale conversion and normalization, before being input into the network. The network output is a predicted score for each frame's clarity, noise level, and color reproduction. During network training, the model parameters are optimized using a backpropagation algorithm to minimize the error between the predicted values ​​and the true labeled scores. To ensure good generalization of the model, data augmentation techniques, such as image rotation, translation, and scaling, are used to expand the training dataset and reduce overfitting. By continuously adjusting the model structure and hyperparameters during training, the deep learning-based low-light video measurement model can output a predicted score corresponding to each video frame quality dimension. A quality-weighted summation is performed on the entire low-light video to ultimately obtain an overall quality prediction score for the low-light video.

[0067] S4: Prediction error loss optimization module, used to perform error loss reinforcement learning on the deep learning-based low-light video measurement model based on the low-light video quality annotation score and the low-light video quality prediction score to generate a low-light video measurement comprehensive model; re-input the low-light video standard frame image set into the low-light video measurement comprehensive model for video quality assessment measurement to output the corresponding low-light video quality comprehensive measurement score.

[0068] In an embodiment of the present invention, after obtaining the low-light video quality annotation score and the low-light video quality prediction score, the low-light video measurement model based on deep learning is further optimized by using the error loss reinforcement learning technology. The core of the error loss reinforcement learning process is to use the low-light video quality annotation score as a supervisory signal to dynamically adjust the difference between the low-light video quality prediction score and the annotation score. First, based on the difference between the prediction score and the annotation score output by the current model, the loss function value is calculated. The loss function value reflects the prediction error of the model in each quality dimension (clarity, noise level, color reproduction). Then, a strong The learning algorithm is optimized, and training is focused on areas with large errors. The model weights are adjusted to increase the model's learning ability for frames with large errors, thereby achieving fine optimization of video quality. This process is carried out in continuous iterations until the model can predict the quality of low-light videos with higher accuracy. Finally, the low-light video standard frame image set is re-input into the optimized low-light video measurement comprehensive model for final video quality evaluation and measurement. The final measurement outputs a comprehensive low-light video quality measurement score, which comprehensively reflects the comprehensive quality of the low-light video in terms of clarity, noise level and color reproduction, and can provide a basis for video quality optimization or enhancement processing.

[0069] Furthermore, the low-light video frame processing module includes the following functions:

[0070] By using video acquisition equipment in different low-light scenes to collect corresponding low-light video data in real time;

[0071] The low-light video data is extracted into single frames at a fixed frame rate of 24 fps to obtain a low-light video single-frame image set;

[0072] Frame image blurriness is measured for a single-frame image set of a low-illumination video to obtain the frame image blurriness of the low-illumination video;

[0073] Performing video frame image filtering on each low-illumination video frame in a low-illumination video single frame image set based on the low-illumination video frame image blurriness to obtain a low-illumination video frame filtered image set;

[0074] The pixel values ​​corresponding to each low-illumination video image frame in the low-illumination video frame filtered image set are normalized to obtain a low-illumination video standard frame image set.

[0075] As an embodiment of the present invention, refer to Figure 2 As shown, Figure 1 Schematic diagram of the functional flow of the low-light video frame processing module. In this embodiment, the low-light video frame processing module includes the following functions:

[0076] S11: using a video capture device to collect corresponding low-light video data in real time under different low-light scenes;

[0077] In an embodiment of the present invention, low-light video data is collected by dedicated video acquisition equipment (such as night vision cameras, low-light cameras, etc.) in different low-light environments. In outdoor scenes at night, different time periods (such as 22:00 in the evening to 2:00 in the morning) are selected for collection to ensure that the collected scenes are representative and low-light. Collection in underground parking lots should be selected in areas with different light intensities, and low-light parking lot images are obtained through video acquisition equipment, especially in places far away from lighting equipment. In indoor low-light environments, the acquisition equipment can be placed in insufficiently lit rooms (such as windowless or dimly lit rooms) to obtain low-light video data. In each scenario, the acquisition equipment needs to be able to maintain a stable frame rate and clarity to ensure that the data quality in different scenarios is consistent, and can cover typical situations in low-light scenarios, and finally obtain corresponding low-light video data.

[0078] S12: extracting single frames from the low-illumination video data at a fixed frame rate of 24 fps to obtain a low-illumination video single-frame image set;

[0079] In an embodiment of the present invention, after the low-light video data is collected, the collected video is processed using a video processing tool to ensure that the frame rate of the video is 24fps (24 frames per second). A separate frame image is extracted every second through an automated script and saved. For each low-light video, all frame data is extracted to obtain a single-frame image set of the low-light video. This process can be done by writing a script in Python to call a video processing library (such as OpenCV), loading the video file through the cv2.VideoCapture method in the code, and using cv2.imwrite to save each frame image as a separate file for subsequent analysis and processing, ultimately obtaining a single-frame image set of the low-light video.

[0080] S13: measuring the frame image blurriness of the low-illumination video single-frame image set to obtain the low-illumination video frame image blurriness;

[0081] In an embodiment of the present invention, blurriness is measured for each frame in a set of single-frame images of a low-illumination video to analyze whether the video frame is blurred. The blurriness measurement can adopt an algorithm based on image gradient, such as the Laplacian operator (Laplacian) or an algorithm based on edge detection. In specific implementation, the cv2.Laplacian() function in OpenCV can be used to calculate the gradient of each frame image, and then obtain the blurriness value of the image. A threshold is set to determine whether the image is a blurred image. When the blurriness is greater than a set threshold, the image is considered to have strong blur. In specific implementation, an automated processing code can be written to calculate the blurriness of each frame image and record the blurriness of each frame image to finally obtain the blurriness of the low-illumination video frame image.

[0082] S14: performing video frame image filtering on each low-illumination video frame in the low-illumination video single frame image set based on the low-illumination video frame image blurriness to obtain a low-illumination video frame filtered image set;

[0083] In an embodiment of the present invention, filtering is performed on each frame image based on a calculation result of the blurriness of a low-illumination video frame image to reduce blurring. The filtering operation may adopt image filtering algorithms such as Gaussian Blur and Median Filter. In specific implementation, the cv2.GaussianBlur() function in OpenCV may be called in the code to process images with higher blurriness, and the size of the Gaussian kernel may be adjusted to adapt to different degrees of blurring. In the filtering process, a stronger filtering process is adopted for frame images with higher blurriness; a weaker filtering intensity is used for frame images with lower blurriness. The obtained low-illumination video frame filtered image set will include the processed image, reduce the blurriness, thereby improving the clarity of the image, and finally obtain the low-illumination video frame filtered image set.

[0084] S15: performing video frame normalization on the pixel values ​​corresponding to each low-illumination video image frame in the low-illumination video frame filtered image set to obtain a low-illumination video standard frame image set.

[0085] In an embodiment of the present invention, based on a set of low-light video frame filtered images, the pixel values ​​of each frame are normalized to ensure that the pixel value range of the image is unified within a standard interval (usually 0 to 255). The image normalization process can be performed by analyzing the pixel distribution of each frame and performing linear or nonlinear normalization on it. In a specific implementation, the maximum pixel value and the minimum pixel value of each frame are first calculated, and the pixel values ​​are normalized using a formula to map their values ​​to the standard range. For example, the following formula is used for normalization: ,in, is the original image pixel, and are the minimum and maximum values ​​of the image, respectively. This step ensures that the visual characteristics of each frame, such as brightness and contrast, are within the same standard range, which facilitates subsequent deep learning model analysis and processing, and ultimately generates a low-light video standard frame image set with high image quality and consistency.

[0086] Furthermore, the different low-light scenes are specifically outdoor scenes at night, underground parking lots, and indoor low-light environment scenes.

[0087] Furthermore, the video quality dimension measurement and annotation module includes the following functions:

[0088] Perform multi-channel decomposition on each frame of video image in the low-illumination video standard frame image set to obtain R, G, and B channel images corresponding to each frame of video image; obtain the grayscale mean and grayscale variance corresponding to each channel through the R, G, and B channel images corresponding to each frame of video image, and measure the channel energy ratio of each frame of video image in the low-illumination video standard frame image set based on the grayscale mean and grayscale variance corresponding to each channel to obtain the channel image energy ratio corresponding to each frame of video image; perform video clarity analysis on each corresponding frame of video image according to the channel image energy ratio corresponding to each frame of video image to obtain the video clarity corresponding to each low-illumination video frame;

[0089] Measuring the noise level of each frame of the low-illumination video standard frame image set to obtain the video noise level corresponding to each low-illumination video frame;

[0090] Performing color restoration measurement on each frame of the low-light video standard frame image set to obtain the video color restoration corresponding to each low-light video frame;

[0091] Based on the video clarity, video noise level, and video color restoration corresponding to each low-light video frame, a video quality dimension measurement calculation formula is used to measure and calculate the quality dimension of the corresponding low-light video to obtain the low-light video quality dimension;

[0092] Based on the low-light video quality dimension, the corresponding low-light video is labeled with a quality dimension system, with a quality dimension of 0%-20% labeled as 1 point, 21%-40% labeled as 2 points, 41%-60% labeled as 3 points, 61%-80% labeled as 4 points, and 81%-100% labeled as 5 points, so as to obtain the low-light video quality labeling score.

[0093] As an embodiment of the present invention, refer to Figure 2 As shown, Figure 1Schematic diagram of the functional flow of the video quality dimension measurement and annotation module. In this embodiment, the video quality dimension measurement and annotation module includes the following functions:

[0094] S21: performing multi-channel decomposition on each frame of video image in the low-illumination video standard frame image set to obtain R, G, and B channel images corresponding to each frame of video image; obtaining the grayscale mean and grayscale variance corresponding to each channel through the R, G, and B channel images corresponding to each frame of video image, and measuring the channel energy ratio of each frame of video image in the low-illumination video standard frame image set based on the grayscale mean and grayscale variance corresponding to each channel to obtain the channel image energy ratio corresponding to each frame of video image; performing video clarity analysis on each corresponding frame of video image according to the channel image energy ratio corresponding to each frame of video image to obtain the video clarity corresponding to each low-illumination video frame;

[0095] In an embodiment of the present invention, each frame in a low-light video standard frame image set is split into RGB (red, green, and blue) channels. In a specific implementation, a computer vision library, such as OpenCV or PIL (Python Imaging Library), is used to perform color space separation on each frame of the image. The RGB information of each frame of the image can be obtained by separating each color channel of the image, where the R channel contains the pixel value of the red channel, the G channel contains the pixel value of the green channel, and the B channel contains the pixel value of the blue channel. The result of the splitting is three single-channel grayscale images, corresponding to the red, green, and blue channel images respectively. In each single-channel image, the pixel value reflects the brightness information of the channel, and each pixel value range of the R, G, and B channels is usually 0 to 255. This process is processed for all video frames to obtain the RGB channel image corresponding to each frame. At the same time, by calculating the grayscale mean and grayscale variance of the R, G, and B channel images of each video frame, the grayscale mean reflects the brightness concentration of the image on the channel, while the grayscale variance measures whether the brightness distribution of the image is uniform. In specific implementation, for the image of each channel, first extract the values ​​of all pixels, calculate the mean (that is, the sum of all pixel values ​​divided by the total number of pixels) and variance (that is, the average of the square of the difference between each pixel value and the mean) of the channel, and then calculate the energy share of the channel based on the mean and variance of each channel. The energy share calculation formula can be performed by standardizing the relationship between the variance and the mean of each channel, thereby obtaining the energy distribution of each frame of the image on each color channel. Finally, the energy share reflects the different color channels in the video frame. In addition, the clarity of each frame of video image is analyzed mainly based on the aforementioned channel energy ratio results. In specific implementation, a weighted clarity evaluation model can be constructed based on the channel energy ratio. This model combines the contribution of different channels in the video image. First, by combining the energy ratio of each channel with the global clarity index of the image (such as image gradient, edge information, etc.), the clarity score of each frame of the image is calculated. Through these calculations, the clarity of each frame of the low-light video can be analyzed to evaluate the clarity level of each frame of the image. The results of the clarity analysis help to understand the visual sharpness and detail presentation of the low-light video image, and finally the video clarity corresponding to each low-light video frame is obtained.

[0096] S22: measuring the noise level of each frame of the low-illumination video standard frame image set to obtain a video noise level corresponding to each low-illumination video frame;

[0097] In an embodiment of the present invention, the noise level of each frame of video image is measured. The calculation of noise generally involves comparing the difference between the real signal and the noise in the image. First, the background information of the image needs to be extracted from each frame of image. Assuming that the background information is the part with less noise, the image can be smoothed by using image filtering technology (such as Gaussian blur) and the difference between the original image and the smoothed image is calculated to obtain the noise. The difference is the noise information. Subsequently, the standard deviation, mean and energy value of the noise can be calculated to measure the noise level of each frame of video image. The higher the noise level, the worse the image quality, which affects the clarity and visual effect of the image. By quantifying the noise level, the noise problem caused by insufficient light in low-light video can be effectively evaluated, and finally the video noise level corresponding to each low-light video frame is obtained.

[0098] S23: measuring the color restoration degree of each frame of the low-illumination video standard frame image set to obtain the video color restoration degree corresponding to each low-illumination video frame;

[0099] In an embodiment of the present invention, the color reproduction degree of each video frame is measured to evaluate the color reproduction accuracy in low-light videos. First, the R, G, and B channels of each frame are extracted and compared with an ideal standard color image. The standard color image can be obtained from a set of standard test images (such as a standard color chart or an artificially synthesized image). Using an image color difference metric (such as the CIEDE2000 color difference formula), the color difference between each color channel of each frame and the corresponding channel of the standard image is calculated. The smaller the color difference, the higher the color reproduction degree of the video image. By calculating the color difference of each frame of the video image and weighting it, a color reproduction score for each frame is obtained, and then the color accuracy of the low-light video is evaluated. The color reproduction degree of each frame of the video is determined based on the color difference score, and ultimately the video color reproduction degree corresponding to each low-light video frame is obtained.

[0100] S24: performing quality dimension measurement calculation on the corresponding low-illumination video based on the video clarity, video noise level, and video color restoration corresponding to each low-illumination video frame using a video quality dimension measurement calculation formula to obtain a low-illumination video quality dimension;

[0101] In an embodiment of the present invention, a suitable video quality dimension measurement calculation formula is formed by combining the total number of low-illumination video frames, the video clarity corresponding to the low-illumination video frames, the video noise level, the video color restoration, the video frame time interval, the time frame parameters, the reference video frame, the noise penalty factor and related parameters to perform quality dimension measurement calculation to quantitatively calculate the quality dimension of each frame of video. In order to improve the calculation accuracy, a deep learning model can be used to automatically evaluate the quality of each video frame, and finally the low-illumination video quality dimension is obtained.

[0102] S25: Based on the low-light video quality dimension, the corresponding low-light video is labeled with a quality dimension system, so that a quality dimension of 0%-20% is labeled as 1 point, 21%-40% is labeled as 2 points, 41%-60% is labeled as 3 points, 61%-80% is labeled as 4 points, and 81%-100% is labeled as 5 points, so as to obtain the low-light video quality labeling score.

[0103] In an embodiment of the present invention, by implementing video quality score annotation according to the quality dimension of low-light video, first, ensure that the calculation result of the quality dimension is within the range of 0%-100%, and then formulate score annotation rules according to the numerical range of the quality dimension. For videos with a quality dimension between 0%-20%, it is marked as 1 point; between 21%-40%, it is marked as 2 points; between 41%-60%, it is marked as 3 points; between 61%-80%, it is marked as 4 points; and between 81%-100%, it is marked as 5 points. In specific implementation, according to According to the obtained video quality dimension value, it is compared with the preset score range. For example, for a frame of video, its quality dimension is 35%, then the video frame can be assigned 2 points according to the labeling rules. In order to ensure the efficiency of the labeling process, an automated scoring system can be developed to automatically assign scores to each frame of video based on the calculation results of the quality dimension, thereby completing the quality labeling of low-light videos. The automated system can be trained using deep learning technology to make the scoring process more accurate and faster, and ultimately obtain the low-light video quality labeling score.

[0104] Furthermore, measuring the noise level of each frame of video image in the low-illumination video standard frame image set includes:

[0105] Performing image local area division on each frame of video image in the low-illumination video standard frame image set to obtain a video local sub-region corresponding to each frame of video image;

[0106] In an embodiment of the present invention, each frame of video image in a low-illumination video standard frame image set is divided into a local area. When implementing this step, each frame of image needs to be divided into several small areas, each of which is called a local area. The specific operation is to use a fixed-size sliding window to scan the video frame and divide the video frame into several sub-areas that are uniform or based on image content features. For each frame of image, the size and step size of the division window are first determined. Generally, common image block sizes such as 16×16 and 32×32 can be selected. The image is divided into multiple overlapping or non-overlapping small blocks by means of a sliding window to ensure that each local area can effectively represent a part of the image. During the division process, the size of each local area can be adjusted according to the spatial resolution of the image and the demand for computing resources. This operation provides basic data support for subsequent texture analysis and noise level calculation, and ultimately the video local sub-area corresponding to each frame of video image is obtained.

[0107] Preferably, the texture fractal dimension of the local sub-region of the video corresponding to each frame of the video image is calculated to obtain the fractal dimension of the local region corresponding to each frame of the video image;

[0108] In an embodiment of the present invention, by extracting texture features from each local area, the box-counting method can be used to calculate the fractal dimension of each local area through fractal theory. Specifically, the core idea of ​​the box counting method is to cover the local area with grids (boxes) of different sizes, and then calculate the minimum number of grids that can cover the local area as the grid size changes. In this way, the complexity of the local area at different scales can be obtained. Finally, the fractal dimension of the local area is obtained by logarithmic regression analysis. For low-light video, the details of the texture are affected by noise. Therefore, the fractal dimension calculated in this way can reflect the complexity of the image texture. The fractal dimension of each local area is closely related to the noise level of the area and the structural characteristics of the image, and finally the fractal dimension of the local area corresponding to each frame of video image is obtained.

[0109] Preferably, a local pixel variance calculation is performed on a local sub-region of the video corresponding to each frame of the video image to obtain a local region pixel variance corresponding to each frame of the video image;

[0110] In an embodiment of the present invention, local pixel variance is calculated for a local sub-region of the video corresponding to each frame of video image. When implementing this step, for each local region, all pixel values ​​in the region are first extracted, and then the variance of the pixel values ​​in the local region is calculated. The specific calculation formula is: local variance = Σ((x_i -μ)^2) / N, where x_i is each pixel value in the region, μ is the average pixel value of the local region, and N is the total number of pixels in the region. The variance can reflect the degree of change of the pixel values ​​in the local region. Therefore, a large variance means that the pixel values ​​in the region change more dramatically and there is a high noise level. By calculating the variance of each local region, the pixel distribution of each region in the video frame can be effectively evaluated, and the degree of noise influence can be inferred, and finally the pixel variance of the local region corresponding to each frame of video image is obtained.

[0111] Preferably, the noise level of each frame of video image in the low-illumination video standard frame image set is measured based on the local region fractal dimension and local region pixel variance corresponding to each frame of video image to obtain the video noise level corresponding to each low-illumination video frame.

[0112] In an embodiment of the present invention, the noise level of each frame of video image in the low-illumination video standard frame image set is measured based on the local region fractal dimension and local region pixel variance corresponding to each frame of video image. When implementing this step, the noise level measurement combines the fractal dimension and pixel variance obtained in the previous two steps. Specifically, using the fractal dimension and variance information of each local region, the features of all local regions are first weighted and summarized to obtain comprehensive features of each frame of video image. These comprehensive features can effectively describe the overall noise level of the image. Then, by establishing a mapping relationship between the noise level and the local region features, regression analysis or machine learning models (such as support vector machines, random forests, etc.) are used to predict the noise level. Images with high noise levels generally have lower fractal dimensions and higher pixel variances, while images with low noise levels have the opposite. Through this method, the noise level of each frame of video image can be accurately evaluated, and ultimately the video noise level corresponding to each low-illumination video frame is obtained.

[0113] Furthermore, the color restoration measurement of each frame of the low-light video standard frame image set includes:

[0114] Calculating the chromaticity coordinates of different color regions corresponding to each frame of the video image in the low-light video standard frame image set to obtain the chromaticity coordinates corresponding to the different color regions in each frame of the video image;

[0115] In an embodiment of the present invention, a low-light video standard frame image set consists of multiple frames of video images, each of which contains multiple different color regions. For each frame of video image, a color segmentation algorithm (such as a threshold-based segmentation method) is first used to divide the image into different color regions. For each pixel in each color region, it is converted from the RGB color space to the CIE XYZ color space. The conversion formula is: ,in is the RGB value of the pixel, which ranges from [0,255]. Then, the chromaticity coordinates (x, y) are calculated based on the CIE XYZ value. The calculation formula is , After calculating the chromaticity coordinates of all pixels in each color area, the average value is taken as the chromaticity coordinate corresponding to the color area. Through this operation, the chromaticity coordinates corresponding to different color areas in each frame of video image are finally obtained.

[0116] Preferably, color saturation of corresponding different color areas in each frame of video image in the low-illumination video standard frame image set is measured based on the chromaticity coordinates corresponding to different color areas in each frame of video image, so as to obtain the color saturation corresponding to different color areas in each frame of video image;

[0117] In an embodiment of the present invention, after obtaining the chromaticity coordinates corresponding to different color regions in each frame of video image, color saturation measurement is performed. First, the chromaticity coordinates (x, y) are converted back to the CIE XYZ color space, and then converted from the CIE XYZ color space to the HSV color space. After conversion to the HSV color space, for each pixel in the color region, its saturation value S is extracted. The saturation S reflects the vividness of the color and has a value range of [0, 1]. For each color region, the average saturation value of all pixels in the region is calculated and used as the color saturation corresponding to the color region. Specifically, suppose there are n pixels in a certain color region, and their saturation values ​​are respectively , then the color saturation of the color area By performing such calculations on different color areas of each frame of video image in the low-illumination video standard frame image set, the color saturation corresponding to different color areas in each frame of video image is finally obtained.

[0118] Preferably, hue distribution statistics are performed on the different color regions corresponding to each frame of the video image in the low-illumination video standard frame image set to obtain the hue distribution corresponding to the different color regions in each frame of the video image;

[0119] In an embodiment of the present invention, after converting the color area of ​​each frame of video image from RGB to HSV color space, the hue value H is focused on. The hue value H represents the type of color and its value range is usually [0,360). For each color area in each frame of video image, its hue value is divided into several intervals, for example, divided into 10 intervals, each interval is 36°, and the number of pixel points in each color area whose hue value falls in each interval is counted. Suppose there are n pixels in a certain color area, and its hue value is counted into k intervals. The number of pixels falling in the jth interval is , then the frequency of this interval The frequency distribution of these intervals constitutes the hue distribution corresponding to the color area. By performing such statistics on the different color areas of each frame of video image in the low-light video standard frame image set, the hue distribution corresponding to the different color areas in each frame of video image is finally obtained.

[0120] Preferably, the color restoration degree of each frame of video image in the low-light video standard frame image set is measured based on the color saturation and hue distribution corresponding to different color areas in each frame of video image to obtain the video color restoration degree corresponding to each low-light video frame.

[0121] In the embodiment of the present invention, for each frame of the low-light video standard frame image set, the color saturation and hue distribution corresponding to different color regions are known. First, weights are set for the color saturation and hue distribution respectively. and ,and For each color region, the difference between the ideal value and the actual measured value of its color saturation is calculated. The difference can be expressed by the absolute value of the difference between the two. At the same time, the difference between the ideal distribution and the actual distribution of its hue distribution is calculated. The Kullback-Leibler divergence can be used to measure the difference between the two distributions. Suppose the ideal color saturation value of a color region is , the actual value is , saturation difference ; The ideal distribution of hue distribution is , the actual distribution is , color tone distribution difference , the color restoration degree of all color areas in each frame of video image is weighted averaged, and the weight can be determined according to the size of the color area , and finally obtain the video color restoration degree corresponding to each low-illumination video frame.

[0122] Furthermore, the video quality dimension measurement calculation formula is specifically as follows:

[0123] ;

[0124] Where, is the low-light video quality dimension, is the total number of low-light video frames, For the Low-light video frames, For the The video clarity corresponding to the low-light video frame, For the The video noise level corresponding to a low-light video frame, For the The video color restoration degree corresponding to the low-light video frame, is the video frame time interval, specifically from arrive timeframe, For in time The corresponding low-light video frame is as follows: is the time frame parameter, is the reference video frame, is the noise penalty factor.

[0125] The present invention obtains a video quality dimension measurement calculation formula by using a specific mathematical model and verification, which is used to measure the quality dimension of the corresponding low-light video. The formula fully considers the low-light video quality dimension. , the total number of low-light video frames , No. Low-light video frames , No. The video clarity corresponding to the low-light video frame , No. The video noise level corresponding to low-light video frames , No. The video color restoration degree corresponding to the low-light video frame , video frame time interval , specifically from arrive Time range, in time The corresponding low-light video frame , time frame parameters , reference video frame , noise penalty factor , according to the low-light video quality dimension The mutual correlation between the above parameters constitutes a functional relationship This formula measures the quality of low-light videos. It also considers multiple quality dimensions, including clarity, noise level, and color reproduction, which are key factors in measuring low-light video quality. The impact of each dimension is factored into the quality assessment, making the results more comprehensive and accurate. By incorporating time intervals and inter-frame differences into the formula, it effectively captures temporal trends in the video. This means that video quality isn't just assessed on a single frame, but rather measures the temporal variation of the entire video sequence, helping to identify the overall quality of the video. The inclusion of a noise penalty factor adjusts the weighting of noise. If a frame has high noise levels, the penalty factor appropriately reduces the quality of that portion of the video, preventing noise from overly influencing the quality score and improving the accuracy of the assessment. By comparing the current frame with a reference frame, the formula performs the assessment on a more standardized basis, reducing bias due to varying video sources or shooting conditions. The reference frame can be a standard-definition frame, making the assessment more relative and comparable. Each factor in the formula (clarity, noise level, color reproduction) is weighted to ensure that each dimension's contribution to the final quality score is reasonable. This weighted calculation can reflect the relative importance of different quality characteristics, thereby obtaining a more accurate quality assessment result. The final low-light video quality dimension scoring system maps the quality dimension to a specific score range (1-5 points). This annotation system not only provides a simple and easy-to-understand scoring standard to help users or systems quickly evaluate video quality, but also provides a basis for subsequent optimization and processing. The division of each score segment makes the evaluation results highly operational and facilitates further quality control and improvement. Since the quality of low-light videos is affected by many factors (such as noise, brightness, clarity, etc.), standardized quality dimension measurement formulas can reduce subjective evaluation differences, improve the consistency of low-light video quality assessment, and make quality scores reproducible.

[0126] Furthermore, the video quality model prediction module includes the following functions:

[0127] A video quality measurement model architecture is constructed using a convolutional neural network, which includes a sharp feature extraction branch, a noise feature extraction branch, and a fully connected neural network layer.

[0128] In an embodiment of the present invention, a deep learning architecture is constructed to measure the quality of low-light videos. The architecture specifically includes a sharp feature extraction branch, a noise feature extraction branch, and multiple fully connected neural network layers. The sharp feature extraction branch is used to extract sharpness information of video frames, while the noise feature extraction branch is used to extract noise information in video frames. The sharp feature extraction branch consists of five 5x5 convolutional layers and five 1x1 pooling layers. The convolutional layers are used to extract spatial features in the image, while the pooling layers are used for dimensionality reduction and feature compression. The noise feature extraction branch consists of five 3x3 convolutional layers and five 1x1 pooling layers. The convolutional layers are small in size to improve sensitivity to noise details. The number of fully connected neural network layers is determined by the number of quality dimensions to be predicted, with each dimension corresponding to a prediction score. The architecture is designed by weightedly fusing the extraction results of sharp features and noise features into a comprehensive feature vector, which is input into the fully connected neural network layer for nonlinear transformation, thereby outputting prediction scores for each quality dimension. The purpose of this process is to quantitatively evaluate video quality through a deep learning model, ensuring that the prediction scores can fully reflect the quality changes of the video, especially the performance under low-light conditions, and ultimately constructing a corresponding video quality measurement model architecture.

[0129] Preferably, the low-illumination video standard frame image set is divided into a training set, a validation set, and a test set according to a division ratio corresponding to 7:2:1;

[0130] In an embodiment of the present invention, a standard frame image set of low-light video is obtained, and the dataset is divided into a training set, a validation set, and a test set according to the requirements of the video quality measurement model. The dataset is divided in a ratio of 7:2:1, of which 70% is used as a training set for model training, 20% is used as a validation set for adjusting hyperparameters and verifying model performance, and 10% is used as a test set for evaluating the performance of the final model. When dividing the dataset, it is ensured that the data distribution of each subset is similar to that of the original dataset, and that video frames of different qualities are evenly distributed in each subset to avoid data bias affecting the effect of model training. During the division process, a random sampling method is used to ensure the fairness of the division, while ensuring the representativeness and diversity of the video frames in each subset. The dataset obtained through this step can provide a reliable data basis for subsequent model training and evaluation.

[0131] Preferably, the training set is input into the video quality measurement model architecture for model training, and stochastic gradient descent is used as an optimization algorithm in combination with the validation set to continuously adjust the corresponding hyperparameters of the model, where the hyperparameters include learning rate, batch size, and number of network layers. At the same time, the test set is used to optimize the model performance to generate a low-light video measurement model based on deep learning, and the prediction score corresponding to each video frame quality dimension is outputted at each fully connected neural network layer;

[0132] In an embodiment of the present invention, a training set is input into a designed deep learning model for training. The model training goal is to minimize the error between the predicted results and the actual quality dimension. Stochastic gradient descent (SGD) is used as an optimization algorithm, and the model weights are continuously adjusted in combination with the training set data. To ensure the adaptability and generalization ability of the model in different situations, a validation set is used to tune the model's hyperparameters. Specifically, the adjusted hyperparameters include the learning rate, batch size, and number of network layers. The learning rate controls the step size of each gradient update, the batch size determines the number of samples used in each training, and the number of network layers affects the complexity of the model. Through experiments on the validation set, the optimal hyperparameter combination is found to improve the training effect and stability of the model. During the training process, a cross-entropy loss function is used to quantify the error between the predicted results and the actual labels. The model parameters are updated through the backpropagation algorithm. The low-light video quality measurement model obtained through training can effectively learn video frame features and generate reasonable prediction scores for different quality dimensions, ultimately outputting the prediction score corresponding to each video frame quality dimension.

[0133] Preferably, based on the prediction score corresponding to each video frame quality dimension, a quality-weighted sum is performed on the corresponding low-illumination video to obtain a low-illumination video quality prediction score.

[0134] In an embodiment of the present invention, the quality of the low-light video is weightedly summed by using the prediction score corresponding to the quality dimension output by the model at each fully connected neural network layer. Specifically, the prediction score of each video frame represents the performance of the frame in a certain quality dimension. After weighted aggregation, the prediction scores of all frames are obtained to obtain the quality prediction score of the entire video. During the weighted summation process, it is necessary to set appropriate weighting coefficients based on the timing, features and performance of the video frames in different quality dimensions. Through this method, it can be ensured that the overall quality prediction of the video can comprehensively consider the quality performance of each frame, while emphasizing the impact of specific dimensions (such as clarity, noise, etc.) on the video quality. The obtained quality prediction score can provide a comprehensive quantitative indicator for the quality evaluation of low-light videos, help further quality optimization and tuning work, and finally predict and measure to obtain the low-light video quality prediction score.

[0135] Furthermore, the video quality measurement model architecture is specifically composed of a clear feature extraction branch, a noise feature extraction branch and multiple fully connected neural network layers, wherein the clear feature extraction branch is composed of 5 5x5 convolutional layers and 5 1x1 pooling layers, and the noise feature extraction branch is composed of 5 3x3 convolutional layers and 5 1x1 pooling layers. The fully connected neural network layer is determined by the total number corresponding to the quality dimension, so as to weightedly fuse the video frame features extracted by each branch into a comprehensive feature vector, and input the comprehensive feature vector into each fully connected neural network layer for nonlinear transformation to output the prediction score corresponding to each quality dimension.

[0136] Furthermore, the prediction error loss optimization module includes the following functions:

[0137] The quality mean square error loss is calculated between the low-light video quality annotation score and the low-light video quality prediction score to obtain the video quality error loss between the model prediction score and the annotation score;

[0138] In an embodiment of the present invention, the quality of low-light videos is labeled and scored, which can be achieved through human eye scoring or relying on existing video quality evaluation methods. These labeled scores will be used as standards for comparison with the scores predicted by the model. The low-light video quality prediction score is obtained by training a deep learning-based model on video frame image data. The quality mean square error loss (MSE) is an effective method for calculating the difference between the predicted score and the labeled score. Specifically, the difference between the labeled quality score of each low-light video frame and the quality score predicted by the model is first calculated, and then the squares of these differences are averaged to obtain the mean square error loss. This loss function can quantify the gap between the model prediction and the actual annotation, and serve as the basis for subsequent optimization, indicating the deviation and error of the model, so as to optimize the low-light video quality prediction model, and finally obtain the video quality error loss between the model prediction score and the labeled score.

[0139] Preferably, error loss reinforcement learning is performed on the low-light video measurement model based on deep learning based on the video quality error loss between the model prediction score and the labeled score, and an L2 regularization term is introduced to constrain the overfitting of the hyperparameters corresponding to the low-light video measurement model to generate a comprehensive low-light video measurement model;

[0140] In an embodiment of the present invention, error loss reinforcement learning is performed on a low-light video measurement model based on deep learning by utilizing the previously calculated mean squared error loss. To further improve model performance, the model needs to be trained using a reinforcement learning strategy. Under the guidance of the loss function, the model gradually improves its predictive ability. Specifically, the model parameters are backpropagated based on the MSE loss, and the network weights are adjusted using an optimization algorithm (such as gradient descent) to minimize the loss function. At this time, to avoid overfitting, an L2 regularization term needs to be introduced. The L2 regularization term increases the sum of squares of the model parameters in the loss function, causing the model to tend to select smaller parameter values ​​during optimization, thereby effectively reducing the occurrence of overfitting. In this way, the trained low-light video measurement model will be more stable and accurate in the video quality assessment process, reducing excessive dependence on specific data. After multiple rounds of training and optimization, a comprehensive model that can accurately assess low-light video quality is obtained, and ultimately a comprehensive low-light video measurement model is generated.

[0141] Preferably, the low-illumination video standard frame image set is re-input into the low-illumination video measurement comprehensive model for video quality assessment measurement to output a corresponding low-illumination video quality comprehensive measurement score.

[0142] In an embodiment of the present invention, a standard frame image set of low-light video will be re-input into a low-light video measurement comprehensive model that has undergone reinforcement learning and regularization processing for quality assessment. The standard frame image set consists of a set of representative video frames covering different low-light scenes and lighting conditions. These images will be passed as input data to the trained low-light video measurement model. The model will analyze based on the image features and video quality prediction mechanism. Specifically, the model will extract key features from each frame of video image, such as noise level, detail loss, color distortion, etc., and then calculate a comprehensive video quality score. The output low-light video quality comprehensive measurement score will reflect the overall quality of the video, including factors such as clarity, noise suppression effect, and brightness adaptation. This score can serve as the basis for subsequent tasks such as low-light video quality optimization and video content analysis. That is, the corresponding scores in each fully connected layer are weighted summed to finally output the corresponding low-light video quality comprehensive measurement score.

[0143] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.

[0144] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.

Claims

1. A low-light video quality measurement system based on deep learning, characterized in that: Includes the following modules: A low-light video frame processing module is used to collect corresponding low-light video data in real time by using a video acquisition device in different low-light scenes, and perform video frame processing on the low-light video data to obtain a low-light video standard frame image set; The video quality dimension measurement and annotation module is used to measure the video index of each frame of video image in the low-light video standard frame image set to obtain the video clarity, video noise level and video color reproduction corresponding to each low-light video frame; Based on the video clarity, video noise level, and video color reproduction corresponding to each low-light video frame, the corresponding low-light video quality dimension measurement and annotation is performed to obtain a low-light video quality annotation score. This includes the following functions: Perform multi-channel decomposition on each frame of video image in the low-illumination video standard frame image set to obtain R, G, and B channel images corresponding to each frame of video image; obtain the grayscale mean and grayscale variance corresponding to each channel through the R, G, and B channel images corresponding to each frame of video image, and measure the channel energy ratio of each frame of video image in the low-illumination video standard frame image set based on the grayscale mean and grayscale variance corresponding to each channel to obtain the channel image energy ratio corresponding to each frame of video image; perform video clarity analysis on each corresponding frame of video image according to the channel image energy ratio corresponding to each frame of video image to obtain the video clarity corresponding to each low-illumination video frame; The noise level of each frame of the low-light video standard frame image set is measured to obtain the video noise level corresponding to each low-light video frame; this includes: Dividing each frame of video image in the low-illumination video standard frame image set into a local image region to obtain a video local sub-region corresponding to each frame of video image; Calculating the texture fractal dimension of the local sub-region of the video corresponding to each frame of the video image to obtain the fractal dimension of the local region corresponding to each frame of the video image; Calculating the local pixel variance of the local sub-region of the video corresponding to each frame of the video image to obtain the local region pixel variance corresponding to each frame of the video image; The noise level of each frame of video image in the low-illumination video standard frame image set is measured based on the local region fractal dimension and the local region pixel variance corresponding to each frame of video image to obtain the video noise level corresponding to each low-illumination video frame; The color reproduction degree of each video frame in the low-light video standard frame image set is measured to obtain the video color reproduction degree corresponding to each low-light video frame; including: Calculating the chromaticity coordinates of different color regions corresponding to each frame of the video image in the low-light video standard frame image set to obtain the chromaticity coordinates corresponding to the different color regions in each frame of the video image; Based on the chromaticity coordinates corresponding to the different color areas in each frame of the video image, the color saturation of the corresponding different color areas in each frame of the video image in the low-light video standard frame image set is measured to obtain the color saturation corresponding to the different color areas in each frame of the video image; Performing hue distribution statistics on different color regions corresponding to each frame of the video image in the low-light video standard frame image set to obtain the hue distribution corresponding to the different color regions in each frame of the video image; The color reproduction degree of each frame of the low-light video standard frame image set is measured based on the color saturation and hue distribution corresponding to different color areas in each frame of the video image to obtain the video color reproduction degree corresponding to each low-light video frame; Based on the video clarity, video noise level, and video color reproduction corresponding to each low-light video frame, a quality dimension measurement and calculation is performed on the corresponding low-light video to obtain the low-light video quality dimension; Based on the low-light video quality dimension, the corresponding low-light video is labeled with a quality dimension system, with a quality dimension of 0%-20% labeled as 1 point, 21%-40% labeled as 2 points, 41%-60% labeled as 3 points, 61%-80% labeled as 4 points, and 81%-100% labeled as 5 points, to obtain the low-light video quality labeling score; The video quality model prediction module is used to construct a video quality measurement model architecture using a convolutional neural network, and to optimize the video quality measurement model architecture using a low-light video standard frame image set to generate a low-light video measurement model based on deep learning, and output a prediction score corresponding to each video frame quality dimension; based on the prediction score corresponding to each video frame quality dimension, a quality-weighted summation is performed on the corresponding low-light video to obtain a low-light video quality prediction score; wherein, the following steps are included: A video quality measurement model architecture is constructed using a convolutional neural network, which includes a sharp feature extraction branch, a noise feature extraction branch, and a fully connected neural network layer. The low-light video standard frame image set is divided into a training set, a validation set, and a test set according to the division ratio of 7:2:1; The training set is fed into the video quality measurement model architecture for model training. Stochastic gradient descent is used as an optimization algorithm in conjunction with the validation set to continuously adjust the model's corresponding hyperparameters, including the learning rate, batch size, and number of network layers. The test set is also used to optimize model performance. This generates a deep learning-based low-light video measurement model, which outputs a predicted score for each video frame quality dimension at each fully connected neural network layer. Based on the prediction scores corresponding to the quality dimensions of each video frame, the corresponding low-light video is quality-weighted summed to obtain the low-light video quality prediction score; The prediction error loss optimization module is used to perform error loss reinforcement learning on the deep learning-based low-light video measurement model based on the low-light video quality annotation score and the low-light video quality prediction score to generate a low-light video measurement comprehensive model; the low-light video standard frame image set is re-input into the low-light video measurement comprehensive model for video quality assessment measurement to output the corresponding low-light video quality comprehensive measurement score.

2. The low-light video quality measurement system based on deep learning according to claim 1, characterized in that The low-light video frame processing module includes the following functions: By using video acquisition equipment to collect corresponding low-light video data in real time under different low-light scenes; The low-light video data is extracted into single frames at a fixed frame rate of 24 fps to obtain a low-light video single-frame image set; Frame image blurriness is measured for a single-frame image set of a low-illumination video to obtain the frame image blurriness of the low-illumination video; Performing video frame image filtering on each low-illumination video frame in a low-illumination video single frame image set based on the low-illumination video frame image blurriness to obtain a low-illumination video frame filtered image set; The pixel values ​​corresponding to each low-illumination video image frame in the low-illumination video frame filtered image set are normalized to obtain a low-illumination video standard frame image set.

3. The low-light video quality measurement system based on deep learning according to claim 2, characterized in that The different low-light scenes specifically include outdoor scenes at night, underground parking lots, and indoor low-light environment scenes.

4. The low-light video quality measurement system based on deep learning according to claim 1, characterized in that The video quality measurement model architecture specifically consists of a clear feature extraction branch, a noise feature extraction branch, and multiple fully connected neural network layers, wherein the clear feature extraction branch consists of 5 5x5 convolutional layers and 5 1x1 pooling layers, and the noise feature extraction branch consists of 5 3x3 convolutional layers and 5 1x1 pooling layers. The fully connected neural network layer is determined by the total number corresponding to the quality dimension, so as to weightedly fuse the video frame features extracted by each branch into a comprehensive feature vector, and input the comprehensive feature vector into each fully connected neural network layer for nonlinear transformation to output the prediction score corresponding to each quality dimension.

5. The low-light video quality measurement system based on deep learning according to claim 1, characterized in that The prediction error loss optimization module includes the following functions: The quality mean square error loss is calculated between the low-light video quality annotation score and the low-light video quality prediction score to obtain the video quality error loss between the model prediction score and the annotation score; Based on the video quality error loss between the model prediction score and the annotated score, the deep learning-based low-light video measurement model is subjected to error loss reinforcement learning. The corresponding hyperparameters of the low-light video measurement model are constrained from overfitting by introducing an L2 regularization term to generate a comprehensive low-light video measurement model. The low-light video standard frame image set is re-input into the low-light video measurement comprehensive model for video quality assessment measurement to output the corresponding low-light video quality comprehensive measurement score.

Citation Information

Patent Citations

  • Video noise detecting method and device

    CN103096117A

  • Highway monitoring video definition detection method based on corner features

    CN104182983A

  • Training method of image quality evaluation model, and image quality evaluation method and device

    CN115661618A

  • Color correction matrix determination method and device, electronic equipment and storage medium

    CN117974811A

  • Marine remote sensing image enhancement method

    CN118674634A