A deep learning-based ultra-high-definition video blur quality evaluation method

By constructing a large-scale blur distortion dataset and a deep learning network, the accuracy and real-time issues of blur quality assessment in ultra-high-definition video in existing technologies are solved, realizing blur quality assessment of ultra-high-definition videos and images, and is applicable to efficient assessment of images and videos with multiple resolutions.

CN116258669BActive Publication Date: 2025-12-19CHINA ELECTRONICS RELIABILITY AND ENVIRONMENTAL TESTING INSTITUTE ((THE FIFTH INSTITUTE OF ELECTRONICS MINISTRY OF INDUSTRY AND INFORMATION TECHNOLOGY) (CHINA SAIBAO LABORATORY) +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211575397.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-08
Publication Date
2025-12-19
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

Existing video quality assessment methods are difficult to effectively assess single types of distortion, especially blur distortion, and cannot meet the real-time and accuracy requirements of ultra-high-definition video. Furthermore, they are not suitable for quality assessment of video images with multiple resolutions.

Method used

A large-scale image and video blur distortion dataset was constructed, and an image blur distortion evaluation and classification network was designed. The network was trained using deep learning methods to achieve blur quality evaluation of ultra-high-definition videos and images. An image gridding random cropping preprocessing module and a convolutional neural network were used for feature extraction to ensure the real-time performance and accuracy of the evaluation.

Benefits of technology

It achieves accurate and real-time blur quality assessment of ultra-high-definition videos and images, and can provide accurate blur distortion level assessment while ensuring speed. It is applicable to videos and images of different resolutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116258669B_ABST
    Figure CN116258669B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on deep learning's ultra-high definition video blur quality evaluation method, comprising the following steps;Step 1, establish ultra-high definition image blur distortion dataset as network training set and verification set;Step 2, construct image blur distortion evaluation classification network;Step 3, train image blur distortion evaluation classification network;Step 4, test ultra-high definition video blur distortion classification;Step 5, evaluate image blur distortion evaluation classification network.The application has made large-scale image blur distortion dataset and ultra-high definition video blur distortion dataset, is specially used for the quality evaluation of ultra-high definition video and image of blur distortion type, while designing image blur distortion evaluation classification network, guarantees the real-time performance and accuracy of ultra-high definition video quality evaluation, solves the problem of slow speed, low accuracy rate of quality evaluation on ultra-high definition video.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of super high definition video blur quality evaluation, and particularly relates to a super high definition video blur quality evaluation method based on deep learning. BACKGROUND

[0002] The main problems of the existing video quality evaluation method and the technical problems to be solved by the application are as follows:

[0003] (1) The existing video quality evaluation method mainly evaluates the quality of video comprehensive distortion, lacks distortion type evaluation of single distortion, and the distortion type mainly of blur is extremely common.

[0004] (2) The existing video quality evaluation method is mainly for low resolution video, and has problems of poor evaluation effect, slow evaluation speed and the like on super high definition video such as 4K, and even due to too many parameters of the neural network, the super high definition video exceeds the memory resource on the existing hardware resource and is difficult to evaluate.

[0005] (3) The existing video quality evaluation method is slow in video evaluation, and cannot meet the requirement of real-time.

[0006] (4) The existing video quality evaluation method is difficult to evaluate the quality of multi-size resolution video images.

[0007] (5) The existing quality evaluation method is difficult to be used for image quality evaluation and video quality evaluation at the same time.

[0008] Shanghai Jiaotong University discloses a super high definition video quality evaluation method and device in its applied patent document "a super high definition video quality evaluation method and device" (application date: March 12, 2020, publication number: CN111385567A, authorization date: January 5, 2021). The method can be used for evaluation of super high definition video and image quality, but it uses a traditional method, lacks a large-scale blur distortion data set, and thus the evaluation effect of the method for blur distortion quality is reduced, and the speed of the method in super high definition video quality evaluation is relatively slow, and real-time evaluation cannot be achieved.

[0009] Xi'an University of Posts and Telecommunications in its application patent document "Full reference image quality evaluation method based on subjective and objective feature fusion" (application date: July 21, 2021, publication number: CN113469998A, authorization date: October 18, 2022) discloses a full reference image quality evaluation method based on subjective and objective feature fusion. The method combines subjective and objective feature fusion for image quality evaluation, but the training data set has less blur distortion, the evaluation performance for blur distortion is not high, and the network is relatively complex for image quality evaluation, the evaluation speed is slow, and it is difficult to be used for ultra-high definition video quality evaluation. SUMMARY

[0010] In order to overcome the shortcomings of the prior art, the purpose of the present application is to provide a deep learning-based ultra-high definition video blur quality evaluation method. A large-scale image blur distortion data set and an ultra-high definition video blur distortion data set are prepared, which are specially used for quality evaluation of ultra-high definition videos and images of blur distortion type. An image blur distortion evaluation classification network is designed to ensure the real-time and accuracy of ultra-high definition video quality evaluation, and to solve the problems of slow speed and low accuracy of quality evaluation on ultra-high definition videos.

[0011] In order to achieve the above purpose, the technical scheme adopted by the present application is:

[0012] A deep learning-based ultra-high definition video blur quality evaluation method, comprising the following steps:

[0013] Step 1, establish an ultra-high definition image blur distortion data set as a network training set and a verification set;

[0014] Step 2, construct an image blur distortion evaluation classification network;

[0015] Step 3, train the image blur distortion evaluation classification network;

[0016] Step 4, test the ultra-high definition video blur distortion classification;

[0017] Step 5, evaluate the image blur distortion evaluation classification network.

[0018] The step 1 specifically comprises:

[0019] Step 1.1, the ultra-high definition image blur distortion dataset is divided into an image blur dataset and an ultra-high definition video blur dataset, and the dataset is obtained in two ways, a synthetic blur dataset and a real blur dataset, the synthetic blur dataset degrades the original image and the ultra-high definition video in a Gaussian blur degradation manner, the original image is from the Waterloo dataset, which is a large-scale image quality evaluation dataset, covering various scenes such as people, plants, animals, sky and buildings, and the original image is degraded using five levels of Gaussian kernels, wherein the five levels of Gaussian blur kernels are [0, 1], [3, 5], [7, 9, 11], [13, 15, 17, 19] and [21, 23, 25, 27, 29, 31, 33, 35, 37], and the degraded Gaussian blur distorted image is obtained;

[0020] Step 1.2, the real blur dataset uses manually captured ultra-high definition images of different scenes, including people, plants, animals, campus, street view, sky and buildings, and randomly sets the focal length of the mobile phone to different degrees of defocus, and then captures the defocus blurred ultra-high definition images in real scenes;

[0021] Step 1.3, the labeling personnel label the image blur distortion class labels of steps 1.1 and 1.2 in a well-lit environment, using the Windows system's built-in picture viewer, the evaluator selects the blur distortion level of the pictures and videos displayed on the display screen, the levels are extremely serious blur distortion, serious blur distortion, obvious blur distortion, slight blur distortion and negligible blur distortion, corresponding to class labels 0, 1, 2, 3 and 4 respectively;

[0022] Step 1.4, for each image, the image blur distortion level recognized by the majority of labeling personnel is selected as the final blur distortion class label of the image;

[0023] Step 1.5, the images and videos obtained in steps 1.1 and 1.2 and their corresponding blur distortion class labels are divided into a network training set and a network validation set in a ratio of 4:1, the training set and the validation set have no intersection and different scene distributions to ensure the scene application generalization performance of the network.

[0024] The step 2 specifically comprises:

[0025] Step 2.1, a preprocessing module of image grid random cropping is constructed, which first divides the input image into 7x7 grids of the same size, then randomly crops a 32x32 image block from each divided grid, and finally splices the 49 32x32 image blocks to form a 224x224 image;

[0026] Step 2.2, a convolution module is constructed, which includes a convolution input layer, a BN layer and a ReLu activation layer;

[0027] Step 2.3, an image blur distortion evaluation classification network is constructed, which is stacked by the convolution module in step 2.2, and a total of eleven convolution modules, a global average pooling layer and a fully connected layer are used, and the structure is in turn: the first convolution module, the second convolution module, the third convolution module, the fourth convolution module, the fifth convolution module, the sixth convolution module, the seventh convolution module, the eighth convolution module, the ninth convolution module, the global average pooling layer, the tenth convolution module, the eleventh convolution module, and the fully connected layer;

[0028] The forward propagation process of the image blur distortion evaluation classification network is as follows: after the input image is input into the image grid random cropping preprocessing module, it is sequentially input into the nine convolution modules to obtain a feature extraction image, the feature extraction image is input into the global average pooling layer to reduce the resolution of the feature extraction image, and then the tenth and eleventh convolution modules and the fully connected layer are input to output a five-classification vector.

[0029] The parameters of each layer of the network are as follows:

[0030] The convolution kernel size of all convolution modules is set to 3x3, the step size of the second, fourth, sixth and ninth convolution modules is set to 2, the step size of the first, third, fifth, seventh, eighth, tenth and eleventh convolution modules is set to 1, and the input channel number of the first to eleventh convolution modules is set to 3, 48, 48, 64, 64, 64, 64, 128, 128, 128, 256, respectively, and the output channel is set to 48, 48, 64, 64, 64, 64, 128, 128, 128, 256, 256, respectively;

[0031] The pooling parameter of the global average pooling layer is set to 1x1;

[0032] The input channel number of the fully connected layer is set to 256, and the output vector size is set to 5.

[0033] The step 3 specifically includes:

[0034] Step 3.1, configure the training environment and install the python library required for network training;

[0035] Step 3.2, set the training hyperparameters, the batch size is set to 32, the initial learning rate is set to 0.1, the weight decay rule is set to learning rate x 0.1 every 35 training periods, the solver is selected as SGD, and the training period epochs is set to 150;

[0036] Step 3.3, training the image blur distortion evaluation classification network using the stochastic gradient descent method, first, the input image data is processed by data augmentation, that is, the input image is processed by the grid random cropping preprocessing module to obtain an image with a resolution of 224x224, and the image is randomly horizontally flipped with a probability p, and the probability p is set to 0.5; then the processed image data is input into the image blur distortion evaluation classification network for forward propagation, the output value and the target value are calculated to obtain the loss function L, the solver SGD is combined with the learning rate lr to perform back propagation and update the network weights; after each epoch, the updated weights are assigned to the classification network, and the validation set is input into the network to calculate the blur distortion test accuracy of the validation set, which helps the network training to avoid overfitting or underfitting. After 150 epochs of training, the loss function L converges, the last updated weights are assigned to the image blur distortion evaluation classification network, and the trained image blur distortion evaluation classification network is obtained; during training, the loss function L is a multi-class cross-entropy loss function:

[0037]

[0038] where N represents the training sample size; M represents the classification category, and M is set to 5; y ij

[0039] represents the symbol function 0 or 1, if the true category of sample i is equal to j, y ij is 1, otherwise y ij is 0; p ij represents the probability that the network output of observation sample i belongs to category j.

[0040] The step 4 grids the ultra-high-definition video frame by frame and randomly crops and reassembles it to 224x224, and inputs it into the image blur distortion evaluation classification network trained in step 3 to output the blur distortion evaluation level frame by frame.

[0041] The step 5 is specifically:

[0042] Step 5.1, establish a super high-definition video blur distortion test set, by randomly setting the focal length of the mobile phone, so that the lens picture appears different degrees of defocus blur, 100 test videos are shot, each video is 10 seconds long, the frame rate is 30fps, the resolution includes 1080p, 2K, 4K, and the scene also covers people, plants and animals, campus, street view, sky, building, to test whether the network can accurately evaluate the blur distortion grade of the super high-definition real defocus blur; the test set is labeled with blur distortion category label frame by frame according to steps 1.3 and 1.4;

[0043] Step 5.2, input each video in the super high-definition video blur distortion test set into the image blur distortion evaluation classification network trained in step 3, output the final video blur distortion objective evaluation category, and calculate the average blur distortion test accuracy of the test set;

[0044] Step 5.3, input the videos with resolutions of 1080p, 2K and 4K into the image blur distortion evaluation classification network, and use the time function of python to calculate the average evaluation time of a frame of video image.

[0045] The evaluation method is applied to image sharpness evaluation function of an automatic imaging system, image and video quality evaluation of automatic screening of imaging results, enhancement of blurred images and videos, and short video recommendation.

[0046] The beneficial effects of the present application are:

[0047] 1) The present application constructs a large-scale super high-definition image and video blur distortion data set, by learning these data sets, the network has a high evaluation accuracy for blur distortion;

[0048] 2) Since the present application divides the super high-definition video image into five categories of extremely serious blur distortion, serious blur distortion, obvious blur distortion, slight blur distortion and negligible blur distortion, covers video blur distortion in various scenes, the trained neural network learns the difference between different blur distortion levels, so that the result of the present application in objective evaluation of video blur distortion quality is more accurate and detailed;

[0049] 3) The present application constructs an image grid random cropping preprocessing module, which randomly crops and splices the image to 224x224, retains global perception and local sensitivity, improves the network accuracy and evaluation speed, and makes the super high-definition video quality evaluation real-time, greatly improves the quality evaluation speed under the premise of ensuring the accuracy;

[0050] 4) The image blur distortion evaluation classification network model constructed by the application can be used for blur quality evaluation of ultra-high definition images and videos at the same time; the network has a small amount of parameters, which further ensures the real-time performance of video quality evaluation; the network can evaluate the quality of images and videos of different resolutions, and has high accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0051] Figure 1 The flowchart of the application is shown in the figure. DETAILED DESCRIPTION

[0052] The application will be further described in detail below with reference to the accompanying drawings.

[0053] As shown in the figure: Figure 1 Step 1, establish an ultra-high definition image blur distortion data set as a network training set and a verification set:

[0054] Step 1.1, the data set of the application is divided into an image blur data set and an ultra-high definition video blur data set, and the data set is mainly obtained by two ways, a synthetic blur data set and a real blur data set. For the synthetic blur data set, the original image and the ultra-high definition video are degraded by using a Gaussian blur degradation method. The source of the original image is the Waterloo data set, which is a large-scale image quality evaluation data set. The database currently contains 4,744 original natural images. The data set covers various scenes such as people, plants, animals, sky, and buildings. Five levels of Gaussian kernels are used to degrade the original natural images, and the five levels of Gaussian blur kernels are [0, 1], [3, 5], [7, 9, 11], [13, 15, 17, 19] and [21, 23, 25, 27, 29, 31, 33, 35, 37]. A total of 23,720 Gaussian blur distorted images are obtained after degradation. The source of the ultra-high definition video is a 4K ultra-high definition video website and manual shooting. The same five levels of Gaussian blur kernels are used for Gaussian degradation, and Gaussian blur distorted ultra-high definition videos are obtained after degradation.

[0055] Step 1.2, the real blur data set adopts manual shooting of ultra-high definition images of different scenes, and the scenes cover people, plants and animals, campus, street view, sky, and buildings. The focal length of the mobile phone is randomly set for each scene to be in different degrees of defocus state, and then the scene is shot to obtain 1000 defocus blur ultra-high definition images in real scenes.

[0056] Step 1.3, eight annotators of the embodiment of the application annotate the image blur distortion class labels of the images in step 1.1 and step 1.2, the embodiment is carried out in a well-lit environment, using the picture viewer of the Windows system, the resolution of the display is 1920x1080, each evaluator selects the blur distortion level of the picture and video displayed on the display screen, the level is extremely serious blur distortion, serious blur distortion, obvious blur distortion, slight blur distortion, and negligible blur distortion, which correspond to class labels 0, 1, 2, 3, and 4, respectively;

[0057] Step 1.4, for each image, the image blur distortion level recognized by the majority of annotators is selected as the final blur distortion class label of the image;

[0058] Step 1.5, the images and videos obtained in step 1.1 and step 1.2 and their corresponding blur distortion class labels are divided into a network training set and a network validation set in a ratio of 4:1, the training set and the validation set have no intersection and different scene distributions to ensure the scene application generalization performance of the network;

[0059] Step 2, construct an image blur distortion evaluation classification network:

[0060] Step 2.1, construct an image grid random cropping preprocessing module, which first uniformly divides the input image into 7x7 grids of the same size, then randomly crops a 32x32 image block from each grid, and then splices the 49 32x32 image blocks to form a 224x224 resolution image;

[0061] Step 2.2, construct a convolution module, which includes a convolution input layer, a BN layer, and a ReLu activation layer;

[0062] Step 2.3, construct an image blur distortion evaluation classification network, which is stacked by the convolution modules in step 2.2, a total of eleven convolution modules, a global average pooling layer, and a fully connected layer, the structure is: first convolution module, second convolution module, third convolution module, fourth convolution module, fifth convolution module, sixth convolution module, seventh convolution module, eighth convolution module, ninth convolution module, global average pooling layer, tenth convolution module, eleventh convolution module, and fully connected layer;

[0063] The forward propagation process of the image blur distortion evaluation classification network is: after the input image is input into the image grid random cropping preprocessing module, it is sequentially input into the nine convolution modules to obtain a feature extraction image, the feature extraction image is input into the global average pooling layer to reduce the resolution of the feature extraction image, and then the tenth and eleventh convolution modules and the fully connected layer are input to output a five-class vector.

[0064] The parameters of each layer of the network are set as follows:

[0065] The convolution kernel size of all convolution modules is set to 3x3, the step size of the second, fourth, sixth and ninth convolution modules is set to 2, the step size of the first, third, fifth, seventh, eighth, tenth and eleventh convolution modules is set to 1, the input channel number of the first to eleventh convolution modules is set to 3, 48, 48, 64, 64, 64, 64, 128, 128, 128, 256, respectively, and the output channel is set to 48, 48, 64, 64, 64, 64, 128, 128, 128, 256, 256, respectively;

[0066] The pooling parameter of the global average pooling layer is set to 1x1;

[0067] The input channel number of the fully connected layer is set to 256, and the output vector size is set to 5.

[0068] Step 3, training image blur distortion evaluation classification network:

[0069] Step 3.1, configure the training environment, and install the python library required for network training;

[0070] Step 3.2, set the training hyperparameters, the batch size is set to 32, the initial learning rate lr is set to 0.1, the weight decay rule is set to learning rate x 0.1 every 35 training periods, the solver is selected as SGD, and the training period epochs is set to 150;

[0071] Step 3.3, use the stochastic gradient descent method to train the image blur distortion evaluation classification network. First, the input image data is processed by data augmentation, that is, the input image is processed by the grid random cropping preprocessing module to obtain an image with a resolution of 224x224, and the image is randomly horizontally flipped with a probability p, and the probability p is set to 0.5; Then the image processed data is input into the image blur distortion evaluation classification network for forward propagation, the output value and the target value are calculated to obtain the loss function L, the solver SGD is combined with the learning rate lr to perform back propagation, and the network weight is updated; After each epoch, the updated weight is assigned to the classification network, and the validation set is input into the network to calculate the blur distortion test accuracy of the validation set, which helps the network training to avoid overfitting or underfitting. After 150 epochs of training, the loss function L converges, and the last updated weight is assigned to the image blur distortion evaluation classification network to obtain the trained image blur distortion evaluation classification network; During training, the loss function L is a multi-class cross-entropy loss function:

[0072]

[0073] Wherein, N represents the training sample size; M represents the classification category, and M is set to 5 in the embodiment; y ij represents a symbol function 0 or 1, if the real category of the sample i is equal to j, y ij is 1, otherwise y ij is 0; p ij represents the probability that the observed sample i belongs to the category j output by the network.

[0074] Step 4, test the blur distortion classification of the ultra-high definition video: the ultra-high definition video is gridized frame by frame, randomly cropped and spliced to 224x224, and then input into the image blur distortion evaluation classification network trained in step 3, and the blur distortion evaluation grade is output frame by frame.

[0075] Step 5, evaluate the image blur distortion evaluation classification network:

[0076] Step 5.1, an ultra-high definition video blur distortion test set is established in the embodiment, the focal length of the mobile phone is randomly set to make the lens picture appear different degrees of defocus blur, 100 test videos are shot, each video has a duration of 10 seconds, a frame rate of 30fps, and a resolution of 1080p, 2K and 4K, and the scenes also cover people, plants and animals, campus, street view, sky and building, etc., so as to test whether the network can accurately evaluate the blur distortion grade of the ultra-high definition real defocus blur; the test set is labeled with blur distortion category labels frame by frame according to steps 1.3 and 1.4;

[0077] Step 5.2, each video in the ultra-high definition video blur distortion test set is input into the trained image blur distortion evaluation classification network, the final video blur distortion objective evaluation category is output, and the average blur distortion test accuracy of the test set is calculated;

[0078] Step 5.3, the videos with resolutions of 1080p, 2K and 4K are respectively input into the image blur distortion evaluation classification network, and the time function of python is used to calculate the average evaluation time of a frame of video image.

[0079] The specific implementation idea of the application is as follows:

[0080] 1) A large number of synthetic and real ultra-high definition image blur data sets are collected to form a training set and a test set, the synthetic blur data set is blurred by the publicly disclosed image data set Waterloo and video data set LIVE_VQC, the real blur data set is composed of images and videos with defocus blur and motion blur actually shot, the image blur data set is mainly used for training, and the image and video data sets are used for testing together;

[0081] 2) The image blur dataset and the ultra-high-definition video blur dataset are divided into five blur distortion levels, namely 0, 1, 2, 3 and 4, wherein 0 is no distortion, 4 is severe distortion blur, and the higher the level value is, the more serious the blur distortion is, and the blur distortion level is refined;

[0082] 3) The present application realizes real-time evaluation of ultra-high-definition video through image gridding random slicing and reorganization, a convolutional neural network, global pooling and a fully connected classification network, inputs the ultra-high-definition video image gridding random slicing and reorganization into the convolutional neural network for feature extraction, unifies the feature vector size through global pooling, and finally outputs the blur distortion classification category through the fully connected layer;

[0083] 4) In the evaluation test phase, the present application collects a large number of synthetic blur distortion and real blur distortion ultra-high-definition video dataset and image dataset to form a test set, and evaluates the accuracy and processing speed of the present application on the image dataset and the ultra-high-definition video dataset respectively as an index for evaluating the performance.

[0084] Application prospect of the present application:

[0085] The blur problem of video and image will affect people's perception, acquisition of information and subsequent processing of video and image, especially in some high-quality image application occasions, such as medical analysis and diagnosis, remote sensing, biometric identification, monitoring system, etc. In addition, with the rapid development of short video field, the server side needs to recommend video content with clearer picture quality to users, therefore, various analysis and processing methods for blurred video and image have been long-term and widely concerned and applied. In recent years, with the rapid development of ultra-high-definition video, it is more necessary to have an algorithm for evaluating the quality of ultra-high-definition video, which can promote the further development of the field of ultra-high-definition video and bring better user experience.

[0086] The deep learning-based ultra-high-definition video blur quality evaluation method based on deep learning of the present application can accurately and quickly evaluate the ultra-high-definition video blur quality, and has very rich application prospect:

[0087] (1) Applied to image sharpness evaluation function of automatic imaging system. In the automatic focusing system based on focusing depth method, there are three important links, namely focusing window selection, image sharpness evaluation function and search algorithm. Among them, the image sharpness evaluation function realizes the quality evaluation of images with different blur degrees, thereby providing a basis for obtaining the final in-focus image;

[0088] (2) Image and video quality evaluation method applied to automatic screening of imaging results. Due to the influence of internal and external factors (such as environmental factors and human factors) of the imaging system, the finally collected image may be a degraded image containing blur problems, and the image and video will also cause distortion in the process of compression, transmission and storage. Therefore, by using effective quality evaluation method to evaluate the image and video, the image and video not meeting the quality requirements are discarded, thereby providing guarantee for subsequent processing;

[0089] (3) Enhancement algorithm applied to blurred images and videos. Image deblurring and video deblurring method as one of the common image and video enhancement algorithms, realizes the deblurring processing of blurred images and videos, and restores the blurred distorted images to clear images;

[0090] (4) Applied to the field of short video recommendation, the video blur quality can be graded, different pushing intensity is given, the video pushing rule is optimized, and the maximum benefit is obtained.

Claims

1. A method for assessing blur quality in ultra-high-definition video based on deep learning, characterized in that, Includes the following steps; Step 1: Establish an ultra-high-definition image blur and distortion dataset as the network training and validation set; Step 2: Construct an image blur distortion evaluation and classification network; Step 3: Train the image blur distortion evaluation and classification network; Step 4: Test the blur and distortion classification of ultra-high-definition video; Step 5: Evaluate the image blur distortion assessment classification network; Step 2 specifically includes: Step 2.1: Construct an image gridding random cropping preprocessing module. This module first divides the input image into 7×7 grids of the same size, then randomly crops a 32×32 image block from each grid, and then stitches together the 49 32×32 image blocks to form a 224×224 resolution image. Step 2.2: Construct a convolutional module, which includes a convolutional input layer, a BN layer, and a ReLU activation layer; Step 2.3: Construct an image blur distortion evaluation and classification network, which is composed of stacked convolutional modules from Step 2.

2. A total of eleven convolutional modules, one global average pooling layer, and a fully connected layer are used. The structure is as follows: first convolutional module, second convolutional module, third convolutional module, fourth convolutional module, fifth convolutional module, sixth convolutional module, seventh convolutional module, eighth convolutional module, ninth convolutional module, global average pooling layer, tenth convolutional module, eleventh convolutional module, and fully connected layer. The forward propagation process of the image blur distortion evaluation and classification network is as follows: After the input image is processed by the image gridding random cropping preprocessing module, it passes through nine convolutional modules in sequence to obtain the feature extraction map. The feature map is then passed through a global average pooling layer to reduce the resolution of the feature map. Finally, it passes through the tenth and eleventh convolutional modules and a fully connected layer to output a five-class vector. The parameters of each layer of the network are as follows: Set the kernel size of all convolutional modules to 3×3, set the stride of the second, fourth, sixth, and ninth convolutional modules to 2, set the stride of the first, third, fifth, seventh, eighth, tenth, and eleventh convolutional modules to 1, set the number of input channels of the first to eleventh convolutional modules to 3, 48, 48, 64, 64, 64, 64, 128, 128, 128, 256 respectively, and set the number of output channels to 48, 48, 64, 64, 64, 64, 128, 128, 128, 256, 256 respectively. Set the pooling parameter of the global average pooling layer to 1×1; Set the number of input channels of the fully connected layer to 256 and the output vector size to 5.

2. The method for evaluating the blur quality of ultra-high-definition video based on deep learning according to claim 1, characterized in that, Step 1 specifically includes: Step 1.1: The ultra-high-definition image blur distortion dataset is divided into an image blur dataset and an ultra-high-definition video blur dataset. The datasets are obtained in two ways: a synthetic blur dataset and a real blur dataset. The synthetic blur dataset uses Gaussian blur degradation to degrade both the original images and the ultra-high-definition videos. The original images are sourced from the Waterloo dataset, a large-scale image quality assessment dataset covering diverse scenes including people, plants, animals, sky, and buildings. Gaussian degradation is performed on the original images using five levels of Gaussian kernels: [0,1], [3,5], [7,9,11], [13,15,17,19], and [21,23,25,27,29,31,33,35,37]. The resulting image is a Gaussian-blurred and distorted image. The ultra-high-definition videos are sourced from 4K ultra-high-definition video websites and manually captured footage. Gaussian degradation is performed using the same five levels of Gaussian kernels as the images, resulting in a Gaussian-blurred and distorted ultra-high-definition video. Step 1.2: The real blur dataset is obtained by manually shooting ultra-high-definition images of different scenes, including people, animals and plants, campus, street scenes, sky, and buildings. The focal length of the mobile phone is randomly set for each scene to make it in different degrees of out-of-focus state, and then the images are taken to obtain out-of-focus blur ultra-high-definition images of real scenes. Step 1.3: The annotators label the images from Steps 1.1 and 1.2 with blur distortion category labels. This is done in a well-lit environment using the image viewer built into the Windows system. The evaluators select the blur distortion level of the images and videos displayed on the screen. The levels are divided into five categories: extremely severe blur distortion, severe blur distortion, obvious blur distortion, slight blur distortion, and negligible blur distortion, corresponding to category labels 0, 1, 2, 3, and 4, respectively. Step 1.4: For each image, select the image blur distortion level agreed upon by the majority of annotators as the final blur distortion category label for that image; Step 1.5: Divide the images and videos obtained in Steps 1.1 and 1.2 and their corresponding blur and distortion category labels into a network training set and a network validation set in a 4:1 ratio. The training set and the validation set have no overlap and different scene distributions to ensure the network's scene application generalization performance.

3. The method for evaluating the blur quality of ultra-high-definition video based on deep learning according to claim 1, characterized in that, Step 3 specifically includes: Step 3.1: Configure the training environment and install the Python libraries required for network training; Step 3.2: Set the training hyperparameters. Set the batch size to 32, the initial learning rate (lr) to 0.1, the weight decay rule to 0.1 every 35 training epochs, the solver to SGD, and the training epochs to 150. Step 3.3: Train the image blur distortion evaluation classification network using stochastic gradient descent. First, perform data augmentation on the input image data. This involves preprocessing the input image using a gridded random cropping module to obtain a 224×224 resolution image, and then randomly horizontally flipping it with probability p (set to 0.5). Next, input the processed image data into the image blur distortion evaluation classification network for forward propagation. Calculate the loss function L between the output and target values. Perform backpropagation using the SGD solver combined with the learning rate lr to update the network weights. After each epoch, assign the updated weights to the classification network. Input the validation set into the network to calculate the validation set blur distortion test accuracy, aiding network training to avoid overfitting or underfitting. After 150 epochs, the loss function L converges. Assign the last updated weights to the image blur distortion evaluation classification network to obtain the trained image blur distortion evaluation classification network. During training, the loss function L is the multi-class cross-entropy loss function. Where N represents the training sample size; M represents the classification categories, and M is set to 5; y ij The sign function represents 0 or 1; if the true class of sample i is equal to j, then y ij If y is 1, otherwise y ij p is 0; ij This represents the probability that the observed sample i in the network output belongs to category j.

4. The method for evaluating the blur quality of ultra-high-definition video based on deep learning according to claim 1, characterized in that, Step 4 involves randomly cropping and re-stitching the ultra-high-definition video frame by frame into a grid, then inputting it into the image blur distortion evaluation and classification network trained in step 3, and outputting the blur distortion evaluation level frame by frame.

5. The method for evaluating the blur quality of ultra-high-definition video based on deep learning according to claim 1, characterized in that, Step 5 specifically involves: Step 5.1: Establish an ultra-high-definition video blur distortion test set. By randomly setting the focal length of the mobile phone, different degrees of blur distortion were caused in the lens image. 100 test videos were shot, each 10 seconds long, at a frame rate of 30fps, with resolutions including 1080p, 2K, and 4K. The scenes also covered people, animals and plants, campus, street scenes, sky, and buildings to test whether the network can accurately assess the blur distortion level of ultra-high-definition real-world blur distortion. The test set was labeled with blur distortion category labels frame by frame according to steps 1.3 and 1.

4. Step 5.2: Input each video in the ultra-high-definition video blur distortion test set into the image blur distortion evaluation and classification network trained in step 3, output the final objective evaluation category of video blur distortion, and calculate the average blur distortion test accuracy of the test set. Step 5.3: Input the 1080p, 2K and 4K resolution videos into the image blur distortion evaluation and classification network respectively, and use the time function of Python to calculate the average evaluation time of one frame of video image.

6. A method for assessing the blur quality of ultra-high-definition video based on deep learning according to any one of claims 1-5, characterized in that, The evaluation method is applied to image sharpness evaluation functions of automatic imaging systems, image and video quality evaluation for automatic screening of imaging results, enhancement of blurred images and videos, and short video recommendation.

Citation Information

Patent Citations

  • Ultra-high-definition video quality evaluation method and device

    CN111385567A

  • Full-reference image quality evaluation method based on subjective and objective feature fusion

    CN113469998A

  • No-reference image quality evaluation method based on double-flow convolutional neural network

    CN111127435A

  • Method and electronic device for processing images that can be played on a virtual device

    US20210374908A1