Video Quality Determination Method, Apparatus, Device, and Storage Medium Without Reference
Through deep feature extraction and space-time normalization processing combined with multi-dimensional analysis of the joint loss module, the accuracy of the video quality evaluation of smart terminals is solved, and more reasonable video quality evaluation results are provided, which improves user experience and subsequent processing decision support.
Patent Information
- Application Number
- CN202210068265.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-20
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-01-20
AI Technical Summary
The existing reference-free video quality evaluation method is insufficient in dealing with distortion and encoding noise in smart terminal videos, and cannot provide efficient and reasonable video quality evaluation results.
The video image frame is compressed through the deep feature extraction network, and after performing time-space normalization processing, the joint loss module is used to perform multi-dimensional feature analysis, and the video quality scores are output, including distortion analysis and subjective analysis modules.
It realizes more accurate video quality evaluation, provides multi-dimensional video quality evaluation results, and improves the visual experience and subsequent processing decision support on the audience.
Smart Images

Figure CN114596259B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of video processing, and in particular, to a method, apparatus, device, and storage medium for determining video quality without reference. Background Art
[0002] In recent years, with the evolution of basic network technologies, the improvement of the hardware performance of mobile devices, and people's pursuit and yearning for a better life, the behavior of watching short videos, live videos, and long video content on intelligent terminals has become widespread. As a carrier of information, the video medium can convey richer information than single text and audio. However, from the generation of video to the viewing by the end user, it involves the conversion of optical signals to digital signals, the scaling process of video frames, video encoding and compression for uploading, etc. Among these, the thermal noise brought by optical signals, the blurring caused by scaling, and the block effect of the encoded picture will all affect the final visual experience of users. In addition, different from the videos shot by professional personnel in traditional media, a large number of videos produced by ordinary users contain defects such as overexposure, underexposure, jitter, or motion blur, and these problems also affect the visual experience of users. Therefore, it is of great significance for user-centered video content providers or related media to determine the quality of videos. Although the quality of each video frame can be determined through manual review, obviously this method is inefficient, there are differences in the subjective vision of different reviewers, and it is impossible to distinguish the quality of video frames with subtle changes. Therefore, a large number of researchers and related practitioners are seeking efficient and accurate automated video quality assessment methods.
[0003] Among them, video quality assessment methods can be divided into full-reference, semi-reference, and no-reference measurement methods. Among them, full-reference and semi-reference require some undistorted original image information, and the prediction results of current mainstream methods have a high correlation with people's subjective visual perception. However, this series of methods requires users to upload uncompressed original image information, which greatly limits the application of this method in large-scale scenarios. On the contrary, no-reference video quality assessment only uses the existing video input to provide a prediction of the subjective video quality score of this video.
[0004] In the prior art, reference-free video quality assessment methods are mainly divided into two types: those based on natural scene statistics and those based on deep features. The former approximates various image characteristics in natural scenes through a generalized skew Gaussian distribution, and usually uses trap filters to extract luminance component features, chrominance component features, as well as related combined features and directional features. However, video media on intelligent terminals often contains multiple superimposed distortions and non-standard shooting methods, and the methods based on natural scene statistics have limited processing capabilities for such problems. The methods based on deep features use advanced neural network models to extract discriminative features on large-scale data sets. These features contain semantic information and intrinsic features of images, but lack an efficient and reasonable determination mechanism for determining the video quality of videos with frame distortions caused by a large number of non-standard shooting methods and coding noises such as block effects introduced by video compression to save network bandwidth. Summary of the Invention
[0005] Embodiments of the present invention provide a reference-free video quality determination method, apparatus, device, and storage medium, which solve the problem that in the prior art, accurate evaluation cannot be performed when determining video quality, and can obtain more reasonable video quality evaluation results, providing good guiding significance for subsequent content distribution by operators.
[0006] In a first aspect, an embodiment of the present invention provides a reference-free video quality determination method, which includes:
[0007] Obtain a video image frame to be evaluated, and perform information compression on the video image frame through a deep feature extraction network to obtain a compressed image;
[0008] Perform spatio-temporal normalization processing on the compressed image to obtain a normalized image;
[0009] Analyze the normalized image through a joint loss module to output a video quality score, where the joint loss module includes at least two feature analysis modules.
[0010] In a second aspect, an embodiment of the present invention further provides a reference-free video quality determination apparatus, including:
[0011] A deep feature extraction module, configured to obtain a video image frame to be evaluated and perform information compression on the video image frame to obtain a compressed image;
[0012] A spatio-temporal feature normalization module, configured to perform spatio-temporal normalization processing on the compressed image to obtain a normalized image;
[0013] A joint loss module, configured to analyze the normalized image to output a video quality score, where the joint loss module includes at least two feature analysis modules.
[0014] In a third aspect, an embodiment of the present invention further provides a no-reference video quality determination device, which includes:
[0015] One or more processors;
[0016] A storage device for storing one or more programs,
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the no-reference video quality determination method described in the embodiments of the present invention.
[0018] In a fourth aspect, an embodiment of the present invention further provides a storage medium storing computer-executable instructions, and the computer-executable instructions are used to execute the no-reference video quality determination method described in the embodiments of the present invention when executed by a computer processor.
[0019] In the embodiments of the present invention, by acquiring a video image frame to be evaluated, compressing the information of the video image frame through a deep feature extraction network to obtain a compressed image, performing spatio-temporal normalization processing on the compressed image to obtain a normalized image, and then analyzing the normalized image through a joint loss module to output a video quality score, wherein the joint loss module includes at least two feature analysis modules, and a more reasonable video quality evaluation result can be obtained. Description of the Drawings
[0020] Figure 1 It is a flowchart of a no-reference video quality determination method provided by an embodiment of the present invention;
[0021] Figure 2 It is a flowchart of a method for analyzing a normalized image through a joint loss module to output a video quality score provided by an embodiment of the present invention;
[0022] Figure 3 It is a flowchart of another method for analyzing a normalized image through a joint loss module to output a video quality score provided by an embodiment of the present invention;
[0023] Figure 4 It is a flowchart of another no-reference video quality determination method provided by an embodiment of the present invention;
[0024] Figure 5 It is a flowchart of another no-reference video quality determination method provided by an embodiment of the present invention;
[0025] Figure 6 It is a flowchart of another no-reference video quality determination method provided by an embodiment of the present invention;
[0026] Figure 7The structural block diagram of a reference - free video quality determination device provided by an embodiment of the present invention;
[0027] Figure 8 The structural schematic diagram of a reference - free video quality determination device provided by an embodiment of the present invention. Detailed implementation manners
[0028] The following further elaborates on the embodiments of the present invention in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the embodiments of the present invention, rather than limiting the embodiments of the present invention. Additionally, it should be noted that for the sake of description, only parts related to the embodiments of the present invention are shown in the drawings, rather than all the structures.
[0029] The terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.
[0030] Figure 1 The flowchart of a reference - free video quality determination method provided by an embodiment of the present invention, which can be used to evaluate the quality of various videos. This method can be executed by computing devices such as servers, smart terminals, laptops, tablets, etc., and specifically includes the following steps:
[0031] Step S101: Obtain the video image frames to be evaluated, and compress the information of the video image frames through a deep feature extraction network to obtain compressed images.
[0032] Among them, the video image frames to be evaluated are the image frames for which the video quality needs to be determined, and they can be one frame of image or multiple frames of images. For example, when determining the video quality of the video images uploaded by users, each frame of the video image or several selected frames of the video image are determined as the video image frames to be evaluated for video quality determination, and the corresponding video quality scores are output.
[0033] In one embodiment, first, the video image frames are compressed by a deep feature extraction network to obtain compressed images. Optionally, the structure of the deep feature extraction network can adopt a network structure pre-trained on the ImageNet large-scale dataset, such as MobileNet, ResNet, or VGG, etc., which includes a plurality of consecutive feature extraction modules and downsampling modules.
[0034] In one embodiment, the video image frames to be evaluated are video frame images of a preset size obtained after whitening processing and size normalization processing. Taking the video image frames to be evaluated as multiple frames of images included in multiple different video data as an example, the input format of the deep feature extraction network is exemplarily denoted as N×D×C×H×W, where N represents the number of videos, D represents the number of video image frames included in each video, C represents the number of channels, and H and W represent the height and width of the preset image size. Defining the number of downsampling times as r, then the output of the corresponding deep feature extraction network can be denoted as N×D×C×H / 2^r×W / 2^r.
[0035] Step S102: Perform spatio-temporal normalization processing on the compressed image to obtain a normalized image.
[0036] In one embodiment, the compressed image is further compressed through spatio-temporal normalization processing and the image features are retained. Optionally, it includes calculating the mean value of the temporal and spatial features of the compressed image and extracting features to obtain a normalized image. The specific formula for the normalization processing used in this embodiment is as follows:
[0037]
[0038] Step S103: Analyze the normalized image through a joint loss module to output a video quality score, where the joint loss module includes at least two feature analysis modules.
[0039] Among them, after obtaining the normalized image through step S102, the normalized image is analyzed through a configured joint loss module to output a video quality score. Among them, the joint loss module is a pre-trained module. The joint loss module includes at least two feature analysis modules, where each feature analysis module can generate a sub-feature score of the corresponding type based on the normalized image.
[0040] In one embodiment, taking the joint loss module including two branches, namely a distortion analysis module and a subjective analysis module, as an example for exemplary illustration, as Figure 2 shown, Figure 2 This is a flowchart of a method for analyzing a normalized image through a joint loss module to output a video quality score provided by an embodiment of the present invention, which specifically includes:
[0041] Step S1031: Perform distortion analysis processing on the normalized image through the distortion analysis module to obtain a predicted output.
[0042] Among them, the distortion analysis module includes one or more feature branches, and each feature branch corresponds to a predicted output value. Taking the clarity branch of the distortion analysis module as an example, it outputs the clarity score of the video image frame.
[0043] Step S1032: Input the predicted output and the spatio-temporal features of the normalized image into the subjective analysis module to output the video quality score.
[0044] Among them, after obtaining the predicted output of the distortion analysis module, input it together with the spatio-temporal features of the normalized image into the subjective analysis module to obtain the final video quality score.
[0045] In another embodiment, the distortion analysis module includes at least two distortion analysis sub-modules. Optionally, taking the distortion analysis module including five distortion analysis sub-modules as an example, each distortion analysis sub-module corresponds to a distortion feature branch, which are exemplarily a noise feature branch, a clarity feature branch, a blurriness feature branch, a brightness feature branch, and a chrominance feature branch. Optionally, the process of analyzing the normalized image through the joint loss module to output the video quality score is as Figure 3 shown, Figure 3 is a flowchart of another method for analyzing the normalized image through the joint loss module to output the video quality score provided by the embodiment of the present invention, which specifically includes:
[0046] Step S1033: Respectively perform distortion analysis processing on the normalized image through each of the distortion analysis sub-modules to obtain their respective corresponding predicted outputs.
[0047] Exemplarily, taking the distortion analysis sub-module including five branches: a noise feature branch, a clarity feature branch, a blurriness feature branch, a brightness feature branch, and a chrominance feature branch as an example, perform distortion analysis processing on the normalized image respectively to obtain a noise feature score, a clarity feature score, a blurriness feature score, a brightness feature score, and a chrominance feature score. Among them, the noise feature branch is used to extract thermal noise, optical noise, and color noise of the video image frame to output a predicted score; the clarity feature branch is used to extract the edge information of the video image frame to output a predicted score; the blurriness feature branch is used to extract the smoothness information of the video image frame to output a predicted score; the brightness feature branch is used to extract the light distribution of the video image frame to output a predicted score; the chrominance feature branch is used to extract the color richness information of the video image frame to output a predicted score. It should be noted that other feature branches can also be trained according to actual needs and added to the joint loss module.
[0048] Optionally, each feature branch can adopt a series of non-linear layers to map the spatio-temporal features in the normalized image to the desired image quality features corresponding to their respective branches. For example, the noise feature branch maps the spatio-temporal features in the normalized image to obtain the noise features of the video image frames; the sharpness feature branch maps the spatio-temporal features in the normalized image to obtain the edge information of the video image frames. The specific mapping formula can be:
[0049]
[0050] where, [W 1 ,W 2 ,W 3 ,W 4 ,W 5 ,B 1 ,B 2 are the optimized parameters obtained through training. Different feature branches correspond to different parameter values. σ(·) represents any non-linear activation function. Exemplarily, it can be the sigmoid function. ReLU(·) is the rectified linear unit function. n is the number of video image frames, j is the identifier of a specific video image frame, and c is the number of image channels.
[0051] Step S1034: Input the prediction output corresponding to each of the distortion analysis sub-modules and the spatio-temporal features of the normalized image into the subjective analysis module to output a video quality score.
[0052] Among them, the noise feature score, sharpness feature score, blurriness feature score, brightness feature score, and chroma feature score obtained in step S1033 are input into the subjective analysis module together with the spatio-temporal features of the normalized image, and the subjective analysis module outputs the final video quality score. Optionally, the subjective analysis module maps the input noise feature score, sharpness feature score, blurriness feature score, brightness feature score, chroma feature score, and the spatio-temporal features of the normalized image to one dimension through a single-layer linear transformation to obtain the final video quality score.
[0053] As can be seen from the above solution, for the video image frames to be evaluated, after compressing the information of the video image frames through the deep feature extraction network to obtain a compressed image, performing spatio-temporal normalization processing on the compressed image to obtain a normalized image, and using the joint loss module to obtain the final video quality score determined in multiple dimensions, it solves the problem of poor rationality of the scores brought about by using a single dimension such as subjective picture quality assessment in the prior art to determine the video quality, making the video quality assessment result more accurate, and providing richer decision-making support for subsequent processing. It is convenient for optimizing the image quality of video images with maximum efficiency in subsequent processing, and further improving the visual experience of viewers at the client side.
[0054] Figure 4A flowchart of another method for determining video quality without reference provided by an embodiment of the present invention provides a specific step of training a joint loss module, such as Figure 4 As shown, specifically including:
[0055] Step S201: Obtain different label values of manually annotated video samples, and train the joint loss module according to the different label values and the corresponding relationship between the label values and the feature analysis module.
[0056] In one embodiment, when the video samples are annotated, the manual scoring of the label values is performed according to multiple different dimensions. Optionally, taking the joint loss module including 6 different feature analysis modules as an example, each feature analysis module corresponds to a feature branch, which are respectively recorded as noise feature branch, clarity feature branch, blur feature branch, brightness feature branch, chromaticity feature branch and subjective feeling branch. When the manual annotation is performed, the 6 dimensions are scored, that is, the noise value, clarity value, blur value, brightness value, color value and subjective value are respectively corresponded. For example, a score range of 1-5 points is performed, where a score of 1 indicates the worst unacceptable, a score of 2 indicates that the bad is acceptable but unwilling to continue browsing, a score of 3 indicates a general viewing experience, a score of 4 indicates a willingness to continue watching, and a score of 5 indicates the expectation to see more similar videos. The specific annotation details exemplarily use the dual-stimulus continuous standard grading method mentioned in the ITU-R BT.500 standard to guide the annotation. It should be noted that the above example is described using six scoring dimensions. In another embodiment, scoring may be performed on at least two of the six dimensions.
[0057] In one embodiment, when training the joint loss module, the objective function may be:
[0058]
[0059] Among them, N is the number of video frame images, i is the number of specific video frame images, and its value range is 1 to N, where F represents the output corresponding to each feature branch, and Y represents the manually labeled label. In the specific training process, L1 norm is used for regression training, and ADAM optimizer is used for parameter optimization, which will not be repeated here. In the training process of the joint loss module, when the output value of the loss function no longer decreases, the training is stopped to obtain a trained joint loss module.
[0060] Step S202: Obtain a video image frame to be evaluated, and compress the information of the video image frame through a deep feature extraction network to obtain a compressed image.
[0061] In one embodiment, before the video image frame is compressed by the depth feature extraction network to obtain a compressed image, it further includes cropping the video image frame to a preset size. Specifically, the center cropping method can be adopted, that is, taking the center of the video image frame as the reference point, cropping the length and width of the preset size to obtain a video image of a fixed size for processing by the depth feature extraction network.
[0062] Step S203: Perform spatio-temporal normalization processing on the compressed image to obtain a normalized image.
[0063] Step S204: Analyze the normalized image through the joint loss module to output a video quality score. The joint loss module includes at least two feature analysis modules.
[0064] As can be seen from the above solution, through the manual annotation of multiple video image quality dimensions and the training of the joint loss module, the finally obtained video quality score is more valuable as a reference, avoiding the problem of poor video quality evaluation effect caused by a single evaluation factor.
[0065] Figure 5 The figure is a flowchart of another no-reference video quality determination method provided by an embodiment of the present invention. Before annotating the video image, it further includes a data cleaning process, such as Figure 5 shown, specifically including:
[0066] Step S301: Obtain the original video data, preprocess the original video data to obtain video information in a unified format, generate a feature frequency histogram corresponding to the video information based on different picture quality characteristics, and filter the original video data based on the feature frequency histogram to obtain video samples.
[0067] Among them, the preprocessing includes converting the original video data from multiple different video image formats into the same format, such as converting it into video data in the YUV420 format. Of course, it can also be converted into image formats such as RGB24 or YUV444. When screening images, consider multiple dimensions of picture quality characteristics for screening. Exemplarily, the picture quality characteristics used for screening include at least one of bit rate, resolution, brightness, chroma, aspect ratio, and contrast. For each collected video segment, perform statistical analysis on the set screened picture quality characteristics frame by frame, and correspondingly use a median filter for noise suppression. For each collected video segment, determine whether to retain the video segment as a sample for manual annotation according to the average score of each frame image obtained.
[0068] Step S302: Obtain different tag values of the manually annotated video samples, and train the joint loss module according to the different tag values and the corresponding relationship between the tag values and the feature analysis modules.
[0069] Step S303: Obtain a video image frame to be evaluated, and perform information compression on the video image frame through a depth feature extraction network to obtain a compressed image.
[0070] Step S304: Perform spatio-temporal normalization processing on the compressed image to obtain a normalized image;
[0071] Step S305: Analyze the normalized image through a joint loss module to output a video quality score. The joint loss module includes at least two feature analysis modules.
[0072] As can be seen from the above, by obtaining the original video data, preprocessing the original video data to obtain video information in a unified format, then generating a feature frequency histogram corresponding to the video information based on different picture quality features, and filtering the original video data based on the feature frequency histogram to obtain video samples, data cleaning can be carried out efficiently and accurately to obtain more reasonable samples that can be manually labeled, avoiding pure color image samples and some meaningless video images as the objects of manual labeling and affecting the labeling efficiency.
[0073] Figure 6 It is a flowchart of another no-reference video quality determination method provided by an embodiment of the present invention, and a scheme for simultaneously outputting multiple types of video quality scores is given. As Figure 6 shown, it specifically includes:
[0074] Step S401: Obtain a video image frame to be evaluated, and perform information compression on the video image frame through a depth feature extraction network to obtain a compressed image;
[0075] Step S402: Perform spatio-temporal normalization processing on the compressed image to obtain a normalized image;
[0076] Step S403: Analyze the normalized image through a joint loss module to output the total score of the video image frame and the sub-feature scores output by each feature analysis module.
[0077] In one embodiment, when determining the video quality score, the joint loss module respectively outputs the sub-feature scores of each feature analysis module and the total score of the video image frame. Specifically, in the joint loss module, after analyzing the normalized image through each feature analysis module to obtain the corresponding sub-feature scores, the corresponding outputs are made to provide multi-dimensional video quality evaluation results for relevant personnel.
[0078] As can be seen from the above, by analyzing the normalized image through the joint loss module to output the total score of the video image frame and the sub-feature scores output by each feature analysis module, the prediction accuracy of the no-reference video quality assessment is significantly improved. At the same time, the quality assessment score of the picture quality can be provided specifically. By providing the influence degree of distortion in different dimensions, it helps to optimize the picture quality with the maximum efficiency in the subsequent link and further improves the visual experience of the audience when watching.
[0079] Figure 7 The following is a structural block diagram of a no-reference video quality determination device provided by an embodiment of the present invention. This device is used to execute the no-reference video quality determination method provided by the above embodiment and has corresponding functional modules and beneficial effects for executing the method. As Figure 7 shown, the device specifically includes: a depth feature extraction module 101, a spatio-temporal feature normalization module 102, and a joint loss module 103, where
[0080] The depth feature extraction module 101 is configured to obtain a video image frame to be evaluated and perform information compression on the video image frame to obtain a compressed image;
[0081] The spatio-temporal feature normalization module 102 is configured to perform spatio-temporal normalization processing on the compressed image to obtain a normalized image;
[0082] The joint loss module 103 is configured to analyze the normalized image to output a video quality score, and the joint loss module includes at least two feature analysis modules.
[0083] As can be seen from the above solution, for the video image frame to be evaluated, after obtaining a compressed image by performing information compression on the video image frame through a depth feature extraction network, performing spatio-temporal normalization processing on the compressed image to obtain a normalized image, and then using the joint loss module to obtain the final video quality score determined in multiple dimensions, it solves the problem of poor score rationality caused by using a single dimension such as subjective picture quality assessment to determine video quality in the prior art, makes the video quality assessment result more accurate, and can provide richer decision support for subsequent processing. It is convenient to optimize the picture quality of video images with the maximum efficiency in subsequent processing and further improves the visual experience of the audience when watching.
[0084] In a possible embodiment, the depth feature extraction network includes a plurality of consecutive feature extraction modules and downsampling modules, and the spatio-temporal feature normalization module 102 is specifically configured to:
[0085] Perform mean calculation and feature extraction on the temporal feature and spatial feature of the compressed image to obtain a normalized image.
[0086] In a possible embodiment, the joint loss module includes a distortion analysis module and a subjective analysis module. The joint loss module 103 is specifically configured to:
[0087] Perform distortion analysis processing on the normalized image through the distortion analysis module to obtain a predicted output;
[0088] Input the predicted output and the spatio-temporal features of the normalized image into the subjective analysis module to output a video quality score.
[0089] In a possible embodiment, the distortion analysis module includes at least two distortion analysis sub-modules. The joint loss module 103 is specifically configured to:
[0090] Perform distortion analysis processing on the normalized image through each of the distortion analysis sub-modules to obtain respective corresponding predicted outputs;
[0091] Input the predicted output corresponding to each of the distortion analysis sub-modules and the spatio-temporal features of the normalized image into the subjective analysis module to output a video quality score.
[0092] In a possible embodiment, the device further includes a model training module 104, which is used for:
[0093] Before analyzing the normalized image through the joint loss module and outputting a video quality score, obtain different tag values of the manually annotated video samples, and train the joint loss module according to the different tag values and the corresponding relationship between the tag values and the feature analysis module.
[0094] In a possible embodiment, the device further includes a data cleaning module 105, which is used for:
[0095] Before obtaining different tag values of the manually annotated video samples, obtain the original video data, preprocess the original video data to obtain video information in a unified format, and generate a feature frequency histogram corresponding to the video information based on different picture quality features;
[0096] Filter the original video data based on the feature frequency histogram to obtain video samples, where the picture quality features include at least one of bit rate, resolution, brightness, chroma, aspect ratio, and contrast.
[0097] In a possible embodiment, the device further includes a manual annotation module 106, which is used for:
[0098] Before obtaining different tag values of the manually annotated video samples, receive and store the tag values of the manual annotation of the current video sample, where the tag values include at least two of a noise value, a color value, a sharpness value, a blurriness value, a brightness value, and a subjective value.
[0099] In a possible embodiment, the joint loss module 103 is specifically configured to:
[0100] Analyze the normalized image through the joint loss module to output the total score value of the video image frame and the sub-feature score values output by each feature analysis module.
[0101] Figure 8 FIG. is a schematic structural diagram of a no-reference video quality determination device provided by an embodiment of the present invention. As Figure 8 shown, the device includes a processor 201, a memory 202, an input device 203, and an output device 204; the number of processors 201 in the device can be one or more. Figure 8 Here, one processor 201 is taken as an example; the processor 201, the memory 202, the input device 203, and the output device 204 in the device can be connected through a bus or other means. Figure 8 Here, connection through a bus is taken as an example. The memory 202, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the no-reference video quality determination method in the embodiment of the present invention. The processor 201 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 202, that is, implements the above-mentioned no-reference video quality determination method. The input device 203 can be used to receive input digital or character information and generate key signal inputs related to the user settings and function controls of the device. The output device 204 can include a display device such as a display screen.
[0102] An embodiment of the present invention further provides a storage medium containing computer-executable instructions, where the computer-executable instructions are used to execute a no-reference video quality determination method described in the above embodiment when executed by a computer processor, specifically including:
[0103] Obtain a video image frame to be evaluated, and compress the information of the video image frame through a deep feature extraction network to obtain a compressed image;
[0104] Perform spatio-temporal normalization processing on the compressed image to obtain a normalized image;
[0105] Analyze the normalized image through a joint loss module to output a video quality score value, where the joint loss module includes at least two feature analysis modules.
[0106] It should be noted that in the above embodiments of the video quality determination device without reference, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present invention.
[0107] Note that the above is only the preferred embodiment of the embodiments of the present invention and the applied technical principles. Those skilled in the art will understand that the embodiments of the present invention are not limited to the specific embodiments described herein, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the embodiments of the present invention. Therefore, although the embodiments of the present invention have been described in more detail through the above embodiments, the embodiments of the present invention are not limited to the above embodiments only. Without departing from the concept of the embodiments of the present invention, more other equivalent embodiments can be included, and the scope of the embodiments of the present invention is determined by the scope of the appended claims.
Claims
1. A video quality determination method without reference, characterized in that, Comprising: Obtain a video image frame to be evaluated, and perform information compression on the video image frame through a depth feature extraction network to obtain a compressed image; Perform spatio-temporal normalization processing on the compressed image to obtain a normalized image; Analyze the normalized image through a joint loss module to output a video quality score. The joint loss module includes a distortion analysis module and a subjective analysis module. The distortion analysis module includes at least two distortion analysis sub-modules. Among them, the distortion analysis sub-module includes a noise feature branch, a sharpness feature branch, a blurriness feature branch, a brightness feature branch, and a chromaticity feature branch. Analyzing the normalized image through the joint loss module to output a video quality score includes: using a non-linear layer through each feature branch to map the spatio-temporal features in the normalized image to the desired image quality features of their respective corresponding branches to obtain feature scores, and inputting the obtained noise feature score, sharpness feature score, blurriness feature score, brightness feature score, and chromaticity feature score, together with the spatio-temporal features of the normalized image, into the subjective analysis module. The subjective analysis module maps the input noise feature score, sharpness feature score, blurriness feature score, brightness feature score, chromaticity feature score, and spatio-temporal features of the normalized image to one dimension through a single-layer linear transformation to obtain the final video quality score.
2. The method for determining video quality without reference according to claim 1, characterized in that The depth feature extraction network includes a plurality of consecutive feature extraction modules and downsampling modules. Performing spatio-temporal normalization processing on the compressed image to obtain a normalized image includes: Perform mean calculation and feature extraction on the temporal features and spatial features of the compressed image to obtain a normalized image.
3. The method for determining video quality without reference according to claim 1, wherein Before analyzing the normalized image through the joint loss module to output a video quality score, it further includes: Obtain different label values of an artificially annotated video sample, and train the joint loss module according to the different label values and the corresponding relationship between the label values and the feature analysis module.
4. The method for determining video quality without reference according to claim 3, wherein Before obtaining different label values of the artificially annotated video sample, it further includes: Obtain original video data, perform preprocessing on the original video data to obtain video information in a unified format, and generate a feature frequency histogram corresponding to the video information based on different image quality features; Filter the original video data based on the feature frequency histogram to obtain video samples, where the image quality features include at least one of bit rate, resolution, brightness, chromaticity, aspect ratio, and contrast.
5. The reference - free video quality determination method according to claim 4, characterized in that, Before obtaining different label values of the artificially annotated video sample, it further includes: Receive and store the label values of the artificial annotation of the current video sample, where the label values include at least two of noise value, color value, sharpness value, blurriness value, brightness value, and subjective value.
6. The method for determining video quality without reference according to any one of claims 1-5, characterized in that Analyzing the normalized image through the joint loss module to output a video quality score includes: Analyze the normalized image through the joint loss module to output the total score of the video image frame, and the sub-feature scores output by each feature analysis module.
7. Video quality determination device without reference, characterized in that, Comprising: A deep feature extraction module, configured to obtain video image frames to be evaluated, and perform information compression on the video image frames to obtain compressed images; A spatio-temporal feature normalization module, configured to perform spatio-temporal normalization processing on the compressed images to obtain normalized images; A joint loss module, configured to analyze the normalized images and output video quality scores. The joint loss module includes a distortion analysis module and a subjective analysis module. The distortion analysis module includes at least two distortion analysis sub-modules. Among them, the distortion analysis sub-module includes a noise feature branch, a sharpness feature branch, a blurriness feature branch, a brightness feature branch, and a chrominance feature branch. Specifically, the joint loss module is configured to: use a non-linear layer through each feature branch to map the spatio-temporal features in the normalized images to the respective desired image quality features in the corresponding branches to obtain feature scores, and input the obtained noise feature scores, sharpness feature scores, blurriness feature scores, brightness feature scores, and chrominance feature scores, together with the spatio-temporal features of the normalized images, into the subjective analysis module. The subjective analysis module maps the input noise feature scores, sharpness feature scores, blurriness feature scores, brightness feature scores, chrominance feature scores, and spatio-temporal features of the normalized images to one dimension through a single-layer linear transformation to obtain the final video quality scores.
8. A no-reference video quality determination device, the device comprising: One or more processors; A storage device, configured to store one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the no-reference video quality determination method according to any one of claims 1-6.
9. A storage medium storing computer-executable instructions, where the computer-executable instructions are used to execute the no-reference video quality determination method according to any one of claims 1-6 when executed by a computer processor.
Citation Information
Patent Citations
No-reference image quality evaluation method and device
CN107123122A