A no-reference video quality assessment method, device, equipment and storage medium

By converting and feature extraction of video frames, long-term dependencies are established, the problem of reference-free video quality evaluation is solved, and more accurate video quality evaluation is achieved.

CN113888502BActive Publication Date: 2025-07-22BEIJING JIUSHI TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111155150.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2025-07-22
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively evaluate the quality of video without references, especially in the absence of references in Internet video browsing, which cannot accurately evaluate the degree of distortion of videos.

Method used

By obtaining video frames, converting them into grayscale images and using LBP operators to generate texture maps, combining neural networks to extract content dependencies and texture features, establish long-term dependencies, and perform weighted calculations to evaluate video quality.

Benefits of technology

It improves the accuracy of reference-free video quality evaluation, simulates the content dependence and time hysteresis effects of the human eye, and improves the evaluation performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113888502B_ABST
    Figure CN113888502B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, device and storage medium for video quality assessment without reference. The present invention obtains a plurality of video frames, performs conversion processing on the video frames to obtain a texture map of each video frame, performs feature extraction processing on the video frames and the texture maps to obtain content-dependent features of each video frame and texture features of each texture map, establishes long-term dependencies based on the content-dependent features and the texture features to obtain frame quality, determines memory quality elements and current quality elements according to the frame quality, and performs weighted calculation processing according to the memory quality elements and the current quality elements to obtain a video quality evaluation result; citing texture features can better fit the distortion sensitivity, and using content-dependent features and establishing long-term dependencies to simulate the content dependence and time lag effect of the human eye is beneficial to improving the performance of the assessment method and can be widely applied to the video field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video, and in particular to a no-reference video quality assessment method, apparatus, device and storage medium. Background Art

[0002] With the rapid development of electronic devices such as mobile phones and tablet computers, people are no longer satisfied with just viewing pictures, but are more inclined to video services. Nowadays, most electronic devices support receiving and playing high-definition videos. Compared with pictures, however, videos are more likely to be distorted during the processes of acquisition, transmission, and storage. Therefore, video quality assessment has become an important research topic in computer vision today. At the same time, in practical applications, the videos that people browse on the Internet do not have any reference videos. Therefore, it is necessary to provide relevant technologies for no-reference video quality assessment.

[0003] Generally, distortion includes natural distortion and synthetic distortion. Natural distortion is a mixture of distortions such as insufficient exposure, overexposure, blurring caused by the movement of the shooter, and compression errors during the shooting process. Synthetic distortion includes white noise, Gaussian blur, JPEG2000, salt-and-pepper noise, global contrast reduction, etc., which are artificially caused distortions. And both natural distortion and synthetic distortion widely appear in videos. Therefore, a no-reference video quality assessment method that can evaluate any type of distortion is needed. Summary of the Invention

[0004] In view of this, in order to solve the above technical problems, the purpose of the present invention is to provide a no-reference video quality assessment method, apparatus, device and storage medium.

[0005] The technical solution adopted in the embodiments of the present invention is:

[0006] A no-reference video quality assessment method, comprising:

[0007] Obtaining a plurality of video frames;

[0008] Performing conversion processing on the video frames to obtain a texture map of each video frame;

[0009] Performing feature extraction processing on the video frames and the texture maps to obtain content-dependent features of each video frame and texture features of each texture map;

[0010] Establishing a long-term dependence according to the content-dependent features and the texture features to obtain frame quality;

[0011] Determining a memory quality element and a current quality element according to the frame quality;

[0012] Perform weighted calculation processing based on the memory quality element and the current quality element to obtain a video quality evaluation result.

[0013] Further, the conversion processing of the video frame to obtain a texture map of each video frame includes:

[0014] Convert the video frame into a grayscale image;

[0015] Convert the grayscale image through an LBP operator to obtain a texture map of each video frame.

[0016] Further, the feature extraction processing of the video frame and the texture map to obtain the content-dependent feature of each video frame and the texture feature of each texture map includes:

[0017] Input the video frame into a neural network for first feature extraction to obtain the content-dependent feature of each video frame;

[0018] Input the texture map into a two-stream network for second feature extraction to obtain the texture feature of each texture map; the network parameters of the two-stream network are the same as those of the neural network.

[0019] Further, establishing a long-term dependence based on the content-dependent feature and the texture feature to obtain the frame quality includes:

[0020] Perform first global average pooling and first global standard deviation pooling on the content-dependent feature;

[0021] Perform second global average pooling and second global standard deviation pooling on the texture feature;

[0022] Perform feature merging on the first global average pooling result, the first global standard deviation pooling result, the second global average pooling result, and the second global standard deviation pooling result to obtain an image feature corresponding to each video frame;

[0023] Establish a long-term dependence based on the image feature to obtain the frame quality.

[0024] Further, establishing a long-term dependence based on the image feature to obtain the frame quality includes:

[0025] Perform dimensionality reduction processing on the image feature;

[0026] Input the dimensionality reduction processing result into a recurrent neural network to establish a long-term dependence to obtain the frame quality.

[0027] Further, determining the memory quality element and the current quality element based on the frame quality includes:

[0028] Determine the current frame from the video frames;

[0029] Take the minimum frame quality corresponding to a preset number of video frames before the current frame as the memory quality element;

[0030] Perform a summation process based on the frame quality corresponding to a preset number of video frames after the current frame, and determine a first weight parameter according to the frame quality corresponding to a preset number of video frames after the current frame and the summation result;

[0031] Perform weighting based on the frame quality corresponding to a preset number of video frames after the current frame and the first weight parameter to obtain the current quality element.

[0032] Furthermore, the weighted calculation process based on the memory quality element and the current quality element to obtain a video quality evaluation result includes:

[0033] Perform weighted calculation according to a second weight parameter, the memory quality element, a third weight parameter, and the current quality element to obtain a subjective frame quality score;

[0034] Calculate an average value according to the subjective frame quality score and the total number of video frames to obtain the video quality evaluation result.

[0035] An embodiment of the present invention further provides a no-reference video quality assessment device, including:

[0036] An acquisition module, configured to acquire a plurality of video frames;

[0037] A conversion module, configured to perform conversion processing on the video frames to obtain a texture map of each video frame;

[0038] A feature extraction module, configured to perform feature extraction processing on the video frames and the texture maps to obtain content-dependent features of each video frame and texture features of each texture map;

[0039] A establishment module, configured to establish long-term dependencies according to the content-dependent features and the texture features to obtain frame quality;

[0040] A determination module, configured to determine a memory quality element and a current quality element according to the frame quality;

[0041] An evaluation module, configured to perform weighted calculation processing according to the memory quality element and the current quality element to obtain a video quality evaluation result.

[0042] An embodiment of the present invention further provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method.

[0043] An embodiment of the present invention further provides a computer-readable storage medium. At least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the method.

[0044] The beneficial effects of the present invention are as follows: By obtaining a plurality of video frames, performing conversion processing on the video frames to obtain a texture map of each video frame, performing feature extraction processing on the video frames and the texture maps to obtain a content-dependent feature of each video frame and a texture feature of each texture map, establishing a long-term dependence according to the content-dependent feature and the texture feature to obtain a frame quality, determining a memory quality element and a current quality element according to the frame quality, and performing weighted calculation processing according to the memory quality element and the current quality element to obtain a video quality evaluation result; Referencing the texture feature can better fit the distortion sensitivity, and using the content-dependent feature and establishing a long-term dependence to simulate the content dependence, texture masking and time lag effects of the human eye is beneficial to improving the performance of the evaluation method and the accuracy of the evaluation. Description of the Drawings

[0045] Figure 1 It is a schematic flowchart of the steps of the method for evaluating the quality of a video without a reference video in a specific embodiment of the present invention;

[0046] Figure 2 It is a schematic diagram comparing the video quality evaluation result and the human subjective evaluation score in a specific embodiment of the present invention;

[0047] Figure 3 It is a schematic diagram of the SROCC results of KoNViD-1k, CVD2014 and LSVQ in a specific embodiment of the present invention;

[0048] Figure 4 It is a schematic diagram of the ratio of each dataset and the number of video frames in the video set in a specific embodiment of the present invention. Detailed Embodiments

[0049] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0050] The terms "first", "second", "third", "fourth", etc. in the specification and claims of this application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0051] Referring to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0052] As Figure 1 shown, an embodiment of the present invention provides a no-reference video quality assessment method, including steps S100 - 5600:

[0053] S100. Obtain a plurality of video frames.

[0054] In the embodiment of the present invention, a video frame refers to a frame extracted from a video. Optionally, the input video can be frame-divided to obtain a plurality of video frames, or directly obtain the video frames obtained by pre-processing the video from the Internet or a storage container, without specific limitation.

[0055] S200. Perform conversion processing on the video frames to obtain a texture map of each video frame.

[0056] Optionally, step S200 is executed by a texture information detection module, including steps S210 - S220:

[0057] S210. Convert the video frame into a grayscale image.

[0058] Specifically, convert each video frame from an RGB image F to a corresponding grayscale image I:

[0059] I(x, y) = 0.1140 * F_B(x, y) + 0.5870 * F_G(x, y) + 0.2989 * F_R(x, y)

[0060] Wherein, I(x, y) represents the result after conversion of the pixel point with coordinates (x, y) in the grayscale image I, F_B(x, y), F_G(x, y), and F_R(x, y) respectively represent the values of the pixel points in the blue (B), green (G), and red (R) channels of the original image (video frame) F at the point (x, y), and the three weighting coefficients 0.1140 (the first weighting coefficient), 0.5870 (the second weighting coefficient), and 0.2989 (the third weighting coefficient) are adjusted according to the human brightness perception system, and may be other values in other embodiments.

[0061] S220. Convert the grayscale image through the LBP operator to obtain the texture map of each video frame.

[0062] In the embodiment of the present invention, the converted grayscale image is converted frame by frame using the LBP operator to obtain the texture map L of each video frame:

[0063]

[0064] Wherein, LBP P,R (x c , y c ) represents the result of the pixel point after conversion of the coordinate point (x c , y c ) in the texture map L, P represents the total number of sampling points, R represents the sampling radius, I(p) and I(c) are respectively the grayscale values of the p1-th sampling point and the point (x c , y c ), sampled from the grayscale image l, s(X) represents the threshold function, X is I(p) - I(c), and is expressed as follows:

[0065]

[0066] The coordinates of the p1-th sampling point are calculated using the following formula:

[0067]

[0068] S300. Perform feature extraction processing on the video frame and the texture map to obtain the content-dependent feature of each video frame and the texture feature of each texture map.

[0069] Optionally, step S300 includes steps S310 - S320:

[0070] S310. Input the video frame into a neural network for the first feature extraction to obtain the content-dependent feature of each video frame.

[0071] In the embodiment of the present invention, assuming that the video has T frames, input T original frames (video frames) into the neural network for the first feature extraction to obtain the content-dependent feature M of each video frame. t1 It should be noted that, in the embodiment of the present invention, the neural network for feature extraction is exemplified by a deep neural network (CNN), such as ResNet-50. In other embodiments, other neural networks can be used, and specific details are not limited. Specifically:

[0072] M t1 = CNN1(F)

[0073] where F represents the original frame (video frame) of the video, and CNN1() is the first feature extraction.

[0074] S320. Input the texture map into a two-stream network for the second feature extraction to obtain the texture feature of each texture map.

[0075] In the embodiment of the present invention, input T texture maps L into the two-stream network for the second feature extraction to obtain the texture feature M of each texture map. t2 Specifically:

[0076] M t2 = CNN2(L)

[0077] where L is the texture map, and CNN2() is the second feature extraction.

[0078] It should be noted that the two-stream network and the neural network have the same network parameters. Optionally, the network parameters include but are not limited to data processing (or preprocessing) related parameters, training process and training related parameters, or network related parameters. For example, data processing (or preprocessing) related parameters include but are not limited to parameters for enriching the database (enrich data), parameters for data generalization processing (feature normalization and scaling), and parameters for BN processing (batch normalization); training process and training related parameters include but are not limited to training momentum, learning rate, decay function, weight initialization, and regularization related methods; network related parameters include but are not limited to classifier selection parameters, number of neurons, number of filters, and number of network layers.

[0079] S400. Establish long-term dependencies based on the content-dependent feature and the texture feature to obtain the frame quality.

[0080] Optionally, step S400 includes steps S410 - S440:

[0081] S410. Perform the first global average pooling and the first global standard deviation pooling on the content-dependent features.

[0082] S420. Perform the second global average pooling and the second global standard deviation pooling on the texture features.

[0083] Specifically:

[0084]

[0085]

[0086]

[0087]

[0088] Among them, is the result of the first global average pooling, is the result of the first global standard deviation pooling, is the result of the second global average pooling, is the second global standard deviation pooling, GP mean1 is the first global average pooling, GP std1 is the first global standard deviation pooling, GP mean2 is the second global average pooling, GP std2 is the second global standard deviation pooling, M t1 is the content-dependent feature, M t2 is the texture feature.

[0089] S430. Perform feature merging on the first global average pooling result, the first global standard deviation pooling result, the second global average pooling result, and the second global standard deviation pooling result to obtain the image feature corresponding to each video frame.

[0090] Specifically, after performing the above global average pooling and the above global standard deviation pooling for dimensionality reduction, the first global average pooling result, the first global standard deviation pooling result, the second global average pooling result, and the second global standard deviation pooling result are merged by a merging function (Concat) to obtain the image feature V corresponding to each video frame. t :

[0091]

[0092] It should be noted that steps S300 and steps S410 - S430 are executed by the feature extraction module.

[0093] S440. Establish long-term dependencies based on the image features to obtain the frame quality.

[0094] It should be noted that steps S440 and steps S500 - S600 are executed by a time lag module. There are long - term dependencies in video quality assessment problems. Generally speaking, people remember frames with poor quality in the past and lower their expectations for subsequent frames. Even if the frame quality has returned to an acceptable level, the task of the time lag module is to establish long - term dependencies of the video by introducing a gated recurrent unit (GRU), and then use a linear function to combine the influence of past frames and subsequent frames on video quality to obtain the final video score (i.e., the video quality evaluation result).

[0095] Specifically, step S440 includes steps S4401 - S4402:

[0096] S4401. Perform dimensionality reduction processing on the image features.

[0097] Optionally, input the image feature V t into a fully - connected layer (FC) for dimensionality reduction processing to remove redundant information.

[0098] S4402. Input the result of dimensionality reduction processing into a recurrent neural network to establish long - term dependencies and obtain the frame quality.

[0099] In the embodiments of the present invention, the recurrent neural network takes GRU (Gate Recurrent Unit) as an example. In other embodiments, other recurrent neural networks can be used, without specific limitation. Specifically, input the results of dimensionality reduction processing FC(V t ) and FC(V t-1 ) of the t - th frame and the (t - 1)-th frame images into GRU to establish long - term dependencies, and then obtain the frame quality (i.e., the frame - level quality score q t ) after dimensionality reduction processing:

[0100] q t = FC(GRU{FC(V t ),GRU{FC(V t-1 )}})

[0101] S500. Determine the memory quality element and the current quality element according to the frame quality.

[0102] Specifically, step S500 includes steps S510 - S540:

[0103] S510. Determine the current frame from the video frames.

[0104] Optionally, assume that the currently processed video frame is the current frame t (i.e., the t - th video frame).

[0105] S520. Use the minimum frame quality corresponding to a preset number of video frames before the current frame as the memory quality element.

[0106] In the embodiment of the present invention, first, a memory quality element x is defined t as the score (minimum frame quality) of the frame with the worst quality among the previous τ (predetermined number) frames of the current frame t:

[0107]

[0108] where V pre ={max(1, t - τ),..., t - 2, t - 1}, representing the value set of the previous τ frames, q n represents the frame quality of the nth frame, q k1 represents the frame quality of the k1th frame.

[0109] S530. Perform a summation process according to the frame qualities corresponding to a predetermined number of video frames after the current frame, and determine a first weight parameter according to the frame qualities corresponding to a predetermined number of video frames after the current frame and the result of the summation process.

[0110] Specifically, use a weighted quality score to assign more weights to the frames with relatively poor quality among the subsequent τ (predetermined number) frames of the current frame t. Optionally, use the softmin function to determine the first weight parameter:

[0111]

[0112] where is the first weight parameter, the denominator represents the result of the summation process, V next ={t, t + 1,..., min(t + τ, T)}, representing the value set of the subsequent τ frames of the current frame t, T represents the total number of video frames, q k2 represents the frame quality of the k2th frame, q j represents the frame quality of the jth frame.

[0113] S540. Perform weighting according to the frame qualities corresponding to a predetermined number of video frames after the current frame and the first weight parameter to obtain the current quality element.

[0114] In the embodiment of the present invention, the current quality element y is defined t :

[0115]

[0116] where is the first weight parameter, q k2 is the frame quality of the k2th frame.

[0117] S600. Perform a weighted calculation process according to the memory quality element and the current quality element to obtain a video quality evaluation result.

[0118] Specifically, step S600 includes steps S610 - S620:

[0119] S610. Perform a weighted calculation based on the second weight parameter, the memory quality element, the third weight parameter, and the current quality element to obtain the subjective frame quality score.

[0120] In the embodiment of the present invention, the memory quality element x t and the current quality element y t are linearly combined and weighted to calculate the subjective frame quality score q′ t :

[0121] q′ t = βx t +(1 - β)y t

[0122] where β is the second weight parameter and 1 - β is the third weight parameter. It should be noted that the second weight parameter β is a hyperparameter that balances the current quality element and the memory quality element, representing the influence of the current quality element and the memory quality element on the final video score.

[0123] S620. Calculate the average value based on the subjective frame quality score and the total number of video frames to obtain the video quality evaluation result.

[0124]

[0125] where Q is the final score, i.e., the video quality evaluation result. The larger Q is, the better the video quality evaluation result. T is the total number of video frames, and q′ t is the subjective frame quality score.

[0126] In the embodiment of the present invention, the no-reference video quality assessment method is tested on three mainstream video sets, including KoNViD-1k, CVD2014, and LSVQ. Among them, LSVQ is the largest video quality assessment data set containing 39,072 videos so far and is representative. At the same time, SROCC and PLCC are used as the comparison measurement indicators for the prediction score and the human subjective evaluation score. 80% of each data set is used as the training set, and 20% is used as the test set. Each result is repeated ten times and the average value is taken. In KoNViD-1k, SROCC is 0.870 and PLCC is 0.860; in CVD2014, SROCC is 0.920 and PLCC is 0.920; in LSVQ, SROCC is 0.842 and PLCC is 0.842. All three data sets have achieved the current state-of-the-art performance. As Figure 2As shown, several videos randomly selected from three datasets are given, and the human subjective evaluation scores and the video quality evaluation results (the predicted values of the present invention) obtained by the no-reference video quality evaluation method of the present invention are compared. It can be seen that the predicted values of the present invention are very close to each other, indicating that the no-reference video quality evaluation method of the present invention has a very good evaluation effect.

[0127] As Figure 3 shown, through a large number of experiments, the parameter selections of three currently mainstream video quality evaluation datasets are analyzed, and the SROCC results of each dataset under different parameters are obtained. For the KoNViD-1k dataset, the model performance is the highest, which is 0.870 when the values of τ and β in the time lag module are 44 and 0.52 respectively; for the CVD2014 dataset, the values of τ and β are 68 and 0.5 respectively when the model performance is the highest, while for the LSVQ dataset, the values of τ and β are 20 and 0.5 respectively. At the same time, through the analysis of the number of frames of each video in different datasets, as Figure 4 shown, it is a schematic diagram of the display ratio and the number of frames of the videos in the video set. It can be found that when the overall number of frames of the dataset is larger, a larger value of τ should be selected. This may be because as the number of video frames increases, people will remember more video information, while the value of β should be kept around 0.5. This may be because the influence of past frames and the next frames on people is almost equal.

[0128] The no-reference video quality evaluation method of the embodiment of the present invention: 1) A new two-stream deep learning network is proposed to extract the texture features and content-dependent features of each frame of the video, and the deep learning network is trained accordingly for the sensitivity to texture region distortion, which better fits the human visual system and improves the sensitivity of the model to distortion information; 2) Different parameters are used for three currently mainstream datasets through a large number of experiments, and a generally applicable parameter selection rule is obtained from them, which can better improve the model performance.

[0129] The embodiment of the present invention also provides a no-reference video quality evaluation device, including:

[0130] An acquisition module, configured to acquire a plurality of video frames;

[0131] A conversion module, configured to perform conversion processing on the video frames to obtain a texture map of each video frame;

[0132] A feature extraction module, configured to perform feature extraction processing on the video frames and the texture maps to obtain the content-dependent features of each video frame and the texture features of each texture map;

[0133] A building module, configured to establish a long-term dependence according to the content-dependent features and the texture features to obtain the frame quality;

[0134] A determination module, configured to determine a memory quality element and a current quality element according to the frame quality;

[0135] An evaluation module, configured to perform weighted calculation processing according to the memory quality element and the current quality element to obtain a video quality evaluation result.

[0136] The content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0137] An embodiment of the present invention further provides an electronic device, which includes a processor and a memory. The memory stores at least one instruction, at least one program, a code set or an instruction set. The at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the no-reference video quality assessment method in the foregoing embodiments. The electronic device in the embodiment of the present invention includes, but is not limited to, any intelligent terminal such as a mobile phone, a tablet computer, a computer, and an in-vehicle computer.

[0138] The content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0139] An embodiment of the present invention further provides a computer-readable storage medium, which stores at least one instruction, at least one program, a code set or an instruction set. The at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the no-reference video quality assessment method in the foregoing embodiments.

[0140] An embodiment of the present invention further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the no-reference video quality assessment method in the foregoing embodiments.

[0141] In the description of the present application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0142] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the relationship between related objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B may be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c may be single or plural.

[0143] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms. The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0144] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. The foregoing storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memories (ROM for short), random access memories (RAM for short), magnetic disks, or optical discs.

[0145] The above, the above embodiments are only used to illustrate the technical solutions of this application, rather than limiting it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of this application.

Claims

1. A no-reference video quality assessment method, characterized in that, Including: Obtain a plurality of video frames; Perform conversion processing on the video frames to obtain a texture map for each of the video frames; Perform feature extraction processing on the video frames and the texture maps to obtain content-dependent features for each of the video frames and texture features for each of the texture maps; Establish long-term dependencies based on the content-dependent features and the texture features to obtain frame quality; Determine a memory quality element and a current quality element according to the frame quality; Perform weighted calculation processing based on the memory quality element and the current quality element to obtain a video quality evaluation result; The performing conversion processing on the video frames to obtain a texture map for each of the video frames includes: Convert the video frames into grayscale images; Convert the grayscale images through an LBP operator to obtain a texture map for each of the video frames; The performing feature extraction processing on the video frames and the texture maps to obtain content-dependent features for each of the video frames and texture features for each of the texture maps includes: Input the video frames into a neural network for first feature extraction to obtain content-dependent features for each of the video frames; Input the texture maps into a two-stream network for second feature extraction to obtain texture features for each of the texture maps; the two-stream network has the same network parameters as the neural network; The establishing long-term dependencies based on the content-dependent features and the texture features to obtain frame quality includes: Perform first global average pooling and first global standard deviation pooling on the content-dependent features; Perform second global average pooling and second global standard deviation pooling on the texture features; Perform feature merging on the results of the first global average pooling, the results of the first global standard deviation pooling, the results of the second global average pooling, and the results of the second global standard deviation pooling through a merging function to obtain image features corresponding to each of the video frames, and the expression of the merging function is: Among them, V t is the image feature, is the first global average pooling result, is the first global standard deviation pooling result, is the second global average pooling result, is the second global standard deviation pooling result; Establish long-term dependencies based on the image features to obtain frame quality.

2. The no-reference video quality assessment method according to claim 1, wherein: The establishing long-term dependencies based on the image features to obtain frame quality includes: Perform dimensionality reduction processing on the image features; Input the dimensionality reduction processing results into a recurrent neural network to establish long-term dependencies to obtain frame quality.

3. The no-reference video quality assessment method according to claim 1, wherein: The determining a memory quality element and a current quality element according to the frame quality includes: Determine a current frame from the video frames; Use the minimum frame quality corresponding to a preset number of video frames before the current frame as the memory quality element; Perform a summation process on the frame qualities corresponding to a preset number of video frames after the current frame, and determine a first weight parameter according to the frame qualities corresponding to a preset number of video frames after the current frame and the summation process result; Perform weighting based on the frame qualities corresponding to a preset number of video frames after the current frame and the first weight parameter to obtain the current quality element.

4. The no-reference video quality assessment method according to claim 1, characterized in that: The performing weighted calculation processing based on the memory quality element and the current quality element to obtain a video quality evaluation result includes: Perform weighted calculation according to a second weight parameter, the memory quality element, a third weight parameter, and the current quality element to obtain a subjective frame quality score; Calculate an average value based on the subjective frame quality score and the total number of video frames to obtain the video quality evaluation result.

5. A no-reference video quality assessment device, characterized in that, Including: An acquisition module for acquiring a plurality of video frames; A conversion module for performing conversion processing on the video frames to obtain a texture map for each of the video frames; A feature extraction module for performing feature extraction processing on the video frames and the texture maps to obtain content-dependent features for each of the video frames and texture features for each of the texture maps; A building module for establishing a long-term dependence based on the content-dependent features and the texture features to obtain a frame quality; A determination module for determining a memory quality element and a current quality element based on the frame quality; An evaluation module for performing weighted calculation processing based on the memory quality element and the current quality element to obtain a video quality evaluation result; The performing conversion processing on the video frames to obtain a texture map for each of the video frames includes: Converting the video frames into grayscale images; Converting the grayscale images through an LBP operator to obtain a texture map for each of the video frames; The performing feature extraction processing on the video frames and the texture maps to obtain content-dependent features for each of the video frames and texture features for each of the texture maps includes: Inputting the video frames into a neural network for first feature extraction to obtain content-dependent features for each of the video frames; Inputting the texture maps into a two-stream network for second feature extraction to obtain texture features for each of the texture maps; the two-stream network has the same network parameters as the neural network; The establishing a long-term dependence based on the content-dependent features and the texture features to obtain a frame quality includes: Performing first global average pooling and first global standard deviation pooling on the content-dependent features; Performing second global average pooling and second global standard deviation pooling on the texture features; Performing feature merging on the results of the first global average pooling, the results of the first global standard deviation pooling, the results of the second global average pooling, and the results of the second global standard deviation pooling through a merging function to obtain an image feature corresponding to each of the video frames, and the expression of the merging function is: Among them, V t is the image feature, is the first global average pooling result, is the first global standard deviation pooling result, is the second global average pooling result, is the second global standard deviation pooling result; Establishing a long-term dependence based on the image features to obtain a frame quality.

6. An electronic device, characterized in that, The electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, At least one instruction, at least one program, a code set, or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • No-reference image quality grading evaluation method and device based on visual fusion features

    CN111507426A