Video quality detection method, device, equipment and medium

By performing multi-level feature extraction and image quality evaluation index stitching on high dynamic range videos, the problem of insufficient accuracy and interpretability in video quality detection in traditional methods is solved, and efficient and accurate video quality evaluation is achieved.

CN121685462APending Publication Date: 2026-03-17MALANSHAN AUDIO & VIDEO LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511862020.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Traditional video quality assessment methods rely on pixel-domain features, resulting in insufficient accuracy and interpretability of detection results, making it difficult to adapt to the complex brightness levels and color performance characteristics of high dynamic range videos.

Method used

By preprocessing the reference video and the distorted video, they are converted into a sequence of target image frames. Multi-level feature maps are extracted frame by frame, including pixel domain, shallow features and deep features. Image quality evaluation indicators are calculated and spliced ​​to form a multi-dimensional quality-aware feature vector. Finally, the vector is input into a quality regression network for aggregation and mean processing.

Benefits of technology

It achieves efficient and accurate quality evaluation of high dynamic range videos, comprehensively captures video quality features, and improves the credibility and persuasiveness of detection results. It is suitable for high dynamic range video coding optimization, streaming media service quality monitoring, and video processing algorithm evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121685462A_ABST
    Figure CN121685462A_ABST
Patent Text Reader

Abstract

The invention discloses a video quality detection method, a video quality detection device, video quality detection equipment and a medium, and relates to the technical field of computer vision, a reference video and a distorted video are preprocessed and converted into a target image frame sequence, and then a multi-level feature map composed of original features, shallow layer features and deep layer features of a pixel domain is extracted frame by frame; and calculating image quality evaluation indexes at each feature level, splicing to form a multi-dimensional quality perception feature vector, inputting the multi-dimensional quality perception feature vector into a quality regression network to obtain single-frame quality data, and performing aggregation mean processing to output video overall quality data. Therefore, the comprehensive feature coverage from the pixel level to the semantic level is realized by utilizing the original information of the pixel domain, the shallow detail texture and the deep semantic information, and the credibility of the detection result is guaranteed through the interpretability and effectiveness of the image quality evaluation index. The method is efficient and accurate, can be widely applied to scenes of high-dynamic-range video coding optimization, streaming media service quality monitoring and the like, and provides a reliable quality basis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a video quality detection method, apparatus, device, and medium. Background Technology

[0002] In high dynamic range (HDR) video quality assessment tasks, traditional methods generally suffer from limitations such as relying on single features and having a one-sided evaluation dimension, making it difficult to adapt to the rich brightness levels and complex color representation characteristics of HDR videos. Most solutions depend solely on pixel-domain features, capturing only the most basic surface information such as brightness distribution and color values, completely ignoring the deep semantic relationships and global structural patterns inherent in the video content. When faced with complex distortion scenarios, the evaluation accuracy is insufficient, making it difficult to balance accuracy, comprehensiveness, and interpretability, and thus failing to meet the high-precision requirements of some practical scenarios. Summary of the Invention

[0003] The purpose of this invention is to provide a video quality detection method, apparatus, device, and medium that can solve the technical problem that traditional video quality evaluation methods rely solely on pixel domain features, resulting in insufficient accuracy and interpretability of detection results.

[0004] To address the aforementioned technical problems, this invention provides a video quality detection method, comprising:

[0005] The reference video and the distorted video are preprocessed and converted into a sequence of target image frames.

[0006] Extract multi-level feature maps containing pixel-domain original features, shallow features, and deep features frame by frame from the target image frame sequence;

[0007] The image quality evaluation index is calculated at each level using the multi-level feature map, and the image quality evaluation indexes of all levels are concatenated to form a multi-dimensional quality-perceived feature vector.

[0008] The quality-aware feature vector is input into the quality regression network to obtain single-frame quality data. After aggregating and averaging the quality data of all frames, the overall video quality data is output.

[0009] To address the aforementioned technical problems, the present invention also provides a video quality detection device, comprising:

[0010] The image frame conversion module is used to preprocess the reference video and the distorted video and convert them into a target image frame sequence;

[0011] A multi-level extraction module is used to extract multi-level feature maps containing pixel-domain original features, shallow features, and deep features from the target image frame sequence frame by frame;

[0012] The index calculation and splicing module is used to calculate the image quality evaluation index at each level using the multi-level feature map, and splice the image quality evaluation indexes of all levels to form a multi-dimensional quality perception feature vector.

[0013] The quality data acquisition module is used to input the quality-aware feature vector into the quality regression network to obtain single-frame quality data, and output the overall video quality data after aggregating and averaging the quality data of all frames.

[0014] To address the aforementioned technical problems, the present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the video quality detection method described above.

[0015] To address the aforementioned technical problems, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the video quality detection method described above.

[0016] As can be seen from the above technical solution, the video quality detection method provided by the present invention includes: First, preprocessing the reference video and the distorted video to convert them into a target image frame sequence; extracting multi-level feature maps containing original pixel domain features, shallow features, and deep features frame by frame from the target image frame sequence; then, calculating image quality evaluation indicators at each level using the multi-level feature maps, and concatenating all the image quality evaluation indicators at all levels to form a multi-dimensional quality-perceived feature vector; finally, inputting the quality-perceived feature vector into a quality regression network to obtain single-frame quality data, and outputting the overall video quality data after aggregating and averaging the quality data of all frames.

[0017] The beneficial effects of this invention are that the video quality detection method provided by this invention, by combining pixel domain, shallow features, and deep features to obtain the difference between the reference video and the distorted video, uses image quality evaluation indicators to measure the difference in both the pixel domain and the feature domain, thus achieving efficient and accurate quality evaluation of high dynamic range videos. The multi-level feature fusion strategy can simultaneously utilize the original information of the pixel domain, the detailed texture information of shallow features, and the semantic information of deep features, forming a comprehensive feature coverage from the pixel level to the semantic level. This allows the model to take into account both structural and texture information at different scales, comprehensively capturing video quality features. At the same time, it fully utilizes the interpretability and effectiveness of image quality evaluation indicators, ensuring the credibility and persuasiveness of the detection results. This method can be widely applied to scenarios such as high dynamic range video coding optimization, streaming media service quality monitoring, and video processing algorithm evaluation, providing accurate and reliable quality basis for related video processing and service optimization work.

[0018] In addition, the present invention also provides a corresponding video quality detection device, electronic device and computer-readable storage medium for the video quality detection method, which have the same or corresponding technical features as the video quality detection method mentioned above, and have the same effect. Attached Figure Description

[0019] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart of a video quality detection method provided in an embodiment of the present invention;

[0021] Figure 2 This is a schematic diagram of the architecture corresponding to the video quality detection method provided in the embodiments of the present invention;

[0022] Figure 3 This is a schematic diagram of the feature map computation architecture provided in an embodiment of the present invention;

[0023] Figure 4 This is a schematic diagram of the video quality detection device provided in an embodiment of the present invention. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0025] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0026] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0027] The specific application environment architecture or specific hardware architecture on which the execution of the video quality detection method depends is described here.

[0028] The embodiments of the present invention provide a video quality detection method, and the method is described in detail in conjunction with the execution flow of the video quality detection method. Figure 1 A flowchart of the video quality detection method provided in the embodiments of the present invention is shown below. Figure 1 As shown, the method includes:

[0029] S101. Preprocess the reference video and the distorted video to convert them into a target image frame sequence.

[0030] It should be noted that the video quality detection method of the present invention can be a full-reference video quality detection method, which quantifies video quality by comparing the differences between the original reference video and the distorted video. The original reference video and the distorted video serve as inputs, and quality data is obtained as output. The results obtained should conform as closely as possible to the subjective perception of the human eye. The quality data here can be a score or other numerical value, and is not limited thereto.

[0031] This invention can be applied to scenarios such as high dynamic range video coding optimization, streaming media service quality monitoring, and video processing algorithm evaluation. High Dynamic Range (HDR) is used to improve the visual quality of images and videos. Its core lies in expanding the range of the image from brightest to darkest (i.e., dynamic range), combined with a wider color gamut, allowing the image to retain more detail in both bright and dark areas, and resulting in smoother, more natural color transitions. This presents a richer sense of detail and realism that is closer to what the human eye sees in the real world. HDR supports a wider color gamut (such as BT.2020) and higher color depth (10-bit or even 12-bit). Different HDR standards correspond to different bit depths, generally ranging from 10 bits (0-1023) or 12 bits (0-4095), while SDR typically has a range of 8 bits (0-255).

[0032] Step S101 involves preprocessing the video and converting it into a sequence of target image frames. This is a fundamental preliminary step in the entire video quality inspection process. It can unify the format and numerical standards of the input data, eliminate the differences between the original reference video and the distorted video in terms of resolution, pixel value range, and storage format, and provide a consistent data foundation for subsequent multi-level feature extraction and difference index calculation.

[0033] S102. Extract multi-level feature maps from the target image frame sequence frame by frame, which contain original pixel domain features, shallow features and deep features.

[0034] It should be noted that typical distortions in high dynamic range (HDR) videos often manifest in fine-grained details such as highlight compression and shadow noise, which are difficult to capture fully using only deep features. Deep features can capture high-level semantics and global structural information, but they are insufficient in terms of detailed texture and local distortion; shallow features retain detailed texture and local information, but are limited in global semantic understanding; pixel-domain features provide raw image information, including the most direct basic information such as brightness and color. Therefore, this invention utilizes a multi-level feature fusion strategy that simultaneously employs pixel-domain, shallow, and deep features, enabling the model to take into account both structural and textural information at different scales, thus comprehensively capturing video quality features.

[0035] Step S102 extracts multi-level feature maps frame by frame, containing original pixel-domain features, shallow features, and deep features. This overcomes the limitations of traditional methods that rely solely on single pixel-domain features. By extracting features in a layered manner, it preserves basic brightness and color information at the pixel level while also uncovering the detailed textures corresponding to shallow features and the global semantic information carried by deep features, forming a complete feature system from the bottom layer to the top layer. Simultaneously, the frame-by-frame processing mode ensures that the quality features of each frame in the video sequence are accurately captured.

[0036] S103. Calculate image quality evaluation indicators at each level using multi-level feature maps, and concatenate all the image quality evaluation indicators at all levels to form a multi-dimensional quality-perceived feature vector.

[0037] In implementation, image quality evaluation metrics are parameters used to quantify the degree of difference between the reference image and the distorted image across multiple dimensions. These dimensions can include mean squared error (MSE), brightness similarity, contrast similarity, and gradient magnitude similarity deviation (GMSD), among others. Image quality evaluation metrics have clear physical meaning and engineering interpretability. This invention uses feature difference measurement methods in both the pixel domain and the feature domain, fully utilizing the interpretability and effectiveness of image quality evaluation metrics to comprehensively capture video quality features.

[0038] Step S103 calculates quality evaluation metrics at three feature levels: pixel domain, shallow layer, and deep layer, transforming quality features at different scales (such as pixel-level numerical deviation, detail texture damage, and semantic structure distortion) into quantifiable parameters. Then, the metrics from all levels are concatenated to construct a multi-dimensional quality-aware feature vector. This hierarchical measurement and integrated approach preserves the quality information of features at each level while achieving feature complementarity, providing comprehensive and effective data input for the subsequent quality regression network to output accurate single-frame quality data.

[0039] S104. Input the quality-aware feature vector into the quality regression network to obtain single-frame quality data. After aggregating and averaging the quality data of all frames, output the overall video quality data.

[0040] Step S104 first inputs the feature vector integrating multi-dimensional quality information into the quality regression network. Through the model's fitting and mapping capabilities, the abstract feature parameters are transformed into intuitive single-frame quality data (such as single-frame quality scores). Then, through the aggregation and mean calculation of the full-frame data (such as scores), a higher-dimensional evaluation from single-frame quality to overall video quality is achieved. This approach takes into account the quality differences between video frames while outputting results that reflect the overall quality level of the video, providing a directly applicable quantitative basis for high dynamic range video quality evaluation.

[0041] The video quality detection method provided in this invention combines pixel-level, shallow-level, and deep-level features to obtain the differences between the reference video and the distorted video. Image quality evaluation metrics are used to measure these differences in both the pixel and feature domains, achieving efficient and accurate quality evaluation of high dynamic range (HDR) videos. The multi-level feature fusion strategy simultaneously utilizes the original information from the pixel domain, the detailed texture information from shallow features, and the semantic information from deep features, forming comprehensive feature coverage from the pixel level to the semantic level. This allows the model to consider both structural and texture information at different scales, comprehensively capturing video quality features. Furthermore, it fully leverages the interpretability and effectiveness of image quality evaluation metrics, ensuring the credibility and persuasiveness of the detection results. This method can be widely applied to scenarios such as HDR video coding optimization, streaming media service quality monitoring, and video processing algorithm evaluation, providing accurate and reliable quality data for related video processing and service optimization work.

[0042] Furthermore, in a specific implementation, in the video quality detection method provided in the embodiments of the present invention, step S101 preprocesses the reference video and the distorted video to convert them into a target image frame sequence. Specifically, this may include: reading the reference video and the distorted video, parsing the original pixel value range of the reference video and the distorted video; normalizing the original pixel value range to adjust the pixel values ​​to a set range; and converting the normalized reference video and the distorted video into image matrices in the form of frame sequences to obtain the target image frame sequence.

[0043] Figure 2 This is a schematic diagram of the architecture corresponding to the video quality detection method provided in this embodiment of the invention. In implementation, as... Figure 2As shown, the input consists of a reference video sequence and a distorted video sequence. Both video sequences undergo preprocessing, including video reading, parsing, and normalization. The video is read as a frame-by-frame image matrix, and then numerical normalization is performed. For example, if the high dynamic range video is read with a 10-bit numerical range (0-1023), it will be uniformly divided by 1023 to ensure the value range is between 0 and 1. Subsequently, frame-by-frame image extraction is performed, converting the video sequence into a target image frame sequence. This eliminates the inherent differences in brightness and color values ​​between different video sources, ensuring data consistency and comparability.

[0044] Furthermore, in a specific implementation, in the video quality detection method provided in the embodiments of the present invention, step S102 extracts a multi-level feature map containing pixel-domain original features, shallow features, and deep features from the target image frame sequence frame by frame. Specifically, this may include: acquiring a single frame image from the target image frame sequence, extracting the original image data of the single frame image as the pixel-domain original features; inputting the single frame image into a backbone network, extracting the corresponding feature map from the first output layer of the backbone network as shallow features to capture image detail texture and local structural information; extracting the corresponding feature map from the second output layer of the backbone network as deep features to capture image global semantics and overall structural information; and integrating the pixel-domain original features, shallow features, and deep features of the single frame image to form a complete multi-level feature map for a single frame.

[0045] In implementation, for each extracted image frame, multi-level feature maps are extracted using a backbone network. The backbone network can be understood as the main feature extraction network. It has a clear hierarchical concept and exhibits a specific pattern as the network depth changes: in the length × width × channel dimension of the output feature map, the length and width gradually decrease with increasing layer depth, while the number of channels continuously increases. For example, an input image of 224×224×3 might output feature maps of sizes such as 112×112×64, 64×64×128, 64×64×256, and 32×32×512 after processing through various layers. Feature extraction is performed based on the output of each layer module. The multi-level feature maps include deep semantic features (extracted from multiple layers of the backbone network) and pixel-domain raw features (original image, with up to 3 channels). For both the reference image and the distorted image, these multi-level feature maps are extracted separately, and then the difference between the reference feature map and the distorted feature map is calculated at each layer. It is important to note that different backbone network models may have different numbers of output layers, and this invention can flexibly adapt to feature extraction with different numbers of layers. In this way, the model can simultaneously utilize the original information of the pixel domain, the detailed texture information of shallow features, and the semantic information of deep features, forming a comprehensive feature coverage from the pixel level to the semantic level, thus fully capturing video quality features. It should be noted that there is no absolutely clear-cut standard for defining shallow and deep features; generally, the first 2-3 layers are classified as shallow, and subsequent layers as deep.

[0046] This paper uses a four-layer feature extraction method as an example, flexibly adapting to the number of output layers of different backbone network models. The extracted multi-level features specifically include three categories: first, pixel-domain features, directly using the original 3-channel RGB image to retain the most basic original information such as brightness and color; second, shallow features, extracted from early layers of the backbone network (such as feature maps with channels C1 and C2 in a four-layer structure), focusing on preserving detailed textures and local information, accurately capturing image detail features; and third, deep features, extracted from deeper layers of the backbone network (such as feature maps with channels C3 and C4 in a four-layer structure), capable of capturing high-level semantic and global structural information, aiding in understanding the overall image framework. This extraction strategy allows the model to utilize rich information from the pixel level to the semantic level simultaneously, providing comprehensive feature representations for subsequent difference calculations. Furthermore, the number of feature extraction layers can be adjusted accordingly for different backbone network layers, ensuring the method's versatility and flexibility.

[0047] Furthermore, in a specific implementation, in the video quality detection method provided in the embodiments of the present invention, step S103 uses multi-level feature maps to calculate image quality evaluation indicators at each level, and concatenates the image quality evaluation indicators of all levels to form a multi-dimensional quality-perceived feature vector. Specifically, this may include: obtaining multi-level feature map sets corresponding to the reference video and the distorted video respectively; calculating multiple image quality evaluation difference indicators for the reference feature map and the distorted feature map at each level respectively; the image quality evaluation difference indicators include mean square error, brightness similarity, contrast similarity, and gradient magnitude similarity deviation; and concatenating the calculation results of each indicator corresponding to all level feature maps in a set order to form a quality-perceived feature vector containing multi-dimensional information from the pixel domain to the semantic level.

[0048] Figure 3 This is a schematic diagram of the feature map computation architecture provided in an embodiment of the present invention. In implementation, as... Figure 3 As shown, this invention calculates image quality evaluation metrics in both the pixel domain and the feature domain (including shallow and deep feature domains) for multi-level feature maps of the reference image and the distorted image. For each level of feature map (including the original pixel domain image and shallow and deep feature maps extracted from the backbone network), image quality evaluation metrics such as mean square error, brightness similarity and contrast similarity in structural similarity metrics, and gradient magnitude similarity deviation are calculated. Then, the image quality evaluation metrics of all levels are concatenated to form a multi-dimensional quality-perceived feature vector. This approach uses feature difference measurement methods in both the pixel domain and the feature domain, fully utilizing the interpretability and effectiveness of image quality evaluation metrics, and achieving comprehensive quality evaluation from the pixel level to the semantic level.

[0049] It should be noted that for each level, such as the matrix of length × width × number of channels (including the feature matrix of the reference image and the feature matrix of the distorted image), after calculation by the image quality evaluation index, it will become a feature vector of 1×1×number of channels. Since the length and width have become 1×1, the stitching of different levels is actually the stitching in the channel dimension, and the dimension itself has been unified.

[0050] For the multi-level feature maps of the reference image and the distorted image, this invention employs an image quality assessment algorithm at each feature level to calculate the differences between them. Specifically, this includes: measuring the numerical differences between the feature maps using mean squared error; calculating brightness similarity based on the feature map mean to reflect the consistency of brightness distribution; using the variance and covariance of the feature maps to solve for contrast similarity to assess the preservation of detail contrast; and quantifying the degree of difference in gradient information using gradient magnitude similarity deviation. Finally, the calculation results of the above quality assessment indicators at all levels are integrated and stitched together to form a multi-dimensional quality-perceived feature vector, providing a comprehensive quantitative basis for difference quantification for subsequent input into the quality regression module for final quality data calculation.

[0051] Furthermore, in specific implementation, in the above steps, multiple image quality evaluation difference indices are calculated for the reference feature maps and distortion feature maps at each level. Specifically, this may include: for the reference feature maps and distortion feature maps at each level, calculating the average of the squared differences of pixel values ​​between the reference feature maps and distortion feature maps through mean operation to obtain the mean square error index result; calculating the mean of the reference feature maps and distortion feature maps respectively to calculate the brightness similarity index result; calculating the variance of the reference feature maps, the variance of the distortion feature maps, and the covariance of the reference feature maps and distortion feature maps respectively to calculate the contrast similarity index result; calculating the horizontal and vertical gradients of the reference feature maps and distortion feature maps respectively using specified gradient operators; calculating the gradient magnitude of the reference feature maps and distortion feature maps respectively based on the gradient values; calculating the gradient magnitude similarity of the reference feature maps and distortion feature maps pixel by pixel, calculating the standard deviation of the gradient magnitude similarity of all pixels, and obtaining the gradient information difference index result.

[0052] In practice, the mean square error index between the reference feature map and the distorted feature map is obtained using the following formula:

[0053] ;

[0054] in, and These are the reference feature and the distortion feature, respectively.

[0055] The brightness similarity S1 index result is calculated using the following formula:

[0056] ;

[0057] in, and c1 represents the mean of the reference feature map and the distorted feature map, respectively, and c1 is a small constant term.

[0058] The contrast similarity S2 index result is calculated using the following formula:

[0059] ;

[0060] in, and These are the variances of the reference feature map and the distorted feature map, respectively. Let c be the covariance of the two, and c2 be the smaller constant term.

[0061] The gradient magnitude similarity deviation (GMSD) of feature maps measures the difference in gradient information between the reference image and the distorted image. The GMSD calculation process includes the following steps:

[0062] Step 1: Calculate the gradient: Calculate the horizontal and vertical gradients of the reference feature map and the distorted feature map using gradient operators (such as the Prewitt operator or the Sobel operator). For the horizontal direction, perform convolution using the horizontal gradient operator (such as [-1,0, 1; -1, 0, 1; -1, 0, 1]); for the vertical direction, perform convolution using the vertical gradient operator (such as [-1, -1, -1; 0, 0, 0; 1, 1, 1]) to obtain the horizontal gradient. and vertical gradient .

[0063] Step 2: Calculate the gradient magnitude: For both the reference feature map and the distorted feature map, calculate the gradient magnitude at each pixel location using the following formula: According to the horizontal gradient and vertical gradient The gradient magnitude of the reference feature map can be obtained. Gradient magnitude of the distorted feature map .

[0064] Step 3: Calculate Gradient Magnitude Similarity (GMS): For each pixel location, calculate the gradient magnitude similarity between the reference feature map and the distorted feature map, using the following formula: , where c is a small constant term used for numerical stability. The value of GMS ranges from [0,1], and a larger value indicates that the gradient magnitudes are more similar.

[0065] Step 4: Calculate the GMSD value: Calculate the standard deviation of the GMS values ​​for all pixel locations, which will be the final GMSD value. The formula is GMSD = std(GMS), where std represents the standard deviation calculation. The smaller the GMSD value, the more similar the reference image and the distorted image are in terms of gradient information, and the better the quality.

[0066] Furthermore, in specific implementation, in the above steps, the calculation results of various indicators corresponding to all hierarchical feature maps are concatenated in a set order to form a quality-perceived feature vector containing multi-dimensional information from the pixel domain to the semantic level. Specifically, this may include: arranging the results in the hierarchical order of pixel domain feature layer, shallow feature layer, and deep feature layer, with each layer arranged in the order of mean square error, brightness similarity, contrast similarity, and gradient magnitude similarity deviation; extracting the calculation results of various indicators from the pixel domain feature layer and organizing them into a one-dimensional vector according to the indicator order; extracting the calculation results of various indicators from the shallow feature layer and organizing them into a one-dimensional vector according to the indicator order, and concatenating them after the pixel domain feature layer vector; extracting the calculation results of various indicators from the deep feature layer and organizing them into a one-dimensional vector according to the indicator order, and concatenating them after the shallow feature layer vector; and performing dimensionality verification on the concatenated vector to determine that no indicators are missing or repeated, thus forming a quality-perceived feature vector containing multi-dimensional information from the pixel domain to the semantic level.

[0067] In implementation, the multi-level feature map difference fusion module takes a reference image and a distorted image (3×H×W) as input after normalization preprocessing, and outputs a multi-dimensional quality feature vector of dimension sum(chns)×K (where chns is the number of channels in each level of the feature map, and K is the number of image quality evaluation metrics calculated for each level). Its working principle is as follows: First, pixel-domain, shallow-layer, and deep-layer features are extracted from the reference image and the distorted image, respectively. Then, traditional image quality evaluation metrics such as mean square error, brightness similarity, contrast similarity, and gradient magnitude similarity deviation are calculated at each level. This achieves feature difference measurement in both the pixel domain and the feature domain, fully utilizing the interpretability and effectiveness of image quality evaluation metrics to comprehensively capture features from the pixel level to the semantic level. Video quality features are ultimately derived by dividing the data into pixel-domain feature layers, shallow feature layers, and deep feature layers. Within each layer, the data is arranged in the order of mean square error, brightness similarity, contrast similarity, and gradient magnitude similarity deviation. The calculation results of each indicator in the pixel-domain feature layer are extracted and organized into a one-dimensional vector. Then, the calculation results of each indicator in the shallow feature layer are extracted, organized into a one-dimensional vector, and concatenated to the pixel-domain feature layer vector. Next, the calculation results of each indicator in the deep feature layer are extracted, organized into a one-dimensional vector, and concatenated to the shallow feature layer vector. The concatenated vectors are then subjected to dimensionality verification to ensure that no indicators are missing or repeated. This results in a quality-aware feature vector containing multi-dimensional information from the pixel domain to the semantic level, providing rich feature representations for subsequent quality regression.

[0068] Furthermore, in a specific implementation, in the video quality detection method provided in the embodiments of the present invention, step S104 inputs the quality-aware feature vector into the quality regression network to obtain single-frame quality data, and outputs the overall video quality data after aggregating and averaging the quality data of all frames. Specifically, this may include: taking the quality-aware feature vector corresponding to a single frame and inputting it into the trained quality regression network; outputting the quality data of a single frame through the linear transformation and feature integration operation of the quality regression network; collecting the quality data of all frames after processing the quality-aware feature vectors of all frames one by one to form a frame-level quality data set; and calculating the arithmetic mean of the frame-level quality data set as the overall video quality data.

[0069] In implementation, this invention first inputs the single-frame quality-aware feature vector, carrying pixel-level to semantic-level quality information, into a trained quality regression network. Through linear transformation and feature integration operations, the abstract feature parameters are mapped into intuitive single-frame quality data. Then, frame-by-frame processing is used to collect full-frame quality data, forming a complete frame-level quality dataset. Finally, the arithmetic mean is used as the overall video quality data, taking into account both individual differences in single-frame quality and objectively reflecting the overall quality level of the video. The entire process is logically coherent, from precise single-frame measurement to comprehensive global evaluation, providing standardized and quantifiable evaluation results for high dynamic range video quality.

[0070] After forming a frame-level quality data set, this invention can perform outlier checks on the frame-level quality data set, removing data that exceeds a set range. Then, the arithmetic mean of the remaining frame-level quality data set is calculated and used as the overall video quality data.

[0071] It should be added that this invention can train the model based on a high dynamic range video quality assessment dataset. The model includes a backbone network and a quality regression network. During the training phase, a subset of frames is randomly sampled from each video, and during the validation phase, a fixed subset of frames is sampled according to a set rule. A mean squared error loss function is used, and hyperparameters such as the learning rate are adjusted through stochastic gradient descent or an adaptive optimizer. All parameters of the backbone network and the quality regression network are jointly trained, and the parameters are iteratively updated through backpropagation, enabling the model to learn effective quality assessment rules.

[0072] In implementation, the model training of this invention adopts an end-to-end training approach, optimizing model parameters by minimizing the difference between predicted quality data and true quality data. The training data uses a high dynamic range video quality assessment dataset, including reference video sequences, distorted video sequences, and their corresponding subjective quality data. To improve data utilization efficiency, N frames are randomly selected from each video during the training phase (this selection can be dynamically chosen based on different datasets and training devices; 16 frames are recommended here). During the validation phase, M frames are sampled at a fixed rate (this can be adjusted based on actual results. For high-precision validation, each frame needs to be sampled; for speed compatibility, 2 frames per second can be sampled) for evaluation.

[0073] This invention uses mean squared error as the loss function to minimize the difference between predicted quality data and actual quality data. The loss function is defined as follows: ,in, For the quality data predicted by the model, This is based on genuine subjective quality data.

[0074] End-to-end training is performed on all parameters of the model, including the backbone network, feature extraction, and quality regression modules. The model parameters are optimized using backpropagation to enable the model to extract effective quality-aware features from multi-level feature maps and accurately predict video quality data. During training, either Stochastic Gradient Descent (SGD) or the Adam optimizer is used to optimize training performance by adjusting hyperparameters such as the learning rate.

[0075] The training process includes the following steps: First, input the reference video sequence and the distorted video sequence to extract multi-level feature maps; then, calculate traditional image quality assessment metrics (mean squared error, brightness similarity, contrast similarity, gradient magnitude similarity deviation, etc.) at each level; next, concatenate all the image quality assessment metrics from all levels to form a multi-dimensional quality-perceived feature vector; finally, calculate the predicted quality data through a quality regression network, calculate the loss with the real quality data, and update the model parameters through backpropagation. Through iterative training, the model can learn effective quality assessment rules.

[0076] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0077] Embodiments of the present invention also provide a video quality detection device. Figure 4This is a schematic diagram of the video quality detection device provided in an embodiment of the present invention. This embodiment is based on functional modules, such as… Figure 4 As shown, the device includes:

[0078] The image frame conversion module 10 is used to preprocess the reference video and the distorted video and convert them into a target image frame sequence;

[0079] The multi-level extraction module 11 is used to extract multi-level feature maps containing pixel domain original features, shallow features and deep features from the target image frame sequence frame by frame;

[0080] The index calculation and splicing module 12 is used to calculate the image quality evaluation index at each level using multi-level feature maps, and splice the image quality evaluation indexes of all levels to form a multi-dimensional quality perception feature vector.

[0081] The quality data acquisition module 13 is used to input the quality-aware feature vector into the quality regression network to obtain single-frame quality data, and to output the overall video quality data after aggregating and averaging the quality data of all frames.

[0082] In the video quality detection device provided in this embodiment of the invention, the differences between the reference video and the distorted video can be obtained through the interaction of the four modules, based on the extracted pixel domain, shallow features, and deep features. Image quality evaluation metrics are used to measure these differences in both the pixel domain and the feature domain, achieving efficient and accurate quality evaluation of high dynamic range videos. The multi-level feature fusion strategy can simultaneously utilize the original information of the pixel domain, the detailed texture information of shallow features, and the semantic information of deep features, forming a comprehensive feature coverage from the pixel level to the semantic level. This allows the model to consider both structural and texture information at different scales, comprehensively capturing video quality features. Simultaneously, it fully utilizes the interpretability and effectiveness of image quality evaluation metrics, ensuring the credibility and persuasiveness of the detection results. This method can be widely applied to scenarios such as high dynamic range video coding optimization, streaming media service quality monitoring, and video processing algorithm evaluation, providing accurate and reliable quality basis for related video processing and service optimization work.

[0083] Since the embodiments of the video quality detection device and the video quality detection method correspond to each other, the descriptions of the features in the embodiments corresponding to the video quality detection device can be found in the relevant descriptions of the embodiments corresponding to the video quality detection method, and will not be repeated here. Furthermore, it has the same beneficial effects as the video quality detection method mentioned above.

[0084] Furthermore, in a specific implementation, in the video quality detection device provided in the embodiments of the present invention, the image frame conversion module 10 can be specifically used to read the reference video and the distorted video, parse the original pixel value range of the parameter video and the distorted video; normalize the original pixel value range so that the pixel values ​​are adjusted to a set range; and convert the reference video and the distorted video after normalization into image matrices in the form of frame sequences to obtain the target image frame sequence.

[0085] Furthermore, in a specific implementation, in the video quality detection device provided in the embodiments of the present invention, the multi-level extraction module 11 can be specifically used to acquire single-frame images in the target image frame sequence, extract the original image data of the single-frame image and use it as the original pixel domain features; input the single-frame image into the backbone network, extract the corresponding feature map from the first output layer of the backbone network as shallow features to capture image detail texture and local structural information; extract the corresponding feature map from the second output layer of the backbone network as deep features to capture image global semantics and overall structural information; integrate the original pixel domain features, shallow features and deep features of the single-frame image to form a complete multi-level feature map of the single frame.

[0086] Furthermore, in a specific implementation, in the video quality detection device provided in the embodiments of the present invention, the index calculation and splicing module 12 can be used to obtain multi-level feature maps corresponding to the reference video and the distorted video respectively; calculate multiple image quality evaluation difference indices for each level of reference feature map and distorted feature map; the image quality evaluation difference indices include mean square error index, brightness similarity index, contrast similarity index and gradient magnitude similarity deviation index; and splice the calculation results of each index corresponding to all levels of feature maps in a set order to form a quality-perceived feature vector containing multi-dimensional information from the pixel domain to the semantic level.

[0087] Furthermore, in a specific implementation, in the video quality detection device provided in the embodiments of the present invention, the quality data acquisition module 13 can be specifically used to extract the quality-aware feature vector corresponding to a single frame and input it into the trained quality regression network; through the linear transformation and feature integration operation of the quality regression network, the quality data of a single frame is output; after processing the quality-aware feature vectors of all frames one by one, the quality data of all frames are collected to form a frame-level quality data set; the arithmetic mean of the frame-level quality data set is calculated and used as the overall video quality data.

[0088] Embodiments of the present invention also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described video quality detection method embodiments.

[0089] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program configured to execute the steps in any of the above-described video quality detection method embodiments when running.

[0090] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0091] Embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described video quality detection method embodiments.

[0092] Embodiments of the present invention also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described video quality detection method embodiments.

[0093] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0094] The present invention has provided a detailed description of a video quality detection method, apparatus, device, and medium. Specific examples have been used to illustrate the principles and implementation methods of the invention. The descriptions of these embodiments are only intended to aid in understanding the method and core ideas of the present invention. It should be noted that those skilled in the art can make various improvements and modifications to the present invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A method of video quality detection, the method comprising: The method comprises the following steps: Preprocessing the reference video and the distorted video to convert them into a target image frame sequence; Extracting a multi-level feature map containing pixel domain original features, shallow features and deep features from the target image frame sequence frame by frame; Calculating an image quality evaluation index on each level of the multi-level feature map, and splicing the image quality evaluation indexes of all levels to form a multi-dimensional quality perception feature vector; Inputting the quality perception feature vector into a quality regression network to obtain single-frame quality data, and outputting video overall quality data after aggregating and averaging the quality data of all frames.

2. The video quality detection method of claim 1, wherein, The preprocessing of the reference video and the distorted video to convert them into a target image frame sequence comprises the following steps: Reading the reference video and the distorted video, and parsing the original pixel value range of the parameter video and the distorted video; Normalizing the original pixel value range to adjust the pixel value to a set interval; Converting the reference video and the distorted video after normalization into image matrices in the form of frame sequences to obtain the target image frame sequence.

3. The video quality detection method of claim 1, wherein, The multi-level feature map containing pixel domain original features, shallow features and deep features extracted from the target image frame sequence frame by frame comprises the following steps: Obtaining a single-frame image in the target image frame sequence, and extracting the original image data of the single-frame image as the pixel domain original features; Inputting the single-frame image into a backbone network, extracting the corresponding feature map from the first output level of the backbone network as the shallow features to capture the image detail texture and local structure information; Extracting the corresponding feature map from the second output level of the backbone network as the deep features to capture the image global semantics and overall structure information; Integrating the pixel domain original features, shallow features and deep features of the single-frame image to form a single-frame complete multi-level feature map.

4. The video quality detection method of claim 1, wherein, The multi-level feature map is used to calculate an image quality evaluation index on each level, and the image quality evaluation indexes of all levels are spliced to form a multi-dimensional quality perception feature vector, which comprises the following steps: Respectively obtaining a multi-level feature map set corresponding to the reference video and the distorted video; Respectively calculating a plurality of image quality evaluation difference indexes for the reference feature map and the distorted feature map of each level; the image quality evaluation difference indexes include a mean square error index, a brightness similarity index, a contrast similarity index and a gradient amplitude similarity deviation index; Splicing the calculation results of each index corresponding to all level feature maps in a set order to form a quality perception feature vector containing multi-dimensional information from the pixel domain to the semantic level.

5. The video quality detection method of claim 4, wherein, Respectively calculating a plurality of image quality evaluation difference indexes for the reference feature map and the distorted feature map of each level, which comprises the following steps: For the reference feature map and the distorted feature map of each level, the average value of the pixel value square difference between the reference feature map and the distorted feature map is obtained by mean operation to obtain the mean square error index result; The mean value of the reference feature map and the distorted feature map is respectively solved to calculate the brightness similarity index result; The variance of the reference feature map, the variance of the distorted feature map and the covariance of the reference feature map and the distorted feature map are respectively solved to calculate the contrast similarity index result; The horizontal and vertical direction gradients of the reference feature map and the distorted feature map are calculated using the specified gradient operator respectively; the gradient amplitudes of the reference feature map and the distorted feature map are calculated based on the gradient values respectively; the gradient amplitude similarity of the reference feature map and the distorted feature map is calculated pixel by pixel, and the standard deviation of the gradient amplitude similarity of all pixels is calculated to obtain the gradient information difference index result.

6. The video quality detection method of claim 5, wherein, The index calculation results of all hierarchical feature maps are spliced in a set order to form a quality perception feature vector containing multi-dimensional information from the pixel domain to the semantic level, including: arranged in the order of the mean square error index, the brightness similarity index, the contrast similarity index and the gradient amplitude similarity deviation index within each level; the index calculation results of the pixel domain feature layer are extracted and arranged as a one-dimensional vector in the order of the indices; the index calculation results of the shallow feature layer are extracted and arranged as a one-dimensional vector in the order of the indices, and spliced after the pixel domain feature layer vector; the index calculation results of the deep feature layer are extracted and arranged as a one-dimensional vector in the order of the indices, and spliced after the shallow feature layer vector; the spliced vector is subjected to dimension checking, and after determining that there is no index omission or repetition, a quality perception feature vector containing multi-dimensional information from the pixel domain to the semantic level is formed.

7. The method of claim 1, wherein, The quality perception feature vector is input into a quality regression network to obtain single-frame quality data, and the quality data of all frames are aggregated and averaged to output video overall quality data, including: taking the quality perception feature vector corresponding to a single frame and inputting it into the trained quality regression network; outputting the quality data of a single frame through linear transformation and feature integration operation of the quality regression network; After processing the quality perception feature vectors of all frames one by one, the quality data of all frames is collected to form a frame-level quality data set; the arithmetic mean of the frame-level quality data set is calculated and used as the video overall quality data.

8. A video quality detection apparatus, characterized by comprising: including: an image frame conversion module for pre-processing the reference video and the distorted video and converting them into a target image frame sequence; a multi-level extraction module for extracting multi-level feature maps containing pixel domain original features, shallow features and deep features from the target image frame sequence frame by frame; an index calculation and splicing module for calculating image quality evaluation indexes on each level using the multi-level feature maps, and splicing the image quality evaluation indexes of all levels to form a multi-dimensional quality perception feature vector; a quality data acquisition module for inputting the quality perception feature vector into a quality regression network to obtain single-frame quality data, and outputting video overall quality data after aggregating and averaging the quality data of all frames.

9. An electronic device, comprising: including: a memory for storing a computer program; a processor for executing the computer program to implement the steps of the video quality detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer program is stored on the computer readable storage medium and executed by the processor to implement the steps of the video quality detection method according to any one of claims 1 to 7.