Video quality detection method and device
By adopting a video quality detection model based on attention mechanism in video quality detection, extracting the temporal and spatial characteristics of the video and determining the quality score, the problem that existing methods cannot effectively reduce detection costs and reflect subjective feelings is solved, and high-accuracy and low-cost video quality detection is achieved.
Patent Information
- Application Number
- CN202510125367.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-30
AI Technical Summary
Existing video quality detection methods cannot effectively reduce detection costs, and cannot fully reflect the subjective visual experience of the human eye.
Using a video quality detection model based on attention mechanism, time domain features and airspace features are extracted from the target video to be detected, and featured by multiple three-dimensional tiles are extracted through the video quality detection model to obtain the spatiotemporal feature map of the video frame sequence, and then the video quality score is determined.
It improves the accuracy of video quality ratings, makes the rating closer to the subjective feelings of the human eye, and greatly reduces the detection cost.
Smart Images

Figure CN120070353A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of computer technology, and in particular, to a method and device for video quality detection. Background Art
[0002] With the rapid development of hardware devices, video, as an important information carrier, affects people's work and life. In addition, with the development of the self-media industry, UGC (user-generated content) also accounts for an increasingly large proportion on various video platforms. However, limited by shooting levels and equipment costs, the video quality of some videos is low, which will affect the viewing experience of users. Therefore, it is necessary to screen videos according to video quality. At the same time, multi-modal foundation large models and AIGC generation models (models that automatically generate various contents through artificial intelligence technology) also need to be trained based on high-quality videos. It can be seen that in many scenarios, there is a need to detect video quality.
[0003] Currently, the methods for video quality detection are mainly divided into reference-based objective quality detection and reference-free subjective quality detection. Among them, reference-based objective quality detection scores videos by comparing and analyzing the differences between the pre- and post-encoded videos. However, the scores obtained by this method cannot fully reflect the subjective visual perception of the human eye. The scoring dimensions of reference-free subjective quality detection are relatively extensive, so the resource consumption is very large and the cost is very high.
[0004] Therefore, it is necessary to provide a low-cost and effective video quality detection solution. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and device for video quality detection, which can ensure accurate detection of video quality while significantly reducing the detection cost.
[0006] The first aspect of this specification provides a method for video quality detection, including:
[0007] Extracting a video frame sequence from a target video to be detected;
[0008] Dividing the video frame sequence into blocks to obtain a plurality of three-dimensional blocks;
[0009] Inputting each three-dimensional block into a video quality detection model. In the video quality detection model, based on the attention mechanism, feature extraction is performed on the plurality of three-dimensional blocks to obtain a spatio-temporal feature map of the video frame sequence; based on the spatio-temporal feature map of the video frame sequence, a quality score of the target video is determined, where the video quality detection model is trained based on training samples, and the training samples are obtained by performing a degradation operation on video samples that meet preset conditions.
[0010] The second aspect of this specification provides a device for video quality detection, including:
[0011] An extraction unit configured to extract a sequence of video frames from a target video to be detected;
[0012] A partitioning unit configured to partition the sequence of video frames to obtain a plurality of three-dimensional tiles;
[0013] An extraction unit configured to input each three-dimensional tile into a video quality detection model, in which, based on an attention mechanism, feature extraction is performed on the plurality of three-dimensional tiles to obtain a spatio-temporal feature map of the sequence of video frames; based on the spatio-temporal feature map of the sequence of video frames, a quality score of the target video is determined, where the video quality detection model is trained based on training samples, and the training samples are obtained by performing a degradation operation on video samples meeting preset conditions.
[0014] The third aspect of this specification provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed in a computer, the computer is made to execute the method described in the first aspect.
[0015] The fourth aspect of this specification provides a computing device, including a memory and a processor, where an executable code is stored in the memory, and when the processor executes the executable code, the method described in the first aspect is implemented.
[0016] The fifth aspect of this specification provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method described in the first aspect are implemented.
[0017] The method for video quality detection provided by one or more embodiments of this specification uses a video quality detection model, based on an attention mechanism, to extract temporal features and spatial features from a target video to be detected, and determines a quality score of the target video based on these two aspects of features. Thus, the accuracy of the quality score can be greatly improved. In addition, since this solution uses a video quality detection model trained based on video data meeting preset conditions to determine the quality score of the target video, the obtained quality score is closer to the subjective human visual perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the technical solutions of the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1Shows a schematic structural diagram of a video quality detection model in an example of this specification;
[0020] Figure 2 Shows a schematic structural diagram of a 3D-swin-Transformer block in an example of this specification;
[0021] Figure 3 Shows a method flowchart for training a video quality detection model according to an embodiment of this specification;
[0022] Figure 4 Shows a schematic method diagram for processing a sample frame sequence using a video quality detection model in an example of this specification;
[0023] Figure 5 Shows a method flowchart for video quality detection according to an embodiment of this specification;
[0024] Figure 6 Shows a schematic diagram of a device for video quality detection according to an embodiment of this specification. Detailed implementation
[0025] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0026] As mentioned above, currently, the methods for video quality detection are mainly divided into reference-based objective quality detection and reference-free subjective quality detection. Among them, reference-based objective quality detection requires an original video as a reference benchmark, and calculates objective metrics such as SSIM (Structural Similarity Index), PSNR (Peak Signal-to-Noise Ratio), or VMAF (Video Multimethod Assessment Fusion) between the detected video and the original video to quickly and efficiently detect the video quality. However, in most cases, the original video is usually unavailable, or the quality of the original video itself is very low, then the corresponding objective metrics lose their meaning. Moreover, as mentioned above, the scores obtained by this method cannot fully reflect the subjective visual perception of the human eye. Reference-free subjective quality detection does not require the provision of the original video, and scores the video by detecting feature information such as video picture quality / audio features, aesthetic features, content features, scene features, and encoding features. In other words, this method needs to score the video from multiple dimensions, while in practice, usually only an overall score of the video is required, and there is no need to care about the scores of multiple dimensions of the video.
[0027] Based on this, this solution innovatively proposes a method for video quality detection, that is, using a video quality detection model, based on the attention mechanism, extracting temporal features and spatial features from the target video to be detected, and determining the quality score of the target video based on these two aspects of features. Thus, the accuracy of the quality score can be greatly improved. In addition, since this solution uses a video quality detection model trained based on video data that meets preset conditions to determine the quality score of the target video, the obtained quality score is closer to the subjective perception of the human eye.
[0028] As follows, first, the structure of the video quality detection model will be described.
[0029] Figure 1 The structural schematic diagram of the video quality detection model in an example of this specification is shown. Figure 1 Among them, the video quality detection model includes multiple feature processing layers. Among them, the first feature processing layer can include a linear embedding layer and several attention blocks; other feature processing layers can include a block fusion layer and several attention blocks. Among them, the several attention blocks included in each feature processing layer process their inputs based on the attention mechanism and a three-dimensional sliding window. The number of several attention blocks between different feature processing layers can be the same or different.
[0030] In addition, several attention blocks in each layer processing layer appear in pairs. That is to say, the number of several attention blocks in each layer feature processing layer is expressed as 2k, where k represents a positive integer, and the value of k can be the same or different between different feature processing layers.
[0031] In one embodiment, the above-mentioned attention block can be implemented as a 3D-swin-Transformer block, and its structure can be as Figure 2 shown. Figure 2 In, the two paired 3D-swin-Transformer blocks both include attention layers. Among them, the attention layer in the previous block adopts a W-MSA (Window Multi-Head Self-Attention) structure, that is, it processes its input based on the multi-head attention algorithm and a conventional three-dimensional sliding window. The attention layer in the latter block adopts an SW-MSA (Shifted Window Multi-Head Self-Attention) structure, that is, it processes its input based on an offset three-dimensional sliding window of the multi-head attention algorithm.
[0032] In addition, in each 3D-swin-Transformer block, a first normalization layer can be further provided before the attention layer for normalizing the input of the 3D-swin-Transformer block; a first residual connection sub-layer is further provided after the attention layer for adding the output of the attention layer and the input of the 3D-swin-Transformer block; a second normalization layer is provided after the first residual connection sub-layer for normalizing the output of the first residual connection sub-layer; an MLP (Multi-Layer Perceptron) is provided after the second normalization layer for further processing its input (the output of the second normalization layer); then a second residual connection sub-layer is further provided after the MLP for normalizing the output of the MLP and the output of the first residual connection sub-layer; then, the second residual connection sub-layer can be connected to the latter 3D-swin-Transformer block, or connected to the latter feature processing layer, or its output is used as the model output.
[0033] It should be noted that Figure 2 the residual connection in makes the network structure of the 3D-swin-Transformer block richer and can prevent gradient disappearance / gradient explosion. In addition, the attention layer can learn richer features through the multi-head attention algorithm. Finally, the conventional three-dimensional sliding window and the offset three-dimensional sliding window can better focus on the internal connection between adjacent features in the time domain and the spatial domain.
[0034] In the case where the above-mentioned attention block is implemented as a 3D-swin-Transformer block, the above-mentioned video quality detection model can also be referred to as a 3D-swin-Transformer model.
[0035] In practice, the above-mentioned video quality detection model can also be implemented as other models based on the Transformer structure, as long as the model can implement feature extraction from both the time domain and the spatial domain based on the attention mechanism and determine the quality score of the video based on the features in these two aspects. This specification does not limit this.
[0036] In addition, Figure 1 This is just an exemplary illustration. In practice, the video quality detection model may also include other network layers such as a pooling layer and an output layer. Among them, the pooling layer is used to perform dimensionality reduction processing on its input, and the output layer is used to output the quality score of the video based on the input.
[0037] The following describes Figure 1 the process of training the shown video quality detection model.
[0038] Figure 3 FIG. shows a flowchart of a method for training a video quality detection model according to an embodiment of the present specification. This method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. It should be noted that this method includes multiple rounds of iteration. Figure 3 FIG. shows the method steps included in the t-th (t is a positive integer) round of iteration. It can be understood that by repeatedly executing the steps shown therein, multi-round iterative updates of the video quality detection model can be achieved. As Figure 3 shown, this method may include the following steps:
[0039] Step S302, obtain a training sample set.
[0040] Among them, the training sample set may include multiple sample frame sequences respectively extracted from multiple video samples, and a single sample frame sequence has a corresponding annotation score.
[0041] In order to enhance the adaptability of the video quality detection model, the above-mentioned multiple video samples may have at least one of the following characteristics: the video sources of the multiple video samples are different. For example, video samples can be collected from websites, social media platforms, etc., or can be shot by oneself; the multiple video samples have different resolutions. For example, their resolutions can be 2k or 4k, etc.; the multiple video samples are selected from different scenes. For example, they can be portraits, animals, life, food, and cartoons, etc.; a single video sample has a uniform video quality level (that is, the video sample maintains a relatively consistent quality level in each part or different time periods).
[0042] In a more specific embodiment, the resolution, saturation, contrast, etc. of the above-mentioned multiple video samples are all greater than the corresponding preset thresholds. For example, the resolution is greater than 2k, that is to say, the above-mentioned multiple video samples are all high-quality videos (such as movie videos) that conform to the subjective human eye perception.
[0043] For the above-mentioned multiple video samples, a degradation operation can be performed on them, and the degradation operation can include one or more of the following: adding noise, blurring, compression, reducing resolution, reducing saturation, and reducing contrast, etc.
[0044] Among them, the above-mentioned adding noise can specifically include adding Gaussian noise, salt-and-pepper noise, etc.; the above-mentioned blurring can specifically include mean blurring, Gaussian blurring, bilateral blurring, etc.; the above-mentioned compression can specifically include video compression, bit-depth compression (such as compressing to 8 bits or 10 bits), etc.; the above-mentioned reducing resolution can be specifically achieved by different interpolation methods; the above-mentioned reducing resolution and reducing saturation can be collectively referred to as color distortion, etc.
[0045] After performing the degradation operation on the above-mentioned video samples, corresponding sample frame sequences can be extracted by skipping frames, that is to say, the image frames in the sample frame sequences are non-consecutive. In this solution, by extracting the sample frame sequences by skipping frames, the amount of data processed by the model can be reduced, and the model processing speed can be improved to a certain extent, the processing efficiency can be improved, and the time performance can be improved.
[0046] Of course, in practice, corresponding sample frame sequences can also be continuously extracted, and this specification does not limit this.
[0047] Finally, for the multiple sample frame sequences respectively extracted from multiple video samples, corresponding quality scores (abbreviated as annotation scores) can be manually annotated, and the annotation scores can specifically be subjective quality scores close to the human eye perception.
[0048] It should be understood that when training a video quality detection model based on the above-mentioned video samples that conform to the subjective human eye perception, it can be ensured that the scoring of the video to be detected by this model is more in line with the subjective human eye perception. In other words, the video quality detection model trained by this solution can also be called a video subjective quality detection model, and the obtained quality scores can also be called subjective quality scores.
[0049] Step S304, use the video quality detection model to process multiple sample frame sequences respectively to obtain the prediction scores of the multiple sample frame sequences.
[0050] Since the processing methods of the video quality detection model for each sample frame sequence are similar, therefore, the following takes any sample frame sequence as an example to illustrate its processing process.
[0051] Figure 4 Schematic diagram showing a method of processing a sample frame sequence using a video quality detection model in an example of this specification. Figure 4 In this method may include the following steps:
[0052] Step S41, partitioning the sample frame sequence to obtain a plurality of three-dimensional patches.
[0053] In one example, the size of the sample frame sequence can be expressed as T×H×W×3, where T represents the number of video frames in the sample frame sequence, H and W respectively represent the height and width of the video frame, and 3 represents the number of channels, which are the RGB channels respectively. In addition, the size of each three-dimensional patch can be expressed as a×b×c, where a, b, and c are all positive integers. Exemplarily, a, b, and c can be equal. Additionally, b and c can be equal, and a can be taken as 1 / 2 of b or c.
[0054] Assume that a takes the value of 2, b and c both take the value of 4. Correspondingly, the preset size is 2×4×4, as Figure 4 shown, for a video frame sequence with a size of T×H×W×3, after partitioning, non-overlapping T / 2×H / 4×W / 4 three-dimensional patches can be obtained, where the feature dimension of each three-dimensional patch is 2×4×4×3 = 96.
[0055] Step S42, inputting the plurality of three-dimensional patches into the video quality detection model. In this video quality detection model, based on the attention mechanism, global and local features of the sample frame sequence are extracted to obtain a spatio-temporal feature map of the sample frame sequence.
[0056] Among them, this feature extraction process includes 4 stages corresponding to 4 feature processing layers:
[0057] Stage 1: In the linear embedding layer, convolution operations are performed on each three-dimensional patch through three-dimensional convolutional kernels, and initial spatio-temporal feature maps of each three-dimensional patch are output, including a temporal feature map and a spatial feature map. For example, the number of three-dimensional convolutional kernels can be 96, the size of a single three-dimensional convolutional kernel is 2×4×4, and the stride is 2×4×4. Then, a spatio-temporal feature map with a size of T / 2×H / 4×W / 4×96 can be obtained.
[0058] Next, in two attention blocks, global and local features are extracted from the above initial spatio-temporal feature map. Specifically, in the two layers of attention layers respectively included in the two attention blocks, a three-dimensional sliding window and an offset sliding window are respectively slid along the temporal dimension and the spatial dimension on the initial spatio-temporal feature map, and global attention calculations are performed on the spatio-temporal feature maps within each sliding window, and an updated spatio-temporal feature map (hereinafter referred to as the updated spatio-temporal feature map) that fuses global and local features is output.
[0059] It should be understood that during the process of extracting features using paired attention blocks, the size of the spatio-temporal feature map remains unchanged. For example, the size of the updated spatio-temporal feature map is still T / 2×H / 4×W / 4×96.
[0060] Stage 2: In the block fusion layer, first combine adjacent spatial domain feature maps in each input spatio-temporal feature map (such as the above-mentioned updated spatio-temporal feature map), then reduce the dimension of the combined spatial domain feature maps while keeping the temporal domain feature maps unchanged, to obtain a target spatio-temporal feature map with a reduced spatial domain dimension. For the spatio-temporal feature map with a size of T / 2×H / 4×W / 4×96 in the previous example, when combining and then pooling and reducing the dimension of adjacent 2×2 spatial domain feature maps, the size of the obtained spatio-temporal feature map can be: T / 2×H / 8×W / 8×192. In this example, the width and height of the temporal domain feature map are both reduced by half compared to before, and in addition, the feature dimension is doubled compared to before.
[0061] Next, in the two attention blocks, extract global features and local features from the above-mentioned target spatio-temporal feature map. Specifically, in the two layers of attention layers included in the two attention blocks respectively, slide a three-dimensional sliding window and an offset sliding window along the temporal dimension and the spatial dimension on the spatio-temporal feature map with a reduced spatial domain dimension, and calculate the global attention for the spatio-temporal feature maps within each sliding window, and output a target spatio-temporal feature map that updates and fuses global and local features.
[0062] After stage 2 is completed, it can enter stage 3 (the corresponding feature processing layer of stage 3 includes 6 attention blocks) and stage 4 (the corresponding feature processing layer of stage 4 includes two attention blocks) in sequence. Among them, after passing through stage 3, the size of the obtained spatio-temporal feature map is: T / 2×H / 16×W / 16×384, and after passing through stage 4, the size of the obtained spatio-temporal feature map is: T / 2×H / 32×W / 32×768.
[0063] In summary, the above 4 stages will gradually reduce the spatial domain dimension of the input spatio-temporal feature map and gradually increase the feature dimension.
[0064] It should be understood that Figure 4 This is only an exemplary illustration. In practice, the above feature extraction process may also include 5 stages or even more. In addition, the number of attention blocks included in each layer of the feature processing layer can also be other numbers, as long as they appear in pairs, and this specification does not make any limitations in this regard.
[0065] Step S43, determine the prediction score of the sample frame sequence based on the spatio-temporal feature map of the sample frame sequence.
[0066] In addition, as described above, the above-mentioned video quality detection model may further include a pooling layer and an output layer. In the pooling layer, a max-pooling operation or an average-pooling operation may be performed on the final spatio-temporal feature map obtained by performing feature extraction processing (i.e., through processing at each stage) through each layer of the feature processing layer, to obtain a downsampled final spatio-temporal feature map. In the output layer, based on the downsampled final spatio-temporal feature map, the prediction score of the sample frame sequence is determined.
[0067] In one embodiment, the above-mentioned output layer may be implemented as a fully connected layer or a multi-layer perceptron MLP, etc.
[0068] Similarly, the prediction scores corresponding to each sample frame sequence in the training sample set can be obtained.
[0069] Returning to Figure 3 in Figure 3 the following steps may further be included:
[0070] Step S306: Determine the prediction loss according to the difference between the prediction scores of multiple sample frame sequences and the annotation scores, and based on this prediction loss, adjust the parameters of the video quality detection model.
[0071] In one embodiment, the mean square error (MSE) loss function may be used to calculate the prediction loss based on the prediction scores and annotation scores of multiple sample frame sequences.
[0072] The formula of the above MSE loss function may be as follows:
[0073]
[0074] where n is the number of sample frame sequences, f(x) is the prediction score of the sample frame sequence, and y is the annotation score of the sample frame sequence.
[0075] After calculating the prediction loss, using the backpropagation method, calculate the update gradients corresponding to the parameters of each layer of the network layer in the video quality detection model, and update the parameters of each layer of the network layer based on them, to obtain the trained video quality detection model.
[0076] After obtaining the trained video quality detection model, the model may further be adjusted according to a preset optimization method, so as to obtain an optimized video quality detection model. Here, the preset optimization method may include one or more of the following: reducing the number of attention heads in the video quality detection model, reducing the network layers of the video quality detection model, etc.
[0077] In addition, for the above-optimized video quality detection model, its format can also be converted into a specified format to obtain a target video quality detection model. Specifically, any model format conversion method in related technologies (such as the conversion method provided by pytorch mentioned later) can be adopted to convert the format of the optimized video quality detection model into a specified format to obtain a target video quality detection model. In some examples, the format of the optimized video quality detection model can be a related format of pytorch, and this specified format can be ONNX (Open Neural Network Exchange). Among them, pytorch is an open-source machine learning library that can implement the definition and training of models. ONNX is a cross-platform model exchange format used to share and use models between different deep learning frameworks. It allows users to train a model in one deep learning framework, convert it into the ONNX format, and then import the model into another framework that supports ONNX for inference.
[0078] It can be understood that the format of the optimized video quality detection model can also be other model formats that can perform model training, and this specified format can also be other formats that can achieve cross-platform model exchange.
[0079] Finally, for the above target video quality detection model, a specified model optimization tool can also be used to process it to reduce the redundant parameters and calculations of the model and obtain a final video quality detection model. In some possible examples, when the above-mentioned specified format is ONNX (Open Neural Network Exchange), this specified model optimization tool can be the ONNX SIN optimization tool. The ONNX SIN optimization tool is a tool specifically used to simplify and optimize models in ONNX format. It eliminates redundant parameters and calculations in the model by performing a series of graph transformation and constant folding operations, thereby reducing the complexity and size of the model and improving the inference efficiency of the model.
[0080] After obtaining the above video quality detection model (or optimized video quality detection model or target video quality detection model or final video quality detection model), its video quality can be detected, and the detection process will be described below.
[0081] Figure 5 FIG. shows a flowchart of a method for video quality detection according to an embodiment of this specification. This method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. Figure 5 In this case, the method may include the following steps:
[0082] Step S502: Extract a sequence of video frames from the target video to be detected.
[0083] For the above-mentioned target video, preprocessing can be performed first. Here, the preprocessing can include one or more of the following: rotating the vertical input video to a horizontal video (the horizontal video accounts for a relatively large proportion), scaling the resolution of the input video to a preset size (for example, 960*540), adjusting the aspect ratio of the input video to a preset ratio (for example, 16:9), converting the black-and-white input video to an RGB three-channel video, etc.
[0084] In the case where the above preprocessing is also performed, the video frame sequence can be extracted by skipping frames from the preprocessed target video, which can improve the detection efficiency of the video.
[0085] Step S504: Divide the video frame sequence into blocks to obtain multiple three-dimensional patches.
[0086] For example, the size of the video frame sequence can be: T×H×W×3, and the size of the three-dimensional patch can be: 2×4×4. In this way, T / 2×H / 4×W / 4 non-overlapping three-dimensional patches can be obtained. Among them, the feature dimension of each three-dimensional patch is 2×4×4×3 = 96.
[0087] Step S506: Input each three-dimensional patch into the video quality detection model. In this video quality detection model, based on the attention mechanism, feature extraction is performed on the video frame sequence to obtain the spatio-temporal feature map of the video frame sequence, and based on the spatio-temporal feature map of the video frame sequence, the quality score of the target video is determined.
[0088] Among them, the process of feature extraction for the video frame sequence is similar to the extraction process of the above sample frame sequence. For example, it can include 4 stages corresponding to 4 feature processing layers. After stage 1, a spatio-temporal feature map with a size of T / 2×H / 4×W / 4×96 can be obtained. Then, after stage 2, a spatio-temporal feature map with a size of T / 2×H / 8×W / 8×192 can be obtained. After that, after stage 3, a spatio-temporal feature map with a size of T / 2×H / 16×W / 16×384 can be obtained. Finally, after stage 4, a spatio-temporal feature map with a size of T / 2×H / 32×W / 32×768 can be obtained.
[0089] After performing the above feature extraction process, in the pooling layer, a maximum pooling operation or an average pooling operation can be performed on the extracted final spatio-temporal feature map to obtain the downsampled final spatio-temporal feature map. And in the output layer, based on the downsampled final spatio-temporal feature map, the quality score of the video frame sequence is determined.
[0090] In summary, in order to achieve the accuracy of existing subjective quality detection methods, this solution constructs a training sample set based on video samples with multiple sources, multiple resolutions, multiple scenarios, and uniform video quality levels. In addition, this solution develops a video quality detection model based on Transformer, which can accurately fit the existing widely used subjective quality detection methods and fully reflect the subjective feelings of the human eye. Finally, this solution optimizes the structure of the video detection model and conducts engineering optimization to achieve real-time detection on the CPU, significantly improving the time performance and reducing the computational cost. All in all, for the video quality detection method proposed in this solution, the relative error of subjective quality detection is about 1%, and the detection cost is reduced by 94% compared with the current commercial detection methods.
[0091] The innovative points of this solution are summarized as follows:
[0092] 1. This solution realizes an efficient video quality detection method, which significantly reduces the detection cost while ensuring the accuracy of subjective visual detection.
[0093] 2. This solution constructs a high-quality training sample set, which contains video data from various scenarios and various qualities in multiple dimensions.
[0094] 3. This solution develops a video quality detection model based on Transformer, which can accurately fit the existing widely used subjective quality detection methods.
[0095] 4. This solution optimizes the structure of the video quality detection model and conducts engineering optimization to achieve real-time detection on the CPU.
[0096] Corresponding to the above video quality detection method, an embodiment of this specification also provides a video quality detection device, as Figure 6 shown. The device may include:
[0097] An extraction unit 602, configured to extract a video frame sequence from a target video to be detected.
[0098] A partitioning unit 604, configured to partition the video frame sequence into blocks to obtain a plurality of three-dimensional patches.
[0099] A determination unit 606, configured to input each three-dimensional patch into the video quality detection model. In the video quality detection model, based on the attention mechanism, feature extraction is performed on the plurality of three-dimensional patches to obtain a spatio-temporal feature map of the video frame sequence, and based on the spatio-temporal feature map of the video frame sequence, a quality score of the target video is determined. Wherein, the video quality detection model is trained based on training samples, and the training samples are obtained by degrading video samples that meet preset conditions.
[0100] In one embodiment, the video quality detection model includes a linear embedding layer and at least one layer of block fusion layers. The determination unit 606 includes:
[0101] A convolution operation sub-module 6062, configured to perform a convolution operation on each three-dimensional tile in the linear embedding layer through a three-dimensional convolutional kernel, and output a first spatio-temporal feature map of each three-dimensional tile. The first spatio-temporal feature map of a single three-dimensional tile includes a temporal feature map and a spatial feature map;
[0102] A fusion sub-module 6064, configured to, in a single layer of block fusion layer, first combine adjacent spatial feature maps in each input spatio-temporal feature map, then perform dimensionality reduction on the combined spatial feature maps, and keep the temporal feature map unchanged, to obtain a second spatio-temporal feature map with reduced spatial dimension.
[0103] In one embodiment, the video quality detection model further includes a plurality of attention blocks respectively corresponding to the linear embedding layer and at least one layer of block fusion layers. The determination unit 606 further includes:
[0104] A calculation sub-module 6066, configured to, in a single attention block, respectively slide a three-dimensional sliding window or an offset sliding window along the temporal dimension and the spatial dimension on a target feature map, and perform global attention calculation on the target feature map within each sliding window, to obtain an updated target feature map that fuses global and local features, where the target feature map is the above-mentioned first spatio-temporal feature map or second spatio-temporal feature map.
[0105] In one embodiment, the video quality detection model further includes a pooling layer. The apparatus further includes:
[0106] A pooling operation unit 608, configured to perform a max-pooling operation or an average-pooling operation on the spatio-temporal feature map of the video frame sequence in the pooling layer, to obtain a spatio-temporal feature map with reduced dimension;
[0107] The determination unit 606 is specifically configured to:
[0108] Determine a quality score of the target video based on the spatio-temporal feature map with reduced dimension.
[0109] In one embodiment, the apparatus further includes:
[0110] A preprocessing unit 610, configured to preprocess the target video;
[0111] The extraction unit 602 is specifically configured to:
[0112] Extract a video frame sequence from the preprocessed target video;
[0113] The above preprocessing includes one or more of the following:
[0114] Rotate the longitudinal input video to a horizontal video;
[0115] Scale the resolution of the input video to a preset size;
[0116] Adjust the aspect ratio of the input video to a preset ratio;
[0117] Convert the black-and-white input video into an RGB three-channel video.
[0118] In one embodiment, the apparatus further includes: a training unit 612;
[0119] The training unit 612 is specifically configured to:
[0120] Obtain a training sample set, which includes a plurality of sample frame sequences respectively extracted from a plurality of video samples, and a single sample frame sequence has a corresponding annotation score;
[0121] Process the plurality of sample frame sequences respectively by using a video quality detection model to obtain prediction scores of the plurality of sample frame sequences;
[0122] Determine a prediction loss according to the difference between the prediction scores and the annotation scores of the plurality of sample frame sequences, and based on the prediction loss, adjust the parameters of the video quality detection model.
[0123] In one embodiment, at least one of the following characteristics is possessed by the above-mentioned plurality of video samples:
[0124] The video sources of the plurality of video samples are different;
[0125] The plurality of video samples have different resolutions;
[0126] The plurality of video samples are selected from different scenes;
[0127] A single video sample has a uniform video quality level.
[0128] In one embodiment, the above-mentioned degradation operations include one or more of the following:
[0129] Adding noise, blurring, compression, reducing resolution, reducing saturation, and reducing contrast.
[0130] The functions of the functional units of the apparatus in the above-mentioned embodiments of this specification can be implemented by the steps of the above-mentioned method embodiments. Therefore, the specific working process of the apparatus provided in an embodiment of this specification will not be repeated here.
[0131] The video quality detection apparatus provided in an embodiment of this specification can ensure accurate detection of video quality while significantly reducing the detection cost.
[0132] According to an embodiment of another aspect, there is also provided a computer-readable storage medium having a computer program stored thereon, which, when executed by a computer, causes the computer to execute the method described in connection with Figure 3 or Figure 5 .
[0133] According to an embodiment of yet another aspect, there is also provided a computing device including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in connection with Figure 3 or Figure 5 is implemented.
[0134] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the medium or device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0135] The steps of the method or algorithm described in connection with the disclosure of this specification can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable hard disk, a CD-ROM, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC. Additionally, the ASIC can be located in a server. Of course, the processor and the storage medium can also exist as discrete components in a server.
[0136] In the 1990s, it was obvious to distinguish whether an improvement in a technology was an improvement in hardware (e.g., improvement in circuit structures such as diodes, transistors, switches, etc.) or an improvement in software (improvement in method processes). However, with the development of technology, many improvements in method processes today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structures by programming the improved method processes into the hardware circuits. Therefore, it cannot be said that an improvement in a method process cannot be implemented with hardware entity modules. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user's programming of the device. Designers can program by themselves to "integrate" a digital system on a piece of PLD without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called Hardware Description Language (HDL), and there is not only one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply making a little logical programming of the method process with the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method process.
[0137] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26k20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0138] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the development of future computer technologies, the computers implementing the functions of the above embodiments can be, for example, personal computers, laptop computers, in-vehicle human-machine interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or any combination of these devices.
[0139] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual device or terminal product is executed, it may be executed in the order of the method shown in the embodiments or the drawings or executed in parallel (for example, in a parallel processor or multi-threaded processing environment, or even in a distributed data processing environment). The terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, product or device comprising a series of elements not only includes those elements but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, there is no exclusion of additional identical or equivalent elements in the process, method, product or device comprising the said elements. For example, if terms such as first and second are used to denote names, they do not denote any particular order.
[0140] For convenience of description, when describing the above device, it is divided into various modules according to functions for separate description. Of course, when implementing one or more of this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0141] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0142] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0144] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0145] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0146] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage, graphene storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0147] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0148] One or more embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0149] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content. In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0150] The above is only the embodiment of one or more embodiments of this specification, and is not used to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims.
Claims
1. A method for video quality detection, comprising: Extracting a video frame sequence from a target video to be detected; Dividing the video frame sequence into blocks to obtain a plurality of three-dimensional image blocks; Inputting each three-dimensional image block into a video quality detection model, in which feature extraction is performed on the multiple three-dimensional image blocks based on an attention mechanism to obtain a spatiotemporal feature map of the video frame sequence; Based on the spatiotemporal feature map of the video frame sequence, the quality score of the target video is determined, wherein the video quality detection model is trained based on training samples, and the training samples are obtained by degrading video samples that meet preset conditions.
2. The method according to claim 1, wherein: The video quality detection model includes a linear embedding layer and at least one block fusion layer; the spatiotemporal feature map of the video frame sequence is obtained, including: In the linear embedding layer, a convolution operation is performed on each of the three-dimensional image blocks through a three-dimensional convolution kernel to output a first spatiotemporal feature map of each of the three-dimensional image blocks; the first spatiotemporal feature map of a single three-dimensional image block includes a temporal feature map and a spatial feature map; In the single-layer block fusion layer, the adjacent spatial feature maps in each input spatiotemporal feature map are first combined, and then the combined spatial feature map is reduced in dimension, while the temporal feature map is kept unchanged to obtain a second spatiotemporal feature map with reduced spatial dimension.
3. The method according to claim 2, wherein: The video quality detection model further includes a plurality of attention blocks corresponding to the linear embedding layer and at least one block fusion layer respectively; In a single attention block, a three-dimensional sliding window or an offset sliding window is slid on a target feature map along the time domain dimension and the spatial dimension, and a global attention calculation is performed on the target feature map in each sliding window to obtain an updated target feature map that integrates global and local features; the target feature map is the first spatiotemporal feature map or the second spatiotemporal feature map.
4. The method according to claim 1, wherein: The video quality detection model also includes a pooling layer; the method also includes: In the pooling layer, a maximum pooling operation or an average pooling operation is performed on the spatiotemporal feature map of the video frame sequence to obtain a spatiotemporal feature map after dimensionality reduction; Determining the quality score of the target video includes: Based on the spatiotemporal feature graph after dimensionality reduction, a quality score of the target video is determined.
5. The method according to claim 1, further comprising: Preprocessing the target video; The step of extracting a video frame sequence from a target video to be detected comprises: Extracting a video frame sequence from the preprocessed target video; The pre-processing includes one or more of the following: Rotate the vertical input video to a horizontal video; Scale the resolution of the input video to a preset size; Adjust the aspect ratio of the input video to the preset ratio; Convert black and white input video to RGB three-channel video.
6. The method according to claim 1, wherein: The video quality detection model is trained by the following steps: Obtaining a training sample set, which includes a plurality of sample frame sequences extracted from a plurality of video samples, wherein each sample frame sequence has a corresponding annotation score; Using the video quality detection model to process the multiple sample frame sequences respectively to obtain prediction scores for the multiple sample frame sequences; A prediction loss is determined according to a difference between the predicted scores and the labeled scores of each of the plurality of sample frame sequences, and a parameter of the video quality detection model is adjusted based on the prediction loss.
7. The method according to claim 6, wherein: The multiple video samples have at least one of the following characteristics: The video sources of the multiple video samples are different; The plurality of video samples have different resolutions; The multiple video samples are selected from different scenes; A single video sample has a uniform video quality level.
8. The method according to claim 1, wherein: The degradation operation includes one or more of the following: Add noise, blur, compress, reduce resolution, desaturate, and lower contrast.
9. A device for video quality detection, comprising: An extraction unit, used for extracting a video frame sequence from a target video to be detected; A dividing unit, used for dividing the video frame sequence into blocks to obtain a plurality of three-dimensional blocks; A determination unit, configured to input each three-dimensional image block into a video quality detection model, in which features of the plurality of three-dimensional image blocks are extracted based on an attention mechanism to obtain a spatiotemporal feature map of the video frame sequence; Based on the spatiotemporal feature map of the video frame sequence, the quality score of the target video is determined, wherein the video quality detection model is trained based on training samples, and the training samples are obtained by degrading video samples that meet preset conditions.
10. A computing device, comprising a memory and a processor, wherein the memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Intelligent diagnosis equipment for audio and video quality
CN122093548A