Video quality detection method, device, equipment and storage medium

CN116389711BActive Publication Date: 2026-08-11PENG CHENG LAB
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-23
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]本发明的主要目的在于提供一种视频质量检测方法、装置、设备及存储介质,旨在解决现有技术中如何全面、多方位且高效对视频的质量进行检测的技术问题

Benefits of technology

[0015] This invention acquires a video to be tested; extracts features from the video according to a preset perceptual classification feature set to obtain multiple perceptual distortion dimensions; determines the overall quality score of the video to be tested based on the multiple perceptual distortion dimensions and a target preference probability regression model; evaluates the video score based on the multiple perceptual distortion dimensions and outputs the dimensional quality score for each perceptual distortion dimension; and completes the video quality detection of the video to be tested based on the overall quality score and the dimensional quality scores for each perceptual distortion dimension. The above method extracts features from the video to be tested based on a preset perceptual classification feature set, obtaining multiple perceptual dimension features under multiple perceptual distortion dimensions. A comprehensive quality score is determined based on these features and a target preference probability regression model. Video scores are then evaluated based on these features, outputting dimensional quality scores for each perceptual distortion dimension. Finally, video quality detection is completed based on the comprehensive quality score and the dimensional quality scores for each perceptual distortion dimension. This method achieves comprehensive and effective video quality detection, providing not only a comprehensive evaluation result but also individual evaluation results for each perceptual distortion dimension. It fully reflects the video quality, effectively displays video distortion, and improves the efficiency and accuracy of video quality detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116389711B_ABST
    Figure CN116389711B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of video inspection technology and discloses a video quality inspection method, apparatus, device, and storage medium. The method includes: acquiring a video to be inspected; extracting features from the video to be inspected according to a preset perceptual classification feature set to obtain multiple perceptual dimension features of multiple perceptual distortion dimensions; determining a comprehensive quality score based on the multiple perceptual dimension features of each perceptual distortion dimension and a target preference probability regression model; evaluating the video score based on the multiple perceptual dimension features of each perceptual distortion dimension and outputting the dimensional quality score of each perceptual distortion dimension; and completing video quality inspection based on the comprehensive quality score and the dimensional quality scores of each perceptual distortion dimension. Through the above method, comprehensive and effective video quality inspection is achieved, obtaining not only a comprehensive evaluation result of the video but also individual evaluation results under each perceptual distortion dimension, fully reflecting the video quality, and improving the efficiency and accuracy of video quality inspection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video inspection technology, and in particular to a video quality inspection method, apparatus, device, and storage medium. Background Technology

[0002] Currently, digital video is widely used throughout society. However, due to quality degradation at each stage of video signal acquisition, compression, transmission, and display, the quality of video directly affects the user's viewing experience and the accuracy of artificial intelligence algorithms in understanding and analyzing video content. Video quality evaluation plays a crucial role in various related fields such as video compression coding, adaptive streaming media, video transmission systems, and preprocessing for AI systems.

[0003] Since the advent of deep learning technology, most current research on no-reference video quality assessment has been based on deep neural networks. However, as the complexity of the network structure and the number of network layers increase, the computational complexity of video quality assessment models also increases, and the computation becomes more time-consuming. On the other hand, the performance of deep learning-based video quality assessment models is highly dependent on the input training data. At present, some traditional quality assessment methods can also achieve high accuracy and save significantly more computational power than deep learning-based methods, and do not rely heavily on training data. However, the modeling of traditional video quality assessment methods is mostly relatively simple. They usually design features based only on specific visual characteristics or select features based on predictive relevance, without fully integrating features from the perspective of the overall human visual system. This results in features that are difficult to comprehensively and effectively represent the quality distortion of images and videos. Summary of the Invention

[0004] The main objective of this invention is to provide a video quality detection method, apparatus, device, and storage medium, aiming to solve the technical problem of how to comprehensively, multi-dimensionally, and efficiently detect the quality of video in the prior art.

[0005] To achieve the above objectives, the present invention provides a video quality detection method, the video quality detection method comprising: Obtain the video to be tested; Based on a preset perceptual classification feature set, feature extraction is performed on the video to be detected to obtain multiple perceptual dimension features with multiple perceptual distortion dimensions. The overall quality score of the video to be detected is determined based on the multiple perceptual dimension features of each perceptual distortion dimension and the target preference probability regression model. Video scores are evaluated based on multiple perceptual dimension features of each perceptual distortion dimension, and the dimension quality score of each perceptual distortion dimension is output. The video quality of the video to be tested is completed based on the overall quality score and the dimensional quality scores of each perceptual distortion dimension.

[0006] Optionally, before determining the overall quality score of the video to be detected based on multiple perceptual dimension features of each perceptual distortion dimension and a target preference probability regression model, the method further includes: Obtain the sample training video set; Construct multiple sample training video pairs based on the sample training video set; Based on the preset perceptual classification feature set, feature extraction is performed on each sample training video pair to obtain the first training feature of the first sample video and the second training feature of the second sample video in each sample training video pair. The initial preference probability regression model is trained based on the first training feature and the second training feature to obtain the target preference probability regression model.

[0007] Optionally, the step of training the initial preference probability regression model based on the first training features and the second training features to obtain the target preference probability regression model includes: Input the first training feature into the first sub-network of the Siamese structure network to obtain the first prediction mean and the first prediction variance; The second training feature is input into the second sub-network of the Siamese structure network to obtain the second prediction mean and the second prediction variance. Input the first prediction mean, the first prediction variance, the second prediction mean, and the second prediction variance into the preference probability model to determine the prediction preference probability of each training video pair. The initial preference probability regression model is adjusted based on the predicted preference probability of each training video pair to obtain the target preference probability regression model.

[0008] Optionally, the step of evaluating video scores based on multiple perceptual dimension features of each perceptual distortion dimension and outputting a dimensional quality score for each perceptual distortion dimension includes: The mean of the first feature is obtained by calculating the mean of multiple perceptual dimension features for each perceptual distortion dimension. The difference between each perceptual dimension feature and the mean of the first feature is calculated to obtain the first decentralized data; Perform matrix calculations on the first decentralized data to determine multiple first eigenvalues ​​and multiple first eigenvectors; The video score is evaluated based on each first feature value and each first feature vector, and the dimensional quality score of each perceptual distortion dimension is output.

[0009] Optionally, the step of evaluating video scores based on each first feature value and each first feature vector, and outputting a dimensional quality score for each perceptual distortion dimension, includes: Sort the first eigenvalues, and determine multiple second eigenvectors and multiple second eigenvalues ​​in each first eigenvector based on the sorting results; The mean of the second feature is obtained by calculating the mean of multiple second feature vectors based on multiple second feature values; The difference between each second feature vector and the mean of the second feature is calculated to obtain the second decentralized data; Matrix calculations are performed based on the second decentralized data to determine multiple third eigenvalues ​​and third eigenvectors; Sort the third feature values ​​and determine the target feature vector from the third feature vectors based on the sorting results; Based on the target feature vector, a video score is evaluated, and the dimensional quality score for each perceptual distortion dimension is output.

[0010] Optionally, the step of evaluating video scores based on the target feature vector and outputting dimensional quality scores for each perceptual distortion dimension includes: The initial space is obtained by constructing a space based on the target feature vector; Based on the initial space, feature transformation is performed on the features of each perceptual dimension to obtain the target space; The dimensional quality score of each perceptual distortion dimension is output based on the target space.

[0011] Optionally, after performing video quality detection on the video to be detected based on the overall quality score and the dimensional quality scores of each perceptual distortion dimension, the method further includes: Obtain the dimensional score threshold for each dimension of perceived distortion; The dimensional quality score of each perceptual distortion dimension is compared with the dimensional score threshold of each perceptual distortion dimension to obtain the first comparison result; The dimensions to be improved in the video to be detected are determined based on the first comparison result; A quality improvement strategy is generated based on the dimensions to be improved.

[0012] Furthermore, to achieve the above objectives, the present invention also proposes a video quality detection device, the video quality detection device comprising: The acquisition module is used to acquire the video to be detected; The extraction module is used to extract features from the video to be detected based on a preset perceptual classification feature set, thereby obtaining multiple perceptual dimension features with multiple perceptual distortion dimensions. The determination module is used to determine the overall quality score of the video to be detected based on the multiple perceptual dimension features of each perceptual distortion dimension and the target preference probability regression model. The evaluation module is used to evaluate video scores based on multiple perceptual dimension features of each perceptual distortion dimension, and output the dimension quality score of each perceptual distortion dimension. The completion module is used to perform video quality detection on the video to be detected based on the overall quality score and the dimensional quality scores of each perceived distortion dimension.

[0013] Furthermore, to achieve the above objectives, the present invention also proposes a video quality detection device, which includes: a memory, a processor, and a video quality detection program stored in the memory and executable on the processor, wherein the video quality detection program is configured to implement the video quality detection method described above.

[0014] In addition, to achieve the above objectives, the present invention also proposes a storage medium storing a video quality detection program, which, when executed by a processor, implements the video quality detection method as described above.

[0015] This invention acquires a video to be tested; extracts features from the video according to a preset perceptual classification feature set to obtain multiple perceptual distortion dimensions; determines the overall quality score of the video to be tested based on the multiple perceptual distortion dimensions and a target preference probability regression model; evaluates the video score based on the multiple perceptual distortion dimensions and outputs the dimensional quality score for each perceptual distortion dimension; and completes the video quality detection of the video to be tested based on the overall quality score and the dimensional quality scores for each perceptual distortion dimension. The above method extracts features from the video to be tested based on a preset perceptual classification feature set, obtaining multiple perceptual dimension features under multiple perceptual distortion dimensions. A comprehensive quality score is determined based on these features and a target preference probability regression model. Video scores are then evaluated based on these features, outputting dimensional quality scores for each perceptual distortion dimension. Finally, video quality detection is completed based on the comprehensive quality score and the dimensional quality scores for each perceptual distortion dimension. This method achieves comprehensive and effective video quality detection, providing not only a comprehensive evaluation result but also individual evaluation results for each perceptual distortion dimension. It fully reflects the video quality, effectively displays video distortion, and improves the efficiency and accuracy of video quality detection. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the structure of a video quality testing device in the hardware operating environment involved in the embodiments of the present invention; Figure 2 This is a flowchart illustrating the first embodiment of the video quality detection method of the present invention; Figure 3 This is a schematic diagram of the overall process of an embodiment of the video quality detection method of the present invention; Figure 4 This is an overall verification comparison diagram of an embodiment of the video quality detection method of the present invention; Figure 5 This is a dimensional verification comparison diagram of an embodiment of the video quality detection method of the present invention; Figure 6 This is a flowchart illustrating the second embodiment of the video quality detection method of the present invention; Figure 7 This is a schematic diagram of the model structure of an embodiment of the video quality detection method of the present invention; Figure 8 This is a structural block diagram of the first embodiment of the video quality detection device of the present invention.

[0017] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0018] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0019] Reference Figure 1 , Figure 1 This is a schematic diagram of the hardware operating environment of the video quality testing device involved in the embodiments of the present invention.

[0020] like Figure 1As shown, the video quality inspection device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0021] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on video quality inspection equipment and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0022] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a network communication module, a user interface module, and a video quality detection program.

[0023] exist Figure 1 In the video quality detection device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the video quality detection device of the present invention can be set in the video quality detection device, and the video quality detection device calls the video quality detection program stored in the memory 1005 through the processor 1001 and executes the video quality detection method provided in the embodiment of the present invention.

[0024] This invention provides a video quality detection method, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of a video quality detection method according to the present invention.

[0025] Video quality inspection methods include the following steps: Step S10: Obtain the video to be tested.

[0026] It should be noted that the execution subject in this embodiment is a terminal device, such as a computer, tablet, or mobile phone. The terminal device is equipped with a video quality inspection system. This system acquires the video to be inspected and inputs it into a target preference probability regression model to determine the overall quality score of the video. It then extracts features from the video based on a preset perceptual classification feature set, obtaining multiple perceptual dimension features for each perceptual distortion dimension. Based on these features, the system evaluates the video score and outputs the dimensional quality score for each perceptual distortion dimension. Finally, it completes the video quality inspection of the video based on the overall quality score and the dimensional quality scores for each perceptual distortion dimension.

[0027] It is understandable that the video to be tested refers to the video that needs to undergo quality testing and evaluation.

[0028] Step S20: Extract features from the video to be detected based on a preset perceptual classification feature set to obtain multiple perceptual dimension features of multiple perceptual distortion dimensions.

[0029] It should be noted that the preset perceptual classification feature set refers to a pre-summarized set of perceptual classification features. In practical applications related to image and video processing, distortion may be introduced at various stages of video acquisition, editing, transmission, and storage, and the factors causing distortion are numerous. However, from the perspective of visual reception perception, distortion can be summarized and classified. In this embodiment, the various stages of image and video processing and their corresponding operations may cause visual distortion effects. Then, the perceived visual effects are integrated, resulting in six distortion measurement dimensions based on visual perception: brightness, color, contrast, sharpness, noise, and temporal motion. Brightness, color, contrast, sharpness, noise, and temporal motion are the multiple perceptual distortion dimensions in this embodiment. The visual effects corresponding to the relevant operations in image and video processing are shown in Table 1.

[0030] Table 1 The categories of visual effects are identified by numbers in the table, with the numbers representing the dimensions of perceptual distortion as follows: 1-brightness, 2-color, 3-contrast, 4-sharpness, 5-noise, 6-temporal motion.

[0031] Understandably, in order to avoid the situation where the feature distribution is messy and a relatively systematic and standardized feature summarization system cannot be formed, in this embodiment, based on multiple perceptual distortion dimensions, the realized features are summarized and integrated to obtain a perceptual classification feature set containing 2096 features. The relevant attributes and descriptions of all features under the six perceptual distortion dimensions are shown in Tables 2 to 7. All features under the six perceptual distortion dimensions are the summarized perceptual classification feature set.

[0032] Table 2 Table 3 In the specific implementation, features are extracted from the video to be detected based on the features listed in the preset perceptual classification feature set, resulting in multiple perceptual dimension features under each perceptual distortion dimension.

[0033] Step S30: Determine the comprehensive quality score of the video to be detected based on the multiple perceptual dimension features of each perceptual distortion dimension and the target preference probability regression model.

[0034] It should be noted that the target preference probability regression model is obtained by training a network including the Siamese structure network and the preference probability model. After inputting all perceptual dimension features under all perceptual distortion dimensions of the video to be detected into the target preference probability model, the target preference probability regression model can output the overall quality score of the video to be detected. The overall quality score is the comprehensive quality score.

[0035] Step S40: Evaluate the video score based on the multiple perceptual dimension features of each perceptual distortion dimension, and output the dimension quality score of each perceptual distortion dimension.

[0036] It should be noted that, in order to achieve comprehensive quality evaluation, unsupervised regression is performed on multiple perceptual dimension features under each perceptual distortion dimension to obtain video quality evaluation scores under each perceptual distortion dimension. These video quality evaluation scores under each perceptual distortion dimension are the dimension quality scores. In this embodiment, to avoid the direct dimensionality reduction regression affecting the accuracy of the results, a two-stage regression strategy based on PCA (Principal Component Analysis) is proposed to achieve regression of unsupervised perceptual classification quality.

[0037] Understandably, in order to obtain an accurate dimension quality score based on multiple perceptual dimension features of each perceptual distortion dimension, the step of evaluating video scores based on multiple perceptual dimension features of each perceptual distortion dimension and outputting the dimension quality score of each perceptual distortion dimension further includes: calculating the mean of multiple perceptual dimension features of each perceptual distortion dimension to obtain a first feature mean; calculating the difference between each perceptual dimension feature and the first feature mean to obtain first decentralized data; performing matrix calculation on the first decentralized data to determine multiple first feature values ​​and multiple first feature vectors; and evaluating video scores based on each first feature value and each first feature vector to output the dimension quality score of each perceptual distortion dimension.

[0038] In the specific implementation, the mean of multiple perceptual dimension features under each perceptual distortion dimension is calculated to obtain the first feature mean. The difference between each perceptual dimension feature and the first feature mean is then calculated to obtain the first decentralized data. Specifically, the process is as follows: Decentralization is performed on all perceptual dimension features under each perceptual distortion dimension of the video to be detected. First, the average of all perceptual dimension features under each perceptual distortion dimension is calculated; this average is the first feature mean. After determining the first feature mean, the difference between each perceptual dimension feature and the first feature mean is calculated, thus completing the decentralization of all perceptual dimension features under each perceptual distortion dimension of the video to be detected, obtaining decentralized data under each perceptual distortion dimension. This decentralized data under each perceptual distortion dimension is the first decentralized data. For example, if there are 9 perceptual dimension features under the brightness dimension, and the mean of these 9 features is 'a', then the first feature mean under the brightness dimension is 'a'.

[0039] It should be noted that matrix calculations are performed on the first decentralized data to determine multiple first eigenvalues ​​and multiple first eigenvectors. The specific process is as follows: calculate the covariance matrix of the first decentralized data for each perceptual distortion dimension, as well as the eigenvalues ​​and eigenvectors of the matrix. The calculated eigenvalues ​​are the first eigenvalues, and the eigenvectors are the first eigenvectors.

[0040] Understandably, after determining multiple first feature values ​​and multiple first feature vectors for each perceptual distortion dimension, regression calculations are performed based on these first feature values ​​and first feature vectors to evaluate video scores under each perceptual distortion dimension, outputting the dimensional quality score for each perceptual distortion dimension. To ensure the accuracy of the output dimensional quality scores, the process of evaluating video scores based on each first feature value and each first feature vector, and outputting the dimensional quality score for each perceptual distortion dimension, further includes: sorting each first feature value; determining multiple second feature vectors and multiple second feature values ​​in each first feature vector based on the sorting results; calculating the mean of multiple second feature vectors based on the multiple second feature values ​​to obtain a second feature mean; calculating the difference between each second feature vector and the second feature mean to obtain second decentralized data; performing matrix calculations based on the second decentralized data to determine multiple third feature values ​​and third feature vectors; sorting each third feature value; determining a target feature vector in each third feature vector based on the sorting results; and evaluating video scores based on the target feature vector to output the dimensional quality score for each perceptual distortion dimension.

[0041] In the specific implementation, each first feature value is sorted, and multiple second feature vectors are determined from the first feature vectors based on the sorting results. The specific process is as follows: the first feature vectors are used to perform the first stage of regression calculation, the first feature values ​​are sorted from largest to smallest, and the feature vectors with feature values ​​greater than 1 are selected to perform the second stage of regression calculation. The selected feature vectors are the second feature vectors.

[0042] It should be noted that the process involves calculating the mean of multiple second feature vectors based on multiple second feature values ​​to obtain the second feature mean. The differences between each second feature vector and the second feature mean are then calculated to obtain the second decentralized data. Matrix calculations are then performed on this second decentralized data to determine multiple third feature values ​​and third feature vectors. Specifically, the above steps are repeated for the selected second feature vectors to decentralize the second feature vectors under each perceptual distortion dimension. First, the average value of all second feature vectors under each perceptual distortion dimension is calculated; this average value is the second feature mean. After determining the second feature mean, the differences between each second feature vector under each perceptual distortion dimension and the second feature mean are calculated. This completes the decentralization of all second feature vectors under each perceptual distortion dimension of the video to be detected, resulting in the decentralized data for all second feature vectors under each perceptual distortion dimension. This decentralized data is the second decentralized data. Calculate the covariance matrix of the second decentralized data for each perceptual distortion dimension, as well as the eigenvalues ​​and eigenvectors of this matrix. The calculated values ​​are the third eigenvalues, and the eigenvectors are the third eigenvectors.

[0043] It is understandable that the third feature values ​​are sorted, and the target feature vector is determined from each third feature vector based on the sorting result. The specific process is as follows: sort the third feature values ​​under each perceptual distortion dimension, determine the feature vector corresponding to the largest feature value, and the feature vector corresponding to the largest feature value is the target feature vector.

[0044] In specific implementation, video quality scores for each perceptual distortion dimension can be evaluated based on the target feature vectors under each perceptual dimension. To ensure the accuracy of the evaluation process, the step of evaluating video scores based on the target feature vectors and outputting the dimensional quality scores for each perceptual distortion dimension further includes: constructing a space based on the target feature vectors to obtain an initial space; performing feature transformation on the features of each perceptual dimension based on the initial space to obtain a target space; and outputting the dimensional quality scores for each perceptual distortion dimension based on the target space.

[0045] It should be noted that after determining the target feature vectors for each perceptual distortion dimension, a space is constructed based on the target feature vectors to obtain an initial space. Feature transformation is then performed on the features of each perceptual dimension based on the initial space to obtain the target space. Finally, the dimensional quality score for each perceptual distortion dimension is output based on the target space. The specific process is as follows: all perceptual dimension features under each perceptual distortion dimension are transformed into the initial space constructed by the target feature vectors to obtain the target space. The video quality score under each perceptual distortion dimension is regressed based on the target space, thereby outputting the dimensional quality score for each perceptual distortion dimension.

[0046] Step S50: Perform video quality detection on the video to be detected based on the overall quality score and the dimensional quality scores of each perceptual distortion dimension.

[0047] It should be noted that the overall quality score reflects the overall quality of the video under test, while the dimensional quality scores for each perceptual distortion dimension reflect the brightness quality in the brightness dimension, the color quality in the color dimension, the contrast quality in the contrast dimension, the sharpness quality in the sharpness dimension, the noise quality in the noise dimension, and the temporal motion quality in the temporal motion dimension, respectively. The video quality is then assessed based on the output overall quality score and the dimensional quality scores for each perceptual distortion dimension.

[0048] Understandably, in order to ensure that video quality can be effectively improved, after completing the video quality detection of the video to be detected based on the comprehensive quality score and the dimensional quality scores of each perceptual distortion dimension, the method further includes: obtaining the dimensional score threshold of each perceptual distortion dimension; comparing the dimensional quality score of each perceptual distortion dimension with the dimensional score threshold of each perceptual distortion dimension to obtain a first comparison result; determining the dimension to be improved of the video to be detected based on the first comparison result; and generating a quality improvement strategy based on the dimension to be improved.

[0049] In practice, the dimension score threshold refers to the critical score corresponding to each perceptual distortion dimension. When the dimension quality score of each perceptual distortion dimension is less than the dimension score threshold, it indicates that the quality of that perceptual distortion dimension urgently needs to be improved.

[0050] It should be noted that the dimensional quality score for each perceptual distortion dimension is compared with its corresponding dimensional quality score to determine the first comparison result. Based on the first and second comparison results, dimensions with scores below a threshold are identified, and these dimensions are those requiring improvement. For example, if the dimensional quality score for the brightness dimension is 30, and the threshold for the brightness dimension is 60, the dimensional quality score for the brightness dimension is less than the threshold for the brightness dimension. Therefore, the dimension requiring improvement in the perceptual distortion dimensions is the brightness dimension.

[0051] Understandably, after determining the dimension to be improved, the factors affecting the dimension to be improved are obtained, and a quality improvement strategy for the video to be tested is generated based on the influencing factors. The influencing factors corresponding to each perceptual distortion dimension are all pre-stored in the terminal device.

[0052] In specific implementations, such as Figure 3 As shown, features are extracted from the video to be detected based on a preset perceptual feature set, resulting in perceptual dimension features under six perceptual distortion dimensions. Unsupervised perceptual distortion classification quality regression is then performed based on these perceptual dimension features to determine the dimension quality score for each perceptual distortion dimension. Finally, the overall quality score of the video to be detected is determined based on the perceptual dimension features under the perceptual distortion dimensions and the target preference probability regression model. The results of the video quality detection method proposed in this embodiment are as follows: Figure 4 As shown, where Figure 4The values ​​in the upper left or lower right corner of the radar chart represent the subjective score label (MOS) of the video to be detected. The radar chart indicates the measurement results of video quality by the method proposed in this embodiment across six perceptual dimensions: brightness, color, contrast, sharpness, noise, and temporal motion. As can be seen from the radar chart, the area enclosed by the measurement results in each perceptual dimension consistently increases with the increase of the subjective score of the sample, and videos with poor quality in each dimension correspond to subjective perception. Furthermore, this embodiment verifies the predictive performance of the dimensional quality scores output in each perceptual distortion dimension by subjectively analyzing specific video samples. Based on the output results, the video sample with the lowest predicted score in each of the six perceptual dimensions (brightness, color, contrast, sharpness, noise, and temporal motion) is selected and subjectively viewed and analyzed. The subjective label (MOS), multi-dimensional measurement results, and video content description of the video are shown in the appendix. Figure 5 As shown in the figure. Subjective verification experiment results show that the dimensional quality scores output in this embodiment can achieve relatively accurate predictions under the six perceptual distortion dimensions.

[0053] This embodiment acquires a video to be tested; extracts features from the video based on a preset perceptual classification feature set to obtain multiple perceptual distortion dimensions; determines the overall quality score of the video based on the multiple perceptual distortion dimensions and a target preference probability regression model; evaluates the video score based on the multiple perceptual distortion dimensions and outputs the dimensional quality score for each perceptual distortion dimension; and completes the video quality detection of the video based on the overall quality score and the dimensional quality scores for each perceptual distortion dimension. The above method extracts features from the video to be tested based on a preset perceptual classification feature set, obtaining multiple perceptual dimension features under multiple perceptual distortion dimensions. A comprehensive quality score is determined based on these features and a target preference probability regression model. Video scores are then evaluated based on these features, outputting dimensional quality scores for each perceptual distortion dimension. Finally, video quality detection is completed based on the comprehensive quality score and the dimensional quality scores for each perceptual distortion dimension. This method achieves comprehensive and effective video quality detection, providing not only a comprehensive evaluation result but also individual evaluation results for each perceptual distortion dimension. It fully reflects the video quality, effectively displays video distortion, and improves the efficiency and accuracy of video quality detection.

[0054] refer to Figure 6 , Figure 6 This is a flowchart illustrating a second embodiment of a video quality detection method according to the present invention.

[0055] Based on the first embodiment described above, the video quality detection method of this embodiment, before step S30, further includes: Step S31: Obtain the sample training video set.

[0056] It should be noted that the sample training set refers to a collection containing a large number of sample videos.

[0057] Step S32: Construct multiple sample training video pairs based on the sample training video set.

[0058] It should be noted that since two inputs are required when training the initial preference probability regression model, the sample videos in the sample training video need to be randomly paired to form multiple sample training video pairs. Each sample training video pair contains two sample videos, namely the first sample video and the second sample video.

[0059] Step S33: Extract features from each training video pair based on the preset perceptual classification feature set to obtain the first training feature of the first sample video and the second training feature of the second sample video in each training video pair.

[0060] It should be noted that, based on the features listed in the preset perceptual classification feature set, feature extraction is performed on the sample videos in each sample training video pair to obtain the first training feature of the first sample video and the second training feature of the second sample video in each sample training video pair.

[0061] Step S34: Train the initial preference probability regression model based on the first training feature and the second training feature to obtain the target preference probability regression model.

[0062] It should be noted that after determining the first and second training features, the initial preference probability regression model is trained to obtain the trained initial preference probability regression model, which is the target preference probability regression model. The initial preference probability regression model is a network that includes a Siamese network structure and a preference probability model.

[0063] Understandably, to ensure the accuracy of the model training process and obtain a high-performance target preference probability regression model, the step of training the initial preference probability regression model based on the first training feature and the second training feature to obtain the target preference probability regression model further includes: inputting the first training feature into the first sub-network of the Siamese structure network to obtain a first prediction mean and a first prediction variance; inputting the second training feature into the second sub-network of the Siamese structure network to obtain a second prediction mean and a second prediction variance; inputting the first prediction mean, the first prediction variance, the second prediction mean, and the second prediction variance into the preference probability model to determine the predicted preference probability of each sample training video pair; and adjusting the initial preference probability regression model based on the predicted preference probability of each sample training video pair to obtain the target preference probability regression model.

[0064] In practical implementation, from a psychological perspective, Thurstone proposed that people's relative judgments statistically approximately follow a Gaussian distribution. In this embodiment, the Thurstone psychological model is used to model the probability of preference for video quality; therefore, the difference in subjective ratings between video sample i and video sample j follows a mean of... The variance is The Gaussian distribution, with respect to the preference probability that video sample i is better than video sample j, is modeled as follows: ,in is the Gaussian cumulative distribution function.

[0065] It should be noted that inputting the first training feature into the first sub-network of the Siamese structure network yields the first predicted mean and the first predicted variance. Inputting the second training feature into the second sub-network of the Siamese structure network yields the second predicted mean and the second predicted variance. Specifically, during the regression process, the first and second training features are input into the Siamese structure network. The two sub-networks with shared weights calculate the corresponding predicted mean and predicted variance respectively. The predicted variance and predicted mean output by the first sub-network are the first predicted variance and the first predicted mean, and the predicted variance and predicted mean output by the second sub-network are the second predicted variance and the second predicted mean.

[0066] Understandably, the first predicted mean, the first predicted variance, the second predicted mean, and the second predicted variance are input into the preference probability model. The preference probability model then fuses the outputs of the two subnets of the Siamese network to obtain the predicted preference probability. The initial preference probability regression model is adjusted based on the predicted preference probabilities of the training video pairs to obtain the target preference probability regression model. That is, the initial preference probability regression model is trained using each training video pair, and the network's fitting objective is the predicted preference probability of the quality between samples.

[0067] In its specific implementation, the structure of the initial preference probability regression model is as follows: Figure 7 As shown, compared to the traditional training method that directly learns subjective score labels, the regression network in this embodiment transforms absolute scores into relative comparisons through ranking comparisons between two Siamese sub-networks, thus effectively avoiding noise in absolute subjective scores during training. In the testing phase, a single sub-network from the Siamese network can be used to regress the input sample features, and the network can predict sample variance in addition to outputting a predicted comprehensive quality score. Based on the initial preference probability regression model, the loss function uses a fidelity loss based on quantum physics as a similarity measure of preference probabilities. In addition, this embodiment can also use different loss calculation methods, considering the binary cross entropy (BCE) loss commonly used in machine learning, and on this basis, to avoid the scale ambiguity problem caused by the joint estimation of mean and variance, a regularization term based on folding loss is introduced.

[0068] In this embodiment, a sample training video set is obtained; multiple sample training video pairs are constructed based on the sample training video set; features are extracted from each sample training video pair according to a preset perceptual classification feature set to obtain the first training feature of the first sample video and the second training feature of the second sample video in each sample training video pair; the initial preference probability regression model is trained based on the first training feature and the second training feature to obtain the target preference probability regression model.

[0069] In addition, refer to Figure 8 This invention also proposes a video quality detection device, which includes: The acquisition module 10 is used to acquire the video to be detected.

[0070] The extraction module 20 is used to extract features from the video to be detected based on a preset perceptual classification feature set, thereby obtaining multiple perceptual dimension features of multiple perceptual distortion dimensions.

[0071] The determination module 30 is used to determine the comprehensive quality score of the video to be detected based on the multiple perceptual dimension features of each perceptual distortion dimension and the target preference probability regression model.

[0072] Evaluation module 40 is used to evaluate video scores based on multiple perceptual dimension features of each perceptual distortion dimension and output the dimension quality score of each perceptual distortion dimension.

[0073] The completion module 50 is used to perform video quality detection on the video to be detected based on the comprehensive quality score and the dimensional quality scores of each perceived distortion dimension.

[0074] This embodiment acquires a video to be tested; extracts features from the video based on a preset perceptual classification feature set to obtain multiple perceptual distortion dimensions; determines the overall quality score of the video based on the multiple perceptual distortion dimensions and a target preference probability regression model; evaluates the video score based on the multiple perceptual distortion dimensions and outputs the dimensional quality score for each perceptual distortion dimension; and completes the video quality detection of the video based on the overall quality score and the dimensional quality scores for each perceptual distortion dimension. The above method extracts features from the video to be tested based on a preset perceptual classification feature set, obtaining multiple perceptual dimension features under multiple perceptual distortion dimensions. A comprehensive quality score is determined based on these features and a target preference probability regression model. Video scores are then evaluated based on these features, outputting dimensional quality scores for each perceptual distortion dimension. Finally, video quality detection is completed based on the comprehensive quality score and the dimensional quality scores for each perceptual distortion dimension. This method achieves comprehensive and effective video quality detection, providing not only a comprehensive evaluation result but also individual evaluation results for each perceptual distortion dimension. It fully reflects the video quality, effectively displays video distortion, and improves the efficiency and accuracy of video quality detection.

[0075] In one embodiment, the determining module 30 is further configured to acquire a sample training video set; Construct multiple sample training video pairs based on the sample training video set; Based on the preset perceptual classification feature set, feature extraction is performed on each sample training video pair to obtain the first training feature of the first sample video and the second training feature of the second sample video in each sample training video pair. The initial preference probability regression model is trained based on the first training feature and the second training feature to obtain the target preference probability regression model.

[0076] In one embodiment, the determining module 30 is further configured to input the first training features into the first sub-network of the Siamese structure network to obtain a first prediction mean and a first prediction variance; The second training feature is input into the second sub-network of the Siamese structure network to obtain the second prediction mean and the second prediction variance. Input the first prediction mean, the first prediction variance, the second prediction mean, and the second prediction variance into the preference probability model to determine the prediction preference probability of each training video pair. The initial preference probability regression model is adjusted based on the predicted preference probability of each training video pair to obtain the target preference probability regression model.

[0077] In one embodiment, the evaluation module 40 is further configured to calculate the mean of multiple perceptual dimension features of each perceptual distortion dimension to obtain a first feature mean. The difference between each perceptual dimension feature and the mean of the first feature is calculated to obtain the first decentralized data; Perform matrix calculations on the first decentralized data to determine multiple first eigenvalues ​​and multiple first eigenvectors; The video score is evaluated based on each first feature value and each first feature vector, and the dimensional quality score of each perceptual distortion dimension is output.

[0078] In one embodiment, the evaluation module 40 is further configured to sort each first feature value and determine a plurality of second feature vectors and a plurality of second feature values ​​in each first feature vector according to the sorting result; The mean of the second feature is obtained by calculating the mean of multiple second feature vectors based on multiple second feature values; The difference between each second feature vector and the mean of the second feature is calculated to obtain the second decentralized data; Matrix calculations are performed based on the second decentralized data to determine multiple third eigenvalues ​​and third eigenvectors; Sort the third feature values ​​and determine the target feature vector from the third feature vectors based on the sorting results; Based on the target feature vector, a video score is evaluated, and the dimensional quality score for each perceptual distortion dimension is output.

[0079] In one embodiment, the evaluation module 40 is further configured to construct a space based on the target feature vector to obtain an initial space; Based on the initial space, feature transformation is performed on the features of each perceptual dimension to obtain the target space; The dimensional quality score of each perceptual distortion dimension is output based on the target space.

[0080] In one embodiment, the completion module 50 is further configured to obtain the dimension score threshold of each perceptual distortion dimension; The dimensional quality score of each perceptual distortion dimension is compared with the dimensional score threshold of each perceptual distortion dimension to obtain the first comparison result; The dimensions to be improved in the video to be detected are determined based on the first comparison result; A quality improvement strategy is generated based on the dimensions to be improved.

[0081] Since this device adopts all the technical solutions of all the above embodiments, it has at least all the beneficial effects brought about by the technical solutions of the above embodiments, which will not be repeated here.

[0082] Furthermore, this embodiment of the invention also proposes a storage medium storing a video quality detection program, which, when executed by a processor, implements the steps of the video quality detection method described above.

[0083] Since this storage medium adopts all the technical solutions of all the above embodiments, it has at least all the beneficial effects brought about by the technical solutions of the above embodiments, which will not be repeated here.

[0084] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this invention. In practical applications, those skilled in the art can select some or all of the workflow to achieve the purpose of this embodiment according to actual needs, and no restrictions are imposed here.

[0085] In addition, for technical details not described in detail in this embodiment, please refer to the video quality detection method provided in any embodiment of the present invention, which will not be repeated here.

[0086] Furthermore, it should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0087] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0088] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0089] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A method of video quality detection, the method comprising: The video quality detection method comprises: obtaining a to-be-detected video; extracting features of the to-be-detected video according to a preset perception classification feature set to obtain multiple perception dimension features of multiple perception distortion dimensions; determining a comprehensive quality score of the to-be-detected video according to the multiple perception dimension features of the multiple perception distortion dimensions and a target preference probability regression model; Before the determining of the comprehensive quality score of the to-be-detected video according to the multiple perception dimension features of the multiple perception distortion dimensions and the target preference probability regression model, the method further comprises: obtaining a sample training video set; constructing multiple sample training video pairs according to the sample training video set; extracting features of each sample training video pair according to the preset perception classification feature set to obtain first training features of a first sample video and second training features of a second sample video in each sample training video pair; training an initial preference probability regression model according to the first training features and the second training features to obtain the target preference probability regression model; The training of the initial preference probability regression model according to the first training features and the second training features to obtain the targeted preference probability regression model comprises: inputting the first training features into a first sub-network of a Siamese structure network to obtain a first predicted mean and a first predicted variance; inputting the second training features into a second sub-network of the Siamese structure network to obtain a second predicted mean and a second predicted variance; inputting the first predicted mean, the first predicted variance, the second predicted mean and the second predicted variance into a preference probability model to determine a predicted preference probability of each sample training video pair; adjusting the initial preference probability regression model according to the predicted preference probability of each sample training video pair to obtain the target preference probability regression model; The preference probability model in the target preference probability regression model is based on Thurstone's law, which models the preference probability P(i>j) of video sample i being better than video sample j as follows: In the formula, Φ is the cumulative distribution function of the standard normal distribution, and μ i and μ j σ represents the predicted mean of the output of the Siamese structural network for video samples i and j, respectively. i ² and σ j ² represents the prediction variance of the corresponding output; performing video score evaluation according to the multiple perception dimension features of the multiple perception distortion dimensions to output a dimension quality score of each perception distortion dimension; completing video quality detection of the to-be-detected video according to the comprehensive quality score and the dimension quality score of each perception distortion dimension.

2. The video quality detection method of claim 1, wherein, The video score evaluation according to the multiple perception dimension features of the multiple perception distortion dimensions to output the dimension quality score of each perception distortion dimension comprises: performing mean calculation on the multiple perception dimension features of each perception distortion dimension to obtain a first feature mean; statistically calculating a difference between each perception dimension feature and the first feature mean to obtain first decentering data; performing matrix calculation on the first decentering data to determine multiple first feature values and multiple first feature vectors; performing video score evaluation according to each first feature value and each first feature vector to output the dimension quality score of each perception distortion dimension.

3. The video quality detection method of claim 2, wherein, The video score evaluation according to each first feature value and each first feature vector to output the dimension quality scores of each perception distortion dimension comprises: sorting each first feature value and determining multiple second feature vectors and multiple second feature values in each first feature vector according to a sorting result; performing mean calculation on the multiple second feature vectors according to the multiple second feature values to obtain a second feature mean; statistically determine a difference between each second feature vector and the second feature mean value to obtain second de-centralized data; perform matrix calculation according to the second de-centralized data to determine a plurality of third feature values and third feature vectors; sort each third feature value and determine a target feature vector from the third feature vectors according to a sorting result; perform video score evaluation according to the target feature vector to output a dimension quality score of each perceptual distortion dimension.

4. The video quality detection method of claim 3, wherein, The video score evaluation according to the target feature vector to output a dimension quality score of each perceptual distortion dimension comprises: perform spatial construction according to the target feature vector to obtain an initial space; perform feature transformation on each perceptual dimension feature according to the initial space to obtain a target space; output a dimension quality score of each perceptual distortion dimension according to the target space.

5. The video quality detection method of any of claims 1 to 4, wherein, After the video quality detection of the to-be-detected video is completed according to the comprehensive quality score and the dimension quality scores of each perceptual distortion dimension, the method further comprises: obtain a dimension score threshold of each perceptual distortion dimension; compare the dimension quality score of each perceptual distortion dimension with the dimension score threshold of each perceptual distortion dimension respectively to obtain a first comparison result; determine a to-be-improved dimension of the to-be-detected video according to the first comparison result; generate a quality improvement strategy according to the to-be-improved dimension.

6. A video quality detection apparatus, characterized by comprising: The video quality detection device comprises: an obtaining module configured to obtain a to-be-detected video; an extracting module configured to perform feature extraction on the to-be-detected video according to a preset perceptual classification feature set to obtain a plurality of perceptual dimension features of a plurality of perceptual distortion dimensions; a determining module configured to determine a comprehensive quality score of the to-be-detected video according to the plurality of perceptual dimension features of each perceptual distortion dimension and a target preference probability regression model; the determining module is further configured to obtain a sample training video set; construct a plurality of sample training video pairs according to the sample training video set; perform feature extraction on each sample training video pair according to the preset perceptual classification feature set to obtain a first training feature of a first sample video and a second training feature of a second sample video in each sample training video pair; and perform model training on an initial preference probability regression model according to the first training feature and the second training feature to obtain the target preference probability regression model; The determining module is further configured to input the first training feature into a first sub-network of a Siamese network to obtain a first predicted mean and a first predicted variance; input the second training feature into a second sub-network of the Siamese network to obtain a second predicted mean and a second predicted variance; input the first predicted mean, the first predicted variance, the second predicted mean and the second predicted variance into a preference probability model to determine a predicted preference probability of each sample training video pair; and adjust an initial preference probability regression model according to the predicted preference probability of each sample training video pair to obtain a target preference probability regression model; wherein the preference probability model in the target preference probability regression model is a model based on Thurstone's law, which models a preference probability P(i>j) that a video sample i is preferred to a video sample j as follows: , wherein Φ is a cumulative distribution function of a standard normal distribution, μ i and μ j are the predicted means output by the Siamese network for the video samples i and j respectively, σ i 2 and σ j 2 are the predicted variances corresponding to the outputs respectively. an evaluating module configured to perform video score evaluation according to the plurality of perceptual dimension features of each perceptual distortion dimension to output a dimension quality score of each perceptual distortion dimension; a completing module configured to complete video quality detection of the to-be-detected video according to the comprehensive quality score and the dimension quality scores of each perceptual distortion dimension.

7. A video quality detection device, characterized by, The device comprises a memory, a processor, and a video quality detection program stored on the memory and executable on the processor, and the video quality detection program is configured to implement the video quality detection method of any one of claims 1 to 5.

8. A storage medium, characterized by The storage medium stores a video quality detection program, and the video quality detection program is executed by the processor to implement the video quality detection method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Video quality evaluation method based on visual saliency area and time-space characteristics

    CN107318014A

  • Transmission spectrum non-destructive quantitative evaluation method of apple watercore

    CN109100323A

  • Quality evaluation method and device and electronic equipment

    CN112950581A