Video picture quality evaluation method, device, equipment and storage medium
By extracting key frames from videos and using a pre-set quality detection model to detect key feature information, the problem of low accuracy in video image quality detection in existing technologies is solved, achieving higher detection accuracy and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING 360 INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2020-12-23
- Publication Date
- 2026-04-17
AI Technical Summary
Current technologies rely on visual inspection for video image quality detection, resulting in low accuracy and a poor user experience.
By extracting key frames from the video to be detected, obtaining key feature information, and using a preset quality detection model for detection, the video detection results are obtained, and finally the video quality is evaluated.
It improves the accuracy of video image quality detection and enhances the user experience.
Smart Images

Figure CN114742741B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video processing technology, and in particular to a method, apparatus, device, and storage medium for evaluating video image quality. Background Technology
[0002] With urbanization, high-definition video footage and stable video streams are particularly important for daily surveillance. To ensure high-definition video footage and stable video streams, quality inspection of video footage (or video images) has become crucial. Current technologies for video footage quality inspection primarily rely on visual inspection, resulting in low accuracy and consequently, a reduced user experience.
[0003] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of this invention is to provide a method, apparatus, device, and storage medium for evaluating video image quality, aiming to solve the technical problem of how to improve the accuracy of video image quality evaluation.
[0005] To achieve the above objectives, the present invention provides a video image quality evaluation method, the video image quality evaluation method comprising:
[0006] Extract key frames to be processed from the video to be detected, and obtain key feature information based on the key frames to be processed;
[0007] The key feature information is detected according to the preset quality detection model to obtain video detection results;
[0008] The video quality of the video to be tested is evaluated based on the video detection results.
[0009] Optionally, before the step of extracting keyframes to be processed from the video to be detected and obtaining key feature information based on the keyframes to be processed, the method further includes:
[0010] Acquire multiple sample videos, and obtain multiple sample images based on the multiple sample videos;
[0011] Multiple sample images are processed separately to obtain sample image feature information and sample video detection results;
[0012] A sample image feature information set is constructed based on the obtained sample image feature information, and a sample video detection result set is constructed based on the obtained sample video detection results;
[0013] The initial residual network model is trained based on the sample image feature information set and the sample video detection result set to obtain the preset quality detection model.
[0014] Optionally, the step of extracting keyframes to be processed from the video to be detected includes:
[0015] The video to be tested is split into frames to obtain multiple images to be tested;
[0016] Extract the image to be processed from multiple images to be detected, and use the image to be processed as the key frame to be processed.
[0017] Optionally, the step of obtaining key feature information based on the keyframe to be processed includes:
[0018] Obtain the pixel information corresponding to the keyframe to be processed;
[0019] The region pixel set is determined based on the pixel information, and the region keyframe is selected from the keyframes to be processed based on the region pixel set.
[0020] Key feature information is obtained from the keyframes of the region.
[0021] Optionally, the step of determining the region pixel set based on the pixel information includes:
[0022] Determine the vertex pixel information based on the pixel information;
[0023] The region pixel set is determined based on the vertex pixel information and the pixel information.
[0024] Optionally, the step of obtaining key feature information based on the keyframes of the region includes:
[0025] Channel processing is performed on the keyframes of the region to obtain a luminance channel image;
[0026] Image texture information is obtained from the brightness channel image, and key feature information is determined based on the image texture information.
[0027] Optionally, the step of obtaining image texture information based on the luminance channel image includes:
[0028] Extract cross-channel invariant information from the brightness channel image;
[0029] The brightness channel image is denoised based on the cross-channel invariant information to obtain a denoised channel image.
[0030] Image texture information is obtained from the denoised channel image.
[0031] Optionally, the video detection result includes image blur, multiple image types, and probability values corresponding to the image types;
[0032] The step of evaluating the video quality of the video to be detected based on the video detection results includes:
[0033] The video quality of the video to be detected is evaluated based on the image blur, multiple image types, and the probability values corresponding to the image types.
[0034] Optionally, the step of evaluating the video quality of the video to be detected based on the image blur, multiple image types, and the probability value corresponding to the image type includes:
[0035] The image blur level is determined based on the blur degree.
[0036] The video quality of the video to be detected is evaluated based on the image blur level, multiple image types, and the probability value corresponding to the image type.
[0037] Optionally, the step of determining the image blur level based on the blur degree includes:
[0038] Based on the blur degree, the corresponding sample image blur level is found from the preset blur level mapping table, and the sample image blur level is used as the image blur level of the video to be detected.
[0039] The preset blur level mapping table includes the correspondence between blur level and blur level of sample image.
[0040] Optionally, the step of evaluating the video quality of the video to be detected based on the image blur level, multiple image types, and the probability value corresponding to the image type includes:
[0041] The target image type is selected from multiple image types based on the probability value;
[0042] The video quality of the video to be detected is evaluated based on the image blur level and the target image type.
[0043] Optionally, the step of selecting the target image type from multiple image types based on the probability value includes:
[0044] Based on the probability values, multiple image types are prioritized and sorted from high to low to obtain the image type sorting result;
[0045] The target image type is selected from multiple image types based on the sorting results of the image types.
[0046] Furthermore, to achieve the above objectives, the present invention also proposes a video image quality evaluation device, the video image quality evaluation device comprising:
[0047] The acquisition module is used to extract key frames to be processed from the video to be detected, and to obtain key feature information based on the key frames to be processed.
[0048] The detection module is used to detect the key feature information according to the preset quality detection model to obtain video detection results;
[0049] The evaluation module is used to evaluate the video quality of the video to be detected based on the video detection results.
[0050] Optionally, the video image quality evaluation device further includes an establishment module;
[0051] The establishment module is used to acquire multiple sample videos and obtain multiple sample images based on the multiple sample videos;
[0052] The establishment module is also used to process multiple sample images respectively to obtain sample image feature information and sample video detection results;
[0053] The establishment module is also used to construct a sample image feature information set based on the obtained sample image feature information, and to construct a sample video detection result set based on the obtained sample video detection results;
[0054] The establishment module is also used to train the initial residual network model based on the sample image feature information set and the sample video detection result set to obtain a preset quality detection model.
[0055] Optionally, the acquisition module is further configured to perform frame splitting on the video to be detected to obtain multiple images to be detected;
[0056] The acquisition module is further configured to extract a processing image from multiple images to be detected, and use the processing image as a processing keyframe.
[0057] Optionally, the acquisition module is further configured to acquire pixel information corresponding to the keyframe to be processed;
[0058] The acquisition module is further configured to determine a set of region pixels based on the pixel information, and select a region keyframe from the keyframe to be processed based on the set of region pixels.
[0059] The acquisition module is also used to obtain key feature information based on the key frames of the region.
[0060] Optionally, the acquisition module is further configured to determine vertex pixel information based on the pixel information;
[0061] The acquisition module is further configured to determine a set of region pixels based on the vertex pixel information and the pixel information.
[0062] Optionally, the acquisition module is further configured to perform channel processing on the keyframes of the region to obtain a luminance channel image;
[0063] The acquisition module is further configured to obtain image texture information based on the brightness channel image, and determine key feature information based on the image texture information.
[0064] Furthermore, to achieve the above objectives, the present invention also proposes a video image quality evaluation device, the device comprising: a memory, a processor, and a video image quality evaluation program stored in the memory and executable on the processor, the video image quality evaluation program being configured to implement the steps of the video image quality evaluation method described above.
[0065] Furthermore, to achieve the above objectives, the present invention also proposes a storage medium storing a video image quality evaluation program, wherein the video image quality evaluation program, when executed by a processor, implements the steps of the video image quality evaluation method as described above.
[0066] This invention first extracts keyframes from the video to be tested and obtains key feature information based on these keyframes. Then, it detects the key feature information using a preset quality detection model to obtain video detection results. Finally, it evaluates the video quality of the video to be tested based on the video detection results. Compared to existing technologies, which do not evaluate the video quality before publishing, resulting in a poor user experience, this invention detects key feature information using a preset quality detection model to obtain video detection results, and then evaluates the video quality of the video to be tested based on these results. This improves the accuracy of video quality detection and thus enhances the user experience. Attached Figure Description
[0067] Figure 1 This is a schematic diagram of the structure of a video image quality evaluation device in the hardware operating environment involved in the embodiments of the present invention;
[0068] Figure 2 This is a flowchart illustrating the first embodiment of the video image quality evaluation method of the present invention;
[0069] Figure 3 This is a flowchart illustrating the second embodiment of the video image quality evaluation method of the present invention;
[0070] Figure 4 This is a flowchart illustrating the third embodiment of the video image quality evaluation method of the present invention;
[0071] Figure 5 This is a structural block diagram of the first embodiment of the video image quality evaluation device of the present invention.
[0072] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0073] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0074] Reference Figure 1 , Figure 1 This is a schematic diagram of the video image quality evaluation device structure in the hardware operating environment involved in the embodiments of the present invention.
[0075] like Figure 1 As shown, the video image quality evaluation device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.
[0076] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the video image quality evaluation device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0077] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a data storage module, a network communication module, a user interface module, and a video image quality evaluation program.
[0078] exist Figure 1In the video quality evaluation device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the video quality evaluation device of the present invention can be set in the video quality evaluation device, and the video quality evaluation device calls the video quality evaluation program stored in the memory 1005 through the processor 1001 and executes the video quality evaluation method provided in the embodiment of the present invention.
[0079] This invention provides a method for evaluating video image quality, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the video image quality evaluation method of the present invention.
[0080] In this embodiment, the video image quality evaluation method includes the following steps:
[0081] Step S10: Extract key frames to be processed from the video to be detected, and obtain key feature information based on the key frames to be processed.
[0082] It is easy to understand that the execution subject of this embodiment can be a video image quality evaluation device with functions such as image processing, data processing, network communication and program execution, or other computer devices with similar functions. This embodiment does not limit it.
[0083] It is understood that the video to be detected can be a live video uploaded by the user, or an entertainment video, etc., and this embodiment does not impose any restrictions.
[0084] The keyframes to be processed are single or multiple images selected by the user from the video to be detected, and this embodiment does not impose any restrictions.
[0085] It should be noted that users can select any image with a black screen or distorted screen from the video to be tested, and then use that image as a keyframe to be processed.
[0086] The steps for extracting keyframes to be processed from the video to be detected can also include: splitting the video to be detected into frames to obtain multiple images to be detected; extracting images to be processed from the multiple images to be detected using an image image detector; the images to be processed being images with black screens or distorted screens; and then using the images to be processed as keyframes to be processed.
[0087] The steps for obtaining key feature information from the keyframe to be processed can be as follows: obtain the pixel information corresponding to the keyframe to be processed, then determine the region pixel set based on the pixel information, select the region keyframe from the keyframe to be processed based on the region pixel set, and finally obtain the key feature information based on the region keyframe.
[0088] The pixel information consists of all pixels corresponding to the keyframe to be processed. Then, the abnormal region image is determined based on the keyframe to be processed. Vertex pixels in the abnormal region image are selected from all pixels in the abnormal region image. The region pixel set is determined based on the vertex pixels and all pixels. The region keyframe is selected from the keyframe to be processed based on the region pixel set. In other words, the region key image is extracted from the keyframe to be processed. This region key image can be the abnormal region image mentioned above. Finally, key feature information is obtained from the region key image.
[0089] It should be noted that key feature information can be image texture information and color information, etc., and this embodiment is not limited thereto.
[0090] Furthermore, in order to accurately obtain key feature information, the steps for obtaining key feature information based on regional keyframes can be as follows: perform channel processing on the regional keyframes to obtain a luminance channel image, then obtain image texture information based on the luminance channel image, and finally determine key feature information based on the image texture information.
[0091] Color information is obtained from the key image of the region. If the color information is black, the key frame to be processed is determined to be a black screen. If the color information is normal, channel processing is required for the key frame of the region. The luminance channel image is selected from multiple channel images. Then, the image texture information is obtained from the luminance channel image. The target texture image information is selected from the texture image information and used as the key feature information.
[0092] The processing method for selecting target texture image information from texture image information can be as follows: there are various messy texture information in the texture image information, and it is necessary to sort out the messy texture information to obtain the target texture information, and finally use the target texture information as key feature information, etc.
[0093] The steps to obtain image texture from the luminance channel image can be as follows: extract cross-channel invariant information from the luminance channel image, perform denoising processing on the luminance channel image based on the cross-channel invariant information to obtain a denoised channel image, and then obtain image texture information based on the denoised channel image.
[0094] The cross-channel invariant information can be image information that remains unchanged at a fixed position in one of the three channels, which can be background information or portrait information, etc. This embodiment does not impose any limitations.
[0095] Assuming that the cross-channel invariant information in the luminance channel image is background information, the luminance channel image is then denoised based on the background information to obtain a background luminance channel image. Finally, background texture information is extracted from the background luminance channel image, and the background texture information is sorted to obtain target background texture information, which is then used as key feature information.
[0096] Assuming that the cross-channel invariant information in the luminance channel image is the image information, the luminance channel image is then denoised based on the image information to obtain the image luminance channel image. Finally, the image texture information is extracted from the image luminance channel image, and the image texture information is sorted to obtain the target image texture information, which is then used as key feature information.
[0097] Step S20: Detect the key feature information according to the preset quality detection model to obtain video detection results.
[0098] The steps for constructing the preset quality detection model are as follows: acquire multiple sample videos, obtain multiple sample images based on the sample videos, process the multiple sample images to obtain sample image feature information and sample video detection results, construct a sample image feature information set based on the obtained sample image feature information, construct a sample video detection result set based on the obtained sample video detection results, and finally train the initial residual network model based on the sample image feature information set and the sample video detection result set to obtain the preset quality detection model, etc.
[0099] In the video frame data, there are some beautified image sets. These beautified image sets cannot be well distinguished from the screen with distorted images. The existing dataset cannot meet the requirements for accurate video detection. Therefore, it is necessary to collect and label the data again, and then use the self-collected dataset to train the model.
[0100] The collected dataset can be divided into three categories: normal (0), distorted screen (1), and black screen (2). The normal dataset can contain 336 data entries, the distorted screen dataset can contain 200 data entries, and the black screen dataset can contain 206 data entries, etc.
[0101] The labeled data can integrate a large amount of video and image data from screenshots to improve the model's compatibility and transferability. The integrated video data contains some interference noise data, including some video components, text information from the video, and installation components.
[0102] During the construction of the preset quality detection model, the network can use a ReNet 18 structure. The input image information can be 3-channel image information with a feature pixel resolution of 360*640 pixels, and the assembly language label is the current image type. The output node is a softmax logistic regression probability classification, with three values, which are the probability values corresponding to the current three categories, etc. Then, the data is fed into the constructed model for training.
[0103] The loss function in model training can be the cross-entropy loss function, the optimization function can be the adaptive gradient algorithm AdaGrad, the model training accuracy is 96.4%, and the model is saved after training. This model is a preset quality detection model, etc.
[0104] Key feature information is input into a preset quality detection model to obtain video detection results, which include image blur, multiple image types and probability values corresponding to each image type.
[0105] The image blur can be 20% or 40%, the image type can be normal, black screen, or distorted screen, and the probability value corresponding to the image type can be 0.7 or 0.3, etc. This embodiment does not impose any limitations.
[0106] Assuming that key feature information is input into a preset quality detection model, it can output blur 70%, normal type 0, black screen type 0.3, and distorted screen type 0.7, etc., and use blur 70%, normal type 0, black screen type 0.3, and distorted screen type 0.7 as video detection results, etc.
[0107] Assuming that the key feature information is input into the preset quality detection model, it can also directly output an image type corresponding to the key frame to be processed, such as a black screen or a distorted screen.
[0108] Step S30: Evaluate the video quality of the video to be detected based on the video detection results.
[0109] The steps for evaluating the video quality of the video under test based on the video detection results can include evaluating the video quality of the video under test based on image blur, multiple image types, and the probability values corresponding to the image types.
[0110] The processing method for evaluating the video quality of the video to be tested based on image blur, multiple image types, and the probability values corresponding to each image type can be as follows: determine the image blur level based on the blur level, which can be low blur level or high blur level, etc., and then evaluate the video quality of the video to be tested based on the blur level, multiple image types, and the probability values corresponding to each image type.
[0111] The steps for determining the blur level of an image based on blur can include: finding the corresponding blur level of a sample image from a preset blur level mapping table based on blur, and using the blur level of the sample image as the blur level of the video to be detected.
[0112] The preset fuzziness level mapping table includes the correspondence between fuzziness and the fuzziness level of the sample image. The preset fuzziness level mapping table contains multiple fuzziness levels and sample image fuzziness levels.
[0113] Assuming the blur is 0-40%, the blur level of the sample image corresponding to the blur is low blur; assuming the blur is 41%-1, the blur level of the sample image corresponding to the blur is high blur, etc. This embodiment does not impose any limitations.
[0114] The steps for evaluating the video quality of the video to be tested based on the blur level, multiple image types, and the probability values corresponding to the image types can be as follows: select the target image type from multiple image types based on the probability values, and then evaluate the video quality of the video to be tested based on the image blur level and the target image type.
[0115] The steps for selecting a target image type from multiple image types based on probability values can be as follows: sort the multiple image types by probability values from high to low to obtain the sorting results, and finally select the target image type from the multiple image types based on the sorting results.
[0116] Assuming key feature information is input into a preset quality detection model, it can output blurriness 70%, normal type 0, black screen type 0.3, and distorted screen type 0.7, etc. Then, based on the blurriness 70%, the blur level is determined as high blur level. Multiple image types are prioritized according to probability values from high to low to obtain the image type ranking result. The image type ranking result is distorted screen type 0.7 - black screen type 0.3 - normal type 0. Based on the image type ranking result, the target image type is distorted screen type 0.7. Finally, based on the high blur level and distorted screen type 0.7, the video picture quality of the video to be detected is evaluated, and the evaluation result can be severe distorted screen, etc.
[0117] Assuming that key feature information is input into a preset quality detection model, it can output blurriness 0, normal type 0, black screen type 1, and distorted screen type 0, etc. Then, based on blurriness 0, the blur level is determined, which is low blur level. Multiple image types are prioritized according to probability values from high to low to obtain the image type sorting result, which is black screen type 1-distorted screen type 0-normal type 0. Based on the image type sorting result, the target image type is black screen type 1. Finally, based on the low blur level and black screen type 1, the video picture quality of the video to be detected is evaluated, and the evaluation result can be black screen, etc.
[0118] In this embodiment, keyframes are first extracted from the video to be detected, and key feature information is obtained based on these keyframes. Then, the key feature information is detected using a preset quality detection model to obtain video detection results. Finally, the video quality of the video to be detected is evaluated based on the video detection results. Compared to existing technologies, which do not evaluate the video quality before publishing, resulting in a poor user experience, this embodiment detects key feature information using a preset quality detection model to obtain video detection results, and then evaluates the video quality of the video to be detected based on these results. This improves the accuracy of video quality detection and thus enhances the user experience.
[0119] refer to Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the video image quality evaluation method of the present invention.
[0120] Based on the first embodiment described above, in this embodiment, step S10 further includes:
[0121] Step S101: Extract the key frame to be processed from the video to be detected, and obtain the pixel information corresponding to the key frame to be processed.
[0122] It is understood that the video to be detected can be a live video uploaded by the user, or an entertainment video, etc., and this embodiment does not impose any restrictions.
[0123] The keyframes to be processed are single or multiple images selected by the user from the video to be detected, and this embodiment does not impose any restrictions.
[0124] It should be noted that users can select any image with a black screen or distorted screen from the video to be tested, and then use that image as a keyframe to be processed.
[0125] The steps for extracting keyframes to be processed from the video to be detected can also include: splitting the video to be detected into frames to obtain multiple images to be detected; extracting images to be processed from the multiple images to be detected using an image image detector; the images to be processed being images with black screens or distorted screens; and then using the images to be processed as keyframes to be processed.
[0126] Pixel information can include all pixels corresponding to the keyframe to be processed.
[0127] Step S102: Determine the region pixel set based on the pixel information, and select the region key frame from the key frame to be processed based on the region pixel set.
[0128] The pixel information consists of all pixels corresponding to the keyframe to be processed. Then, the abnormal region image is determined based on the keyframe to be processed. Vertex pixels in the abnormal region image are selected from all pixels based on the abnormal region image. The region pixel set is determined based on the vertex pixels and all pixels. The region keyframe is selected from the keyframe to be processed based on the region pixel set. In other words, the region key image is extracted from the keyframe to be processed. This region key image can be the abnormal region image mentioned above.
[0129] Step S103: Obtain key feature information based on the key frames of the region.
[0130] Key feature information is obtained from key images of the region. The key feature information may be image texture information and color information, etc., and this embodiment does not limit it.
[0131] Furthermore, in order to accurately obtain key feature information, the steps for obtaining key feature information based on regional keyframes can be as follows: perform channel processing on the regional keyframes to obtain a luminance channel image, then obtain image texture information based on the luminance channel image, and finally determine key feature information based on the image texture information.
[0132] Color information is obtained from the key image of the region. If the color information is black, the key frame to be processed is determined to be a black screen. If the color information is normal, channel processing is required for the key frame of the region. The luminance channel image is selected from multiple channel images. Then, the image texture information is obtained from the luminance channel image. The target texture image information is selected from the texture image information and used as key feature information.
[0133] The processing method for selecting target texture image information from texture image information can be as follows: there are various messy texture information in the texture image information, and it is necessary to sort out the messy texture information to obtain the target texture information, and finally use the target texture information as key feature information, etc.
[0134] The steps to obtain image texture from the luminance channel image can be as follows: extract cross-channel invariant information from the luminance channel image, perform denoising processing on the luminance channel image based on the cross-channel invariant information to obtain a denoised channel image, and then obtain image texture information based on the denoised channel image.
[0135] The cross-channel invariant information can be image information that remains unchanged at a fixed position in one of the three channels, which can be background information or portrait information, etc. This embodiment does not impose any limitations.
[0136] Assuming that the cross-channel invariant information in the luminance channel image is background information, the luminance channel image is then denoised based on the background information to obtain a background luminance channel image. Finally, background texture information is extracted from the background luminance channel image, and the background texture information is sorted to obtain target background texture information, which is then used as key feature information.
[0137] Assuming that the cross-channel invariant information in the luminance channel image is the image information, the luminance channel image is then denoised based on the image information to obtain the image luminance channel image. Finally, the image texture information is extracted from the image luminance channel image, and the image texture information is sorted to obtain the target image texture information, which is then used as key feature information.
[0138] In this embodiment, firstly, keyframes to be processed are extracted from the video to be detected, and the pixel information corresponding to the keyframes is obtained. Then, a set of pixel points in a region is determined based on the pixel point information, and keyframes in that region are selected from the keyframes to be processed based on the set of pixel points in that region. Finally, key feature information is obtained based on the keyframes in that region, thereby enabling accurate acquisition of key feature information.
[0139] refer to Figure 4 , Figure 4 This is a flowchart illustrating the third embodiment of the video image quality evaluation method of the present invention.
[0140] Based on the first embodiment described above, in this embodiment, step S30 further includes:
[0141] Step S301: The video detection result includes image blur, multiple image types and probability values corresponding to the image types. The video quality of the video to be detected is evaluated based on the image blur, multiple image types and probability values corresponding to the image types.
[0142] The processing method for evaluating the video quality of the video to be tested based on image blur, multiple image types, and the probability values corresponding to each image type can be as follows: determine the image blur level based on the blur level, which can be low blur level or high blur level, etc., and then evaluate the video quality of the video to be tested based on the blur level, multiple image types, and the probability values corresponding to each image type.
[0143] The steps for determining the blur level of an image based on blur can be as follows: find the corresponding blur level of a sample image from a preset blur level mapping table based on blur, and use the blur level of the sample image as the blur level of the video to be detected.
[0144] The preset fuzziness level mapping table includes the correspondence between fuzziness and the fuzziness level of the sample image. The preset fuzziness level mapping table contains multiple fuzziness levels and sample image fuzziness levels.
[0145] Assuming the blur is 0-40%, the blur level of the sample image corresponding to the blur is low blur; assuming the blur is 41%-1, the blur level of the sample image corresponding to the blur is high blur, etc. This embodiment does not impose any limitations.
[0146] The steps for evaluating the video quality of the video to be tested based on the blur level, multiple image types, and the probability values corresponding to the image types can be as follows: select the target image type from multiple image types based on the probability values, and then evaluate the video quality of the video to be tested based on the image blur level and the target image type.
[0147] The steps for selecting a target image type from multiple image types based on probability values can be as follows: sort the multiple image types by probability values from high to low to obtain the sorting results, and finally select the target image type from the multiple image types based on the sorting results.
[0148] Assuming key feature information is input into a preset quality detection model, it can output blurriness 70%, normal type 0, black screen type 0.3, and distorted screen type 0.7, etc. Then, based on the blurriness 70%, the blur level is determined as high blur level. Multiple image types are prioritized according to probability values from high to low to obtain the image type ranking result. The image type ranking result is distorted screen type 0.7 - black screen type 0.3 - normal type 0. Based on the image type ranking result, the target image type is distorted screen type 0.7. Finally, based on the high blur level and distorted screen type 0.7, the video picture quality of the video to be detected is evaluated, and the evaluation result can be severe distorted screen, etc.
[0149] Assuming that key feature information is input into a preset quality detection model, it can output blurriness 0, normal type 0, black screen type 1, and distorted screen type 0, etc. Then, based on blurriness 0, the blur level is determined, which is low blur level. Multiple image types are prioritized according to probability values from high to low to obtain the image type sorting result, which is black screen type 1-distorted screen type 0-normal type 0. Based on the image type sorting result, the target image type is black screen type 1. Finally, based on the low blur level and black screen type 1, the video picture quality of the video to be detected is evaluated, and the evaluation result can be black screen, etc.
[0150] In this embodiment, the video quality of the video to be detected is evaluated based on image blur, multiple image types, and the probability values corresponding to the image types, thereby improving the accuracy of video quality detection.
[0151] Reference Figure 5 , Figure 5 This is a structural block diagram of the first embodiment of the video image quality evaluation device of the present invention.
[0152] like Figure 5 As shown, the video image quality evaluation device proposed in this embodiment of the invention includes:
[0153] The acquisition module 5001 is used to extract key frames to be processed from the video to be detected, and to obtain key feature information based on the key frames to be processed.
[0154] The detection module 5002 is used to detect the key feature information according to the preset quality detection model to obtain video detection results;
[0155] The evaluation module 5003 is used to evaluate the video quality of the video to be detected based on the video detection results.
[0156] In this embodiment, keyframes are first extracted from the video to be detected, and key feature information is obtained based on these keyframes. Then, the key feature information is detected using a preset quality detection model to obtain video detection results. Finally, the video quality of the video to be detected is evaluated based on the video detection results. Compared to existing technologies, which do not evaluate the video quality before publishing, resulting in a poor user experience, this embodiment detects key feature information using a preset quality detection model to obtain video detection results, and then evaluates the video quality of the video to be detected based on these results. This improves the accuracy of video quality detection and thus enhances the user experience.
[0157] Furthermore, the video image quality evaluation device also includes an establishment module;
[0158] The establishment module is used to acquire multiple sample videos and obtain multiple sample images based on the multiple sample videos;
[0159] The establishment module is also used to process multiple sample images respectively to obtain sample image feature information and sample video detection results;
[0160] The establishment module is also used to construct a sample image feature information set based on the obtained sample image feature information, and to construct a sample video detection result set based on the obtained sample video detection results;
[0161] The establishment module is also used to train the initial residual network model based on the sample image feature information set and the sample video detection result set to obtain a preset quality detection model.
[0162] Furthermore, the acquisition module 5001 is also used to perform frame splitting processing on the video to be detected to obtain multiple images to be detected;
[0163] The acquisition module 5001 is further configured to extract a processing image from multiple images to be detected, and use the processing image as a processing keyframe.
[0164] Furthermore, the acquisition module 5001 is also used to acquire pixel information corresponding to the key frame to be processed;
[0165] The acquisition module 5001 is further configured to determine a set of region pixels based on the pixel information, and select a region key frame from the key frame to be processed based on the set of region pixels.
[0166] The acquisition module 5001 is also used to obtain key feature information based on the key frame of the region.
[0167] Furthermore, the acquisition module 5001 is also used to determine vertex pixel information based on the pixel information;
[0168] The acquisition module 5001 is further configured to determine a set of region pixels based on the vertex pixel information and the pixel information.
[0169] Furthermore, the acquisition module 5001 is also used to perform channel processing on the key frames of the region to obtain a brightness channel image;
[0170] The acquisition module 5001 is further configured to obtain image texture information based on the brightness channel image, and determine key feature information based on the image texture information.
[0171] Furthermore, the acquisition module 5001 is also used to extract cross-channel invariant information from the brightness channel image;
[0172] The acquisition module 5001 is further configured to perform denoising processing on the brightness channel image based on the cross-channel invariant information to obtain a denoised channel image.
[0173] The acquisition module 5001 is also used to obtain image texture information based on the denoised channel image.
[0174] Furthermore, the video detection result includes image blur, multiple image types, and probability values corresponding to the image types;
[0175] The evaluation module 5003 is also used to evaluate the video quality of the video to be detected based on the image blur, multiple image types and the probability value corresponding to the image type.
[0176] Furthermore, the evaluation module 5003 is also used to determine the image blur level based on the blur degree;
[0177] The evaluation module 5003 is also used to evaluate the video quality of the video to be detected based on the image blur level, multiple image types and the probability value corresponding to the image type.
[0178] Furthermore, the evaluation module 5003 is also used to look up the corresponding sample image blur level from the preset blur level mapping table according to the blur degree, and use the sample image blur level as the image blur level of the video to be detected;
[0179] The preset blur level mapping table includes the correspondence between blur level and blur level of sample image.
[0180] Furthermore, the evaluation module 5003 is also used to select a target image type from multiple image types based on the probability value;
[0181] The evaluation module 5003 is also used to evaluate the video quality of the video to be detected based on the image blur level and the target image type.
[0182] Furthermore, the evaluation module 5003 is also used to prioritize and sort multiple image types from high to low according to the probability values to obtain image type sorting results;
[0183] The evaluation module 5003 is also used to select a target image type from multiple image types based on the image type sorting result.
[0184] Other embodiments or specific implementations of the video image quality evaluation device of the present invention can be referred to the above-described method embodiments, and will not be repeated here.
[0185] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0186] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0187] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory / random access memory, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0188] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A method for evaluating video image quality, characterized in that, The video image quality evaluation method includes: Extract key frames to be processed from the video to be detected, and obtain key feature information based on the key frames to be processed; The key feature information is detected according to a preset quality detection model to obtain video detection results; The video quality of the video to be detected is evaluated based on the video detection results. The key feature information includes image texture information and / or color information; The step of extracting key frames to be processed from the video to be detected and obtaining key feature information based on the key frames to be processed includes: The video to be tested is split into frames to obtain multiple images to be tested; Extract the key frames that show black screen or distorted screen from the image to be detected; The region pixel set is determined based on the pixel information of the key frame to be processed, and the region key frame is extracted from the key frame to be processed based on the region pixel set. Key feature information is obtained based on the keyframes of the region.
2. The method as described in claim 1, characterized in that, Before the step of extracting key frames to be processed from the video to be detected and obtaining key feature information based on the key frames to be processed, the method further includes: Acquire multiple sample videos, and obtain multiple sample images based on the multiple sample videos; Multiple sample images are processed separately to obtain sample image feature information and sample video detection results; A sample image feature information set is constructed based on the obtained sample image feature information, and a sample video detection result set is constructed based on the obtained sample video detection results; The initial residual network model is trained based on the sample image feature information set and the sample video detection result set to obtain the preset quality detection model.
3. The method as described in claim 1, characterized in that, The step of obtaining key feature information based on the keyframe to be processed includes: Obtain the pixel information corresponding to the keyframe to be processed; The region pixel set is determined based on the pixel information, and the region keyframe is selected from the keyframes to be processed based on the region pixel set. Key feature information is obtained from the keyframes of the region.
4. The method as described in claim 3, characterized in that, The step of determining the region pixel set based on the pixel information includes: Determine the vertex pixel information based on the pixel information; The region pixel set is determined based on the vertex pixel information and the pixel information.
5. The method as described in claim 3, characterized in that, The step of obtaining key feature information based on the keyframes of the region includes: Channel processing is performed on the keyframes of the region to obtain a luminance channel image; Image texture information is obtained from the brightness channel image, and key feature information is determined based on the image texture information.
6. The method as described in claim 5, characterized in that, The step of obtaining image texture information based on the brightness channel image includes: Extract cross-channel invariant information from the brightness channel image; The brightness channel image is denoised based on the cross-channel invariant information to obtain a denoised channel image. Image texture information is obtained from the denoised channel image.
7. The method according to any one of claims 1-6, characterized in that, The video detection results include image blur, multiple image types, and probability values corresponding to the image types; The step of evaluating the video quality of the video to be detected based on the video detection results includes: The video quality of the video to be detected is evaluated based on the image blur, multiple image types, and the probability values corresponding to the image types.
8. The method as described in claim 7, characterized in that, The step of evaluating the video quality of the video to be detected based on the image blur, multiple image types, and the probability value corresponding to the image type includes: The image blur level is determined based on the blur degree. The video quality of the video to be detected is evaluated based on the image blur level, multiple image types, and the probability value corresponding to the image type.
9. The method as described in claim 8, characterized in that, The step of determining the image blur level based on the blur degree includes: Based on the blur degree, the corresponding sample image blur level is found from the preset blur level mapping table, and the sample image blur level is used as the image blur level of the video to be detected. The preset blur level mapping table includes the correspondence between blur level and blur level of sample image.
10. The method as described in claim 8, characterized in that, The step of evaluating the video quality of the video to be detected based on the image blur level, multiple image types, and the probability value corresponding to the image type includes: The target image type is selected from multiple image types based on the probability value; The video quality of the video to be detected is evaluated based on the image blur level and the target image type.
11. The method as described in claim 10, characterized in that, The step of selecting the target image type from multiple image types based on the probability value includes: Based on the probability values, multiple image types are prioritized and sorted from high to low to obtain the image type sorting result; The target image type is selected from multiple image types based on the sorting results of the image types.
12. A video image quality evaluation device, characterized in that, The video image quality evaluation device includes: The acquisition module is used to extract key frames to be processed from the video to be detected, and to obtain key feature information based on the key frames to be processed. The detection module is used to detect the key feature information according to a preset quality detection model in order to obtain video detection results; The evaluation module is used to evaluate the video quality of the video to be detected based on the video detection results. The key feature information includes image texture information and / or color information; The step of extracting key frames to be processed from the video to be detected and obtaining key feature information based on the key frames to be processed includes: The video to be tested is split into frames to obtain multiple images to be tested; Extract the key frames that show black screen or distorted screen from the image to be detected; The region pixel set is determined based on the pixel information of the key frame to be processed, and the region key frame is extracted from the key frame to be processed based on the region pixel set. Key feature information is obtained based on the keyframes of the region.
13. The apparatus as claimed in claim 12, characterized in that, The video image quality evaluation device also includes an establishment module; The establishment module is used to acquire multiple sample videos and obtain multiple sample images based on the multiple sample videos; The establishment module is also used to process multiple sample images respectively to obtain sample image feature information and sample video detection results; The establishment module is also used to construct a sample image feature information set based on the obtained sample image feature information, and to construct a sample video detection result set based on the obtained sample video detection results; The establishment module is also used to train the initial residual network model based on the sample image feature information set and the sample video detection result set to obtain a preset quality detection model.
14. The apparatus as claimed in claim 13, characterized in that, The acquisition module is also used to acquire pixel information corresponding to the key frame to be processed; The acquisition module is further configured to determine a set of region pixels based on the pixel information, and select a region keyframe from the keyframe to be processed based on the set of region pixels. The acquisition module is also used to obtain key feature information based on the key frames of the region.
15. The apparatus as claimed in claim 14, characterized in that, The acquisition module is further configured to determine vertex pixel information based on the pixel information; The acquisition module is further configured to determine a set of region pixels based on the vertex pixel information and the pixel information.
16. The apparatus as claimed in claim 14, characterized in that, The acquisition module is also used to perform channel processing on the key frames of the region to obtain a brightness channel image; The acquisition module is further configured to obtain image texture information based on the brightness channel image, and determine key feature information based on the image texture information.
17. A video image quality evaluation device, characterized in that, The device includes: a memory, a processor, and a video quality evaluation program stored in the memory and executable on the processor, the video quality evaluation program being configured to implement the steps of the video quality evaluation method as described in any one of claims 1-11.
18. A storage medium, characterized in that, The storage medium stores a video image quality evaluation program, which, when executed by a processor, implements the steps of the video image quality evaluation method as described in any one of claims 1-11.
Citation Information
Patent Citations
Color image denoising method and system thereof
CN102156964A
Model training method, video category detection method and device, electronic device and computer readable medium
CN110119757A
Video quality evaluation method and device and model training method and device
CN110837842A