Model training method, video quality assessment method, device, equipment and medium

The model training method addresses the limitations of existing video quality assessment methods by using a 3D convolutional neural network and additional modules to enhance feature extraction and boundary detection, achieving improved accuracy and generalization in video quality evaluation.

JP7799816B2Active Publication Date: 2026-01-15ZTE CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024514541
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-09
Filing Date
2022-09-01
Publication Date
2026-01-15
Estimated Expiration
2042-09-01

AI Technical Summary

Technical Problem

Current video quality assessment methods, such as PSNR, SSIM, and VMAF, suffer from incomplete feature extraction and motion information loss, leading to inaccurate results, making them unsuitable for precise video quality evaluation.

Method used

A model training method that includes obtaining training video data, determining MOS values, and training an initial video quality assessment model using a 3D convolutional neural network and additional modules to extract features and detect boundaries, ensuring convergence through separate parameter and hyperparameter optimization with independent datasets.

Benefits of technology

The method enhances video quality assessment accuracy by fully extracting image features, accurately detecting boundaries, and improving the generalization ability of the model, resulting in more precise quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007799816000004
    Figure 0007799816000004
  • Figure 0007799816000005
    Figure 0007799816000005
  • Figure 0007799816000006
    Figure 0007799816000006
Patent Text Reader

Abstract

The present disclosure provides a model training method for video quality assessment, including: obtaining training video data including reference video data and distorted video data; determining an MOS value, which is an average opinion value, of each of the training video data; and training a preset initial video quality assessment model based on the training video data and its MOS value until a convergence condition is reached, to obtain a final video quality assessment model. The present disclosure further provides a video quality assessment method, an apparatus, an equipment and a medium.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This disclosure claims priority to patent application No. 202111055446.0 filed with the China Patent Office on September 9, 2021, and the entire contents of that Chinese patent application are incorporated by reference.

[0002] [Technical field] The present disclosure relates to the field of image processing technology, but is not limited thereto. [Background technology]

[0003] With the advent of the 5G era, video applications such as live streaming, short videos, and video calls are becoming increasingly popular. In this age of the Internet where everything relies on video, the ever-increasing data traffic poses serious challenges to the stability of video service systems. How to correctly assess video quality has become a major bottleneck restricting the development of various technologies. It can even be said that video quality assessment is the most fundamental and most important problem in the audio-video field, and needs to be solved urgently. Summary of the Invention

[0004] The present disclosure provides a model training method for video quality assessment, a video quality assessment method, a model training apparatus for video quality assessment, a video quality assessment apparatus, an electronic device, and a computer storage medium.

[0005] In a first aspect, the present disclosure provides a model training method for video quality assessment, including the steps of obtaining training video data comprising reference video data and distorted video data; determining an MOS value, which is an average opinion value, of each of the training video data; and training a preset initial video quality assessment model based on the training video data and its MOS value until a convergence condition is reached, to obtain a final video quality assessment model.

[0006] In another aspect, the present disclosure provides a video quality assessment method, comprising a step of processing video data to be assessed based on a final quality assessment model obtained by training using any one of the methods described herein, and obtaining a quality assessment score for the video data to be assessed.

[0007] In another aspect, the present disclosure provides a model training apparatus for video quality assessment, comprising: an acquisition module configured to acquire training video data comprising reference video data and distorted video data; a processing module configured to determine a MOS value, which is a mean opinion value, of each of the training video data; and a training module configured to train a preset initial video quality assessment model based on the training video data and its MOS value until a convergence condition is reached, and obtain a final video quality assessment model.

[0008] In another aspect, the present disclosure provides a video quality assessment device comprising an assessment module arranged to process video data to be assessed based on a final quality assessment model obtained by training using the model training method for video quality assessment, and to obtain a quality assessment score for the video data to be assessed.

[0009] In another aspect, the present disclosure provides an electronic device comprising one or more processors and a storage device having one or more programs stored therein, the one or more programs, when executed by the one or more processors, causing the one or more processors to implement any one of the model training methods for video quality assessment described herein.

[0010] In another aspect, the present disclosure provides an electronic device comprising one or more processors and a storage device having one or more programs stored therein, the one or more programs, when executed by the one or more processors, causing the one or more processors to implement any one of the video quality assessment methods described herein.

[0011] In another aspect, the present disclosure provides a computer storage medium having stored thereon a computer program that, when executed by a processor, implements any one of the model training methods for video quality assessment described herein.

[0012] In another aspect, the present disclosure provides a computer storage medium having stored thereon a computer program that, when executed by a processor, implements any one of the video quality assessment methods described herein. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 is a flowchart of a model training method for video quality assessment provided by the present disclosure. [Figure 2] FIG. 2 is a flowchart for training an initial video quality assessment model provided by this disclosure. [Figure 3] FIG. 3 is a schematic diagram of a 3D convolutional neural network provided by the present disclosure. [Figure 4] FIG. 4 is a flowchart of DenseNet (Densely connected convolutional networks) provided by the present disclosure. [Figure 5] FIG. 5 is a flowchart of the attention mechanism network provided by the present disclosure. [Figure 6] FIG. 6 is a flowchart of a hierarchical convolutional network provided by the present disclosure. [Figure 7] FIG. 7 is a flowchart of the initial video quality assessment model provided by this disclosure. [Figure 8a] FIG. 8a is a schematic diagram of the 3D-PVQA method provided by the present disclosure. [Figure 8b] FIG. 8b is a screenshot of the reference video data and the distorted video data provided by the present disclosure. [Figure 9] FIG. 9 is a flowchart for determining the MOS value, which is the mean opinion value of each training video data provided by the present disclosure. [Figure 10] FIG. 10 is a flowchart of the video quality assessment method provided by the present disclosure. [Figure 11] FIG. 11 is a module schematic diagram of a model training apparatus for video quality assessment provided by the present disclosure. [Figure 12] FIG. 12 is a module schematic diagram of a video quality assessment device provided by the present disclosure. [Figure 13] FIG. 13 is a schematic diagram of an electronic device provided by the present disclosure. [Figure 14] FIG. 14 is a schematic diagram of a computer storage medium provided by the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0014] Hereinafter, exemplary embodiments will be described in detail with reference to the drawings, but the exemplary embodiments may be embodied in different forms and should not be construed as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and that those skilled in the art will be able to fully appreciate the scope of the present disclosure.

[0015] As used in this disclosure, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0016] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the disclosure. As used in this disclosure, the singular forms "a," "an," and "the" are intended to include the plural unless the context clearly dictates otherwise. It will be further understood that the use of the terms "comprising" and / or "consisting of" herein refers to the presence of said features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof.

[0017] The embodiments described herein may be described with reference to plan views and / or cross-sectional views in accordance with idealized schematic diagrams of the present disclosure. Therefore, the exemplary illustrations may be changed depending on manufacturing techniques and / or tolerances. Therefore, the embodiments are not limited to those shown in the drawings, but include modifications of configurations formed based on manufacturing processes. Therefore, the regions illustrated in the drawings have exemplary attributes, and the shapes of the regions illustrated in the drawings exemplify the specific shapes of the regions of elements, but are not intended to be limiting.

[0018] All terms (including technical and scientific terms) used in this disclosure, unless otherwise limited, have the same meaning as commonly understood by those skilled in the art. It will be further understood that those terms defined in common usage dictionaries will have a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be construed as having an ideal or overly formal meaning unless expressly limited herein.

[0019] Current industry video quality assessment methods can be divided into two categories: subjective and objective. Subjective methods rely on observers to intuitively judge video quality. While accurate, they are relatively complex and their results are easily affected by various factors, making them unsuitable for direct industrial application. Therefore, objective methods based on artificial intelligence (AI) are generally used, which are easy to implement. However, currently, the schemes developed using these techniques, such as Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measurement (SSIM), and Video Multi-Method Assessment Fusion (VMAF), have limited final results. Therefore, more accurate video quality assessment remains a challenging issue that must be addressed urgently.

[0020] Currently, common video quality assessment schemes such as PSNR, SSIM, and VMAF still have problems such as incomplete feature extraction and unclear boundary distinction, resulting in poor final results. This disclosure proposes to preset an initial video quality assessment model to fully extract features and accurately detect boundaries within an image, obtain reference video data and distorted video data, and use the reference video data, distorted video data, and their MOS values ​​(MOS--Mean Opinion Score) to train the initial video quality assessment model and obtain a final video quality assessment model, thereby improving the accuracy of video quality assessment.

[0021] As shown in FIG. 1, the present disclosure provides a model training method for video quality assessment, which may include the following steps S11 to S13.

[0022] In step S11, training video data comprising reference video data and distorted video data is obtained.

[0023] In step S12, the MOS value of each training video data is determined. In step S13, the preset initial video quality assessment model is trained based on the training video data and its MOS value until a convergence condition is reached, and a final video quality assessment model is obtained.

[0024] Here, the reference video data can be regarded as standard video data, and the reference video data and distorted video data can be obtained from the open-source datasets LIVE, CSIQ, and IVP, and the self-developed dataset CENTER. The MOS value is a numerical value used to represent the quality of video data. While the video data in the open-source datasets LIVE, CSIQ, and IVP typically contain corresponding MOS values, the video data in the self-developed dataset CENTER does not. Therefore, it is necessary to determine the MOS value for each training video data. The MOS values ​​can be obtained directly from the training video data in the open-source datasets LIVE, CSIQ, and IVP to generate corresponding MOS values ​​for the training video data from the self-developed dataset CENTER. Of course, corresponding MOS values ​​can also be generated directly for all training video data. During the training of the initial video quality assessment model, when a convergence condition is reached, it can be determined that the model has already met the needs of video quality assessment. At this point, training is stopped, and a final video quality assessment model is obtained.

[0025] As can be seen from steps S11 to S13 above, the model training method for video quality assessment provided by the present disclosure presets an initial video quality assessment model for fully extracting image features and accurately detecting boundaries in images, obtains training video data including reference video data and distorted video data, and simultaneously trains the initial video quality assessment model using the reference video data and the distorted video data to obtain a final video quality assessment model, which can clearly distinguish between distorted video data and undistorted video data, i.e., reference video data, and ensure the independence and diversity of the video data for training the model. The final video quality assessment model obtained by training the initial video quality assessment model can fully extract image features and accurately detect boundaries in images, and can be directly used to assess the quality of the video data to be assessed, thereby improving the accuracy of video quality assessment.

[0026] Generally, when the network architecture is determined, the model has two parts that affect the final performance of the model: one part is the model parameters such as weights and biases, and the other part is the model hyperparameters such as the learning rate and the number of network layers. Optimizing the model parameters and hyperparameters using the same training data may lead to absolute overfitting of the model. Therefore, two independent datasets can be used to separately optimize the parameters and hyperparameters of the initial video quality assessment model.

[0027] Correspondingly, as shown in FIG. 2, in some embodiments, the step of training a preset initial video quality assessment model based on training video data and its MOS value until a convergence condition is reached (i.e., as described in step S13) may include the following steps S131 and S132:

[0028] In step S131, a training set and a validation set are determined based on the preset ratio and the training video data, and the intersection of the training set and the validation set is an empty set.

[0029] In step S132, the parameters of the initial video quality assessment model are adjusted based on the training set and the MOS value of each video data in the training set, and the hyperparameters of the initial video quality assessment model are adjusted based on the validation set and the MOS value of each video data in the validation set until a convergence condition is reached.

[0030] The present disclosure does not specifically limit the preset ratio. For example, the training data may be divided into a training set and a validation set at a ratio of 6:4. Of course, the preset ratio may also be other ratios, such as 8:2 or 5:5. To easily evaluate the generalization ability of the final video quality assessment model, the training video data may be determined as a training set, a validation set, and a test set. For example, the training video data may be divided into a training set, a validation set, and a test set at a ratio of 6:2:2, with the intersection between each of the training set, the validation set, and the test set being an empty set. After the division, the training set and the validation set are used to train an initial video quality assessment model to obtain a final video quality assessment model, and the test set is used to evaluate the generalization ability of the final video quality assessment model. Note that the larger the test set data, the longer it takes to evaluate the generalization ability of the final video quality assessment model using the test set. However, the more video data used to train the initial video quality assessment model, the higher the accuracy of the final video quality assessment model. In order to further improve the efficiency and accuracy of video quality assessment, the number of training video data can be appropriately increased, the ratio of the training set to the validation set in the training video data can be appropriately increased, and in the training video data, the training set, validation set, and test set can be divided in other ratios, such as 10:1:1.

[0031] As can be seen from the above steps S131 and S132, the model training method for video quality assessment provided by the present disclosure determines a training set and a validation set whose intersection is an empty set based on a preset ratio and training video data, adjusts the parameters of the initial video quality assessment model using the training set and the MOS value of each video data in the training set, and adjusts the hyperparameters of the initial video quality assessment model using the validation set and the MOS value of each video data in the validation set. When the convergence condition is reached, a final video quality assessment model that can sufficiently extract image features and detect boundaries within images with high accuracy is obtained, thereby improving the accuracy of video quality assessment.

[0032] Training a preset initial video quality assessment model based on training video data and its MOS value is a deep learning-based model training process, which corresponds to using the MOS value of the training video data as a reference and striving to make the model output result constantly approach the MOS value. If the difference between the model output result and the MOS value is small, it can be considered that the model meets the needs of video quality assessment.

[0033] Correspondingly, the convergence condition includes that the evaluation error rate of each video data in the training set and the validation set does not exceed a preset threshold, and the evaluation error rate is calculated by the following formula: E=(|S-Mos|) / Mos E is the evaluation error rate of the current video data, S is the evaluation score of the current video data output by the initial quality evaluation model after adjusting the parameters and hyperparameters, Mos is the MOS value of the current video data.

[0034] For any video data, the MOS value of the current video data has already been determined in advance. After inputting the current video data into the initial quality assessment model after adjusting the parameters and hyperparameters, the initial quality assessment model after adjusting the parameters and hyperparameters outputs an assessment score S for the current video data, and the assessment error rate E for the current video data can be calculated. If the error assessment rates of each video data in the training set and each video data in the validation set do not exceed the preset threshold, the difference between the assessment result output by the model and the MOS value is small, indicating that the model has already met the needs for video quality assessment. At this time, training can be stopped. Note that the present disclosure does not specifically limit the preset threshold, and the preset threshold may be, for example, 0.28, 0.26, 0.24, etc.

[0035] Currently, common video quality assessment schemes such as PSNR, SSIM, and VMAF still have the problem of motion information loss, resulting in poor final results. In the present disclosure, a preset initial video quality assessment model may include a 3D convolutional neural network for extracting motion information to improve the accuracy of video quality assessment. Correspondingly, in some embodiments, the initial video quality assessment model includes a 3D convolutional neural network for extracting motion information of image frames.

[0036] As shown in Figure 3, Figure 3 is a schematic diagram of a 3D convolutional neural network provided by the present disclosure. The 3D convolutional neural network can stack multiple consecutive image frames into a cube and then apply a 3D convolution kernel to the cube. In the 3D convolutional neural network architecture, each feature map in a convolutional layer (shown in the right half of Figure 3) is connected to multiple adjacent consecutive frames in the previous layer (shown in the left half of Figure 3), thereby capturing motion information between consecutive image frames.

[0037] Extracting motion information of image frames using only a 3D convolutional neural network cannot provide a complete assessment of video data. Accordingly, in some embodiments, the initial video quality assessment model may further include an attention model, a data fusion processing module, a global pooling module, and a fully connected layer, where the attention model, the data fusion processing module, the 3D convolutional neural network, the global pooling module, and the fully connected layer are cascaded in sequence.

[0038] In some embodiments, the attention model comprises a cascaded multi-input network, a 2D convolution module, a DenseNet (Densely connected convolutional networks), a downsampling processing module, a hierarchical convolution network, an upsampling processing module, and an attention mechanism network, wherein the DenseNet (Densely connected convolutional networks) comprises at least two cascaded dense convolution modules, and the dense convolution module comprises four cascaded densely connected convolution layers.

[0039] As shown in Figure 4, Figure 4 is a schematic diagram of DenseNet (Densely connected convolutional networks) provided by the present disclosure. DenseNet (Densely connected convolutional networks) includes at least two cascaded dense convolutional modules, each of which includes four cascaded densely connected convolutional layers. The input of each dense convolutional layer is a fusion of feature maps from all previous dense convolutional layers of the current dense convolutional module. After pooling in each layer of the encoder, the feature maps pass through one dense convolutional module. Each pass through a dense convolutional module performs a BN (Batch Normalization) operation, a ReLU (Rectified Linear Unit) active function operation, and a convolutional convolution operation.

[0040] In some embodiments, the attention mechanism network comprises a cascaded attention convolution module, a linear correction unit active module, a non-linear active module, and an attention upsampling processing module.

[0041]

number

[0042] In some embodiments, the hierarchical convolutional network comprises a first hierarchical network, a second hierarchical network, a third hierarchical network, and a fourth upsampling processing module, wherein the first hierarchical network comprises a cascaded first downsampling processing module and a first hierarchical convolution module, the second hierarchical network comprises a cascaded second downsampling processing module, a second hierarchical convolution module, and a second upsampling processing module, and the third hierarchical network comprises a cascaded global pooling module, a third hierarchical convolution module, and a third upsampling processing module, wherein the first hierarchical convolution module is also cascaded with the second downsampling processing module, the first hierarchical convolution module and the second upsampling processing module are cascaded with the fourth upsampling processing module, and the fourth upsampling processing module and the third upsampling processing module are also cascaded with the third hierarchical convolution module.

[0043] Here, the first downsampling processing module and the second downsampling processing module are all configured to perform downsampling processing on the data, and the second upsampling processing module, the third upsampling processing module and the fourth upsampling processing module are all used to perform upsampling processing on the data.

[0044]

number

[0045] In some embodiments, a first layer convolution module may perform a Conv 5*5 operation (i.e., a 5*5 convolution operation) on the data, a second layer convolution module may perform a Conv 3*3 operation (i.e., a 3*3 convolution operation) on the data, and a third layer convolution module may perform a Conv 1*1 operation (i.e., a 1*1 convolution operation) on the data. It will be appreciated that the same convolution module may be used to perform the Conv 5*5 operation, the Conv 3*3 operation, and the Conv 1*1 operation, respectively.

[0046] As shown in Figure 7, Figure 7 is a flowchart of the initial video quality assessment model provided by the present disclosure, where the initial video quality assessment model may include a multi-input network, a 2D convolution module, a DenseNet (Densely connected convolutional networks), a downsampling processing module, a hierarchical convolution network, an upsampling processing module, an attention mechanism network, a data fusion processing module, a 3D convolutional neural network, a global pooling module, and a fully connected layer.

[0047] The video quality assessment model provided by the present disclosure may be referred to as a 3D-PVQA (3 Dimensions Pyramid Video Quality Assessment) model and a 3D-PVQA method. In step S132, each video data in the training set and each video data in the validation set are divided into distorted video data and residual video data, which are input into the 3D-PVQA model, i.e., residual multi-input and distorted multi-input. The residual video data is obtained by processing the distorted video data and reference video data using residual frames. The multi-input network outputs the input data as two groups of data: the first group of data is initial input data, and the second group of data is data obtained by reducing the initial input data by 1x according to the size of the data frame.

[0048] Taking the distorted multi-input (Distorted-Multi-Input) in the lower half as an example, the multi-input network outputs two groups of data. The first group of data is processed through a 2D convolution module, then input to a DenseNet (Densely Connected Convolutional Networks) for processing, and then input to a downsampling processing module for processing. The second group of data is processed through a 2D convolution module, then concatenated with the output of the downsampling processing module, and then input back to the DenseNet (Densely Connected Convolutional Networks) for processing. At this time, part of the output of the DenseNet (Densely Connected Convolutional Networks) is input back to the downsampling processing module for processing, and the output of the downsampling processing module is input to a hierarchical convolutional network for processing. The output of the hierarchical convolutional network, along with the rest of the DenseNet (Densely Connected Convolutional Networks), is input to an attention mechanism network for processing. The data fusion processing module performs data fusion processing on the output result of the residual video data obtained by the attention mechanism network processing and the output result of the distorted video data, and the output of the data fusion processing module is input into two 3D convolutional neural networks, and the 3D convolutional neural networks output a perceptibility threshold of the lost frame, and the perceptibility threshold of the lost frame and the residual data frame obtained from the residual frame are subjected to matrix multiplication processing, and finally input into the global pooling module and the fully connected layer for processing, and output a quality evaluation score of the video data.

[0049] The same modules can be used repeatedly, and although two first layer convolutional modules, two second layer convolutional modules, and three third layer convolutional modules are shown in Figure 6, it should be understood that this does not mean that the hierarchical convolutional network has two first layer convolutional modules, two second layer convolutional modules, and three third layer convolutional modules. The downsampling processing module and the downsampling processing module in the hierarchical convolutional network may be the same downsampling processing module or different downsampling processing modules, and the upsampling processing module, the upsampling processing module in the hierarchical convolutional network, and the attention upsampling processing module in the attention mechanism network may be the same upsampling processing module or different upsampling processing modules.

[0050] As shown in Figure 8a, the training video data can be divided into a training set, a validation set, and a test set based on a preset ratio. The training set can be input to the 3D-PVQA model for training, the validation set can be input to the 3D-PVQA model for validation, and the test set can be input to the 3D-PVQA model for testing, all of which will yield corresponding quality assessment scores. As previously shown, the test set can be used to evaluate the generalization ability of the final video quality assessment model. As shown in Figure 8b, the left side is a screenshot of the reference video data, and the right side is a screenshot of the distorted video data. Table 1 below shows the MOS values ​​of the video data and the corresponding quality assessment scores of the video data output by the 3D-PVQA model.

[0051] [Table 1]

[0052] As shown in FIG. 9, in some embodiments, the step of determining the MOS value, which is the mean opinion value, of each training video data (ie, step S12) may include the following steps S121 to S124.

[0053] In step S121, the training video data are divided into groups, each group having one reference video data and multiple distorted video data, and the resolution of each video data in each group is the same, and the frame rate of each video data in each group is the same.

[0054] In step S122, each distorted video data in each group is classified. In step S123, each distorted video data of each classification in each group is graded.

[0055] In step S124, the MOS value of each training video data is determined based on the grouping, classification and grading of each training video data.

[0056] Here, when classifying each distorted video data in each group, the distorted video data can be divided into different categories of distorted video data such as packet loss type distortion and coding type distortion, and when grading each distorted video data of each category in each group, the distorted video data can be divided into three different distortion levels: mild, medium, and severe.

[0057] After the training video data are grouped, classified, and graded, each group has one reference video data and multiple distorted video data, the multiple distorted video data belong to different categories, and the distorted video data in each category belong to different distortion levels. Based on the reference video data in each group, the MOS value of each training video data can be determined using the Subjective Assessment Method for Video Quality evaluation (SAMVIQ) method and the grouping, classification, and grading situations.

[0058] As shown in FIG. 10, the present disclosure further provides a video quality assessment method, which may include the following step S21.

[0059] In step S21, the video data to be evaluated is processed based on the final quality evaluation model obtained by training using the model training method for video quality evaluation, and a quality evaluation score for the video data to be evaluated is obtained.

[0060] By presetting an initial video quality assessment model for fully extracting image features and accurately detecting boundaries within an image, obtaining training video data including reference video data and distorted video data, and simultaneously training the initial video quality assessment model using the reference video data and distorted video data to obtain a final video quality assessment model, the independence and diversity of the video data used to train the model are ensured, and the final video quality assessment model can be directly used to assess the quality of the video data to be assessed, thereby improving the accuracy of video quality assessment.

[0061] Based on the same technical idea, the present disclosure further provides a model training apparatus for video quality assessment, as shown in FIG. 11 , which may include an acquisition module 101 , a processing module 102 and a training module 103 .

[0062] The acquisition module 101 is arranged to acquire training video data comprising reference video data and distorted video data.

[0063] The processing module 102 is arranged to determine a MOS value for each training video data.

[0064] The training module 103 is configured to train a preset initial video quality assessment model based on the training video data and its MOS value until a convergence condition is reached, and obtain a final video quality assessment model.

[0065] In some embodiments, the training module 103 is configured to determine a training set and a validation set based on a preset ratio and training video data, where the intersection of the training set and the validation set is an empty set, adjust parameters of an initial video quality assessment model based on the training set and the MOS value of each video data in the training set, and adjust hyperparameters of the initial video quality assessment model based on the validation set and the MOS value of each video data in the validation set until a convergence condition is reached.

[0066] In some embodiments, the convergence condition includes that neither the evaluation error rate of each video data in the training set nor the validation set exceeds a preset threshold, and the evaluation error rate is calculated by the following formula: E=(|S-Mos|) / Mos E is the evaluation error rate of the current video data, S is the evaluation score of the current video data output by the initial quality evaluation model after adjusting the parameters and hyperparameters, Mos is the MOS value of the current video data.

[0067] In some embodiments, the initial video quality assessment model comprises a 3D convolutional neural network for extracting motion information of image frames.

[0068] In some embodiments, the initial video quality assessment model further comprises an attention model, a data fusion processing module, a global pooling module, and a fully connected layer, wherein the attention model, the data fusion processing module, the 3D convolutional neural network, the global pooling module, and the fully connected layer are cascaded in sequence.

[0069] In some embodiments, the attention model comprises a cascaded multi-input network, a 2D convolution module, a DenseNet (Densely connected convolutional networks), a downsampling processing module, a hierarchical convolution network, an upsampling processing module, and an attention mechanism network, wherein the DenseNet (Densely connected convolutional networks) comprises at least two cascaded dense convolution modules, and the dense convolution module comprises four cascaded densely connected convolution layers.

[0070] In some embodiments, the attention mechanism network comprises a cascaded attention convolution module, a linear correction unit active module, a non-linear active module, and an attention upsampling processing module.

[0071] In some embodiments, the hierarchical convolutional network comprises a first hierarchical network, a second hierarchical network, a third hierarchical network, and a fourth upsampling processing module, wherein the first hierarchical network comprises a cascaded first downsampling processing module and a first hierarchical convolution module, the second hierarchical network comprises a cascaded second downsampling processing module, a second hierarchical convolution module, and a second upsampling processing module, and the third hierarchical network comprises a cascaded global pooling module, a third hierarchical convolution module, and a third upsampling processing module, wherein the first hierarchical convolution module is also cascaded with the second downsampling processing module, the first hierarchical convolution module and the second upsampling processing module are cascaded with the fourth upsampling processing module, and the fourth upsampling processing module and the third upsampling processing module are also cascaded with the third hierarchical convolution module.

[0072] In some embodiments, the processing module 102 is configured to group each training video data, each group comprising one reference video data and a plurality of distorted video data, each video data in each group having the same resolution and the same frame rate, classify each video data in each group, grade each video data of each classification in each group, and determine a MOS value for each training video data based on the grouping, classification, and grading of each training video data.

[0073] Based on the same technical idea, as shown in FIG. 12, the present disclosure further provides a video quality assessment device, which includes an assessment module 201 configured to process the video data to be assessed based on the final quality assessment model obtained by training using the model training method for video quality assessment, and obtain a quality assessment score for the video data to be assessed.

[0074] In addition, as shown in FIG. 13 , the disclosed embodiments further provide an electronic device, comprising one or more processors 301 and a storage device 302 storing one or more programs, and when the one or more programs are executed by the one or more processors 301, causing the one or more processors 301 to realize at least one of the model training method for video quality assessment according to each of the embodiments and the video quality assessment method according to each of the embodiments.

[0075] In addition, as shown in FIG. 14, an embodiment of the present disclosure further provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, realizes at least one of the model training method for video quality assessment according to each of the embodiments and the video quality assessment method according to each of the embodiments.

[0076] Those skilled in the art will understand that all or some of the steps in the methods and functional modules / units of the apparatuses disclosed above may be implemented as software, firmware, hardware, or a suitable combination thereof. In hardware embodiments, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components. For example, one physical component may have multiple functions, and one function or step may be performed by several physical components working together. Some or all of the physical components may be implemented as software executed by a processor such as a central processing unit, digital signal processor, or microprocessor, as hardware, or as an integrated circuit such as an application-specific integrated circuit. Such software may be distributed on computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic memory, or any other medium that can be used to store desired information and that can be accessed by a computer. Additionally, those skilled in the art will know that communication media typically include computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and can include any information delivery media.

[0077] Exemplary embodiments are disclosed herein, and although specific terms are employed, they are used and should be interpreted in a general and descriptive sense only, and not for purposes of limitation. It will be apparent to those skilled in the art that, in some instances, features, characteristics, and / or elements described in connection with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless expressly noted otherwise. Accordingly, those skilled in the art will recognize that various changes in form and detail may be made without departing from the scope of the present disclosure, as defined by the appended claims.

Claims

1. obtaining training video data comprising reference video data and distorted video data; determining a mean opinion score (MOS) for each of the training video data; training a preset initial video quality assessment model based on the training video data and its MOS value until a convergence condition is reached, to obtain a final video quality assessment model; training a preset initial video quality assessment model based on the training video data and its MOS value until a convergence condition is reached; determining a training set and a validation set based on a preset ratio and the training video data, wherein an intersection of the training set and the validation set is an empty set; adjusting parameters of the initial video quality assessment model based on the training set and the MOS value of each video data in the training set, and adjusting hyperparameters of the initial video quality assessment model based on the validation set and the MOS value of each video data in the validation set until a convergence condition is reached; The convergence condition includes that the evaluation error rate of each video data in the training set and the validation set does not exceed a preset threshold, and the evaluation error rate is calculated by the following formula: E = (|S-Mos|) / Mos E is the evaluation error rate of the current video data, S is the evaluation score of the current video data output by the initial video quality assessment model after adjusting the parameters and hyperparameters; Mos is the MOS value of the current video data Model training methods for video quality assessment.

2. The initial video quality assessment model comprises a 3D convolutional neural network for extracting motion information of image frames. The method of claim 1.

3. The initial video quality assessment model further includes an attention model, a data fusion processing module, a global pooling module, and a fully connected layer, and the attention model, the data fusion processing module, the 3D convolutional neural network, the global pooling module, and the fully connected layer are sequentially cascaded. The method of claim 2.

4. The attention model includes a cascaded multi-input network, a 2D convolution module, a DenseNet (Densely connected convolutional networks), a downsampling processing module, a hierarchical convolution network, an upsampling processing module, and an attention mechanism network. The DenseNet (Densely connected convolutional networks) includes at least two cascaded dense convolution modules, and the dense convolution module includes four cascaded densely connected convolution layers. The method of claim 3.

5. The attention mechanism network includes a cascaded attention convolution module, a linear correction unit active module, a nonlinear active module, and an attention upsampling processing module. The method of claim 4.

6. The hierarchical convolutional network comprises a first hierarchical network, a second hierarchical network, a third hierarchical network, and a fourth upsampling processing module, the first hierarchical network comprising a first downsampling processing module and a first hierarchical convolution module connected in cascade, the second hierarchical network comprising a second downsampling processing module, a second hierarchical convolution module, and a second upsampling processing module connected in cascade, the third hierarchical network comprising a global pooling module, a third hierarchical convolution module, and a third upsampling processing module connected in cascade, the first hierarchical convolution module is also cascaded with the second downsampling processing module, the first hierarchical convolution module and the second upsampling processing module are cascaded with the fourth upsampling processing module, and the fourth upsampling processing module and the third upsampling processing module are also cascaded with the third hierarchical convolution module. The method of claim 4.

7. The step of determining a MOS value, which is a mean opinion value of each of the training video data, includes: a step of dividing the training video data into groups, each group including one reference video data and a plurality of distorted video data, and each video data in each group has the same resolution and the same frame rate; classifying each video data in each group; Grading each video data of each classification in each group; determining a MOS value for each of the training video data based on the grouping, classification, and grading of the training video data. The method of claim 1.

8. and processing the video data to be evaluated based on the final quality assessment model obtained by training using the method according to any one of claims 1 to 7, to obtain a quality assessment score for the video data to be evaluated. Video quality assessment methods.

9. an acquisition module arranged to acquire training video data comprising reference video data and distorted video data; a processing module arranged to determine a mean opinion score (MOS) value for each of said training video data; a training module configured to train a preset initial video quality assessment model based on the training video data and its MOS value until a convergence condition is reached, to obtain a final video quality assessment model; training a preset initial video quality assessment model based on the training video data and its MOS value until a convergence condition is reached; determining a training set and a validation set based on a preset ratio and the training video data, wherein an intersection of the training set and the validation set is an empty set; adjusting parameters of the initial video quality assessment model based on the training set and the MOS value of each video data in the training set, and adjusting hyperparameters of the initial video quality assessment model based on the validation set and the MOS value of each video data in the validation set until a convergence condition is reached; The convergence condition includes that the evaluation error rate of each video data in the training set and the validation set does not exceed a preset threshold, and the evaluation error rate is calculated by the following formula: E = (|S-Mos|) / Mos E is the evaluation error rate of the current video data, S is the evaluation score of the current video data output by the initial video quality assessment model after adjusting the parameters and hyperparameters; Mos is the MOS value of the current video data Model training apparatus for video quality assessment.

10. and an evaluation module configured to process video data to be evaluated based on a final quality evaluation model obtained by training using the method for model training for video quality evaluation according to any one of claims 1 to 7, and to obtain a quality evaluation score for the video data to be evaluated. Video quality assessment device.

11. one or more processors; a storage device storing one or more programs; The one or more programs, when executed by the one or more processors, cause the one or more processors to implement the method for training a model for video quality assessment according to any one of claims 1 to 7. electronic equipment.

12. one or more processors; a storage device storing one or more programs; The one or more programs, when executed by the one or more processors, cause the one or more processors to implement the video quality assessment method of claim 8. electronic equipment.

13. A computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the method for training a model for video quality assessment according to any one of claims 1 to 7. A computer-readable storage medium.

14. A computer-readable storage medium having stored thereon a computer program, the program implementing the video quality assessment method of claim 8 when executed by a processor. A computer-readable storage medium.

Citation Information

Patent Citations

  • Video quality evaluation method and device, electronic equipment and storage medium

    CN110751649A

  • Image quality estimation apparatus, method, and program

    JP2007306105A

  • Image quality estimation device, image quality estimation method, and image quality estimation program

    JP2014130427A

  • Automated segmentation using full-layer convolutional networks

    JP2020510463A

  • Method of detecting at least one element of interest visible in input image by means of convolutional neural network

    JP2021089730A