A method for determining the quality of video images for video conferencing
By constructing a knowledge distillation teacher subnet and a student subnet, the problem of extracting deep features on small-scale rich data sets in the existing technology is solved, and the generalization ability of video image quality evaluation model is realized in video conferencing applications.
Patent Information
- Application Number
- CN202210126393.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-10
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-02-10
AI Technical Summary
The prior art is difficult to learn how to extract deep features highly correlated with video image quality scores from small-scale but very rich datasets, resulting in poor generalization ability in video conferencing applications.
Build a knowledge distillation teacher subnet and a student subnet, extract high-dimensional features from data sets with rich image content through knowledge distillation technology, and improve the generalization ability of the student subnet without increasing the computational complexity.
It realizes the generalization ability of the video image quality evaluation model without increasing the computational complexity, and can better adapt to video conference scenarios with rich image content.
Smart Images

Figure CN114785978B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video image quality evaluation, and particularly to a method for determining the quality of video images for video conferencing. Background Art
[0002] Since the outbreak of the COVID-19 pandemic, video conferencing, as a real-time video communication method, has become an important means for individuals to maintain close contact with society. Video conferencing can help us continue to work and study during the pandemic, improving work and learning efficiency during the pandemic. In the application of video conferencing, visual information needs to be compressed and transmitted before being received by the end user, inevitably introducing unpredictable distortions and causing losses in video image quality. In order for the end user to obtain a high-quality visual experience, it is necessary to evaluate the quality of video images and adjust the relevant parameters of the encoder and transmission channel according to the evaluation results. Since the ultimate receptor of video is usually the human eye, the subjective evaluation of video image quality by the human eye is considered the most accurate method for evaluating video image quality. Although the subjective image quality evaluation technology directly participated by humans is accurate and reliable, it is difficult to meet the real-time requirements of applications such as video conferencing due to its high time consumption. Therefore, there is an urgent need in the prior art for an objective image quality evaluation technology that can monitor and feedback the quality of video images in real time.
[0003] The video objective quality evaluation method refers to an objective evaluation method that automatically and quickly scores the video quality by designing a mathematical model. According to the degree of dependence on the reference video image, video objective quality evaluation is divided into three categories: full reference, partial reference, and no reference. Since it is difficult to obtain reference video images in most practical applications, the no-reference video image quality evaluation technology in video objective quality evaluation technology has been the most widely used. The no-reference video image quality evaluation technology aims to design an algorithm that can quickly and automatically predict the perceptual quality of video images without using any information of the reference video image, so as to simulate the perception of video image quality by the human eye. In digital video applications related to digital multimedia, the no-reference video objective image quality evaluation technology plays an important role in quality detection on the server side and terminal quality experience, that is, according to the video image evaluation, feedback the quality information of the video image, dynamically adjust the video encoder parameters and transmission channel parameters on the server side, improve the perceptual quality of the video image at the receiving end, and give the end user a high-quality visual experience.
[0004] In the prior art, deep learning has been widely applied in the field of no-reference video image quality assessment, making it possible to jointly optimize feature extraction and quality regression. However, the methods in the prior art still have deficiencies. It is difficult to learn from a small-scale but very image-content-rich training set how to extract deep features highly correlated with quality scores, and it is difficult to generalize well to video conferencing applications with very rich image content. Summary of the Invention
[0005] In view of the above problems, the present invention is proposed to provide a method for determining the quality of video images for video conferencing that overcomes the above problems or at least partially solves the above problems.
[0006] According to one aspect of the present invention, there is provided a method for determining the quality of video images for video conferencing, the determination method comprising:
[0007] Construct a knowledge distillation teacher sub-network;
[0008] Construct a knowledge distillation student sub-network;
[0009] Obtain an image quality assessment data set with rich image content;
[0010] Construct a training set and a test set according to the image quality assessment data set, and the training set further includes corresponding quality score labels;
[0011] Perform data preprocessing on the training set and the test set to obtain a preprocessed data set;
[0012] Generate image blocks of video frames to be evaluated according to the preprocessed data set;
[0013] Use the trained student sub-network to predict the quality assessment scores of multiple image blocks of the video frames to be evaluated;
[0014] Calculate the average of multiple quality assessment scores to obtain the quality assessment score of the video to be evaluated.
[0015] Optionally, the construction of the knowledge distillation teacher sub-network specifically includes:
[0016] Build a 7-layer knowledge distillation teacher sub-network, and the structure is successively: the first convolution calculation unit, the second convolution calculation unit, the third convolution calculation unit, the fourth convolution calculation unit, the fifth convolution calculation unit, the first fully connected layer, the second fully connected layer; the second to fifth convolution calculation units adopt a bottleneck structure, and each bottleneck structure is composed of three cascaded convolution layers;
[0017] The first convolutional calculation unit consists of only one convolutional layer, with an input channel number of 64, an output channel number of 128, a convolutional kernel size of 7×7, and a stride of 2; the numbers of bottleneck structures in the second to fourth convolutional calculation units are 3, 4, 6, and 3 respectively, and the convolutional kernel sizes of the convolutional layers in each bottleneck structure are set to 1×1, 3×3, and 1×1 respectively; the input channel number of the first fully connected layer is 128, and the output channel number is 64; the input channel number of the second fully connected layer is 64, and the output channel number is 1.
[0018] Optionally, the construction of the knowledge distillation student sub-network specifically includes:
[0019] Build a 10-layer knowledge distillation student sub-network, and its structure is successively: the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, the fifth convolutional layer, the sixth convolutional layer, the seventh convolutional layer, the eighth convolutional layer, the first fully connected layer, the second fully connected layer;
[0020] The input channel number of the first convolutional layer is 3, the output channel number is 48, the convolutional kernel size is 3×3, and the stride is 1; the input channel number of the second convolutional layer is 48, the output channel number is 48, the convolutional kernel size is 3×3, and the stride is 2; the input channel number of the third convolutional layer is 48, the output channel number is 64, the convolutional kernel size is 3×3, and the stride is 1; the input channel number of the fourth convolutional layer is 64, the output channel number is 64, the convolutional kernel size is 3×3, and the stride is 2; the input channel number of the fifth convolutional layer is 64, the output channel number is 64, the convolutional kernel size is 3×3, and the stride is 1; the input channel number of the sixth convolutional layer is 64, the output channel number is 64, the convolutional kernel size is 3×3, and the stride is 1; the input channel number of the seventh convolutional layer is 64, the output channel number is 128, the convolutional kernel size is 3×3, and the stride is 1; the input channel number of the third convolutional layer is 128, the output channel number is 128, the convolutional kernel size is 3×3, and the stride is 1; the input channel number of the first fully connected layer is 128, and the output channel number is 64; the input channel number of the second fully connected layer is 64, and the output channel number is 1.
[0021] Optionally, the construction of the training set and the test set according to the image quality evaluation data set specifically includes:
[0022] Select at least 1000 reference-free natural images with different image contents from the natural image quality evaluation data set to form a sample set;
[0023] Randomly divide 80% of the reference-free natural images to form a training set, and the remaining 20% of the reference-free natural images form a test set.
[0024] Optionally, the data preprocessing of the training set and the test set to obtain a preprocessed data set specifically includes:
[0025] Normalize and block each image in the training set and the test set in sequence;
[0026] The blocking process uses a sliding window of size 112×112, and slides and blocks each image in the training set and the test set in the order of first row then column, first left then right, with a sliding step of 80;
[0027] For the teacher sub-network, perform supervised training. The image blocks obtained after blocking the same image all use the quality score label of the corresponding image as the quality score label of the image block for supervised training;
[0028] For the student sub-network, perform supervised training. The image blocks obtained by blocking the same image use the predicted score of the teacher sub-network for the image block as the quality score as the label for supervised training.
[0029] Optionally, after generating the image blocks of the video frames to be evaluated according to the preprocessed data set, it further includes:
[0030] After the student sub-network is trained, divide the image blocks of the video frames to be evaluated into multiple image blocks.
[0031] Optionally, the loss function used for the supervised training of the teacher sub-network is
[0032]
[0033] where represents the loss function of the teacher sub-network, f(·) represents the distorted image of the training set the predicted quality score of the image quality output by the teacher sub-network, and S represents the distorted image quality score label.
[0034] Optionally, the loss function used for the supervised training of the student sub-network is:
[0035]
[0036] where represents the loss function of the student sub-network, f(·) represents the distorted image of the training set the predicted quality score output by the well-trained teacher sub-network is used as the pseudo label of the quality score of the corresponding distorted image of the student sub-network, and g(·) represents the distorted image the predicted quality score output by the student sub-network.
[0037] Optionally, the training parameters for the supervised training are as follows: set the initial learning rate of the teacher sub-network to 2e-5, set the initial learning rate of the student sub-network to 1e-4, set the batch size to 64, set the weight decay to 5e-4, and set the number of training iterations to 60.
[0038] A method for determining the quality of video images for video conferencing provided by the present invention, the determination method includes: constructing a knowledge distillation teacher sub-network; constructing a knowledge distillation student sub-network; obtaining an image quality evaluation data set with rich image content; constructing a training set and a test set according to the image quality evaluation data set, and the training set also includes corresponding quality score labels; performing data preprocessing on the training set and the test set to obtain a preprocessed data set; generating video frame image blocks to be evaluated according to the preprocessed data set; using the trained student sub-network to predict the quality evaluation scores of multiple video frame image blocks to be evaluated; taking the average of multiple quality evaluation scores to obtain the quality evaluation score of the video to be evaluated. It can learn how to extract depth features more relevant to quality scores from a small-scale but very rich image content data set, and improve the generalization ability without increasing the computational complexity. Utilize the characteristics that the complex model has strong feature extraction ability but weak real-time performance, and the simplified model has weak feature extraction ability but strong real-time performance.
[0039] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below. Brief Description of the Drawings
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 It is a flowchart of a method for determining the quality of video images for video conferencing provided by an embodiment of the present invention. Detailed Embodiments
[0042] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0043] In the embodiments of the specification, claims and drawings of the present invention, the terms "comprising", "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a series of steps or units are included.
[0044] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments.
[0045] The object of the present invention is to address the deficiencies of the above-mentioned existing technologies, and propose a no-reference video image quality evaluation method for video conferencing based on knowledge distillation. By taking advantage of the strong feature extraction ability but weak real-time performance of complex models, and the weak feature extraction ability but strong real-time performance of lightweight models, a knowledge distillation network is used to fully exploit the relative advantages of complex and lightweight models, and solve the problem of poor generalization ability of lightweight models for video conferencing scenarios with rich image content.
[0046] The idea to achieve the object of the present invention is as follows: A teacher sub-network module with a higher model complexity is constructed to extract high-dimensional features highly relevant to image quality from a dataset with rich image content, and then the high-dimensional features are input into a fully connected layer to jointly optimize feature extraction and quality regression. After the teacher sub-network obtains a high test accuracy, the quality score predicted by the teacher sub-network from the distorted images in the training set is used as the pseudo-label of the quality score of the distorted images in the training set for the student sub-network module with a lower model complexity. Under the guidance of the pseudo-label of the quality score, the joint optimization of feature extraction and pseudo-label quality score regression is achieved, so that the student sub-network can learn the advanced generalization ability of the teacher sub-network for the quality evaluation dataset with rich content, and solve the problem of poor generalization ability of lightweight models for video conferencing scenarios with rich image content.
[0047] As Figure 1 shown, to achieve the above object, the specific steps of the present invention are as follows:
[0048] (1) Construct a knowledge distillation teacher sub-network:
[0049] (1a) Build a 7-layer knowledge distillation teacher sub-network, whose structure is successively: the first convolution calculation unit, the second convolution calculation unit, the third convolution calculation unit, the fourth convolution calculation unit, the fifth convolution calculation unit, the first fully connected layer, the second fully connected layer; the second to fifth convolution calculation units adopt a bottleneck structure, and each bottleneck structure consists of three cascaded convolution layers.
[0050] (1b) The first convolutional computing unit consists of only one convolutional layer, with 64 input channels, 128 output channels, a convolutional kernel size of 7×7, and a stride of 2; the numbers of bottleneck structures in the second to fourth convolutional computing units are 3, 4, 6, and 3 respectively, and the convolutional kernel sizes of the convolutional layers in each bottleneck structure are set to 1×1, 3×3, and 1×1 respectively; the first fully connected layer has 128 input channels and 64 output channels; the second fully connected layer has 64 input channels and 1 output channel;
[0051] (2) Construct a knowledge distillation student sub-network
[0052] (2a) Build a 10-layer knowledge distillation student sub-network, and its structure is as follows: the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, the fifth convolutional layer, the sixth convolutional layer, the seventh convolutional layer, the eighth convolutional layer, the first fully connected layer, and the second fully connected layer;
[0053] (2b) The first convolutional layer has 3 input channels, 48 output channels, a convolutional kernel size of 3×3, and a stride of 1; the second convolutional layer has 48 input channels, 48 output channels, a convolutional kernel size of 3×3, and a stride of 2; the third convolutional layer has 48 input channels, 64 output channels, a convolutional kernel size of 3×3, and a stride of 1; the fourth convolutional layer has 64 input channels, 64 output channels, a convolutional kernel size of 3×3, and a stride of 2; the fifth convolutional layer has 64 input channels, 64 output channels, a convolutional kernel size of 3×3, and a stride of 1; the sixth convolutional layer has 64 input channels, 64 output channels, a convolutional kernel size of 3×3, and a stride of 1; the seventh convolutional layer has 64 input channels, 128 output channels, a convolutional kernel size of 3×3, and a stride of 1; the third convolutional layer has 128 input channels, 128 output channels, a convolutional kernel size of 3×3, and a stride of 1; the first fully connected layer has 128 input channels and 64 output channels; the second fully connected layer has 64 input channels and 1 output channel.
[0054] (3) Construct a training set and a test set based on an image quality evaluation data set with rich image content, and the training set also includes corresponding quality score labels;
[0055] Select at least 1000 reference-free natural images with different image contents from the natural image quality evaluation data set to form a sample set, and randomly divide 80% of the reference-free natural images to form a training set, and the remaining 20% of the reference-free natural images form a test set.
[0056] (4) Data preprocessing
[0057] (4a) Perform normalization processing and block processing on each image in the training set and the test set in turn;
[0058] (4b) The block processing uses a sliding window of size 112×112, and slides and divides each image in the training set and the test set in the order of first row then column, first left then right, with a sliding step of 80;
[0059] (4c) For the teacher sub-network, the image blocks obtained after dividing the same image are all supervised and trained using the quality score label of this image as the quality score label of the image blocks; for the student sub-network, the image blocks obtained by dividing the same image are supervised and trained using the predicted score of the teacher sub-network for the image blocks as the pseudo quality score label.
[0060] (5) Generate the image blocks of the video frames to be evaluated
[0061] After the student sub-network is trained, each frame of the video to be evaluated is segmented into several image blocks according to the above-mentioned block method.
[0062] (6) Use the trained student sub-network to predict the quality evaluation scores Q of several image blocks of each frame of the image, and then calculate the mean value of the quality evaluation scores Q of all the image blocks of the video to be evaluated. The obtained average value is the quality evaluation score of the video to be evaluated.
[0063] The effect of the present invention is further illustrated by combining simulation experiments:
[0064] Simulation experiment conditions
[0065] The hardware platform for the simulation experiment of the present invention is: the processor is Intel(R)Xeon(R)CPU E5-2630 v4@2.20GHz, and the graphics card is NVIDIA GeForce GTX 2080Ti.
[0066] The software platform used in the simulation experiment of the present invention is: Ubuntu 18.04.3LTS operating system, Python3.5.2, Numpy 1.14.0, Pytorch 1.4.0 deep learning framework. The input images used in the simulation experiment of the present invention are natural images with complex and variable content in a simulated video conference, which are sourced from the publicly available image quality evaluation database LIVE Inthe Wild Image Quality Challenge (LIVEC).
[0067] The LIVEC database includes 1169 distorted images with different image contents, and its image format is bmp or jpg format.
[0068] Simulation content and its result analysis:
[0069] The simulation experiment of the present invention uses the present invention to perform no-reference image quality evaluation on 1,169 distorted images with different contents from the publicly available image quality evaluation database LIVEC, so as to simulate the no-reference image quality evaluation in a video conferencing scenario with complex and variable image contents.
[0070] In the simulation experiment, the publicly available image quality evaluation database used refers to:
[0071] The LIVEC database refers to the image quality evaluation database proposed by D. Ghadiyaram et al. in "Massive online crowdsourced study of subjective and objective picture quality[J]. IEEE Transactions on Image Processing, 25(1): 372–387, 2015", abbreviated as the LIVEC publicly available database.
[0072] The simulation experiment of the present invention uses two metrics, the Spearman rank-order correlation coefficient SROCC (Spearman rank-order correlation cofficient) and the Pearson linear correlation coefficient PLCC (Pearson linear correlation coefficient), to evaluate the video image quality evaluation effects of the no-reference video image quality evaluation method based on knowledge distillation after introducing the teacher sub-network and the no-reference video image quality evaluation method with only the student sub-network, respectively. The specific evaluation method is that the two methods are trained and tested using the same training set and test set, and the PLCC and SROCC values are calculated for the quality prediction scores of the two methods for N samples in the test set and the quality label scores corresponding to the test samples.
[0073] (1) The Spearman rank-order correlation coefficient SROCC, SROCC ∈ [-1, 1], is used to measure the monotonicity of the algorithm prediction. The higher its value, the more the evaluation result of the no-reference image quality evaluation method being evaluated can reflect the quality of the image. The expression is
[0074]
[0075] where d i represents the difference between the score predicted by the model for the i-th test image and the true score. N is the total number of samples in the test set.
[0076] (2) Pearson linear correlation coefficient (PLCC), which is mainly used to measure the accuracy of algorithm prediction. The higher its value, the closer the evaluation result of the no-reference image quality evaluation method being evaluated is to the subjective quality evaluation score of humans. The expression is:
[0077]
[0078] where s i and represent the true subjective quality score and the predicted subjective quality score of the i-th image. and represent the mean values of s i and . N is the number of samples in the test set.
[0079] The simulation results are shown in Table 1.
[0080] Table 1. Comparison table of evaluation results between the present invention and the no-reference video image quality evaluation method with only the student sub-network
[0081]
[0082] As can be seen from Table 1, for the publicly available image quality database of LIVE C which includes 1169 distorted images with different contents, both the Spearman rank correlation coefficient (SROCC) and the Pearson linear correlation coefficient (PLCC) of the evaluation results of the present invention are higher than those of the video image quality evaluation effect of the no-reference video image quality evaluation method with only the student sub-network.
[0083] The simulation experiment results effectively prove that the present invention can improve the generalization ability of the student sub-network model without increasing the computational complexity.
[0084] Beneficial effects:
[0085] The present invention evaluates the quality of video images under the condition of no original video images by using the trained student sub-network network framework.
[0086] The present invention uses the model compression technology based on knowledge distillation, enabling the trained student sub-network to improve the generalization ability without increasing the model complexity.
[0087] The above specific implementation manners further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for determining the quality of video images for video conferencing, characterized in that, The determination method includes: Construct a knowledge distillation teacher sub-network; Construct a knowledge distillation student sub-network; Obtain an image quality evaluation data set with rich image content; Construct a training set and a test set according to the image quality evaluation data set, and the training set also includes corresponding quality score labels; Perform data preprocessing on the training set and the test set to obtain a preprocessed data set, including: Perform normalization processing and block processing on each image in the training set and the test set in sequence; The block processing uses a sliding window with a size of 112×112, and slides and blocks each image in the training set and the test set in the order of first row then column, first left then right, with a sliding step of 80; For the teacher sub-network, perform supervised training. For the image blocks obtained after block processing of the same image, the quality score label of the corresponding image is used as the quality score label of the image blocks for supervised training; For the student sub-network, perform supervised training. For the image blocks obtained by block processing of the same image, the predicted score of the teacher sub-network for the image blocks is used as the quality score label for supervised training; Generate image blocks of video frames to be evaluated according to the preprocessed data set; Use the trained student sub-network to predict the quality evaluation scores of multiple image blocks of the video frames to be evaluated; Calculate the average of multiple quality evaluation scores to obtain the quality evaluation score of the video to be evaluated.
2. The method for determining the quality of video images for video conferencing according to claim 1, characterized in that, The construction of the knowledge distillation teacher sub-network specifically includes: Build a 7-layer knowledge distillation teacher sub-network, and the structure is in turn: the first convolutional calculation unit, the second convolutional calculation unit, the third convolutional calculation unit, the fourth convolutional calculation unit, the fifth convolutional calculation unit, the first fully connected layer, the second fully connected layer; the second to fifth convolutional calculation units adopt bottleneck structures, and each bottleneck structure consists of three cascaded convolutional layers; The first convolutional calculation unit consists of only one convolutional layer, with an input channel number of 64, an output channel number of 128, a convolutional kernel size of 7×7, and a stride of 2; the number of bottleneck structures of the second to fourth convolutional calculation units are 3, 4, and 6 respectively, and the convolutional kernel sizes of the convolutional layers in each bottleneck structure are set to 1×1, 3×3, and 1×1 respectively; the input channel number of the first fully connected layer is 128, and the output channel number is 64; the input channel number of the second fully connected layer is 64, and the output channel number is 1.
3. The method for determining the quality of video images for video conferencing according to claim 1, characterized in that, The construction of the knowledge distillation student sub-network specifically includes: Build a 10-layer knowledge distillation student sub-network, and its structure is in turn: the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, the fifth convolutional layer, the sixth convolutional layer, the seventh convolutional layer, the eighth convolutional layer, the first fully connected layer, the second fully connected layer; The number of input channels of the first convolutional layer is 3, the number of output channels is 48, the convolutional kernel size is 3×3, and the stride is 1; the number of input channels of the second convolutional layer is 48, the number of output channels is 48, the convolutional kernel size is 3×3, and the stride is 2; the number of input channels of the third convolutional layer is 48, the number of output channels is 64, the convolutional kernel size is 3×3, and the stride is 1; the number of input channels of the fourth convolutional layer is 64, the number of output channels is 64, the convolutional kernel size is 3×3, and the stride is 2; the number of input channels of the fifth convolutional layer is 64, the number of output channels is 64, the convolutional kernel size is 3×3, and the stride is 1; the number of input channels of the sixth convolutional layer is 64, the number of output channels is 64, the convolutional kernel size is 3×3, and the stride is 1; the number of input channels of the seventh convolutional layer is 64, the number of output channels is 128, the convolutional kernel size is 3×3, and the stride is 1; the number of input channels of the eighth convolutional layer is 128, the number of output channels is 128, the convolutional kernel size is 3×3, and the stride is 1; the number of input channels of the first fully connected layer is 128, and the number of output channels is 64; the number of input channels of the second fully connected layer is 64, and the number of output channels is 1.
4. The method for determining the quality of video images for video conferencing according to claim 1, characterized in that, The construction of the training set and the test set according to the image quality evaluation data set specifically includes: Select at least 1000 reference-free natural images with different image contents from the natural image quality evaluation data set to form a sample set; Randomly divide 80% of the reference-free natural images to form a training set, and the remaining 20% of the reference-free natural images form a test set.
5. The method for determining the quality of video images for video conferencing according to claim 1, characterized in that, After generating the video frame image blocks to be evaluated according to the preprocessed data set, it further includes: After the student sub-network is trained, the video frame image blocks to be evaluated are divided into multiple image blocks.
6. The method for determining the quality of video images for video conferencing according to claim 1, characterized in that, The loss function used for the supervised training of the teacher sub-network is Among them, represents the loss function of the teacher sub-network, and f(·) represents the distorted images in the training set is the predicted quality score of the image quality output by the teacher sub-network, and S represents the distorted image is the quality score label.
7. The method for determining the quality of video images for video conferencing according to claim 1, characterized in that, The loss function used for the supervised training of the student sub-network is: Among them, represents the loss function of the student sub-network, and f(·) represents the distorted images in the training set The predicted quality score output by the fully trained teacher sub-network is used as the pseudo-label of the quality score of the corresponding distorted image of the student sub-network, and g(·) represents the distorted image The predicted quality score output by the student sub-network.
8. A method for determining the video image quality for video conferencing according to claim 1, characterized in that The training parameters of the supervised training are: set the initial learning rate of the teacher sub-network to 2e-5, set the initial learning rate of the student sub-network to 1e-4, set the batch size to 64, set the weight decay to 5e-4, and set the number of training iterations to 60.
Citation Information
Patent Citations
Knowledge distillation-based model training method and device
CN112101526A
Cross-modal image aesthetics quality evaluation method based on knowledge distillation
CN112613303A
Non-reference image quality evaluation method based on deep feature transfer learning
CN113421237A