Video Quality Assessment Method, Its Device, Equipment, and Medium

By using a pre-trained video quality evaluation model, combined with image feature extraction, feature manifestation and classification network, the problem of insufficient robustness of existing models under different video content is solved, and a more accurate and stable video quality evaluation is achieved.

CN114881971BActive Publication Date: 2025-06-24GUANGZHOU HUANJU SHIDAI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210502154.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-09
Publication Date
2025-06-24
Estimated Expiration
2042-05-09

AI Technical Summary

Technical Problem

The existing reference-free video quality evaluation model is not robust enough to effectively evaluate different video content, resulting in a significant decrease in the effectiveness of the model when inserting sports-independent advertising videos during live broadcast.

Method used

The video quality evaluation model that is pre-trained to the converging state is adopted. Image feature extraction network is extracted through the image feature extraction network, the feature manifestation network is pooled for pooling operations to improve the significant feature weight, the classification network is classified and mapped to obtain quality scores, and the accuracy of feature extraction and mapping is enhanced through the attention layer and the global pooling layer.

Benefits of technology

It improves the accuracy and robustness of the video quality evaluation model, can effectively evaluate it under different video content, and reduces the effectiveness of the model due to content differences during live broadcast.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114881971B_ABST
    Figure CN114881971B_ABST
Patent Text Reader

Abstract

The present application relates to a video quality assessment method, apparatus, device, and medium in the field of image recognition technology. The method includes: obtaining a plurality of consecutive image frames in a video stream and formatting them into standardized data; extracting multi-channel image feature information corresponding to the standardized data by using an image feature extraction network in a video quality assessment model that has been pre-trained to a convergent state; performing pooling operations on the image feature information in different ways by a feature manifestation network of the assessment model to enhance the weights of the significant features between channels in the image feature information and obtain significant feature information; and performing classification mapping on the significant feature information by a classification network of the assessment model to obtain quality scores corresponding to the plurality of image frames. The present application can achieve high-robustness video quality assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image recognition, and in particular, to a video quality evaluation method, a corresponding device, a computer device, and a computer-readable storage medium. Background Art

[0002] Video quality evaluation is a method that has been studied in academia and industry since the emergence of video, mainly divided into reference-based evaluation and reference-free evaluation. The reference-based evaluation compares the effect differences frame by frame according to the given reference video, so that a relatively accurate evaluation score can be obtained relatively easily. The video reference-free evaluation is a method of giving a score for the subjective experience of the quality of the current video. According to the score of the reference-free evaluation, quality indicators such as the bit rate, frame rate, and resolution of network transmission can be dynamically adjusted to achieve the purpose of controlling costs. Through adjustment, the subjective video viewing experience of users is basically not affected. In this way, both user needs are met and network transmission resources can be saved indirectly. It can be seen that it can bring considerable economic benefits.

[0003] The difficulty of the reference-free evaluation method is very great. Traditional algorithms can only play a very limited role in reference-free video evaluation. Until the rapid development of artificial intelligence in recent years, the reference-free video evaluation method based on CNN (Convolutional Neural Network) has been taken seriously and developed rapidly, becoming the main research and application direction. The industry tries to ensure the adaptability and high availability of the model through many different technical solutions, but most of them do not consider the inevitable differences in video content that the model may face when dealing with different video inputs, and cannot perform corresponding reference-free evaluations for different video contents, resulting in insufficient robustness of the model. For example, applying a reference-free video evaluation model for sports events to a sports event live broadcast and inserting an advertisement video unrelated to sports during the live broadcast will greatly reduce the effectiveness of the model at this time, thus causing the robustness to deteriorate.

[0004] Based on the deficiencies of traditional technologies, the present application conducts corresponding explorations. Summary of the Invention

[0005] The purpose of the present application is to solve at least one of the above problems and provide a video quality evaluation method, a corresponding device, a computer device, and a computer-readable storage medium.

[0006] To meet the various purposes of the present application, the following technical solutions are adopted:

[0007] On the one hand, to meet one of the purposes of the present application, a video quality evaluation method is provided, including the following steps:

[0008] Obtain a continuous plurality of image frames in the video stream and format them into standardized data;

[0009] Extract multi-channel image feature information corresponding to the standardized data by using the image feature extraction network in the video quality assessment model pre-trained to the convergence state;

[0010] Pool the image feature information in different ways by the feature manifestation network of the evaluation model to enhance the weight of the significant features between channels in the image feature information and obtain significant feature information;

[0011] Perform classification mapping on the significant feature information by the classification network of the evaluation model to obtain the quality scores corresponding to the multiple image frames.

[0012] In a further embodiment, pooling the image feature information in different ways by the feature manifestation network of the evaluation model to enhance the weight of the significant features between channels in the image feature information and obtain significant feature information includes the following steps:

[0013] Divide the image feature information into two paths and input them into the pooling layers of the two branches of the feature manifestation network respectively to perform different types of pooling operations to obtain corresponding sampled feature information;

[0014] In the two branches, apply attention layers to extract weights from the corresponding sampled feature information respectively to enhance the numerical gap between the significant features and non-significant features therein and obtain corresponding weighted feature information;

[0015] Concatenate the weighted feature information of the two branches and perform normalization to obtain significant feature information.

[0016] In a further embodiment, performing classification mapping on the significant feature information by the classification network of the evaluation model to obtain the quality scores corresponding to the multiple image frames includes the following steps:

[0017] Input the significant feature information into a global pooling layer to perform a pooling operation to obtain corresponding pooled feature information;

[0018] Output the pooled feature information to a first fully connected layer for classification mapping to obtain a predicted quality score;

[0019] Output the pooled feature information to a second fully connected layer for classification mapping to obtain the confidence corresponding to the quality score;

[0020] Determine the quality scores with confidence higher than a preset threshold as the quality scores corresponding to the multiple image frames.

[0021] In an extended embodiment, after the step of obtaining the quality scores corresponding to the multiple image frames, the following steps are included:

[0022] Adjust the encoding frame rate of the encoder of the video stream according to the quality score.

[0023] In a further embodiment, before the step of obtaining a plurality of consecutive image frames in the video stream, the following steps are included:

[0024] Iteratively train the video quality assessment model using the training samples in a preset dataset until it converges.

[0025] In a further embodiment, iteratively training the video quality assessment model using the training samples in a preset dataset until it converges includes the following steps:

[0026] Call a single training sample from the preset dataset to train the video quality assessment model. The training sample includes a plurality of consecutive image frames constituting video data, and is labeled with the corresponding subjective quality score;

[0027] Calculate the model loss value based on the subjective quality score. The model loss value is the sum of the first loss value corresponding to the quality score obtained by the assessment model and the second loss value of the confidence of the quality score;

[0028] Judge whether the model loss value reaches a preset threshold. When it reaches the preset threshold, it is determined that the assessment model has converged and the training is terminated; otherwise, perform gradient update on the assessment model and continue to perform iterative training using the next training sample.

[0029] In a further embodiment, calculating the model loss value based on the subjective quality score includes the following steps:

[0030] Calculate the loss value between the predicted quality score corresponding to each image frame in the training sample and the pre-labeled subjective quality score, and obtain the mean absolute error of each loss value as the first loss value;

[0031] Calculate the cross-entropy loss of the confidence of the quality score to obtain the second loss value. The second loss value is the mean of the logarithms of the corrected prediction probabilities obtained by matching the positive sample prediction probabilities corresponding to each image frame in the training sample with the confidence as the weight;

[0032] Calculate the sum of the first loss value and the second loss value to obtain the model loss value corresponding to the training sample.

[0033] On the other hand, a video quality assessment device provided to meet one of the purposes of the present application includes an image processing module, a feature extraction module, a weight extraction module, and a score output module, where: The image processing module is used to obtain a continuous plurality of image frames in the video stream and format them into standardized data; The feature extraction module is configured to extract multi-channel image feature information corresponding to the standardized data by using an image feature extraction network in a video quality assessment model that has been pre-trained to a converged state; The weight extraction module is configured to perform pooling operations on the image feature information in different ways by a feature manifestation network of the assessment model to enhance the weights of the significant features between channels in the image feature information and obtain significant feature information; The score output module is configured to perform classification mapping on the significant feature information by a classification network of the assessment model to obtain the quality scores corresponding to the plurality of image frames.

[0034] In a further embodiment, the weight extraction module includes: Two-way pooling sub-module, used to divide the image feature information into two paths and input them into the pooling layers of two branches of the feature manifestation network respectively to perform different types of pooling operations to obtain corresponding sampled feature information; Weighted feature information acquisition sub-module, used to apply attention layers to the corresponding sampled feature information in the two branches respectively to extract weights to enhance the numerical gap between the significant features and non-significant features therein and obtain corresponding weighted feature information; Normalization sub-module, used to splice the weighted feature information of the two branches and perform normalization to obtain significant feature information.

[0035] In a further embodiment, the score output module includes: Global pooling sub-module, used to input the significant feature information into a global pooling layer to perform pooling operations to obtain corresponding pooled feature information; Quality score acquisition sub-module, used to output the pooled feature information to a first fully connected layer for classification mapping to obtain a predicted quality score; Confidence level acquisition sub-module, used to output the pooled feature information to a second fully connected layer for classification mapping to obtain the confidence level corresponding to the quality score; Quality score determination sub-module, used to determine the quality score with a confidence level higher than a preset threshold as the quality score corresponding to the plurality of image frames.

[0036] In an extended embodiment, after the score output module, there is further included: An encoding frame rate adjustment module, used to adjust the encoding frame rate of the encoder of the video stream according to the quality score.

[0037] In a further embodiment, before the image processing module, there is further included: An iterative training sub-module, used to perform iterative training on the video quality assessment model by using training samples in a preset dataset and train it to a converged state.

[0038] In a further embodiment, the iterative training sub-module includes: an implementation training unit for training the video quality assessment model by calling a single training sample from a preset data set, where the training sample includes a plurality of consecutive image frames constituting video data and is labeled with a corresponding subjective quality score; a loss value unit for calculating a model loss value based on the subjective quality score, where the model loss value is the sum of a first loss value corresponding to the quality score obtained by the assessment model and a second loss value of the confidence of the quality score; an iterative training judgment unit for judging whether the model loss value reaches a preset threshold, and when it reaches the preset threshold, determining that the assessment model has converged and terminating the training; otherwise, performing gradient update on the assessment model and continuing to perform iterative training with the next training sample.

[0039] In a further embodiment, the loss value unit includes: a first loss value sub-unit for calculating the loss value between the predicted quality score corresponding to each image frame in the training sample and the pre-labeled subjective quality score, and obtaining the mean absolute error of each of the loss values as the first loss value; a second loss value sub-unit for obtaining a second loss value by using the cross-entropy loss of the confidence of the quality score, where the second loss value is the mean of the logarithm sum of the corrected prediction probabilities obtained by matching the positive sample prediction probabilities corresponding to each image frame in the training sample with the confidence as the weight; a model loss value sub-unit for calculating the sum of the first loss value and the second loss value to obtain the model loss value corresponding to the training sample.

[0040] In another aspect, a computer device provided to meet one of the purposes of the present application includes a central processing unit and a memory, and the central processing unit is used to call and run a computer program stored in the memory to execute the steps of the video quality assessment method described in the present application.

[0041] In another aspect, a computer-readable storage medium provided to meet another purpose of the present application stores a computer program implemented according to the video quality assessment method in the form of computer-readable instructions, and when the computer program is called and run by a computer, it executes the steps included in the method.

[0042] Compared with the prior art, the present application has multiple advantages, including at least the following aspects:

[0043] In this application, a plurality of temporally consecutive image frames are obtained by decoding a video stream, formatted and converted into standardized data, and input into a video quality assessment model that has been pre-trained to a convergent state to extract corresponding multi-channel image feature information. Subsequently, different pooling operations are performed to enhance the weights of the significant features between channels in the image feature information, obtaining significant feature information. Furthermore, classification mapping is performed on the significant feature information to obtain quality scores corresponding to the plurality of image frames. It can be understood that this application can obtain various advantages, including but not limited to:

[0044] On the one hand, the model performs different pooling operations on the extracted image feature information, highly manifesting the corresponding significant features and effectively hiding the corresponding secondary features, guiding the model to highly focus on the necessary features in the image feature information, which helps to improve the accuracy and reliability of the model.

[0045] On the other hand, the video quality assessment model has a simple structure, is easy to implement and train, and is convenient to be deployed on the media stream server corresponding to the network live broadcast on the e-commerce platform to provide services for evaluating the quality of the live video stream. This enables, without affecting the user's viewing experience, accurately evaluating reliable quality scores and correspondingly adjusting the picture quality of the live video stream, saving the bandwidth for transmitting the video stream, stably playing the live video stream, ensuring the smoothness and real-time nature of the live video stream playback, enhancing the user experience, and increasing user stickiness. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The above and / or additional aspects and advantages of this application will become apparent and easier to understand from the following description of the embodiments in conjunction with the drawings, where:

[0047] Figure 1 is a schematic flowchart of a typical embodiment of the video quality assessment method of this application;

[0048] Figure 2 is a schematic structural diagram showing an exemplary structure of the video quality assessment model in an embodiment of this application;

[0049] Figure 3 is a schematic structural diagram showing an exemplary structure of the feature manifestation network in the video quality assessment model in an embodiment of this application;

[0050] Figure 4 is a schematic flowchart showing different pooling operations performed by the feature manifestation network of the video quality assessment model in an embodiment of this application;

[0051] Figure 5 is a schematic flowchart showing the classification mapping performed by the classification network of the video quality assessment model in an embodiment of this application;

[0052] Figure 6Flow schematic diagram of an extended embodiment of the present application;

[0053] Figure 7 In the embodiment of the present application, it is a flow schematic diagram of pre-training the video quality assessment model until convergence;

[0054] Figure 8 In the embodiment of the present application, it is a flow schematic diagram of implementing iterative training of the video quality assessment model;

[0055] Figure 9 In the embodiment of the present application, it is a flow schematic diagram of obtaining the loss value of the video quality assessment model;

[0056] Figure 10 It is a principle block diagram of the video quality assessment device of the present application;

[0057] Figure 11 It is a structural schematic diagram of a computer device adopted by the present application. Detailed implementation manners

[0058] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application.

[0059] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0060] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the field to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0061] Those skilled in the art of the present technology can understand that the "client", "terminal", and "terminal device" used herein include both devices with wireless signal receivers that only have the ability to receive and do not have the ability to transmit, and devices with receiving and transmitting hardware that have the receiving and transmitting hardware capable of two-way communication on a two-way communication link. Such devices may include: cellular or other communication devices such as personal computers, tablet computers, etc., which have a single-line display or a multi-line display or cellular or other communication devices without a multi-line display; PCS (Personal Communications Service), which can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which may include a radio frequency receiver, pager, Internet / intranet access, web browser, notepad, calendar, and / or GPS (Global Positioning System) receiver; conventional laptop and / or palm computers or other devices, which are conventional laptop and / or palm computers or other devices with and / or including a radio frequency receiver. The "client", "terminal", and "terminal device" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to run locally, and / or run in a distributed form at any other location on the earth and / or in space. The "client", "terminal", and "terminal device" used herein can also be a communication terminal, an Internet access terminal, a music / video playback terminal, for example, it can be a PDA, MID (Mobile Internet Device), and / or a mobile phone with music / video playback function, or it can also be a smart TV, a set-top box, etc.

[0062] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer, and is a hardware device with the necessary components disclosed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.

[0063] It should be noted that the concept of "server" as used in this application can similarly be extended to apply to server clusters. According to the network deployment principles understood by those skilled in the art, the servers should be logically divided. Physically, these servers can either be independent of each other but can be invoked through interfaces, or integrated into a single physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility and should not be restricted by this when implementing the network deployment method of this application.

[0064] One or several technical features of this application, unless expressly specified, can either be deployed on a server and accessed by a client remotely invoking the online service interface provided by the server, or directly deployed and run on the client for access.

[0065] The neural network models cited or potentially cited in this application, unless expressly specified, can either be deployed on a remote server and remotely invoked on the client, or deployed on a client capable of handling the device and directly invoked. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operating resources and avoid over-occupying the client's hardware operating resources.

[0066] All kinds of data involved in this application, unless expressly specified, can either be remotely stored on a server or stored on a local terminal device, as long as it is suitable for being invoked by the technical solution of this application.

[0067] Those skilled in the art should be aware that although the various methods of this application are described based on the same concept and thus show commonality with each other, unless otherwise specified, these methods can all be executed independently. Similarly, for the various embodiments disclosed in this application, they are all proposed based on the same inventive concept. Therefore, for concepts with the same expression, as well as concepts that are only appropriately transformed for convenience although the concept expressions are different, they should be equivalently understood.

[0068] For the various embodiments to be disclosed in this application, unless expressly indicated that there is a mutually exclusive relationship between them, the relevant technical features involved in each embodiment can be cross-combined to flexibly construct new embodiments, as long as this combination does not deviate from the creative spirit of this application and can meet the requirements in the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.

[0069] Please refer to Figure 1 , in a typical embodiment of the video quality assessment method of this application, it includes the following steps:

[0070] Step S1100: Obtain a continuous plurality of image frames in the video stream and format them into standardized data;

[0071] The video quality assessment model of the present application will be used to predict a no-reference score for a video stream. Accordingly, in one embodiment, the video quality assessment model of the present application can be deployed on a media server corresponding to a webcast in an e-commerce platform to provide corresponding services. The service can be to predict a no-reference score for a video stream generated by a webcast in an e-commerce platform. The video stream is composed of a series of image frames that are continuous in time. The corresponding multiple image frames can be obtained by decoding the video stream. For example, the image data type corresponding to the image frame is unit8, which needs to be formatted, converted to the float data type, and then divided by 255 to be normalized to the corresponding value range of [0, 1], so as to obtain the corresponding normalized data.

[0072] Step S1200: Extract multi-channel image feature information corresponding to the normalized data by using the image feature extraction network in the video quality assessment model that has been pre-trained to a converged state;

[0073] An exemplary example of the network structure of the video quality assessment model is as follows Figure 2 shown, which includes a backbone module 200 and a classification network 203. The backbone module 200 includes an image feature extraction network 201 and a feature manifestation network 202. The classification network 203 includes a global pooling layer 204, a prediction module 205, and a confidence module 206.

[0074] In one embodiment, the image feature extraction network 201 can preferably be ResNet (Residual Network)-50. ResNet-50 includes 50 conv2d operations, has a relatively deep network hierarchy and strong network performance. Secondly, compared with a general CNN (Convolutional Neural Network), ResNet-50 adds a residual network. The residual network adds an identity mapping, skips the operation of one layer or multiple layers of the network. At the same time, in the process of backpropagation, through the shortcut connection, the low-level network is directly connected to the high-level network, and the gradient of the high-level network can be directly transmitted to the low-level network, solving the problem of gradient disappearance caused by the network being too deep in the backpropagation process. In addition, ResNet-50 is a commonly used backbone network for deep learning image processing, and can achieve good results in related image processing such as image classification and object detection. Accordingly, the normalized data is input into ResNet-50, and multi-channel image feature information corresponding to the normalized data is extracted. The specific number of channels is determined by the network structure of ResNet-50. According to the disclosure of this embodiment, those skilled in the art can also select a corresponding neural network as the image feature extraction network based on prior knowledge or experimental experience.

[0075] Step S1300: The feature manifestation network of the evaluation model performs pooling operations on the image feature information in different ways to increase the weights of the significant features between channels in the image feature information, and obtain significant feature information;

[0076] In one embodiment, the feature manifestation network 202 of the evaluation model may preferably use CA (Channel Attention), which is connected to the image feature extraction network 201. Accordingly, CA performs attention on the image feature information between the corresponding channels of the image feature extraction network to which it is connected. The attention is to learn the weight distribution for each channel in the channel dimension, and then apply the weight distribution to the image feature information. Here, the weight distribution can be learned by performing pooling operations on the image feature information in different ways. Further, according to the weight distribution, the image feature information extracted by the image feature extraction network is weighted, increasing the weights of the significant features between channels in the image feature information and reducing the weights of the non-significant features between channels in the image feature information. Furthermore, the significant feature information output by CA is obtained. The different ways of pooling operations will be further revealed in subsequent embodiments of this part, and this step will not be elaborated for the time being. According to the disclosure of this embodiment, those skilled in the art can also flexibly select an appropriate neural network as the feature manifestation network based on prior knowledge or experimental experience.

[0077] Step S1400: The classification network of the evaluation model performs classification mapping based on the significant feature information to obtain the quality scores corresponding to the multiple image frames.

[0078] Referring to the disclosure of the network structure of the video quality evaluation model in Step S1200, in one embodiment, the global pooling layer 204 in the classification network 203 is connected to the feature manifestation network 202 in the backbone module 200, and is respectively connected to the prediction module 205 and the confidence module 206 in the classification network 203. Accordingly, the global pooling layer 204 receives the significant feature information output by the feature manifestation network, performs global average pooling operation on it to spatially average the significant feature information, and then unfolds it into a one-dimensional vector in the channel direction, and then inputs it into the prediction module 205 and the confidence module 206 respectively. First, the prediction module 205 predicts the quality scores corresponding to the multiple image frames. Second, the confidence module 206 obtains the reliability of the quality scores. The specific implementation of the classification network will be further revealed in subsequent embodiments of this part, and this step will not be elaborated for the time being.

[0079] Please refer to Figure 4, in a further embodiment, in step S1300, the feature manifestation network of the evaluation model performs pooling operations on the image feature information in different ways to enhance the weights of the significant features between channels in the image feature information and obtain significant feature information, including the following steps:

[0080] Step S1310: Divide the image feature information into two paths and respectively input them into the pooling layers of the two branches of the feature manifestation network to perform different types of pooling operations to obtain corresponding sampled feature information;

[0081] The feature display network is CA, and its structural exemplary example is Figure 3 as shown, which includes a max pooling layer 300, an average pooling layer 301, an attention layer 302, and a Sigmoid function 303. The max pooling layer 300 and the average pooling layer 301 are connected to their corresponding attention layers 302, and the attention layer 302 is a convolutional layer.

[0082] The feature manifestation network receives the image feature information input by the image feature network connected to it, divides the image feature information into two paths, and respectively inputs them into the pooling layers of the two branches of the feature manifestation network, namely the max pooling layer 300 and the average pooling layer 301, and then performs different types of pooling operations. First, perform max pooling operation in the max pooling layer 300, calculate and record the position of the corresponding maximum value in the image feature information, so as to better retain the features corresponding to texture and contour in the image feature information and appropriately reduce the influence of useless information. Second, perform average pooling operation in the average pooling layer, average the values in the corresponding feature map of the image feature information through the corresponding filter, so as to better retain the overall features in the image feature information and highlight the features corresponding to the background. Accordingly, obtain the corresponding sampled feature information of the two pooling layers.

[0083] Step S1320: In the two branches, respectively apply the attention layer to extract the weights of the corresponding sampled feature information to enhance the numerical gap between the significant features and the non-significant features therein and obtain the corresponding weighted feature information;

[0084] Further, the attention layers 302 connected correspondingly after the pooling layers of the two branches respectively receive the corresponding sampled feature information of the two pooling layers, and then respectively perform further weight extraction through the corresponding convolutional layers. The convolutional layer includes multiple specific parameters, such as the convolutional kernel size, the number of convolutional kernels, the parameters inside the convolutional kernel, the stride, etc. Those skilled in the art can flexibly set them according to prior knowledge or experimental data here. Then, the corresponding weight distribution is learned. Further, according to the weight distribution, the sampled feature information is weighted, the weight value of the corresponding significant feature is increased, and at the same time, the weight value of the corresponding non-significant feature is decreased, so as to obtain the corresponding weighted feature information.

[0085] Step S1330: Concatenate the weighted feature information of the two branches and then perform normalization to obtain significant feature information.

[0086] The extracted feature information of the two branches is concatenated correspondingly and input into the Sigmoid function 303 for normalization to obtain the output significant feature information corresponding to the channels corresponding to the image feature information. The Sigmoid function can be flexibly set by those skilled in the art according to prior knowledge or experimental data.

[0087] In this embodiment, by using CA as the feature display network, its corresponding attention mechanism learns the weight distribution, so as to perform corresponding weight division on the image feature information, increase the weight of the corresponding significant feature, and then obtain the corresponding significant feature information. It can be seen that the accuracy of feature recognition of the video quality assessment model can be significantly improved, laying a foundation for the subsequent prediction module and confidence module of the model, and making the quality score predicted by the model more reliable.

[0088] Please refer to Figure 5 , in a further embodiment, step S1400: The classification network of the evaluation model performs classification mapping according to the significant feature information to obtain the quality scores corresponding to the multiple image frames, including the following steps:

[0089] Step S1410: Input the significant feature information into the global pooling layer to perform a pooling operation to obtain the corresponding pooled feature information;

[0090] Referring to the disclosure of the network structure of the video quality assessment model in step S1200, in one embodiment, the global pooling layer 204 receives the significant feature information output by the feature manifestation network 202 connected to it, and performs global average pooling operation therein, so as to perform spatial averaging on the significant feature information. Then, it is expanded into a one-dimensional vector in the channel direction corresponding to the significant feature information as the pooled feature information.

[0091] Step S1420: Output the pooling feature information to the first fully connected layer for classification mapping to obtain a predicted quality score.

[0092] Refer to the disclosure of the network structure of the video quality assessment model in step S1200. In one embodiment, the first fully connected layer in the prediction module 205 receives the pooling feature information output by the global pooling layer 204 connected thereto. The first fully connected layer is a 2-layer FC (fully connected) layer, and the recommended parameters are 1280 and 512 respectively. Accordingly, the first fully connected layer performs a linear transformation and corresponding classification mapping to map to the corresponding classification space to obtain a predicted quality score.

[0093] Step S1430: Output the pooling feature information to the second fully connected layer for classification mapping to obtain the confidence level corresponding to the quality score.

[0094] Refer to the disclosure of the network structure of the video quality assessment model in step S1200. In one embodiment, the confidence module 206 receives the pooling feature information output by the global pooling layer 204 connected thereto. The second fully connected layer is a 3-layer FC (fully connected) layer, and the recommended parameters are 512, 128, and 256 respectively. Accordingly, the second fully connected layer performs a linear transformation and corresponding classification mapping to map to the corresponding classification space to obtain the confidence level corresponding to the quality score.

[0095] Step S1440: Determine the quality scores corresponding to the multiple image frames as the quality scores with confidence levels higher than a preset threshold.

[0096] It is allowed to set a preset threshold for comparing with the confidence level to determine that the quality score corresponding to the confidence level higher than it is the quality score corresponding to the multiple image frames. The specific value of the threshold here can be flexibly set by those skilled in the art according to business requirements and experimental data.

[0097] In this embodiment, the classification network predicts the corresponding quality scores for the input significant feature information and obtains the corresponding confidence levels, and then uses the appropriate confidence level corresponding quality scores as the quality scores corresponding to the image frames. It can be seen that due to obtaining reliable confidence levels, it is possible to determine whether the quality scores predicted by the video quality assessment model are valid based on the confidence levels, and quality scores with high robustness can be obtained, which greatly improves the robustness and reliability of the model, enables prediction of the image frames corresponding to video streams with different bitrates and / or resolutions, accurately predicts reliable quality scores, and can achieve good overall benefits.

[0098] Please refer to Figure 6, in an extended embodiment, after step S1400 of obtaining quality scores corresponding to the multiple image frames, the following steps are included:

[0099] Step S1500, adjusting the encoding frame rate of the encoder of the video stream according to the quality scores.

[0100] In one embodiment, in the webcast scenario of an e-commerce platform, a live user submits a video stream to a media server corresponding to the webcast of the e-commerce platform. After decoding it, image frames in YUV or RGB format are obtained. At this time, through the video quality assessment model of the present application deployed inside the media server corresponding to the webcast, if quality scores corresponding to multiple temporally consecutive image frames of the video stream that meet the confidence level required by the service and quality scores that do not meet the confidence level required by the service are predicted. Among them, for the image frames corresponding to the quality scores that meet the confidence level required by the service, the encoding frame rate of the encoder of the video stream is adjusted to correspondingly reduce the bit rate and / or reduce the resolution to degrade the image quality. For the image frames corresponding to the quality scores that do not meet the confidence level required by the service, the encoding frame rate of the encoder of the video stream is adjusted to correspondingly increase or maintain the original higher bit rate and / or resolution to ensure the image quality. Then, the adjusted video stream is pushed by the media server to the devices of the viewer users in the live broadcast room so that the devices can render the video stream for playback accordingly. The media server is set up by the e-commerce platform to provide corresponding online services for its webcast users and can perform corresponding encoding and decoding on the video stream.

[0101] In another embodiment, in the webcast scenario of a terminal device, an imaging unit acquires data frames corresponding to image frames. One path is to draw the data frame texture for rendering and display on the image user interface of the live user, and the other path is to convert the data frame into an image frame in YUV or RGB format through corresponding decoding operations. At this time, through the video quality assessment model of the present application deployed inside the media server corresponding to the webcast, if quality scores corresponding to multiple temporally consecutive image frames of the video stream that meet the confidence level required by the service and quality scores that do not meet the confidence level required by the service are predicted. Among them, for the image frames corresponding to the quality scores that meet the confidence level required by the service, the encoding frame rate of the encoder of the video stream is adjusted to correspondingly reduce the bit rate and / or reduce the resolution to degrade the image quality. For the image frames corresponding to the quality scores that do not meet the confidence level required by the service, the encoding frame rate of the encoder of the video stream is adjusted to correspondingly increase or maintain the original higher bit rate and / or resolution to ensure the image quality. Then, the adjusted video stream is pushed by the media server to the devices of the viewer users in the live broadcast room so that the devices can render the video stream for playback accordingly.

[0102] In this embodiment, it is disclosed that the online deployment of the video quality assessment model of the present application is applied to the video stream of the network live broadcast, so that the quality score carrying confidence can be predicted according to the model for the video stream, and accordingly, the encoding frame rate of the encoder of the video stream can be adjusted. On the one hand, the adjusted video stream can not only meet the video quality required by the service, but also save the bandwidth for transmitting the video stream. On the other hand, the quality score combined with confidence is more reliable, which can effectively avoid incorrect quality scores predicted due to insufficient model robustness, so as to avoid playback accidents caused by incorrect adjustment of the encoding frame rate and output of the corresponding video stream based on the incorrect quality score alone.

[0103] Please refer to Figure 7 , in a further embodiment, before the step S1100 of obtaining consecutive multiple image frames in the video stream, the following steps are included:

[0104] Step S1000: Perform iterative training on the video quality assessment model using the training samples in the preset dataset until it converges.

[0105] It can be achieved by collecting videos corresponding to certain amounts of different video quality metrics. The video quality metrics can be subjective human ratings, bitrate, resolution, frame rate, etc. Then, the videos are decoded to obtain corresponding multiple image frames. For example, the image data type corresponding to the image frames is unit8, which needs to be formatted, converted to the float data type, and then divided by 255 to be normalized to the corresponding value range of [0, 1]. The formatted multiple image frames are used as the training samples in the dataset. Further, the number of training samples in the dataset is expanded by means of data augmentation, so as to prevent model overfitting to a certain extent. Data Augmentation refers to changing the features of the samples according to prior knowledge under the condition of keeping the sample labels unchanged, so that the newly generated samples also conform to or approximately conform to the true distribution of the data. The ways of data augmentation can be mirroring, rotation, scaling, adjusting brightness, contrast, Gaussian noise, Mosaic, Mixup, Cutout, CutMix, etc. According to the above disclosure, those skilled in the art can flexibly implement the preset dataset.

[0106] In one embodiment, reference may be made to the disclosure of the network structure of the video quality assessment model in step S1200. Accordingly, a single training sample in a preset data set is input into the video quality assessment model. Through the image feature extraction network and the feature manifestation network of the model, the significant feature information corresponding to the training sample is learned and extracted. Then, it is input into the classification network of the model to make a corresponding classification mapping, predict the corresponding quality score, and obtain the confidence level corresponding to this quality score. Further, two loss values corresponding to the quality score predicted by the model and its corresponding confidence level are calculated. Accordingly, it is determined whether the preset threshold is reached. When it reaches, it is determined that the model has been trained to convergence, and the training is terminated; otherwise, the evaluation model is updated by gradient, and the iterative training is continued with the next training sample. The specific training process of the video quality assessment model will be further disclosed in subsequent embodiments, and this step will not be elaborated for the time being.

[0107] Please refer to Figure 8 , in a further embodiment, step S1000: Iteratively train the video quality assessment model with the training samples in a preset data set until it reaches a convergence state, including the following steps:

[0108] Step S1010: Call a single training sample from a preset data set to train the video quality assessment model. The training sample includes a continuous plurality of image frames constituting video data, and is labeled with the corresponding subjective quality score.

[0109] The preset data set can be implemented with reference to step S1000. In addition, each training sample in the data set is labeled with the corresponding subjective quality score. The subjective quality score can be obtained by subjectively scoring the video quality of a continuous plurality of image frames of the video data corresponding to the training sample by multiple persons and calculating the average value of these subjective scores. Accordingly, a single training sample is called from the preset data set and input into the video quality assessment model for training. In one embodiment, the video quality assessment model is as disclosed in the network structure of the video quality assessment model in step S1200. The image feature extraction network and the feature manifestation network of the video quality assessment model learn and extract the significant feature information corresponding to the training sample, and then input it into the classification network of the model to make a corresponding classification mapping, predict the corresponding quality score, and obtain the confidence level corresponding to this quality score.

[0110] Step S1020: Calculate the model loss value based on the subjective quality score. The model loss value is the sum value of the first loss value corresponding to the quality score obtained by the evaluation model and the second loss value of the confidence level of this quality score.

[0111] Further, call the first loss function to calculate the first loss value corresponding to the quality score obtained by the evaluation model based on the subjective quality score. Call the second loss function, and use the cross-entropy loss of the confidence of the quality score obtained by the evaluation model as the second loss value. Furthermore, use the sum of the first loss value and the second loss value as the loss value of the evaluation model. Specifically, the first loss function and the second loss function can be flexibly implemented by those skilled in the art according to prior knowledge or experimental data, and can also be implemented according to the further disclosure of some subsequent embodiments. For the time being, this step will not be elaborated here.

[0112] Step S1030: Determine whether the model loss value reaches a preset threshold. When it reaches the preset threshold, it is determined that the evaluation model has converged, and the training is terminated; otherwise, perform gradient update on the evaluation model, and continue to perform iterative training using the next training sample.

[0113] It is allowed to set a preset threshold for determining whether the evaluation model has converged. When it reaches or exceeds the preset threshold, it indicates that the model has been trained to convergence, and the training can be terminated accordingly. In addition, when it does not reach the preset threshold, it indicates that the model has not been fully fitted and has not converged, and the gradient update can be performed on the evaluation model accordingly. Update the weight parameters corresponding to some or all of the networks in the image feature extraction network, feature visualization network, and classification network of the evaluation model, and then perform iterative training using the next training sample.

[0114] In this embodiment, the training process of the quality evaluation model is disclosed. It can be seen that under the supervised training of the subjective quality score, the model is capable of evaluating the quality of each image frame corresponding to the video, obtaining a quality score with confidence. Furthermore, when the model is put into use, it can be deployed to the server. The server receives the video data and applies the model to quickly and accurately predict the quality score with confidence of each image frame corresponding to the video data. Then, the reliable quality score can be determined according to the confidence that meets the business requirements. On the premise of meeting the user's overall visual perception requirements for the video, the image quality of the corresponding image frames can be appropriately adjusted to save the bandwidth for transmitting video data.

[0115] Please refer to Figure 9 , in a further embodiment, step S1020: Calculate the model loss value based on the subjective quality score, including the following steps:

[0116] Step S1021: Calculate the loss value between the predicted quality score of each image frame in the training sample and the pre-annotated subjective quality score, and obtain the mean absolute error of each loss value as the first loss value;

[0117] Call the first loss function. According to the predicted quality scores corresponding to each image frame in the training samples and the pre-annotated subjective quality scores, calculate the absolute value of the error between the two, and then obtain the average value of the absolute value corresponding to the total number of image frames in the training samples to obtain the first loss value. The first loss function is exemplified as follows:

[0118]

[0119] Where: Loss mos is the first loss value, n is the total number of image frames in the training samples, y out is the quality score predicted by the model for a corresponding single image frame, y gt is the subjective quality score for a corresponding single image frame.

[0120] Step S1022: Calculate the cross-entropy loss of the confidence of the quality score to obtain the second loss value. The second loss value is the mean of the logarithmic sum of the corrected prediction probabilities obtained by matching the predicted positive sample prediction probabilities corresponding to each image frame in the training samples with the confidence as the weight;

[0121] Call the second loss function, calculate the cross-entropy loss of the confidence of the quality score to obtain the second loss value. The second loss function is exemplified as follows:

[0122]

[0123] Where, Loss conf is the second loss value, n is the total number of image frames in the training samples, λ is the adjustment parameter, p c is the corrected prediction probability obtained by matching the predicted positive sample prediction probabilities corresponding to each image frame in the training samples with the confidence as the weight, exemplified as follows:

[0124] p c = c * (1 - |y out - y gt |)

[0125] Where, c is the confidence corresponding to the training samples, y out is the quality score predicted by the model for a corresponding single image frame, y gt is the subjective quality score for a corresponding single image frame.

[0126] As disclosed above, those skilled in the art should know that for each image frame of the training samples, taking 1 - |y out - y gt | as the positive sample output probability corresponding to the model, and then obtaining the corrected prediction probability p cIn addition, when calculating the cross-entropy loss corresponding to the training samples using the second loss function of the exemplary example, in order to prevent the model from learning that the Loss obtained when c = 1 is the optimal, resulting in incorrect confidence, accordingly, a regularization term regarding C is added. It is recommended that the λ in Loss conf is 0.1. conf In Loss

[0127] Step S1023: Calculate the sum of the first loss value and the second loss value to obtain the model loss value corresponding to the training sample.

[0128] According to the exemplary examples corresponding to the first and second loss functions in steps S1021 - 1022, calculate the sum of the first loss value Loss mos and the second loss value Loss conf to obtain the model loss value corresponding to the training sample.

[0129] In this embodiment, the first and second loss values corresponding to the quality score and its corresponding confidence obtained by the model predicting the training sample are calculated through the first and second loss functions, and then the first and second loss values are added as the loss value of the model. Compared with using only the first loss value as the loss value of the model, there is a correction by the second loss value, which controls the inaccuracy of the first loss value caused by insufficient model robustness and affects the accuracy of model prediction, thereby improving the overall accuracy, robustness, and reliability of the model.

[0130] Please refer to Figure 10 , a video quality assessment device provided to meet one of the purposes of this application is a functional embodiment of the video quality assessment method of this application. The device includes an image processing module 1100, a feature extraction module 1200, a weight extraction module 1300, and a score output module 1400, where: The image processing module 1100 is used to obtain a continuous plurality of image frames in the video stream and format them into standardized data; the feature extraction module 1200 is configured to extract multi-channel image feature information corresponding to the standardized data using the image feature extraction network in the video quality assessment model pre-trained to a convergent state; the weight extraction module 1300 is configured to perform pooling operations on the image feature information in different ways by the feature manifestation network of the assessment model to enhance the weights of the significant features between channels in the image feature information and obtain significant feature information; the score output module 1400 is configured to perform classification mapping on the significant feature information by the classification network of the assessment model to obtain the quality scores corresponding to the plurality of image frames.

[0131] In a further embodiment, the weight extraction module 1300 includes: two-way pooling sub-modules, configured to divide the image feature information into two paths and respectively input them into the pooling layers of two branches of the feature manifestation network to perform different types of pooling operations, so as to obtain corresponding sampled feature information; a weight-enhanced feature information acquisition sub-module, configured to respectively apply attention layers to the corresponding sampled feature information in the two branches to extract weights, so as to enhance the numerical gap between the significant features and the non-significant features therein, and obtain corresponding weight-enhanced feature information; a normalization sub-module, configured to splice the weight-enhanced feature information of the two branches and then perform normalization to obtain significant feature information.

[0132] In a further embodiment, the scoring output module 1400 includes: a global pooling sub-module, configured to input the significant feature information into a global pooling layer to perform a pooling operation, so as to obtain corresponding pooled feature information; a quality score acquisition sub-module, configured to output the pooled feature information to a first fully-connected layer for classification mapping to obtain a predicted quality score; a confidence level acquisition sub-module, configured to output the pooled feature information to a second fully-connected layer for classification mapping to obtain the confidence level corresponding to the quality score; a quality score determination sub-module, configured to determine the quality score with a confidence level higher than a preset threshold as the quality score corresponding to the multiple image frames.

[0133] In an extended embodiment, after the scoring output module 1400, there is further included: an encoding frame rate adjustment module, configured to adjust the encoding frame rate of the encoder of the video stream according to the quality score.

[0134] In a further embodiment, before the image processing module 1100, there is further included: an iterative training sub-module, configured to perform iterative training on the video quality assessment model by using training samples in a preset dataset until it converges.

[0135] In a further embodiment, the iterative training sub-module includes: an implementation training unit, configured to call a single training sample from a preset dataset to perform training on the video quality assessment model, where the training sample includes a plurality of consecutive image frames constituting video data and is labeled with a corresponding subjective quality score; a loss value unit, configured to calculate a model loss value based on the subjective quality score, where the model loss value is the sum of a first loss value corresponding to the quality score obtained by the assessment model and a second loss value of the confidence level of the quality score; an iterative training judgment unit, configured to judge whether the model loss value reaches a preset threshold. When it reaches the preset threshold, it is determined that the assessment model has converged and the training is terminated; otherwise, gradient update is performed on the assessment model, and iterative training is continued by using the next training sample.

[0136] In a further embodiment, the loss value unit includes: a first loss value sub-unit, configured to calculate a loss value between a predicted quality score corresponding to each image frame in the training sample and a pre-annotated subjective quality score, and obtain an average absolute error of each of the loss values as a first loss value; a second loss value sub-unit, configured to calculate a cross-entropy loss of the confidence of the quality score to obtain a second loss value, where the second loss value is an average value of the logarithm sum of the corrected prediction probabilities obtained by matching the predicted positive sample probabilities corresponding to each image frame in the training sample with the confidence as a weight; and a model loss value sub-unit, configured to calculate a sum value of the first loss value and the second loss value to obtain a model loss value corresponding to the training sample.

[0137] To solve the above technical problems, an embodiment of the present application further provides a computer device. As Figure 11 shown, it is a schematic internal structure diagram of the computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database may store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a video quality assessment method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device may store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the video quality assessment method of the present application. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art can understand that Figure 11 the structure shown in

[0138] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. Figure 10 In this embodiment, the processor is used to execute the specific functions of each module and its sub-modules in

[0139] The memory stores program codes and various types of data required to execute the above modules or sub-modules. The network interface is used for data transmission between a user terminal and a server. The memory in this embodiment stores program codes and data required to execute all modules / sub-modules in the video quality assessment device of the present application. The server can call the program codes and data of the server to execute the functions of all sub-modules. The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the video quality assessment method according to any embodiment of the present application.

[0140] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments of the present application can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0141] In summary, the video quality assessment model of the present application has a simple structure and is easy to implement and train. It is convenient to be deployed on the media stream server corresponding to the network live broadcast of the e-commerce platform to provide a service for evaluating the quality of the live video stream, so that it can accurately evaluate a reliable quality score and adjust the picture quality of the live video stream accordingly without affecting the user's viewing experience, save the bandwidth for transmitting the video stream, stably play the live video stream, ensure the smoothness and real-time nature of the live video stream playback, improve the user experience, and increase user stickiness.

[0142] Secondly, the video quality assessment model can output a reliable confidence level, based on which it can be judged whether the quality score predicted by the model is valid, avoiding excessive prediction errors, improving the robustness of the model, and at the same time being able to correct the quality score predicted by the model according to the confidence level, improving the accuracy and reliability of the model.

[0143] In addition, the quality score carrying the confidence level output by the video quality assessment model can accurately evaluate the quality of the video stream, evaluate a corresponding reliable quality score, and the model has better robustness, effectively coping with various force majeure situations of real video streams, and can achieve overall good benefits.

[0144] Those skilled in the art of the present technology can understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in the present application can be alternated, changed, combined, or deleted. Further, the other steps, measures, and solutions in the various operations, methods, and processes discussed in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and solutions in the prior art that are the same as those disclosed in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted.

[0145] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A video quality assessment method, characterized in that, Including the following steps: Obtain a continuous plurality of image frames in the video stream and format them into standardized data; Extract multi-channel image feature information corresponding to the standardized data by using the image feature extraction network in the video quality assessment model that has been pre-trained to a convergent state; The feature manifestation network of the assessment model performs pooling operations on the image feature information in different ways to enhance the weights of the significant features between channels in the image feature information and obtain significant feature information; The classification network of the assessment model performs classification mapping according to the significant feature information to obtain the quality scores corresponding to the plurality of image frames; Among them, before the step of obtaining a continuous plurality of image frames in the video stream, it includes: Performing iterative training on the video quality assessment model by using the training samples in a preset dataset and training it to a convergent state, including: Calling a single training sample from the preset dataset to perform training on the video quality assessment model, where the training sample includes a continuous plurality of image frames constituting video data and is labeled with the corresponding subjective quality score; Calculating the model loss value based on the subjective quality score, where the model loss value is the sum value of the first loss value corresponding to the quality score obtained by the assessment model and the second loss value of the confidence of the quality score; Judging whether the model loss value reaches a preset threshold. When it reaches the preset threshold, it is determined that the assessment model has converged and the training is terminated; otherwise, gradient update is performed on the assessment model and iterative training is continued by using the next training sample; The calculating the model loss value based on the subjective quality score includes: Calculating the loss value between the predicted quality score corresponding to each image frame in the training sample and the pre-labeled subjective quality score, and obtaining the mean absolute error of each loss value as the first loss value; Calculating the cross-entropy loss of the confidence of the quality score to obtain the second loss value. The second loss value is the negative mean value of the logarithm of the corrected prediction probability obtained after the positive sample prediction probability corresponding to each image frame in the training sample matches the confidence as the weight, minus the regularization term corresponding to the confidence; Calculating the sum value of the first loss value and the second loss value to obtain the model loss value corresponding to the training sample.

2. The video quality assessment method according to claim 1, wherein The feature manifestation network of the assessment model performs pooling operations on the image feature information in different ways to enhance the weights of the significant features between channels in the image feature information and obtain significant feature information, including the following steps: Dividing the image feature information into two paths and respectively inputting them into the pooling layers of the two branches of the feature manifestation network to perform different types of pooling operations to obtain corresponding sampled feature information; In the two branches, respectively applying an attention layer to extract weights for the corresponding sampled feature information to enhance the numerical gap between the significant features and the non-significant features therein and obtain corresponding weighted feature information; Concatenating the weighted feature information of the two branches and performing normalization to obtain significant feature information.

3. The video quality assessment method according to claim 1, wherein The classification network of the evaluation model performs classification mapping based on the significant feature information to obtain the quality scores corresponding to the multiple image frames, including the following steps: Input the significant feature information into the global pooling layer to perform pooling operation, and obtain the corresponding pooled feature information; Output the pooled feature information to the first fully connected layer for classification mapping to obtain the predicted quality score; Output the pooled feature information to the second fully connected layer for classification mapping to obtain the confidence level corresponding to the quality score; Determine the quality scores with confidence levels higher than the preset threshold as the quality scores corresponding to the multiple image frames.

4. The video quality assessment method according to any one of claims 1 to 3, characterized in that, After the step of obtaining the quality scores corresponding to the multiple image frames, the following steps are included: Adjust the encoding frame rate of the encoder of the video stream according to the quality scores.

5. A video quality assessment device, characterized in that, Including: An image processing module, configured to obtain multiple consecutive image frames in the video stream and format them into standardized data; A feature extraction module, configured to extract multi-channel image feature information corresponding to the standardized data by using an image feature extraction network in a video quality evaluation model pre-trained to a convergent state; A weight extraction module, configured to perform pooling operations on the image feature information in different ways by the feature manifestation network of the evaluation model to enhance the weights of the significant features between channels in the image feature information and obtain significant feature information; A score output module, configured to perform classification mapping based on the significant feature information by the classification network of the evaluation model to obtain the quality scores corresponding to the multiple image frames; Among them, before the step of obtaining multiple consecutive image frames in the video stream, it includes: Performing iterative training on the video quality evaluation model by using training samples in a preset dataset and training it to a convergent state, including: Calling a single training sample from the preset dataset to perform training on the video quality evaluation model, where the training sample includes multiple consecutive image frames constituting video data and is labeled with corresponding subjective quality scores; Calculating the model loss value based on the subjective quality score, where the model loss value is the sum of the first loss value corresponding to the quality score obtained by the evaluation model and the second loss value of the confidence level of the quality score; Judging whether the model loss value reaches a preset threshold. When it reaches the preset threshold, it is determined that the evaluation model has converged and the training is terminated; otherwise, gradient update is performed on the evaluation model and iterative training is continued by using the next training sample; The calculating the model loss value based on the subjective quality score includes: Calculating the loss value between the predicted quality score corresponding to each image frame in the training sample and the pre-labeled subjective quality score, and obtaining the mean absolute error of each loss value as the first loss value; Calculating the cross-entropy loss of the confidence level of the quality score to obtain the second loss value, where the second loss value is the negative mean of the logarithm of the corrected prediction probability obtained after the positive sample prediction probability corresponding to each image frame in the training sample matches the confidence level as the weight minus the regularization term corresponding to the confidence level; Calculate the sum of the first loss value and the second loss value to obtain the model loss value corresponding to the training sample.

6. A computer device, comprising a central processing unit and a memory, characterized in that, The central processing unit is configured to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, It stores a computer program implemented according to the method according to any one of claims 1 to 4 in the form of computer-readable instructions. When the computer program is called and run by a computer, it executes the steps included in the corresponding method.