Video quality evaluation model training method and device, quality evaluation method and device and medium
By training a video quality assessment model and utilizing a distortion feature extraction module and a prediction network to evaluate video distortion, the problem of assessing video quality loss was solved, enabling a comprehensive assessment and localization of video quality issues, thereby improving user experience and platform health.
Patent Information
- Application Number
- CN202511134446.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies struggle to effectively assess and pinpoint the image quality loss introduced during the processing and transcoding of user-generated content videos, impacting user experience and the health of the platform ecosystem.
By training a video quality assessment model using a distortion feature extraction module, a gate structure, and a prediction network, the model extracts video distortion features, evaluates distortion weights and degrees, generates quality scores, and trains the model to improve assessment accuracy.
It enables a comprehensive assessment and identification of video quality issues, guiding subsequent video optimization efforts to improve user experience and the health of the platform ecosystem.
Smart Images

Figure CN120976834A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, in particular, to a video quality evaluation model training method, a video quality evaluation method, a device, an electronic device, a computer program product and a storage medium. BACKGROUND
[0002] With the vigorous development of user generated content (UGC) video platforms, there are significant differences in the source quality of videos uploaded by a large number of users. At the same time, in order to adapt to different terminals and network environments, the platform needs to perform complex processing and transcoding link operations on the videos uploaded by users, which inevitably introduces additional picture quality loss, resulting in a decrease in video definition. These definition problems directly affect the consumption experience and satisfaction of users on the platform, and further challenge the consumption ecological health of the platform. Therefore, effectively monitoring the definition status and trend of videos at each link from production (user upload), processing enhancement, transcoding to final consumption is crucial to guarantee user experience and optimize the platform ecology. Based on this, there is an urgent need in the art for a quality evaluation method that can effectively evaluate video quality and locate quality problems. SUMMARY
[0003] The present disclosure provides a video quality evaluation model training method, a video quality evaluation method, a device, an electronic device, a computer program product and a storage medium.
[0004] According to a first aspect of the embodiments of the present disclosure, a video quality evaluation model training method is provided, comprising: obtaining a training data set comprising a plurality of video quality training samples; each of the training samples at least comprises: a video sample and a sample quality score; the sample quality score is used to represent the quality score of the corresponding video sample; performing feature extraction on the video sample through a distortion feature extraction module to obtain a video distortion feature; performing distortion weight evaluation on the video distortion feature through a to-be-trained gate structure to obtain a distortion weight value; performing distortion degree evaluation on the video distortion feature through a prediction network to obtain a distortion score value; determining a video quality score according to the distortion weight value and the distortion score value; generating a quality score loss value according to the video quality score and the sample quality score; the quality score loss value is used to represent the difference degree between the video quality score and the sample quality score; training a to-be-trained video quality evaluation model according to the quality score loss value to obtain a video quality evaluation model; the video quality evaluation model at least comprises: the distortion feature extraction module, the gate structure and the prediction network.
[0005] In some example embodiments of the present disclosure, the feature extraction on the video sample by the distortion feature extraction module to obtain video distortion features comprises: feature extraction on the video sample by at least one distortion sub-module to obtain at least one distortion sub-feature; each of the distortion sub-modules corresponds to a distortion type; wherein the distortion feature extraction module comprises the at least one distortion sub-module; and the video distortion features comprise the at least one distortion sub-feature.
[0006] In some example embodiments of the present disclosure, the training sample further comprises a first sample distortion sub-score corresponding to a first distortion type; the method further comprises: feature extraction on the video sample by a first distortion sub-module to be trained to obtain a first distortion sub-feature; distortion degree evaluation on the first distortion sub-feature by the prediction network to obtain a first distortion sub-score; generation of a first distortion sub-score loss value according to the first distortion sub-score and the first sample distortion sub-score; and training of the first distortion sub-module to be trained according to the first distortion sub-score loss value to obtain a first distortion sub-module; wherein the first distortion sub-module belongs to the at least one distortion sub-module.
[0007] In some example embodiments of the present disclosure, the distortion weight value comprises at least one distortion sub-weight value corresponding to the distortion type; the distortion score value comprises at least one distortion sub-score value corresponding to the distortion type; and the determination of the video quality score according to the distortion weight value and the distortion score value comprises: obtaining a corresponding video quality sub-score according to the corresponding distortion sub-weight value and distortion sub-score value based on the distortion type; and determining the video quality score according to each of the video quality sub-scores.
[0008] In some example embodiments of the present disclosure, the determination of the video quality score according to the distortion weight value and the distortion score value comprises: fusion perception of the video distortion features, the distortion weight value and the distortion score value by a multi-layer perception to obtain the video quality score.
[0009] According to a second aspect of the embodiments of the present disclosure, a video quality evaluation method is provided, comprising: obtaining a video to be evaluated; inputting the video to be evaluated into a video quality evaluation model; the video quality evaluation model is obtained by training according to any one of the video quality evaluation model training methods; and evaluating the video to be evaluated by the video quality evaluation model to obtain a video quality evaluation result.
[0010] In some example embodiments of the present disclosure, the evaluating the to-be-evaluated video by the video quality evaluation model to obtain a video quality evaluation result comprises: performing feature extraction on the to-be-evaluated video by a distortion feature extraction module to obtain video distortion features; performing distortion weight evaluation on the video distortion features by a gate structure to obtain distortion weight values; performing distortion degree evaluation on the video distortion features by a prediction network to obtain distortion score values; and generating the video quality evaluation result according to the distortion weight values and the distortion score values.
[0011] In some example embodiments of the present disclosure, the performing feature extraction on the to-be-evaluated video by the distortion feature extraction module to obtain video distortion features comprises: performing feature extraction on the to-be-evaluated video by at least one distortion submodule to obtain at least one distortion sub-feature; each of the distortion submodules corresponds to a distortion type; and the distortion feature extraction module comprises the at least one distortion submodule, and the video distortion features comprise the at least one distortion sub-feature.
[0012] In some example embodiments of the present disclosure, the distortion weight values comprise at least one distortion sub-weight value corresponding to the distortion type, the distortion score values comprise at least one distortion sub-score value corresponding to the distortion type, and the generating the video quality evaluation result according to the distortion weight values and the distortion score values comprises: obtaining a corresponding video quality sub-score based on the distortion type according to the corresponding distortion sub-weight value and the distortion sub-score value, and determining a video quality score according to each of the video quality sub-scores.
[0013] In some example embodiments of the present disclosure, the generating the video quality evaluation result according to the distortion weight values and the distortion score values comprises: performing fusion perception on the video distortion features, the distortion weight values and the distortion score values by a multi-layer perception machine to obtain a video quality score.
[0014] According to a third aspect of the embodiments of the present disclosure, a video quality evaluation model training apparatus is provided, comprising: a sample acquisition module configured to acquire a training data set comprising a plurality of video quality training samples; each of the training samples comprises at least: a video sample and a sample quality score; the sample quality score is used to represent the quality score of the corresponding video sample; a distortion feature extraction module configured to perform feature extraction on the video sample through the distortion feature extraction module to obtain a video distortion feature; a gate structure module configured to perform distortion weight evaluation on the video distortion feature through a to-be-trained gate structure to obtain a distortion weight value; a prediction network module configured to perform distortion degree evaluation on the video distortion feature through a prediction network to obtain a distortion score value; a video quality score module configured to determine a video quality score according to the distortion weight value and the distortion score value; a loss value module configured to generate a quality score loss value according to the video quality score and the sample quality score; the quality score loss value is used to represent the difference degree between the video quality score and the sample quality score; a model training module configured to train a to-be-trained video quality evaluation model according to the quality score loss value to obtain a video quality evaluation model; the video quality evaluation model comprises at least: the distortion feature extraction module, the gate structure and the prediction network.
[0015] According to a fourth aspect of the embodiments of the present disclosure, a video quality evaluation apparatus is provided, comprising: an evaluation request acquisition module configured to acquire a to-be-evaluated video; an evaluation request input module configured to input the to-be-evaluated video into a video quality evaluation model; the video quality evaluation model is obtained by training through any one of the video quality evaluation model training methods; a quality evaluation module configured to evaluate the to-be-evaluated video through the video quality evaluation model to obtain a video quality evaluation result.
[0016] According to a fifth aspect of the embodiments of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing executable instructions of the processor; wherein the processor is configured to execute the executable instructions to implement any one of the video quality evaluation model training methods or any one of the video quality evaluation methods.
[0017] According to a sixth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, when instructions in the computer readable storage medium are executed by a processor of an electronic device, the electronic device can execute any one of the video quality evaluation model training methods or any one of the video quality evaluation methods.
[0018] According to a seventh aspect of the embodiments of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, performs any one of the video quality evaluation model training method or any one of the video quality evaluation method.
[0019] The video quality evaluation model training method provided by the embodiments of the present disclosure extracts features of a video sample through a distortion feature extraction module to obtain video distortion features; evaluates distortion weights of the video distortion features through a gate structure to obtain distortion weight values; evaluates distortion degrees of the video distortion features through a prediction network to obtain distortion score values; determines a video quality score according to the distortion weight values and the distortion score values; generates a quality score loss value according to the video quality score and a sample quality score; and trains a to-be-trained video quality evaluation model according to the quality score loss value to obtain a video quality evaluation model. The video quality evaluation model trained by the method evaluates the influence of each distortion type on the quality problem of the video sample from different dimensions of influence weights and distortion degrees, so that the video quality problem can be more comprehensively evaluated and positioned to guide subsequent video optimization.
[0020] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting to the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0021] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0022] Figure 1 A schematic diagram of an exemplary system architecture to which the method of the embodiments of the present disclosure can be applied is shown.
[0023] Figure 2 is a flowchart of a video quality evaluation model training method according to an exemplary embodiment.
[0024] Figure 3 is a flowchart of a distortion sub-module training method according to an example.
[0025] Figure 4 is a flowchart of a video quality evaluation method according to an example.
[0026] Figure 5 is a flowchart of a video quality evaluation model evaluation process according to an example.
[0027] Figure 6 is a schematic diagram of a video quality assessment model according to an example.
[0028] Figure 7 is a block diagram of a video quality assessment model training apparatus according to an example embodiment.
[0029] Figure 8 is a block diagram of a video quality assessment apparatus according to an example embodiment.
[0030] Figure 9 is a structural schematic diagram of an electronic device suitable for implementing the example embodiments of the present disclosure according to an example embodiment. DETAILED DESCRIPTION
[0031] Example embodiments now will be described more fully hereinafter with reference to the accompanying drawings. Example embodiments, however, can be implemented in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of example embodiments to those skilled in the art. Like reference numerals refer to like elements throughout the several views and the description of the figures.
[0032] The features, structures, or characteristics described in connection with the present disclosure can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the present disclosure. One skilled in the relevant art will recognize, however, that the techniques described herein can be practiced without one or more of the specific details, or with other methods, components, materials, and so forth. In other instances, well-known structures, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0033] The accompanying drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification. The drawings illustrate examples of the present disclosure and, together with the description, serve to explain principles of the present disclosure. In the drawings:
[0034] The flow diagrams illustrated in the drawings, are examples only and are not necessarily meant to include all of the steps and / or actions, nor are they necessarily meant to be performed in the order illustrated. For example, some steps can be combined or partially combined, and the order of execution can be changed based on the actual implementation.
[0035] In this specification, the terms "one", "a", "an", "said", and "the" are used to indicate that there is at least one of the elements / components / etc.; the terms "comprise", "include" and "have" are used to indicate an open-ended inclusion in such a way that additional elements / components / etc. can be present in addition to the listed elements / components / etc.; the terms "first", "second" and "third" and the like are used only as labels, not as a numerical limitation of their objects.
[0036] Figure 1 A schematic diagram showing an exemplary system architecture to which the method of the embodiments of the present disclosure can be applied is shown.
[0037] As Figure 1 shown, the system architecture can include a server 101, a network 102, a terminal device 103, a terminal device 104, and a terminal device 105. The network 102 is a medium to provide a communication link between the terminal device 103, the terminal device 104, or the terminal device 105 and the server 101. The network 102 can include various connection types, such as wired, wireless communication links, or fiber optic cables, and the like.
[0038] The server 101 can be a server that provides various services, such as a background management server that provides support for devices operated by a user using the terminal device 103, the terminal device 104, or the terminal device 105. The background management server can analyze and process received request data and the like, and feed back the processing results to the terminal device 103, the terminal device 104, or the terminal device 105.
[0039] The terminal device 103, the terminal device 104, and the terminal device 105 can be a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a wearable smart device, a virtual reality device, an augmented reality device, and the like, but are not limited thereto.
[0040] It should be understood that Figure 1 the number of terminal devices 103, terminal devices 104, terminal devices 105, networks 102, and servers 101 in
[0041] In the following, the various steps of the method in the example embodiments of the present disclosure will be described in more detail in conjunction with the accompanying drawings and examples.
[0042] Figure 2 is a flowchart of a video quality assessment model training method according to an example embodiment. Figure 6is a schematic diagram of a video quality evaluation model according to an example. Figure 2 The method provided by the embodiments can be executed by any electronic device, for example, the terminal device in the foregoing Figure 1 , or the server in the foregoing Figure 1 , or the terminal device and the server in the foregoing Figure 1 jointly execute, but the disclosure does not limit this.
[0043] In step S210, a training data set including a plurality of video quality training samples is obtained; each of the training samples at least includes a video sample and a sample quality score; the sample quality score is used to represent the quality score of the corresponding video sample.
[0044] In the embodiments of the disclosure, a training data set including a plurality of video quality training samples is obtained. Each of the training samples at least includes a video sample and a sample quality score. The sample quality score is used to represent the quality score of the corresponding video sample as a label value in model training. The sample quality score can be manually scored or obtained by other quality evaluation modules.
[0045] In the exemplary embodiments, the quality problem of the video can be caused by multiple reasons. In order to attribute the video quality problem caused by various reasons in the video quality evaluation process, the subsequent video optimization is guided. The disclosure pre-summarizes and forms several distortion types causing the video quality problem. For example, compression distortion, blur, sharpening, noise, exposure, etc. In each training sample, not only the sample quality score representing the overall quality of the video sample is included, but also the sample distortion sub-score corresponding to each of the above distortion types is further included. The sample distortion sub-score is used to represent the quality score of the video sample in the corresponding distortion type dimension. Similarly, the sample distortion sub-score can also be manually scored or obtained by other quality evaluation sub-modules.
[0046] Table 1 is an example of the sample quality score and the sample distortion sub-score of the training sample. As shown in the table, the "overall clarity" is the sample quality score, which represents the sample quality score of the overall quality of the video sample. The "compression distortion, blur, sharpening, noise, exposure" is the sample distortion sub-score, which respectively represents the quality score of the video sample in the distortion type dimension of compression distortion, blur, sharpening, noise, exposure, etc. For each score, a scoring method of 1-5 is used for labeling.
[0047] It should be pointed out that each of the training samples should at least include the sample quality score representing the overall quality of the video sample. However, for the sample distortion sub-score, only the sample distortion sub-score corresponding to part of the distortion type dimension can be included, or there is no sample distortion sub-score, which should be considered within the protection scope of the disclosure.
[0048] Clarity issues Labelling criteria Overall clarity Scored 1-5 for severity Compression distortion Scored 1-5 for severity Blurring Scored 1-5 for severity Sharpening Scored 1-5 for severity Noise Scored 1-5 for severity Exposure Scored 1-5 for severity
[0049] Table 1
[0050] In step S220, the video sample is subjected to feature extraction by a distortion feature extraction module to obtain video distortion features.
[0051] In the embodiments of the present disclosure, as shown in Figure 6 , a video sample is input into a distortion feature extraction module. The distortion feature extraction module can employ a neural network and is pre-trained to extract video distortion features related to video quality from the video sample. The distortion feature extraction module can employ a network structure such as ResNet, MobileNet, or ViT. The distortion feature extraction module can be pre-trained or can employ an existing video feature extraction module to achieve the related functions, which are not limited in the present disclosure.
[0052] In exemplary embodiments, the distortion feature extraction module can include several distortion sub-modules. Each distortion sub-module corresponds to one of the aforementioned distortion types and is used to extract distortion sub-features of the corresponding distortion type dimension from the video sample. The distortion feature extraction module extracts features from the video sample through the above-mentioned distortion sub-modules to obtain corresponding distortion sub-features. The distortion sub-features are used to represent the distortion features of the video sample in the corresponding distortion type dimension. The aforementioned video distortion features include the distortion sub-features.
[0053] As shown in Figure 6 , the distortion feature extraction module includes at least a blur sub-module N blur , a sharpening sub-module N sharp , a noise sub-module N noise , an exposure sub-module N expose , and a compression distortion sub-module N blocky , which correspond to the extraction of distortion features in the dimensions of blur, sharpening, noise, exposure, and compression distortion, respectively. The above-mentioned distortion sub-modules extract features from the video sample to obtain blur sub-features f blur , sharpening sub-features f sharp , noise sub-features f noise , exposure sub-features f expose , and compression distortion sub-features f blocky , respectively. The video distortion features include at least the blur sub-features f blur , the sharpening sub-features f sharp , the noise sub-features f noise , the exposure sub-features f expose , and the compression distortion sub-features f blockyThe distortion sub-modules corresponding to different distortion types can be used to extract the distortion sub-features of the corresponding distortion types, so as to accurately locate the video quality problem and guide the subsequent video optimization and provide a data basis.
[0054] In step S230, the distortion weight value is obtained by evaluating the video distortion feature through the to-be-trained gating structure.
[0055] In the embodiments of the present disclosure, as shown in Figure 6 The video distortion feature is input into the to-be-trained gating structure (Gating Module, GM). The gating structure is used to evaluate the distortion weight of each distortion type in the quality problem of the video sample according to the video distortion feature, and obtain a distortion weight value. The distortion weight value is used to represent the influence weight of each distortion type on the overall quality of the video sample. The gating structure can adopt an MLP, a Transformer, an LSTM, or the like.
[0056] In the example embodiments, the distortion weight value can include at least one distortion sub-weight value corresponding to each distortion type. The distortion sub-weight value is used to represent the influence weight of the corresponding distortion type dimension on the overall quality of the video sample.
[0057] As shown in Figure 6 After the distortion weight evaluation by the gating structure, the distortion sub-weight values a1, a2, a3, a4,..., aN corresponding to the blur, sharpening, noise, exposure, compression distortion, and the like are obtained. m .
[0058] In step S240, the distortion degree evaluation of the video distortion feature is performed through the prediction network, and a distortion score value is obtained.
[0059] In the embodiments of the present disclosure, as shown in Figure 6 The video distortion feature is input into the prediction network (Predict Network, PN). The prediction network is a model architecture based on deep learning, which is used to evaluate the distortion degree of each distortion type in the quality problem of the video sample according to the video distortion feature, and obtain a distortion score value. The distortion score value is used to represent the distortion degree of each distortion type on the overall quality of the video sample. The prediction network is a pre-trained model.
[0060] As shown in Figure 6 After the distortion degree evaluation by the prediction network, the distortion sub-score values Sblur, Sharpening, Snoise, SExposure, SCompression corresponding to the blur, sharpening, noise, exposure, compression distortion, and the like are obtained. blur sharp , and the like.noise , an exposure sub-score value S expose and a compression distortion sub-score value S blocky .
[0061] It should be pointed out that the distortion weight value obtained by the door structure represents the influence weight corresponding to each distortion type, and the distortion score value obtained by the prediction network represents the distortion degree corresponding to each distortion type, which are not the same. As shown in FIG. 8, wherein the pie chart is used to represent the distortion weight value corresponding to each distortion type, and the score value (Blur Score, Sharp Score, Noise Score, Expose Score, Blocky Score) above the pie chart is used to represent the distortion score value corresponding to each distortion type. The present disclosure can more comprehensively evaluate and locate the video quality problem by evaluating the influence of each distortion type on the video sample quality problem from different dimensions of influence weight and distortion degree. Figure 6
[0062] In step S250, a video quality score is determined according to the distortion weight value and the distortion score value.
[0063] In the embodiment of the present disclosure, the influence of each distortion type on the video sample quality problem is evaluated from different dimensions of influence weight and distortion degree according to the distortion weight value and the distortion score value. According to the distortion weight value and the distortion score value, the video quality score of the video sample is determined comprehensively from different angles of influence weight and distortion degree. The video quality score is used to represent the quality score of the overall quality of the video sample by the video quality evaluation model, which corresponds to the sample quality score of the training sample.
[0064] In the exemplary embodiment, as shown in FIG. 9, wherein the quality score (Quality Score) corresponds to the video quality score and the sample quality score, and represents the quality score of the overall quality of the video sample. Figure 6
[0065] It should be pointed out that the video quality score can be determined by various calculation methods according to the distortion weight value and the distortion score value, and the present disclosure does not limit the specific calculation method.
[0066] In step S260, a quality score loss value is generated according to the video quality score and the sample quality score; the quality score loss value is used to represent the difference between the video quality score and the sample quality score.
[0067] In the embodiments of the present disclosure, the video quality score is taken as a predicted value, the sample quality score is taken as a true value, and a quality score loss value is generated according to the video quality score and the sample quality score. The quality score loss value is used to represent the difference between the video quality score and the sample quality score. Through the quality score loss value, the video quality evaluation model can be trained so that the score of the model is closer and closer to the sample quality score of the training sample.
[0068] It should be pointed out that there can be various ways to design the related quality score loss value, and those skilled in the art should consider that the design schemes of the score loss function provided by the existing technical means according to actual needs are within the protection scope of the present disclosure.
[0069] In step S270, the video quality evaluation model to be trained is trained according to the quality score loss value, and a video quality evaluation model is obtained; the video quality evaluation model at least includes the distortion feature extraction module, the gate structure and the prediction network.
[0070] In the embodiments of the present disclosure, the video quality evaluation model is trained according to the quality score loss value generated in the foregoing, so that the video quality score output by the model is closer and closer to the sample quality score of the training sample. The video quality evaluation model includes but is not limited to the gate structure and / or MLP in the video quality evaluation model. When the video quality evaluation model meets a preset convergence condition, the video quality evaluation model after parameter adjustment is a target video quality evaluation model.
[0071] In the exemplary embodiments, as shown in Figure 6 The video quality evaluation model at least includes the distortion feature extraction module, the gate structure and the prediction network. The model parameters of the distortion feature extraction module and the prediction network can be determined by pre-training. In the training process, the model parameters in the distortion feature extraction module and the prediction network are fixed, and only the gate structure to be trained is adjusted. That is, the gate structure is adjusted based on the quality score loss value, so as to adjust the distortion sub-weight values corresponding to each distortion type, and then the corresponding video quality score is affected based on the distortion weight values, so that the video quality score is closer and closer to the sample quality score.
[0072] The video quality evaluation model training method provided by the embodiments of the present disclosure extracts features of a video sample through a distortion feature extraction module to obtain video distortion features; evaluates distortion weights of the video distortion features through a gate structure to obtain distortion weight values; evaluates distortion degrees of the video distortion features through a prediction network to obtain distortion score values; determines a video quality score according to the distortion weight values and the distortion score values; generates a quality score loss value according to the video quality score and a sample quality score; and trains a to-be-trained video quality evaluation model according to the quality score loss value to obtain a video quality evaluation model. The video quality evaluation model trained by the method evaluates the influence of each distortion type on the quality of a video sample from different dimensions of influence weights and distortion degrees, so that the video quality problem can be more comprehensively evaluated and positioned to guide subsequent video optimization.
[0073] Figure 3 is a flowchart of a distortion submodule training method according to an example. As shown in Figure 3 The pre-training process of the aforementioned distortion feature extraction module can include the following steps in the embodiments of the present disclosure.
[0074] In step S310, a first distortion submodule to be trained extracts features of the video sample to obtain first distortion sub-features.
[0075] In the embodiments of the present disclosure, the distortion feature extraction module can include a plurality of distortion submodules. Each distortion submodule corresponds to one of the aforementioned distortion types and is used to extract distortion sub-features of the corresponding distortion type dimension from a video sample. Before training the gate structure, each distortion submodule can be pre-trained to adjust parameters of each distortion submodule. Each distortion submodule can be constructed using a neural network, such as a ResNet, MobileNet, ViT, or other network structure. The first distortion submodule is any one of the distortion submodules in the distortion feature extraction module. The first distortion submodule corresponds to the first distortion type.
[0076] The video sample is a training sample including a first sample distortion sub-score corresponding to the first distortion type. By screening the training sample including the first sample distortion sub-score from the training data set, the training data set for training the first distortion submodule is obtained.
[0077] In the embodiments of the present disclosure, the video sample is input into the first distortion submodule. The first distortion submodule is used to extract video distortion features corresponding to the first distortion type from the video sample to obtain first distortion sub-features.
[0078] In step S320, the prediction network evaluates the distortion degree of the first distortion sub-features to obtain a first distortion sub-score.
[0079] In the embodiments of the present disclosure, the first distortion sub-feature is input into a prediction network to obtain a corresponding first distortion sub-score. The first distortion sub-score is used to represent the distortion degree of the first distortion type on the overall quality of the video sample. The prediction network is the prediction network described in step S240, and will not be described here.
[0080] In step S330, a first distortion sub-score loss value is generated according to the first distortion sub-score and the first sample distortion sub-score.
[0081] In the embodiments of the present disclosure, the first distortion sub-score loss value is generated according to the first distortion sub-score and the corresponding first sample distortion sub-score in the training sample. The first distortion sub-score loss value is used to represent the difference between the first distortion sub-score and the first sample distortion sub-score.
[0082] In step S340, the first distortion sub-module to be trained is trained according to the first distortion sub-score loss value to obtain a first distortion sub-module.
[0083] In the embodiments of the present disclosure, the first distortion sub-module to be trained is trained according to the first distortion sub-score loss value generated as described above, so that the distortion sub-score output by the first distortion sub-module is closer and closer to the sample distortion sub-score of the training sample. When the first distortion sub-module meets the preset convergence condition, the first distortion sub-module after parameter adjustment is the target first distortion sub-module.
[0084] In the above manner, each distortion sub-module in the distortion feature extraction module can be trained to obtain each distortion sub-module in the distortion feature extraction module.
[0085] The video quality evaluation model training method provided in the embodiments of the present disclosure sets each distortion sub-module in the distortion feature extraction module for each distortion type, and pre-trains each distortion sub-module according to the corresponding sample distortion sub-score in the training sample. By pre-training each distortion sub-module, distortion sub-features can be extracted for each distortion type, so as to accurately locate the video quality problem, guide subsequent video optimization, and provide a data basis.
[0086] As described above, the video quality score can be determined by a plurality of calculation methods according to the distortion weight value and the distortion score value. Some feasible calculation methods are specifically given below.
[0087] In some example embodiments, the distortion weight value includes at least one distortion sub-weight value corresponding to each distortion type. The distortion score value includes at least one distortion sub-score value corresponding to each distortion type. The video quality score can be determined by the following steps.
[0088] According to the corresponding distortion sub-weight value and the distortion sub-score value, a corresponding video quality sub-score is obtained based on the distortion type;
[0089] According to each video quality sub-score, the video quality score is determined.
[0090] In the embodiments of the present disclosure, according to the corresponding distortion sub-weight value and the distortion sub-score value, a video quality sub-score corresponding to the distortion type is obtained based on the corresponding distortion type. For example, the video quality sub-score is obtained by multiplying the distortion sub-weight value and the distortion sub-score value. According to the video quality sub-score corresponding to each distortion type, the video quality score is determined. For example, each video quality sub-score can be summed or weightedly summed to obtain the video quality score.
[0091] In some example embodiments, as shown in FIG. 1, the video quality score can be determined by the following steps. Figure 6
[0092] The video distortion feature, the distortion weight value and the distortion score value are fused and perceived by a multilayer perceptron to obtain the video quality score.
[0093] In the embodiments of the present disclosure, since the distortion in the video is relatively complex, simply multiplying the distortion sub-weight value and the distortion sub-score value cannot accurately reflect the relationship between the influence weight and the distortion degree. In order to better predict the video quality score, the present disclosure fuses and perceives various parameters by a multilayer perceptron (MLP). The multilayer perceptron is a kind of feedforward neural network, which is composed of a fully connected input layer, a hidden layer and an output layer, and can map multiple input data sets to a single output data set. The video distortion feature, the distortion weight value and the distortion score value are fused and perceived by the multilayer perceptron, so as to better fuse various parameters and obtain the video quality score.
[0094] In example embodiments, the video distortion feature, the distortion weight value and the distortion score value can be respectively input into the multilayer perceptron. The multilayer perceptron fuses and perceives them to obtain the video quality score.
[0095] In example embodiments, the corresponding distortion sub-weight value and the distortion sub-score value can be multiplied first, and then the product obtained and the video distortion feature can be input into the multilayer perceptron. The multilayer perceptron fuses and perceives the product and the video distortion feature to obtain the video quality score.
[0096] The video quality evaluation model training method provided in the embodiments of the present disclosure can determine the video quality score through different calculation methods. The video quality sub-score can be obtained by multiplying the distortion sub-weight value and the distortion sub-score value to meet the demand for calculation efficiency. The video quality score can be obtained by fusing the video distortion feature, the distortion weight value and the distortion score value through the multi-layer perception, so as to better fuse various parameters.
[0097] Figure 4 FIG. 1 is a flowchart of a video quality evaluation method according to an example. Figure 6 FIG. 2 is a schematic diagram of a video quality evaluation model according to an example. Figure 4 The method provided in the embodiments can be executed by any electronic device, for example, the terminal device in the above Figure 1 , or the server in the above Figure 1 , or the terminal device and the server in the above Figure 1 jointly execute, but the present disclosure does not limit this.
[0098] In step S410, a video to be evaluated is obtained.
[0099] In the embodiments of the present disclosure, a video quality evaluation request of a user is obtained. The video quality evaluation request at least includes a video to be evaluated.
[0100] In step S420, the video to be evaluated is input into a video quality evaluation model; the video quality evaluation model is obtained by training the video quality evaluation model training method described above.
[0101] In step S430, the video to be evaluated is evaluated by the video quality evaluation model to obtain a video quality evaluation result.
[0102] In the embodiments of the present disclosure, the video to be evaluated in the video quality evaluation request is input into the video quality evaluation model obtained by training the video quality evaluation model training method described above. The video quality evaluation model at least includes a distortion feature extraction module, a gate structure and a prediction network. The video quality evaluation model evaluates the video to be evaluated through the distortion feature extraction module, the gate structure and the prediction network to obtain a video quality evaluation result.
[0103] In the example embodiments, the video quality evaluation result includes a distortion weight value and a distortion score value.
[0104] In the example embodiments, the video quality evaluation result includes a video quality score.
[0105] In the example embodiments, the video quality evaluation result includes a distortion weight value, a distortion score value and a video quality score.
[0106] Figure 5 is a flowchart of a video quality evaluation model evaluation process according to an example. As shown in Figure 5 the foregoing step S430, the foregoing step S430 can include the following steps.
[0107] In step S510, the distortion feature extraction module is used to extract features of the video to be evaluated to obtain video distortion features.
[0108] In step S520, the video distortion features are evaluated by the gate structure to obtain distortion weight values.
[0109] In step S530, the video distortion features are evaluated by the prediction network to obtain distortion score values.
[0110] In step S540, the video quality evaluation result is generated according to the distortion weight values and the distortion score values.
[0111] In the embodiment of the present disclosure, the steps S510, S520, S530 and S540 in the video quality evaluation model evaluation process are similar to the steps S220, S230, S240 and S250 in the video quality evaluation model training method shown in the foregoing Figure 2 , which will not be repeated here.
[0112] The video quality evaluation method provided by the embodiment of the present disclosure extracts features of the video to be evaluated by the distortion feature extraction module to obtain video distortion features, evaluates the video distortion features by the gate structure to obtain distortion weight values, evaluates the video distortion features by the prediction network to obtain distortion score values, and determines the video quality score according to the distortion weight values and the distortion score values. The video quality evaluation method evaluates the influence of each distortion type on the video quality problem to be evaluated from different dimensions of the influence weight and the distortion degree, so that the video quality problem can be more comprehensively evaluated and positioned to guide the subsequent video optimization.
[0113] In the example embodiment, the step S510 can include the following steps.
[0114] Each of the at least one distortion submodule is used to extract features of the video to be evaluated to obtain at least one distortion sub-feature; each of the distortion submodules corresponds to a distortion type.
[0115] The distortion feature extraction module includes the at least one distortion submodule; and the video distortion features include the at least one distortion sub-feature.
[0116] In the embodiments of the present disclosure, the step is similar to the process of extracting features of the video sample by the distortion feature extraction module to obtain the video distortion feature in the foregoing video quality evaluation model training method, and will not be repeatedly introduced here.
[0117] As described above, the video quality score can be determined by various calculation methods according to the distortion weight value and the distortion score value. Some feasible calculation methods are specifically given below.
[0118] In some example embodiments, the distortion weight value includes at least one distortion sub-weight value corresponding to each distortion type. The distortion score value includes at least one distortion sub-score value corresponding to each distortion type. The step S540 can be determined by the following steps.
[0119] According to the corresponding distortion sub-weight value and distortion sub-score value, a corresponding video quality sub-score is obtained based on the distortion type.
[0120] According to each video quality sub-score, the video quality score is determined.
[0121] In some example embodiments, as shown in Figure 6 The step S540 can be determined by the following steps.
[0122] The video quality score is obtained by fusing perception of the video distortion feature, the distortion weight value and the distortion score value through a multi-layer perception machine.
[0123] In the embodiments of the present disclosure, the calculation method of the video quality score is similar to the calculation method of the video quality score in the foregoing video quality evaluation model training method, and will not be repeatedly introduced here.
[0124] The video quality evaluation method provided by the embodiments of the present disclosure can determine the video quality score by different calculation methods. The video quality sub-score can be obtained by multiplying the distortion sub-weight value and the distortion sub-score value to meet the demand for calculation efficiency. The video quality score can be obtained by fusing perception of the video distortion feature, the distortion weight value and the distortion score value through a multi-layer perception machine, so as to better fuse various parameters.
[0125] The following is an apparatus embodiment of the present disclosure, which can be used to execute the method embodiments of the present disclosure. For details not disclosed in the apparatus embodiments of the present disclosure, please refer to the method embodiments of the present disclosure.
[0126] Figure 7 is a block diagram of a video quality evaluation model training apparatus according to an example embodiment. Refer to Figure 7The device 700 can include a sample acquisition module 710, a distortion feature extraction module 720, a gate structure module 730, a prediction network module 740, a video quality score module 750, a loss value module 760, and a model training module 770.
[0127] The sample acquisition module 710 is configured to acquire a training data set including a plurality of video quality training samples; each of the training samples at least includes a video sample and a sample quality score; the sample quality score is used to represent the quality score of the corresponding video sample.
[0128] The distortion feature extraction module 720 is configured to perform feature extraction on the video sample by the distortion feature extraction module to obtain a video distortion feature.
[0129] The gate structure module 730 is configured to perform distortion weight evaluation on the video distortion feature by a to-be-trained gate structure to obtain a distortion weight value.
[0130] The prediction network module 740 is configured to perform distortion degree evaluation on the video distortion feature by a prediction network to obtain a distortion score value.
[0131] The video quality score module 750 is configured to determine a video quality score according to the distortion weight value and the distortion score value.
[0132] The loss value module 760 is configured to generate a quality score loss value according to the video quality score and the sample quality score; the quality score loss value is used to represent the difference degree between the video quality score and the sample quality score.
[0133] The model training module 770 is configured to train a to-be-trained video quality evaluation model according to the quality score loss value to obtain a video quality evaluation model; the video quality evaluation model at least includes the distortion feature extraction module, the gate structure, and the prediction network.
[0134] In some exemplary embodiments of the present disclosure, the distortion feature extraction module 720 is further configured to perform feature extraction on the video sample by at least one distortion submodule respectively to obtain at least one distortion sub-feature; each of the distortion submodules corresponds to one distortion type; wherein the distortion feature extraction module includes the at least one distortion submodule; and the video distortion feature includes the at least one distortion sub-feature.
[0135] In some example embodiments of the present disclosure, the training sample further comprises a first sample distortion sub-score corresponding to the first distortion type; the distortion feature extraction module 720 is further configured to perform feature extraction on the video sample by a to-be-trained first distortion sub-module respectively, to obtain first distortion sub-features; perform distortion degree evaluation on the first distortion sub-features by the prediction network, to obtain a first distortion sub-score; generate a first distortion sub-score loss value according to the first distortion sub-score and the first sample distortion sub-score; and train the to-be-trained first distortion sub-module according to the first distortion sub-score loss value, to obtain a first distortion sub-module; wherein the first distortion sub-module belongs to the at least one distortion sub-module.
[0136] In some example embodiments of the present disclosure, the distortion weight value comprises at least one distortion sub-weight value corresponding to the distortion type; the distortion score value comprises at least one distortion sub-score value corresponding to the distortion type; and the video quality score module 750 is further configured to obtain a corresponding video quality sub-score according to the corresponding distortion sub-weight value and distortion sub-score value based on the distortion type, and determine the video quality score according to each video quality sub-score.
[0137] In some example embodiments of the present disclosure, the video quality score module 750 is further configured to perform fusion perception on the video distortion features, distortion weight values and distortion score values by a multi-layer perception machine, to obtain the video quality score.
[0138] Figure 8 is a block diagram of a video quality evaluation device according to an example embodiment. Referring to Figure 8 The device 800 can comprise an evaluation request acquisition module 810, an evaluation request input module 820 and a quality evaluation module 830.
[0139] The evaluation request acquisition module 810 is configured to acquire a to-be-evaluated video.
[0140] The evaluation request input module 820 is configured to input the to-be-evaluated video into a video quality evaluation model; the video quality evaluation model is obtained by training through any one of the video quality evaluation model training methods.
[0141] The quality evaluation module 830 is configured to evaluate the to-be-evaluated video by the video quality evaluation model, to obtain a video quality evaluation result.
[0142] In some example embodiments of the present disclosure, the quality evaluation module 830 is further configured to perform feature extraction on the video to be evaluated by a distortion feature extraction module to obtain video distortion features; perform distortion weight evaluation on the video distortion features by a gate structure to obtain distortion weight values; perform distortion degree evaluation on the video distortion features by a prediction network to obtain distortion score values; and generate the video quality evaluation result according to the distortion weight values and the distortion score values.
[0143] In some example embodiments of the present disclosure, the quality evaluation module 830 is further configured to perform feature extraction on the video to be evaluated by at least one distortion sub-module to obtain at least one distortion sub-feature; each of the distortion sub-modules corresponds to a distortion type; wherein the distortion feature extraction module comprises the at least one distortion sub-module; and the video distortion features comprise the at least one distortion sub-feature.
[0144] In some example embodiments of the present disclosure, the quality evaluation module 830 is further configured to the distortion weight values comprise at least one distortion sub-weight value corresponding to the distortion type; the distortion score values comprise at least one distortion sub-score value corresponding to the distortion type; based on the distortion type, obtain a corresponding video quality sub-score according to the corresponding distortion sub-weight value and the distortion sub-score value; and determine a video quality score according to each of the video quality sub-scores.
[0145] In some example embodiments of the present disclosure, the quality evaluation module 830 is further configured to perform fusion perception on the video distortion features, the distortion weight values and the distortion score values by a multi-layer perception machine to obtain a video quality score.
[0146] As to the apparatus in the above-mentioned embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and thus will not be described here in detail.
[0147] The electronic device 900 according to such an embodiment of the present disclosure will be described below with reference to Figure 9 Figure 9 The display electronic device 900 is merely an example, and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0148] As shown in Figure 9 The electronic device 900 is in the form of a general computing device. The components of the electronic device 900 can include, but are not limited to, the above-mentioned at least one processing unit 910, the above-mentioned at least one storage unit 920, a bus 930 connecting different system components including the storage unit 920 and the processing unit 910, and a display unit 940.
[0149] The storage unit stores program codes which can be executed by the processing unit 910, so that the processing unit 910 performs the steps described in the above "Exemplary Methods" section according to various exemplary embodiments of the present disclosure. For example, the processing unit 910 can perform various steps as shown in Figure 2
[0150] For another example, the electronic device can implement various steps as shown in Figure 2 or Figure 4
[0151] The storage unit 920 can include a readable medium in the form of volatile storage unit, such as a random access memory (RAM) 921 and / or a cache memory 922, and can further include a read-only memory (ROM) 923.
[0152] The storage unit 920 can further include program / utility 924 having a set of program modules 925 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which can include an implementation of a network environment, or a combination thereof.
[0153] The bus 930 can represent one or more of several types of bus structures, including a storage unit bus or bus controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of a variety of bus structures.
[0154] The electronic device 900 can also communicate with one or more external devices 970 such as a keyboard or pointing device, a Bluetooth device, etc.; other devices such as a storage device or an external effects device; and / or one or more devices that enable a user to interact with the electronic device 900; and / or one or more devices that enable the electronic device 900 to communicate with one or more other computing devices. Such communication can be via an input / output (I / O) interface 950. Still yet, the electronic device 900 can communicate with one or more networks such as a local area network (LAN), a general wide area network (WAN), and / or a public network such as the Internet, via a network adapter 960. As depicted, the network adapter 960 communicates with the other components of the electronic device 900 via the bus 930. It should be appreciated that although the network adapter 960 is depicted as a single component, the network adapter 960 can include any number of components, such as a plurality of network adapters 960. Further, it should be appreciated that the bus 930 can be implemented using any type of communications fabric and that the
[0155] Through the above description of the embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the disclosure can be embodied in the form of a software product. The software product can be stored in a nonvolatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) or a network, and includes a number of instructions to make a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) execute the method according to the embodiments of the disclosure.
[0156] In the example embodiments, a computer readable storage medium including instructions, such as a memory including instructions, is also provided. The instructions can be executed by a processor of an apparatus to complete the above method. Optionally, the computer readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0157] In the example embodiments, a computer program product including a computer program is also provided. The computer program is executed by a processor to implement the method in the above embodiments.
[0158] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the disclosure disclosed herein. This application is intended to cover any variations, uses or adaptations of the disclosure that are deemed to fall within the general principles of the disclosure and include examples of the disclosure that are apparent to those skilled in the art. The specification and examples are to be regarded as illustrative only and the true scope and spirit of the disclosure is indicated by the following claims.
[0159] It should be understood that the present disclosure is not limited to the precise structures described and shown in the drawings and that various modifications and changes can be made without departing from the scope of the present disclosure. The scope of the present disclosure is limited only by the claims that follow.
Claims
1. A method for training a video quality assessment model, characterized in that, The method comprises the following steps: obtaining a training data set comprising a plurality of video quality training samples; each of the training samples comprises at least: a video sample and a sample quality score; the sample quality score is used to represent the quality score of the corresponding video sample; extracting features of the video sample through a distortion feature extraction module to obtain video distortion features; evaluating distortion weights of the video distortion features through a to-be-trained gate structure to obtain distortion weight values; evaluating distortion degrees of the video distortion features through a prediction network to obtain distortion score values; determining a video quality score according to the distortion weight values and the distortion score values; generating a quality score loss value according to the video quality score and the sample quality score; the quality score loss value is used to represent the difference between the video quality score and the sample quality score; training a to-be-trained video quality evaluation model according to the quality score loss value to obtain a video quality evaluation model; the video quality evaluation model comprises at least: the distortion feature extraction module, the gate structure and the prediction network.
2. The method of claim 1, wherein, The method comprises the following steps: extracting features of the video sample through at least one distortion submodule to obtain at least one distortion sub-feature; each of the distortion submodules corresponds to a distortion type; wherein the distortion feature extraction module comprises the at least one distortion submodule; the video distortion features comprise the at least one distortion sub-feature.
3. The method of claim 2, wherein, The training sample further comprises a first sample distortion sub-score corresponding to a first distortion type; the method further comprises the following steps: extracting features of the video sample through a to-be-trained first distortion submodule to obtain first distortion sub-features; evaluating distortion degrees of the first distortion sub-features through the prediction network to obtain first distortion sub-scores; generating a first distortion sub-score loss value according to the first distortion sub-scores and the first sample distortion sub-score; training the to-be-trained first distortion submodule according to the first distortion sub-score loss value to obtain a first distortion submodule; wherein the first distortion submodule belongs to the at least one distortion submodule.
4. The method of claim 2, wherein, The distortion weight values comprise at least one distortion sub-weight value corresponding to the distortion type; the distortion score values comprise at least one distortion sub-score value corresponding to the distortion type; The method comprises the following steps: based on the distortion type, obtaining corresponding video quality sub-scores according to the corresponding distortion sub-weight values and distortion sub-score values; determining the video quality score according to each of the video quality sub-scores.
5. The method of claim 2, wherein, The method comprises the following steps: fusing perception of the video distortion features, the distortion weight values and the distortion score values through a multi-layer perception machine to obtain the video quality score.
6. A method of video quality assessment, characterized by, The method comprises the following steps: obtaining a to-be-evaluated video; inputting the video to be evaluated into a video quality evaluation model; the video quality evaluation model is obtained by training the video quality evaluation model training method in any one of claims 1 to 5; evaluating the video to be evaluated by the video quality evaluation model to obtain a video quality evaluation result.
7. The method of claim 6, wherein, The evaluation of the video to be evaluated by the video quality evaluation model to obtain a video quality evaluation result comprises: extracting features of the video to be evaluated by a distortion feature extraction module to obtain video distortion features; evaluating distortion weights of the video distortion features by a gate structure to obtain distortion weight values; evaluating distortion degrees of the video distortion features by a prediction network to obtain distortion score values; generating the video quality evaluation result according to the distortion weight values and the distortion score values.
8. The method of claim 7, wherein, The feature extraction of the video to be evaluated by the distortion feature extraction module to obtain video distortion features comprises: extracting features of the video to be evaluated by at least one distortion submodule to obtain at least one distortion sub-feature; each distortion submodule corresponds to one distortion type; The distortion feature extraction module comprises the at least one distortion submodule; the video distortion features comprise the at least one distortion sub-feature.
9. The method of claim 8, wherein, The distortion weight values comprise at least one distortion sub-weight value corresponding to the distortion type; the distortion score values comprise at least one distortion sub-score value corresponding to the distortion type; The generation of the video quality evaluation result according to the distortion weight values and the distortion score values comprises: based on the distortion type, obtaining corresponding video quality sub-scores according to corresponding distortion sub-weight values and distortion sub-score values; determining a video quality score according to each video quality sub-score.
10. The method of claim 8, wherein, The generation of the video quality evaluation result according to the distortion weight values and the distortion score values comprises: fusing perception of the video distortion features, the distortion weight values and the distortion score values by a multi-layer perception machine to obtain a video quality score.
11. A video quality assessment model training apparatus, characterized in that, It comprises: a sample acquisition module configured to acquire a training data set comprising a plurality of video quality training samples; each training sample comprises at least a video sample and a sample quality score; the sample quality score is used to represent the quality score of the corresponding video sample; a distortion feature extraction module configured to extract features of the video sample by a distortion feature extraction module to obtain video distortion features; a gate structure module configured to evaluate distortion weights of the video distortion features by a to-be-trained gate structure to obtain distortion weight values; a prediction network module configured to evaluate distortion degrees of the video distortion features by a prediction network to obtain distortion score values; a video quality score module configured to determine a video quality score according to the distortion weight values and the distortion score values; a loss value module configured to generate a quality score loss value according to the video quality score and the sample quality score; the quality score loss value is used to represent the difference between the video quality score and the sample quality score. The model training module is configured to train the video quality assessment model to be trained based on the quality score loss value, thereby obtaining the video quality assessment model; the video quality assessment model includes at least the distortion feature extraction module, a gate structure, and a prediction network.
12. A video quality assessment apparatus characterized by comprising: include: The evaluation request acquisition module is configured to acquire the video to be evaluated; An evaluation request input module is configured to input the video to be evaluated into a video quality evaluation model; the video quality evaluation model is obtained by training the video quality evaluation model training method as described in any one of claims 1 to 5. The quality assessment module is configured to evaluate the video to be evaluated using the video quality assessment model to obtain a video quality assessment result.
13. An electronic device, comprising: include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the executable instructions to implement the video quality assessment model training method as described in any one of claims 1 to 5 or the video quality assessment method as described in any one of claims 6 to 10.
14. A computer-readable storage medium, wherein instructions in the computer-readable storage medium, when executed by a processor of an electronic device, enable the electronic device to perform the video quality assessment model training method as claimed in any one of claims 1 to 5 or the video quality assessment method as claimed in any one of claims 6 to 10.
15. A computer program product comprising a computer program, characterized in that, When executed by a processor, the computer program implements the video quality assessment model training method as described in any one of claims 1 to 5 or the video quality assessment method as described in any one of claims 6 to 10.