Video quality evaluation method, electronic equipment, storage medium and program product

By extracting global and fragmented content features from the video through the evaluation model, the problem of the single dimension of video quality evaluation in the existing technology is solved, and more accurate video quality evaluation is achieved.

CN121509646APending Publication Date: 2026-02-10CHINA MOBILE GRP BEIJING +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511694361.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In existing technologies, video quality assessment algorithms rely on a single dimension when evaluating video quality, making it difficult to accurately assess video quality.

Method used

The evaluation model extracts global content features and fragmented content features from the target video, calculates the quality scores of the global content features and fragmented content features respectively, and then sums them by weight to output the overall quality score of the target video.

Benefits of technology

It enables the evaluation of video quality from multiple dimensions, resulting in a more accurate assessment of video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509646A_ABST
    Figure CN121509646A_ABST
Patent Text Reader

Abstract

The invention discloses a video quality evaluation method, electronic equipment, a storage medium and a program product, belongs to the technical field of computer vision, and is used for improving the evaluation accuracy of video quality. The method comprises the following steps: acquiring a to-be-evaluated target video; and inputting the target video into an evaluation model, and outputting a first quality score of the target video through the evaluation model, the evaluation model being used for determining the first quality score according to a first global content feature and a first fragment content feature of the target video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer vision technology, specifically relating to a video quality assessment method, electronic device, storage medium, and program product. Background Technology

[0002] Currently, advancements in video calling technology have led to a greater preference for video calls in social interactions, and the quality and stability of these calls significantly impact user experience. However, current video quality assessment algorithms tend to rely on a single dimension for evaluation. For instance, they often only assess static indicators like "image sharpness," making it difficult to consider different dimensions and calculate appropriate quality scores. This results in inaccurate video quality assessments. Summary of the Invention

[0003] This application provides a method, electronic device, storage medium, and program product for evaluating video quality, which can solve the problem of difficulty in accurately evaluating video quality.

[0004] In a first aspect, embodiments of this application provide a method for evaluating video quality. The method includes: acquiring a target video to be evaluated; inputting the target video into an evaluation model; and outputting a first quality score of the target video through the evaluation model. The evaluation model is used to determine the first quality score based on a first global content feature and a first fragment content feature of the target video.

[0005] Secondly, embodiments of this application provide a video quality evaluation device, which includes: an acquisition module for acquiring a target video to be evaluated; and an evaluation module for inputting the target video into an evaluation model and outputting a first quality score of the target video through the evaluation model, wherein the evaluation model is used to determine the first quality score based on a first global content feature and a first fragment content feature of the target video.

[0006] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0007] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0008] Fifthly, embodiments of this application provide a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the steps of the method described in the first aspect.

[0009] In this embodiment of the application, the target video to be evaluated is obtained; the target video is input into the evaluation model, and the evaluation model outputs the first quality score of the target video. The evaluation model is used to determine the first quality score based on the first global content feature and the first fragment content feature of the target video. The first quality score of the target video can be calculated from two aspects: the score of the global content feature and the quality score of the fragment content feature. The video quality of the target video is evaluated from multiple dimensions, which can more accurately evaluate the video quality of the target video. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart illustrating a video quality assessment method provided in an embodiment of this application; Figure 2 This is a schematic diagram illustrating the training process of an evaluation model and the video quality evaluation process of a target video, as provided in an embodiment of this application. Figure 3 This is a schematic diagram of the structure of a video quality evaluation device provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0013] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0014] The following description, in conjunction with the accompanying drawings, details the video quality evaluation method, electronic device, storage medium, and program product provided in this application through specific embodiments and application scenarios.

[0015] Figure 1 This application illustrates an embodiment of a video quality evaluation method, which can be performed by an electronic device. In other words, the method can be performed by software or hardware installed in the electronic device, and includes the following steps: Step S101: Obtain the target video to be evaluated.

[0016] Video over New Radio (ViNR) technology leverages the high bandwidth and low latency of 5G to provide high-definition video call capabilities. During high-definition video calls, video image quality assessment has a significant impact on user experience. In this application, the target video whose quality needs to be assessed can be a ViNR video.

[0017] Step S102: Input the target video into the evaluation model, and output the first quality score of the target video through the evaluation model.

[0018] In this embodiment, the target video can be input into an evaluation model, which extracts the first global content feature and the first fragment content feature of the target video. A quality score is then applied to the first global content feature and the first fragment content feature to obtain the evaluation model's quality score for the first global content feature and the quality score for the first fragment content feature. Based on these two quality scores, the first quality score of the target video can be finally output. For example, the first quality score of the target video can be calculated by summing the quality scores of the first global content feature and the first fragment content feature, or by weighted summation. In one embodiment, the first global content feature of the target video may include at least one of the following: content-related features, imaging effect, color, and lighting effects. The first fragment content feature includes at least one of the following: sharpness, focus, noise, motion blur, flash exposure, video compression, and video latency.

[0019] The video quality assessment method provided in this application involves acquiring a target video to be assessed; inputting the target video into an assessment model; and outputting a first quality score for the target video through the assessment model. The assessment model is used to determine the first quality score based on the first global content feature and the first fragment content feature of the target video. The first quality score of the target video can be calculated from two aspects: the score of the global content feature and the quality score of the fragment content feature. By assessing the video quality of the target video from multiple dimensions, the video quality of the target video can be assessed more accurately.

[0020] In one embodiment, before acquiring the target video to be evaluated, the method further includes: inputting a training video into an evaluation model; outputting a second quality score of a second global content feature of the training video through the evaluation model; outputting a third quality score of a second fragment content feature of the training video through the evaluation model; and optimizing the evaluation model based on the second quality score and the third quality score.

[0021] Figure 2 This illustration shows a schematic diagram of the training process for an evaluation model and the video quality evaluation process for a target video, as provided in an embodiment of this application. The following is in conjunction with... Figure 2This application describes the training of the evaluation model and the video quality evaluation of the target video in the embodiments of this application. Before performing video quality evaluation on the target video, the evaluation model needs to be trained. A training video for training the evaluation model can be obtained. In this application, the training video needs to be preprocessed, adjusting its width and height to 512*512 pixels to ensure it has at least 64 frames. Each frame of the training video is standardized by channel, i.e., subtracting the mean from each of the three RGB channels and then dividing by the standard deviation. The means are [0.485, 0.456, 0.406], and the standard deviations are [0.229, 0.224, 0.225]. The preprocessed training video is input into the evaluation model. The evaluation model outputs a second quality score for the second global content feature of the training video; it also outputs a third quality score for the second fragment content feature of the training video. The evaluation model can then be optimized based on the second and third quality scores.

[0022] The second quality score of the second global content feature of the training video is obtained by the evaluation model through evaluating the second global content feature of the training video, which integrates the global dimensions of the training video, including content, imaging quality, color, and lighting effects. The third quality score of the second fragment content feature of the training video is obtained by the evaluation model through evaluating the second fragment content feature of the training video, which includes fragmented and temporal dimensions such as sharpness, focus, noise, motion blur, flash exposure, video compression, and video latency. After obtaining the second quality score of the second global content feature and the third quality score of the second fragment content feature of the training data evaluated by the evaluation model, the evaluation model can be optimized based on the second and third quality scores.

[0023] In one embodiment, the second quality score of the second global content feature of the training video is output by the evaluation model, including: extracting the second global content feature of the training video by the global content quality evaluation module of the evaluation model; evaluating the second global content feature by the global content quality evaluation module and outputting the second quality score.

[0024] This application embodiment designs a global content quality evaluation module for the evaluation model. In this embodiment, four frames of the training video with the same duration can be extracted first and used as input to the global content quality evaluation module. This application embodiment designs a feature extraction network module that can extract features from the 4x512x512x3 training video input, extracting the second global content features of the training video. In this embodiment, a global mean pooling module can be added to the global content quality evaluation module to map the second global content features to features with a dimension of 1x2048, and finally feed them into the fully connected layer of the global content quality evaluation module to evaluate the second quality score of the second global content features.

[0025] The quality score for global content features is primarily evaluated by considering the content quality score of the extracted video frames within their own image range. This method, with its relatively low frame rate, is mainly used to capture relatively static semantic information in space, thereby assessing the overall content quality. Due to the small number of frames processed, a larger network backbone can be designed to extract global content features without incurring significant computational costs.

[0026] In one embodiment, evaluating a third quality score of the second fragment content features of a training video by an evaluation model includes: extracting multiple video frame images from the training video, dividing them into multiple blocks at equal intervals, and extracting fragment regions from each block; stitching the fragment regions extracted from the multiple video frame images into a fragment image; extracting second fragment content features from the fragment image using a fragment content quality evaluation module of the evaluation model; and evaluating the second fragment content features using the fragment content quality evaluation module to output a third quality score.

[0027] This application embodiment designs a fragment content quality assessment module. This embodiment can extract 16 video frames of equal length from the training video. Each video frame is processed as follows: the extracted video frame is divided into 16x16 blocks at equal intervals, and a 16x16 fragment region is taken from the center of each block. After stitching the fragment regions extracted from each frame, the original video frame size is reduced from 512x512 to 256x256 fragment images. These fragment images are used as input to the fragment content quality assessment module.

[0028] This application's embodiments design a feature extraction network module to extract second fragment content features from the input 16x256x256x3 fragmented image. The fragment content quality assessment module considers the temporal technical quality of each fragment after image fragmentation, achieving a high frame rate for the video. Furthermore, the fragmentation design not only reduces the input resolution but also eliminates concerns about spatial semantics, primarily capturing motion information that changes significantly over time and the quality of small fragments themselves. Due to the large number of frames processed, a smaller network backbone can be designed for feature extraction. In this application's embodiments, when the extracted second fragment content features are downsampled to feature maps of sizes 64*64, 32*32, and 16*16 by the fragment content quality assessment module, the extracted second fragment content features are connected to the global content quality assessment module to form feature merging. The time-related features captured by the fragmented content quality assessment module, such as sharpness, focus, noise, motion blur, flash exposure, video compression, and video latency, are supplemented into the feature representation of the global content quality assessment module. This allows the module to perceive the technical quality information in the time dimension while assessing global content and spatial semantics, thus forming a more robust feature representation.

[0029] This application embodiment adds a global mean pooling module to the fragmented content quality assessment module, mapping the second fragmented content features to features of dimension 1x512, and finally feeding them into a fully connected layer to evaluate the second quality score of the second fragmented content. This application embodiment employs a design with different sampling frequencies and the idea of ​​fragmented images to reduce the computational resources consumed by video quality assessment, allowing for deployment on hardware devices with less computing power.

[0030] In one embodiment, optimizing the evaluation model based on the global content quality score and the fragmented content quality score includes: optimizing the evaluation model based on a second quality score, a third quality score, a first pre-labeled score of the second global content feature, a second pre-labeled score of the second fragmented content feature, and a loss function.

[0031] In this embodiment, the absolute average error loss is used as the loss function to evaluate the training of the model, as shown in the following formula:

[0032] Where n is the number of training videos input in each batch. It is an absolute value function. This is the second quality score evaluated by the global content quality assessment module. The first pre-annotated score is the second global content feature annotation of the training video. This is the third quality score assessed by the fragmented content quality assessment module. The second pre-labeled score is used to annotate the second fragment content features of the training video.

[0033] In this embodiment, a training dataset can be constructed based on the training data. In the training dataset, a first pre-labeled score can be pre-annotated to the second global content features of the training data, and a second pre-labeled score can be pre-annotated to the second fragment content features. The annotation format is as follows:

[0034] In this embodiment, the CLIP model can be used for prediction followed by annotation, and the corresponding score is calculated after assigning weight ratios to each dimension. It should be noted that in this embodiment, the process of training the evaluation model can also be the process of evaluating the first quality score of the target video to be tested.

[0035] In one embodiment, the first quality score of the target video is output through the evaluation model, including: evaluating the first global content feature of the target video through the global content quality evaluation module of the evaluation model and outputting a fourth quality score; evaluating the first fragment content feature of the target video through the fragment content quality evaluation module of the evaluation model and outputting a fifth quality score; and outputting the first quality score based on the fourth quality score, the fifth quality score, the first weighted weight corresponding to the first global content feature, and the second weighted weight corresponding to the first fragment content feature.

[0036] In this embodiment, after training the evaluation model, the network weight parameters of the global content quality evaluation module and the fragment content quality evaluation module can be saved. After inputting the target video into the evaluation model, the global content quality evaluation module of the evaluation model can evaluate the first global content feature of the target video and output a fourth quality score. The fragment content quality evaluation module of the evaluation model can evaluate the first fragment content feature of the target video and output a fifth quality score. Then, the fourth and fifth quality scores are summed in a weighted manner to obtain the first quality score VMOS of the target video, as shown in the following formula:

[0037] It is the fourth mass fraction. In the fifth mass fraction, This is the weighting factor for the fourth mass fraction. The weighting factor for the fifth mass fraction is , and .

[0038] This application embodiment can calculate the quality score of a target video from two aspects: the score of its global content features and the score of its fragmented content features. This multi-dimensional evaluation of the target video's quality allows for a more accurate assessment. It should be noted that the video quality evaluation process is the same as the process of training an evaluation model using training data; therefore, it will not be repeated here to avoid repetition.

[0039] Optionally, the network structure of the fragmented content quality assessment module can be consistent with the backbone of the global content quality assessment module, but the size and number of channels can be designed differently. Different network backbones can also be selected. The principle is that the global content quality assessment module should use a large network structure, while the fragmented content quality assessment module should use a lightweight network structure. The following is a design scheme with the same backbone but different specific parameters:

[0040] It should be noted that the video quality evaluation method provided in this application embodiment can be executed by a video quality evaluation device or a control module within that device for executing the video quality evaluation method. This application embodiment uses the execution of the video quality evaluation method by a video quality evaluation device as an example to illustrate the video quality evaluation device provided in this application embodiment.

[0041] Figure 3 This is a schematic diagram of the structure of a video quality evaluation device according to an embodiment of this application. Figure 3 As shown, the video quality evaluation device 300 includes an acquisition module 310 and an evaluation module 320.

[0042] The acquisition module 310 is used to acquire the target video to be evaluated; the evaluation module 320 is used to input the target video into the evaluation model and output the first quality score of the target video through the evaluation model, wherein the evaluation model is used to determine the first quality score based on the first global content feature and the first fragment content feature of the target video.

[0043] In one embodiment, the evaluation module 320 is further configured to input the training video into the evaluation model; output a second quality score of the second global content feature of the training video through the evaluation model; output a third quality score of the second fragment content feature of the training video through the evaluation model; and optimize the evaluation model based on the second quality score and the third quality score.

[0044] In one embodiment, the evaluation module 320 is used to extract the second global content feature of the training video through the global content quality evaluation module of the evaluation model; evaluate the second global content feature through the global content quality evaluation module, and output the second quality score.

[0045] In one embodiment, the evaluation module 320 is configured to extract multiple video frame images from the training video, divide them into multiple blocks at equal intervals, extract fragment regions from each block, stitch the fragment regions extracted from the multiple video frame images into a fragment image, extract the second fragment content feature from the fragment image through the fragment content quality evaluation module of the evaluation model, evaluate the second fragment content feature through the fragment content quality evaluation module, and output the third quality score.

[0046] In one embodiment, the evaluation module 320 is configured to optimize the evaluation model based on the second quality score, the third quality score, the first pre-labeled score of the second global content feature, the second pre-labeled score of the second fragmented content feature, and a loss function.

[0047] In one embodiment, the evaluation module 320 is configured to evaluate the first global content feature of the target video through the global content quality evaluation module of the evaluation model and output a fourth quality score; evaluate the first fragment content feature of the target video through the fragment content quality evaluation module of the evaluation model and output a fifth quality score; and output the first quality score based on the fourth quality score, the fifth quality score, the first weighted weight corresponding to the first global content feature, and the second weighted weight corresponding to the first fragment content feature.

[0048] In one embodiment, the first global content feature includes at least one of the following: the content of the target video, imaging effect, color, and lighting effect; the first fragment content feature includes at least one of the following: sharpness, focus, noise, motion blur, flash exposure, video compression, and video delay.

[0049] The video quality evaluation device in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.

[0050] The video quality evaluation device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0051] The video quality evaluation device provided in this application embodiment can achieve Figures 1 to 2 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0052] Optionally, such as Figure 4As shown in the illustration, this application embodiment also provides an electronic device 400, including a processor 401 and a memory 402. The memory 402 stores a program or instructions that can run on the processor 401. When the program or instructions are executed by the processor 401, they perform the following: acquiring a target video to be evaluated; inputting the target video into an evaluation model; and outputting a first quality score of the target video through the evaluation model. The evaluation model is used to determine the first quality score based on a first global content feature and a first fragment content feature of the target video.

[0053] In one embodiment, before acquiring the target video to be evaluated, a training video is input into the evaluation model; a second quality score of a second global content feature of the training video is output by the evaluation model; a third quality score of a second fragment content feature of the training video is output by the evaluation model; and the evaluation model is optimized based on the second quality score and the third quality score.

[0054] In one embodiment, the second global content feature of the training video is extracted by the global content quality assessment module of the evaluation model; the second global content feature is evaluated by the global content quality assessment module, and the second quality score is output.

[0055] In one embodiment, multiple video frame images are extracted from the training video and divided into multiple blocks at equal intervals. Fragment regions are extracted from each block. The fragment regions extracted from the multiple video frame images are stitched together to form a fragment image. The fragment content quality assessment module of the evaluation model extracts the second fragment content feature from the fragment image. The fragment content quality assessment module evaluates the second fragment content feature and outputs the third quality score.

[0056] In one embodiment, the evaluation model is optimized based on the second quality score, the third quality score, the first pre-labeled score of the second global content feature, the second pre-labeled score of the second fragmented content feature, and a loss function.

[0057] In one embodiment, the first global content feature of the target video is evaluated by the global content quality evaluation module of the evaluation model, and a fourth quality score is output; the first fragment content feature of the target video is evaluated by the fragment content quality evaluation module of the evaluation model, and a fifth quality score is output; the first quality score is output based on the fourth quality score, the fifth quality score, the first weighted weight corresponding to the first global content feature, and the second weighted weight corresponding to the first fragment content feature.

[0058] In one embodiment, the first global content feature includes at least one of the following: the content of the target video, imaging effect, color, and lighting effect; the first fragment content feature includes at least one of the following: sharpness, focus, noise, motion blur, flash exposure, video compression, and video delay.

[0059] The specific execution steps can be found in the various steps of the above-described video quality assessment method embodiment, and can achieve the same technical effect. To avoid repetition, they will not be repeated here.

[0060] It should be noted that the electronic devices in the embodiments of this application include: servers, terminals, or other devices besides terminals.

[0061] The above electronic device structure does not constitute a limitation on the electronic device. An electronic device may include more or fewer components than illustrated, or combine certain components, or arrange them differently. For example, an input unit may include a Graphics Processing Unit (GPU) and a microphone, and a display unit may use a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar display panels. User input units include at least one of a touch panel and other input devices. A touch panel is also called a touchscreen. Other input devices may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be elaborated further here.

[0062] Memory can be used to store software programs and various data. Memory can primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area can store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, memory can include volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (Synchlink DRAM, SLDRAM), and direct memory bus RAM (DRRAM).

[0063] The processor may include one or more processing units; optionally, the processor integrates an application processor and a modem processor, wherein the application processor mainly handles operations related to the operating system, user interface, and applications, while the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor.

[0064] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described video quality evaluation method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0065] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as ROM, RAM, magnetic disk, or optical disk.

[0066] This application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, which, when executed by a computer, cause the computer to perform various processes of the above-described video quality evaluation method embodiments and achieve the same technical effect. To avoid repetition, these will not be described again here.

[0067] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0068] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0069] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for evaluating video quality, characterized in that, include: Obtain the target video to be evaluated; The target video is input into the evaluation model, and the evaluation model outputs a first quality score for the target video. The evaluation model is used to determine the first quality score based on the first global content feature and the first fragment content feature of the target video.

2. The method according to claim 1, characterized in that, Before acquiring the target video to be evaluated, the following is also included: The training video is input into the evaluation model; The evaluation model outputs a second quality score for the second global content feature of the training video. The evaluation model outputs a third quality score for the second fragment content features of the training video; The evaluation model is optimized based on the second quality score and the third quality score.

3. The method according to claim 2, characterized in that, The second quality score of the second global content feature of the training video output by the evaluation model includes: The second global content feature of the training video is extracted through the global content quality assessment module of the evaluation model; The second global content feature is evaluated by the global content quality assessment module, and the second quality score is output.

4. The method according to claim 2, characterized in that, The third quality score of the second fragment content features of the training video output by the evaluation model includes: The training video is divided into multiple blocks by extracting multiple video frame images at equal intervals, and fragment regions are extracted from each block. Fragmented regions extracted from multiple video frame images are stitched together to form a fragmented image; The fragment content quality assessment module of the evaluation model extracts the second fragment content features from the fragment image; The fragment content quality assessment module evaluates the characteristics of the second fragment content and outputs the third quality score.

5. The method according to claim 2, characterized in that, The optimization of the evaluation model based on the global content quality score and the fragmented content quality score includes: The evaluation model is optimized based on the second quality score, the third quality score, the first pre-labeled score of the second global content feature, the second pre-labeled score of the second fragmented content feature, and the loss function.

6. The method according to claim 1, characterized in that, The step of outputting a first quality score for the target video through the evaluation model includes: The first global content feature of the target video is evaluated by the global content quality evaluation module of the evaluation model, and a fourth quality score is output. The fragment content quality assessment module of the assessment model evaluates the first fragment content features of the target video and outputs a fifth quality score. The first quality score is output based on the fourth quality score, the fifth quality score, the first weighted weight corresponding to the first global content feature, and the second weighted weight corresponding to the first fragmented content feature.

7. The method according to claim 1, characterized in that, The first global content feature includes at least one of the following: The target video's content, imaging effects, color, and lighting effects; The first fragment content features include at least one of the following: sharpness, focus, noise, motion blur, flash exposure, video compression, and video delay.

8. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the video quality evaluation method as described in any one of claims 1-7.

9. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the video quality evaluation method as described in any one of claims 1-7.

10. A computer program product, characterized in that, The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions that, when executed by a computer, cause the computer to perform the steps of the video quality evaluation method as described in any one of claims 1-7.