Video quality evaluation method and device, computer equipment and storage medium

By performing frame extraction processing and multimodal evaluation model analysis on the video, multi-dimensional video quality rating information is generated, which solves the problems of incomplete and unstable video quality evaluation in the existing technology and achieves efficient and accurate video quality evaluation.

CN120634985APending Publication Date: 2025-09-12BEIJING KINGSOFT CLOUD NETWORK TECH CO LTD

Patent Information

Application Number
CN202510705169.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing video quality assessment methods have limitations in subjective experience, multimodal content understanding, and cross-scene adaptability, making it difficult to achieve a comprehensive and stable evaluation of the overall video quality.

Method used

By extracting frames from the target video, an image sequence is obtained and input into a trained multimodal large model or lightweight visual quality assessment model to extract the index scoring information of multiple video quality indicators. The video scoring information is generated by combining weighted average or nonlinear combination, supporting global and local area evaluation.

Benefits of technology

It achieves efficient, comprehensive and accurate evaluation of multiple quality dimensions of videos, and generates representative and explainable comprehensive video scoring results, which are suitable for content optimization and intelligent recommendation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634985A_ABST
    Figure CN120634985A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a video quality evaluation method and device, computer equipment and a storage medium, and the method comprises the steps: carrying out the frame extraction of a target video, and obtaining an image sequence; inputting the image sequence into a trained target model, and outputting index scoring information corresponding to each video quality index of the target video through the target model to obtain multiple pieces of index scoring information; and determining video scoring information of the target video according to the multiple pieces of index scoring information. Therefore, the scoring information can be generated for the quality indexes of the multiple dimensions of the video through the target model, the video scoring information is further generated according to the scoring information of the different quality indexes, automatic scoring of the video based on the multiple quality indexes of the video is achieved, and therefore the overall quality of the video is evaluated efficiently, comprehensively and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of video processing technology, and in particular to a video quality assessment method, apparatus, computer equipment, and storage medium. Background Art

[0002] Image / Video Quality Assessment (IQA) is a crucial technology for ensuring the quality of video content dissemination. It is widely used in scenarios such as video compression optimization, image enhancement, content recommendation, and quality control. Existing assessment methods are generally categorized into three types: full-reference, semi-reference, and no-reference. No-reference video quality assessment, which does not rely on the original video, offers greater flexibility and practicality in practical applications.

[0003] Currently, mainstream methods for quality scoring rely on image feature extraction, statistical analysis, or convolutional neural networks. However, these methods still have significant limitations in terms of subjective experience, understanding multimodal content, and cross-scenario adaptability. Traditional algorithms struggle to handle complex distortions and diverse content. Furthermore, existing deep learning models often lack interpretability and multimodal support, making it difficult to comprehensively assess overall video quality, resulting in unstable and incomplete evaluation results. Therefore, improving the accuracy of video quality assessment has become a pressing issue. Summary of the Invention

[0004] In view of this, in order to solve the above technical problems or part of the technical problems, embodiments of the present invention provide a video quality assessment method, apparatus, computer equipment and storage medium.

[0005] In a first aspect, an embodiment of the present invention provides a video quality assessment method, comprising:

[0006] Perform frame extraction on the target video to obtain an image sequence;

[0007] Inputting the image sequence into a trained target model, so as to output the indicator scoring information corresponding to each video quality indicator of the target video through the target model, thereby obtaining a plurality of indicator scoring information;

[0008] The video rating information of the target video is determined based on the plurality of indicator rating information.

[0009] In a possible implementation, the video quality index includes: a plurality of picture perception indexes and a plurality of technical indexes;

[0010] Outputting the indicator scoring information corresponding to each video quality indicator of the target video through the target model includes:

[0011] Outputting the index score and index evaluation corresponding to each of the picture perception indexes and each of the technical indicators through the target model;

[0012] The indicator score and the indicator evaluation are used as the indicator scoring information.

[0013] In a possible implementation, when the video rating information is a global rating for the target video, determining the video rating information of the target video based on the plurality of indicator rating information includes:

[0014] Obtaining an indicator weight corresponding to each of the video quality indicators;

[0015] Determine the video score of the target video according to the indicator weight and indicator score corresponding to each of the video quality indicators;

[0016] Determine the video rating of the target video according to the indicator evaluation corresponding to each of the video quality indicators;

[0017] The video score and the video evaluation are used as the video rating information.

[0018] In one possible implementation, when the video rating information is a local area rating for the target video, the method further includes:

[0019] Identifying a target region of each image in the image sequence;

[0020] Inputting the image sequence into a trained target model, so as to output regional indicator scoring information corresponding to each regional quality indicator of the target region through the target model, thereby obtaining a plurality of regional indicator scoring information;

[0021] The regional scoring information of each target area is determined according to the plurality of regional index scoring information.

[0022] In one possible implementation, determining the video score information of the target video based on the plurality of indicator score information includes:

[0023] When the index score of any of the video quality indicators is less than a first threshold, generating a visual explanation according to the index evaluation corresponding to the video quality indicator;

[0024] The video scoring information is determined based on the indicator scoring information and the visual interpretation.

[0025] In one possible implementation, the method further includes:

[0026] When the video score in the video rating information is less than a second threshold, determining a target video quality index from the plurality of video quality indexes, wherein the index score of the target video quality index is less than a third threshold;

[0027] Determining a video adjustment strategy corresponding to the target video quality indicator;

[0028] The target video is adjusted according to the video adjustment strategy.

[0029] In one possible implementation, before performing frame extraction processing on the target video to obtain an image sequence, the method further includes:

[0030] Obtain the playback platform to which the target video belongs;

[0031] The type of the video rating information is determined according to the playback platform, where the type includes: a global rating for the target video, or a local area rating for the target video.

[0032] In a second aspect, an embodiment of the present invention provides a video quality scoring device, comprising:

[0033] The video processing module is used to extract frames from the target video to obtain an image sequence;

[0034] An information output module is used to input the image sequence into a trained target model to output the indicator scoring information corresponding to each video quality indicator of the target video through the target model to obtain multiple indicator scoring information;

[0035] The determination module is used to determine the video score information of the target video according to the multiple indicator score information.

[0036] In a third aspect, an embodiment of the present invention provides a computer device, comprising: a processor and a memory, wherein the processor is configured to execute a video quality scoring program stored in the memory to implement the video quality assessment method described in any one of the first aspects above.

[0037] In a fourth aspect, an embodiment of the present invention provides a storage medium storing one or more programs, which can be executed by one or more processors to implement the video quality assessment method described in any one of the first aspects above.

[0038] The video quality assessment solution provided by an embodiment of the present invention extracts frames from a target video to obtain an image sequence; this image sequence is input into a trained target model, which then outputs indicator scoring information corresponding to each video quality indicator of the target video, thereby obtaining multiple indicator scoring information; and a video scoring information for the target video is determined based on the multiple indicator scoring information. Thus, the target model can generate scoring information for quality indicators across multiple dimensions of the video, and further generate video scoring information based on the scoring information for different quality indicators, thereby achieving automated scoring of videos based on multiple quality indicators, thereby efficiently, comprehensively, and accurately assessing the overall quality of the video. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 A schematic diagram of a flow chart of a video quality assessment method provided by an embodiment of the present invention;

[0040] Figure 2 A schematic diagram of a flow chart of another video quality assessment method provided by an embodiment of the present invention;

[0041] Figure 3 A schematic diagram of a video quality indicator provided by an embodiment of the present invention;

[0042] Figure 4 A schematic diagram of an indicator score and indicator evaluation provided by an embodiment of the present invention;

[0043] Figure 5 A schematic structural diagram of a video quality assessment device provided by an embodiment of the present invention;

[0044] Figure 6 A schematic structural diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0046] To facilitate understanding of the embodiments of the present invention, specific embodiments will be further explained below with reference to the accompanying drawings. The embodiments do not limit the embodiments of the present invention.

[0047] Figure 1 A flow chart of a video quality assessment method provided by an embodiment of the present invention is shown as follows: Figure 1As shown, the method specifically includes:

[0048] S11. Perform frame extraction processing on the target video to obtain an image sequence.

[0049] The video quality assessment method provided in an embodiment of the present invention is applied to a computer device, which may include but is not limited to: a server, a desktop computer, a tablet computer, etc. Specifically, a target model is used to generate scoring information for multiple frames of video images based on quality indicators in multiple dimensions, and video scoring information is further generated based on the scoring information of different quality indicators.

[0050] In this embodiment, key frames are extracted from the target video to form an image sequence to ensure coverage of different scenes (light / dark, static / moving, near / far); each frame of the image is preprocessed, including resolution unification, color space standardization (such as unification to RGB), noise removal, etc.

[0051] Specifically, the target video is the video whose quality needs to be assessed. Representative key frames are extracted from the target video to reduce redundancy and improve assessment efficiency. Specifically, a frame extraction interval can be set (for example, extracting one frame every 1 second or using scene change detection to extract key frames). Video frame information is read using a video processing tool (such as OpenCV), and frames are extracted according to a set strategy. This outputs a preliminary image frame sequence.

[0052] Furthermore, the image is preprocessed to unify the input format and eliminate interference caused by inconsistent image quality. First, resize the image to the size required by the model (for example, 224×224). Then, normalize the image to convert pixel values ​​from [0, 255] to [0, 1] or [-1, 1]. Color channel normalization: Adjust the channel order according to the model requirements (for example, BGR to RGB) and subtract the mean and divide by the standard deviation. If the original image size ratio is inconsistent, center cropping or zero padding is used.

[0053] The preprocessed key frames are saved in the form of image sequences for subsequent input into the multimodal video quality assessment model.

[0054] S12. Input the image sequence into the trained target model, so that the target model outputs the indicator scoring information corresponding to each video quality indicator of the target video, and obtains multiple indicator scoring information.

[0055] In this embodiment, each frame of the image sequence represents the visual content of the target video at a certain moment, ensuring that the sample has temporal representativeness and content diversity. The image sequence is input into the target model, and the target model can adopt a trained multimodal large model or a lightweight visual quality assessment model (such as CNN, Vision Transformer, or Swin-T+MLP and other structures), whose function is to extract the semantics, content, texture and other quality-related features of the image. The model can set multiple parallel output heads in the last layer, and each output head corresponds to a specific video quality indicator (for example, clarity, color accuracy, noise, artifacts, picture stability, aesthetic perception, content integrity, etc.) to form a multi-dimensional scoring output. If inter-frame consistency is considered, a temporal modeling module (such as LSTM, Transformer Encoder) can also be introduced to perform inter-frame feature fusion to improve the stability and overall sense of the score.

[0056] High-level visual features are extracted from each frame and input into various indicator scoring branches. Each branch contains several linear layers, normalization layers, and score normalization modules, and outputs a normalized score (such as 01 or 15 points) and a text evaluation corresponding to different scores for each video quality indicator. The text evaluation can be pre-set according to the evaluation requirements (for example, if the video quality indicator is clarity and the score is 01, the text evaluation is: the image in the video is very blurry, all details are lost, and only a small amount of information is recognized). The model's final output is a collection of indicator scoring information corresponding to multiple indicators. These scores can be frame-level averages, frame-level maximum / minimum values, or weighted aggregation using an attention mechanism to form a more representative indicator score.

[0057] This enables independent assessment of multiple video quality dimensions, without the need for reference videos, achieving true reference-free quality perception. Each score can be used for subsequent weighted fusion to derive an overall video quality score, or individually to identify specific quality issues, supporting content optimization and intelligent recommendations.

[0058] S13. Determine video rating information of the target video based on multiple indicator rating information.

[0059] In this embodiment, multi-dimensional indicator scoring information is received from the model output, and the weight of each indicator is preset or obtained through training and learning (for example: clarity: 0.25, color performance: 0.15, noise level: 0.10). The weight can be adjusted according to different application scenarios (such as video compression evaluation, platform recommendation optimization, etc.).

[0060] Use weighted average to combine multiple indicator scores and calculate an overall quality score:

[0061] The total video score = ∑(metric score × corresponding weight). Fusion can also employ learned nonlinear combinations (such as MLP) or fusion models with attention mechanisms to improve scoring accuracy and robustness. A unified comprehensive score is output as the final quality score for the target video. Meanwhile, the textual evaluation of each metric can be used to generate a comprehensive textual evaluation of the target video as a whole through the large model. Sub-scores can also be output for subsequent quality analysis, anomaly detection, and recommendation ranking. This ensures that all video quality dimensions are properly considered, ultimately generating a representative and interpretable comprehensive video score.

[0062] The video quality assessment method provided by an embodiment of the present invention extracts frames from a target video to obtain an image sequence; inputs the image sequence into a trained target model, which then outputs indicator scoring information corresponding to each video quality indicator of the target video, thereby obtaining multiple indicator scoring information; and determines a video scoring information for the target video based on the multiple indicator scoring information. Thus, the target model can generate scoring information for quality indicators across multiple dimensions of the video, and further generate video scoring information based on the scoring information for different quality indicators, thereby achieving automated scoring of videos based on multiple quality indicators, thereby efficiently, comprehensively, and accurately assessing the overall quality of the video.

[0063] Figure 2 A flow chart of another video quality assessment method provided by an embodiment of the present invention is shown as follows: Figure 2 As shown, the method specifically includes:

[0064] S21. Obtain the playback platform to which the target video belongs; and determine the type of video rating information according to the playback platform.

[0065] In this embodiment, different video quality scoring methods need to be adopted for different video playback platforms to obtain different types of video scoring information. The video scoring information may include a global score for the target video, or a local area score for the target video. The global score represents the overall score of the target video, and the local area score represents the score of the key area in the target video screen (for example, the area where the character's face is located, or the area where the target object is located).

[0066] Specifically, the types of video rating information corresponding to different playback platforms are pre-set. For example, the playback platforms may include: live broadcast platforms need to score the quality of key areas (faces or products) in the target video screen to ensure the quality of key areas that users need to watch during live broadcasts; video playback platforms need to globally score the target video to ensure the overall quality of the video.

[0067] S22: extract frames from the target video to obtain an image sequence.

[0068] In this embodiment, similar to step S11, please refer to Figure 1 For the sake of brevity, the relevant content will not be elaborated here.

[0069] S23. Input the image sequence into the trained target model, and output the indicator score and indicator evaluation corresponding to each picture perception indicator and each technical indicator through the target model; and use the indicator score and indicator evaluation as indicator scoring information.

[0070] In this embodiment, if Figure 3 FIG. 1 is a schematic diagram of a video quality indicator provided by an embodiment of the present invention. Figure 3 The video quality indicators may include: picture perception indicators (five first-level indicators, each of which contains multiple second-level indicators) and technical indicators (four first-level indicators, each of which contains multiple second-level indicators). Picture perception indicators include: 1. Clarity: resolution, detail, and edge sharpness; 2. Color: color reproduction, saturation, and color cast; 3. Dynamic range: contrast, dark area clarity, and highlight area clarity; 4. Theme perception: subject integrity, subject prominence, and subject recognizability; 5. Content experience: composition accuracy, subject clarity, and content richness. Technical indicators include: 1. Noise control: low-light noise, random noise, and aliasing noise; 2. False color: moiré, color banding, and color bleeding; 3. Texture fidelity: texture clarity, texture naturalness, texture continuity, and texture integrity; 4. Compression artifacts: blocking, ringing, color banding, and mosaic. Each video quality indicator can output the corresponding indicator score and indicator evaluation through the target model, where the indicator score represents the score of the quality of the current indicator, and the indicator evaluation represents the textual evaluation of the current indicator.

[0071] Specifically, the target model can use a large multimodal model (such as a large vision + text model), and can also equip images with OCR-recognized subtitles, scene text, or speech recognition results for content understanding. A multi-branch deep neural network is used, with each branch responsible for scoring different indicator dimensions; the backbone network is responsible for basic image feature extraction (such as ResNet and Swin Transformer); the branch networks include: 1. Picture perception module: evaluates five categories: clarity, color performance, dynamic range, theme perception, and content experience; 2. Technical indicator module: evaluates four categories: noise control, false color, texture fidelity, and compression artifacts.

[0072] Each sub-indicator under each dimension (more than 30 items in total) corresponds to an independent scoring output node; each output node returns two contents: 1. Indicator score: It can be a floating-point value, generally 0-1 or 0-100, used to represent the quality score of the item. 2. Indicator evaluation: It can be a natural language label (such as "good clarity", "slight color cast", etc.) automatically generated based on the score interval and the model's attention map. The model can be supervised by the label data (or expert scoring data) in the training set. Further, the indicator scores and indicator evaluations of all sub-indicators are combined into a structured output to obtain the indicator scoring information of each indicator. For example Figure 4 The figure shows a schematic diagram of an index score and index evaluation provided by an embodiment of the present invention, wherein 1.5-4.5 represent the index scores of the clarity index, and the text after each index score represents the index evaluation of clarity corresponding to the different index scores.

[0073] S24. When the video rating information is a global rating for the target video, obtain the indicator weight corresponding to each video quality indicator; determine the video score of the target video based on the indicator weight and indicator score corresponding to each video quality indicator; determine the video evaluation of the target video based on the indicator evaluation corresponding to each video quality indicator; and use the video score and video evaluation as the video rating information.

[0074] In this embodiment, when it is determined in step S21 that a global score is currently required for the target video, a weight is set for each video quality indicator based on the predefined indicator importance. The weight can be statically configured (e.g., based on ITU recommended standards or expert experience) or dynamically generated (e.g., adaptively adjusted based on video type), and all weights are summed to 1.

[0075] The weighted average method is used to calculate the global video score of the target video based on each indicator score and indicator weight. Video score = ∑i = 1n (indicator score i × indicator weight i), where i represents the i-th video quality indicator. The evaluation labels of each indicator (such as "good", "dark", "blurred", etc.) are collected, and a comprehensive evaluation of the entire video is generated using a rule engine or language model. For example: if multiple high-weight indicators are evaluated as "excellent", the comprehensive evaluation is "excellent overall picture quality"; if one or two key indicators (such as clarity, dynamic range) are evaluated poorly, then "the picture is slightly blurred, and the dynamic performance needs to be improved" is given. Finally, the video score (a numerical value) and the video evaluation (a natural language description) are combined to form the "video rating information" of the target video, which is used for output display or subsequent applications (such as review, recommendation, enhancement, etc.).

[0076] S25. When the video scoring information is a local area score for the target video, the target area of ​​each image in the image sequence is identified; the image sequence is input into a trained target model to output regional indicator scoring information corresponding to each regional quality indicator of the target area through the target model to obtain multiple regional indicator scoring information; and the regional scoring information of each target area is determined based on the multiple regional indicator scoring information.

[0077] In this embodiment, for each frame in the image sequence, a preprocessing model (such as an object detection or semantic segmentation network) is used to identify the local area that the current user wants to evaluate as the target area, such as the face area (the sensitive area of ​​the human eye), the main object (such as a person, a vehicle, or a product), the action area (an area of ​​intense movement), and specific scene elements (such as subtitles, logos, etc.). Each frame can be divided into several "target areas", each of which has location information (bounding box or mask) for subsequent model recognition.

[0078] Input the entire frame image and the "target area" it contains into the trained target model. The model analyzes each target area independently, extracts local features and evaluates the quality of the area. For each target area, the model outputs its corresponding various types of "regional quality index scoring information", including: regional index scores (such as the clarity score, noise score, etc. of the area) and regional index evaluations (such as "clear", "with artifacts", "underexposure", etc.). This step is similar to generating index scores and index evaluations in step S23. For details, please refer to the relevant description of S23. Among them, the regional quality index can be the same as the video quality index, and can also include but is not limited to: clarity, color performance, edge sharpness, regional noise, local compression artifacts, whether the face area is distorted, etc.

[0079] Multiple quality indicators of the same area are aggregated and analyzed to generate the overall scoring information of the area (regional score + regional evaluation), which is similar to the global scoring: regional score = ∑i = 1m (regional indicator score i × regional indicator weight i). Combined with the regional evaluation content, a text description is generated, such as: "The face area is clear, the color is saturated, and there is no obvious compression artifact", "There is slight motion blur in the action area", forming complete local regional scoring information. The scoring information of multiple target areas in each frame is combined to form the "local regional scoring information" of the entire image sequence; it can be used for subsequent tasks such as local enhancement suggestions, content quality diagnosis, and fine-grained recommendation.

[0080] In a possible implementation, determining the video rating information of the target video based on the multiple indicator rating information further includes:

[0081] When the indicator score of any video quality indicator is less than a first threshold, a visual explanation is generated according to the indicator evaluation corresponding to the video quality indicator; and video score information is determined according to the indicator score information and the visual explanation.

[0082] In this embodiment, the index scores of each generated video quality index are traversed and judged; when any index score is lower than a preset first threshold, it is considered that there is an obvious quality problem in the index, and the visual explanation mechanism is triggered. Based on the corresponding indicator evaluation (such as "blurry," "noise," or "color distortion"), the affected areas requiring special attention are identified. Visual explanations are generated using the following methods: Heatmap overlay: Based on the model's internal attention mechanism or feature response map, a high-response heatmap of the area related to the quality issue is generated and overlaid on the original image. Key area annotation: The problematic area (such as blurred faces or distorted backgrounds) is framed on the original image and labeled with text descriptions. Text explanation generation: Quality diagnostic text is automatically generated for the frame or area, such as: "The central face area of ​​the image is noticeably blurred, affecting the viewing experience." Visual explanation content for the indicator is appended to the existing indicator scoring information. Video scoring information is defined as consisting of the following components: an indicator score set (such as scores for nine dimensions), an indicator evaluation set (such as "bright colors" and "artifacts present"), and a visual explanation set (such as "noise exists in frame 25, as annotated on the heatmap below"). This structure improves the interpretability and user trust of the score, while providing direct guidance for subsequent optimizations (such as video restoration and content filtering).

[0083] In one possible implementation, when the video score in the video rating information is less than a second threshold, a target video quality indicator is determined from multiple video quality indicators, and the indicator score of the target video quality indicator is less than a third threshold; a video adjustment strategy corresponding to the target video quality indicator is determined; and the target video is adjusted according to the video adjustment strategy.

[0084] In this embodiment, the video score in the video rating information is obtained, and it is determined whether the video score is lower than a preset second threshold. If it is lower than the second threshold, it indicates that the overall video quality does not meet the requirements, and the quality optimization process is triggered. Multiple video quality indicators are traversed to find the indicator item with an indicator score lower than the third threshold, which is defined as the target video quality indicator. For each target video quality indicator, a predefined video adjustment strategy library is queried; for example: low clarity applies a super-resolution enhancement algorithm; color distortion applies color correction or automatic white balance adjustment; excessive noise applies a temporal filter or image denoising algorithm; severe artifacts apply codec optimization or intra-frame reconstruction technology; the strategy library can be based on empirical rules or constructed through training data.

[0085] Based on the determined adjustment strategy, the corresponding video processing module is called to perform regional or overall adjustments to the target video. Based on the visual interpretation of the target indicators, the adjustment area can be located and targeted optimization can be performed. The adjusted video quality is then re-evaluated to form a closed loop.

[0086] The video quality assessment method provided by the embodiment of the present invention combines picture perception indicators and technical indicators to realize a multimodal scoring mechanism, which not only evaluates physical properties such as clarity and noise, but also comprehensively analyzes content quality to provide a score that is closer to human subjective perception. At the same time, based on the large model, it can adapt to data sets of different distortion types and different fields, and can effectively evaluate both low-level visual damage (such as blur, color cast) and high-level semantic impact (such as poor content quality, pornography and violence, etc.). At the same time, the significance analysis distortion positioning is introduced to enable users to intuitively understand which areas affect the final quality score. The scoring basis is transparent, which enhances user trust and is suitable for high-demand scenarios such as content review and quality control.

[0087] Figure 5 A schematic diagram of the structure of a video quality assessment device provided by an embodiment of the present invention is shown in FIG. Figure 5 As shown, the device specifically includes:

[0088] The video processing module 51 is used to perform frame extraction processing on the target video to obtain an image sequence;

[0089] An information output module 52 is configured to input the image sequence into a trained target model, so as to output the indicator scoring information corresponding to each video quality indicator of the target video through the target model, thereby obtaining a plurality of indicator scoring information;

[0090] The determination module 53 is configured to determine the video score information of the target video according to the plurality of indicator score information.

[0091] In a possible implementation, the information output module is specifically configured to output, through the target model, an indicator score and an indicator evaluation corresponding to each of the picture perception indicators and each of the technical indicators;

[0092] The indicator score and the indicator evaluation are used as the indicator scoring information.

[0093] In a possible implementation, the determining module is specifically configured to obtain an indicator weight corresponding to each of the video quality indicators;

[0094] Determine the video score of the target video according to the indicator weight and indicator score corresponding to each of the video quality indicators;

[0095] Determine the video rating of the target video according to the indicator evaluation corresponding to each of the video quality indicators;

[0096] The video score and the video evaluation are used as the video rating information.

[0097] In one possible implementation, the recognition module 54 is configured to recognize a target region in each image in the image sequence;

[0098] The information output module is further configured to input the image sequence into a trained target model, so as to output regional indicator scoring information corresponding to each regional quality indicator of the target region through the target model, thereby obtaining a plurality of regional indicator scoring information;

[0099] The determining module is further configured to determine the regional scoring information of each target area based on the plurality of regional indicator scoring information.

[0100] In one possible implementation, the determining module is specifically configured to generate a visual explanation based on an indicator evaluation corresponding to the video quality indicator when the indicator score of any of the video quality indicators is less than a first threshold;

[0101] The video scoring information is determined based on the indicator scoring information and the visual interpretation.

[0102] In one possible implementation, the determination module is further configured to, when the video score in the video rating information is less than a second threshold, determine a target video quality indicator from the plurality of video quality indicators, wherein the indicator score of the target video quality indicator is less than a third threshold;

[0103] Determining a video adjustment strategy corresponding to the target video quality indicator;

[0104] The target video is adjusted according to the video adjustment strategy.

[0105] In a possible implementation, the acquisition module 55 is configured to acquire the playback platform to which the target video belongs;

[0106] The determination module is further configured to determine a type of the video rating information according to the playback platform, where the type includes: a global rating for the target video, or a local area rating for the target video.

[0107] The video quality assessment device provided in this embodiment can be as follows Figure 5 The device shown in , can perform the following Figure 1-2 All steps of the video quality assessment method in order to achieve Figure 1-2 For details on the technical effects of the video quality assessment method shown, please refer to Figure 1-2 For the sake of brevity, the relevant description will not be repeated here.

[0108] Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of the present invention is provided. Figure 6 The computer device 600 shown includes: at least one processor 601, memory 602, at least one network interface 604 and other user interfaces 603. The various components in the computer device 600 are coupled together via a bus system 605. It is understood that the bus system 605 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 605 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 605 is not described in detail. Figure 6 Various buses are labeled as bus system 605.

[0109] The user interface 603 may include a display, a keyboard, or a pointing device (eg, a mouse, a trackball, a touchpad, or a touch screen).

[0110] It is understood that the memory 602 in the embodiment of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDRSDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 602 described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0111] In some embodiments, the memory 602 stores the following elements, executable units, or data structures, or a subset thereof, or an extended set thereof: an operating system 6021 and application programs 6022 .

[0112] The operating system 6021 includes various system programs, such as a framework layer, a core library layer, and a driver layer, for implementing various basic services and handling hardware-based tasks. Application programs 6022 include various application programs, such as a media player and a browser, for implementing various application services. Programs implementing the methods of the embodiments of the present invention may be included in application programs 6022.

[0113] In an embodiment of the present invention, by calling a program or instruction stored in the memory 602, specifically, a program or instruction stored in the application 6022, the processor 601 is configured to execute the method steps provided in each method embodiment, for example, including:

[0114] Perform frame extraction on the target video to obtain an image sequence;

[0115] Inputting the image sequence into a trained target model, so as to output the indicator scoring information corresponding to each video quality indicator of the target video through the target model, thereby obtaining a plurality of indicator scoring information;

[0116] The video rating information of the target video is determined based on the plurality of indicator rating information.

[0117] In a possible implementation, the target model outputs an indicator score and an indicator evaluation corresponding to each of the picture perception indicators and each of the technical indicators;

[0118] The indicator score and the indicator evaluation are used as the indicator scoring information.

[0119] In a possible implementation, obtaining an indicator weight corresponding to each of the video quality indicators;

[0120] Determine the video score of the target video according to the indicator weight and indicator score corresponding to each of the video quality indicators;

[0121] Determine the video rating of the target video according to the indicator evaluation corresponding to each of the video quality indicators;

[0122] The video score and the video evaluation are used as the video rating information.

[0123] In one possible implementation, identifying a target region of each image in the image sequence;

[0124] Inputting the image sequence into a trained target model, so as to output regional indicator scoring information corresponding to each regional quality indicator of the target region through the target model, thereby obtaining a plurality of regional indicator scoring information;

[0125] The regional scoring information of each target area is determined according to the plurality of regional index scoring information.

[0126] In one possible implementation, when the index score of any of the video quality indicators is less than a first threshold, generating a visual explanation based on the index evaluation corresponding to the video quality indicator;

[0127] The video scoring information is determined based on the indicator scoring information and the visual interpretation.

[0128] In one possible implementation, when the video score in the video rating information is less than a second threshold, determining a target video quality indicator from the plurality of video quality indicators, wherein the indicator score of the target video quality indicator is less than a third threshold;

[0129] Determining a video adjustment strategy corresponding to the target video quality indicator;

[0130] The target video is adjusted according to the video adjustment strategy.

[0131] In one possible implementation, obtaining a playback platform to which the target video belongs;

[0132] The type of the video rating information is determined according to the playback platform, where the type includes: a global rating for the target video, or a local area rating for the target video.

[0133] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 601. Processor 601 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in processor 601 or by software instructions. The above processor 601 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present invention can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software units in the decoding processor. The software units can be located in storage media well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 602 , and the processor 601 reads the information in the memory 602 and completes the steps of the above method in combination with its hardware.

[0134] It is understood that the embodiments described herein may be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSP devices, DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or a combination thereof.

[0135] For software implementation, the technology described herein can be implemented by a unit that performs the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0136] The computer device provided in this embodiment may be Figure 6 The computer device shown in , can execute Figure 1-2 All steps of the video quality assessment method in order to achieve Figure 1-2 For details on the technical effects of the video quality assessment method shown, please refer to Figure 1-2 For the sake of brevity, the relevant description will not be repeated here.

[0137] An embodiment of the present invention further provides a storage medium (computer-readable storage medium). The storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; and the memory may also include a combination of the aforementioned types of memory.

[0138] When one or more programs in the storage medium can be executed by one or more processors, the video quality assessment method executed on the device side can be implemented.

[0139] The processor is configured to execute a video quality assessment program stored in the memory to implement the following steps of a video quality assessment method performed on a device side:

[0140] Perform frame extraction on the target video to obtain an image sequence;

[0141] Inputting the image sequence into a trained target model, so as to output the indicator scoring information corresponding to each video quality indicator of the target video through the target model, thereby obtaining a plurality of indicator scoring information;

[0142] The video rating information of the target video is determined based on the plurality of indicator rating information.

[0143] In a possible implementation, the target model outputs an indicator score and an indicator evaluation corresponding to each of the picture perception indicators and each of the technical indicators;

[0144] The indicator score and the indicator evaluation are used as the indicator scoring information.

[0145] In a possible implementation, obtaining an indicator weight corresponding to each of the video quality indicators;

[0146] Determine the video score of the target video according to the indicator weight and indicator score corresponding to each of the video quality indicators;

[0147] Determine the video rating of the target video according to the indicator evaluation corresponding to each of the video quality indicators;

[0148] The video score and the video evaluation are used as the video rating information.

[0149] In one possible implementation, identifying a target region of each image in the image sequence;

[0150] Inputting the image sequence into a trained target model, so as to output regional indicator scoring information corresponding to each regional quality indicator of the target region through the target model, thereby obtaining a plurality of regional indicator scoring information;

[0151] The regional scoring information of each target area is determined according to the plurality of regional index scoring information.

[0152] In one possible implementation, when the index score of any of the video quality indicators is less than a first threshold, generating a visual explanation based on the index evaluation corresponding to the video quality indicator;

[0153] The video scoring information is determined based on the indicator scoring information and the visual interpretation.

[0154] In one possible implementation, when the video score in the video rating information is less than a second threshold, determining a target video quality indicator from the plurality of video quality indicators, wherein the indicator score of the target video quality indicator is less than a third threshold;

[0155] Determining a video adjustment strategy corresponding to the target video quality indicator;

[0156] The target video is adjusted according to the video adjustment strategy.

[0157] In one possible implementation, obtaining a playback platform to which the target video belongs;

[0158] The type of the video rating information is determined according to the playback platform, where the type includes: a global rating for the target video, or a local area rating for the target video.

[0159] Professionals should also be further aware that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0160] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0161] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A video quality assessment method, characterized in that: include: Perform frame extraction on the target video to obtain an image sequence; Inputting the image sequence into a trained target model, so as to output the indicator scoring information corresponding to each video quality indicator of the target video through the target model, thereby obtaining a plurality of indicator scoring information; The video rating information of the target video is determined based on the plurality of indicator rating information.

2. The method according to claim 1, characterized in that The video quality indicators include: multiple picture perception indicators and multiple technical indicators; Outputting the indicator scoring information corresponding to each video quality indicator of the target video through the target model includes: Outputting the index score and index evaluation corresponding to each of the picture perception indexes and each of the technical indicators through the target model; The indicator score and the indicator evaluation are used as the indicator scoring information.

3. The method according to claim 2, characterized in that When the video rating information is a global rating for the target video, determining the video rating information of the target video based on the plurality of indicator rating information includes: Obtaining an indicator weight corresponding to each of the video quality indicators; Determine the video score of the target video according to the indicator weight and indicator score corresponding to each of the video quality indicators; Determine the video rating of the target video according to the indicator evaluation corresponding to each of the video quality indicators; The video score and the video evaluation are used as the video rating information.

4. The method according to claim 1, wherein When the video rating information is a local area rating for the target video, the method further includes: Identifying a target region of each image in the image sequence; Inputting the image sequence into a trained target model, so as to output regional indicator scoring information corresponding to each regional quality indicator of the target region through the target model, thereby obtaining a plurality of regional indicator scoring information; The regional scoring information of each target area is determined according to the plurality of regional index scoring information.

5. The method according to claim 2, characterized in that The determining the video score information of the target video according to the plurality of indicator score information includes: When the index score of any of the video quality indicators is less than a first threshold, generating a visual explanation according to the index evaluation corresponding to the video quality indicator; The video scoring information is determined based on the indicator scoring information and the visual interpretation.

6. The method according to claim 1, characterized in that The method further comprises: When the video score in the video rating information is less than a second threshold, determining a target video quality index from the plurality of video quality indexes, wherein the index score of the target video quality index is less than a third threshold; Determining a video adjustment strategy corresponding to the target video quality indicator; The target video is adjusted according to the video adjustment strategy.

7. The method according to claim 2, characterized in that Before performing frame extraction processing on the target video to obtain an image sequence, the method further includes: Obtain the playback platform to which the target video belongs; The type of the video rating information is determined according to the playback platform, where the type includes: a global rating for the target video, or a local area rating for the target video.

8. A video quality scoring device, characterized in that: include: The video processing module is used to extract frames from the target video to obtain an image sequence; An information output module is used to input the image sequence into a trained target model to output the indicator scoring information corresponding to each video quality indicator of the target video through the target model to obtain multiple indicator scoring information; The determination module is used to determine the video score information of the target video according to the multiple indicator score information.

9. A computer device, characterized in that: include: A processor and a memory, wherein the processor is configured to execute a video quality scoring program stored in the memory to implement the video quality assessment method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the video quality assessment method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Quality evaluation method and device and electronic equipment

    CN112950581A

  • Video quality evaluation method and device, equipment, storage medium and program product

    CN116471262A

  • Image quality evaluation method and system, electronic equipment and storage medium

    CN119671977A

  • Video quality assessment model training method, quality assessment method, device and medium

    CN119763018A

Cited By

  • Multi-mode video content retrieval method and device based on shot frame sampling and medium

    CN121256087A

  • User generated video quality evaluation method, program, equipment and storage medium

    CN121415306A