Video quality determination method and device, equipment and storage medium
By extracting key image blocks from video frames and performing underlying distortion and semantic features extraction, combined with video evaluation model, the problem of poor video quality evaluation in the prior art is solved, and a more accurate and efficient video quality evaluation is achieved.
Patent Information
- Application Number
- CN202410099897.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-23
- Publication Date
- 2025-07-25
AI Technical Summary
The existing video quality evaluation model without reference cannot effectively extract and analyze the video feature, resulting in poor video evaluation results.
Key image blocks that meet the motion and texture conditions are extracted from the video frame of the target video, and the underlying distortion characteristics and semantic features are extracted, and the video quality scores are determined based on the feature extraction network of the video evaluation model.
The accuracy and efficiency of video quality evaluation are improved, and the video quality is comprehensively and accurately determined by comprehensively considering the underlying distortion characteristics and semantic characteristics.
Smart Images

Figure CN120378646A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method, apparatus, device, and storage medium for determining video quality. Background Art
[0002] In recent years, with the rapid development of digital media and Internet technologies, video content has become an indispensable part of people's daily lives. Since video signals may suffer quality degradation at various stages of acquisition, compression, transmission, and display, the quality of video directly affects the conveyance effect of video content.
[0003] As an important evaluation method, no-reference video quality assessment (NR-VQA) plays an increasingly important role in the field of video quality assessment due to its characteristic of not relying on the original undistorted video, especially in situations where the original, undistorted video cannot be obtained as a reference.
[0004] Existing deep learning-based video evaluation models used for no-reference video quality assessment usually cannot effectively extract and analyze video features, resulting in poor video evaluation effects. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, apparatus, device, and storage medium for determining video quality that can improve the effect of video evaluation.
[0006] In a first aspect, this application provides a method for determining video quality. The method includes:
[0007] Extracting key image blocks that meet the motion condition and texture condition from video frames of the target video;
[0008] Extracting a global mosaic image from video frames of the target video;
[0009] Performing low-level distortion feature extraction on the global mosaic image to obtain low-level distortion features of the target video;
[0010] Performing semantic feature extraction on the key image blocks to obtain semantic features of the target video that include target texture information and motion information;
[0011] Determining a quality score of the target video based on the low-level distortion features and the semantic features.
[0012] In a second aspect, this application further provides a device for determining video quality. The device includes:
[0013] A key image block extraction module, configured to extract key image blocks that meet the motion condition and texture condition from video frames of the target video;
[0014] A global stitching image extraction module, configured to extract a global stitching image from video frames of the target video;
[0015] A low-level distortion feature extraction module, configured to extract low-level distortion features from the global stitching image to obtain the low-level distortion features of the target video;
[0016] A semantic feature extraction module, configured to extract semantic features from the key image blocks to obtain semantic features of the target video including target texture information and motion information;
[0017] A quality score determination module, configured to determine a quality score of the target video based on the low-level distortion features and the semantic features.
[0018] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented:
[0019] Extract key image blocks that meet motion conditions and texture conditions from video frames of a target video;
[0020] Extract a global stitching image from video frames of the target video;
[0021] Extract low-level distortion features from the global stitching image to obtain the low-level distortion features of the target video;
[0022] Extract semantic features from the key image blocks to obtain semantic features of the target video including target texture information and motion information;
[0023] Determine a quality score of the target video based on the low-level distortion features and the semantic features.
[0024] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0025] Extract key image blocks that meet motion conditions and texture conditions from video frames of a target video;
[0026] Extract a global stitching image from video frames of the target video;
[0027] Extract low-level distortion features from the global stitching image to obtain the low-level distortion features of the target video;
[0028] Extract semantic features from the key image blocks to obtain semantic features of the target video that include target texture information and motion information;
[0029] Determine the quality score of the target video based on the underlying distortion features and the semantic features.
[0030] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the following steps:
[0031] Extract key image blocks that meet the motion condition and texture condition from the video frames of the target video;
[0032] Extract a global stitched image from the video frames of the target video;
[0033] Extract underlying distortion features from the global stitched image to obtain the underlying distortion features of the target video;
[0034] Extract semantic features from the key image blocks to obtain semantic features of the target video that include target texture information and motion information;
[0035] Determine the quality score of the target video based on the underlying distortion features and the semantic features.
[0036] The above method, apparatus, computer device, storage medium, and computer program product for determining video quality extract key image blocks that meet the motion condition and texture condition from the video frames of the target video; key image blocks are usually the places where the video quality loss is most obvious. By extracting semantic features from the key image blocks, semantic features of the target video that include target texture information and motion information are obtained. When determining the video quality based on the semantic features subsequently, the accuracy and efficiency of video quality determination can be improved; a global stitched image is extracted from the video frames of the target video. The global stitched image carries the global loss information of the target video. By extracting underlying distortion features from the global stitched image, the underlying distortion features of the target video are obtained. When determining the video quality based on the underlying distortion features subsequently, the accuracy of video quality determination can be improved, and a large amount of calculation required for directly analyzing the original video frames is avoided during the extraction of the underlying distortion features, and at the same time, the efficiency of video quality determination is further improved; the quality score of the target video is determined based on the underlying distortion features and the semantic features. When determining the video quality, the underlying distortion features and the semantic features are comprehensively considered, that is, the feature dimensions for determining the video quality are more comprehensive and accurate, thereby further improving the accuracy of video quality determination. Description of the Drawings
[0037] Figure 1It is an application environment diagram of a method for determining video quality in an embodiment;
[0038] Figure 2 It is a schematic flowchart of a method for determining video quality in an embodiment;
[0039] Figure 3 It is a schematic diagram of a block processing process in an embodiment;
[0040] Figure 4 It is a schematic diagram of the distribution of sample videos in an embodiment;
[0041] Figure 5 It is a schematic flowchart of the steps for constructing sample videos in an embodiment;
[0042] Figure 6 It is a schematic flowchart of a method for determining video quality in another embodiment;
[0043] Figure 7 It is a schematic diagram of the structure of a video evaluation model in an embodiment;
[0044] Figure 8 It is a schematic flowchart of the steps for training a video evaluation model in an embodiment;
[0045] Figure 9 It is a schematic diagram of the performance of a video evaluation model in an embodiment;
[0046] Figure 10 It is a schematic diagram of the performance of a video evaluation model in another embodiment;
[0047] Figure 11 It is a schematic diagram of the performance of a video evaluation model in another embodiment;
[0048] Figure 12 It is a schematic diagram of the performance of a video evaluation model in another embodiment;
[0049] Figure 13 It is a structural block diagram of a device for determining video quality in an embodiment;
[0050] Figure 14 It is a structural block diagram of a device for determining video quality in an embodiment;
[0051] Figure 15 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0052] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0053] The video quality determination method provided by the embodiments of this application relates to technologies such as machine learning in artificial intelligence, where:
[0054] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0055] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models, also known as large models or foundation models, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0056] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Pre-trained models are the latest development results of deep learning, integrating the above technologies.
[0057] The video quality determination method provided by the embodiments of this application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set separately, integrated on the server 104, placed on the cloud or other devices. The method for determining the video quality can be executed independently by the terminal 102 or the server 104, or executed collaboratively by the terminal 102 and the server 104. In some embodiments, the method for determining the video quality is executed by the terminal 102. The terminal 102 extracts key image blocks that meet the motion condition and texture condition from the video frames of the target video; extracts a global stitched image from the video frames of the target video; extracts low-level distortion features from the global stitched image to obtain the low-level distortion features of the target video; extracts semantic features from the key image blocks to obtain semantic features of the target video including target texture information and motion information; determines the quality score of the target video based on the low-level distortion features and the semantic features.
[0058] Among them, the terminal can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, portable wearable devices, and network devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The network devices can be routers, switches, firewalls, load balancers, network storage devices, network adapters, etc.
[0059] The server 104 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.
[0060] In one embodiment, the method for determining the video quality is executed by a computer device, and the computer device can be, for example, Figure 1 the terminal 102 or the server 104 shown. As Figure 2 shown, the method may include the following steps:
[0061] S202, extract key image blocks that meet the motion condition and texture condition from the video frames of the target video.
[0062] Among them, the target video is the video to be evaluated for quality. Specifically, it can be a video obtained from a movie or a TV program for evaluating post-production and encoding quality, or a video obtained from a video platform for evaluating the impact of compression and transmission on video quality, or a video obtained by an acquisition device for evaluating the video quality of the acquisition device.
[0063] The motion condition and the texture condition are the criteria for screening image blocks in video frames. In video processing and analysis, they are used to determine which image blocks are important for understanding and evaluating video content; the motion condition is used to identify regions with significant motion in video frames; the texture condition is used to screen out regions that are significant in terms of texture features. The key image blocks can be the image blocks in the video frame that meet the combined conditions of the motion condition and the texture condition, or the image blocks in the video frame that meet both the motion condition and the texture condition simultaneously.
[0064] Specifically, the computer device can perform frame extraction on the target video to obtain the video frames of the target video. For any one video frame, it performs block processing on it to obtain the respective image blocks in the video frame, and conducts content analysis on each image block to obtain the content analysis result. Based on the content analysis result, it selects the key image blocks that meet the motion condition and the texture condition from the respective image blocks of the video frame. The number of video frames of the target video extracted is at least one.
[0065] Reference Figure 3 to the schematic diagram of the block processing process shown in Figure 3 (A) in it is the original image of a certain video frame, Figure 3 (B) in it is the respective image blocks obtained after performing block processing on the video frame.
[0066] S204, extract the global mosaic image from the video frames of the target video.
[0067] The global mosaic image refers to an image that carries the overall visual content in the video frame to which it belongs. Specifically, it can be a mosaic image formed by combining the content or features in different regions of the video frame.
[0068] Specifically, the computer device can perform frame extraction on the target video to obtain the video frames of the target video. For any one video frame, it performs block processing on it to obtain the respective image blocks in the video frame, and extracts image sub-blocks from each image block, and splices the extracted image sub-blocks to obtain the global mosaic image of the video frame. By performing the above processing on each of the extracted video frames, the global mosaic images corresponding to each of the extracted video frames can be obtained. The number of video frames of the target video extracted is at least one.
[0069] S206. Extract the low-level distortion features from the globally stitched image to obtain the low-level distortion features of the target video.
[0070] Among them, the low-level distortion features include the distortions in the image caused by reasons such as compression, transmission errors, insufficient sampling, etc., such as blurring, noise, blocking effect, color distortion, etc.
[0071] Specifically, after the computer device obtains the globally stitched image corresponding to the extracted video frame, it can input each globally stitched image into the video evaluation model according to the frame sequence of each extracted video frame, and extract the low-level distortion features through the feature extraction network of the video evaluation model to obtain the low-level distortion features of the target video.
[0072] Among them, the video evaluation model is a model for evaluating the quality of video content. The video rating model in the embodiments of the present application belongs to a no-reference model (No Reference Model, NR), that is, the video evaluation model does not require the original uncompressed or processed video as a reference when evaluating and analyzing the video. The video evaluation model can be a machine learning model, specifically a model obtained by training using deep learning technology. The video evaluation model can include a feature extraction network, and the feature extraction network is used to extract feature information from the input data to evaluate the video quality.
[0073] S208. Extract the semantic features of the key image blocks to obtain the semantic features of the target video including the target texture information and motion information.
[0074] Among them, the semantic features refer to high-level features that can describe the content of the image. The semantic features including the target texture information and motion information refer to high-level features that describe the surface features (texture) and dynamic characteristics (motion) of the objects in the video. The target texture information can also be called spatial texture information.
[0075] Specifically, after the computer device obtains the key image blocks corresponding to each extracted video frame, it can input each key image block into the video evaluation model according to the frame sequence of each extracted video frame, and extract the semantic features through the feature extraction network of the video evaluation model to obtain the semantic features of the target video.
[0076] Among them, the video evaluation model is a model for evaluating the quality of video content. The video rating model in the embodiments of the present application belongs to a no-reference model (No Reference Model, NR), that is, the video evaluation model does not require the original uncompressed or processed video as a reference when evaluating and analyzing the video. The video evaluation model can be a machine learning model, specifically a model obtained by training using deep learning technology. The video evaluation model can include a feature extraction network, and the feature extraction network is used to extract feature information from the input data to evaluate the video quality.
[0077] S210. Determine the quality score of the target video based on the underlying distortion features and semantic features.
[0078] Among them, the quality score is a predicted score regarding the quality of the target video, reflecting the overall quality of the target video.
[0079] Specifically, after obtaining the underlying distortion features and semantic features of the target video, the computer device can input the underlying distortion features and semantic features of the target video into the prediction network of the video evaluation model. The prediction network processes the underlying distortion features and semantic features and outputs the quality score of the target video.
[0080] In one embodiment, the prediction network can be a multi-layer perceptron structure (MLP). The process of obtaining the quality score of the target video by processing the underlying distortion features and semantic features through the prediction network can be expressed as follows:
[0081] y = sigmod(MLP(F S + F D ))
[0082] Among them, y represents the predicted quality score, F S represents the semantic features, F D represents the underlying distortion features, MLP represents the multi-layer perceptron, and sigmod is the activation function.
[0083] In one embodiment, the computer device can also perform feature fusion on the underlying distortion features and semantic features of the target video to obtain fused features, and determine the quality score of the target video based on the fused features. Among them, feature fusion can be achieved in various ways, including but not limited to simple concatenation, weighted average, feature mapping, and transformation, etc.
[0084] In the above-mentioned method for determining video quality, the computer device extracts key image blocks that meet motion conditions and texture conditions from the video frames of the target video; the key image blocks are usually the places where the video quality loss is most obvious. By extracting semantic features from the key image blocks, semantic features of the target video containing target texture information and motion information are obtained. When the video quality is subsequently determined based on the semantic features, the accuracy and efficiency of the video quality determination can be improved; a global spliced image is extracted from the video frames of the target video, and the global spliced image carries the global loss information of the target video. The underlying distortion features are extracted from each global spliced image to obtain the underlying distortion features of the target video. When the video quality is subsequently determined based on the underlying distortion features, the accuracy of the video quality determination can be improved. When the underlying distortion features are extracted, the large amount of calculations required when directly analyzing the original video frames are avoided, and the efficiency of the video quality determination is further improved; the quality score of the target video is determined based on the underlying distortion features and the semantic features. When the video quality is determined, the underlying distortion features and the semantic features are comprehensively considered, that is, the feature dimensions of the video quality determination are more comprehensive and accurate, thereby further improving the accuracy of the video quality determination.
[0085] In one embodiment, the process of extracting key image blocks that meet motion conditions and texture conditions from video frames of a target video by a computer device includes the following steps: extracting at least two first video frames from the target video according to a first frame interval; performing block processing on each first video frame to obtain at least two image blocks corresponding to each first video frame; and selecting key image blocks that meet motion conditions and texture conditions from the at least two image blocks of each first video frame.
[0086] The first frame interval is an inter-frame interval of video frames extracted from the target video, which specifies how many frames should be spaced between the first video frames selected for analysis in the continuous video frame sequence of the target video.
[0087] Block processing refers to dividing the video frame into several smaller image regions or blocks, each of which can be analyzed separately.
[0088] Specifically, the computer device performs frame extraction processing on the target video through a video processing application. During the frame extraction process, a video frame is extracted from the target video at every first frame interval, thereby obtaining at least two first video frames. After completing the frame extraction process on the target video, a first video frame sequence consisting of at least two first video frames can be obtained. For any first video frame in the first video frame sequence, the first video frame is block-processed according to a first block scheme to obtain at least two image blocks of the first video frame, and content analysis is performed on each image block to obtain a content analysis result. Based on the content analysis result, key image blocks that meet motion conditions and texture conditions are selected from each image block of the video frame.
[0089] The first block scheme defines the size, shape, and segmentation type of the image blocks obtained by segmentation. The segmentation type includes uniform segmentation and non-uniform segmentation. In the embodiments of the present application, uniform segmentation can be adopted. The shape of the image blocks can be rectangular, and the size of the image blocks can be 224*224.
[0090] In the above embodiments, the computer device extracts at least two first video frames from the target video at a first frame interval, which can reduce the number of subsequent processed video frames while retaining enough video frames to evaluate the video quality, thereby improving the efficiency of video quality determination. Each first video frame is separately block-processed to obtain at least two image blocks corresponding to each first video frame. From the at least two image blocks of each first video frame, key image blocks that meet the motion condition and texture condition are selected. Thus, the key image blocks that are most likely to exhibit quality loss in the video frames can be identified and extracted. When determining the video quality based on the key image blocks subsequently, the amount of calculation is further reduced while ensuring the accuracy of video quality determination, which can further improve the efficiency of video quality determination.
[0091] In one embodiment, the process of the computer device selecting key image blocks that meet the motion condition and texture condition from the at least two image blocks of each first video frame includes the following steps: extracting content information from the at least two image blocks of each first video frame respectively to obtain the content information of each image block; based on the content information of each image block, selecting key image blocks that meet the motion condition and texture condition from the at least two image blocks of each first video frame.
[0092] Among them, the content information includes information such as the motion characteristics and texture characteristics of the image blocks.
[0093] Specifically, for any one first video frame, after the computer device obtains each image block of the first video frame, it uses a content information extraction algorithm to extract the content information of each image block respectively, obtains the content information of each image block, and selects the N image blocks with the richest content information from each image block according to the content information. The first N image blocks are the key image blocks that meet the motion condition and texture condition selected from the first video frame.
[0094] Selecting the N image blocks with the richest content information from each image block according to the content information can specifically be: sorting each image block from largest to smallest according to the richness of the content information to obtain a sorting result, and selecting the image blocks ranked in the top N based on the sorting result.
[0095] In the above embodiments, the computer device extracts the content information of each of at least two image blocks of each first video frame to obtain the content information of each image block; based on the content information of each image block, key image blocks that meet the motion condition and the texture condition are selected from at least two image blocks of each first video frame, so that the area in the video frame that is most likely to exhibit quality loss can be determined. Thus, subsequent video quality analysis can be based only on the determined area that is most likely to exhibit quality loss, avoiding a comprehensive analysis of the entire video frame, and thereby improving the accuracy and efficiency of video quality determination.
[0096] In one embodiment, the process of the computer device extracting the content information of each of at least two image blocks of each first video frame to obtain the content information of each image block includes the following steps: extracting the texture information of each of at least two image blocks of each first video frame to obtain the spatial texture information of each image block; in the prior video frame of the first video frame, obtaining the target image block corresponding to each image block; determining the motion change information of each image block based on the difference between each image block and the corresponding target image block; and fusing the spatial texture information and the motion change information to obtain the content information of each image block.
[0097] The prior video frame of the first video frame refers to at least one video frame that is adjacent to the first video frame and is before the first video frame in the target video. For example, if the first video frame is the 10th frame in the target video, the prior video frame of this first video frame can be the 9th frame in the target video.
[0098] The target image block corresponding to each image block in the first video frame is the image block in the prior video frame that corresponds to the position of each image block respectively. Specifically, after obtaining the prior video frame, the prior video frame can be block-processed using the first block scheme to obtain at least two image blocks of the prior video frame. For any one image block in the first video frame, the image block in the prior video frame that is in the same position as this image block is determined as the target image block of this image block.
[0099] Specifically, for any image block of a certain first video frame, the computer device performs texture analysis on it using a preset texture analysis algorithm to obtain the spatial texture information of the image block; and obtains the prior video frame of the first video frame, and performs block processing on the prior video frame using a first block scheme to obtain at least two image blocks of the prior video frame. For any image block in the first video frame, the image block in the prior video frame that is in the same position as the image block is determined as the target image block of the image block, and a preset motion analysis algorithm is used to determine the difference between the image block and the corresponding target image block, and the motion change information of the image block is determined according to the difference, and the first weight corresponding to the spatial texture information and the second weight corresponding to the motion change information are obtained, and the spatial texture information and the motion change information are fused according to the first weight and the second weight to obtain the content information of the image block; the above processing is performed on each image block of each first video frame, so that the content information of each image block can be obtained.
[0100] Among them, the preset texture analysis algorithm can specifically be algorithms such as gray-level co-occurrence matrix (GLCM), local binary pattern (LBP), or Sobel edge detection. The preset motion analysis algorithm can specifically be algorithms such as frame difference method, optical flow method, etc.
[0101] For example, for the i-th image block in the first video frame (i.e., the current first video frame) with frame sequence n in the target video frame, the Sobel edge detection algorithm is used to perform texture detection on the i-th image block to obtain the spatial texture information of the i-th image block. This process can be expressed as follows:
[0102] SI i = std(Sobel(Frame i ))
[0103] Among them, SI i represents the spatial texture information of the i-th image block, which can also be called the spatial texture index. Frame i represents the i-th image block. Sobel(Frame i ) represents using the Sobel operator to perform edge detection on the i-th image block. std(Sobel(Frame i )) represents applying the standard deviation to the result of the Sobel operator, that is, calculating the degree of dispersion of the edge intensity values. A high SI i value usually indicates that the image block has more complex or more obvious texture features, and a low SI i value usually indicates that the image block is relatively smooth or the texture change is not obvious.
[0104] The frame difference method is used to detect the motion transformation information of the i-th image block to obtain the motion change information of the i-th image block. This process can be expressed as follows:
[0105]
[0106] Among them, TI i represents the motion change information of the i-th image block, represents the i-th image block in the first video frame (i.e., the current first video frame) with frame sequence n in the target video frame, represents the i-th image block in the previous video frame of the first video frame with frame sequence n, represents the pixel difference of the i-th image block between the n-th frame and the (n - 1)-th frame, represents the result of applying the standard deviation to the inter-frame difference, that is, the change degree of the i-th image block between adjacent frames. A high TI i usually indicates that the image block changes greatly over time, such as fast movement or significant scene changes, and a low TI i usually indicates that the image block does not change much between these two frames.
[0107] Fuse the spatial texture information and motion change information of the i-th image block to obtain the content information of the i-th image block. This process can be expressed as follows:
[0108] STI i = α * SI i + (1 - α) * TI i
[0109] Among them, STI i represents the content information of the i-th image block, SI i represents the spatial texture information of the i-th image block, TI i represents the motion change information of the i-th image block, α represents the first weight corresponding to the spatial texture information, and (1 - α) represents the second weight corresponding to the motion change information.
[0110] After obtaining the content information STI of each image block in the first video frame with frame sequence n, M image blocks with the largest STI values can be selected from each image block. These M image blocks with the largest STI values are the key image blocks. M can take the value of 3. The selection process can be specifically expressed as follows:
[0111] Patch1, Patch2...Patch m = argmax(STI i )
[0112] Among them, Patch1, Patch2...Patch m represents the m image blocks obtained by block processing, argmax(STI i) indicates identifying the image block with the maximum STI value.
[0113] In the above embodiment, the computer device extracts texture information from at least two image blocks of each first video frame respectively to obtain spatial texture information of each image block; obtains a target image block corresponding to each image block in a previous video frame of the first video frame; determines motion change information of each image block based on the difference between each image block and the corresponding target image block; and fuses the spatial texture information and the motion change information to obtain content information of each image block, that is, the content information of each image block comprehensively considers static features (such as texture and details) and dynamic changes, so that the area of the video frame that is most likely to show quality loss can be accurately determined based on the content information of each image block, thereby improving the accuracy and efficiency of video quality determination.
[0114] In one embodiment, a process in which a computer device extracts a global stitched image from video frames of a target video includes the following steps: extracting at least two second video frames from the target video at a second frame interval; performing block processing on the second video frames to obtain at least two image blocks corresponding to each second video frame; performing sub-block sampling in each image block according to sampling parameters corresponding to each image block to obtain image sub-blocks corresponding to each image block; and performing stitching processing on image sub-blocks belonging to the same second video frame to obtain global stitched images corresponding to each second video frame.
[0115] The second frame interval is the frame interval between video frames extracted from the target video, which specifies how many frames should be between the second video frames selected for analysis in the continuous video frame sequence of the target video. The second frame interval can be the same as the first frame interval or different, and the second frame interval can be an integer multiple of the first frame interval, for example, the first frame interval is 1 frame and the second frame interval is 3 frames.
[0116] The sampling parameters specify a specific way of extracting information from each image block of the second video frame, which may specifically include a sampling position and a sampling size. The sampling position specifies a specific position for extracting a sub-block from an image block, such as the center or upper left corner of the image block. The sampling size specifies the size of the extracted sub-block, which may be a fixed size or a size dynamically determined based on specific features of the image block.
[0117] The sampling parameters corresponding to the image blocks at the same position in each second video frame are the same. For example, in the first image block of the first second video frame, a 32*32 image sub-block is selected from the upper right corner of the image block. In the first image blocks of subsequent other second video frames, a 32*32 image sub-block is selected from the upper right corner of the first image block. In the second image block of the first second video frame, a 32*32 image sub-block is selected from the center of the image block. In the second image blocks of subsequent other second video frames, a 32*32 image sub-block is selected from the center of the second image block.
[0118] It can be understood that the image blocks at the same position have visual continuity between different second videos. Sampling the image blocks at the same position with the same sampling parameters can enable the global stitched images of each second video frame to carry the continuous features of the target video in time, so that more accurate underlying distortion features of the target video can be extracted based on the global stitched images of each second video frame.
[0119] Specifically, the computer device performs frame extraction on the target video through a video processing application. During the frame extraction process, a video frame is extracted from the target video every second frame interval, so as to obtain at least two second video frames. After completing the frame extraction of the target video, a second video frame sequence composed of at least two second video frames can be obtained. For any second video frame in the video frame sequence, the second video frame is block-processed according to the second block scheme to obtain at least two image blocks of the second video frame. For any image block, sub-block sampling is performed in the image block according to the corresponding sampling position and sampling size, so as to obtain the image sub-block corresponding to the image block. By performing the above processing on each image block in the second video frame, the image sub-block corresponding to each image block can be obtained, and the image sub-blocks in the second video frame are stitched to obtain the global stitched image of the second video frame. By performing the above processing on each second video frame, the global stitched images corresponding to each second video frame can be obtained.
[0120] In one embodiment, the process of extracting the global stitched image from the video frames of the target video can also be referred to as Grid Uniform Sampling (GUS), and this process can be expressed as follows:
[0121] Fragment = GUS Spatial (Frame)
[0122] Among them, Fragment represents the global stitched image, Frame represents the second video frame, and GUS SpatialIndicates uniform partition fragment sampling.
[0123] In the above embodiments, the computer device can reduce the number of subsequent processed video frames by extracting at least two second video frames from the target video at the second frame interval, while retaining sufficient video frames to evaluate the video quality, thereby improving the efficiency of video quality determination; performing block processing on the second video frames respectively to obtain at least two image blocks corresponding to each second video frame; performing sub-block sampling in each image block according to the sampling parameters corresponding to each image block to obtain image sub-blocks corresponding to each image block; performing splicing processing on the image sub-blocks belonging to the same second video frame to obtain global splicing images corresponding to each second video frame. Thus, when determining the video quality based on the global splicing images subsequently, while ensuring the accuracy of video quality determination, the computational amount is further reduced, and the efficiency of video quality determination can be further improved.
[0124] In one embodiment, the process of the computer device extracting the underlying distortion features of the target video from the global splicing images includes the following steps: extracting the underlying distortion features of the target video by passing the global splicing images through the underlying branch of the video evaluation model; the process of the computer device extracting the semantic features of the key image blocks to obtain the semantic features including the target texture information and motion information of the target video includes the following steps: extracting the semantic features of the key image blocks by passing the key image blocks through the semantic branch of the video evaluation model to obtain the semantic features including the target texture information and motion information of the target video.
[0125] Among them, the feature extraction network of the video evaluation model includes an underlying branch and a semantic branch. The underlying branch is used to process the global splicing images corresponding to the second video frame sequence. The underlying branch can be a network architecture based on Transformer (encoder), for example, it can be a Video Swin Transformer (video sliding window encoder) network. The Video Swin Transformer network is used to process time series data and can capture the motion patterns in the video and the dynamic relationships between objects.
[0126] The semantic branch is used to process the key image blocks corresponding to the first video frame sequence. Specifically, the semantic branch can be a network architecture based on ResNet (Residual Network), for example, it can be ResNet-18. ResNet-18 is an effective deep convolutional neural network that can extract complex high-level features from images and solve the degradation problem in the training of deep networks by using residual connections, enabling the network to learn deeper features.
[0127] Specifically, after the computer device obtains the key image blocks corresponding to the first video frame and the global stitched image corresponding to the second video frame, it can input the global stitched image into the underlying branch of the video evaluation model, and process each global stitched image through each network layer of the underlying branch to obtain the underlying distortion features of the target video; and input the key image blocks into the semantic branch of the video evaluation model, and process each key image through each network layer of the semantic branch to obtain the semantic features of the target video that include the target texture information and motion information.
[0128] In one embodiment, the semantic branch of the video evaluation model can be ResNet-18. The process of obtaining the semantic features of the target video that include the target texture information and motion information by extracting semantic features from the key image blocks through the semantic branch can be expressed as follows:
[0129] F S = ResNet(Input P )
[0130] Input P = Concatenate(Patch1, Patch2...Patch n*m )
[0131] Where F S represents the semantic features, Input P represents the input data of the semantic branch, that is, the key image block sequence, Patch1, Patch2...Patch n*m represents different key image blocks, and ResNet is the residual network, which is used to process the input data.
[0132] In one embodiment, the underlying branch of the video evaluation model can be Video Swin Transformer. The process of obtaining the underlying distortion features of the target video by extracting the underlying distortion features from the global stitched image through the underlying branch can be expressed as follows:
[0133] F D = Swin(Input F )
[0134] Input F = GUS(Video)
[0135] Where F D represents the underlying distortion features, Input F represents the input data of the underlying branch, that is, the global stitched image sequence, Swin represents the underlying branch, and GUS represents uniform fragmentation sampling of the target video.
[0136] In the above embodiments, the computer device extracts the underlying distortion features of the global stitched image through the underlying branch of the video evaluation model to obtain the underlying distortion features of the target video; and extracts the semantic features of the key image blocks through the semantic branch of the video evaluation model to obtain the semantic features of the target video including the target texture information and motion information, so that the relevant features beneficial to quality determination of the target video can be accurately and quickly extracted, and then the overall quality of the target video can be effectively determined, improving the accuracy and efficiency of video quality determination.
[0137] In one embodiment, the video evaluation model further includes a feature fusion network and a prediction network. The process by which the computer device determines the quality score of the target video based on the underlying distortion features and semantic features includes the following steps: performing feature fusion on the underlying distortion features and semantic features through the feature fusion network to obtain fused features; and processing the fused features through the prediction network to obtain the quality score of the target video.
[0138] Among them, the feature fusion network is used to integrate features from different levels, specifically, it can combine the underlying distortion features of the target video and the high-level semantic features.
[0139] The prediction network can be specifically constructed based on a deep learning architecture, such as a convolutional neural network (CNN), a recurrent neural network (RNN), or a fully connected neural network, and is used to make decisions based on the input features, such as inputting the quality score of the target video.
[0140] Specifically, after the computer device obtains the underlying distortion features of the target video through the underlying branch and the semantic features of the target video through the semantic branch, it can input the underlying distortion features and semantic features into the feature fusion network of the video evaluation model, and process the underlying distortion features and semantic features through each network layer of the feature fusion network to achieve feature fusion between the underlying distortion features and semantic features, obtain fused features, and input the obtained fused features into the prediction network of the video evaluation model, and process the fused features through each network layer of the prediction network to obtain the quality score of the target video.
[0141] Among them, feature fusion can be achieved in various ways, including but not limited to simple concatenation, weighted average, feature mapping, and transformation.
[0142] In the above embodiments, the computer device performs feature fusion on the underlying distortion features and semantic features through the feature fusion network to obtain fused features, and processes the fused features through the prediction network to obtain the quality score of the target video. Combining the underlying distortion features and semantic features not only improves the comprehensiveness of the evaluation, but also improves the accuracy and reliability of the evaluation, thus improving the accuracy of video quality determination.
[0143] In one embodiment, the method for determining the above video quality further includes a process of training a video evaluation model, and this process specifically includes the following steps: Extract sample key image blocks that meet the motion condition and texture condition from the video frames of the sample video; Extract sample global mosaic images from the video frames of the sample video; Through the underlying branch of the initial video evaluation model, extract underlying distortion features from each sample global mosaic image to obtain the sample underlying distortion features of the sample video; Through the semantic branch of the initial video evaluation model, extract semantic features from each sample key image block to obtain sample semantic features of the sample video that include target texture information and motion information; Determine the predicted quality score of the sample video based on the sample underlying distortion features and the sample semantic features; Optimize the initial video evaluation model based on the predicted quality score to obtain the video evaluation model.
[0144] Among them, the sample video is training data for training the video evaluation model. These videos provide the necessary data to enable the model to learn how to accurately evaluate video quality. Specifically, it can be a video obtained after compressing the original high-quality video. The video obtained after compression can specifically be videos of different quality levels, so that the model can learn to recognize different quality features.
[0145] Specifically, the computer device can perform frame extraction on the sample video to obtain the first sample video frame of the sample video. For any one of the first sample video frames, perform block processing on it to obtain each image block in the first sample video frame, and perform content analysis on each image block. Based on the content analysis results, select sample key image blocks that meet the motion condition and texture condition from each image block in the first sample video frame; it can perform frame extraction on the sample video to obtain the second sample video frame of the sample video. For any one of the second sample video frames, perform block processing on it to obtain each image block in the second sample video frame, and extract image sub-blocks from each image block, and splice the extracted image sub-blocks to obtain the sample global spliced image of the second sample video frame. By performing the above processing on each of the extracted second sample video frames, the sample global spliced image corresponding to each of the extracted second sample video frames can be obtained; input each sample global spliced image into the underlying branch of the initial video evaluation model, and process each sample global spliced image through each network layer of the underlying branch to obtain the sample underlying distortion feature of the sample video; and input the sample key image blocks into the semantic branch of the initial video evaluation model, and process each sample key image through each network layer of the semantic branch to obtain the sample semantic feature of the sample video including target texture information and motion information; input the sample underlying distortion feature and the sample semantic feature of the sample video into the prediction network of the initial video evaluation model, process the sample underlying distortion feature and the sample semantic feature through the prediction network, output the predicted quality score of the sample video, and determine the training loss value based on the predicted quality score, and optimize the initial video evaluation model based on the training loss value to obtain the trained video evaluation model.
[0146] In one embodiment, the process by which the computer device extracts sample key image blocks that meet the motion condition and texture condition from the video frames of the sample video includes the following steps: extract the first sample video frames from the sample video at the first frame interval; perform block processing on each of the first sample video frames to obtain at least two image blocks corresponding to each of the first sample video frames; select sample key image blocks that meet the motion condition and texture condition from the at least two image blocks of each of the first sample video frames.
[0147] It can be understood that the specific implementation manner in which the model selects sample key image blocks that meet the motion condition and texture condition from the sample video during the training process may be the same as or similar to the implementation manner in which the model selects key image blocks that meet the motion condition and texture condition from the target video during the aforementioned model application process.
[0148] In one embodiment, the process by which a computer device extracts a sample global stitching image from video frames of a sample video includes the following steps: extracting second sample video frames from the sample video at a second frame interval; performing a block processing on each of the second sample video frames to obtain at least two image blocks corresponding to each of the second sample video frames; performing sub-block sampling on each of the image blocks according to the sampling parameters corresponding to each of the image blocks to obtain image sub-blocks corresponding to each of the image blocks; and performing a stitching process on the image sub-blocks belonging to the same second sample video frame to obtain a sample global stitching image corresponding to each of the second sample video frames.
[0149] It can be understood that the specific implementation manner in which the model selects a sample global stitching image from a sample video during the training process may be the same as or similar to the implementation manner in which the model selects a global stitching image from a target video during the aforementioned model application process.
[0150] In one embodiment, the initial video evaluation model further includes a feature fusion network and a prediction network. The process by which a computer device determines the predicted quality score of a sample video based on the sample low-level distortion features and the sample semantic features further includes the following steps: performing feature fusion on the sample low-level distortion features and the sample semantic features through the feature fusion network to obtain sample fusion features; and processing the sample fusion features through the prediction network to obtain the predicted quality score of the sample video.
[0151] It can be understood that the specific implementation manner in which the model determines the quality score of a target video based on the low-level distortion features and the semantic features during the training process may be the same as or similar to the implementation manner in which the model determines the predicted quality score of a sample video based on the sample low-level distortion features and the sample semantic features during the aforementioned model application process.
[0152] In the above embodiment, the computer device extracts sample key image blocks that meet the motion condition and the texture condition from the video frames of the sample video; extracts a sample global stitching image from the video frames of the sample video; extracts sample low-level distortion features of the sample video by performing low-level distortion feature extraction on each of the sample global stitching images through the low-level branch of the initial video evaluation model; extracts sample semantic features of the sample video, which include target texture information and motion information, by performing semantic feature extraction on each of the sample key image blocks through the semantic branch of the initial video evaluation model; determines the predicted quality score of the sample video based on the sample low-level distortion features and the sample semantic features; and optimizes the initial video evaluation model based on the predicted quality score, so that a video evaluation model capable of accurately evaluating the video quality can be obtained. Furthermore, when subsequently determining the quality of a target video based on the video evaluation model, the accuracy of video quality determination can be improved.
[0153] In one embodiment, the process of the computer device optimizing the initial video evaluation model based on the predicted quality score to obtain the video evaluation model includes the following steps: determining the training loss value based on the predicted quality score; adjusting the parameters of the initial video evaluation model based on the training loss value until the training stop condition is met, thereby obtaining the video evaluation model.
[0154] Among them, the training loss value is used to measure the performance of the model on the training data to guide the optimization process of the model parameters. The training stop condition is a condition used to determine when to end the model training, which can prevent overfitting and ensure that the model reaches sufficient performance. Specifically, it can be a loss threshold condition, a performance threshold condition, an iteration number condition, or a custom condition. The loss threshold condition can specifically be setting a specific loss function threshold. When the loss value of the model drops below this threshold, it is considered that the model has been sufficiently optimized and the training can be stopped; the performance threshold condition can specifically be based on the performance of the model on the validation set, such as accuracy, recall, etc. If the model performance no longer improves significantly or reaches a predetermined performance level, the training is stopped; the iteration number condition can specifically be setting the maximum number of iterations for training. Regardless of the model performance, the training stops after reaching this iteration number; the custom condition can specifically be customizing a set of rules for stopping training according to specific tasks or requirements. This set of rules can be a combination of the above conditions or conditions based on other specific metrics.
[0155] Specifically, after the computer device obtains the predicted quality score of the sample video, it inputs the predicted quality score of the sample video into a preset loss function, calculates the predicted quality score through this loss function to obtain the training loss value corresponding to the target video, and uses the backpropagation algorithm to adjust the parameters of the initial video evaluation model based on the training loss value. Through multiple iterations, the training loss value is gradually reduced until the training stop condition is met, thereby obtaining the trained video evaluation model.
[0156] In the above embodiment, the computer device determines the training loss value based on the predicted quality score. The training loss value reflects the gap between the model prediction and the actual video quality. By adjusting the parameters of the initial video evaluation model based on the training loss value, the direction of model optimization can be guided by accurately quantifying the prediction error, ensuring that the model learns to more accurately evaluate the video quality until the training stop condition is met, and obtaining the video evaluation model. The training process ensures the sufficiency and stability of the model training, thereby obtaining a video evaluation model that can accurately evaluate the video quality and has stable performance. Subsequently, when determining the quality of the target video based on the video evaluation model, the accuracy of the video quality determination can be improved.
[0157] In one embodiment, the process by which a computer device determines a training loss value based on a predicted quality score includes the following steps: obtaining the label quality score of a sample video, the reference label quality score and the reference predicted quality score of a reference sample video; determining a ground-truth loss value based on the label quality score and the predicted quality score; determining a ranking loss value according to the reference label quality score, the reference predicted quality score, the label quality score and the predicted quality score; and determining the training loss value based on the ground-truth loss value and the ranking loss value.
[0158] Among them, the ground-truth loss value is used to characterize the difference between the prediction result and the true value. In the embodiments of the present application, the prediction result is the predicted quality score of the sample video, and the true value is the label quality score of the sample video. The label quality score of the sample video may specifically be the result obtained by manually scoring the sample video.
[0159] The reference sample video is another sample video in the training dataset other than the currently evaluated sample video and for which the predicted quality score has been output by the initial video model.
[0160] The ranking loss value is used to measure the logical consistency or monotonicity of the model's prediction results. Suppose there are two video samples A and B, with their true quality scores being 80 and 70 respectively, and the quality scores predicted by the model being 75 and 85 respectively. The order predicted by the model does not match the true order (the model believes that B has a higher quality), so this ranking loss value will be greater than 0, indicating that the model needs to be adjusted to improve the accuracy of its prediction.
[0161] Specifically, after obtaining the predicted quality score of the sample video, the computer device may further obtain the label quality score of the sample video, determine the difference between the predicted quality score and the label quality score, and determine the ground-truth loss value based on this difference. In addition, the computer device may also obtain the label quality score and the predicted quality score of the reference sample video, compare the label quality score of the reference sample video with the label quality score of the sample video to obtain a first comparison result, compare the predicted quality score of the reference sample video with the predicted quality score of the sample video to obtain a second comparison result, determine the ranking loss value based on the first comparison result and the second comparison result, obtain the ground-truth loss weight corresponding to the ground-truth loss value and the ranking loss weight corresponding to the ranking loss value, and determine the training loss value according to the ground-truth loss weight, the ground-truth loss value, the ranking loss weight and the ranking loss value.
[0162] In one embodiment, the following relationship holds among the ground-truth loss value, the predicted quality score of the sample video, and the label quality score of the sample video:
[0163]
[0164] Among them, L MAE represents the ground-truth loss value, N represents that there are a total of N sample videos, y iRepresents the predicted quality score of the i-th sample video, Represents the labeled quality score of the i-th sample video.
[0165] In one embodiment, the following relationships are satisfied among the ranking loss value, the predicted quality score of the sample video, the labeled quality score of the sample video, the predicted quality score of the reference sample video, and the labeled quality score of the reference sample video:
[0166]
[0167]
[0168]
[0169] Wherein, L rank Represents the ranking loss value, Represents the monotonic difference value between the i-th sample video and the j-th sample video, Represents the labeled quality score of the i-th sample video, Represents the labeled quality score of the j-th sample video, y i Represents the predicted quality score of the i-th sample video, y j Represents the predicted quality score of the j-th sample video.
[0170] In one embodiment, the following relationships are satisfied among the training loss value, the ground-truth loss value, and the ranking loss value:
[0171] L = L MAE + ω * L rank
[0172] Wherein, L represents the training loss value, L MAE Represents the ground-truth loss value, L_rank represents the ranking loss value, and ω represents the adjustment coefficient.
[0173] In the above embodiments, the computer device obtains the labeled quality score of the sample video, the reference labeled quality score and the reference predicted quality score of the reference sample video; determines the ground-truth loss value based on the labeled quality score and the predicted quality score, and the ground-truth loss value reflects the difference between the model prediction and the real data. Minimizing this loss can improve the accuracy of the model's predicted quality score; determines the ranking loss value according to the reference labeled quality score, the reference predicted quality score, the labeled quality score, and the predicted quality score to evaluate the ranking accuracy of the model among different samples; determines the training loss value based on the ground-truth loss value and the ranking loss value. By integrating these two losses, it is possible to simultaneously optimize the accuracy of the model in predicting the quality of a single video and the accuracy of ranking the quality among multiple videos, ensuring that the model is not only accurate in predicting specific scores but also consistent when comparing different videos.
[0174] In one embodiment, before obtaining the label quality score of the sample video, the reference label quality score and the reference prediction quality score of the reference sample video, the computer device may further obtain the manual annotation quality score of the sample video and the corresponding scoring correlation information; perform data screening on the manual annotation quality score based on the scoring correlation information to obtain the screened manual annotation quality score; determine the label quality score of the sample video based on the screened manual annotation quality score.
[0175] Among them, the scoring correlation information is other information involved in the manual quality evaluation of the sample video, which affects the accuracy of the corresponding manual annotation quality score. Specifically, it may include information such as viewing duration and the scoring tendency of the scorers. The viewing duration can reflect the attention degree of the scorers to the video content. Completing the viewing of the video usually means more accurate and comprehensive scoring. If a scorer only watches a small part of the video, he may miss the key elements affecting the overall quality evaluation of the video, resulting in an inaccurate manual annotation quality score. The scoring tendency of the scorers refers to the consistent pattern shown by the scorers during the quality evaluation. The scoring tendency may cause the scoring result to deviate from the objective standard. For example, a scorer with a malicious scoring tendency may, due to some prejudice or purpose, deliberately give an untrue manual annotation quality score.
[0176] Specifically, for any sample video, a preset number of scorers can be used to watch the video. After each scorer watches the sample video, they can give the corresponding manual annotation quality score. The computer device can obtain the scoring correlation information of the scorers regarding the sample video, determine the reliability of the manual annotation quality score given by each scorer according to the scoring correlation information, screen out the manual annotation quality scores whose reliability meets the reliability condition from the manual annotation quality scores, and determine the average value of the screened manual annotation quality scores, and determine the obtained average value as the label quality score of the sample video.
[0177] For example, for sample video A, 30 scorers are used to watch sample video A. After each scorer watches sample video A, according to the pre-established scoring criteria, they give their manual annotation quality scores. The computer device can collect the viewing duration of each scorer watching the sample video, determine the viewing completion rate based on the viewing duration and the total duration of sample video A, and determine the manual annotation quality scores with a viewing completion rate reaching 80% as the manual annotation quality scores that meet the reliability condition, and determine the average value of the screened manual annotation quality scores, and determine the obtained average value as the label quality score of the sample video.
[0178] For another example, for sample video A, 30 raters are asked to watch the sample video A. After each rater watches the sample video A, they give their manual annotation quality scores according to a pre-established scoring standard. The computer device can collect the manual annotation quality scores of each rater for other sample videos. For any one rater, based on their manual annotation quality scores for other sample videos, it is determined whether the rater has a tendency to give malicious scores, so as to obtain the scoring tendency of each rater. Then, among the manual annotation quality scores corresponding to sample video A, the manual annotation quality scores without the tendency to give malicious scores are screened out, and the average value of the screened manual annotation quality scores is determined. The obtained average value is determined as the label quality score of the sample video.
[0179] In one embodiment, the pre-established scoring standard can be the ITU-R BT.500 standard. When rating, raters can use a seven / five-level comparison scale to measure the relative quality of the target video's picture quality. For each rater, the total viewing time of the target video is controlled within 30 minutes to avoid fatigue. As Figure 4 shown is the distribution of the label quality scores of sample videos in the training dataset in one embodiment, where the horizontal axis represents the label quality score (Mean Opinion Score, MOS), and the vertical axis represents the number of sample videos. It can be seen from the figure that the distribution of sample videos with different label quality scores in the training dataset is relatively uniform.
[0180] In the above embodiment, the computer device obtains the manual annotation quality scores of the sample video and the corresponding scoring correlation information; based on the scoring correlation information, data screening is performed on the manual annotation quality scores to obtain the screened manual annotation quality scores, so as to exclude those unreliable scores. Based on the screened manual annotation quality scores, the label quality score of the sample video is determined. Subsequently, when training the model based on the label quality score of the sample video, the model can learn video quality assessment on a more accurate basis, so as to obtain a video evaluation model that can accurately evaluate video quality and has stable performance.
[0181] In one embodiment, the method for determining the above video quality further includes the process of constructing sample videos for training, and this process includes the following steps: obtaining the original video and at least two constant speed factor intervals corresponding thereto; respectively selecting a constant speed factor in each constant speed factor interval to obtain at least two constant speed factors; determining the encoder and resolution corresponding to the target constant speed factor among the at least two constant speed factors; and compressing and encoding the original video by the encoder according to the resolution and the target constant speed factor to obtain the sample video.
[0182] Among them, the original video can specifically be an uncompressed high-quality video.
[0183] The constant rate factor refers to the Constant Rate Factor (CRF). The constant rate factor is an encoding setting that specifies the trade-off between quality and compression during video encoding. The lower the value of the CRF, the higher the quality of the output video and the larger the file size; the higher the value of the CRF, the lower the video quality, but the file size decreases.
[0184] Specifically, after the computer device obtains the original video, it can sequentially select the constant rate factor range required for the current round from H constant rate factor ranges, randomly select a value from this constant rate factor range as the target constant rate factor, randomly select an encoder from the encoder library, randomly select a resolution from the resolution library, and determine the selected encoder and resolution as the encoder and resolution corresponding to the target constant rate factor. Then, compress and encode the original video according to the selected resolution and target constant rate factor through the selected encoder to obtain a sample video; then sequentially select the constant rate factor range required for the next round from multiple constant rate factor ranges, randomly select a value from this constant rate factor range as the target constant rate factor, randomly select an encoder from the encoder library, randomly select a resolution from the resolution library, and determine the selected encoder and resolution as the encoder and resolution corresponding to the target constant rate factor. Compress and encode the original video according to the selected resolution and target constant rate factor through the selected encoder to obtain the next sample video; repeat the above process until a sample video corresponding to each constant rate factor range is obtained, that is, H sample videos with different compression situations can be obtained for an original video. Here, H is a positive integer greater than or equal to 2.
[0185] Taking the following scenario as an example, the process of generating the sample video is described. Specifically, the original video can be obtained from the data source. The data source can be Professionally Generated Content (PGC), User Generated Content (UGC), etc. Obtain a sufficient number of original videos from the data source. The sufficient number of original videos can cover a sufficient number of video scene distributions. There are four encoders provided in the encoder library, specifically x264, x265, AV1, and VP9. There are five resolutions provided in the resolution library, specifically 270p, 480p, 540p, 720p, and 1080p. In addition, eight CRF ranges are provided, specifically [0,10), [10,16), [16,22), [22,28), [22,34), [34,40), [40,46), and [46,52). For any original video, 8 sample videos can be encoded. Refer to Figure 5The flowchart shown, and the encoding process specifically includes the following steps:
[0186] Sequentially select the k-th constant speed factor (CRF) interval from 7 constant speed factor intervals, randomly select a value from this constant speed factor interval as the target constant speed factor, randomly select an encoder from the encoder library, randomly select a resolution from the resolution library, and determine the encoder and resolution corresponding to the target constant speed factor for the selected encoder and resolution. Compress and encode the original video according to the selected resolution and target constant speed factor through the selected encoder to obtain a sample video; let k = k + 1, and return to execute the step of sequentially selecting the k-th constant speed factor interval from 7 constant speed factor intervals until a sample video corresponding to each constant speed factor interval is obtained, that is, 8 sample videos with different compression situations can be obtained for an original video. Among them, k is an integer between 1 and 6.
[0187] In the above embodiment, the computer device obtains the original video and at least two corresponding constant speed factor intervals; respectively select constant speed factors in each constant speed factor interval to obtain at least two constant speed factors; determine the encoder and resolution corresponding to the target constant speed factor among at least two constant speed factors; compress and encode the original video according to the resolution and target constant speed factor through the encoder, so that sample videos of different quality levels can be obtained, ensuring that the training data set is both diverse and comprehensive, covering various video types from high quality to low quality. Thus, when training the model based on the sample videos subsequently, the model can learn and adapt to videos of various quality levels, improving the accuracy and stability of the video evaluation model.
[0188] In one example, as Figure 6 shown, a method for determining video quality is also provided. Taking the example that this method is applied to a computer device, it includes the following steps:
[0189] S602, Extract at least two first video frames from the target video according to the first frame interval.
[0190] S604, Perform block processing on each first video frame to obtain at least two image blocks corresponding to each first video frame.
[0191] S606, Extract texture information from at least two image blocks of each first video frame to obtain the spatial texture information of each image block.
[0192] S608, In the prior video frames of the first video frame, obtain the target image blocks corresponding to each image block.
[0193] S610, Determine the motion change information of each image block based on the difference between each image block and the corresponding target image block.
[0194] S612, fuse the spatial texture information and the motion change information to obtain the content information of each image block.
[0195] S614, based on the content information of each image block, select key image blocks that meet the motion condition and the texture condition from at least two image blocks of each first video frame.
[0196] S616, extract at least two second video frames from the target video at a second frame interval.
[0197] S618, perform block processing on each second video frame to obtain at least two image blocks corresponding to each second video frame.
[0198] S620, perform sub-block sampling in each image block according to the sampling parameters corresponding to each image block to obtain image sub-blocks corresponding to each image block.
[0199] Among them, the sampling parameters corresponding to the image blocks in the same position in each second video frame are the same.
[0200] S622, perform splicing processing on the image sub-blocks belonging to the same second video frame to obtain a global spliced image corresponding to each second video frame.
[0201] S624, through the underlying branch of the video evaluation model, extract the underlying distortion features of each global spliced image to obtain the underlying distortion features of the target video.
[0202] S626, through the semantic branch of the video evaluation model, extract semantic features from each key image block to obtain semantic features of the target video including target texture information and motion information.
[0203] S628, perform feature fusion on the underlying distortion features and the semantic features through a feature fusion network to obtain fusion features.
[0204] S630, process the fusion features through a prediction network to obtain the quality score of the target video.
[0205] This application also provides an application scenario, which implements the above method for determining video quality through a video evaluation model. Refer to Figure 7 the structural schematic diagram of the video evaluation model shown, and the method for determining the above video quality specifically includes the following steps:
[0206] Step 1, perform frame extraction on the target video to obtain the first video frame and the second video frame.
[0207] Step 2.1, perform block processing on the first video frame according to the first block scheme to obtain at least two image blocks corresponding to each first video frame; select key image blocks that meet the motion condition and texture condition from the at least two image blocks of each first video frame.
[0208] Step 2.2, perform block processing on the second video frame according to the second block scheme to obtain at least two image blocks corresponding to each second video frame; perform sub-block sampling on each image block according to the sampling parameters corresponding to each image block to obtain image sub-blocks corresponding to each image block; wherein, the sampling parameters corresponding to the image blocks in the same position in each second video frame are the same; perform splicing processing on the image sub-blocks belonging to the same second video frame to obtain the global spliced image corresponding to each second video frame.
[0209] Step 3.1, input the key image blocks into the ResNet-18 branch of the video evaluation model to extract semantic features containing target texture information and motion information.
[0210] Step 3.2, input the global spliced image into the Video Swin Transformer branch of the video evaluation model to extract the underlying distortion features of the target video.
[0211] Step 4, input the semantic features and the underlying distortion features into the feature fusion network of the video evaluation model, and perform feature fusion on the underlying distortion features and the semantic features through the feature fusion network to obtain fused features.
[0212] Step 5, input the fused features into the prediction network of the video evaluation model, and process the fused features through the prediction network to obtain the quality score of the target video.
[0213] The video evaluation model is pre-trained on a large scale based on sample videos. Referring to Figure 8 the flowchart shown, the training process includes the following steps:
[0214] Step 1, construct sample videos.
[0215] Specifically, obtain the original video and at least two constant speed factor intervals corresponding thereto; select constant speed factors in each constant speed factor interval respectively to obtain at least two constant speed factors; determine the encoder and resolution corresponding to the target constant speed factor among the at least two constant speed factors; compress and encode the original video by the encoder according to the resolution and the target constant speed factor to obtain the sample video.
[0216] Step 2, score the sample videos.
[0217] By manually viewing the sample videos and scoring them, the manual annotation quality scores of the sample videos and the corresponding scoring correlation information are obtained; based on the scoring correlation information, data screening is performed on the manual annotation quality scores to obtain the screened manual annotation quality scores; based on the screened manual annotation quality scores, the label quality scores of the sample videos are determined.
[0218] Step 3, model structure design.
[0219] The structure of the initial video evaluation model is as Figure 8 shown, including a bottom layer branch, a semantic branch, a feature fusion network, and a prediction network.
[0220] Step 4, model training.
[0221] Sample key image blocks that meet the motion condition and texture condition are extracted from the video frames of the sample videos; sample global stitching images are extracted from the video frames of the sample videos; through the bottom layer branch of the initial video evaluation model, bottom layer distortion features of each sample global stitching image are extracted to obtain the sample bottom layer distortion features of the sample videos; through the semantic branch of the initial video evaluation model, semantic features of each sample key image block are extracted to obtain the sample semantic features of the sample videos that include target texture information and motion information; based on the sample bottom layer distortion features and the sample semantic features, the prediction quality scores of the sample videos are determined; the label quality scores of the sample videos, the reference label quality scores of the reference sample videos, and the reference prediction quality scores are obtained; based on the label quality scores and the prediction quality scores, the true value loss values are determined; based on the reference label quality scores, the reference prediction quality scores, the label quality scores, and the prediction quality scores, the ranking loss values are determined; based on the true value loss values and the ranking loss values, the training loss values are determined; based on the training loss values, the parameters of the initial video evaluation model are adjusted until the training stop condition is met, and a video evaluation model is obtained.
[0222] In addition, the training method provided in this application is respectively applied to train video evaluation models on two different data sets, and at the same time, models are also trained using other existing methods on the two different data sets. The corresponding results are shown in the following table:
[0223]
[0224] Among them, SRCC (Spearman Rank Correlation Coefficient) is the Spearman rank correlation coefficient, which is used to measure the rank correlation between two variables, that is, their monotonic association; PLCC (Pearson Linear Correlation Coefficient) is the Pearson linear correlation coefficient, which is a statistic for measuring the strength of the linear relationship between two variables; MSE (Mean Squared Error) is the mean squared error, which is a common statistical indicator for measuring the difference between the predicted value and the actual observed value. It can be seen from the above table that the performance of the video evaluation model trained by the training method provided in this application is better than that of the models trained by other existing methods.
[0225] In addition, referring to Figure 9 and Figure 10 respectively shown prediction quality score-compression ratio distribution diagrams, Figure 9 is the distribution of the prediction quality scores of the sample videos with different compression ratios (QP) corresponding to the original video A trained by the video evaluation model trained by the training method provided in this application, Figure 10 is the distribution of the prediction quality scores of the sample videos with different compression ratios (QP) corresponding to the original video B trained by the video evaluation model trained by the training method provided in this application. The horizontal axis represents the compression ratio and the vertical axis represents the prediction quality score. Combining Figure 9 and Figure 10 it can be seen that the prediction quality score decreases as the compression ratio increases, which is in line with the change of the reduced picture quality.
[0226] Referring to Figure 11 and Figure 12 respectively shown true value-predicted value comparison diagrams, Figure 11 is the comparison of the prediction quality score and the label quality score of the sample videos with different compression situations (CRF) corresponding to the original video C trained by the video evaluation model trained by the training method provided in this application, Figure 12 is the comparison of the prediction quality score and the label quality score of the sample videos with different compression situations (CRF) corresponding to the original video D trained by the video evaluation model trained by the training method provided in this application. Combining Figure 11 and Figure 12 it can be seen that the change curve of the prediction quality score of the sample video fits closely with the change curve of the label quality score, indicating that the accuracy of the prediction quality score is relatively high and is in line with the human eye perception result.
[0227] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0228] Based on the same inventive concept, an embodiment of the present application also provides a video quality determination device for implementing the video quality determination method described above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the video quality determination device provided below can refer to the limitations on the video quality determination method in the above text, and will not be repeated here.
[0229] In one embodiment, as Figure 13 shown, a video quality determination device is provided, including: a key image block extraction module 1302, a global mosaic image extraction module 1304, a low-level distortion feature extraction module 1306, a semantic feature extraction module 1308, and a quality score determination module 1310, where:
[0230] The key image block extraction module 1302 is configured to extract key image blocks that meet the motion condition and texture condition from the video frames of the target video.
[0231] The global mosaic image extraction module 1304 is configured to extract a global mosaic image from the video frames of the target video.
[0232] The low-level distortion feature extraction module 1306 is configured to extract low-level distortion features from the global mosaic image to obtain the low-level distortion features of the target video.
[0233] The semantic feature extraction module 1308 is configured to extract semantic features from the key image blocks to obtain semantic features of the target video including target texture information and motion information.
[0234] The quality score determination module 1310 is configured to determine the quality score of the target video based on the low-level distortion features and semantic features.
[0235] In the above embodiments, by extracting key image blocks that meet the motion condition and texture condition from the video frames of the target video; the key image blocks are usually the places where the video quality loss is most obvious. By extracting semantic features from the key image blocks, semantic features of the target video including target texture information and motion information are obtained. When determining the video quality based on the semantic features subsequently, the accuracy and efficiency of video quality determination can be improved; a global stitching image is extracted from the video frames of the target video, and the global stitching image carries the global loss information of the target video. By extracting the underlying distortion features from the global stitching image, the underlying distortion features of the target video are obtained. When determining the video quality based on the underlying distortion features subsequently, the accuracy of video quality determination can be improved, and a large amount of calculation required for directly analyzing the original video frames is avoided during the extraction of the underlying distortion features, and the efficiency of video quality determination is further improved; the quality score of the target video is determined based on the underlying distortion features and semantic features. When determining the video quality, both the underlying distortion features and semantic features are comprehensively considered, that is, the feature dimensions for video quality determination are more comprehensive and accurate, thereby further improving the accuracy of video quality determination.
[0236] In one embodiment, the key image block extraction module 1302 is further configured to: extract at least two first video frames from the target video at a first frame interval; perform block processing on each of the first video frames to obtain at least two image blocks corresponding to each of the first video frames; select key image blocks that meet the motion condition and texture condition from the at least two image blocks of each of the first video frames.
[0237] In one embodiment, the key image block extraction module 1302 is further configured to: extract content information of each of the at least two image blocks of each of the first video frames to obtain the content information of each image block; based on the content information of each image block, select key image blocks that meet the motion condition and texture condition from the at least two image blocks of each of the first video frames.
[0238] In one embodiment, the key image block extraction module 1302 is further configured to: extract texture information of each of the at least two image blocks of each of the first video frames to obtain the spatial texture information of each image block; obtain the target image block corresponding to each image block in the prior video frame of the first video frame; determine the motion change information of each image block based on the difference between each image block and the corresponding target image block; fuse the spatial texture information and the motion change information to obtain the content information of each image block.
[0239] In one embodiment, the global stitching image extraction module 1304 is further configured to: extract at least two second video frames from the target video at a second frame interval; perform block processing on the second video frames respectively to obtain at least two image blocks corresponding to each second video frame; perform sub-block sampling on each image block according to the sampling parameters corresponding to each image block to obtain image sub-blocks corresponding to each image block; wherein, the sampling parameters corresponding to the image blocks at the same position in each second video frame are the same; perform stitching processing on the image sub-blocks belonging to the same second video frame to obtain the global stitching image corresponding to each second video frame.
[0240] In one embodiment, the underlying distortion feature extraction module 1306 is further configured to: extract underlying distortion features of the target video by the underlying branch of the video evaluation model from the global stitching image; the semantic feature extraction module 1308 is further configured to: extract semantic features of the key image blocks by the semantic branch of the video evaluation model to obtain semantic features of the target video including target texture information and motion information.
[0241] In one embodiment, the video evaluation model further includes a feature fusion network and a prediction network; the underlying distortion feature extraction module 1306 is further configured to: perform feature fusion on the underlying distortion features and the semantic features through the feature fusion network to obtain fused features; process the fused features through the prediction network to obtain the quality score of the target video.
[0242] In one embodiment, as Figure 14 shown, the apparatus further includes a model training module 1312, configured to: extract sample key image blocks that meet the motion condition and the texture condition from the video frames of the sample video; extract sample global stitching images from the video frames of the sample video; extract sample underlying distortion features of the sample video by the underlying branch of the initial video evaluation model from each sample global stitching image; extract sample semantic features of the sample video including target texture information and motion information by the semantic branch of the initial video evaluation model from each sample key image block; determine the predicted quality score of the sample video based on the sample underlying distortion features and the sample semantic features; optimize the initial video evaluation model based on the predicted quality score to obtain the video evaluation model.
[0243] In one embodiment, the model training module 1312 is further configured to: determine a training loss value based on the predicted quality score; adjust the parameters of the initial video evaluation model based on the training loss value until the training stop condition is met to obtain the video evaluation model.
[0244] In one embodiment, the model training module 1312 is further configured to: obtain the label quality score of the sample video, the reference label quality score and the reference prediction quality score of the reference sample video; determine the ground-truth loss value based on the label quality score and the prediction quality score; determine the ranking loss value according to the reference label quality score, the reference prediction quality score, the label quality score and the prediction quality score; and determine the training loss value based on the ground-truth loss value and the ranking loss value.
[0245] In one embodiment, the model training module 1312 is further configured to: obtain the manual annotation quality score of the sample video and the corresponding scoring correlation information; perform data screening on the manual annotation quality score based on the scoring correlation information to obtain the screened manual annotation quality score; and determine the label quality score of the sample video based on the screened manual annotation quality score.
[0246] In one embodiment, the model training module 1312 is further configured to: obtain the original video and at least two constant speed factor intervals corresponding thereto; select a constant speed factor in each constant speed factor interval to obtain at least two constant speed factors; determine the encoder and the resolution corresponding to the target constant speed factor among the at least two constant speed factors; and perform compression encoding on the original video by the encoder according to the resolution and the target constant speed factor to obtain the sample video.
[0247] Each module in the above video quality determination device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0248] In one embodiment, the computer device can be a terminal or a server. In this embodiment, taking the computer device as a terminal as an example for illustration, its internal structure diagram can be as Figure 15As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements a method for determining video quality. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0249] Those skilled in the art can understand that Figure 15 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0250] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0251] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0252] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0253] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0254] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.
[0255] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0256] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for determining video quality, characterized in that, The method includes: extracting key image blocks that meet the motion condition and texture condition from the video frames of the target video; extracting a global stitching image from the video frames of the target video; performing low-level distortion feature extraction on the global stitching image to obtain the low-level distortion features of the target video; performing semantic feature extraction on the key image blocks to obtain semantic features of the target video that include target texture information and motion information; determining the quality score of the target video based on the low-level distortion features and the semantic features.
2. The method according to claim 1, wherein The extracting key image blocks that meet the motion condition and texture condition from the video frames of the target video includes: extracting at least two first video frames from the target video at a first frame interval; performing block processing on each of the first video frames to obtain at least two image blocks corresponding to each of the first video frames; selecting key image blocks that meet the motion condition and texture condition from the at least two image blocks of each of the first video frames.
3. The method according to claim 2, wherein The selecting key image blocks that meet the motion condition and texture condition from the at least two image blocks of each of the first video frames includes: performing content information extraction on the at least two image blocks of each of the first video frames to obtain the content information of each of the image blocks; selecting key image blocks that meet the motion condition and texture condition from the at least two image blocks of each of the first video frames based on the content information of each of the image blocks.
4. The method according to claim 3, wherein The performing content information extraction on the at least two image blocks of each of the first video frames to obtain the content information of each of the image blocks includes: performing texture information extraction on the at least two image blocks of each of the first video frames to obtain the spatial texture information of each of the image blocks; acquiring the target image blocks corresponding to each of the image blocks in the prior video frame of the first video frame; determining the motion change information of each of the image blocks based on the difference between each of the image blocks and the corresponding target image blocks; fusing the spatial texture information and the motion change information to obtain the content information of each of the image blocks.
5. The method according to claim 1, characterized in that, The extracting a global stitching image from the video frames of the target video includes: extracting at least two second video frames from the target video at a second frame interval; performing block processing on each of the second video frames to obtain at least two image blocks corresponding to each of the second video frames; performing sub-block sampling on each of the image blocks according to the sampling parameters corresponding to each of the image blocks to obtain image sub-blocks corresponding to each of the image blocks; wherein, the sampling parameters corresponding to the image blocks in the same position in each of the second video frames are the same; performing stitching processing on the image sub-blocks belonging to the same second video frame to obtain the global stitching image corresponding to each of the second video frames.
6. The method according to claim 1, wherein The performing low-level distortion feature extraction on the global stitching image to obtain the low-level distortion features of the target video includes: performing low-level distortion feature extraction on the global stitching image through the low-level branch of the video evaluation model to obtain the low-level distortion features of the target video; Performing semantic feature extraction on the key image blocks to obtain semantic features of the target video including target texture information and motion information, including: Performing semantic feature extraction on the key image blocks through the semantic branch of the video evaluation model to obtain semantic features of the target video including target texture information and motion information.
7. The method according to claim 6, characterized in that, The video evaluation model further includes a feature fusion network and a prediction network; determining the quality score of the target video based on the underlying distortion features and the semantic features, including: Performing feature fusion on the underlying distortion features and the semantic features through the feature fusion network to obtain fused features; Processing the fused features through the prediction network to obtain the quality score of the target video.
8. The method according to claim 6, wherein The method further includes: Extracting sample key image blocks that meet the motion condition and texture condition from the video frames of the sample video; extracting a sample global mosaic image from the video frames of the sample video; Performing underlying distortion feature extraction on the sample global mosaic image through the underlying branch of the initial video evaluation model to obtain sample underlying distortion features of the sample video; Performing semantic feature extraction on the sample key image blocks through the semantic branch of the initial video evaluation model to obtain sample semantic features of the sample video including target texture information and motion information; Determining a predicted quality score of the sample video based on the sample underlying distortion features and the sample semantic features; optimizing the initial video evaluation model based on the predicted quality score to obtain the video evaluation model.
9. The method according to claim 8, wherein Optimizing the initial video evaluation model based on the predicted quality score to obtain the video evaluation model, including: Determining a training loss value based on the predicted quality score; Adjusting the parameters of the initial video evaluation model based on the training loss value until a training stop condition is met to obtain the video evaluation model.
10. The method according to claim 9, characterized in that, Determining the training loss value based on the predicted quality score, including: Obtaining the label quality score of the sample video, the reference label quality score of the reference sample video, and the reference predicted quality score; Determining a ground truth loss value based on the label quality score and the predicted quality score; Determining a ranking loss value according to the reference label quality score, the reference predicted quality score, the label quality score, and the predicted quality score; Determining the training loss value based on the ground truth loss value and the ranking loss value.
11. The method according to claim 10, wherein Before obtaining the label quality score of the sample video, the reference label quality score of the reference sample video, and the reference predicted quality score, the method further includes: Obtaining the manually annotated quality score of the sample video and the corresponding scoring correlation information; Performing data screening on the manually annotated quality score based on the scoring correlation information to obtain the screened manually annotated quality score; Determining the label quality score of the sample video based on the screened manually annotated quality score.
12. The method according to claim 8, wherein The method further includes: Obtaining an original video and at least two constant speed factor intervals corresponding thereto; Selecting a constant speed factor in each of the constant speed factor intervals to obtain at least two constant speed factors; Determine the encoder and resolution corresponding to the target constant speed factor among the at least two constant speed factors; Compress and encode the original video by the encoder according to the resolution and the target constant speed factor to obtain the sample video.
13. A device for determining video quality, characterized in that, The device includes: A key image block extraction module, configured to extract key image blocks that meet the motion condition and texture condition from video frames of a target video; A global stitching image extraction module, configured to extract a global stitching image from video frames of the target video; A low-level distortion feature extraction module, configured to extract low-level distortion features of the global stitching image to obtain low-level distortion features of the target video; A semantic feature extraction module, configured to extract semantic features of the key image blocks to obtain semantic features of the target video including target texture information and motion information; A quality score determination module, configured to determine a quality score of the target video based on the low-level distortion features and the semantic features.
14. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
16. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.