Video quality evaluation method and device
By performing regional segmentation and weighted scores on UGC videos, the problem that traditional methods are difficult to accurately evaluate UGC video quality is solved, and video quality evaluation is more in line with human eye perception.
Patent Information
- Application Number
- CN202510069517.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional video quality evaluation methods are difficult to accurately evaluate the true quality of UGC videos, especially in the case of diverse video content and complex shooting environments.
By segmenting the target video in the area, the quality prediction scores for each area are obtained, and the weight of the area is determined based on these scores, and weighted scores are performed to improve the accuracy of video quality evaluation.
This method can dynamically adjust its contribution to the overall quality assessment results according to the importance of the region, ensuring that high-quality areas have a greater impact on the final score, and the impact of low-quality areas is effectively suppressed, and the results are more in line with the human eye's visual perception of video quality.
Smart Images

Figure CN119991600A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of image processing technology, and in particular, to a video quality assessment method and device. Background Art
[0002] With the development of the Internet and social platforms, the number of user-generated content (UGC) videos has increased dramatically. These videos are widely used in various scenarios such as social media, short video platforms, online education, news reports, etc., covering a variety of forms from professional production to personal creation. In order to review and recommend UGC videos and improve the user viewing experience, it is usually necessary to evaluate the video quality of UGC videos.
[0003] UGC videos are usually highly diverse and unstructured. They may contain various backgrounds, scene switching, and different shooting angles. At the same time, due to the wide range of creators, UGC videos usually face challenges such as instability and complex shooting environments. These factors lead to large differences in the quality of UGC videos. When evaluating the quality of such videos, traditional methods are difficult to accurately evaluate the video quality based on the true quality of the video and the actual visual perception of the human eye. There is an urgent need to introduce a video quality evaluation method to provide more accurate and reliable evaluation results for the growing and diverse types of video content. Summary of the invention
[0004] The embodiments of this specification provide a video quality assessment solution, which uses the region segmentation results of the target video and the corresponding quality prediction scores for weighted scoring, thereby effectively improving the accuracy of video quality assessment.
[0005] In a first aspect, an embodiment of the present specification provides a video quality assessment method, including: obtaining a target video to be quality assessed, the target video including multiple continuous video frames; performing region segmentation on a first video frame in the target video to determine multiple regions in the first video frame; obtaining a quality prediction score for each region, and determining a weight corresponding to the region based on the quality prediction score of the region, the weight being positively correlated with the quality prediction score; determining the quality score of the target video based on the quality prediction score of each region and the corresponding weight.
[0006] In some embodiments, a first video frame in a target video is segmented into regions to determine multiple regions in the first video frame, including: determining a supporting image corresponding to the target video, the supporting image having a label mask, the label mask being used to annotate the semantic categories to which each pixel in the supporting image belongs, the semantic categories in the label mask including the semantic categories to which the regions in the first video frame of the target video belong; inputting the first video frame and the supporting image into a small sample semantic segmentation network to obtain multiple regions in the first video frame, the multiple regions belonging to different semantic categories, and the small sample semantic segmentation network is trained by a small sample learning task.
[0007] In some embodiments, a first video frame and a supporting image are input into a small sample semantic segmentation network to obtain multiple regions in the first video frame, including: inputting the supporting image and the first video frame into a backbone network in the small sample semantic segmentation network respectively, and extracting multiple levels of first feature representations corresponding to the supporting image and multiple levels of second feature representations corresponding to the first video frame respectively; determining first video frame features and alignment features based on the first feature representation and the second feature representation; fusing an initial query vector and the alignment features to obtain a query vector; performing mask prediction based on the first video frame features and the query vector to determine a mask prediction result of the first video frame, wherein different masks in the mask prediction result correspond to different regions in the first video frame; performing category prediction based on the query vector to determine a category prediction result of the first video frame, wherein the category prediction result includes the semantic category to which the mask belongs.
[0008] In some embodiments, based on the first feature representation and the second feature representation, determining the first video frame feature includes: calculating a weighted prototype of the supporting image based on the low-level first feature representation and the label mask of the supporting image, the weighted prototype represents different weights assigned to different pixels in the first feature representation according to the label mask; calculating a similarity feature between the supporting image and the first video frame based on the high-level first feature representation and the high-level second feature representation; obtaining the first video frame feature based on the weighted prototype, the similarity feature and the low-level second feature representation.
[0009] In some embodiments, after calculating the similarity features of the supporting image and the first video frame based on the high-level first feature representation and the high-level second feature representation, the method further includes: performing a maximum pooling operation on the similarity features to obtain the importance weight of each pixel in the similarity features; and determining the weighted similarity features based on the importance weight of each pixel in the similarity features.
[0010] In some embodiments, determining the alignment feature based on the first feature representation and the second feature representation includes: determining the alignment feature based on a feature alignment result of the high-level first feature representation and the high-level second feature representation.
[0011] In some embodiments, a small sample semantic segmentation network is trained by: obtaining a small sample training set, the small sample training set includes multiple small sample learning tasks, the small sample learning tasks include paired sample support images and sample query images, the sample support images have support labels, the support labels are used to mark the semantic categories to which each pixel in the sample support images belongs, and the sample query images have query labels, the query labels are used to mark the semantic categories to which each pixel in the sample query images belongs; the sample support images and the sample query images are respectively input into the backbone network in the small sample segmentation network, and the first sample features of multiple levels corresponding to the sample support images and the multiple first sample features corresponding to the sample query images are respectively extracted. The first sample feature and the second sample feature of the layer are used as the second sample feature; based on the first sample feature and the second sample feature, the sample query image feature and the sample alignment feature are determined; the initial sample query vector and the sample alignment feature are fused to obtain the sample query vector; mask prediction is performed based on the sample query image feature and the sample query vector to determine the mask prediction samples of the sample query image, and the mask prediction samples include multiple sample masks; category prediction is performed based on the sample query vector to determine the category prediction samples of the sample mask in the sample query image, and the category prediction samples include the semantic category to which the sample mask belongs; based on the mask prediction samples, the category prediction samples and the query label, the small sample initial segmentation network is trained to obtain the small sample semantic segmentation network.
[0012] In some embodiments, obtaining the quality prediction score of each region includes: obtaining the spatiotemporal characteristics of the target video; determining the quality prediction scores of each pixel in each first video frame corresponding to the spatiotemporal characteristics based on the spatiotemporal characteristics; and determining the quality prediction score of each region based on the quality prediction scores of the pixels in each region.
[0013] In some embodiments, obtaining the spatiotemporal features of a target video includes: preprocessing the target video to obtain preprocessed video data, wherein the amount of the preprocessed video data is smaller than that of the target video; and extracting spatiotemporal features from the preprocessed video data to obtain the spatiotemporal features of the target video.
[0014] In some embodiments, the target video is preprocessed to obtain preprocessed video data, including: segmenting the video frames in the target video to obtain multiple grids; for each grid, selecting a pixel block corresponding to the grid; dividing multiple continuous video frames in the target video to obtain multiple video time periods; sampling each video time period to determine the continuous sampling frames corresponding to the video time period; for each continuous sampling frame, determining the sampling segment corresponding to the pixel block; and splicing the sampling segments to obtain preprocessed video data.
[0015] In some embodiments, the weight corresponding to the region is determined based on the quality prediction score of the region, including: for a first video frame, calculating the sum of the quality prediction scores of each region in the first video frame to obtain a sum of scores; for each region in the first video frame, determining the weight corresponding to the region based on the ratio of the quality prediction score of the region to the sum of scores.
[0016] In some embodiments, the quality score of the target video is determined based on the quality prediction score and the corresponding weight of each region, including: for each first video frame, according to the weight corresponding to each region in the first video frame, the quality prediction scores of each region are aggregated to obtain the quality prediction score of the first video frame; according to the quality prediction scores of multiple first video frames in the target video, the quality score of the target video is determined.
[0017] In some embodiments, determining a weight corresponding to a region based on a quality prediction score of the region includes: determining a semantic category to which the region belongs; and determining a weight corresponding to the region based on the quality prediction score of the region and the semantic category to which the region belongs.
[0018] In a second aspect, an embodiment of the present specification provides a video quality assessment device, including: a video acquisition unit, configured to acquire a target video to be quality assessed, the target video including a plurality of continuous video frames;
[0019] A region segmentation unit is configured to perform region segmentation on a first video frame in a target video to determine a plurality of regions in the first video frame;
[0020] a score acquisition unit, configured to acquire a quality prediction score of each region, and determine a weight corresponding to the region based on the quality prediction score of the region, wherein the weight is positively correlated with the quality prediction score;
[0021] The score determination unit is configured to determine the quality score of the target video according to the quality prediction score of each region and the corresponding weight.
[0022] In a third aspect, an embodiment of the present specification provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in any implementation manner in the first aspect is implemented.
[0023] In the solution provided in the above-mentioned embodiments of the present specification, multiple regions in the first video frame are obtained through region segmentation and the quality prediction score of each region is obtained, and then a corresponding weight is assigned to each region according to its quality prediction score. This weighting mechanism can dynamically adjust the contribution of the region to the overall quality assessment score according to its importance, ensuring that the high-quality region has a greater impact on the final score and the impact of the low-quality region is effectively suppressed. This is in line with the human eye's perception habit of visual quality, that is, paying more attention to areas with higher clarity, so that the overall score is more in line with the actual viewing experience of the video. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only the multiple embodiments disclosed in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0025] Figure 1 is a flowchart of a video quality assessment method in an embodiment of this specification;
[0026] Figure 2 is a flowchart of the training process of a small sample semantic segmentation network in an embodiment of this specification;
[0027] Figure 3 It is a framework diagram of a small semantic segmentation network in an embodiment of this specification;
[0028] Figure 4 is a framework diagram of video quality assessment in an embodiment of this specification;
[0029] Figure 5 is another flow chart of the video quality assessment method in the embodiment of this specification;
[0030] Figure 6 It is a structural diagram of a video quality assessment device in an embodiment of this specification. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be described clearly and completely below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0032] As mentioned above, for UGC videos, in order to optimize video content review, recommendation mechanism and improve user viewing experience, it is necessary to evaluate their quality. In traditional technologies, there are mainly the following evaluation methods:
[0033] The first method is manual evaluation, which involves people manually rating videos based on their actual viewing experience. However, since there is no objective standard, different people may give very different scores, which is not only inefficient but also difficult to meet the processing needs of large-scale video content.
[0034] The second method is to extract the visual features of the video for evaluation based on the traditional video quality assessment model. These models usually focus on different image quality indicators in the video during the evaluation, such as image detail loss, noise, blur, artifacts and color distortion. For complex and changeable UGC videos with great quality differences, these models often cannot fully consider the importance of different areas in the video in the human eye based on human visual perception. For example, when there is a subject in the video, the human eye tends to focus on the clarity of the subject and not the details of the background area. The traditional video quality assessment will give equal weight to the image quality of the subject and the background area. The quality assessment results are difficult to reflect the actual visual perception of the video quality by the human eye.
[0035] To this end, the embodiment of this specification proposes a video quality assessment method. In the process of video quality assessment, firstly, a plurality of regions are obtained by segmenting the first video frame in the target video and obtaining the quality prediction score of each region, and then a corresponding weight is assigned to each region according to the quality prediction score. This weighting mechanism can dynamically adjust the contribution of the region to the overall quality assessment result according to its importance, ensuring that the high-quality region that the human eye pays more attention to has a greater impact on the final score, while the impact of the low-quality region that the human eye will ignore is effectively suppressed, which is in line with the human eye's perception of visual quality, making the overall score more in line with people's actual viewing experience of the video.
[0036] The video quality assessment method provided by this embodiment is described below.
[0037] Figure 1 This is a flow chart of the video quality assessment method in the embodiment of this specification. The process can be performed by any device, platform or device cluster with computing and processing capabilities, including steps S101-S104 as shown below.
[0038] like Figure 1 As shown, in step S101, a target video to be quality evaluated is obtained.
[0039] The target video includes a plurality of continuous video frames. The present embodiment does not limit the method for obtaining the target video. For example, a UGC video shot or edited by a user and uploaded by the user may be obtained.
[0040] Next, in step S102, the first video frame in the target video is segmented into regions to determine a plurality of regions in the first video frame.
[0041] For ease of description, in this embodiment, the video frame in the target video that undergoes region segmentation processing is referred to as a first video frame. The first video frame may be any one or more of a plurality of continuous video frames in the target video.
[0042] The region refers to the area occupied by objects of different categories in the video frame, for example, the area where objects such as people, animals, texts, various objects and background are located.
[0043] In practice, edge detection may be used to detect edges in the first video frame and segment them to obtain different regions, or semantic segmentation may be used to detect regions belonging to different semantic categories in the first video frame. This embodiment does not limit the image processing method used to implement region segmentation.
[0044] Then, in step S103, the quality prediction score of each region is obtained, and the weight corresponding to the region is determined based on the quality prediction score of the region.
[0045] Specifically, when obtaining the quality prediction score of each region, the quality prediction score of each pixel in the first video frame can be obtained by first using the image quality assessment model, and then the quality prediction score of each region can be obtained by combining the quality prediction scores of each pixel; or for each region, the image quality assessment model can be used to score the region, thereby obtaining the quality prediction scores of different regions. This embodiment does not limit the specific method for obtaining the quality prediction score of each region.
[0046] In order to capture the spatiotemporal characteristics of the target video and thus accurately evaluate the video quality, as an implementation method, when obtaining the quality prediction score of each region, the spatiotemporal characteristics of the target video can be obtained, and based on the spatiotemporal characteristics, the quality prediction scores of each pixel in each first video frame corresponding to the spatiotemporal characteristics are determined; and according to the quality prediction scores of the pixels in each region, the quality prediction score of each region is determined.
[0047] When acquiring the spatiotemporal features of the target video, a feature extraction network such as a spatiotemporal graph convolutional network and a 3D convolutional neural network may be used to extract features from multiple continuous video frames of the target video to obtain the spatiotemporal features.
[0048] Exemplarily, the feature extraction network of this embodiment can use 3D Swin Transformer. In the process of local feature extraction, feature sampling can be performed through a local window of fixed size, so that the network can effectively capture image details and textures while maintaining the structure and spatial information of the image. In order to better capture the timing information in the video, a continuous frame sampling strategy based on the time dimension can be introduced in the 3D Swin Transformer. This strategy ensures the retention of video time changes by sampling continuous frames, and encodes features in the spatiotemporal dimension through a shifted window operation (SHifted Windows), thereby enhancing the network's ability to model global features, so that the network can simultaneously capture local details and global information of the video, thereby effectively processing rapidly changing dynamic scenes in the video.
[0049] In the feature extraction process after the video data of the target video is input into the 3D Swin Transformer, at each layer of the network, local features and global features are gradually fused and refined, which improves the expressive power of the extracted features and enables the network to extract more accurate spatiotemporal information at different levels, thereby further optimizing the representation of spatiotemporal features through multi-level feature fusion. The combination of the above local window and shift operations not only reduces the computational complexity, but also effectively improves the efficiency of feature extraction and avoids the high computational consumption of the traditional global self-attention mechanism. 3DSwin Transformer can accurately capture the spatiotemporal features in the video while ensuring computational efficiency, significantly improving the accuracy and real-time performance of the video quality assessment task.
[0050] As an implementation method, before extracting the spatiotemporal features of the target video, the target video can be preprocessed to obtain preprocessed video data. The amount of data of the preprocessed video data is smaller than that of the target video. The preprocessing is used to reduce the amount of data that needs to be processed by the feature extraction network and speed up the video processing. Then, the spatiotemporal features of the preprocessed video data are extracted to obtain the spatiotemporal features of the target video.
[0051] For example, a preprocessed video can be obtained by efficiently selecting and sampling key video frames and global information in the target video, or a video summary can be automatically generated based on the target video through processing by a neural network model, thereby selecting video clips with representative or key information and reducing the content that needs to be processed.
[0052] As an implementation method, preprocessing can be a local, small-size grid sampling of the target video to avoid the high computational complexity of full-frame sampling while retaining important quality-related details in the video. Specifically, the following processing steps are included: segmenting the video frames in the target video to obtain multiple grids; for each grid, selecting the pixel blocks corresponding to the grid; dividing multiple continuous video frames in the target video to obtain multiple video time periods; sampling each video time period to determine the continuous sampling frames corresponding to the video time period; for each continuous sampling frame, determining the sampling segments corresponding to the pixel blocks; splicing the sampling segments to obtain preprocessed video data.
[0053] Taking the GMS (Global Memory Sampling) sampling method in Fast-VQA (Video Quality Assessment) as an example, each video frame is first divided into multiple G f ×G f To maintain the local texture features of the video frame, for each grid, a grid of size S is randomly selected inside the grid. f ×S f Then, we continue to sample the time dimension of the target video to further improve the subsequent computational efficiency. Assuming that the target video contains continuous T frames, we divide the T frames in the target video into G t In each video time segment, select continuous frames for sampling and determine the time length corresponding to the video time segment as S t The continuous sampling frames can be regarded as the spatiotemporal grid g k,i,j , where k represents the kth video time segment, i and j are the position indexes in the spatial grid, representing the specific position of the patcH divided in each video frame to ensure that the temporal characteristics of the video are preserved. At the same time, the consistency of the corresponding mini-patcH in the spatiotemporal dimension is maintained, and the corresponding mini-patcH is obtained from each spatiotemporal grid g. k,i,j The position corresponding to mini-patcH is continuously sampled to obtain a size of T f ×S f ×S f Mini-cube(MC k,i,j ), and all the sampled segments are spliced along the time dimension to obtain the preprocessed video data. In order to ensure the timing consistency between video frames, the processing processes of all the above frames are strictly aligned in time.
[0054] In this way, through the preprocessing step, we extract a mini-cube (MC) with a small amount of data from the target video but which can reflect the global and local information of the target video. k,i,j ), this method not only reduces the computational complexity, but also avoids excessive computation and loss of spatial information, and can retain the local texture features in the image and ensure the integrity of detail information.
[0055] Extracting spatiotemporal features from the preprocessed video can obtain the spatiotemporal features f fast Next, we will continue to analyze the spatiotemporal features f fast , determine the spatiotemporal characteristics f fast The quality prediction scores of the pixels in each corresponding first video frame and the process of determining the quality prediction score of each region are described according to the quality prediction scores of the pixels in each region.
[0056] Specifically, we can analyze the spatiotemporal features f fast Up-sampling is performed and it is mapped into a feature with the same size as the video frame, and then the quality score of the feature is predicted. For example, the feature is input into a pre-trained quality score prediction network, and the quality prediction score of each pixel in each first video frame corresponding to the feature is output. It should be noted that, in this embodiment, the first video frame is a video frame in the target video corresponding to the spatiotemporal feature, that is, a continuous sampling frame of the above preprocessing process.
[0057] After determining the quality prediction score of each pixel, the quality prediction score of the pixels in each region may be averaged to obtain the quality prediction score of the region. In other examples, the quality prediction score of the region may be determined by taking the median, mode, or other methods. This implementation effectively combines spatial and temporal information by mapping spatiotemporal features to the size of the video frame and aligning them with the region segmentation results to accurately evaluate video quality.
[0058] After obtaining the quality prediction score of the region, the weight corresponding to the region is determined based on the quality prediction score of the region.
[0059] Among them, the weight of the region is positively correlated with the quality prediction score. Usually different regions in an image have different visual significance. The human eye system will pay more attention to the regions with better quality. Different regions in the image are not equally weighted in the human eye. They should also be distinguished when calculating the quality score of the image. Therefore, it is proposed to weight the regions based on the quality score, giving higher weights to the regions with better quality that the human eye pays more attention to, and vice versa, giving lower weights to the regions with poor quality that the human eye tends to ignore.
[0060] For example, weights may be assigned in proportion to the quality prediction scores of different regions, or a nonlinear mapping function may be constructed to map the quality prediction scores of different regions into weights.
[0061] As an implementation method, for the first video frame, the sum of the quality prediction scores of each region in the first video frame can be calculated to obtain the sum of the scores; for each region in the first video frame, the weight corresponding to the region can be determined based on the ratio of the quality prediction score of the region and the sum of the scores.
[0062] Specifically, for each first video frame, a weight is assigned to each region based on the quality prediction score of each region by the following formula: Among them, w i is the weight of region i, is the quality prediction score of the region, and N is the total number of regions. In this way, regions with higher quality receive larger weights, while regions with lower quality receive smaller weights.
[0063] Continue to view Figure 1 In step S104, the quality score of the target video is determined according to the quality prediction score of each region and the corresponding weight.
[0064] As an implementation method, for each first video frame, the quality prediction scores of each area in the first video frame are aggregated according to the weight corresponding to each area to obtain the quality prediction score of the first video frame; and the quality score of the target video is determined according to the quality prediction scores of multiple first video frames in the target video.
[0065] For example, for each first video frame, weighted processing may be performed according to the quality prediction score of each region and the corresponding weight to obtain the quality prediction score of the first video frame, and then the quality score of the target video may be determined according to the average value or sum of the quality prediction scores of multiple first video frames.
[0066] In other implementations, all regions in the target video may be weighted according to their quality prediction scores and corresponding weights to determine the quality score of the target video. Finally, the overall quality evaluation Q of the video is total is the sum of all region weighted scores: This weighting mechanism ensures that the key areas in the video contribute more to the evaluation results and improves the accuracy of video quality assessment.
[0067] In the solution provided by the above embodiment, multiple regions in the first video frame are obtained through region segmentation and the quality prediction score of each region is obtained, and then a corresponding weight is assigned to each region according to its quality prediction score. This weighting mechanism can dynamically adjust the contribution of the region to the overall quality assessment score according to its importance, ensuring that high-quality regions have a greater impact on the final score and the impact of low-quality regions is effectively suppressed. This is in line with the human eye's perception habits of visual quality, that is, more attention is paid to areas with higher clarity, so that the overall score is more in line with the actual viewing experience of the video.
[0068] In fact, when watching different types of videos, the audience's focus is also different. For example, when watching a landscape video, the audience pays more attention to the landscape area, and when watching a dance video, the audience pays more attention to the area where the human body is located. Considering that the human eye pays attention to different areas in different types of videos, the semantic category to which the area belongs can be considered when assigning weights to different areas. When the above embodiment determines the weight corresponding to the area based on the quality prediction score of the area, it can be determined that the semantic category to which the area belongs, and the weight corresponding to the area is determined based on the quality prediction score of the area and the semantic category to which it belongs.
[0069] In practice, determining the semantic category to which a region belongs can be performed by obtaining the semantic categories to which different regions belong while obtaining multiple regions in a semantic segmentation manner during region segmentation.
[0070] When determining the weight corresponding to the region based on the quality prediction score of the region and the semantic category to which it belongs, a mapping relationship between the type of the target video and different semantic categories and their corresponding weight coefficients may be established in advance. When applied, the weight coefficients of the semantic categories to which different regions belong may be determined according to the mapping relationship. For example, the weight determined based on the quality prediction score of the region in the above step S104 may be multiplied by the weight coefficient determined based on the semantic category to obtain the final weight corresponding to the region. For example, for dance videos, the weight coefficient of the semantic category "person" is set to 1.5, and the weight coefficient of the semantic category "background" is set to 0.5. For landscape videos, the weight coefficient of the semantic category "background" is set to 1.5, and the weight coefficient of the semantic category "person" is set to 0.5. It is also possible to give a higher weight to a key semantic category and a lower weight to other semantic categories. For example, the key semantic category may be "face", "specific object", etc., so that when evaluating the video quality, special attention is paid to the key areas (such as the face of a person, a specific object or a close-up) in the target video, thereby providing a more detailed quality evaluation result.
[0071] Traditional semantic segmentation methods require a large number of manually labeled samples to improve the accuracy of semantic segmentation. In addition, in scenarios where UGC video content is extremely diverse, it takes a lot of time to collect and label different types of samples. As an implementation method, small sample semantic segmentation can be used to identify different areas of the first video frame, thereby avoiding the reliance on a large number of labeled samples in traditional methods.
[0072] In this implementation, it is first necessary to obtain a small sample semantic segmentation network through training of a small sample learning task. When semantically segmenting the first video frame, determine the support image corresponding to the target video, wherein the support image serves as a guide for the small sample semantic segmentation network, and the small sample semantic segmentation network understands the basic features of different semantic categories from the support image. The support image has a label mask, and the label mask is used to mark the semantic category to which each pixel in the support image belongs. The semantic category in the label mask includes the semantic category to which the area in the first video frame of the target video belongs. For example, when there is a human body in the first video frame, there should also be a human body in the support image, and the pixels corresponding to the position of the human body are marked with the semantic category "human body" in the label mask.
[0073] Then, the first video frame and the supporting image are input into a small sample semantic segmentation network to obtain multiple regions in the first video frame, and the multiple regions belong to different semantic categories.
[0074] Among them, the small sample semantic segmentation network segments the area in the query image belonging to the semantic category according to the semantic categories annotated in a small number of supporting images, and inputs the first video frame and the supporting image into the small sample semantic segmentation network. Based on the features of the supporting image, the prototype of the corresponding semantic category is generated. The small sample semantic segmentation network will match the prototypes of the first video frame and the supporting image, and segment the first video frame into multiple regions according to the similarity, and each region corresponds to a semantic category.
[0075] Unlike traditional machine learning methods that require a large amount of labeled data to train models, small sample learning can still train a model with good performance using limited sample data. In this way, the model trained with a small amount of labeled data (i.e., small sample learning task) can obtain excellent region segmentation capabilities and can accurately identify important areas in the video (such as people, objects, background, etc.) under limited labeled data conditions.
[0076] The following is an explanation of the training process of the small sample semantic segmentation network used in the embodiments of this specification.
[0077] The training process can be performed by any device, platform or device cluster with computing and processing capabilities, including steps S201-S207 as shown below. Figure 3The framework diagram of the small semantic segmentation network shown introduces the training process.
[0078] like Figure 2 As shown, in step S201, a small sample training set is obtained.
[0079] First, select a small number of samples from the labeled data as the support set. For example, randomly select 1 to 5 images from the labeled data for each semantic category as sample support images in the support set. The sample support image has a support label, which is used to annotate the semantic category to which each pixel in the sample support image belongs. The support label contains complete pixel-level annotation information and is used to guide the model to learn the features of each category.
[0080] Next, we select a query set from the labeled data, and the sample query images in the query set The sample query image has a query label, which is used to annotate the semantic category to which each pixel in the sample query image belongs and is used to calculate the training loss of the model.
[0081] Subsequently, the sample images in the support set and the query set are paired. For example, each pair of pairing tasks includes a sample support image and a sample query image, and the semantic category in the sample query image exists in the sample support image, such as Figure 3 As shown in the figure, both the sample support image and the sample query image contain the semantic category "person". This task pairing method generates multiple small sample learning tasks, ensuring that the model can learn effectively with the support of limited labeled data.
[0082] In order to increase the diversity of images, data augmentation techniques can be applied to sample images in the support set and query set, including operations such as image rotation, flipping, cropping, and color transformation. Through these augmentation methods, the diversity of training data is expanded, thereby improving the robustness of the model. During the augmentation process, ensure that the labels of sample images in the support set and query set are consistent to avoid interference with label information, that is, no matter what kind of augmentation operation is performed on the image in the support set or the sample image in the query set, the label information should always remain correct. For example, if a pixel in a sample support image in the support set represents the label "dog", it should still represent the label "dog" after rotation or cropping.
[0083] Through the above processing, a small sample training set is effectively constructed, which can support video quality assessment tasks based on small sample learning and significantly improve the learning and generalization capabilities of the model with a small amount of labeled data.
[0084] Next, in step S202, the sample support image and the sample query image are respectively input into the backbone network in the small sample segmentation network, and the first sample features of multiple levels corresponding to the sample support image and the second sample features of multiple levels corresponding to the sample query image are respectively extracted.
[0085] Specifically, the sample supports images and a sample query image Features can be extracted through a weight-sharing backbone network. Each layer of the backbone network extracts features corresponding to that layer of the network. For example, the first sample features of the four layers can be obtained. and four levels of second sample features Where l is the hierarchical index of the feature.
[0086] Next, in step S203, based on the first sample feature and the second sample feature, a sample query image feature and a sample alignment feature are determined.
[0087] In practice, this step can be performed by a region alignment module in a small sample segmentation network, which can align the first sample features and the second sample features at different levels to obtain sample query image features and sample alignment features, respectively.
[0088] As an implementation method, the sample query image feature can be obtained by prototype alignment using the low-level first sample feature and the second sample feature. The low level here refers to the network level close to the input end in the backbone network, for example, l=1 and l=2 are low levels, and the network level close to the output end in the backbone network is a high level, for example, l=3 is a high level. The prototype alignment process can be based on the low-level first sample feature and the support label of the sample support image, and the weighted prototype of the sample support image is calculated. The weighted prototype represents the different weights assigned to different pixels in the first sample feature according to the support label; based on the high-level first sample feature and the high-level second sample feature, the similarity feature of the sample support image and the sample query image is calculated; based on the weighted prototype, the similarity feature and the low-level second sample feature, the sample query image feature is obtained.
[0089] Specifically, first, the low-level features (layer 1 and layer 2) are selected and the low-level features are reduced in dimension through 1×1 convolution to obtain the sample support image features after dimension reduction. and sample query image features To reduce the computational complexity.
[0090] Next, based on the mask M in the support label of the sample support image s , calculate the weighted prototype of the sample support image:
[0091]
[0092] Among them, (x, y) is the coordinate position of the pixel in the feature, and I is the indicator function, which is used to determine whether the pixel (x, y) in the feature is in the mask area. If it is in the mask area, the output value is 1, otherwise the output value is 0, thereby calculating the weighted prototype of the sample support image.
[0093] In order to align the important regions in the sample query image with the regions of the corresponding semantic categories in the sample support image, the pixel-level similarity feature s of the sample support image and the sample query image can be calculated. p ∈R hw×hw :
[0094]
[0095] in, It is the third-layer feature of the sample support image and the sample query image. Sim stands for Similarity, which is a similarity measurement function used to measure the similarity between pixels.
[0096] In addition, in order to further determine the important areas in the similarity features and help the model focus on the core semantic features, thereby improving its performance and generalization ability in different scenarios, the similarity features can also be subjected to a maximum pooling operation to obtain the importance weight of each pixel in the similarity feature, thereby assigning an importance weight to each pixel.
[0097] Specifically, for the similarity feature s p The maximum pooling operation can be performed by selecting a sliding window and then selecting the maximum value of the similarity features corresponding to each pixel in the window. In this way, the most significant similarity feature in each window area can be extracted and used as the importance weight of each pixel in the window. According to the importance weight of each pixel in the similarity feature, the pixels in the feature are weighted to determine the weighted similarity feature. For example, s p =Max(s p , dim=0), an importance weight is assigned to the feature corresponding to each pixel to indicate its importance.
[0098] In order to concatenate the weighted prototype, similarity features and low-level second sample features, the above features need to be converted to the same size. For example, the weighted prototype can be expanded, the similarity features can be reshaped, and the expanded weighted prototype p∈R C×h×w , similarity features after reshaping and the second sample characteristics Splicing is performed in the channel dimension, where C is the number of channels, that is, the number of semantic categories, to obtain the fusion feature:
[0099]
[0100] Then, the fusion feature Input to the pixel decoder for upsampling to obtain the final sample query image features:
[0101]
[0102] Among them, PDecoder (Pixel-Decoder) is a pixel decoder, which obtains the final sample query image features by upsampling the fusion features. The sample query image features combine the attention weighting of important areas, the weighted prototype information of the sample support image, and the fusion of multi-level features, which can effectively improve the semantic understanding and image matching capabilities of the model, strengthen the important areas in the sample query image, and reduce the interference of irrelevant areas. This enables the model to process the sample query image more accurately while ensuring computational efficiency, thereby improving the overall performance.
[0103] As an implementation method, the sample alignment feature can be obtained by aligning the high-level first sample feature and the second sample feature. The feature alignment process can be to perform feature splicing on the high-level first sample feature and the high-level second sample feature, and perform feature transformation and mapping on the spliced features, so that the number of features can be gradually reduced.
[0104] Specifically, feature alignment can be performed by a meta-alignment module, which is composed of a stack of multiple alignment blocks, each of which includes a transformer layer and an alignment layer. s 3 and Reshape into R hw×C shape, and then aligned and concatenated as the input of the meta-alignment module Where L = hw, It is sent to the transformer layer of the first alignment block. In each alignment block, the transformer layer exchanges the information of the first sample feature and the second sample feature to obtain the fused transformation feature:
[0105]
[0106] Among them, i is the index of the alignment block, and TLayer is Transformer Layers.
[0107] For each alignment block, after completing the TLayer calculation, Re-divide into the features of the support part (i.e. the part corresponding to the first sample feature) and the features of the query part (i.e. the part corresponding to the second sample feature), and use linear mapping to reduce the number of feature regions in the query part:
[0108]
[0109] After completing the calculation of multiple alignment blocks, we use the features of the query part output by the last alignment block as the sample alignment features. Used in the subsequent Meta-Mask prediction module.
[0110] As an implementation method, an auxiliary loss function can be added here to constrain the accuracy of the output of the region alignment module. Specifically, the aligned feature f can be calculated a With f s 3 The similarity between Then, the most relevant f for each pixel is assigned according to the similarity s. a With f s 3 The concatenated features are then processed using a 1×1 convolution to generate pixel predictions for the semantic category of each pixel in the query image. Through the following auxiliary loss function L aux To constrain the output:
[0111]
[0112] in, is the true label of the i-th pixel in the sample support image, is the predicted value of the model.
[0113] This loss function constrains the aligned features f a Better prediction of pixel labels in the support image is used to filter out features in the sample query image that are most relevant to the sample support image, forcing the model to pay more attention to different areas of the sample support image, thereby improving the accuracy of alignment.
[0114] Through the region alignment module, the features of the sample query image and the sample support image can be effectively aligned, and the important regions can be automatically focused on to optimize the feature learning process. Through multi-layer feature fusion, similarity calculation and pixel-level attention mechanism, this module significantly improves the segmentation accuracy and video quality assessment effect in small sample learning.
[0115] Next, in step S204, the initial sample query vector and the sample alignment feature are fused to obtain a sample query vector.
[0116] In practice, this step is handled by the Meta-Mask module in the small sample segmentation network, which is used to decode the sample alignment features and generate the final output prediction.
[0117] First, initialize N learnable initial sample query vectors N is a custom value, and it is aligned with the sample feature f a Fusion is performed and processed by the Transformer decoder TDecoder to generate an updated sample query vector This process completes the encoding of the global information of sample alignment features.
[0118] Next, in step S205 , mask prediction is performed based on the sample query image features and the sample query vector to determine a mask prediction sample of the sample query image.
[0119] Among them, the mask prediction sample contains multiple sample masks. Sample query vector is fed into two independent linear branches to predict the mask and semantic category respectively. In the mask branch, the linear layer Mapping to mask embedding and through the pixel decoder output Multiply to get the mask prediction sample In each sample mask, pixels with a value of 0 represent the background area, and pixels with a value of 1 represent the foreground area.
[0120] In step S206, category prediction is performed based on the sample query vector to determine a category prediction sample of the sample mask in the sample query image.
[0121] Among them, the category prediction sample contains the semantic category to which the sample mask belongs. In the category branch, the query vector is mapped to class prediction Where C represents the total number of semantic categories, and here we get the semantic category to which the foreground area of each sample mask belongs.
[0122] Finally, in step S207, the small sample initial segmentation network is trained based on the mask prediction samples, the category prediction samples and the query label to obtain a small sample semantic segmentation network.
[0123] To optimize the mask and category prediction, the cross-entropy loss function is used to calculate the classification loss L cls , the mask loss L is calculated using focal loss and dice loss mask :
[0124]
[0125] L mask =L fl +L dl (9)
[0126] Among them, the focus loss L fl for:
[0127]
[0128] Dice loss L dl for:
[0129]
[0130] in, For the model to query pixels in the sample image The predicted semantic category, M j The true mask value in the sample query image.
[0131] Finally, through the comprehensive loss function L total To optimize the entire network:
[0132] L total =λ cls L cls +λ msk L mask +λ aux L aux (12)
[0133] Among them, λ cls , msk , aux are the weighting coefficients for classification, masking and auxiliary losses, respectively.
[0134] In other embodiments, L cls and L mask Perform weighted summation to obtain the comprehensive loss function.
[0135] The training framework of the small sample semantic segmentation network provided in the embodiments of this specification utilizes a small amount of labeled data to train the model to obtain region segmentation capability, and adopts feature alignment and mask fusion to improve the semantic segmentation capability of the model. The trained model can efficiently perform region-level semantic segmentation under limited labeled samples, so that video quality assessment can still accurately capture key area information in images and videos under small sample conditions, avoiding the reliance on a large number of labeled samples and manual feature extraction in traditional methods.
[0136] Next, combine Figure 4 The video quality assessment framework diagram introduces the complete process of video quality assessment. Figure 5 This is another flow chart of the video quality assessment method in the embodiment of this specification. The method can be executed by any device, platform or device cluster with computing and processing capabilities, including steps S501-S507 as shown below.
[0137] like Figure 5 As shown, in step S501, a target video to be quality evaluated is obtained.
[0138] See also Figure 4 ,The target video is a video of a man playing tennis, and the semantic category it contains is mainly "person".
[0139] Next, in step S502, the target video is preprocessed to obtain preprocessed video data.
[0140] For the explanation of step S502, please refer to the relevant description of step S103 in the previous text, which will not be repeated here.
[0141] Continuing, in step S503, a supporting image corresponding to the target video is determined, and the first video frame and the supporting image are input into a small sample semantic segmentation network to obtain multiple regions in the first video frame.
[0142] In practice, one or more video frames may be selected from the target video as support images, and the support images may be annotated pixel-wise to obtain label masks. In other embodiments, images that also contain the semantic category "person" may be selected as support images.
[0143] Specifically, the first video frame can be a video frame in the target video corresponding to the preprocessed video data in the previous step, and the first video frame in the supporting image and the target video is input into Figure 4 The backbone network in the small and medium sample semantic segmentation network extracts the first feature representations of multiple levels corresponding to the support image and the second feature representations of multiple levels corresponding to the first video frame; determines the first video frame features and alignment features based on the first feature representations and the second feature representations; fuses the initial query vector and the alignment features to obtain the query vector; and then Figure 4 The mask fusion module in the method performs: performing mask prediction based on the first video frame feature and the query vector to determine the mask prediction result of the first video frame, where different masks in the mask prediction result correspond to different areas in the first video frame; performing category prediction based on the query vector to determine the category prediction result of the first video frame, where the category prediction result includes the semantic category to which the mask belongs.
[0144] As an implementation method, based on the first feature representation and the second feature representation, determining the first video frame feature can be performed by Figure 4The prototype calculation module and the pixel decoder module in the method are executed, specifically, based on the low-level first feature representation and the label mask of the support image, the weighted prototype of the support image is calculated, and the weighted prototype represents different weights assigned to different pixels in the first feature representation according to the label mask; based on the high-level first feature representation and the high-level second feature representation, the similarity feature between the support image and the first video frame is calculated; based on the weighted prototype, the similarity feature and the low-level second feature representation, the first video frame feature is obtained.
[0145] As an implementation method, after calculating the similarity features of the supporting image and the first video frame based on the high-level first feature representation and the high-level second feature representation, it also includes: performing a maximum pooling operation on the similarity features to obtain the importance weight of each pixel in the similarity features; and determining the weighted similarity features based on the importance weight of each pixel in the similarity features.
[0146] As an implementation method, based on the first feature representation and the second feature representation, the alignment feature can be determined by Figure 4 The regional alignment module in the embodiment is executed, including: determining the alignment feature based on the feature alignment result of the high-level first feature representation and the high-level second feature representation.
[0147] Among them, the explanation of step S503 is similar to the processing process of steps S202-S206 in the previous text, and will not be repeated here. The small sample segmentation network used in this embodiment can be trained by the training method of the above embodiment. Unlike the training stage, in the application stage of the small sample segmentation network, the first video frame has no labeled label.
[0148] It should be noted that this embodiment does not limit the execution order of steps S503 to S505, which may be executed sequentially or in parallel.
[0149] In step S504, spatiotemporal features are extracted from the preprocessed video data to obtain spatiotemporal features of the target video.
[0150] like Figure 4As shown, the pre-processed video data is input into the framework of 3D Swin Transformer, and 3D Swin Transformer contains m 3D Blocks, each of which performs the following operations: local self-attention calculation, using the window method to perform self-attention calculation in each block; layer-by-layer stacking, multiple 3D Blocks are stacked together, and through the stacking of different layers, the model can gradually extract higher-level features; Shifted Windows, using the method of shifting windows, so that the windows between adjacent blocks can overlap each other, thereby enhancing the transfer of features. For a detailed explanation of step S504, please refer to the relevant description in the previous text, which will not be repeated here.
[0151] In step S505, based on the spatiotemporal features, a quality prediction score of each pixel in each first video frame corresponding to the spatiotemporal features is determined.
[0152] In step S506, the quality prediction score of each region is determined according to the quality prediction score of the pixels in each region, and the weight corresponding to the region is determined based on the quality prediction score of the region. Thereafter, step S507 may be performed.
[0153] In step S507, the quality score of the target video is determined according to the quality prediction score of each region and the corresponding weight.
[0154] Figure 5 The solution provided by the corresponding embodiment uses a weighting mechanism based on the region segmentation results. In the process of video quality assessment, the quality prediction score of each region is first obtained through the region segmentation ability of small sample learning, and then the corresponding weight is assigned to each region according to the prediction score. This weighting mechanism can dynamically adjust the contribution of the region to the overall quality assessment result according to its importance, ensuring that the high-quality region has a greater impact on the final score and the impact of the low-quality region is effectively suppressed. Moreover, by aligning the spatiotemporal features with the region segmentation results, the spatial and temporal information in the target video is effectively combined to accurately assess the video quality. This method not only reduces the computational complexity and avoids excessive calculation and loss of spatial information, but also can achieve efficient and accurate video quality assessment through alignment and weighting mechanisms, which is particularly suitable for dynamic video scenes and short video fields.
[0155] Figure 6 Schematic diagram of the structure of the video quality assessment device in the embodiment of this specification. The device can be applied to any device, platform or device cluster with computing and processing capabilities. The device includes:
[0156] A video acquisition unit 601 is configured to acquire a target video to be quality evaluated, wherein the target video includes a plurality of continuous video frames;
[0157] A region segmentation unit 602 is configured to perform region segmentation on a first video frame in the target video to determine a plurality of regions in the first video frame;
[0158] A score acquisition unit 603 is configured to acquire a quality prediction score of each of the regions, and determine a weight corresponding to the region based on the quality prediction score of the region, wherein the weight is positively correlated with the quality prediction score;
[0159] The score determination unit 604 is configured to determine the quality score of the target video according to the quality prediction score of each of the regions and the corresponding weight.
[0160] In some embodiments, the region segmentation unit 602 is specifically configured to: determine a supporting image corresponding to the target video, the supporting image has a label mask, the label mask is used to annotate the semantic category to which each pixel in the supporting image belongs, and the semantic category in the label mask includes the semantic category to which the region in the first video frame of the target video belongs; input the first video frame and the supporting image into a small sample semantic segmentation network to obtain multiple regions in the first video frame, and the multiple regions belong to different semantic categories. The small sample semantic segmentation network is trained by a small sample learning task.
[0161] In some embodiments, the region segmentation unit 602 can be further configured to: input the supporting image and the first video frame into the backbone network in the small sample semantic segmentation network, respectively, and extract the first feature representation of multiple levels corresponding to the supporting image and the second feature representation of multiple levels corresponding to the first video frame, respectively; determine the first video frame features and alignment features based on the first feature representation and the second feature representation; fuse the initial query vector and the alignment features to obtain the query vector; perform mask prediction based on the first video frame features and the query vector to determine the mask prediction result of the first video frame, where different masks in the mask prediction result correspond to different regions in the first video frame; perform category prediction based on the query vector to determine the category prediction result of the first video frame, where the category prediction result includes the semantic category to which the mask belongs.
[0162] In some embodiments, the region segmentation unit 602 can be further configured to: calculate a weighted prototype of the support image based on the low-level first feature representation and the label mask of the support image, the weighted prototype represents different weights assigned to different pixels in the first feature representation according to the label mask; calculate a similarity feature between the support image and the first video frame based on the high-level first feature representation and the high-level second feature representation; and obtain the first video frame feature based on the weighted prototype, the similarity feature and the low-level second feature representation.
[0163] In some embodiments, the region segmentation unit 602 may be further configured to: perform a maximum pooling operation on the similarity feature to obtain an importance weight of each pixel in the similarity feature; and determine a weighted similarity feature according to the importance weight of each pixel in the similarity feature.
[0164] In some embodiments, the region segmentation unit 602 may be further configured to: determine an alignment feature based on a feature alignment result of a high-level first feature representation and a high-level second feature representation.
[0165] In some embodiments, a small sample semantic segmentation network is trained by: obtaining a small sample training set, the small sample training set includes multiple small sample learning tasks, the small sample learning tasks include paired sample support images and sample query images, the sample support images have support labels, the support labels are used to mark the semantic categories to which each pixel in the sample support images belongs, and the sample query images have query labels, the query labels are used to mark the semantic categories to which each pixel in the sample query images belongs; the sample support images and the sample query images are respectively input into the backbone network in the small sample segmentation network, and the first sample features of multiple levels corresponding to the sample support images and the multiple first sample features corresponding to the sample query images are respectively extracted. The first sample feature and the second sample feature of the layer are used as the second sample feature; based on the first sample feature and the second sample feature, the sample query image feature and the sample alignment feature are determined; the initial sample query vector and the sample alignment feature are fused to obtain the sample query vector; mask prediction is performed based on the sample query image feature and the sample query vector to determine the mask prediction samples of the sample query image, and the mask prediction samples include multiple sample masks; category prediction is performed based on the sample query vector to determine the category prediction samples of the sample mask in the sample query image, and the category prediction samples include the semantic category to which the sample mask belongs; based on the mask prediction samples, the category prediction samples and the query label, the small sample initial segmentation network is trained to obtain the small sample semantic segmentation network.
[0166] In some embodiments, the score acquisition unit 603 is further configured to: acquire the spatiotemporal characteristics of the target video; determine the quality prediction score of each pixel in each first video frame corresponding to the spatiotemporal characteristics based on the spatiotemporal characteristics; and determine the quality prediction score of each region based on the quality prediction scores of the pixels in each region.
[0167] In some embodiments, the score acquisition unit 603 is further configured to: preprocess the target video to obtain preprocessed video data, wherein the amount of data of the preprocessed video data is smaller than that of the target video; extract spatiotemporal features from the preprocessed video data to obtain spatiotemporal features of the target video.
[0168] In some embodiments, the score acquisition unit 603 is further configured to: segment the video frames in the target video to obtain multiple grids; for each grid, select the pixel block corresponding to the grid; divide the multiple continuous video frames in the target video to obtain multiple video time periods; sample each video time period to determine the continuous sampling frames corresponding to the video time period; for each continuous sampling frame, determine the sampling segment corresponding to the pixel block; splice the sampling segments to obtain preprocessed video data.
[0169] In some embodiments, the score acquisition unit 603 is further configured to: for the first video frame, calculate the sum of the quality prediction scores of each region in the first video frame to obtain a sum of scores; for each region in the first video frame, determine the weight corresponding to the region based on the ratio of the quality prediction score of the region to the sum of scores.
[0170] In some embodiments, the score determination unit 604 is further configured to, for each first video frame, aggregate the quality prediction scores of each region according to the weight corresponding to each region in the first video frame to obtain the quality prediction score of the first video frame; and determine the quality score of the target video according to the quality prediction scores of multiple first video frames in the target video.
[0171] In some embodiments, the score acquisition unit 603 is further configured to: determine the semantic category to which the region belongs; and determine the weight corresponding to the region based on the quality prediction score of the region and the semantic category to which the region belongs.
[0172] The present specification also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the following Figures 1 to 5 Describe the method.
[0173] The embodiment of the present specification also provides a computing device, including a memory and a processor, wherein the memory stores an executable code, and when the processor executes the executable code, the following is implemented: Figures 1 to 5 Describe the method.
[0174] The embodiments of the present specification also provide a computer program product, including a computer program / instruction, which is executed by a processor to implement the following Figures 1 to 5 Describe the steps of the method.
[0175] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the multiple embodiments disclosed in this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0176] In some cases, the actions or steps described in the claims may be performed in a different order than in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0177] The specific implementation methods described above further illustrate in detail the purposes, technical solutions and beneficial effects of the multiple embodiments disclosed in this specification. It should be understood that the above description is only the specific implementation methods of the multiple embodiments disclosed in this specification, and is not intended to limit the protection scope of the multiple embodiments disclosed in this specification. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solutions of the multiple embodiments disclosed in this specification should be included in the protection scope of the multiple embodiments disclosed in this specification.
Claims
1. A video quality assessment method, the method comprising: Acquire a target video to be quality assessed, wherein the target video includes a plurality of continuous video frames; Performing region segmentation on a first video frame in the target video to determine a plurality of regions in the first video frame; Obtaining a quality prediction score for each of the regions, and determining a weight corresponding to the region based on the quality prediction score of the region, wherein the weight is positively correlated with the quality prediction score; The quality score of the target video is determined according to the quality prediction score of each of the regions and the corresponding weight.
2. The method according to claim 1, wherein: The performing region segmentation on the first video frame in the target video to determine a plurality of regions in the first video frame includes: Determine a supporting image corresponding to the target video, wherein the supporting image has a label mask, wherein the label mask is used to mark the semantic category to which each pixel in the supporting image belongs, and the semantic category in the label mask includes the semantic category to which the region in the first video frame of the target video belongs; The first video frame and the supporting image are input into a small sample semantic segmentation network to obtain multiple regions in the first video frame, where the multiple regions belong to different semantic categories. The small sample semantic segmentation network is trained through a small sample learning task.
3. The method according to claim 2, wherein: The step of inputting the first video frame and the supporting image into a small sample semantic segmentation network to obtain a plurality of regions in the first video frame includes: Inputting the supporting image and the first video frame into the backbone network in the small sample semantic segmentation network respectively, and extracting first feature representations of multiple levels corresponding to the supporting image and second feature representations of multiple levels corresponding to the first video frame respectively; Determining first video frame features and alignment features based on the first feature representation and the second feature representation; Fusing the initial query vector and the alignment feature to obtain a query vector; Performing mask prediction based on the first video frame feature and the query vector to determine a mask prediction result of the first video frame, where different masks in the mask prediction result correspond to different regions in the first video frame; A category prediction is performed based on the query vector to determine a category prediction result of the first video frame, where the category prediction result includes a semantic category to which the mask belongs.
4. The method according to claim 3, wherein: The determining of a first video frame feature based on the first feature representation and the second feature representation includes: Based on the first feature representation at a low level and the label mask of the support image, a weighted prototype of the support image is calculated, wherein the weighted prototype represents different weights assigned to different pixels in the first feature representation according to the label mask; Based on the first high-level feature representation and the second high-level feature representation, calculate and obtain a similarity feature between the supporting image and the first video frame; Based on the weighted prototype, the similarity feature and the second feature representation at a low level, the first video frame feature is obtained.
5. The method according to claim 4, wherein: After calculating the similarity feature between the supporting image and the first video frame based on the first feature representation at a high level and the second feature representation at a high level, the method further includes: Performing a maximum pooling operation on the similarity feature to obtain an importance weight of each pixel in the similarity feature; The weighted similarity feature is determined according to the importance weight of each pixel in the similarity feature.
6. The method according to claim 3, wherein: The determining of the alignment feature based on the first feature representation and the second feature representation includes: An alignment feature is determined based on a feature alignment result of the first feature representation at a high level and the second feature representation at a high level.
7. The method according to claim 3, wherein: The small sample semantic segmentation network is trained in the following way: Acquire a small sample training set, wherein the small sample training set includes multiple small sample learning tasks, wherein the small sample learning task includes a paired sample support image and a sample query image, wherein the sample support image has a support label, wherein the support label is used to mark the semantic category to which each pixel in the sample support image belongs, and the sample query image has a query label, wherein the query label is used to mark the semantic category to which each pixel in the sample query image belongs; Inputting the sample support image and the sample query image into the backbone network in the small sample segmentation network respectively, and extracting first sample features of multiple levels corresponding to the sample support image and second sample features of multiple levels corresponding to the sample query image respectively; Determine a sample query image feature and a sample alignment feature based on the first sample feature and the second sample feature; Fusing the initial sample query vector and the sample alignment feature to obtain a sample query vector; Performing mask prediction based on the sample query image feature and the sample query vector to determine a mask prediction sample of the sample query image, wherein the mask prediction sample includes a plurality of sample masks; Performing category prediction based on the sample query vector to determine a category prediction sample of the sample mask in the sample query image, wherein the category prediction sample includes a semantic category to which the sample mask belongs; Based on the mask prediction samples, the category prediction samples and the query label, the small sample initial segmentation network is trained to obtain the small sample semantic segmentation network.
8. The method according to claim 1, wherein: The obtaining of the quality prediction score of each region comprises: Acquire the spatiotemporal characteristics of the target video; Based on the spatiotemporal feature, determining a quality prediction score of each pixel in each of the first video frames corresponding to the spatiotemporal feature; The quality prediction score of each of the regions is determined according to the quality prediction scores of the pixels in each of the regions.
9. The method according to claim 8, wherein: The acquiring the spatiotemporal features of the target video includes: Preprocessing the target video to obtain preprocessed video data, wherein the amount of data of the preprocessed video data is smaller than that of the target video; The spatiotemporal features of the preprocessed video data are extracted to obtain the spatiotemporal features of the target video.
10. The method according to claim 9, wherein: The preprocessing of the target video to obtain preprocessed video data includes: Segmenting the video frames in the target video to obtain a plurality of grids; For each of the networks, a pixel block corresponding to the grid is selected; Dividing a plurality of continuous video frames in the target video to obtain a plurality of video time periods; Sampling is performed for each video time period to determine continuous sampling frames corresponding to the video time period; For each segment of the continuous sampling frames, determining a sampling segment corresponding to the pixel block; The sampled segments are spliced together to obtain the preprocessed video data.
11. The method according to claim 1, wherein: The determining the weight corresponding to the region based on the quality prediction score of the region includes: For the first video frame, calculating the sum of the quality prediction scores of the regions in the first video frame to obtain a sum of scores; For each of the regions in the first video frame, a weight corresponding to the region is determined according to a ratio of a quality prediction score of the region to the sum of the scores.
12. The method according to claim 1, wherein: Determining the quality score of the target video according to the quality prediction score of each region and the corresponding weight includes: For each of the first video frames, aggregating quality prediction scores of the respective regions according to a weight corresponding to each region in the first video frame to obtain a quality prediction score of the first video frame; The quality score of the target video is determined according to the quality prediction scores of the plurality of the first video frames in the target video.
13. The method according to claim 1, wherein: The determining the weight corresponding to the region based on the quality prediction score of the region includes: determining the semantic category to which the region belongs; A weight corresponding to the region is determined based on the quality prediction score of the region and the semantic category to which the region belongs.
14. A video quality assessment device, comprising: A video acquisition unit, configured to acquire a target video to be quality evaluated, wherein the target video includes a plurality of continuous video frames; A region segmentation unit is configured to perform region segmentation on a first video frame in the target video to determine a plurality of regions in the first video frame; a score acquisition unit, configured to acquire a quality prediction score of each of the regions, and determine a weight corresponding to the region based on the quality prediction score of the region, wherein the weight is positively correlated with the quality prediction score; The score determination unit is configured to determine the quality score of the target video according to the quality prediction score of each of the regions and the corresponding weight.
15. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 14 is implemented.