Comprehensive blind evaluation method for cartoon video quality

By establishing a multi-task self-supervised learning network model, combining space-time and line features, and using text-guided quality perception modules, the accuracy and generalization of cartoon video quality evaluation in the existing technology is solved, and efficient and reliable evaluation of cartoon video quality is achieved.

CN120164062APending Publication Date: 2025-06-17ANHUI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510122329.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing video quality evaluation algorithms are difficult to accurately evaluate cartoon video quality, especially in the absence of reference videos, and the evaluation method of natural video cannot be applied to cartoon videos.

Method used

A comprehensive blind evaluation method is adopted to obtain the spatiotemporal and line characteristics of the distorted cartoon video frame sequence, and combine the text-guided quality perception module to establish a multi-task self-supervised learning network model to train to obtain a comprehensive evaluation of video quality.

Benefits of technology

Accurate and reliable evaluation of cartoon video quality is achieved, the problem of insufficient labeled data sets is overcome, and the generalization of the model is improved, making it more in line with human visual perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164062A_ABST
    Figure CN120164062A_ABST
Patent Text Reader

Abstract

The invention discloses a comprehensive blind evaluation method for cartoon video quality, which comprises the following steps of: training a multi-task self-supervised learning network model through a distorted cartoon video frame sequence based on a time sequence to obtain the weight and bias of the multi-task self-supervised learning network model; taking the labeled cartoon video frame sequence as a network parameter of the cartoon video quality comprehensive blind evaluation model, and then performing fine adjustment on the cartoon video quality comprehensive blind evaluation model according to the labeled cartoon video frame sequence; and carrying out quality analysis on the to-be-evaluated cartoon video based on the fine-tuned cartoon video quality comprehensive blind evaluation model so as to evaluate the cartoon video needing to be evaluated. According to the method, the model is trained through the multi-task self-supervised learning network model based on self-supervised learning, the problem that an unnatural video mark data set is insufficient is solved, and the cartoon video quality is comprehensively evaluated in combination with cartoon content video characteristics, so that the cartoon video quality is improved. Therefore, the reliability of cartoon video quality evaluation and the generalization of cartoon videos are improved, and the cartoon video quality evaluation is more in line with human visual perception.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cartoon video quality assessment, and particularly to a comprehensive blind assessment method for cartoon video quality. Background Art

[0002] With the rapid development of the cartoon industry, the processing of cartoon images and videos has gradually received more attention, such as cartoon video interpolation, cartoon-style video generation, and early cartoon restoration. However, the lack of accurate video quality assessment algorithms (VQA) for cartoon content may have a negative impact on these related processing tasks. In addition, there are a large number of cartoons from various periods on the Internet. Due to the influence of production technology, storage, compression, and transmission, the quality of these cartoon contents is uneven. Therefore, cartoon video quality assessment is of great significance to both the industrial and academic fields.

[0003] VQA methods are generally divided into subjective VQA and objective VQA. Subjective VQA refers to that evaluators directly score the video quality. Although this method is simple and accurate, it is time-consuming and laborious, and is impractical for large-scale data requirements. Therefore, objective VQA has attracted the strong interest of many researchers. According to the degree of dependence on the reference video, objective VQA methods can be divided into full reference (FR), partial reference (RR), and no reference (NR). Due to the lack of reference videos in practical applications, NR-VQA has greater challenges and potential in the VQA field. Recently, learning-based methods that can automatically extract deep hidden features from videos have become the mainstream. However, for videos with non-natural content such as cartoons, there is still a lack of large-scale labeled VQA datasets for training, which brings great challenges to the development of learning-based NR-VQA methods for cartoon content videos.

[0004] Current VQA methods mainly focus on natural videos. However, due to the obvious differences between cartoons, which have clear lines, smooth color blocks, and flat backgrounds, and natural videos, directly applying natural VQA methods cannot accurately evaluate the quality of cartoon videos. In the field of image quality assessment (IQA), some researchers have proposed IQA methods for cartoon image quality assessment. Since IQA methods ignore the temporal information between frames, these cartoon IQA methods also have poor generalization in the cartoon video scenario. The accuracy of evaluating the quality of cartoon videos is low. In addition, it has been studied that other similar cartoon videos cannot be directly evaluated using existing natural VQA models, such as computer graphics (CG) videos and screen content (SC) videos. Summary of the Invention

[0005] The present invention discloses a comprehensive blind assessment method for cartoon video quality to overcome the above technical problems.

[0006] To achieve the above object, the technical solution of the present invention is as follows:

[0007] A comprehensive blind evaluation method for the quality of cartoon videos, comprising the following steps:

[0008] S1: Obtain distorted cartoon video frames from the cartoon video and the corresponding distorted cartoon video, and perform preprocessing on them. Based on the preprocessed distorted cartoon video frames, obtain a sequence of distorted cartoon video frames in chronological order;

[0009] S2: Establish a line perception spatio-temporal feature extraction module and a text-guided quality perception module;

[0010] The line perception spatio-temporal feature extraction module is used to obtain the spatio-temporal features of the sequence of distorted cartoon video frames and the line features of the sequence of distorted cartoon video frames according to the sequence of distorted cartoon video frames;

[0011] The text-guided quality perception module is used to obtain the aesthetic features of the sequence of distorted cartoon video frames and the texture features of the sequence of distorted cartoon video frames for the sequence of distorted cartoon video frames;

[0012] S3: According to the line perception spatio-temporal feature extraction module and the text-guided quality perception module, establish a multi-task self-supervised learning network model for cartoon videos, and based on the loss function of the multi-task self-supervised learning network model for cartoon videos and the sequence of distorted cartoon video frames, train the multi-task self-supervised learning network model for cartoon videos to obtain the weights and biases of the trained multi-task self-supervised learning network model for cartoon videos;

[0013] S4: Establish a comprehensive blind evaluation model for the quality of cartoon videos, and fine-tune the comprehensive blind evaluation model for the quality of cartoon videos according to the weights, biases of the trained multi-task self-supervised learning network model for cartoon videos, the loss function of the comprehensive blind evaluation model for the quality of cartoon videos, and the sequence of labeled cartoon video frames;

[0014] S5: According to the fine-tuned comprehensive blind evaluation model for the quality of cartoon videos, obtain the quality score of the cartoon video to be evaluated, so as to evaluate the cartoon video to be evaluated.

[0015] Further, the line perception spatio-temporal feature extraction module includes: a first backbone network, a first global average pooling layer, a Sobel convolutional layer, a batch normalization layer, an L2 normalization layer, a Sigmoid activation layer, a line extractor, and a second global average pooling layer;

[0016] The first backbone network is used to obtain the deep spatio-temporal features of the sequence of distorted cartoon video frames according to the sequence of distorted cartoon video frames;

[0017] The input end of the first global average pooling layer is connected to the output end of the backbone network, and is used to obtain the spatio-temporal features of the distorted cartoon video frame sequence according to the depth spatio-temporal features of the cartoon video;

[0018] The input end of the Sobel convolution layer is connected to the output end of the backbone network, and is used to obtain the line-sensitive spatio-temporal features of the distorted cartoon video frame sequence according to the distorted cartoon video frame sequence;

[0019] The input end of the batch normalization layer is connected to the output end of the Sobel convolution layer, and is used to obtain the normalized line-sensitive spatio-temporal features of the distorted cartoon video frame sequence according to the line-sensitive spatio-temporal features of the distorted cartoon video frame sequence;

[0020] The input end of the L2 normalization layer is connected to the output end of the batch normalization layer, and is used to obtain the scaled line-sensitive spatio-temporal features of the distorted cartoon video frame sequence according to the normalized line-sensitive spatio-temporal features of the distorted cartoon video frame sequence;

[0021] The input end of the Sigmoid activation layer is connected to the output end of the L2 normalization layer, and is used to obtain the line perception weight map according to the scaled line-sensitive spatio-temporal features of the distorted cartoon video frame sequence;

[0022] The input end of the multiplication module is respectively connected to the output end of the backbone network and the output end of the Sigmoid activation layer, and is used to obtain the line-sensitive spatio-temporal feature map according to the depth spatio-temporal features of the cartoon video and the line perception weight map;

[0023] The line extractor is connected to the output end of the multiplication module, and is used to obtain the line features of the distorted cartoon video frame sequence according to the line-sensitive spatio-temporal feature map;

[0024] The input end of the second global average pooling layer is connected to the output end of the line extractor, and is used to obtain the line features of the distorted cartoon video frame sequence according to the line features of the distorted cartoon video frame sequence.

[0025] Further, the line extractor includes a first convolutional layer, a second convolutional layer, a GELU activation layer, a third convolutional layer, and an addition module;

[0026] The input end of the first convolutional layer is connected to the output end of the multiplication module, and is used to obtain shallow line features according to the line-sensitive spatio-temporal feature map;

[0027] The input end of the second convolutional layer is connected to the output end of the first convolutional layer, and is used to obtain shallow quality-sensitive line features according to the shallow line features;

[0028] The input end of the GELU activation layer is connected to the output end of the second convolutional layer, and is used to obtain smooth shallow quality-sensitive line features according to the shallow quality-sensitive line features;

[0029] The input end of the third convolutional layer is connected to the output end of the GELU activation layer, and is used to obtain deep quality-sensitive line features according to the smooth shallow quality-sensitive line features;

[0030] The input ends of the addition module are respectively connected to the output end of the third convolutional layer and the output end of the first convolutional layer, and are used to obtain the line features of the distorted cartoon-like video frame sequence V according to the deep quality-sensitive line features and the shallow line features.

[0031] Further, the text-guided quality perception module includes a local binary layer, a downsampling layer, a linear layer, a backbone network image encoder, a backbone network text encoder, a cosine similarity calculation module, and a texture cross-attention mechanism module;

[0032] The local binary layer is used to obtain the texture information map of the distorted cartoon-like video frame sequence according to the distorted cartoon-like video frame sequence;

[0033] The downsampling layer is used to obtain the scaled distorted cartoon-like video frame sequence according to the distorted cartoon-like video frame sequence;

[0034] The input ends of the backbone network image encoder are respectively connected to the output ends of the local binary layer and the downsampling layer, and are used to obtain texture information image features and aesthetic information image features according to the texture information map of the distorted cartoon-like video frame sequence and the scaled distorted cartoon-like video frame sequence;

[0035] The backbone network text encoder is used to obtain texture text features and aesthetic text features according to the texture prompt words and aesthetic prompt words;

[0036] The input ends of the cosine similarity calculation module are respectively connected to the output ends of the backbone network image encoder and the backbone network text encoder, and are used to obtain frame-level texture features and aesthetic features of the distorted cartoon-like video frame sequence according to the texture information image features, aesthetic information image features, texture text features, and aesthetic text features;

[0037] The input end of the linear layer is connected to the output end of the first global average pooling layer, and is used to obtain dimension-reduced spatio-temporal features according to the spatio-temporal features of the distorted cartoon-like video frame sequence;

[0038] The input ends of the texture cross-attention mechanism module are respectively connected to the output ends of the cosine similarity calculation module and the linear layer, and are used to obtain the texture features of the distorted cartoon video frame sequence V according to the frame-level texture features and the dimensionality-reduced spatio-temporal features.

[0039] Further, the multi-task self-supervised learning network model for the cartoon video includes: a first series fusion module, a second series fusion module, a first prediction head module, a second prediction head module, and a third prediction head module;

[0040] The input ends of the first series fusion module are respectively connected to the output ends of the line-aware spatio-temporal feature extraction module and the text-guided quality perception module, and are used to obtain the comprehensive distortion features of the distorted cartoon video frame sequence according to the spatio-temporal features, line features, and texture features of the distorted cartoon video frame sequence.

[0041] The input end of the first prediction head module is connected to the output end of the first series fusion module, and is used to obtain the distortion type loss of the distorted cartoon video frame sequence according to the comprehensive distortion features of the distorted cartoon video frame sequence.

[0042] The input end of the second prediction head module is connected to the output end of the first series fusion module, and is used to obtain the distortion level loss of the distorted cartoon video frame sequence according to the comprehensive distortion features of the distorted cartoon video frame sequence.

[0043] The input end of the second series fusion module is connected to the output end of the text-guided quality perception module, and is used to obtain the comprehensive quality features of the distorted cartoon video frame sequence according to the aesthetic features and the comprehensive distortion features of the distorted cartoon video frame sequence.

[0044] The input end of the third prediction head module is connected to the output end of the second series fusion module, and is used to obtain the pairwise quality ranking loss of the distorted cartoon video frame sequence according to the comprehensive quality features of the distorted cartoon video frame sequence.

[0045] Further, the loss function of the multi-task self-supervised learning network model for the cartoon video includes a distortion type loss function, a distortion level loss function, and a pairwise quality ranking loss function;

[0046] The calculation formula of the loss function of the multi-task self-supervised learning network model for the cartoon video is as follows:

[0047] L total =L rank +α×L type +β×L level

[0048] Where: L total is the loss function of the multi-task self-supervised learning network model for cartoon videos, that is, the joint loss function of the distortion type loss, the distortion level loss, and the pairwise quality ranking loss; both α and β are weight coefficients; L rank represents the value of the pairwise quality ranking; L type represents the value of the distortion type loss of the cartoon video frame; L level represents the value of the distortion level loss of the cartoon video frame;

[0049] Among them,

[0050]

[0051] Among them, P represents the pre-training batch size; F CE is the cross-entropy loss function, T represents the predicted distortion type; T GT represents the true label of the distortion type; D represents the predicted distortion degree; D GT represents the true label of the distortion degree;

[0052]

[0053] j represents the quality of the distorted cartoon video frame sequence at the j-th distortion level; R i represents the quality of the distorted cartoon video frame sequence at the i-th distortion level; ε is a hyperparameter.

[0054] Furthermore, the comprehensive blind evaluation model for cartoon video quality includes: a line-aware spatio-temporal feature extraction module, a text-guided quality perception module, and a quality prediction module;

[0055] The input end of the quality prediction module is respectively connected to the output ends of the line-aware spatio-temporal feature extraction module and the text-guided quality perception module, and is used to obtain the quality score of the distorted cartoon video frame sequence according to the spatio-temporal feature F ST of the distorted cartoon video frame sequence V, the line feature of the distorted cartoon video frame sequence V, the aesthetic feature of the distorted cartoon video frame sequence V, and the texture feature of the distorted cartoon video frame sequence V.

[0056] Furthermore, the formula for the loss function of the comprehensive blind evaluation model for cartoon video quality is as follows:

[0057]

[0058] Where: B is the fine-tuning batch size; Q GT is the true quality score of the cartoon video;

[0059] PLCC(·) is the Pearson linear correlation coefficient; Q is the predicted quality score of the cartoon video; L Q is the value of the loss function of the comprehensive blind evaluation model for the quality of cartoon videos.

[0060] Beneficial effects: A comprehensive blind evaluation method for the quality of cartoon videos according to the present invention obtains a distorted cartoon video frame sequence based on time sequence, trains a multi-task self-supervised learning network model for cartoon videos of a line perception spatio-temporal feature extraction module and a text-guided quality perception module, obtains the weights and biases of the multi-task self-supervised learning network model for cartoon videos, and uses them as the initial network parameters of the comprehensive blind evaluation model for the quality of cartoon videos. Then, according to the labeled cartoon video frame sequence, the comprehensive blind evaluation model for the quality of cartoon videos is fine-tuned; based on the fine-tuned comprehensive blind evaluation model for the quality of cartoon videos, the quality analysis of the cartoon video to be evaluated is obtained, and the evaluation of the cartoon video to be evaluated is realized. The present invention trains the model through a multi-task self-supervised learning network model based on self-supervised learning, overcomes the problem of insufficient non-natural video labeled data sets, and comprehensively evaluates the quality of cartoon videos in combination with the characteristics of cartoon content videos, thereby improving the reliability of cartoon video quality evaluation and the generalization ability in cartoon videos, making the evaluation of cartoon video quality more in line with human visual perception. Description of the Drawings

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0062] Figure 1 is the flow chart of the comprehensive blind evaluation method for the quality of cartoon videos of the present invention;

[0063] Figure 2 is the schematic structural diagram of the comprehensive blind evaluation model for the quality of cartoon videos in the embodiments of the present invention;

[0064] Figure 3 is the structural block diagram of the multi-task self-supervised learning network model for cartoon videos in the embodiments of the present invention;

[0065] Figure 4 is the structural block diagram of the line perception spatio-temporal feature extraction module in the embodiments of the present invention;

[0066] Figure 5 is the structural block diagram of the text-guided quality perception module in the embodiments of the present invention;

[0067] Figure 6 The quality evaluation effect diagram of the cartoon video in the embodiment of the present invention;

[0068] Figure 7 The quality evaluation effect diagram of the computer graphics video in the embodiment of the present invention;

[0069] Figure 8 The quality evaluation effect diagram of the screen content video in the embodiment of the present invention;

[0070] Figure 9 The sequence diagram of the evaluation method in the embodiment of the present invention. Detailed implementation manners

[0071] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0072] This embodiment introduces a comprehensive blind evaluation method for the quality of cartoon videos, as Figure 1 and Figure 9 shown, including the following steps:

[0073] S1: According to the cartoon video and the distorted cartoon video corresponding to the cartoon video, obtain the distorted cartoon video frames and perform preprocessing, so as to obtain a sequence of distorted cartoon video frames based on the time sequence according to the preprocessed distorted cartoon video frames;

[0074] Specifically, in this embodiment, through high-quality original cartoon videos, multiple distorted cartoon videos with different distortion types and degrees are obtained through conventional distortion enhancement strategies, and then the distorted cartoon video frames are obtained, so as to expand the data volume in the dataset and perform preprocessing on the distorted cartoon video frames. Among them, obtaining video frames from a video is a conventional operation in the field and will not be described in detail here. The distortion enhancement strategy is a means of expanding the dataset of self-supervised learning technology. The dataset generated by distortion enhancement has clear labels such as distortion types and degrees, which can facilitate multi-task training in self-supervised learning.

[0075] Specifically, the preprocessing method for the distorted cartoon video frames is as follows:

[0076] S11: Obtain a distorted cartoon video corresponding to the cartoon video through a distortion enhancement strategy; that is, artificially synthesize distortions on the original cartoon video to generate corresponding distorted cartoon videos of different types and degrees.

[0077] S12: Uniformly extract frames from the distorted cartoon video; obtain multiple initial distorted cartoon video frames.

[0078] S13: Crop each of the multiple initial distorted cartoon video frames to obtain multiple distorted cartoon video frames with the same resolution.

[0079] Specifically, crop the nth initial distorted cartoon video frame I n , n = (1, 2,.., N). The resolution of the cropped distorted cartoon video frame is H×W. Then the total size of all the cropped distorted cartoon video frames is N×3×H×W, where n is the index of the distorted cartoon video frame; N is the total number of distorted cartoon video frames; H is the height of the distorted cartoon video frame; W is the width of the distorted cartoon video frame; 3 represents the three RGB channels.

[0080] S2: Establish a line-aware spatio-temporal feature extraction module and a text-guided quality perception module.

[0081] The line-aware spatio-temporal feature extraction module is used to obtain the spatio-temporal feature F ST of the distorted cartoon video frame sequence V and the line feature of the distorted cartoon video frame sequence V.

[0082] The text-guided quality perception module is used to obtain the aesthetic feature and the texture feature of the distorted cartoon video frame sequence V for the distorted cartoon video frame sequence.

[0083] Preferably, the line-aware spatio-temporal feature extraction module includes: a first backbone network, a first global average pooling layer, a Sobel convolutional layer, a batch normalization layer, an L2 normalization layer, a Sigmoid activation layer, a line extractor, and a second global average pooling layer.

[0084] The first backbone network is used to obtain the deep spatio-temporal feature F VST ;

[0085] The input end of the first global average pooling layer is connected to the output end of the backbone network, and is used to obtain the spatio-temporal feature F VST of the distorted cartoon video frame sequence V according to the deep spatio-temporal feature F ST ;

[0086] Specifically, when inputting a cartoon video with a lot of redundant frames and high resolution, it will cause the model training to be very slow and consume a lot of computer resources. Therefore, in this embodiment, the input cartoon video is preprocessed, that is, frames are evenly sampled and clipped, and the extracted frames form a video frame sequence to replace the input cartoon video. The overall spatio-temporal feature here is actually the overall of all the extracted frames, but represents the overall spatio-temporal feature of the input cartoon.

[0087] The input end of the Sobel convolutional layer is connected to the output end of the backbone network, and is used to obtain the line-sensitive spatio-temporal feature of the distorted cartoon video frame sequence V according to the distorted cartoon video frame sequence.

[0088] The input end of the batch normalization layer is connected to the output end of the Sobel convolutional layer, and is used to obtain the normalized line-sensitive spatio-temporal feature of the distorted cartoon video frame sequence V according to the line-sensitive spatio-temporal feature of the distorted cartoon video frame sequence V.

[0089] The input end of the L2 normalization layer is connected to the output end of the batch normalization layer, and is used to obtain the scaled line-sensitive spatio-temporal feature of the distorted cartoon video frame sequence V according to the normalized line-sensitive spatio-temporal feature of the distorted cartoon video frame sequence V.

[0090] The input end of the Sigmoid activation layer is connected to the output end of the L2 normalization layer, and is used to obtain the line perception weight map according to the scaled line-sensitive spatio-temporal feature of the distorted cartoon video frame sequence V.

[0091] The input end of the multiplication module is respectively connected to the output end of the backbone network and the output end of the Sigmoid activation layer, and is used to obtain the line-sensitive spatio-temporal feature map according to the deep spatio-temporal feature F of the cartoon video VST and the line perception weight map.

[0092] The line extractor is connected to the output end of the multiplication module, and is used to obtain the line feature of the distorted cartoon video frame sequence V according to the line-sensitive spatio-temporal feature map.

[0093] The input end of the second global average pooling layer is connected to the output end of the line extractor, and is used to obtain the line feature of the distorted cartoon video frame sequence V according to the line feature of the distorted cartoon video frame sequence V.

[0094] Preferably, the line extractor includes a first convolutional layer, a second convolutional layer, a GELU activation layer, a third convolutional layer, and an addition module;

[0095] The input end of the first convolutional layer is connected to the output end of the multiplication module, and is used to obtain the shallow line feature according to the line-sensitive spatio-temporal feature map.

[0096] The input end of the second convolutional layer is connected to the output end of the first convolutional layer, and is used to obtain shallow quality-sensitive line features according to the shallow line features;

[0097] The input end of the GELU activation layer is connected to the output end of the second convolutional layer, and is used to obtain smooth shallow quality-sensitive line features according to the shallow quality-sensitive line features;

[0098] The input end of the third convolutional layer is connected to the output end of the GELU activation layer, and is used to obtain deep quality-sensitive line features according to the smooth shallow quality-sensitive line features;

[0099] The input ends of the addition module are respectively connected to the output end of the third convolutional layer and the output end of the first convolutional layer, and are used to obtain the line features of the distorted cartoon video frame sequence V according to the deep quality-sensitive line features and the shallow line features.

[0100] Specifically, the video contains rich static spatial information and dynamic temporal information. Effectively extracting the spatio-temporal features of cartoon video frames is a key prerequisite for processing video-related tasks. At the same time, sharp and clear lines are one of the main components of cartoon videos. Since the human visual system is highly sensitive to edge line information, the line quality in cartoon videos plays a crucial role in the quality assessment of cartoon videos. Therefore, in this embodiment, a line-aware spatio-temporal feature extraction module is established to specifically capture the spatio-temporal features of cartoon videos and represent information related to line quality, such as Figure 4 shown.

[0101] Specifically, in this embodiment, the Video Swin Transformer-Tiny network is used as the first backbone network, and the network is used in four stages to extract the deep spatio-temporal features F of the input distorted cartoon video frame sequence V VST , and then the global average pooling layer GAP is used to obtain the overall spatio-temporal feature F ST , and the formula used is as follows:

[0102] F ST = GAP(f vst (V))

[0103] where, f vst represents the Video Swin Transformer-Tiny network; F ST represents the overall spatio-temporal network obtained by the global average pooling layer; GAP(·) represents the global average pooling operation;

[0104] Specifically, in this embodiment, a learnable Sobel convolutional layer is also set to enhance the depth spatio-temporal feature F of the distorted cartoon-like video frame sequence V VST of the line information. First, the Sobel convolutional layer is used to determine the sharp line pixels of the input feature map, and then batch normalization BN and L2 normalization are performed through the batch normalization layer and the L2 normalization layer in sequence. Finally, the Sigmoid activation function in the Sigmoid activation layer is used to obtain the line perception weight map, and the output features of the backbone network are re-weighted element-wise to obtain the line-sensitive spatio-temporal feature map F l , and the specific formula is as follows:

[0105] F l = Sobel(F VST )

[0106] where Sobel(·) is the Sobel convolutional layer; F l represents the line-sensitive spatio-temporal feature map; F VST represents the depth spatio-temporal feature of the distorted cartoon-like video frame sequence V; V represents the distorted cartoon-like video frame sequence;

[0107] Specifically, in this embodiment, a line extractor is also designed to further process and extract the depth line quality-sensitive features, and then the second global average pooling layer is used to obtain the overall line quality feature F L , and the specific formula is as follows:

[0108] F L = GAP(E line (F l ))

[0109] where E line (·) is the line extractor, which first extracts the line information through a 1×1 convolutional layer, then captures the deep line quality-related information through two 3×3 convolutional layers, applies GELU activation in the middle, and finally introduces a residual connection operation.

[0110] Preferably, the text-guided quality perception module includes a local binary layer, a downsampling layer, a linear layer, a backbone network image encoder, a backbone network text encoder, a cosine similarity calculation module, and a texture cross-attention mechanism module;

[0111] The local binary layer is used to obtain the texture information map of the distorted cartoon-like video frame sequence according to the distorted cartoon-like video frame sequence;

[0112] The downsampling layer is used to obtain the scaled distorted cartoon-like video frame sequence according to the distorted cartoon-like video frame sequence;

[0113] The input end of the backbone network image encoder is respectively connected to the output ends of the local binary layer and the downsampling layer, and is used to obtain texture information image features and aesthetic information image features according to the texture information map of the distorted cartoon-like video frame sequence and the scaled distorted cartoon-like video frame sequence;

[0114] The backbone network text encoder is used to obtain texture text features and aesthetic text features according to texture prompts and aesthetic prompts;

[0115] The input end of the cosine similarity calculation module is respectively connected to the output ends of the backbone network image encoder and the backbone network text encoder, and is used to obtain frame-level texture features and aesthetic features of the distorted cartoon-like video frame sequence V according to the texture information image features, aesthetic information image features, texture text features and aesthetic text features;

[0116] The input end of the linear layer is connected to the output end of the first global average pooling layer, and is used to obtain a dimensionality-reduced spatio-temporal feature according to the spatio-temporal feature F of the distorted cartoon-like video frame sequence V ST , to obtain a dimensionality-reduced spatio-temporal feature;

[0117] The input ends of the texture cross-attention mechanism module are respectively connected to the output ends of the cosine similarity calculation module and the linear layer, and are used to obtain the texture features of the distorted cartoon-like video frame sequence V according to the frame-level texture features and the dimensionality-reduced spatio-temporal features.

[0118] Specifically, cartoon-like videos lack complex natural textures. Therefore, analyzing the texture features of cartoon-like videos is beneficial to perceiving pseudo-textures caused by distortions such as compression or noise. Moreover, the aesthetic quality of cartoon content and artistic style has a huge impact on its overall quality. Through research, CLIP, as a basic cross-modal model, has shown excellent performance in many visual tasks. It can guide the extraction of subjective and fine-grained quality information from images and videos through specially designed text prompts. Therefore, in this embodiment, a text-guided quality perception module is established to utilize the powerful function of the pre-trained CLIP to characterize the aesthetic and texture quality of cartoon-like videos by designing appropriate prompts for aesthetic and texture features, as Figure 5 shown.

[0119] Specifically, in this embodiment, through the local binary descriptor LBP of local binary and downsampling operations, each frame I of the input cartoon-like video n is respectively converted into a texture image and a downsampled image The downsampling operation reduces the impact of distortion on the extraction of aesthetic features, and through preprocessing with LBP, rich texture information can be obtained;

[0120] Specifically, the texture prompt and aesthetic prompt in this embodiment are manually set by those skilled in the art according to specific requirements, and are used to guide the extraction of texture features and the extraction of aesthetic features from both content and artistic style aspects in sequence. The specific format of the text is as follows:

[0121] Texture text: This is a <ctx>photo with <init>texture.

[0122] Content text: This is a <ctx>photo with <init>content.

[0123] Stylistic text: This is a(an) <init>aesthetic <ctx>photo.

[0124] Among them, <init>Is the descriptive term filled in correspondingly, <ctx>It is a text marker, initialized with the letter X and optimized during training. The specific description words are presented in Table 1;

[0125]

[0126] Specifically, the text-guided quality-aware module network in this embodiment uses the multi-modal large model CLIP as the backbone network. and are input into the CLIP image encoder, and the text prompt is input into the CLIP text encoder. By means of CLIP, the matching value between the image and the corresponding prompt is calculated, and then the Sigmoid activation function is used to normalize to obtain the texture and aesthetic features of each input frame. Finally, the features of each frame are concatenated and fused to form the texture feature F CT and aesthetic feature F A of the input video. The specific formula is as follows:

[0127]

[0128] where, is the texture feature of the video frame obtained by the image and the texture text prompt through CLIP; is the aesthetic feature of each input frame calculated by the image and the two groups of text prompts of content and style through CLIP; n is the index of the video frame; N is the total number of video frames;

[0129] Specifically, in this embodiment, a texture cross-attention mechanism is set to capture the temporal texture information between frames. The spatio-temporal feature F ST is mapped into F CT of the same dimension through a linear layer, and then used as the key K ST and value V ST , and F CT is used as the query Q T . Finally, the attention mechanism is used to obtain the texture feature F T of the input cartoon class. The specific formula is as follows:

[0130] Q T =F CT W Q ,K ST =FC(F ST )W K ,V ST =FC(F ST )W V

[0131]

[0132] where, W Q 、W Q and W Q represent the corresponding parameter matrices, FC(·) and C represent the linear layer and the feature dimension of K ST respectively; Softmax represents the Softmax activation function;

[0133] S3: According to the line-aware spatio-temporal feature extraction module and the text-guided quality perception module, establish a multi-task self-supervised learning network model for cartoon videos, and based on the loss function of the multi-task self-supervised learning network model for cartoon videos and the distorted cartoon video frame sequence, train the multi-task self-supervised learning network model for cartoon videos to obtain the weights and biases of the trained multi-task self-supervised learning network model for cartoon videos;

[0134] Preferably, the multi-task self-supervised learning network model for cartoon videos includes: a first series fusion module, a second series fusion module, a first prediction head module, a second prediction head module, and a third prediction head module;

[0135] The input end of the first series fusion module is respectively connected to the output ends of the line-aware spatio-temporal feature extraction module and the text-guided quality perception module, and is used to obtain the comprehensive distortion feature of the distorted cartoon video frame sequence according to the spatio-temporal feature F ST of the distorted cartoon video frame sequence V, the line feature of the distorted cartoon video frame sequence V, and the texture feature of the distorted cartoon video frame sequence V;

[0136] The input end of the first prediction head module is connected to the output end of the first series fusion module, and is used to obtain the distortion type loss of the distorted cartoon video frame sequence V according to the comprehensive distortion feature of the distorted cartoon video frame sequence;

[0137] The input end of the second prediction head module is connected to the output end of the first series fusion module, and is used to obtain the distortion level loss of the distorted cartoon video frame sequence V according to the comprehensive distortion feature of the distorted cartoon video frame sequence;

[0138] The input end of the second series fusion module is connected to the output end of the text-guided quality perception module, and is used to obtain the comprehensive quality feature of the distorted cartoon video frame sequence according to the aesthetic feature of the distorted cartoon video frame sequence V and the comprehensive distortion feature of the distorted cartoon video frame sequence;

[0139] The input end of the third prediction head module is connected to the output end of the second series fusion module, and is used to obtain the pairwise quality ranking loss of the distorted cartoon video frame sequence V according to the comprehensive quality feature of the distorted cartoon video frame sequence.

[0140] Preferably, the loss function of the multi-task self-supervised learning network model for the cartoon-like video includes a distortion type loss function, a distortion level loss function, and a pairwise quality ranking loss function;

[0141] The calculation formula of the loss function of the multi-task self-supervised learning network model for the cartoon-like video is as follows:

[0142] Specifically, in this embodiment, the weighted sum of the three auxiliary task losses is used as the joint loss function L total for self-supervised pre-training, and the specific formula is as follows:

[0143] L total = L rank + α × L type + β × L level

[0144] In the formula: L total is the loss function of the multi-task self-supervised learning network model for the cartoon-like video, that is, the joint loss function of the distortion type loss, the distortion level loss, and the pairwise quality ranking loss; α and β are both weight coefficients, which are positive values and are used as weight parameters to adjust the relative importance of the three components in the joint loss function; in this example, the values of α and β are both 0.6; L rank represents the value of the pairwise quality ranking; L type represents the value of the distortion type loss of the cartoon-like video frame; L level represents the value of the distortion level loss of the cartoon-like video frame;

[0145] Among them,

[0146]

[0147]

[0148] Among them, P represents the pre-training batch size; F CE is the cross-entropy loss function, T represents the predicted distortion type; T GT represents the true label of the distortion type; D represents the predicted distortion degree; D GT represents the true label of the distortion degree;

[0149] Specifically, for the pairwise quality ranking auxiliary task, this embodiment introduces a ranking loss L rank to optimize the model for pairwise quality ranking, and the specific formula is as follows:

[0150]

[0151] Among them, i and j both represent the distortion level indices, i ≠ j; L represents the total number of distortion degrees; R j denotes the quality of the distorted cartoon - like video frame sequence at the j - th predicted distortion level; R i denotes (the quality of the distorted cartoon - like video frame sequence at the i - th distortion level; ε is a hyper - parameter; in this example, its value is 0.05;

[0152] Specifically, input the distorted cartoon - like video frame sequence into the multi - task self - supervised learning network model of the cartoon - like video. Based on the loss function of the multi - task self - supervised learning network model of the cartoon - like video, which includes a distortion - type loss function, a distortion - level loss function, and a pairwise quality ranking loss function, train the multi - task self - supervised learning network model of the cartoon - like video. Among them, through the distortion - type loss function and the distortion - level loss function of the cartoon - like video frames, perform self - supervised pre - training on the model. For the prediction - assisted tasks of the distortion type and the distortion level, use the cross - entropy loss as the loss function \(L\) for predicting the distortion type and the degree of distortion type 、\(L\) level ;

[0153] Specifically, in this embodiment, perform self - supervised pre - training on the multi - task self - supervised learning network model of the cartoon - like video. Use the backpropagation method to calculate the gradients of the parameters in the model, and use the stochastic gradient descent method to update the parameters in the model. Save the parameters of the final multi - task self - supervised learning network model of the cartoon - like video for extracting spatio - temporal, line, aesthetic, and texture features.

[0154] Specifically, since the typical characteristics of cartoons are clear lines, smooth color blocks, and lack of natural texture. Cartoon - like videos have this inherent prior knowledge, which is different from natural videos with diverse content and complex textures. In addition, as an art form, the aesthetic quality of the content and art style of cartoons is also an important factor affecting the overall quality of cartoon - like videos. Therefore, this embodiment constructs a comprehensive blind evaluation model for the quality of cartoon - like videos. At the same time, for the research on non - natural video quality assessment methods, such as cartoons, computer graphics, and screen content, there is still a lack of large - scale labeled samples to train deep - learning models, which will lead to model over - fitting and reduce the generalization of the model. Therefore, this embodiment adopts a self - supervised learning strategy to use pseudo - labels generated from large - scale unlabeled data for model pre - training to alleviate this problem, and formulates different auxiliary tasks according to different purposes to build a multi - task self - supervised framework, as Figure 3 shown.

[0155] S4: Establish a comprehensive blind evaluation model for the quality of cartoon videos, and fine-tune the comprehensive blind evaluation model for the quality of cartoon videos according to the weights and biases of the multi-task self-supervised learning network model of the trained cartoon videos, the loss function of the comprehensive blind evaluation model for the quality of cartoon videos, and the labeled sequence of cartoon video frames; among them, the labeled sequence of cartoon video frames is a video for which the scores of the distorted cartoon video frame sequences have been marked.

[0156] Preferably, the comprehensive blind evaluation model for the quality of cartoon videos includes: a line-aware spatio-temporal feature extraction module, a text-guided quality perception module, and a quality prediction module, as Figure 2 shown;

[0157] The input end of the quality prediction module is respectively connected to the output ends of the line-aware spatio-temporal feature extraction module and the text-guided quality perception module, and is used to obtain the quality score of the distorted cartoon video frame sequence according to the spatio-temporal feature F ST of the distorted cartoon video frame sequence V, the line feature of the distorted cartoon video frame sequence V, the aesthetic feature of the distorted cartoon video frame sequence V, and the texture feature of the distorted cartoon video frame sequence V.

[0158] Preferably, the expression formula of the loss function of the comprehensive blind evaluation model for the quality of cartoon videos is as follows:

[0159]

[0160] In the formula: B is the fine-tuning batch size; Q GT is the true quality score of the cartoon video;

[0161] PLCC(·) is the Pearson linear correlation coefficient; Q is the predicted quality score of the cartoon video; L Q is the value of the loss function of the comprehensive blind evaluation model for the quality of cartoon videos.

[0162] Specifically, in this embodiment, the preprocessed sequence of cartoon video frames with true quality labels is input into the comprehensive blind evaluation model for the quality of cartoon videos to obtain the predicted quality score Q of the cartoon video;

[0163] Specifically, in this embodiment, for the comprehensive blind evaluation model for the quality of cartoon videos, fine-tuning is performed on the labeled data set. The weights and biases of the trained multi-task self-supervised learning network model of the cartoon videos are used as the initial parameters of the comprehensive blind evaluation model for the quality of cartoon videos, and the backpropagation method is used to calculate the gradients of the parameters in the model, and the stochastic gradient descent method is used to update the parameters in the model. The model parameters when the PLCC index is the best are saved, and the comprehensive blind evaluation model for the quality of cartoon videos is fine-tuned to evaluate the quality of cartoon videos.

[0164] S5: Obtain the quality score of the cartoon video to be evaluated according to the fine-tuned comprehensive blind evaluation model of cartoon video quality, so as to evaluate the cartoon video to be evaluated.

[0165] Specifically, in this embodiment, a line-aware spatio-temporal feature extraction module and a text-guided quality perception module are established, and based on them, a multi-task self-supervised learning network model for cartoon videos and a comprehensive blind evaluation model for cartoon video quality are established;

[0166] Specifically, in this embodiment, by setting the use of the line-aware spatio-temporal feature extraction module and the text-guided quality perception module, the spatio-temporal feature F of the input cartoon video is extracted ST , line feature F L , texture feature F T and aesthetic feature F A are concatenated and fused to obtain the final feature F representing the quality of the cartoon video Q , and the formula is as follows:

[0167] F Q = Concat(F ST , F L , F T , F A )

[0168] Then map F Q to the quality score Score of the input cartoon video, and the specific formula is as follows:

[0169] Score = f out (F Q )

[0170] where f out represents the quality prediction module, which first deeply extracts the quality-sensitive features of the cartoon video through a linear layer, and then processes them through the GELU activation function and predicts the quality score of the cartoon video through a linear layer.

[0171] Specifically, in this embodiment, by establishing a multi-task self-supervised learning network model for cartoon videos, a multi-task self-supervised learning framework for cartoon videos is formed: using the line-aware spatio-temporal feature extraction module and the text-guided quality perception module, the distortion-sensitive feature F of the input cartoon video is extracted D , ranking quality feature F R , and the formula forms are the same, and the specific representation is as follows:

[0172] F D = Concat(F ST , F L , F T )

[0173] Specifically, the prediction head structure M in this embodiment pred is used for the auxiliary task. The prediction head structure M pred passes through a linear layer, then is processed by a GELU activation function, and then passes through another linear layer;

[0174] For the auxiliary tasks of self-supervised pre-training for distortion aspects, two tasks of distortion type prediction and distortion level prediction are used: F D respectively through the prediction head structure and to predict the probability T of the distortion type and the distortion degree D. The formula is as follows:

[0175]

[0176] wherein, represents the first prediction head module for predicting the distortion type; represents the second prediction head module for predicting the distortion level;

[0177] For the auxiliary tasks of self-supervised pre-training for comprehensive quality aspects, a pairwise quality ranking task is used, and the ranking quality feature F R calculates the ranking quality R through the third prediction head module for predicting the pairwise quality ranking loss. The formula is as follows:

[0178]

[0179] Specifically, the cartoon video is input into the pre-trained comprehensive blind evaluation model for cartoon video quality, the model is fine-tuned, and the model parameters are saved for the evaluation of cartoon video quality. The cartoon video to be evaluated is input into the fine-tuned comprehensive blind evaluation model for cartoon video quality to obtain its quality score.

[0180] Specifically, the multi-task self-supervised learning framework for cartoon videos in this embodiment includes three auxiliary tasks, namely pairwise quality ranking, distortion type prediction, and distortion level prediction. Because compared with single-task learning, multi-task learning can enhance the feature representation and generalization ability of the model by sharing features and knowledge among multiple tasks. Compared with directly evaluating the quality score of a single video, the human visual system (HVS) is more inclined to evaluate the relative quality of two videos. Therefore, using the quality ranking task to encourage the model to rank the video quality helps to learn quality-sensitive features. In addition, since different distortion types and levels in cartoon videos will significantly affect the perceived quality, the distortion type and level prediction tasks enable the model to focus on distortion information, thereby learning and understanding the quality changes of cartoon videos.

[0181] Such as Figure 6 As shown, it shows a sequence of cartoon video frames and the evaluated quality scores. Clearly, the quality on the right is lower than that on the left, and the higher the score, the better the quality. From Figure 6 it can be seen that the comprehensive blind evaluation method for the quality of cartoon-like videos in this embodiment can generate quality scores consistent with human visual perception. In addition, in Figure 7 and Figure 8 , the proposed model is shown to evaluate the quality of computer graphics videos and screen content videos respectively. The comprehensive blind evaluation model for the quality of cartoon-like videos in this embodiment can also predict reasonable quality scores for these videos with visual characteristics similar to cartoons.

[0182] Specifically, in the line perception spatio-temporal feature extraction module, text-guided quality perception module, multi-task self-supervised learning network model for cartoon-like videos, and comprehensive blind evaluation model for the quality of cartoon-like videos built in this embodiment, each module used is an existing technology in the field. This embodiment only uses them to achieve the functions to be achieved in this embodiment.

[0183] Specifically, by considering that compared with natural videos, cartoon-like videos have fewer complex local textures and clear lines, and the important influence of the aesthetic perception of cartoons as an art form on video quality, this embodiment designs a multi-angle evaluation model for the quality of cartoon-like videos, which comprehensively evaluates the quality of cartoon-like videos from distortion, line quality, texture quality, and aesthetics.

[0184] This embodiment also explores the application of multi-modal large models in the cartoon field, and effectively guides the cross-modal large model to extract the texture and aesthetic features of cartoon-like videos by designing a series of text prompts.

[0185] This embodiment adopts a multi-task self-supervised learning strategy to pre-train the designed model, overcomes the problem of insufficient labeled samples for cartoon-like videos, enables the model to learn distortion-sensitive features from spatio-temporal, line, and texture through distortion type and basic distortion prediction tasks, and formulates a paired quality ranking task to let the model initially comprehensively learn information related to the quality of cartoon-like videos.

[0186] After the multi-task self-supervised learning network model for the cartoon-like videos in this embodiment completes self-supervised pre-training, the comprehensive blind evaluation model for the quality of cartoon-like videos is fine-tuned by using its network parameters as the initial network parameters of the comprehensive blind evaluation model for the quality of cartoon-like videos. It can not only enable the fine-tuned comprehensive blind evaluation model for the quality of cartoon-like videos to be used to effectively evaluate the quality of cartoon videos, but also be used for the quality evaluation of computer graphics videos and screen content videos, improving the accuracy and robustness of the model for evaluating the quality of cartoon-like content videos and being more in line with human visual perception.

[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.< / ctx> < / init> < / ctx> < / init> < / init> < / ctx> < / init> < / ctx>

Claims

1. A comprehensive blind evaluation method for cartoon video quality, characterized in that: The steps include: S1: According to a cartoon video and a distorted cartoon video corresponding to the cartoon video, a distorted cartoon video frame is obtained and preprocessed, so as to obtain a distorted cartoon video frame sequence based on a time sequence according to the preprocessed distorted cartoon video frame; S2: Establishing line-aware spatiotemporal feature extraction module and text-guided quality perception module; The line-aware spatiotemporal feature extraction module is used to obtain the spatiotemporal features of the distorted cartoon video frame sequence and the line features of the distorted cartoon video frame sequence according to the distorted cartoon video frame sequence; The text-guided quality perception module is used for the distorted cartoon video frame sequence to obtain aesthetic features of the distorted cartoon video frame sequence and texture features of the distorted cartoon video frame sequence; S3: establishing a multi-task self-supervised learning network model for cartoon videos according to the line-aware spatiotemporal feature extraction module and the text-guided quality perception module, and training the multi-task self-supervised learning network model for cartoon videos based on the loss function of the multi-task self-supervised learning network model for cartoon videos and the distorted cartoon video frame sequence, so as to obtain weights and biases of the trained multi-task self-supervised learning network model for cartoon videos; S4: establishing a comprehensive blind assessment model for cartoon video quality, and fine-tuning the comprehensive blind assessment model for cartoon video quality according to the weights and biases of the trained multi-task self-supervised learning network model for cartoon videos, the loss function of the comprehensive blind assessment model for cartoon video quality, and a labeled cartoon video frame sequence; S5: According to the fine-tuned cartoon video quality comprehensive blind evaluation model, the quality score of the cartoon video to be evaluated is obtained, so as to evaluate the cartoon video to be evaluated.

2. A comprehensive blind evaluation method for cartoon video quality according to claim 1, characterized in that: The line-aware spatiotemporal feature extraction module includes: a first backbone network, a first global average pooling layer, a Sobel convolution layer, a batch normalization layer, an L2 normalization layer, a Sigmoid activation layer, a line extractor, and a second global average pooling layer; The first backbone network is used to obtain deep spatiotemporal features of the distorted cartoon-like video frame sequence according to the distorted cartoon-like video frame sequence; The input end of the first global average pooling layer is connected to the output end of the backbone network, and is used to obtain the spatiotemporal features of the distorted cartoon video frame sequence according to the deep spatiotemporal features of the cartoon video; The input end of the Sobel convolution layer is connected to the output end of the backbone network, and is used to obtain the line-sensitive spatiotemporal features of the distorted cartoon video frame sequence according to the distorted cartoon video frame sequence; The input end of the batch normalization layer is connected to the output end of the Sobel convolution layer, and is used to obtain the normalized line-sensitive spatiotemporal features of the distorted cartoon-like video frame sequence according to the line-sensitive spatiotemporal features of the distorted cartoon-like video frame sequence; The input end of the L2 normalization layer is connected to the output end of the batch normalization layer, and is used to obtain the scaled line-sensitive spatiotemporal features of the distorted cartoon-like video frame sequence according to the normalized line-sensitive spatiotemporal features of the distorted cartoon-like video frame sequence; The input end of the Sigmoid activation layer is connected to the output end of the L2 normalization layer, and is used to obtain a line perception weight map according to the scaled line sensitive spatiotemporal features of the distorted cartoon video frame sequence; The input end of the multiplication module is connected to the output end of the backbone network and the output end of the Sigmoid activation layer respectively, and is used to obtain a line-sensitive spatiotemporal feature map according to the deep spatiotemporal features of the cartoon video and the line-perception weight map; The line extractor is connected to the output end of the multiplication module, and is used to obtain the line features of the distorted cartoon-like video frame sequence according to the line-sensitive spatiotemporal feature map; The input end of the second global average pooling layer is connected to the output end of the line extractor, and is used to obtain the line features of the distorted cartoon video frame sequence according to the line features of the distorted cartoon video frame sequence.

3. A comprehensive blind evaluation method for cartoon video quality according to claim 2, characterized in that: The line extractor includes a first convolutional layer, a second convolutional layer, a GELU activation layer, a third convolutional layer, and an addition module; The input end of the first convolutional layer is connected to the output end of the multiplication module, and shallow line features are obtained according to the line-sensitive spatiotemporal feature map; The input end of the second convolutional layer is connected to the output end of the first convolutional layer, and is used to obtain shallow quality-sensitive line features according to shallow line features; The input end of the GELU activation layer is connected to the output end of the second convolutional layer, and is used to obtain a smooth shallow quality-sensitive line feature according to the shallow quality-sensitive line feature; The input end of the third convolutional layer is connected to the output end of the GELU activation layer, and is used to obtain the deep quality-sensitive line features according to the smoothed shallow quality-sensitive line features; The input end of the adding module is connected to the output end of the third convolutional layer and the output end of the first convolutional layer respectively, and is used to obtain the line features of the distorted cartoon video frame sequence V according to the deep quality-sensitive line features and the shallow line features.

4. A comprehensive blind evaluation method for cartoon video quality according to claim 1, characterized in that: The text-guided quality perception module includes a local binary layer, a downsampling layer, a linear layer, a backbone network image encoder, a backbone network text encoder, a cosine similarity calculation module, and a texture cross-attention mechanism module; The local binary layer is used to obtain a distorted cartoon-like video frame sequence texture information map according to the distorted cartoon-like video frame sequence; The downsampling layer is used to obtain a scaled distorted cartoon video frame sequence according to the distorted cartoon video frame sequence; The input end of the backbone network image encoder is connected to the output ends of the local binary layer and the downsampling layer respectively, and is used to obtain texture information image features and aesthetic information image features according to the distorted cartoon video frame sequence texture information map and the scaled distorted cartoon video frame sequence; The backbone network text encoder is used to obtain texture text features and aesthetic text features according to texture prompt words and aesthetic prompt words; The input end of the cosine similarity calculation module is connected to the output end of the backbone network image encoder and the output end of the backbone network text encoder respectively, and is used to obtain frame-level texture features and aesthetic features of distorted cartoon video frame sequences according to texture information image features, aesthetic information image features, texture text features and aesthetic text features; The input end of the linear layer is connected to the output end of the first global average pooling layer, and is used to obtain reduced-dimensionality spatiotemporal features according to the spatiotemporal features of the distorted cartoon-like video frame sequence; The input end of the texture cross-attention mechanism module is respectively connected to the output end of the cosine similarity calculation module and the linear layer to obtain the texture features of the distorted cartoon video frame sequence V according to the frame-level texture features and the reduced-dimensional spatiotemporal features.

5. The comprehensive blind evaluation method for cartoon video quality according to claim 1, characterized in that: The multi-task self-supervised learning network model for cartoon videos includes: a first series fusion module, a second series fusion module, a first prediction head module, a second prediction head module, and a third prediction head module; The input end of the first series fusion module is connected to the output ends of the line perception spatiotemporal feature extraction module and the text-guided quality perception module respectively, and is used to obtain the comprehensive distortion feature of the distorted cartoon video frame sequence according to the spatiotemporal feature of the distorted cartoon video frame sequence, the line feature of the distorted cartoon video frame sequence and the texture feature of the distorted cartoon video frame sequence; The input end of the first prediction head module is connected to the output end of the first series fusion module, and is used to obtain the distortion type loss of the distorted cartoon video frame sequence according to the comprehensive distortion characteristics of the distorted cartoon video frame sequence; The input end of the second prediction head module is connected to the output end of the first series fusion module, and is used to obtain the distortion level loss of the distorted cartoon video frame sequence according to the comprehensive distortion characteristics of the distorted cartoon video frame sequence; The input end of the second serial fusion module is connected to the output end of the text-guided quality perception module, and is used to obtain the comprehensive quality features of the distorted cartoon video frame sequence according to the aesthetic features of the distorted cartoon video frame sequence and the comprehensive distortion features of the distorted cartoon video frame sequence; The input end of the third prediction head module is connected to the output end of the second serial fusion module, and is used to obtain the pairwise quality ranking loss of the distorted cartoon video frame sequence according to the comprehensive quality characteristics of the distorted cartoon video frame sequence.

6. A comprehensive blind evaluation method for cartoon video quality according to claim 1, characterized in that: The loss functions of the multi-task self-supervised learning network model for cartoon videos include a distortion type loss function, a distortion level loss function, and a pairwise quality ranking loss function; The calculation formula of the loss function of the multi-task self-supervised learning network model of the cartoon video is as follows: L total =L rank +α×L type +β×L level Where: L total is the loss function of the multi-task self-supervised learning network model for cartoon videos, i.e., the joint loss function of distortion type loss, distortion level loss, and pairwise quality ranking loss; α and β are both weight coefficients; L rank The value representing the pairwise quality ranking; L type The value representing the distortion type loss of cartoon-like video frames; L level A value representing the distortion level loss of cartoon-like video frames; in, Where P represents the pre-training batch size; F CE is the cross entropy loss function, T represents the predicted distortion type; T GT Indicates the true label of the distortion type; D indicates the predicted degree of distortion; D GT The true label indicating the degree of distortion; Where i, j represent the distortion level index, i≠j; L represents the total distortion level; R j R represents the quality of the distorted cartoon-like video frame sequence at the predicted j-th distortion level; i represents the quality of the distorted cartoon-like video frame sequence at the i-th distortion level; ε is a hyperparameter.

7. A comprehensive blind evaluation method for cartoon video quality according to claim 1, characterized in that: The cartoon video quality comprehensive blind assessment model comprises: a line-aware spatiotemporal feature extraction module, a text-guided quality perception module and a quality prediction module; The input end of the quality prediction module is connected to the output end of the line perception spatiotemporal feature extraction module and the text-guided quality perception module respectively, and is used to extract the spatiotemporal features F of the distorted cartoon video frame sequence V according to the spatiotemporal features F of the distorted cartoon video frame sequence V. ST , line features of the distorted cartoon video frame sequence V, aesthetic features of the distorted cartoon video frame sequence V, and texture features of the distorted cartoon video frame sequence V, to obtain the quality score of the distorted cartoon video frame sequence.

8. A comprehensive blind evaluation method for cartoon video quality according to claim 1, characterized in that: The loss function of the comprehensive blind evaluation model for cartoon video quality is expressed as follows: Where: B is the fine-tuning batch size; Q GT is the real quality score of cartoon videos; PLCC(·) is the Pearson linear correlation coefficient; Q is the predicted quality score of cartoon videos; L Q is the value of the loss function of the comprehensive blind assessment model for cartoon video quality.