No-reference video quality evaluation method based on large language model
By converting human subjective video quality scores into text-level scores and using large language models to extract features from video content and text information, the problem of insufficient generalization ability of video quality evaluation methods in the existing technology is solved, and accurate evaluation of new types of video content and robust visual scoring are achieved.
Patent Information
- Application Number
- CN202510157406.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-23
AI Technical Summary
When facing new and different types of video content, the existing video quality evaluation methods have poor generalization capabilities and are difficult to accurately adapt to and evaluate these new content, resulting in their performance degradation when dealing with different scoring scenarios.
Using a reference-free video quality evaluation method based on a large language model, a video quality evaluation data set without reference-level scores is constructed by converting continuous scores in the human subjective video quality scoring process into discrete text-level scoring, and using a large language model to extract features from video content and text information to simulate the human visual output text-level scoring.
The generalization ability of the model is improved, allowing it to predict accurate and robust visual ratings when facing new and different types of video content, and improves the accuracy and adaptability of video quality evaluation.
Smart Images

Figure CN120032298A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large language models, and in particular, to a reference-free video quality evaluation method based on large language models. Background Art
[0002] With the explosive growth of online video content, there is an increasing demand for accurate video quality assessors to robustly evaluate the quality scores of videos in various types of visual content. Although existing methods have been able to achieve significant accuracy on specific datasets, they often struggle to accurately adapt to and evaluate new content when faced with new and different types of content due to their limited ability to handle complex scoring factors, resulting in poor generalization ability.
[0003] Xidian University disclosed a reference-free video quality evaluation method based on three-dimensional spatio-temporal feature decomposition in its patent literature "Reference-free Video Quality Evaluation Method Based on Three-dimensional Spatio-temporal Feature Decomposition" (Patent Application No.: CN202010944337.3; Publication No.: CN112085102B). This method decomposes a video into spatial and temporal features. The spatial features are extracted by a convolutional neural network, and the temporal features are extracted by an optical flow method or a 3D convolutional neural network. Then the extracted spatio-temporal features are fused to form comprehensive features. Finally, a regression model (such as support vector regression or neural network) is used to train the fused features to predict the video quality score. The disadvantage of this method is that the extraction of temporal features depends on the optical flow method or 3D convolutional neural network, but in fast-moving or complex scenes, the optical flow method may not be able to accurately capture motion information, resulting in inaccurate extraction of temporal features, and thus it cannot accurately simulate the human quality perception process and the prediction result accuracy is not high.
[0004] Recently, large language models have attracted much attention due to the great potential they have shown in many related fields. They have strong generalization ability and can adapt to diverse and complex data environments. Using large language models for video quality assessment has good potential. Therefore, how to teach large language models to subjectively score videos like humans is an urgent challenge currently faced.
[0005] The difficulty in solving the above problems is that existing scoring models often struggle to accurately adapt to and evaluate new content when faced with new and different types of video content due to their limited ability to handle complex scoring factors, resulting in poor generalization ability. Using large language models for video quality assessment has good potential. How to teach large language models to subjectively evaluate videos like humans is an urgent challenge currently faced.
[0006] To this end, the present invention (project: Research and Verification of Key Technology System of Intelligent Computing Network, project number: 2022ZD0115300) proposes a reference-free video quality evaluation method based on a large language model, and develops an effective reference-free video quality evaluation method for a large language model, so that the large language model can perform subjective video quality evaluation on the video like a human, and has strong generalization when facing new and different types of video content, and can predict accurate and robust visual scores. Summary of the invention
[0007] The present invention provides a no-reference video quality assessment method based on a large language model to solve the problem that the video quality assessment method has poor generalization ability and usually encounters performance degradation when processing different scoring scenarios (such as mixing multiple data sets).
[0008] The technical solution of the present invention is as follows: The method for no-reference video quality assessment based on a large language model of the present invention comprises the following steps: S1. converting continuous scores (MOS values) in the process of human subjective video quality scoring into discrete text-level scores; S2. constructing a video quality assessment dataset with no-reference text-level scores; S3. constructing a large language model, constructing a visual encoder, and constructing a text encoder; S4. fine-tuning the large language model, formatting the text-level scores into instruction-response pairs, and performing visual instruction adjustment on the LLaMA2 model; and S5. model inference, converting the text-level scores output by model inference into MOS values.
[0009] Optionally, in the above-mentioned large language model-based no-reference video quality assessment method, in step S1, the scores are converted into rating levels using equidistant intervals, the range between the highest score Z and the lowest score z of the continuous score MOS value is uniformly divided into five different intervals, and the scores of each interval are assigned to respective levels: , if (1) Where s is the continuous fraction MOS value, is a standard text rating level defined by the ITU.
[0010] Optionally, in the above-mentioned reference-free video quality assessment method based on a large language model, in step S2, a video quality assessment dataset is constructed, each data sample includes a set of image sequences, a text-level score, and a subjective quality assessment score; the RGB values of the image sequences are normalized to the interval [0,1]; and the subjective quality assessment scores corresponding to all image sequences are converted into discrete text-level scores.
[0011] Optionally, in the above-mentioned no-reference video quality assessment method based on a large language model, in step S3, for the input text and video, the input information needs to be preprocessed first: the video frame information needs to first pass through a visual encoder composed of a Vision Transformer, and the visual encoder encodes the image into tokens of length 1024; then the tokens are input into a visual feature extractor composed of multiple Swin Transformers, and the visual feature extractor reduces the number of tokens for each image from 1024 to 64.
[0012] Optionally, in the above-mentioned large language model-based no-reference video quality assessment method, in step S4, the dialogue format of the training task is defined, and the image sequence features are represented as , the text-level score of the video is expressed as <level>After the large language model inputs the preprocessed image sequence features and text features, it outputs the text level score for the video. The sample dialogue format of the training task is as follows: enter: Rate the quality of the video ; Output: The quality of the video is <level>; Label: The quality of the video is < good > The loss function used in training is the cross entropy loss function: (2) in and The true label and predicted label corresponding to the i-th category, n is the number of categories; During the training process, the image sequence and text in the training set are input into the image encoder and text encoder in turn to extract the image sequence features and text features respectively. The features are then sent to the LLaMA2 model for training, the cross entropy loss is calculated, and the LLaMA2 model is fine-tuned through back propagation.
[0013] Optionally, in the above-mentioned large language model-based no-reference video quality assessment method, in step S5, the final score is predicted based on probability-based large language model reasoning, the human scoring process is simulated, and the text level score output by reasoning is converted into a MOS value by weighted averaging; first, a reverse mapping T from the text level score to the score is defined as follows: (3) in is the text level score, i is the MOS value; Simulating the human scoring process, the MOS value is the conversion score and frequency of each text level score. The weighted average is calculated as: (4) Since the prediction of a large language model is a probability distribution over all possible text-level ratings, denoted as X, in Maximize softmax on the result to get the probability of each text level score ,all The sum of is 1, The definition is as follows: (5) Finally, the score predicted by the large language model is defined as : (6).
[0014] According to the technical solution of the present invention, the beneficial effects produced are: By imitating the visual scoring process of human scoring and post-processing, the present invention proposes a video quality assessment method using text discrete level scoring, which is used to train large language models and improve the generalization ability of the model.
[0015] In order to better understand and illustrate the concept, working principle and effect of the present invention, the present invention is described in detail below through specific embodiments in conjunction with the accompanying drawings: BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the specific implementation of the present invention or the technical solution in the prior art, the drawings required for use in the specific implementation or the description of the prior art are briefly introduced below.
[0017] Figure 1 A flowchart of a no-reference video quality assessment method based on a large language model provided by the present invention; Figure 2 This is a model structure diagram of the no-reference video quality assessment method based on a large language model provided by the present invention. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical method and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and specific examples. These examples are only illustrative and not limiting of the present invention.
[0019] The non-reference video quality evaluation method based on a large language model of the present invention teaches the large language model to use text-level scoring to evaluate video quality like humans, rather than just using MOS scores to teach the large language model. The method of the present invention converts the continuous score (MOS value) in the process of human subjective video quality scoring into a discrete text-level score, then processes the video content and its corresponding text instructions as input, uses a large language model to extract features from these video content and text information, simulates human vision, and outputs a text-level score; the whole process does not require a reference image, only the video content and text information are needed for non-reference visual scoring.
[0020] like Figure 1 and Figure 2 As shown, the large language model-based no-reference video quality assessment method of the present invention comprises the following steps: S1. Convert the continuous score (MOS value) in the process of human subjective video quality rating into discrete text-level rating; In step S1, the continuous score (MOS value) in the process of human subjective video quality rating is converted into a discrete text-level rating (e.g., excellent, good, average, poor, and extremely poor). In the past, video quality evaluation tasks required the use of continuous scores (MOS values) for training. In the present invention, the continuous score (MOS value) needs to be converted into a discrete text-level rating so that the large language model can learn to use the rating to predict video quality. This process is adopted because in daily life, when asked to make an evaluation, people tend to respond with qualitative adjectives (e.g., excellent, good, average, bad, and extremely poor) rather than numerical ratings (e.g., 8.15, 1.56, and 4.67). Therefore, using text-level ratings for video quality evaluation tasks and utilizing this innate ability of humans (providing qualitative adjectives) can minimize human cognitive load and improve the results of subjective visual evaluation.
[0021] First, the continuous score (MOS value) needs to be converted into a discrete text-level score. Since adjacent levels in human ratings are essentially equidistant, the present invention uses equidistant intervals to convert the score into a rating level. Specifically, the range between the highest MOS score (Z) and the lowest score (z) is uniformly divided into five different intervals, and the scores in each interval are assigned to their respective levels: , if (1) Where s is the MOS value, is a standard text rating level defined by the ITU.
[0022] S2. Construct a video quality assessment dataset without reference text-level ratings; Specifically, a dataset is constructed, in which each data sample includes a set of image sequences, a text-level score, and a subjective MOS value; the RGB values of the image sequences are normalized to the interval [0,1]; and the subjective quality evaluation continuous scores (subjective MOS values) corresponding to all image sequences are converted into discrete text-level scores.
[0023] S3. Build a large language model, build a visual encoder, and build a text encoder; The large language model structure of the present invention is based on the recently released open source LLaMA2-7B model, which has excellent visual perception and good language understanding capabilities. For input text and video, the input information needs to be preprocessed first: the video frame information needs to pass through a visual encoder, which is composed of a Vision Transformer, which encodes the image into a token of length 1024; then the token is input into a visual feature extractor, which is composed of multiple Swin Transformers, which reduces the number of tokens per picture from 1024 to 64. The LLaMA2 model supports a context length of 2048, which allows the present invention to put together an image sequence of up to 40 pictures as a video input to the large language model during supervised fine-tuning.
[0024] For the input text, it needs to go through a text encoder. The text encoder first uses a Tokenizer to divide the text into a sequence of vocabulary units, and then uses Embedding to represent these vocabulary units as vectors. The main function of a Tokenizer is to divide the text into vocabulary units (tokens). This process usually includes segmenting words, punctuation marks, numbers, etc. in the text, and removing some noise or redundant information. The word segmenter used in the present invention is WordPiece-Tokenizer, a letter-based word segmentation method that builds vocabulary by learning letter-level vocabulary units and is suitable for processing vocabulary of indefinite length. Embedding is the process of mapping vocabulary units to vector space. Word embedding technology usually uses a neural network model to map each vocabulary unit to a fixed-length vector, so that the vector representation of each vocabulary unit can capture its semantics and contextual information. The Embedding used in the present invention is Word2Vec, a natural language processing (NLP) technology for generating word vectors, launched by Google in 2013.
[0025] Subsequently, the features of the input text and image obtained after preprocessing will be sent together to the LLaMA2 model for training.
[0026] S4. Fine-tune a large language model to format text-level scores into command-response pairs for visual command fine-tuning on the LLaMA2 model; The dialogue format of the training task is defined. The image sequence features are represented as , the text-level score of the video is expressed as <level>After the large language model inputs the preprocessed image sequence features and text features, it outputs the text level score for the video. The sample dialogue format of the training task is as follows: enter: Rate the quality of the video .
[0027] Output: The quality of the video is <level>.
[0028] Label: The quality of the video is < good >.
[0029] The loss function used in training is the cross entropy loss function: (2) in and The true label and predicted label corresponding to the i-th category, n is the number of categories.
[0030] During the training process, the image sequence and text in the training set are input into the image encoder and text encoder in turn to extract the image sequence features and text features respectively. The features are then sent to the LLaMA2 model for training, the cross entropy loss is calculated, and the LLaMA2 model is fine-tuned through back propagation.
[0031] S5. Model inference, converting the text-level score output by model inference into a MOS value.
[0032] After training, model inference is performed, and the text-level score output by the model inference needs to be converted into a MOS value. The present invention proposes a large-scale language model inference based on probability to predict the final score, simulates the human scoring process, and converts the text-level score output by the inference into a MOS value by weighted average. The present invention first defines the reverse mapping T from the text-level score to the score as follows: (3) in is the text level score, i is the MOS value. For example, good is converted to a score of 4, and bad is converted to a score of 1.
[0033] Simulating the human scoring process, the MOS value is the conversion score and frequency of each text level score. The weighted average is calculated as: (4) Since the prediction of a large language model is a probability distribution over all possible text-level ratings (denoted as X), the present invention Maximize softmax on the result to get the probability of each text level score ,all The sum of is 1, The definition is as follows: (5) Finally, the score predicted by the large language model is defined as : (6) The effect of the present invention is further described below in conjunction with simulation experiments: (1) Simulation experiment conditions: The hardware platform of the simulation experiment of the present invention is: the processor is Intel (R) Xeon (R) Silver 4310 CPU, the main frequency is 2.10 GHz, the memory is 128 GB, and the graphics card is NVIDIA -A100.
[0034] The software platform of the simulation experiment of the present invention is: CentOS7.6 operating system, Pytorch2.0.1 framework, Python3.9.
[0035] The present invention uses the large-scale LSVQ dataset (containing 28,056 videos) for training and two test sets, LSVQtest and LSVQ1080p (LSVQ's official dataset internal test subset), for testing. LSVQtest consists of 7,400 videos of different resolutions, ranging from 240P to 720P, while LSVQ1080p consists of 3,600 1080P high-resolution videos.
[0036] Under the above simulation experimental conditions, the method of the present invention and three baseline methods Simple-VQA, PVQ, and TLVQM are used to perform video quality evaluation on the LSVQtest and LSVQ1080p datasets, predict the quality scores of the videos, and calculate their respective Spearman rank correlation coefficients SRCC and Pearson linear correlation coefficients PLCC. The results are shown in Table 1.
[0037] Table 1. Comparison of evaluation results of various methods As can be seen from Table 1, the Spearman rank correlation coefficient SRCC and Pearson linear correlation coefficient PLCC of the present invention on two video quality assessment datasets are mostly higher than those of the original baseline method. Simulation experiments show that the method based on the large language model of the present invention has better video quality assessment effect.
[0038] The above description is the best embodiment based on the concept and working principle of the invention. The above embodiment should not be understood as limiting the protection scope of the present claims, and other implementations and combinations of implementations according to the concept of the present invention belong to the protection scope of the present invention.< / level> < / level> < / level> < / level>
Claims
1. A no-reference video quality assessment method based on a large language model, characterized in that: The following steps are involved: S1. Convert the continuous score MOS value in the process of human subjective video quality rating into discrete text-level rating; S2. Construct a video quality assessment dataset without reference text-level ratings; S3. Build a large language model, build a visual encoder, and build a text encoder; S4. Fine-tuning the large language model to format text-level scores into command-response pairs for visual command adjustment on the LLaMA2 model; and S5. Model inference, converting the text-level score output by model inference into a MOS value.
2. The method for no-reference video quality assessment based on a large language model according to claim 1, characterized in that: In step S1, the scores are converted into rating levels using equidistant intervals, the range between the highest score Z and the lowest score z of the continuous score MOS value is uniformly divided into five different intervals, and the scores in each interval are assigned to respective levels: , if (1) Where s is the continuous fraction MOS value, is a standard text rating level defined by the ITU.
3. The method for no-reference video quality assessment based on a large language model according to claim 1, characterized in that: In step S2, the video quality assessment dataset is constructed, each data sample includes a set of image sequences, a text level score, and a subjective quality assessment score; the RGB values of the image sequences are normalized to the [0,1] interval; and the subjective quality assessment scores corresponding to all the image sequences are converted into discrete text level scores.
4. The method for no-reference video quality assessment based on a large language model according to claim 1, characterized in that: In step S3, for input text and video, the input information needs to be preprocessed first: the video frame information needs to first pass through a visual encoder composed of a Vision Transformer, which encodes the image into tokens of length 1024; then the tokens are input into a visual feature extractor composed of multiple Swin Transformers, which reduces the number of tokens for each image from 1024 to 64.
5. The method for no-reference video quality assessment based on a large language model according to claim 1, characterized in that: In step S4, the dialogue format of the training task is defined, and the image sequence features are represented as , the text level score of the video is represented as < LEVEL >, and the large language model inputs the preprocessed image sequence features and text features, and outputs the text level score of the video. The example dialogue format of the training task is as follows: enter: Rate the quality of the video ; Output: The quality of the video is < LEVEL >; Label: The quality of the video is < good > The loss function used in training is the cross entropy loss function: (2) in and The true label and predicted label corresponding to the i-th category, n is the number of categories; During the training process, the image sequence and text in the training set are input into the image encoder and the text encoder in turn, and the image sequence features and the text features are extracted respectively. Then the features are sent to the LLaMA2 model for training, the cross entropy loss is calculated, and the LLaMA2 model is fine-tuned through back propagation.
6. The method for no-reference video quality assessment based on a large language model according to claim 1, characterized in that: In step S5, the large language model inference based on probability is used to predict the final score, simulating the process of human scoring, and converting the text-level score output by the inference into a MOS value by weighted average; first, the reverse mapping T from the text-level score to the score is defined as follows: (3) in is the text level score, i is the MOS value; Simulating the human scoring process, the MOS value is the conversion score and frequency of each text level score. The weighted average is calculated as: (4) Since the prediction of the large language model is a probability distribution over all possible text-level ratings, denoted as X, in Maximize softmax on the result to get the probability of each text level score ,all The sum of is 1, The definition is as follows: (5) Finally, the score predicted by the large language model is defined as : (6)。
Citation Information
Patent Citations
A No-Reference Video Quality Assessment Method Based on 3D Spatiotemporal Feature Decomposition
CN112085102B