Video quality evaluation model training method, video quality evaluation method, video quality evaluation device and video quality evaluation equipment

By constructing a target frame-level feature matrix and learning the temporal features between video frames, the problem of ignoring the temporal characteristics of video in existing technologies is solved, and more accurate video quality assessment is achieved.

CN120953877APending Publication Date: 2025-11-14BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511062901.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing video quality assessment methods ignore the temporal characteristics of videos, leading to inaccurate assessment results.

Method used

By acquiring video quality training samples, the video frames are evaluated using an image quality assessment model, frame-level features are extracted, and a target frame-level feature matrix is ​​constructed based on the frame-level features and temporal relationships. The video quality assessment model is then trained to learn the temporal feature influence between frame-level quality scores.

Benefits of technology

This improves the accuracy of video quality assessment, better reflects the impact of video temporal characteristics on quality, and enhances the accuracy of assessment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953877A_ABST
    Figure CN120953877A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video quality evaluation model training method and device, a video quality evaluation method and device, electronic equipment, a storage medium and a program product. Comprising the steps that a training data set of a plurality of video quality training samples is acquired, and the video quality training samples comprise video samples and video quality labels; evaluating each video frame in the video sample by using the image quality evaluation model to obtain a frame-level quality score corresponding to each video frame; based on the time sequence and the frame-level quality score of each video frame, obtaining a target frame-level feature matrix; performing quality evaluation on the target frame-level feature matrix through a to-be-trained video quality evaluation model to obtain a video quality score; and training a to-be-trained video quality evaluation model according to the video quality score and the video quality label to obtain the video quality evaluation model. The video quality evaluation model can learn the influence of the time sequence characteristics between the frame-level quality scores on the video quality, so that the evaluation result of the video quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more specifically, to a video quality assessment model training method, a video quality assessment method, an apparatus, an electronic device, a storage medium, and a computer program product. Background Technology

[0002] With the rapid development of video streaming, social media, and video conferencing, users' expectations for the quality of video content are increasing.

[0003] In related technologies, video quality assessment typically involves calculating the quality score of each video frame individually or by sampling the frames, and then calculating the average or median of the quality scores of all video frames as the overall quality score of the video.

[0004] The above quality assessment methods ignore the temporal characteristics of videos and fail to reflect the impact of video temporal characteristics on quality, resulting in poor assessment results. Summary of the Invention

[0005] This disclosure provides a video quality assessment model training method, a video quality assessment method, an apparatus, an electronic device, a storage medium, and a computer program product. The model can learn the impact of temporal features between frame-level quality scores on video quality, thereby improving the video quality assessment results.

[0006] According to a first aspect of the present disclosure, a method for training a video quality assessment model is provided, comprising: acquiring a training dataset of multiple video quality training samples, wherein the video quality training samples include corresponding video samples and video quality labels; evaluating each video frame in the video samples using an image quality assessment model for each video sample in the training dataset to obtain a frame-level quality score corresponding to each video frame; obtaining a target frame-level feature matrix based on the temporal and frame-level quality scores of each video frame; performing quality assessment on the target frame-level feature matrix using the video quality assessment model to be trained to obtain a video quality score; and training the video quality assessment model to be trained based on the video quality score and the video quality label to obtain a video quality assessment model.

[0007] In some exemplary embodiments of this disclosure, obtaining a target frame-level feature matrix based on the temporal sequence and frame-level quality score of each video frame includes: extracting features from each video frame to obtain frame-level features corresponding to each video frame; and obtaining a target frame-level feature matrix based on the temporal sequence of each video frame, according to the frame-level features and frame-level quality score.

[0008] In some exemplary embodiments of this disclosure, feature extraction is performed on each video frame to obtain frame-level features corresponding to each video frame, including: multi-dimensional feature extraction is performed on each video frame to obtain frame-level features corresponding to each video frame in the corresponding dimension; based on the temporal sequence of each video frame, a target frame-level feature matrix is ​​obtained according to the frame-level features and frame-level quality scores, including: based on the temporal sequence of each video frame, a target frame-level feature matrix is ​​obtained according to the frame-level features and frame-level quality scores of each video frame in the corresponding dimension.

[0009] In some exemplary embodiments of this disclosure, the frame-level features include frame-level motion features, frame-level attention region features, and frame-level semantic features. Multi-dimensional feature extraction is performed on each video frame to obtain the frame-level features corresponding to each video frame in the corresponding dimension, including: calculating the frame-level motion features of each video frame using an optical flow field algorithm; extracting the frame-level attention region features of each video frame using an attention region detection model; and extracting the frame-level semantic features of each video frame using a semantic encoder.

[0010] In some exemplary embodiments of this disclosure, a target frame-level feature matrix is ​​obtained based on the temporal sequence of each video frame, according to frame-level features and frame-level quality scores, including: concatenating frame-level features and frame-level quality scores based on the temporal sequence of each video frame to obtain a first frame-level feature matrix; performing inter-frame difference calculation on the first frame-level feature matrix to obtain a difference feature matrix; and using the difference feature matrix as the target frame-level feature matrix.

[0011] In some exemplary embodiments of this disclosure, after performing inter-frame difference calculation on the first frame-level feature matrix to obtain the difference matrix features, the method further includes: concatenating the difference feature matrix with the first frame-level feature matrix to obtain the target frame-level feature matrix.

[0012] In some exemplary embodiments of this disclosure, a target frame-level feature matrix is ​​obtained based on the temporal sequence of each video frame, according to frame-level features and frame-level quality scores, including: performing abnormal frame detection on each video frame to obtain an abnormal identifier value corresponding to each video frame, wherein the abnormal identifier value is used to identify whether the corresponding video frame is an abnormal frame; and obtaining the target frame-level feature matrix based on the temporal sequence of each video frame, according to frame-level features, frame-level quality scores and abnormal identifier values.

[0013] In some exemplary embodiments of this disclosure, obtaining a target frame-level feature matrix based on the temporal and frame-level quality scores of each video frame includes: determining the target video length corresponding to the training dataset; obtaining a second frame-level feature matrix based on the temporal and frame-level quality scores of each video frame; and if the length of a video sample is less than the length of the target video, padding is used to fill the second frame-level feature matrix to obtain the target frame-level feature matrix.

[0014] In some exemplary embodiments of this disclosure, the target frame-level feature matrix is ​​obtained by filling the second frame-level feature matrix with padding values, including: generating a mask vector based on the temporal sequence of each video frame, wherein the mask vector is used to identify whether the position of the corresponding video frame is obtained by padding; and concatenating the padded second frame-level feature matrix with the mask vector to obtain the target frame-level feature matrix.

[0015] According to a second aspect of the present disclosure, a video quality assessment method is provided, comprising: acquiring a video to be assessed; inputting the video to be assessed into a video quality assessment model; wherein the video quality assessment model is trained by the video quality assessment model training method as described in the first aspect above, the video quality assessment model is used to assess the quality of the video to be assessed, and obtain a quality score for the video to be assessed; and outputting the quality score for the video to be assessed.

[0016] According to a third aspect of the present disclosure, a video quality assessment model training apparatus is provided, comprising: a sample acquisition module configured to acquire a training dataset of multiple video quality training samples, wherein the video quality training samples include corresponding video samples and video quality labels; a frame-level quality assessment module configured to evaluate each video frame in the video samples using an image quality assessment model for each video sample in the training dataset, thereby obtaining a frame-level quality score corresponding to each video frame; a target frame-level feature matrix determination module configured to obtain a target frame-level feature matrix based on the temporal sequence and frame-level quality scores of each video frame; a video quality score determination module configured to perform quality assessment on the target frame-level feature matrix using a video quality assessment model to be trained, thereby obtaining a video quality score; and a model training module configured to train the video quality assessment model to be trained based on the video quality scores and video quality labels, thereby obtaining a video quality assessment model.

[0017] According to a fourth aspect of the present disclosure, a video quality assessment apparatus is provided, comprising: an assessment video acquisition module configured to acquire a video to be assessed; an assessment video input module configured to input the video to be assessed into a video quality assessment model; wherein the video quality assessment model is trained using the video quality assessment model training method as described in the first aspect above, and the video quality assessment model is used to assess the quality of the video to be assessed to obtain a quality score for the video to be assessed; and a quality score output module configured to output the quality score for the video to be assessed.

[0018] According to a fifth aspect of the present disclosure, an electronic device is provided, characterized in that it includes: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the executable instructions to implement either a video quality assessment model training method or either a video quality assessment method.

[0019] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform either a video quality assessment model training method or either a video quality assessment method.

[0020] According to a seventh aspect of the present disclosure, a computer program product is provided, including a computer program that is executed by a processor as either a video quality assessment model training method or a video quality assessment method.

[0021] The video quality assessment model training method provided in this disclosure evaluates each video frame in a video sample using an image quality assessment model, obtaining a frame-level quality score for each video frame. Based on the temporal and frame-level quality scores of each video frame, a target frame-level feature matrix is ​​obtained. The video quality assessment model is then trained using the target frame-level feature matrix and the video quality labels corresponding to the video samples. This method trains the video quality assessment model using the target frame-level feature matrix. Since the target frame-level feature matrix includes the temporal features of each video frame, the video quality assessment model can learn the influence of the temporal features between frame-level quality scores on video quality. This allows the video quality assessment model to more accurately assess video quality, thereby improving the video quality assessment results.

[0022] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0024] Figure 1 A schematic diagram of an exemplary system architecture for which the video quality assessment model training method of the present disclosure can be applied is shown.

[0025] Figure 2 This is a flowchart illustrating a video quality assessment model training method according to an exemplary embodiment.

[0026] Figure 3 This is a flowchart illustrating a method for determining a target frame-level feature matrix according to an exemplary embodiment.

[0027] Figure 4 This is a flowchart illustrating another method for determining the target frame-level feature matrix according to an exemplary embodiment.

[0028] Figure 5 This is a flowchart illustrating yet another method for determining a target frame-level feature matrix according to an exemplary embodiment.

[0029] Figure 6 This is a flowchart illustrating another method for determining a target frame-level feature matrix according to an exemplary embodiment.

[0030] Figure 7 This is a flowchart illustrating another method for determining a target frame-level feature matrix according to an exemplary embodiment.

[0031] Figure 8 This is a flowchart illustrating a video quality assessment method according to an exemplary embodiment.

[0032] Figure 9 This is a block diagram illustrating a video quality assessment model training apparatus according to an exemplary embodiment.

[0033] Figure 10 This is a block diagram illustrating a video quality assessment apparatus according to an exemplary embodiment.

[0034] Figure 11 This is a schematic diagram illustrating the structure of an electronic device suitable for implementing exemplary embodiments of the present disclosure, according to an exemplary embodiment. Detailed Implementation

[0035] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted.

[0036] The features, structures, or characteristics described in this disclosure can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more specific details omitted, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0037] The accompanying drawings are merely illustrative of this disclosure, and the same reference numerals in the drawings denote the same or similar parts, thus omitting repeated descriptions of them. Some block diagrams shown in the drawings do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in at least one hardware module or integrated circuit, or in different network and / or processor devices and / or microcontroller devices.

[0038] The flowchart shown in the accompanying drawings is merely illustrative and does not necessarily include all content and steps, nor does it require execution in the described order. For example, some steps may be broken down, while others may be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0039] In this specification, the terms “a,” “an,” “the,” “the,” and “at least one” are used to indicate the presence of at least one element / component / etc.; the terms “comprising,” “including,” and “having” are used to indicate an open-ended inclusion and to mean that there may be other elements / components / etc. in addition to the listed elements / components / etc.; the terms “first,” “second,” and “third,” etc., are used only as markings and are not a limitation on the number of objects.

[0040] Figure 1 A schematic diagram of an exemplary system architecture to which the methods of embodiments of this disclosure can be applied is shown.

[0041] like Figure 1 As shown, the system architecture may include server 101, network 102, terminal device 103, terminal device 104, and terminal device 105. Network 102 serves as the medium for providing a communication link between terminal device 103, terminal device 104, or terminal device 105 and server 101. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0042] Server 101 can be a server that provides various services, such as a back-end management server that supports the devices operated by users using terminal devices 103, 104, or 105. The back-end management server can analyze and process received requests and other data, and feed back the processing results to terminal devices 103, 104, or 105.

[0043] Terminal devices 103, 104, and 105 can be smartphones, tablets, laptops, desktop computers, smart speakers, wearable smart devices, virtual reality devices, augmented reality devices, etc., but are not limited to these.

[0044] It should be understood that Figure 1 The number of terminal devices 103, 104, 105, network 102, and server 101 in the diagram is merely illustrative. Server 101 can be a single physical server, a server cluster consisting of multiple servers, or a cloud server. Depending on actual needs, it can have any number of terminal devices, networks, and servers.

[0045] The steps of the method in the exemplary embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings and examples.

[0046] Figure 2 This is a flowchart illustrating a video quality assessment model training method according to an exemplary embodiment. Figure 2 The method provided in the embodiments can be executed by any electronic device, such as the one described above. Figure 1 Terminal devices in or Figure 1 The server in, or Figure 1 The terminal devices and servers in the process execute together, but this disclosure does not limit this. For example... Figure 2 As shown, the video quality assessment model training method provided in this embodiment mainly includes steps S210-S250.

[0047] In step S210, a training dataset of multiple video quality training samples is obtained. The video quality training samples include corresponding video samples and video quality labels.

[0048] In this embodiment of the disclosure, in order to train the video quality assessment model to be trained, a training dataset related to video quality assessment is first obtained. This training dataset includes multiple video quality training samples. Each video quality training sample includes at least one set of video samples and video quality labels, with a corresponding relationship between video samples and video quality labels within the same set. The video quality labels are determined by multiple manually labeled quality labels.

[0049] In an exemplary embodiment, to adapt to different video quality assessment scenarios, the aforementioned video quality label may include a quality score or a rating level. The rating level is a classification and evaluation of video quality using non-numerical levels and descriptive language, such as "Excellent, Good, Average, Poor" or "Clear, Fairly Clear, Blurry, Severely Blurry." The quality score is a precise measurement of video quality using specific numerical values, reflecting the degree of quality.

[0050] In one possible implementation, the rating results of multiple people watching the video sample are obtained. If the rating results are in rating levels, the median of multiple rating levels is obtained as the video quality label of the video sample. If the rating results are in quality scores, the mean of multiple quality scores is obtained as the video quality label of the video sample.

[0051] In step S220, for each video sample in the training dataset, the image quality assessment model is used to evaluate each video frame in the video sample to obtain the frame-level quality score corresponding to each video frame.

[0052] An image quality assessment model can be understood as an algorithmic model used to evaluate the quality of a single video frame, and the output may include a quantized score or a rating level. Furthermore, this image quality assessment model is a no-reference model, meaning it does not require a reference image and scores the video frame directly by analyzing its statistical characteristics. Examples of image quality assessment models include, but are not limited to: No-Reference Image Quality Assessment (NR-IQA) models or Blind / Referenceless Image Spatial Quality Evaluator (BRISQUE).

[0053] Frame-level quality scoring can be understood as the quality quantification result of an image quality assessment model for a single video frame, reflecting the quality level of that video frame when it exists independently. Frame-level quality scoring only reflects the quality of a single video frame and does not consider the impact of temporal continuity between video frames (such as flickering or stuttering) on ​​video quality. Frame-level quality scoring includes quantization scores or rating levels.

[0054] In an exemplary embodiment, for each video sample in the training dataset, each video sample is processed frame by frame or frame by frame along the timeline to extract multiple video frames from the video sample. For each video frame, the video frame is input into an image quality assessment model, which performs a quality assessment on the input single video frame to obtain a frame-level quality score corresponding to that video frame.

[0055] In step S230, the target frame-level feature matrix is ​​obtained based on the temporal and frame-level quality scores of each video frame.

[0056] The temporal sequence of each video frame can be understood as the order in which the video frames in the video sample are arranged in the time dimension. The essence of video is a continuous sequence of video frames. The temporal sequence of video frames can be represented by the timestamp of each video frame (such as 0.03 milliseconds in the 1st second, 0.05 milliseconds in the 2nd second) or the frame number (such as frame 1, frame 2, frame N). The temporal sequence of video frames reflects the sequential relationship between video frames in the time dimension. For example, frame 1 and frame 2 are adjacent, and frame 1 comes before frame 2.

[0057] In one possible implementation, the temporal information of video frames is integrated with frame-level quality scores to form a feature matrix that reflects the correlation between time and frame-level quality scores. Specifically, based on the temporal identifier (timestamp or frame number) of each video frame, the chronological order of the video frames is determined, and the corresponding frame-level quality scores are matched. A matrix is ​​designed where the rows represent feature types, i.e., the rows of the matrix are frame-level quality scores, and the columns of the matrix are strictly arranged according to the chronological order of the video frames. The frame-level quality scores are used as matrix elements to fill the corresponding positions, forming the target frame-level feature matrix. The target frame-level feature matrix simultaneously retains the quality score of individual video frames and the quality variation patterns over time.

[0058] For example, the video sample has 6 consecutive frames, numbered 1 to 6 in chronological order. Each video frame has a corresponding frame-level quality score of 90, 88, 89, 70, 72, and 85. Arranging these quality scores in the order of the video frames forms the target frame-level feature matrix. Each column of the target frame-level feature matrix corresponds to a video frame, and the first row represents the frame-level quality score. Therefore, the target frame-level feature matrix is ​​a 1-row, 6-column feature matrix, i.e., [90, 88, 89, 70, 72, 85].

[0059] In step S240, the target frame-level feature matrix is ​​evaluated using the video quality assessment model to be trained, and a video quality score is obtained.

[0060] The video quality assessment model to be trained can be understood as an algorithmic model that has not yet been trained and is used to predict the overall quality of the video. The video quality assessment model to be trained is used to learn the temporal pattern of frame-level quality scores from the input target frame-level feature matrix, and finally outputs a quantified quality score, which is the video quality rating.

[0061] Optionally, the video quality assessment model to be trained includes a temporal model. A temporal model is a model that can process time-series data and can consider the temporal order and dependencies of the input data. The video quality assessment model to be trained includes, but is not limited to: Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Temporal Convolutional Network (TCN), Transformer structure models, etc.

[0062] In one possible implementation, an LSTM model is chosen as the video quality assessment model to be trained. The input layer dimension of the LSTM model matches the number of columns in the target frame-level feature matrix. The intermediate layers capture the temporal relationships between video frames through recurrent units or convolutional kernels, and the output layer is a single neuron that outputs the video quality score. The target frame-level feature matrix is ​​input into the LSTM model, and the input, intermediate, and output layers of the LSTM model process the target frame-level feature matrix sequentially. The output layer then outputs the video quality score for that video sample.

[0063] In step S250, the video quality assessment model to be trained is trained based on the video quality score and video quality label to obtain the video quality assessment model.

[0064] A video quality assessment model is a model that has been trained and optimized to output a quality score based on the input video. It can be directly used for the quality assessment of new videos.

[0065] In this embodiment, a loss function is used to calculate the loss value between the video quality score and the video quality label. This loss value reflects the current prediction error of the video quality assessment model being trained. Based on this loss value, the model parameters of the video quality assessment model being trained are adjusted using the backpropagation algorithm, so that the loss value gradually decreases.

[0066] The above training process needs to be carried out iteratively on the training dataset. Each iteration updates the parameters using different batches of video samples until the loss value of the video quality assessment model to be trained on the validation set is less than a set threshold. At this point, the video quality assessment model to be trained is optimized and becomes a video quality assessment model that can directly assess the quality of input videos.

[0067] In this embodiment, the perception of video quality depends not only on the quality score of individual video frames but also on the temporal characteristics between video frames. For example, temporal characteristics such as quality abrupt changes in consecutive frames and periodic jitter have a far greater impact on overall video quality than the simple summation of the quality scores of individual video frames. The target frame-level feature matrix preserves the temporal order of video frames, organizing the temporal features of individual video frames and frame-level quality scores into a structured matrix form, enabling the model to capture the temporal dependencies between frame-level quality scores and video frames. By transforming scattered frame-level quality scores into a holistic feature matrix containing a temporal dimension, the model can learn the patterns of video quality changes over time, rather than simply relying on the summation or average of isolated individual video frame quality scores, thus better aligning with the human eye's perception of overall video quality.

[0068] In this embodiment, the target frame-level feature matrix is ​​used to train the video quality assessment model. Since the target frame-level feature matrix includes the temporal features of each video frame, the video quality assessment model can learn the impact of the temporal features between frame-level quality scores on the overall video quality. That is, the model learns the mapping rule between frame-level quality scores and the overall video quality score, implicitly expressing the aggregation rule of human judgment. This enables the video quality assessment model to more accurately assess video quality, thereby improving the video quality assessment results.

[0069] Based on the above examples, this embodiment optimizes the target frame-level feature matrix determination process in step S230, as follows: Figure 3 As shown, the optimized target frame-level feature matrix determination process includes steps S310-S330.

[0070] In step S310, the image quality assessment model is used to evaluate each video frame in the video sample to obtain a frame-level quality score corresponding to each video frame.

[0071] The frame-level quality score determination process provided in this embodiment is the same as the frame-level quality score determination process in embodiment S230 above. For details, please refer to the description in the above embodiment.

[0072] In step S320, feature extraction is performed on each video frame to obtain frame-level features corresponding to each video frame.

[0073] Feature extraction refers to extracting quantitative information from video frames that reflects their quality or content characteristics. It primarily transforms video frames into machine-understandable feature data. Frame-level features are feature vectors or values ​​obtained through feature extraction that describe the attributes of a single video frame. Frame-level features include, but are not limited to: sharpness features, noise features, semantic features, attention region features, motion features, etc.

[0074] In one exemplary implementation, for each video frame, image processing algorithms or pre-trained feature extraction models are used to extract features, resulting in frame-level features corresponding to each video frame. For example, edge detection operators are used to extract edge sharpness features of the video frame; or noise features are extracted through statistical pixel value distribution; or higher-order semantic features in the video frame are extracted using intermediate layers of a pre-trained convolutional neural network or a semantic encoder.

[0075] In step S330, based on the temporal sequence of each video frame, the target frame-level feature matrix is ​​obtained according to the frame-level features and frame-level quality score.

[0076] In this embodiment, the video frames are arranged in chronological order, and the frame-level features and frame-level quality scores of the video frames are arranged in chronological order to form a target frame-level feature matrix. The columns of the target frame-level feature matrix correspond to a single video frame, the rows of the target frame-level feature matrix correspond to the frame-level features and frame-level quality scores, and the matrix elements are the frame-level feature values ​​and frame-level quality scores of each video frame.

[0077] For example, if a video sample has 10 video frames, and each video frame contains one frame-level feature and one frame-level quality score, then the target frame-level feature matrix is ​​a two-dimensional feature matrix with 2 rows and 10 columns. The first row contains the frame-level quality scores arranged in the temporal order of the video frames, and the second row contains the frame-level features arranged in the temporal order of each video frame. This ensures that the target frame-level feature matrix fully preserves the temporal relationship, frame-level quality score, and frame-level features of the video frames.

[0078] In this embodiment, frame-level features corresponding to each video frame are introduced into the target frame-level feature matrix, enabling the model to learn the evolution rules of frame-level features between video frames, and thus more accurately determine the impact of temporal characteristics on the overall video quality.

[0079] Based on the above examples, this embodiment optimizes the target frame-level feature matrix determination process in step S230, as follows: Figure 4 As shown, the optimized target frame-level feature matrix determination process includes steps S410-S430.

[0080] In step S410, the image quality assessment model is used to evaluate each video frame in the video sample to obtain a frame-level quality score corresponding to each video frame.

[0081] The frame-level quality score determination process provided in this embodiment is the same as the frame-level quality score determination process in embodiment S230 above. For details, please refer to the description in the above embodiment.

[0082] In step S420, multi-dimensional feature extraction is performed on each video frame to obtain the frame-level features corresponding to each video frame in the corresponding dimension.

[0083] Multi-dimensional feature extraction can be understood as extracting features from multiple different dimensions of a video frame to comprehensively characterize its quality attributes. Frame-level features include frame-level motion features, frame-level attention region features, and frame-level semantic features.

[0084] For each video frame, features are extracted from multiple dimensions to obtain the frame-level features corresponding to each video frame in the corresponding dimensions.

[0085] In one possible implementation, multi-dimensional feature extraction is performed on each video frame to obtain frame-level features corresponding to each video frame in the corresponding dimension. This includes: calculating the frame-level motion features of each video frame using an optical flow field algorithm; extracting the frame-level attention region features of each video frame using an attention region detection model; and extracting the frame-level semantic features of each video frame using a semantic encoder.

[0086] Among them, the optical flow field algorithm is an algorithm that calculates the pixel motion trajectory between consecutive video frames to quantify the speed and direction of an object's motion on the image plane, and is used to capture the displacement relationship between video frames. Frame-level motion features refer to the features obtained based on the optical flow field algorithm that describe the motion state between the current video frame and adjacent video frames, including but not limited to: the magnitude, direction distribution, and motion region proportion of motion vectors. Frame-level motion features are used to reflect the intensity or smoothness of the motion of objects in the video.

[0087] A Region of Interest (ROI) is the most informative and attention-grabbing area in a video frame. For example, in portrait videos, this includes the face and hands; in traffic videos, it includes vehicles, pedestrians, and traffic lights. An attention region detection model is an algorithm based on deep learning (such as Convolutional Neural Networks (CNNs) and Transformers) used to automatically identify attention regions in video frames and calculate their area proportions. Furthermore, this area proportion is used as a frame-level attention region feature. Frame-level attention region features are the characteristics output by the attention region detection model that characterize salient regions within a single video frame, such as the area proportion of a face or the position of a vehicle, reflecting the quality of the salient regions in the video frame.

[0088] A semantic encoder is a model that can transform image content into high-level semantic features, which can be used to extract abstract meanings such as object categories and scene types from video frames. Frame-level semantic features refer to the features output by the semantic encoder that describe the meaning of the content of a single video frame, such as the quantified representation of abstract features like "contains a face" or "belongs to an indoor scene," reflecting the semantic attributes of the video frame's content.

[0089] In one exemplary implementation, an optical flow field algorithm is used to process consecutive video frames, calculating the motion vector of each pixel in adjacent frames, and then converting the vector information into frame-level motion features using statistical methods. For example, the average magnitude, directional entropy, and percentage of moving pixels of all motion vectors are calculated to ultimately form the frame-level motion features of the video frame.

[0090] In one exemplary implementation, an attention region detection model is used to analyze video frames, identify face regions in the video frames, calculate the area ratio of the face regions to the entire video frame, and obtain frame-level attention region features.

[0091] In one exemplary implementation, a pre-trained semantic encoder is used to extract features from a single video frame, and the feature vectors from the intermediate or output layers of the model are taken as frame-level semantic features. For example, the encoder identifies objects such as "car" and "road" in the video frame, transforming the semantic information into a fixed-dimensional vector. Each element in the vector corresponds to the confidence level of different semantic categories, forming frame-level semantic features.

[0092] In step S430, based on the temporal sequence of each video frame, the target frame-level feature matrix is ​​obtained according to the frame-level features and frame-level quality scores of each video frame in the corresponding dimension.

[0093] Arrange the video frames in temporal order, and then arrange the frame-level features and frame-level quality scores of each video frame in chronological order to form a target frame-level feature matrix. Each column of the target frame-level feature matrix corresponds to a single video frame, and each row corresponds to different dimensions of frame-level features and frame-level quality scores. The matrix elements are the frame-level feature values ​​and frame-level quality scores of each video frame in different dimensions. For example, assuming the feature dimension is M and the video length is N frames, an M×N frame-level feature matrix for that video frame can be obtained. It should be noted that the frame-level quality score is one of the feature dimensions.

[0094] For example, if a video sample includes 10 frames, and each video frame has 3 dimensions of frame-level features and 1 frame-level quality score, then the matrix is ​​a 4-row, 10-column feature matrix. The first row is the frame-level quality score arranged in the temporal order of the video frames, the second row is the frame-level operational features arranged in the temporal order of each video frame, the third row is the frame-level attention region features arranged in the temporal order of each video frame, and the fourth row is the frame-level semantic features arranged in the temporal order of each video frame. This ensures that the target frame-level feature matrix fully preserves the temporal relationship and multi-dimensional quality information of the video frames.

[0095] In this embodiment, multi-dimensional feature extraction captures the specific attribute features of video frames in different dimensions (such as operation, attention region, semantics, etc.). Based on the temporal relationship of video frames, it integrates multi-dimensional attribute features with frame-level quality scores, so that the target frame-level feature matrix not only fully preserves the quality details of individual video frames in multiple dimensions, but also naturally embeds the temporal dependencies between video frames through temporal arrangement, providing the model with rich quality signals that are closer to human subjective perception, thereby enabling the model to more accurately evaluate the overall quality of the video.

[0096] Based on the above examples, this embodiment optimizes the target frame-level feature matrix determination process in step S230, as follows: Figure 5 As shown, the optimized target frame-level feature matrix determination process includes steps S510-S550.

[0097] In step S510, the image quality assessment model is used to evaluate each video frame in the video sample to obtain a frame-level quality score corresponding to each video frame.

[0098] In step S520, feature extraction is performed on each video frame to obtain frame-level features corresponding to each video frame.

[0099] The frame-level quality score determination process and frame-level feature extraction process provided in this embodiment are the same as those in embodiments S310-S320 above. For details, please refer to the description in the above embodiments.

[0100] In step S530, based on the temporal sequence of each video frame, the frame-level features and frame-level quality scores are concatenated to obtain the first frame-level feature matrix.

[0101] In one possible implementation, based on the temporal sequence of each video frame, a first frame-level feature matrix is ​​obtained according to the frame-level features and frame-level quality scores of each video frame in the corresponding dimension. Specifically, all frames are arranged in temporal order, and the frame-level features and frame-level quality scores of the corresponding dimensions of the video frames are arranged in chronological order to form the first frame-level feature matrix. The columns of the first frame-level feature matrix correspond to a single video frame, and the rows of the first frame-level feature matrix correspond to frame-level features and frame-level quality scores in different dimensions. The matrix elements are the frame-level features and frame-level quality scores of each video frame in the corresponding dimension.

[0102] In step S540, inter-frame difference calculation is performed on the first frame-level feature matrix to obtain the difference feature matrix.

[0103] Inter-frame difference calculation refers to the quantization of the differences in feature values ​​across adjacent frames or frames at specified intervals based on the temporal relationship of video frames. This is used to capture the magnitude and trend of feature changes over time. Further, difference calculation includes first-order difference calculation and / or second-order difference calculation. First-order difference calculation involves directly interpolating the feature values ​​of consecutive video frames to quantify the magnitude of feature changes between adjacent frames, reflecting the instantaneous rate of change of features over time. Second-order difference calculation involves further interpolating adjacent first-order difference results based on the first-order difference results, quantifying the magnitude of the first-order difference changes, reflecting the trend of feature change rate, i.e., the acceleration of feature change.

[0104] The difference feature matrix is ​​a new matrix calculated through inter-frame difference. The elements of the difference feature matrix are the feature differences of the corresponding video frames in the same dimension in the first frame-level feature matrix. The matrix structure is consistent with the original matrix and is used to characterize the dynamic changes of features between video frames.

[0105] In this embodiment, the first-order inter-frame difference calculation is performed on the first-frame-level feature matrix to obtain the difference feature matrix, or the second-order inter-frame difference calculation is performed on the first-frame-level feature matrix to obtain the difference feature matrix.

[0106] In one possible implementation, for the frame-level features (including frame-level quality scores) arranged in chronological order in the first frame-level feature matrix, the difference between the two is calculated in each feature dimension. The differences in all dimensions are integrated into a first-order difference feature vector, which is then arranged in chronological order to form a first-order difference feature matrix. This matrix is ​​used to capture direct changes in features between frames, such as instantaneous increases or decreases in motion intensity and abrupt changes in attention regions.

[0107] In one possible implementation, the first-order difference feature matrix is ​​used as input, and adjacent first-order difference feature vectors are subjected to the same-dimensional difference operation again to generate second-order difference feature vectors, which are then arranged in temporal order to form a second-order difference feature matrix. Second-order difference calculation, by quantifying the magnitude of change in the first-order difference, can capture the trend of feature changes, such as a slowdown in the rate of motion intensity increase or an acceleration in the change of attention region. This further uncovers the dynamic evolution of video features in the temporal dimension, providing a more refined quantitative basis for temporal feature analysis.

[0108] In this embodiment, after performing inter-frame difference calculation on the first frame-level feature matrix to obtain the difference feature matrix, the target frame-level feature matrix can be determined based on the difference feature matrix. Determining the target frame-level feature matrix based on the difference feature matrix may include steps S550a or S550b.

[0109] In step S550a, the difference feature matrix is ​​used as the target frame-level feature matrix.

[0110] In this embodiment, the difference feature matrix is ​​directly used as the target frame-level feature matrix, and then used as the input to the video quality assessment model to be trained.

[0111] This application embodiment quantifies the differences between adjacent video frames in terms of features and scoring dimensions through inter-frame difference calculation, capturing the dynamic changes of video frames in the temporal dimension. The difference feature matrix is ​​then used as the target frame-level feature matrix, ensuring that the target frame-level feature matrix inherits the relevant information of the quality features in the original matrix while enhancing the temporal dynamic features through difference operations. This allows the target frame-level feature matrix to more comprehensively characterize the dual attributes of video quality in the spatial dimension (single-frame features) and the temporal dimension (inter-frame changes), providing richer input signals for subsequent video quality assessment models. This enables the model to perceive the evolution trend of frame-level features and quality scores over time, thereby improving the accuracy and robustness of the assessment results, making it suitable for time-sensitive video quality assessment scenarios.

[0112] In step S550b, the difference feature matrix is ​​concatenated with the first frame-level feature matrix to obtain the target frame-level feature matrix.

[0113] In one possible implementation, if the first frame-level feature matrix is ​​M×N and the difference feature matrix is ​​M×N, the frames need to be aligned first during splicing, and then merged according to the direction of the frame-level features to form a target frame-level feature matrix with a dimension of (2M×N).

[0114] In this matrix, the first M rows retain the static eigenvalues ​​of the original first-frame-level feature matrix, while the last M rows contain the dynamic changes of the difference feature matrix. This simultaneously covers the inherent features of a single video frame and the inter-frame change trends, providing a more comprehensive feature input for video quality assessment. This enables the model to more accurately perceive the evolution trend of frame-level features and quality scores over time, thereby improving the accuracy of the assessment results.

[0115] Based on the above examples, this embodiment optimizes the target frame-level feature matrix determination process in step S230, as follows: Figure 6 As shown, the optimized target frame-level feature matrix determination process includes steps S610-S640.

[0116] In step S610, the image quality assessment model is used to evaluate each video frame in the video sample to obtain a frame-level quality score corresponding to each video frame.

[0117] In step S620, feature extraction is performed on each video frame to obtain frame-level features corresponding to each video frame.

[0118] The frame-level quality score determination process and frame-level feature extraction process provided in this embodiment are the same as those in embodiments S310-S320 above. For details, please refer to the description in the above embodiments.

[0119] In step S630, abnormal frame detection is performed on each video frame to obtain an abnormal identifier value corresponding to each video frame. The abnormal identifier value is used to identify whether the corresponding video frame is an abnormal frame.

[0120] Abnormal frames are identified by algorithms or models as frames in a video that deviate from normal quality or content patterns. Examples include sudden blurring, screen tearing, and frozen frames. Anomaly flags are used to quantify whether a video frame is abnormal. For example, 0 represents a normal frame, and 1 represents an abnormal frame.

[0121] In one possible implementation, an anomaly detection algorithm is used to analyze video frames. For example, by calculating the deviation of features such as edge gradients and motion vectors of a frame from those of a normal frame, if the deviation exceeds a preset threshold, it is determined to be an abnormal frame and assigned an anomaly flag value of 1; otherwise, it is a normal frame and the anomaly flag value is 0.

[0122] In step S640, based on the temporal sequence of each video frame, the target frame-level feature matrix is ​​obtained according to the frame-level features, frame-level quality score, and anomaly identification value.

[0123] The video frames are arranged in chronological order. The frame-level features, frame-level quality scores, and anomaly identification values ​​of each video frame are arranged in chronological order to form a target frame-level feature matrix. The columns of the target frame-level feature matrix correspond to a single video frame, and the rows of the target frame-level feature matrix correspond to the frame-level features, frame-level quality scores, and anomaly identification values. The matrix elements are the frame-level feature values, frame-level quality scores, and anomaly identification values ​​of each video frame.

[0124] For example, if a video sample consists of 10 frames, and each video frame has 3 dimensions of frame-level features, 1 frame-level quality score, and 1 anomaly label value, then the matrix is ​​a 5-row, 10-column two-dimensional feature matrix. The first row is the frame-level quality score arranged in the temporal order of the video frames, the second row is the frame-level operational features arranged in the temporal order of each video frame, the third row is the frame-level attention region features arranged in the temporal order of each video frame, the fourth row is the frame-level semantic features arranged in the temporal order of each video frame, and the fifth row is the anomaly label value arranged in the temporal order of each video frame.

[0125] In this embodiment, the target frame-level feature matrix integrates anomaly identification values, frame-level features, and frame-level quality scores in a temporal sequence. This allows the target frame-level feature matrix to simultaneously contain the frame-level features, frame-level quality scores, and anomaly identification values ​​of the video frames. This preserves the details of individual video frames while reflecting the temporal impact of anomalies, enabling the model to learn from a small number of significant anomalies and improve the accuracy and robustness of the evaluation.

[0126] Based on the above examples, this embodiment optimizes the target frame-level feature matrix determination process in step S230, as follows: Figure 7 As shown, the optimized target frame-level feature matrix determination process includes steps S710-S730.

[0127] In step S710, the target video length corresponding to the training dataset is determined.

[0128] The target video length can be understood as a standardized video length for the training dataset, used to regulate the input length of training samples to ensure consistent sample structure during model training. For example, the target video length is the maximum video length, which can be determined by the number of frames included in the video.

[0129] In one possible implementation, the number of frames in all video samples in the training dataset is counted, and the maximum value of the sample length is selected as the target video length to avoid cropping the video samples, which could lead to the loss of important content.

[0130] In step S720, a second frame-level feature matrix is ​​obtained based on the temporal and frame-level quality scores of each video frame.

[0131] In one possible implementation, the frame-level quality scores of each video frame are integrated based on their temporal sequence to obtain a second frame-level feature matrix. Specifically, all frames are arranged in temporal order, and the frame-level quality scores of the video frames are arranged in chronological order to form the second frame-level feature matrix. The columns of the second frame-level feature matrix correspond to individual video frames, the rows correspond to the frame-level quality scores, and the matrix elements are the frame-level quality scores of each video frame.

[0132] In step S730, if the length of the video sample is less than the length of the target video, padding values ​​are used to fill the second frame-level feature matrix to obtain the target frame-level feature matrix.

[0133] The padding values ​​are used to supplement the preset values ​​of the number of columns in the second-frame-level feature matrix. For example, the padding values ​​can be 0, the mean, or a specific marker value. When the video sample length is insufficient, the padding is used in the second-frame-level feature matrix to match the target video length.

[0134] For example, if the video sample length is 8 frames and the target video length is 10 frames, then calculate T=2, add 2 padding values ​​to the end of the second frame-level feature matrix, adjust the dimension of the target frame-level feature matrix to 5×10, and form a standardized matrix that conforms to the target video length, where 5 is the feature dimension.

[0135] In this embodiment, model training requires consistent input data dimensions. If the video samples vary greatly in length, direct input will lead to dimension mismatch, causing training errors or parameter update chaos. By setting a target length and padding, the feature matrix dimensions of all samples are made uniform, providing the model with structurally consistent input and ensuring stable training.

[0136] In one possible implementation, the target frame-level feature matrix is ​​obtained by filling the second frame-level feature matrix with padding values, including steps S731 and S732.

[0137] In step S731, a mask vector is generated based on the temporal sequence of each video frame. The mask vector is used to identify whether the position of the corresponding video frame is obtained by padding. In step S732, the padded second frame-level feature matrix is ​​concatenated with the mask vector to obtain the target frame-level feature matrix.

[0138] Based on the temporal sequence of video frames, the positions of the original video frames are assigned a first-type label, such as 0, representing no padding; the positions of the padded frames are assigned a second-type label, such as 1, representing padding. For example, if the original length of the video sample is 8 frames and the target video length is 10 frames, and 2 frames need to be padded, then the mask vector is [0, 0, ..., 0 (8), 1, 1], and the length of the mask vector is the same as the length of the target video.

[0139] If the second frame-level feature matrix after padding is M×N and the mask vector is 1×N, then the dimension of the target frame-level feature matrix after stitching is (M+1)×N, that is, an additional row is added to the feature column of each video frame to identify whether it is a padding frame.

[0140] In this embodiment, the padding value is an artificial numerical value introduced to unify the input length of the model. It does not contain any valid information. If the model mistakenly treats it as a real frame feature, it will lead to learning bias. Using mask vectors to mark the padding position enables the model to distinguish between real video frames and padding values. This automatically weakens the weight of the padding part during training, reduces the interference of invalid information on parameter updates, and improves the accuracy of feature learning.

[0141] Based on all the above embodiments, this disclosure provides an example of a video quality assessment model training method, which mainly includes the following process:

[0142] Obtain B video samples, organize multiple users to watch each of the B video samples, and assign scores. For each video sample, use the average score from all users as the video quality label. Construct the correspondence between each video sample and its corresponding video quality label to form the training dataset. In other words, the training dataset includes B corresponding video samples and their video quality labels.

[0143] An image quality recognition model is used to score each video frame of the video sample to obtain the frame-level quality score of each video frame; an optical flow field algorithm is used to calculate the frame-level motion features of each video frame; an attention region detection model is used to extract the frame-level attention region features of each video frame; and a semantic encoder is used to extract the frame-level semantic features of each video frame.

[0144] The video frames are arranged chronologically, and their frame-level features and quality scores are arranged in chronological order to form a target frame-level feature matrix. Each column of the target frame-level feature matrix corresponds to a single video frame, and each row corresponds to different dimensions of frame-level features and quality scores. The matrix elements are the frame-level feature values ​​and quality scores for each video frame in different dimensions. Specifically, the first row contains frame-level quality scores arranged chronologically, the second row contains frame-level operational features arranged chronologically, the third row contains frame-level attention region features arranged chronologically, and the fourth row contains frame-level semantic features arranged chronologically. This ensures that the target frame-level feature matrix fully preserves the temporal relationships and multi-dimensional quality information of the video frames.

[0145] Furthermore, the frame-level features and frame-level quality scores of each dimension of the video frame are arranged in chronological order to form a target frame-level feature matrix. Inter-frame difference calculation is performed on the target frame-level feature matrix to obtain the target frame-level feature matrix.

[0146] Furthermore, abnormal frame detection is performed on each video frame to obtain an abnormal identifier value corresponding to each video frame. The abnormal identifier value is used to identify whether the corresponding video frame is an abnormal frame. The video frames are arranged in chronological order, and the abnormal identifier value of each video frame is added to the above target frame-level feature matrix in chronological order to form an M×N feature vector.

[0147] Since the videos vary in length, models that support variable-length sequences, such as LSTM / Transformer, can be used for training. For video feature samples of different lengths in the same batch, the length of the videos in the same batch can be padded to the length of the target video, and a mask vector can be added to indicate whether the video frame at the corresponding position is obtained by padding, in order to assist model training.

[0148] The sample size for each batch is B×(M+1)×N maxWhere B is the number of samples in the training batch, and N max , represents the number of the longest video frames in this batch, and M+1 represents the regular M-dimensional feature plus the mask vector.

[0149] Using B×(M+1)×N max The feature vectors are used to train the video quality assessment model to obtain the video quality assessment model.

[0150] Figure 8 This is a flowchart illustrating a video quality assessment method according to an exemplary embodiment. Figure 8 The method provided in the embodiments can be executed by any electronic device, such as the one described above. Figure 1 Terminal devices in or Figure 1 The server in, or Figure 1 The terminal devices and servers in the process execute together, but this disclosure does not limit this. For example... Figure 8 As shown, the video quality assessment method provided in this embodiment includes steps S810-S830.

[0151] In step S810, the video to be evaluated is acquired.

[0152] In this embodiment of the disclosure, a user's video quality assessment request is obtained. This video quality assessment request includes at least one element: the video to be assessed. The video to be assessed is the target video content to be evaluated.

[0153] In step S820, the video to be evaluated is input into the video quality evaluation model. The video quality evaluation model is trained using any of the video quality evaluation model training methods. The video quality evaluation model is used to evaluate the video to be evaluated and obtain a quality score for the video to be evaluated.

[0154] In step S830, the quality score of the video to be evaluated is output.

[0155] In this embodiment of the disclosure, the video to be evaluated in the video quality evaluation request is input into a video quality evaluation model trained by any of the aforementioned video quality evaluation model training methods. This video quality evaluation model can perform video quality evaluation on the video to be evaluated based on the model parameters adjusted during the aforementioned training process. The video quality evaluation model can consider the impact of temporal characteristics between frame-level quality scores on the overall video quality, thereby improving the video quality evaluation results.

[0156] The video quality assessment method provided in this disclosure inputs the video to be assessed into a video quality assessment model trained by the aforementioned video quality assessment model training method. Since this video quality assessment model can learn the mapping pattern between frame-level quality scores and the overall video quality score, it implicitly expresses the aggregation pattern of human judgment. Therefore, the quality score obtained based on this video quality assessment model can consider the impact of the temporal characteristics between frame-level quality scores on the overall video quality, thereby improving the video quality assessment results.

[0157] The following are embodiments of the apparatus disclosed herein, which can be used to execute embodiments of the method disclosed herein. For details not disclosed in the apparatus embodiments of this disclosure, please refer to the embodiments of the method disclosed herein.

[0158] Figure 9 This is a block diagram illustrating a video quality assessment model training apparatus according to an exemplary embodiment. (Refer to...) Figure 9 The device 900 may include: a sample acquisition module 910, a frame-level quality assessment module 920, a target frame-level feature matrix determination module 930, a video quality score determination module 940, and a model training module 950.

[0159] The sample acquisition module 910 is configured to acquire a training dataset of multiple video quality training samples, which includes corresponding video samples and video quality labels. The frame-level quality assessment module 920 is configured to evaluate each video frame in the video samples using an image quality assessment model to obtain a frame-level quality score for each video frame. The target frame-level feature matrix determination module 930 is configured to obtain a target frame-level feature matrix based on the temporal sequence and frame-level quality scores of each video frame. The video quality score determination module 940 is configured to perform quality assessment on the target frame-level feature matrix using the video quality assessment model to be trained to obtain a video quality score. The model training module 950 is configured to train the video quality assessment model to be trained based on the video quality score and video quality labels to obtain a video quality assessment model.

[0160] In some exemplary embodiments of this disclosure, the target frame-level feature matrix determination module 930 is further configured to extract features from each video frame to obtain frame-level features corresponding to each video frame; and to obtain a target frame-level feature matrix based on the temporal sequence of each video frame, according to the frame-level features and the frame-level quality score.

[0161] In some exemplary embodiments of this disclosure, the target frame-level feature matrix determination module 930 is further configured to perform multi-dimensional feature extraction on each video frame to obtain frame-level features corresponding to each video frame in the corresponding dimension; and to obtain a target frame-level feature matrix based on the temporal sequence of each video frame, according to the frame-level features and frame-level quality scores, including: obtaining a target frame-level feature matrix based on the temporal sequence of each video frame, according to the frame-level features and frame-level quality scores of each video frame in the corresponding dimension.

[0162] In some exemplary embodiments of this disclosure, the frame-level features include frame-level motion features, frame-level attention region features, and frame-level semantic features. The target frame-level feature matrix determination module 930 is further configured to calculate the frame-level motion features of each video frame using an optical flow field algorithm; extract the frame-level attention region features of each video frame using an attention region detection model; and extract the frame-level semantic features of each video frame using a semantic encoder.

[0163] In some exemplary embodiments of this disclosure, the target frame-level feature matrix determination module 930 is further configured to concatenate frame-level features and frame-level quality scores based on the temporal sequence of each video frame to obtain a first frame-level feature matrix; perform inter-frame difference calculation on the first frame-level feature matrix to obtain a difference feature matrix; and use the difference feature matrix as the target frame-level feature matrix.

[0164] In some exemplary embodiments of this disclosure, the target frame-level feature matrix determination module 930 is further configured to concatenate the difference feature matrix with the first frame-level feature matrix to obtain the target frame-level feature matrix.

[0165] In some exemplary embodiments of this disclosure, the target frame-level feature matrix determination module 930 is further configured to perform abnormal frame detection on each video frame to obtain an abnormal identification value corresponding to each video frame, wherein the abnormal identification value is used to identify whether the corresponding video frame is an abnormal frame; and to obtain the target frame-level feature matrix based on the temporal sequence of each video frame, according to the frame-level features, the frame-level quality score and the abnormal identification value.

[0166] In some exemplary embodiments of this disclosure, the target frame-level feature matrix determination module 930 is further configured to determine the target video length corresponding to the training dataset; obtain a second frame-level feature matrix based on the temporal and frame-level quality scores of each video frame; and if the length of the video sample is less than the target video length, fill the second frame-level feature matrix with padding values ​​to obtain the target frame-level feature matrix.

[0167] In some exemplary embodiments of this disclosure, the target frame-level feature matrix determination module 930 is further configured to generate a mask vector based on the temporal sequence of each video frame, the mask vector being used to identify whether the position of the corresponding video frame is filled; and to concatenate the filled second frame-level feature matrix with the mask vector to obtain the target frame-level feature matrix.

[0168] Figure 10 This is a block diagram illustrating a video quality assessment apparatus according to an exemplary embodiment. (Refer to...) Figure 10 The device 1000 may include: a video acquisition module 1010, a video input module 1020, and a quality scoring output module 1030.

[0169] The video acquisition module 1010 is configured to acquire the video to be evaluated; the video input module 1020 is configured to input the video to be evaluated into the video quality evaluation model; wherein the video quality evaluation model is trained by the video quality evaluation model training method as described in any one of claims 1 to 9, and the video quality evaluation model is used to evaluate the video to be evaluated to obtain a quality score for the video to be evaluated; the quality score output module 1030 is configured to output the quality score for the video to be evaluated.

[0170] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0171] The following reference Figure 11 To describe an electronic device 1100 according to such an embodiment of the present disclosure. Figure 11 The electronic device 1100 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.

[0172] like Figure 11 As shown, the electronic device 1100 is manifested in the form of a general-purpose computing device. The components of the electronic device 1100 may include, but are not limited to: at least one processing unit 1110, at least one storage unit 1120, a bus 1130 connecting different system components (including storage unit 1120 and processing unit 1110), and a display unit 1140.

[0173] The storage unit stores program code, which can be executed by the processing unit 1110 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processing unit 1110 can perform actions such as... Figure 2 The steps shown.

[0174] For example, electronic devices can achieve such Figure 2 or Figure 8 The steps shown.

[0175] Storage unit 1120 may include readable media in the form of volatile storage units, such as random access memory (RAM) 1121 and / or cache memory 1122, and may further include read-only memory (ROM) 1123.

[0176] Storage unit 1120 may also include a program / utility 1124 having a set (at least one) program module 1125, such program module 1125 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0177] Bus 1130 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0178] Electronic device 1100 can also communicate with one or more external devices 1170 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 1100, and / or with any device that enables electronic device 1100 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 1150. Furthermore, electronic device 1100 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 1160. As shown, network adapter 1160 communicates with other modules of electronic device 1100 via bus 1130. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 1100, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0179] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0180] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions that can be executed by a processor of the device to perform the described method. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0181] In an exemplary embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods described above.

[0182] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0183] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for training a video quality assessment model, characterized in that, include: Obtain a training dataset of multiple video quality training samples, wherein the video quality training samples include video samples and video quality labels that are in a corresponding relationship; For each video sample in the training dataset, an image quality assessment model is used to evaluate each video frame in the video sample to obtain a frame-level quality score corresponding to each video frame. Based on the temporal sequence of each video frame and the frame-level quality score, the target frame-level feature matrix is ​​obtained; The target frame-level feature matrix is ​​evaluated using a video quality assessment model to be trained, and a video quality score is obtained. The video quality assessment model is trained based on the video quality score and the video quality label to obtain the video quality assessment model.

2. The video quality assessment model training method according to claim 1, characterized in that, The process of obtaining the target frame-level feature matrix based on the temporal sequence of each video frame and the frame-level quality score includes: Feature extraction is performed on each video frame to obtain frame-level features corresponding to each video frame; Based on the temporal sequence of each video frame, the target frame-level feature matrix is ​​obtained according to the frame-level features and the frame-level quality score.

3. The video quality assessment model training method according to claim 2, characterized in that, The step of extracting features from each video frame to obtain frame-level features corresponding to each video frame includes: Multi-dimensional feature extraction is performed on each video frame to obtain the frame-level features corresponding to each video frame in the corresponding dimension; The step of obtaining the target frame-level feature matrix based on the temporal sequence of each video frame, according to the frame-level features and the frame-level quality score, includes: Based on the temporal sequence of each video frame, and according to the frame-level features of each video frame in the corresponding dimension and the frame-level quality score, the target frame-level feature matrix is ​​obtained.

4. The video quality assessment model training method according to claim 3, characterized in that, The frame-level features include frame-level motion features, frame-level attention region features, and frame-level semantic features. The step of extracting multi-dimensional features from each video frame to obtain frame-level features corresponding to each video frame in the corresponding dimension includes: The frame-level motion features of each video frame are calculated using an optical flow field algorithm; Frame-level attention region features of each video frame are extracted using an attention region detection model; The semantic encoder is used to extract the frame-level semantic features of each video frame.

5. The video quality assessment model training method according to claim 2, characterized in that, The step of obtaining the target frame-level feature matrix based on the temporal sequence of each video frame, according to the frame-level features and the frame-level quality score, includes: Based on the temporal sequence of each video frame, the frame-level features and the frame-level quality score are concatenated to obtain the first frame-level feature matrix; Inter-frame difference calculation is performed on the first frame-level feature matrix to obtain the difference feature matrix; The difference feature matrix is ​​used as the target frame-level feature matrix.

6. The video quality assessment model training method according to claim 5, characterized in that, After performing inter-frame difference calculation on the first frame-level feature matrix to obtain the difference matrix features, the method further includes: The difference feature matrix is ​​concatenated with the first frame-level feature matrix to obtain the target frame-level feature matrix.

7. The video quality assessment model training method according to claim 2, characterized in that, The step of obtaining the target frame-level feature matrix based on the temporal sequence of each video frame, according to the frame-level features and the frame-level quality score, includes: Anomaly detection is performed on each video frame to obtain anomaly identifier values ​​corresponding to each video frame. The anomaly identifier values ​​are used to identify whether the corresponding video frame is an anomaly frame. Based on the temporal sequence of each video frame, the target frame-level feature matrix is ​​obtained according to the frame-level features, the frame-level quality score, and the anomaly identifier value.

8. The video quality assessment model training method according to claim 1, characterized in that, The process of obtaining the target frame-level feature matrix based on the temporal sequence of each video frame and the frame-level quality score includes: Determine the target video length corresponding to the training dataset; Based on the temporal sequence of each video frame and the frame-level quality score, the first frame-level feature matrix is ​​obtained; If the length of the video sample is less than the length of the target video, then padding values ​​are used to fill the second frame-level feature matrix to obtain the target frame-level feature matrix.

9. The video quality assessment model training method according to claim 8, characterized in that, The step of filling the second frame-level feature matrix with padding values ​​to obtain the target frame-level feature matrix includes: Based on the timing of each video frame, a mask vector is generated, which is used to identify whether the position of the corresponding video frame is filled. The second frame-level feature matrix after padding is concatenated with the mask vector to obtain the target frame-level feature matrix.

10. A video quality assessment method, characterized in that, include: Obtain the video to be evaluated; The video to be evaluated is input into the video quality assessment model, wherein the video quality assessment model is trained by the video quality assessment model training method as described in any one of claims 1 to 9, and the video quality assessment model is used to perform quality assessment on the video to be evaluated to obtain a quality score for the video to be evaluated. Output the quality score of the video to be evaluated.

11. A video quality assessment model training device, characterized in that, include: The sample acquisition module is configured to acquire a training dataset of multiple video quality training samples, wherein the video quality training samples include video samples and video quality labels that are in a corresponding relationship. The frame-level quality assessment module is configured to evaluate each video frame in the video sample using an image quality assessment model for each video sample in the training dataset, and obtain a frame-level quality score corresponding to each video frame. The target frame-level feature matrix determination module is configured to obtain the target frame-level feature matrix based on the temporal sequence of each video frame and the frame-level quality score; The video quality scoring module is configured to perform quality assessment on the target frame-level feature matrix using a video quality assessment model to be trained, and obtain a video quality score. The model training module is configured to train the video quality assessment model to be trained based on the video quality score and the video quality label, thereby obtaining the video quality assessment model.

12. A video quality assessment device, characterized in that, include: The video acquisition module is configured to acquire the video to be evaluated; A video input module is configured to input the video to be evaluated into a video quality assessment model; wherein the video quality assessment model is trained by the video quality assessment model training method as described in any one of claims 1 to 9, and the video quality assessment model is used to evaluate the video to be evaluated to obtain a quality score for the video to be evaluated; The quality score output module is configured to output a quality score for the video to be evaluated.

13. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the executable instructions to implement the video quality assessment model training method as described in any one of claims 1 to 9 or the video quality assessment method as described in claim 10.

14. A computer-readable storage medium, wherein instructions in the computer-readable storage medium, when executed by a processor of an electronic device, enable the electronic device to perform the video quality assessment model training method as claimed in any one of claims 1 to 9 or the video quality assessment method as claimed in claim 10.

15. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the video quality assessment model training method as described in any one of claims 1 to 9 or the video quality assessment method as described in claim 10.