Video quality evaluation method and device
By using a target video quality assessment model, which assesses video quality based on regional attention information, the problem of existing technologies not considering user visual perception is solved, and a more reasonable, reliable and accurate video quality assessment is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies fail to effectively consider the characteristics of user visual perception in video quality assessment, resulting in unreasonable, unreliable, and inaccurate assessment results.
The target video quality assessment model evaluates quality based on regional attention information of video frames, focusing on key regions rather than all regions. It uses deep learning models such as Swin Transformer and U-Net to extract features and combines global features and regional attention for quality assessment.
This improves the reliability and accuracy of video quality assessment, conforms to the characteristics of user visual perception, and enhances the rationality and accuracy of assessment results.
Smart Images

Figure CN121789015A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and in particular to a video quality assessment method. It also relates to a video quality assessment method applied to cloud-side devices, a video quality assessment device, a video quality assessment device located on a cloud-side device, a computing device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] With the popularization of fifth-generation communication technology, video has become the mainstream information carrier, and it is widely used in many fields such as short videos, social platforms, office platforms, smart healthcare, and streaming media.
[0003] Generally, video quality directly impacts user experience, making its evaluation crucial. It's worth noting that users exhibit significant differences in their attention span across different areas of a video. For instance, in historical lecture videos, users tend to focus more on the speaker and key areas like subtitles, while paying less attention to areas like the background. Therefore, the visual perception characteristics of users should be considered when evaluating video quality. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a video quality assessment method. One or more embodiments of this specification also relate to a video quality assessment method applied to cloud-side devices, a video quality assessment device, a video quality assessment device located on a cloud-side device, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a video quality assessment method is provided, comprising: Obtain the video to be evaluated, wherein the video to be evaluated includes multiple video frames to be evaluated; The video to be evaluated is input into the target video quality assessment model to obtain the quality assessment result corresponding to the video to be evaluated. The quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be evaluated. The regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be evaluated. The regional attention information represents the importance of each region in the video frame to be evaluated to visual perception. The quality assessment result represents the quality assessment information of the video to be evaluated based on visual perception.
[0006] According to a second aspect of the embodiments of this specification, a video quality assessment method is provided, applied to a cloud-side device, comprising: The receiving end device sends a video to be evaluated, wherein the video to be evaluated includes multiple video frames to be evaluated; The video to be evaluated is input into the target video quality assessment model to obtain the quality assessment result corresponding to the video to be evaluated. The quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be evaluated. The regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be evaluated. The regional attention information represents the importance of each region in the video frame to be evaluated to visual perception. The quality assessment result represents the quality assessment information of the video to be evaluated based on visual perception. The quality assessment results are sent to the end-side device.
[0007] According to a third aspect of the embodiments of this specification, a video quality assessment apparatus is provided, comprising: The acquisition module is configured to acquire the video to be evaluated, wherein the video to be evaluated includes multiple video frames to be evaluated; The first input module is configured to input the video to be evaluated into a target video quality assessment model to obtain a quality assessment result corresponding to the video to be evaluated. The quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be evaluated. The regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be evaluated. The regional attention information characterizes the importance of each region in the video frame to be evaluated to visual perception. The quality assessment result characterizes the quality assessment information of the video to be evaluated based on visual perception.
[0008] According to a fourth aspect of the embodiments of this specification, a video quality assessment apparatus is provided, located on a cloud-side device, comprising: The receiving module is configured to receive a video to be evaluated sent by the end-side device, wherein the video to be evaluated includes multiple video frames to be evaluated; The second input module is configured to input the video to be evaluated into a target video quality assessment model to obtain a quality assessment result corresponding to the video to be evaluated. The quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be evaluated. The regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be evaluated. The regional attention information represents the importance of each region in the video frame to be evaluated to visual perception. The quality assessment result represents the quality assessment information of the video to be evaluated based on visual perception. The sending module is configured to send the quality assessment results to the end-side device.
[0009] According to a fifth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above method.
[0010] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0011] According to a seventh aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0012] The video quality assessment method provided in this specification can acquire a video to be assessed, wherein the video to be assessed includes multiple video frames to be assessed. The video to be assessed is input into a target video quality assessment model to obtain a quality assessment result corresponding to the video to be assessed. The quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be assessed. The regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be assessed. The regional attention information characterizes the importance of each region in the video frame to visual perception. The quality assessment result characterizes the quality assessment information of the video to be assessed based on visual perception.
[0013] One embodiment of this specification can perform video quality assessment based on the regional attention information of different regions in each video frame. This allows the video quality assessment to focus on key regions in the video frame, rather than all regions in the video frame, which conforms to visual perception characteristics and improves the reliability and accuracy of video quality assessment. Attached Figure Description
[0014] Figure 1 This is a flowchart illustrating a video quality assessment method provided in one embodiment of this specification; Figure 2 This is a schematic diagram of the structure of a target video quality assessment model provided in one embodiment of this specification; Figure 3 This is a schematic diagram of the structure of a target video quality assessment model provided in one embodiment of this specification; Figure 4 This is a schematic diagram of the structure of a target video quality assessment model provided in one embodiment of this specification; Figure 5This is a schematic diagram of the structure of a target video quality assessment model provided in one embodiment of this specification; Figure 6 This is a flowchart illustrating a training method for an initial video quality assessment model provided in one embodiment of this specification; Figure 7 This is a flowchart illustrating a video quality assessment method for cloud-side devices, provided in one embodiment of this specification. Figure 8 This is a schematic diagram of the structure of a video quality assessment device provided in one embodiment of this specification; Figure 9 This is a schematic diagram of the structure of a video quality assessment device located on a cloud-side device according to one embodiment of this specification; Figure 10 This is an architecture diagram of a video quality assessment system provided in one embodiment of this specification; Figure 11 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0015] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0016] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0017] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0018] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0019] The technical solutions provided in this application can employ deep learning models with relatively large parameter scales. However, this large model is merely an example; this application does not limit the number of model parameters supported by the deep learning model used, aiming to meet actual needs. The deep learning models involved in this application can be artificial intelligence-based language models (LM) or multimodal models (MM).
[0020] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0021] Video Quality Assessment (VQA): In this specification, VQA refers to the subjective visual quality score of a video predicted by an algorithm, which is consistent with the visual perception characteristics of the observer.
[0022] The Human Visual System (HVS) refers to the entire visual perception system in which the human eye and brain work together. Characteristics of visual perception include the differentiated attention paid to different areas of an image.
[0023] Regional importance refers to the degree to which different spatial regions within a video frame contribute to the overall perceived visual quality. Distortion in highly important regions has a greater impact on user experience than distortion in less important regions. Regional importance can be determined by factors such as visual attractiveness, semantic content (e.g., faces, text), motion states, or task relevance.
[0024] Swin Transformer: A hierarchical visual Transformer model based on window self-attention, which performs well in image feature extraction.
[0025] No-Reference Video Quality Assessment (NR-VQA): This refers to assessing the quality of a distorted video solely by analyzing the distorted video itself, without the availability of an original, undistorted reference video. The video quality assessment method using a video quality assessment model provided in this manual falls under the NR-VQA category.
[0026] In video-related fields, accurate video quality assessment is crucial for improving user experience and video service quality. Considering the visual perception characteristics of the human visual system—that is, the significant differences in attention paid to different areas within a video—user visual perception characteristics should not be ignored when assessing video quality.
[0027] Based on this, the video quality assessment method provided in this specification can acquire a video to be assessed, wherein the video to be assessed includes multiple video frames to be assessed. Then, a target video quality assessment model can be used to obtain the quality assessment result corresponding to the video to be assessed. The quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be assessed. The regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be assessed. The regional attention information characterizes the importance of each region in the video frame to visual perception. The quality assessment result characterizes the quality assessment information of the video to be assessed based on visual perception.
[0028] The above method assesses video quality based on the regional attention information of different areas in each video frame. This allows the assessment to focus on key areas within the video frame, rather than all areas, which aligns with visual perception characteristics and improves the reliability and accuracy of video quality assessment.
[0029] This specification provides a video quality assessment method. One or more embodiments of this specification also relate to a video quality assessment method applied to a cloud-side device, a video quality assessment apparatus, a video quality assessment apparatus located on a cloud-side device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0030] See Figure 1 , Figure 1 A flowchart of a video quality assessment method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0031] Step 102: Obtain the video to be evaluated, wherein the video to be evaluated includes multiple video frames to be evaluated.
[0032] It should be noted that the subject executing the technical solution in this specification can be any computing device with computing capabilities, such as a server or terminal, and this specification does not impose any specific restrictions on it.
[0033] In one or more embodiments of this specification, the computing device can acquire a video to be evaluated. The video to be evaluated includes multiple video frames to be evaluated.
[0034] In one or more embodiments of this specification, the computing device may provide a video quality assessment service, allowing other computing devices to send videos to be assessed to the computing device, which then receives the videos. Alternatively, the computing device may collect videos from its own applications as videos to be assessed for video quality evaluation. This specification does not limit the method of acquiring the videos to be assessed.
[0035] It should be noted that the multiple video frames to be evaluated are a continuous sequence of image frames with a specific frame rate. This manual does not limit the frame rate of the video to be evaluated, nor does it limit the number of video frames to be evaluated included in the video to be evaluated. These can be set according to actual needs.
[0036] Step 104: Input the video to be evaluated into the target video quality assessment model to obtain the quality assessment result corresponding to the video to be evaluated. The quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be evaluated. The regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be evaluated. The regional attention information represents the importance of each region in the video frame to be evaluated to visual perception. The quality assessment result represents the quality assessment information of the video to be evaluated based on visual perception.
[0037] As described in the background section, users exhibit significant differences in their attention to different areas within a video. Therefore, to ensure that the video quality assessment results better align with users' visual perception characteristics, thereby improving the reliability, rationality, and accuracy of the video quality assessment, in one or more embodiments of this specification, for each video frame to be evaluated, the corresponding area attention information can be determined.
[0038] The regional attention information represents the importance of each region in the video frame to be evaluated to visual perception. In one or more embodiments of this specification, the regional attention information corresponding to each video frame to be evaluated can be obtained through the regional attention prediction network in the target video quality assessment model.
[0039] It should be noted that regional attention information can include regional attention values for each region in the video frame to be evaluated. These regional attention values represent the degree of attention users pay to different regions, or in other words, the importance of different regions to users' visual perception. Alternatively, regional attention information can include importance values for different regions in the video frame to be evaluated; that is, this regional attention information can be regional importance information, which represents the importance of different regions to users' visual perception. In one or more embodiments of this specification, after obtaining the regional attention information corresponding to the video frame to be evaluated through the target video quality assessment model, the quality assessment result output by the target video quality assessment model can be obtained based on the regional attention information corresponding to each video frame to be evaluated.
[0040] The quality assessment result represents the quality assessment information of the video to be evaluated based on visual perception features; in other words, the quality assessment result is the assessment information that evaluates the merits of the video to be evaluated based on the user's visual perception dimension. The quality assessment result can reflect the degree of matching between the video to be evaluated and the user's visual perception.
[0041] In one or more embodiments of this specification, regional attention information includes regional attention values, and the quality assessment result is a quality assessment score. The video to be evaluated can first be input into the target video quality assessment model to obtain the regional attention values corresponding to the video frames to be evaluated. Then, the regional attention values corresponding to each video frame to be evaluated can be input into the target video quality assessment model to obtain the quality assessment score corresponding to that video.
[0042] In one or more embodiments of this specification, when obtaining the regional attention information corresponding to each video frame to be evaluated through the target video quality assessment model, it can be obtained through the regional attention prediction network in the target video quality assessment model. When obtaining the quality assessment score, it can be obtained based on the attention information corresponding to each video frame to be evaluated through the evaluation result prediction network in the target video quality assessment model. That is to say, the target video quality assessment model in this specification includes a regional attention prediction network and an evaluation result prediction network, such as... Figure 2 As shown, Figure 2 This diagram illustrates the structure of a target video quality assessment model according to one embodiment of this specification. The model includes a regional attention prediction network to generate regional attention values for each region within the video frame to be assessed, and an assessment result prediction network to perform regression prediction based on the regional attention values corresponding to each video frame to generate a quality assessment score for the video. Therefore, when determining the regional attention information for each video frame to be assessed, the video can be input into the regional attention prediction network of the target video quality assessment model to obtain the regional attention information for each video frame.
[0043] It should be noted that the regional attention prediction network in this target video quality assessment model can generate regional attention values for each region in the video frame to be evaluated. These regional attention values conform to the user's visual perception characteristics. In other words, the regional attention prediction network can generate larger regional attention values for regions with high user attention and relatively smaller regional attention values for regions with low user attention.
[0044] It should be noted that this target video quality assessment model is pre-trained based on sample videos and their corresponding quality assessment labels. The quality assessment labels are obtained by at least one user evaluating the quality of the sample videos. When training the initial target video quality assessment model, the sample videos can be input into a region attention prediction network to obtain the region attention information corresponding to each sample video frame. Then, the region attention information corresponding to each sample video frame is input into the quality assessment prediction network to obtain the predicted quality assessment result for the sample video. Based on the predicted quality assessment result and the quality assessment label, the loss can be calculated, and the initial target quality assessment model can be trained based on the loss until the training stopping condition is met, thus obtaining the target video quality assessment model.
[0045] Based on the foregoing, the trained target video quality assessment model possesses the ability to generate regional attention information corresponding to video frames, and also the ability to generate quality assessment results based on this regional attention information. Therefore, during the quality assessment of the video to be evaluated, the regional attention prediction network can generate the regional attention information corresponding to the video frames to be evaluated, and then the quality assessment prediction network can generate the quality assessment results based on this regional attention information.
[0046] It should be noted that when the region attention prediction network generates region attention information for each region in the video frame to be evaluated, the regions in the video frame to be evaluated can be divided according to objects, such as people, subtitles, backgrounds, etc. Therefore, the regions in the video frame to be evaluated can be divided into person regions, subtitle regions, background regions, etc. Furthermore, person regions can be further divided into facial regions, limb regions, etc., and background regions can be divided into key background regions and non-key background regions. Key background regions can be regions associated with the user or background regions that change over time, while non-key background regions can be fixed background regions, etc. The specific method of dividing the regions in the video frame to be evaluated and how to obtain the region attention information for each region in the video frame to be evaluated is not specifically limited here; this is a capability learned by the region attention prediction network through the aforementioned training process. Of course, the regions in the video frame to be evaluated can also be divided according to pixels or blocks.
[0047] Based on the above method, when conducting video quality assessment, the regional attention information corresponding to the video frame to be assessed is first determined by the target video quality assessment model. Then, based on the regional attention information, the video quality assessment result corresponding to the video to be assessed is obtained through the target video quality assessment model. When generating the quality assessment result, the target video quality assessment model focuses on the key regions in the video frame based on the regional attention information, rather than focusing on all regions in the video frame. In other words, different regions correspond to different regional attention values, which conforms to the characteristics of visual perception and improves the reliability and accuracy of video quality assessment.
[0048] Typically, when using machine learning models for video quality assessment, the results often have poor correlation with users' visual perception characteristics; in other words, these characteristics are not taken into account. Specifically, in reality, users differentiate their visual attention based on semantic importance, such as focusing more on facial features and subtitles, rather than the background. Machine learning models for video quality assessment, however, treat all regions in a video frame with equal weight, ignoring user visual perception features. This leads to the machine learning model being overly sensitive to minor distortions in non-critical areas, such as the background where users pay less attention, while failing to react well to severe distortions in truly important areas, such as facial features. Consequently, the quality assessment results are unreasonable, unreliable, and inaccurate. Therefore, in one or more embodiments of this specification, when training the initial target video quality assessment model to obtain the target video quality assessment model, the labels used are obtained by at least one user performing quality assessments on sample videos.
[0049] In one or more embodiments of this specification, the regional attention information corresponding to each video frame to be evaluated can be obtained through a target video quality assessment model, and the quality assessment result corresponding to the video to be evaluated can be obtained. The target video quality assessment model is pre-trained based on sample videos and their corresponding quality assessment labels, and the quality assessment labels are obtained by at least one user evaluating the sample videos. The target video quality assessment model provided in this specification predicts the regional attention information corresponding to the video frames to be evaluated and uses this as input data to obtain the quality assessment result. This allows the target video quality assessment model to focus on areas relatively important to the user's visual perception during the quality assessment result acquisition stage, making the target video quality model more sensitive to distortion changes in these areas. This results in quality assessment results that better match the user's visual perception characteristics, improving the rationality, reliability, and accuracy of the quality assessment results.
[0050] The target video quality assessment model in this specification may further include a global feature extraction network. In one or more embodiments of this specification, before inputting the regional attention information corresponding to each video frame to be evaluated into the assessment result prediction network to obtain the quality assessment result corresponding to the video to be evaluated, the model further includes: The video to be evaluated is input into the global feature extraction network to obtain the global features corresponding to each video frame to be evaluated. The region attention information corresponding to each video frame to be evaluated is input into the evaluation result prediction network to obtain the quality evaluation result corresponding to the video to be evaluated, including: The regional attention information and global features corresponding to each video frame to be evaluated are input into the evaluation result prediction network to obtain the quality evaluation result corresponding to the video to be evaluated.
[0051] By using the above method, a global feature extraction network is used to extract global features of the video to be evaluated. The global features are then combined with regional attention information to predict the subsequent quality assessment results. This can provide more information to the target video quality assessment model and improve the accuracy of the quality assessment results generated by the target video quality assessment model.
[0052] like Figure 3 As shown, Figure 3 This is a schematic diagram illustrating the structure of a target video quality assessment model provided in one embodiment of this specification. As can be seen, the target video quality assessment model includes a regional attention prediction network, a global feature extraction network, and an assessment result prediction network. The global feature extraction network includes a feature extraction subnetwork and a global perception subnetwork.
[0053] In one or more embodiments of this specification, the video to be evaluated is input into the global feature extraction network to obtain global features corresponding to each video frame to be evaluated, including: The video to be evaluated is input into the feature extraction subnet to obtain the first feature map corresponding to each video frame to be evaluated. The first feature map corresponding to each video frame to be evaluated is input into the global perception subnet to obtain the global features corresponding to each video frame to be evaluated.
[0054] Here, the first feature map refers to the spatial features of each video frame to be evaluated, i.e., spatial features. The feature extraction subnet can be a model or network used to extract the spatial features of the video frames to be evaluated.
[0055] In one or more embodiments of this specification, the feature extraction subnet can be a Swing Transformer. By inputting the video to be evaluated into the Swing Transformer, fine-grained spatial features corresponding to each video frame to be evaluated can be obtained, which can also obtain the first feature map corresponding to each video frame to be evaluated, as shown in Formula 1 below.
[0056] Formula 1 in, This represents the first feature map output by the Swing Transformer, and Swing-Backbone() represents the processing by the Swing Transformer. This represents the t-th frame of the video to be evaluated.
[0057] In practical applications, the first feature map corresponding to each video frame to be evaluated can also be input into the global perception subnetwork to obtain the global features corresponding to each video frame to be evaluated. This global perception subnetwork can perform global average pooling on the first feature map corresponding to each video frame to be evaluated to capture global quality features and obtain the global features corresponding to each video frame to be evaluated, as shown in Formula 2 below.
[0058] Formula 2 in, This represents the global feature corresponding to the t-th video frame to be evaluated. Let t be the first feature map corresponding to the video frame to be evaluated obtained by Equation 1, and GlobalAveragePooling() represents the processing of the global perception subnet.
[0059] Furthermore, such as Figure 4 As shown, Figure 4 This is a schematic diagram illustrating the structure of a target video quality assessment model provided in one embodiment of this specification. As can be seen, the target video quality assessment model includes a regional attention prediction network, a global feature extraction network, and an assessment result prediction network. The global feature extraction network includes a feature extraction subnetwork and a global perception subnetwork, and the regional attention prediction network includes an attention prediction subnetwork and a focus perception subnetwork.
[0060] In one or more embodiments of this specification, the video to be evaluated is input into the regional attention prediction network to obtain regional attention information corresponding to each video frame to be evaluated, including: The video to be evaluated is input into the attention prediction subnet to obtain the second feature map corresponding to each video frame to be evaluated. The second feature map and the first feature map corresponding to each video frame to be evaluated are input into the focus perception subnet to obtain the regional attention information corresponding to each video frame to be evaluated.
[0061] The second feature map refers to the importance map corresponding to the video frame to be evaluated. This importance map represents the degree of contribution of each spatial region in the video frame to visual perception.
[0062] In one or more embodiments of this specification, the attention prediction subnetwork can generate a pixel-level importance map for each video frame to be evaluated. That is, the attention prediction subnetwork can generate a corresponding importance value for each pixel (i.e., each region) in the video frame to be evaluated. This importance value refers to the degree of importance of the pixel's location to video quality evaluation based on visual perception. This importance map is the second feature map. Then, for each video frame to be evaluated, the first and second feature maps corresponding to that video frame can be input into the focus perception subnetwork. The focus perception subnetwork then performs weighted processing on the first feature map based on the second feature map to obtain the regional attention features corresponding to the video to be evaluated. These regional attention features are the regional attention values, or regional attention information.
[0063] Specifically, the focus-aware subnet can be used to perform biased feature encoding on regions in the video to be evaluated that contribute significantly to visual perception. Specifically, the focus-aware subnet uses the first feature map as a weight mask, then performs weighted processing on the second feature map (i.e., element-wise multiplication) to obtain the weighted second feature map corresponding to the video frame to be evaluated. Finally, global average pooling is performed on the weighted second feature map to obtain the focus feature vector, which is the region attention feature, as shown in Formulas 3 and 4 below.
[0064] Formula 3 in, The weighted second feature map representing the t-th video frame to be evaluated. The first feature map corresponding to the t-th video frame to be evaluated is obtained using Formula 1. The second feature map represents the video frame to be evaluated in frame t. This represents the Hadamard product, which is an element-wise multiplication operation.
[0065] It should be noted that when performing the multiplication calculation of the first feature map and the second feature map, it should be ensured that... and The spatial dimensions are the same, in and When the spatial dimensions are different, upsampling or downsampling can be used to make... and The spatial dimensions are the same. In one or more embodiments of this specification, it is possible to... Perform downsampling so that the next sampled Space dimensions and The spatial dimensions are the same. Of course, it is also possible to... Perform upsampling so that the upsampled Space dimensions and Since the spatial dimensions are the same, this instruction manual does not impose specific restrictions on this.
[0066] Formula 4 in, The focus feature vector representing the t-th video frame to be evaluated. The second feature map is a weighted representation of the video frame to be evaluated in frame t. GlobalAveragePooling() represents global average pooling.
[0067] The above method uses the second feature map as a weight to weight the first feature map, which can amplify the response of high-importance or high-attention regions in the global features. This enhances the representation of high-importance or high-attention regions in the global features, conforms to the characteristics of user visual perception, and can improve the rationality, reliability and accuracy of subsequent quality assessment.
[0068] Furthermore, in one or more embodiments of this specification, the attention prediction subnet may include encoder paths and decoder paths, i.e., downsampling paths and upsampling paths, wherein the encoder path includes sub-encoders of multiple scales, the decoder path includes sub-decoders of multiple scales, and the scales of the encoder path correspond one-to-one with the scales of the decoder path. Specifically, the attention prediction subnet may be a lightweight U-Net network.
[0069] The encoder path can extract multi-scale coded feature maps corresponding to the video frames to be evaluated through convolution and downsampling operations. For the sub-encoder of the j-th scale, the output feature map can be represented by Formula 5, which is shown below.
[0070] Formula 5 in, Let j be the encoded feature map output by the sub-encoder at the j-th scale. This is the encoded feature map output by the sub-encoder at the (j-1)th scale. Pool(ReLU(Conv())) represents the processing of the sub-encoder at the j-th scale.
[0071] The decoder path can recover the resolution and fuse feature detail information based on the multi-scale encoded feature map through upsampling and skip connection operations to obtain the multi-scale decoded feature map. Then, the multi-scale decoded feature map can be normalized through convolutional layers and the Sigmoid activation function to obtain the second feature map corresponding to the video to be evaluated, as shown in Formulas 6 and 7 below.
[0072] Formula Six in, Let j be the decoded feature map output by the sub-decoder at the j-th scale. Let be the encoded feature map output by the (j+1)th scale sub-encoder, and let Conv(Concat(Up(),)) represent the processing of the j-th scale sub-decoder. This is the encoded feature map output by the sub-encoder at the j-th scale.
[0073] Formula 7 in, The second feature map represents the video frame to be evaluated in frame t. () indicates the processing of a 1x1 convolutional layer, and Sigmoid() indicates the processing of the Sigmoid activation function. This represents the multi-scale decoding feature map corresponding to the t-th video frame to be evaluated.
[0074] It should be understood that the above description of the attention prediction subnet, which may include encoder and decoder paths, is only a simplified description. The working principle of the encoder and decoder paths in the attention prediction subnet is the same as that of the U-Net network. In other words, the attention prediction subnet in this specification is a U-Net network or a lightweight U-Net network. The specific data flow or workflow in the U-Net network and the lightweight U-Net network is a relatively mature technology, and this specification will not elaborate on it in detail.
[0075] Furthermore, the target video quality assessment model includes a regional attention prediction network, a global feature extraction network, and an assessment result prediction network, which includes a feature fusion subnetwork and a quality prediction subnetwork.
[0076] In one or more embodiments of this specification, the regional attention information corresponding to each video frame to be evaluated and the global features corresponding to each video frame to be evaluated are input into the evaluation result prediction network to obtain the quality evaluation result corresponding to the video to be evaluated, including: The regional attention information and global features corresponding to each video frame to be evaluated are input into the feature fusion subnet to obtain the fused features corresponding to each video frame to be evaluated. The fusion features corresponding to each video frame to be evaluated are input into the quality prediction subnet to obtain the quality evaluation result corresponding to the video to be evaluated.
[0077] In one or more embodiments of this specification, the feature fusion subnet is used to fuse global features and regional attention information, i.e., regional attention features.
[0078] In practical applications, frame-level feature fusion can be performed first, that is, fusing the global features and regional attention features corresponding to the video frame to be evaluated. Specifically, for each video frame to be evaluated, the regional attention features and global features corresponding to that video frame can be fused through a feature fusion subnet to obtain the fused features corresponding to that video frame. This allows the fused features for each video frame to be evaluated to be obtained.
[0079] In one or more embodiments of this specification, for each video frame to be evaluated, the regional attention features and global features corresponding to the video frame to be evaluated can be spliced together through a feature fusion subnet to achieve fusion, as shown in Formula 8 below.
[0080] Formula 8 in, Characterize the fusion features corresponding to the t-th video frame to be evaluated. Characterize the region attention features corresponding to the t-th video frame to be evaluated. represents the global features corresponding to the video frame to be evaluated in frame t, and Concat() represents the splicing process of the feature fusion subnet.
[0081] In practical applications, before inputting the fused features corresponding to each video frame to be evaluated into the temporal feature extraction subnet, learnable positional codes can be added to each fused feature to obtain the fused features after adding positional codes. This incorporates sequence order information, as shown in Formula 9 below.
[0082] Formula Nine in, Let P be the set of fused features corresponding to each video frame to be evaluated, where P is the position code. For each fused feature after adding position encoding.
[0083] In one or more embodiments of this specification, in order to capture dynamic patterns of quality changes over time, such as quality abrupt changes or quality fluctuations, a temporal feature extraction subnetwork may be set up. This temporal feature subnetwork captures the dynamic patterns of video frame quality changes over time. That is, the evaluation result prediction network also includes a temporal feature extraction subnetwork. Figure 5 As shown, Figure 5This is a schematic diagram of the structure of a target video quality assessment model provided in one embodiment of this specification. As can be seen, the target video quality assessment model includes a regional attention prediction network, a global feature extraction network, and an assessment result prediction network. The assessment result prediction network includes a feature fusion subnetwork, a temporal feature extraction subnetwork, and a quality prediction subnetwork.
[0084] Before inputting the fused features corresponding to each video frame to be evaluated into the quality prediction subnet to obtain the quality evaluation result corresponding to the video to be evaluated, the method further includes: The fused features corresponding to each video frame to be evaluated are input into the temporal feature extraction subnet to obtain the video features corresponding to the video to be evaluated. The fused features corresponding to each video frame to be evaluated are input into the quality prediction subnet to obtain the quality evaluation results corresponding to the video to be evaluated, including: The video features are input into the quality prediction subnet to obtain the quality assessment result corresponding to the video to be evaluated.
[0085] In one or more embodiments of this specification, the temporal feature extraction subnetwork can be a lightweight Transformer encoder to achieve temporal dynamic modeling of the fused features corresponding to each video frame to be evaluated. In this Transformer encoder, a multi-head self-attention mechanism allows the fused features corresponding to each video frame to be evaluated to interact with the fused features corresponding to other video frames to be evaluated, thereby learning global temporal dependencies. Furthermore, through feedforward networks and layer normalization processing, an output sequence can be obtained, improving the accuracy of subsequent quality assessment.
[0086] In one or more embodiments of this specification, the temporal feature extraction subnet may also perform average pooling on each fused feature after adding position encoding, thereby generating video features corresponding to the video to be evaluated, as shown in Formula 10 below.
[0087] Formula 10 in, The video features are defined by LN(), which performs layer normalization, and MHA(), which performs multi-head self-attention. The Mean() function represents the aggregation process for adding location-encoded fused features. Specifically, it can merge the values of multiple locations or features into a single value.
[0088] In practical applications, video features can be input into a quality prediction subnet to obtain the quality assessment result corresponding to the video to be evaluated. In one or more embodiments of this specification, the quality prediction subnet is used to perform regression prediction based on the video features of the video to be evaluated to obtain a quality assessment value, i.e., the quality assessment result. The quality prediction subnet can be a multilayer perceptron regression head, as shown in Formula 11 below.
[0089] Formula Eleven Where Q is the quality assessment value, i.e., the quality assessment result, and MLP() is the processing of the multilayer perceptron regression head, i.e., the quality prediction subnet. These are the video features corresponding to the video to be evaluated.
[0090] Using the methods described above, the target video quality assessment model provided in this specification predicts the regional attention information corresponding to the video frame to be evaluated, introduces a point-sensing network for importance graph weighting, and then uses this as input data to obtain the quality assessment results. This allows the target video quality assessment model to focus on regions relatively important to user visual perception during the quality assessment result acquisition stage, making the model more sensitive to distortion changes in these regions. This results in quality assessment results that better match user visual perception characteristics, improving the rationality, reliability, and accuracy of the quality assessment results. Furthermore, the global perception subnetwork, as a feature supplement, ensures that the target video quality assessment model does not lose its ability to perceive global distortions such as background quality, overall color cast, and blockiness, further improving the accuracy of the quality assessment results.
[0091] Furthermore, compared to complex and computationally expensive 3D convolutional networks such as SlowFast and I3D, the target video quality assessment model provided in this manual can effectively balance model structure and model accuracy, simplifying the model structure while ensuring the accuracy of the quality assessment results.
[0092] This manual also provides a method for training an initial video quality assessment model. For example... Figure 6 As shown, Figure 6 This is a flowchart illustrating a training method for an initial video quality assessment model, provided as an embodiment of this specification.
[0093] Step 602: Obtain the sample video and the corresponding quality assessment label of the sample video, wherein the sample video includes multiple sample video frames.
[0094] In one or more embodiments of this specification, the sample video and the corresponding quality assessment label of the sample video can be obtained from a sample database, such as KoNViD-1k or LIVE-VQC.
[0095] In one or more embodiments of this specification, obtaining the quality assessment label corresponding to the sample video includes: Obtain quality assessment values from at least two users who have evaluated the quality of the sample video. The average of all quality assessment values is used as the quality assessment label.
[0096] In other words, the computing device that trains the initial video quality assessment model can display each sample video to at least two users, or send the sample video to the computing devices corresponding to at least two users. Then, the users can use their corresponding computing devices to perform quality assessments on the sample video, that is, mark the quality assessment value corresponding to the sample video. The quality assessment value is then sent to the computing device that trains the initial video quality assessment model through the user's corresponding computing device. The computing device that trains the initial video quality assessment model can obtain the quality assessment values of the sample video performed by at least two users, and can calculate the average value of each quality assessment value. Then, the average value is used as the quality assessment label corresponding to the sample video.
[0097] It should be noted that this instruction manual does not limit the number of sample videos.
[0098] In practical applications, data augmentation techniques such as cropping and horizontal flipping can be used to generate more sample videos, thereby increasing the diversity of training data and improving the robustness of the trained target video quality assessment model.
[0099] Since the quality assessment labels used to train the initial video quality assessment model are obtained by different users assessing the quality of sample videos, the trained video quality assessment model can learn the patterns of different users assessing the quality of sample videos, predict the importance of each region in the sample video frame to the user's visual perception, and generate the final quality assessment result based on the importance of each region to the user's visual perception.
[0100] Step 604: Input the sample video into the regional attention prediction network in the initial video quality assessment model to obtain the sample regional attention information corresponding to each sample video frame.
[0101] Step 606: Input the attention information of the sample area into the evaluation result prediction network in the initial video quality assessment model to obtain the prediction result corresponding to the sample video.
[0102] Step 608: Based on the prediction results and the quality assessment labels, train the initial video quality assessment model until the model training stops, and obtain the target video quality assessment model.
[0103] In one or more embodiments of this specification, a loss can be determined based on the prediction results and quality assessment labels, and an initial video quality assessment model can be trained based on the loss.
[0104] It should be noted that this specification does not impose specific restrictions on this loss; it may be the mean square error loss or the L1 loss, which can be determined according to actual needs.
[0105] In one or more embodiments of this specification, the model training stopping conditions include, but are not limited to, the loss being less than a preset loss, the number of sample videos used being greater than a preset number, the number of iterations being greater than a preset number, etc.
[0106] It should be understood that the structure of the initial video quality assessment model can be consistent with the above, that is, it also includes a global feature extraction network, which may include a feature extraction subnetwork and a global perception subnetwork. The regional attention prediction network may include an attention prediction subnetwork and a focus perception subnetwork. The evaluation result prediction network may also include a temporal feature extraction subnetwork. When the structure of the initial video quality assessment model is as described above, these networks should be adjusted during training. The specific data transfer process will not be described here.
[0107] Using the above method, the trained target video quality assessment model can predict the regional attention information corresponding to the video frames to be evaluated. A point-sensing network is introduced for importance graph weighting, which is then used as input data to obtain the quality assessment results. This allows the target video quality assessment model to focus on regions relatively important to user visual perception during the quality assessment result acquisition stage. This makes the target video quality model more sensitive to distortion changes in these regions, resulting in quality assessment results that better match user visual perception characteristics and improving the reasonableness, reliability, and accuracy of the quality assessment results. Furthermore, the global perception subnetwork, as a feature supplement, ensures that the target video quality assessment model does not lose its ability to perceive global distortions such as background quality, overall color cast, and blockiness, further improving the accuracy of the quality assessment results.
[0108] Furthermore, compared to complex and computationally expensive 3D convolutional networks such as SlowFast and I3D, the target video quality assessment model provided in this manual can effectively balance model structure and model accuracy, simplifying the model structure while ensuring the accuracy of the quality assessment results.
[0109] Corresponding to the above method embodiments, this specification also provides a video quality assessment method, applied to cloud-side devices. Figure 7 A flowchart illustrating a video quality assessment method for cloud-side devices, as provided in one embodiment of this specification, is shown. Figure 7 As shown, it includes: Step 702: Receive the video to be evaluated sent by the receiving end-side device, wherein the video to be evaluated includes multiple video frames to be evaluated.
[0110] Step 704: Input the video to be evaluated into the target video quality assessment model to obtain the quality assessment result corresponding to the video to be evaluated. The quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be evaluated. The regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be evaluated. The regional attention information represents the importance of each region in the video frame to be evaluated to visual perception. The quality assessment result represents the quality assessment information of the video to be evaluated based on visual perception.
[0111] Step 706: Send the quality assessment results to the end-side device.
[0112] It should be noted that the specific process of determining the regional attention information corresponding to each video frame to be evaluated and obtaining the quality assessment result corresponding to the video to be evaluated based on the regional attention information corresponding to each video frame to be evaluated is consistent with that described in the aforementioned video quality assessment method, and will not be repeated here.
[0113] Based on the above method, a video to be evaluated can be obtained, comprising multiple video frames to be evaluated. The video to be evaluated is then input into a target video quality assessment model to obtain a quality assessment result. This result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame. The regional attention information is determined by the target video quality assessment model based on at least one region within each video frame. The regional attention information characterizes the importance of each region in the video frame to visual perception. The quality assessment result represents the quality assessment information of the video based on visual perception. Therefore, the above method, when performing video quality assessment, can focus on key regions within a video frame, rather than all regions within the video frame. In other words, different regions correspond to different regional attention values, which aligns with visual perception characteristics and improves the reliability and accuracy of video quality assessment.
[0114] Corresponding to the above method embodiments, this specification also provides an embodiment of a video quality assessment device. Figure 8 A schematic diagram of a video quality assessment device according to one embodiment of this specification is shown. Figure 8 As shown, the device includes: The acquisition module 802 is configured to acquire a video to be evaluated, wherein the video to be evaluated includes multiple video frames to be evaluated; The first input module 804 is configured to input the video to be evaluated into a target video quality assessment model to obtain a quality assessment result corresponding to the video to be evaluated. The quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be evaluated. The regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be evaluated. The regional attention information characterizes the importance of each region in the video frame to be evaluated to visual perception. The quality assessment result characterizes the quality assessment information of the video to be evaluated based on visual perception.
[0115] Optionally, the first input module 804 is further configured to input the video to be evaluated into the regional attention prediction network to obtain regional attention information corresponding to each video frame to be evaluated. The first input module 804 is further configured to input the regional attention information corresponding to each video frame to be evaluated into the evaluation result prediction network to obtain the quality evaluation result corresponding to the video to be evaluated.
[0116] Optionally, the target video quality assessment model further includes a global feature extraction network; The first input module 804 is further configured to input the video to be evaluated into the global feature extraction network to obtain global features corresponding to each video frame to be evaluated; The first input module 804 is further configured to input the regional attention information corresponding to each video frame to be evaluated and the global features corresponding to each video frame to be evaluated into the evaluation result prediction network to obtain the quality evaluation result corresponding to the video to be evaluated.
[0117] Optionally, the global feature extraction network includes a feature extraction subnetwork and a global perception subnetwork; The first input module 804 is further configured to input the video to be evaluated into the feature extraction subnetwork to obtain a first feature map corresponding to each video frame to be evaluated; and input the first feature map corresponding to each video frame to be evaluated into the global perception subnetwork to obtain global features corresponding to each video frame to be evaluated.
[0118] Optionally, the regional attention prediction network includes an attention prediction subnetwork and a focus perception subnetwork; The first input module 804 is further configured to input the video to be evaluated into the attention prediction subnet to obtain a second feature map corresponding to each video frame to be evaluated. The second feature map and the first feature map corresponding to each video frame to be evaluated are input into the focus perception subnet to obtain the regional attention information corresponding to each video frame to be evaluated.
[0119] Optionally, the evaluation result prediction network includes a feature fusion subnetwork and a quality prediction subnetwork; The first input module 804 is further configured to input the regional attention information and global features corresponding to each video frame to be evaluated into the feature fusion subnetwork to obtain the fusion features corresponding to each video frame to be evaluated; and input the fusion features corresponding to each video frame to be evaluated into the quality prediction subnetwork to obtain the quality evaluation result corresponding to the video to be evaluated.
[0120] Optionally, the evaluation result prediction network further includes a temporal feature extraction subnetwork; The first input module 804 is further configured to input the fusion features corresponding to each video frame to be evaluated into the temporal feature extraction subnet to obtain the video features corresponding to the video to be evaluated; The first input module 804 is further configured to input the video features into the quality prediction subnet to obtain the quality assessment result corresponding to the video to be evaluated.
[0121] Optionally, the device further includes a training module configured to acquire sample videos and corresponding quality assessment labels, wherein the sample videos include multiple sample video frames; input the sample videos into the regional attention prediction network in the initial video quality assessment model to obtain sample regional attention information corresponding to each sample video frame; input the sample regional attention information into the evaluation result prediction network in the initial video quality assessment model to obtain the prediction result corresponding to the sample videos; and train the initial video quality assessment model based on the prediction result and the quality assessment labels until the model training stopping condition is met to obtain the target video quality assessment model.
[0122] Optionally, the training module is further configured to obtain quality assessment values from at least two users who evaluate the sample video; and to use the average of the quality assessment values as the quality assessment label.
[0123] Based on the above apparatus, a video to be evaluated can be acquired, wherein the video to be evaluated includes multiple video frames to be evaluated; the video to be evaluated is input into a target video quality evaluation model to obtain a quality evaluation result corresponding to the video to be evaluated, wherein the quality evaluation result is determined by the target video quality evaluation model based on the regional attention information corresponding to each video frame to be evaluated, the regional attention information is determined by the target video quality evaluation model based on at least one region in the video frame to be evaluated, the regional attention information characterizes the importance of each region in the video frame to be evaluated to visual perception, and the quality evaluation result characterizes the quality evaluation information of the video to be evaluated based on visual perception.
[0124] As can be seen, the above method can focus on key areas in video frames rather than all areas in the video frame when conducting video quality assessment, which is in line with the characteristics of visual perception and improves the reliability and accuracy of video quality assessment.
[0125] The above is a schematic scheme of a video quality assessment device according to this embodiment. It should be noted that the technical solution of this video quality assessment device and the technical solution of the video quality assessment method described above belong to the same concept. For details not described in detail in the technical solution of the video quality assessment device, please refer to the description of the technical solution of the video quality assessment method described above.
[0126] Corresponding to the above method embodiments, this specification also provides an embodiment of a video quality assessment device located on a cloud-side device. Figure 9 This specification illustrates a schematic diagram of a video quality assessment device located on a cloud-side device according to one embodiment. Figure 9 As shown, the device includes: The receiving module 902 is configured to receive a video to be evaluated sent by the end-side device, wherein the video to be evaluated includes multiple video frames to be evaluated. The second input module 904 is configured to input the video to be evaluated into a target video quality assessment model to obtain a quality assessment result corresponding to the video to be evaluated. The quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be evaluated. The regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be evaluated. The regional attention information characterizes the importance of each region in the video frame to be evaluated to visual perception. The quality assessment result characterizes the quality assessment information of the video to be evaluated based on visual perception. The sending module 906 is configured to send the quality assessment result to the end-side device.
[0127] Based on the above-described device, the cloud-side device can acquire a video to be evaluated, wherein the video to be evaluated includes multiple video frames to be evaluated; the video to be evaluated is input into a target video quality assessment model to obtain a quality assessment result corresponding to the video to be evaluated, wherein the quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be evaluated, the regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be evaluated, the regional attention information characterizes the importance of each region in the video frame to be evaluated for visual perception, and the quality assessment result characterizes the quality assessment information of the video to be evaluated based on visual perception.
[0128] As can be seen, the above method can focus on key areas in video frames rather than all areas in the video frame when conducting video quality assessment, which is in line with the characteristics of visual perception and improves the reliability and accuracy of video quality assessment.
[0129] The above is a schematic scheme of a video quality assessment device located on a cloud-side device according to this embodiment. It should be noted that the technical solution of this video quality assessment device located on a cloud-side device belongs to the same concept as the technical solution of the video quality assessment method located on a cloud-side device described above. For details not described in detail in the technical solution of the video quality assessment device located on a cloud-side device, please refer to the description of the technical solution of the video quality assessment method located on a cloud-side device described above.
[0130] See Figure 10 , Figure 10 This specification illustrates an architecture diagram of a video quality assessment system according to one embodiment of the present specification. The video quality assessment system may include a client 100 and a server 200. Client 100 is used to send a video to be evaluated to server 200, wherein the video to be evaluated includes multiple video frames to be evaluated; Server 200 is used to receive the video to be evaluated, input the video to be evaluated into a target video quality evaluation model, and obtain the quality evaluation result corresponding to the video to be evaluated. The quality evaluation result is determined by the target video quality evaluation model based on the regional attention information corresponding to each frame of the video to be evaluated. The regional attention information is determined by the target video quality evaluation model based on at least one region in the frame of the video to be evaluated. The regional attention information characterizes the importance of each region in the frame of the video to be evaluated to visual perception. The quality evaluation result characterizes the quality evaluation information of the video to be evaluated based on visual perception. Server 200 sends the quality evaluation result corresponding to the video to be evaluated to client 100. Client 100 is also used to receive the quality assessment result corresponding to the video to be evaluated sent by server 200.
[0131] Applying the scheme of the embodiments in this specification, the server 200 can obtain a video to be evaluated, wherein the video to be evaluated includes multiple video frames to be evaluated; the video to be evaluated is input into a target video quality assessment model to obtain a quality assessment result corresponding to the video to be evaluated, wherein the quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be evaluated, the regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be evaluated, the regional attention information characterizes the importance of each region in the video frame to visual perception, and the quality assessment result characterizes the quality assessment information of the video to be evaluated based on visual perception. It can be seen that the above method, when performing video quality assessment, can focus on key regions in the video frame, rather than focusing on all regions in the video frame, thereby improving the reliability and accuracy of video quality assessment.
[0132] The video quality assessment system may include multiple clients 100 and a server 200. Clients 100 can be referred to as edge devices, and server 200 can be referred to as cloud devices. Multiple clients 100 can establish communication connections through server 200. In the video quality assessment scenario, server 200 is used to provide video quality assessment services between multiple clients 100. Each client 100 can act as a sender or receiver, communicating through server 200.
[0133] Users can interact with server 200 through client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In a video quality assessment scenario, users can publish data streams to server 200 through client 100, and server 200 can generate quality assessment results for the video to be assessed based on the data stream, and push the quality assessment results for the video to be assessed to other clients that have established communication.
[0134] In this system, client 100 and server 200 establish a connection via a network. The network provides the medium for communication between client 100 and server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. Data transmitted by client 100 may need to undergo encoding, transcoding, compression, or other processing before being published to server 200.
[0135] Client 100 can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. Client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by server 200, such as a real-time communication (RTC) SDK. Client 100 can be deployed on a computing device and depends on the device or certain apps on the device to run. The computing device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. Various other types of applications can also be configured on the computing device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.
[0136] Server 200 may include servers providing various services, such as servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that server 200 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0137] It is worth noting that the video quality assessment method provided in the embodiments of this specification is generally executed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server, thereby executing the video quality assessment method provided in the embodiments of this specification. In other embodiments, the video quality assessment method provided in the embodiments of this specification may also be executed jointly by the client and the server.
[0138] Figure 11A structural block diagram of a computing device 1100 according to an embodiment of this application is shown. The components of the computing device 1100 include, but are not limited to, a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 1130, and a database 1150 is used to store data.
[0139] The computing device 1100 also includes an access device 1140, which enables the computing device 1100 to communicate via one or more networks 1160. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 1140 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0140] In one embodiment of this application, the aforementioned components of the computing device 1100 and Figure 11 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 11 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.
[0141] The computing device 1100 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1100 can also be a mobile or stationary server.
[0142] The processor 1120 is configured to execute a computer program / instruction that, when executed by the processor, implements the steps of the video quality assessment method described above.
[0143] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the video quality assessment method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the video quality assessment method described above.
[0144] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the video quality assessment method described above.
[0145] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. In particular, the computer-readable storage medium embodiments are described simply because they are substantially similar to the video quality assessment method embodiments; relevant details can be found in the descriptions of the video quality assessment method embodiments.
[0146] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the video quality assessment method described above.
[0147] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the video quality assessment method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the video quality assessment method described above.
[0148] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0149] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0150] It should be noted that the above description describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0151] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0152] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A video quality assessment method, comprising: Obtain the video to be evaluated, wherein the video to be evaluated includes multiple video frames to be evaluated; The video to be evaluated is input into the target video quality assessment model to obtain the quality assessment result corresponding to the video to be evaluated. The quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be evaluated. The regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be evaluated. The regional attention information represents the importance of each region in the video frame to be evaluated to visual perception. The quality assessment result represents the quality assessment information of the video to be evaluated based on visual perception.
2. The method as described in claim 1, wherein the target video quality assessment model includes a regional attention prediction network and an assessment result prediction network; Based on at least one region in the video frame to be evaluated, determine the region attention information corresponding to each video frame to be evaluated, including: The video to be evaluated is input into the regional attention prediction network to obtain the regional attention information corresponding to each video frame to be evaluated. The quality assessment result of the video to be evaluated is determined based on the regional attention information corresponding to each video frame to be evaluated, including: The regional attention information corresponding to each video frame to be evaluated is input into the evaluation result prediction network to obtain the quality evaluation result corresponding to the video to be evaluated.
3. The method as described in claim 2, wherein the target video quality assessment model further includes a global feature extraction network; Before inputting the regional attention information corresponding to each video frame to be evaluated into the evaluation result prediction network to obtain the quality evaluation result corresponding to the video to be evaluated, the method further includes: The video to be evaluated is input into the global feature extraction network to obtain the global features corresponding to each video frame to be evaluated. The region attention information corresponding to each video frame to be evaluated is input into the evaluation result prediction network to obtain the quality evaluation result corresponding to the video to be evaluated, including: The regional attention information and global features corresponding to each video frame to be evaluated are input into the evaluation result prediction network to obtain the quality evaluation result corresponding to the video to be evaluated.
4. The method as described in claim 3, wherein the global feature extraction network comprises a feature extraction subnetwork and a global perception subnetwork; The video to be evaluated is input into the global feature extraction network to obtain the global features corresponding to each video frame to be evaluated, including: The video to be evaluated is input into the feature extraction subnet to obtain the first feature map corresponding to each video frame to be evaluated. The first feature map corresponding to each video frame to be evaluated is input into the global perception subnet to obtain the global features corresponding to each video frame to be evaluated.
5. The method as described in claim 4, wherein the regional attention prediction network comprises an attention prediction subnetwork and a focus perception subnetwork; The video to be evaluated is input into the regional attention prediction network to obtain regional attention information corresponding to each video frame to be evaluated, including: The video to be evaluated is input into the attention prediction subnet to obtain the second feature map corresponding to each video frame to be evaluated. The second feature map and the first feature map corresponding to each video frame to be evaluated are input into the focus perception subnet to obtain the regional attention information corresponding to each video frame to be evaluated.
6. The method of claim 3, wherein the evaluation result prediction network comprises a feature fusion subnetwork and a quality prediction subnetwork; The regional attention information and global features corresponding to each video frame to be evaluated are input into the evaluation result prediction network to obtain the quality evaluation result corresponding to the video to be evaluated, including: The regional attention information and global features corresponding to each video frame to be evaluated are input into the feature fusion subnet to obtain the fused features corresponding to each video frame to be evaluated. The fusion features corresponding to each video frame to be evaluated are input into the quality prediction subnet to obtain the quality evaluation result corresponding to the video to be evaluated.
7. The method of claim 6, wherein the evaluation result prediction network further includes a temporal feature extraction subnetwork; Before inputting the fused features corresponding to each video frame to be evaluated into the quality prediction subnet to obtain the quality evaluation result corresponding to the video to be evaluated, the method further includes: The fused features corresponding to each video frame to be evaluated are input into the temporal feature extraction subnet to obtain the video features corresponding to the video to be evaluated. The fused features corresponding to each video frame to be evaluated are input into the quality prediction subnet to obtain the quality evaluation results corresponding to the video to be evaluated, including: The video features are input into the quality prediction subnet to obtain the quality assessment result corresponding to the video to be evaluated.
8. The method of claim 2, wherein the target video quality assessment model is obtained using the following method: Obtain sample videos and their corresponding quality assessment tags, wherein, The sample video includes multiple sample video frames; The sample videos are input into the regional attention prediction network in the initial video quality assessment model to obtain the sample regional attention information corresponding to each sample video frame. The attention information of the sample area is input into the evaluation result prediction network in the initial video quality assessment model to obtain the prediction result corresponding to the sample video. Based on the prediction results and the quality assessment labels, an initial video quality assessment model is trained until the model training stops, thus obtaining the target video quality assessment model.
9. The method as described in claim 8, wherein obtaining the quality assessment label corresponding to the sample video includes: Obtain quality assessment values from at least two users who have evaluated the quality of the sample video. The average of all quality assessment values is used as the quality assessment label.
10. A video quality assessment method, applied to cloud-side devices, comprising: The receiving end device sends a video to be evaluated, wherein the video to be evaluated includes multiple video frames to be evaluated; The video to be evaluated is input into the target video quality assessment model to obtain the quality assessment result corresponding to the video to be evaluated. The quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be evaluated. The regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be evaluated. The regional attention information represents the importance of each region in the video frame to be evaluated to visual perception. The quality assessment result represents the quality assessment information of the video to be evaluated based on visual perception. The quality assessment results are sent to the end-side device.
11. A video quality assessment device, comprising: The acquisition module is configured to acquire the video to be evaluated, wherein the video to be evaluated includes multiple video frames to be evaluated; The first input module is configured to input the video to be evaluated into a target video quality assessment model to obtain a quality assessment result corresponding to the video to be evaluated. The quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be evaluated. The regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be evaluated. The regional attention information characterizes the importance of each region in the video frame to be evaluated to visual perception. The quality assessment result characterizes the quality assessment information of the video to be evaluated based on visual perception.
12. A video quality assessment device, located on a cloud-side device, comprising: The receiving module is configured to receive a video to be evaluated sent by the end-side device, wherein the video to be evaluated includes multiple video frames to be evaluated. The second input module is configured to input the video to be evaluated into a target video quality assessment model to obtain a quality assessment result corresponding to the video to be evaluated. The quality assessment result is determined by the target video quality assessment model based on the regional attention information corresponding to each video frame to be evaluated. The regional attention information is determined by the target video quality assessment model based on at least one region in the video frame to be evaluated. The regional attention information represents the importance of each region in the video frame to be evaluated to visual perception. The quality assessment result represents the quality assessment information of the video to be evaluated based on visual perception. The sending module is configured to send the quality assessment results to the end-side device.
13. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 10.
14. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.
15. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.