Ultrahigh-definition video visual quality evaluation method based on adaptive space fusion and time selection

By adopting adaptive spatial fusion and time selection methods in ultra-high-definition video quality evaluation, the problems of dimensional disasters and high computing resource requirements in the prior art are solved, and efficient video quality evaluation is achieved.

CN120151501APending Publication Date: 2025-06-13SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510085274.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art faces dimensional disasters and high computing resource requirements when evaluating the visual quality of ultra-high-definition video.

Method used

Using a method based on adaptive spatial fusion and time selection, the spatial adaptive fusion network fusion block-by-block feature is obtained, and the features of several time positions are selected through the time adaptive selection network for series connection, and mapped to the prediction score of video quality.

Benefits of technology

It effectively avoids dimensional disasters caused by information loss, and adaptively selects a small number of time positions, reducing the computing resource requirements and improving the efficiency of video quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151501A_ABST
    Figure CN120151501A_ABST
Patent Text Reader

Abstract

The invention discloses an ultra-high-definition video visual quality assessment method based on adaptive space fusion and time selection. The method comprises the following steps: acquiring an ultra-high-definition video to be assessed; performing intra-frame processing on each frame of the ultra-high-definition video through a spatial adaptive fusion network so as to fuse the block-by-block features and obtain frame-by-frame features; selecting a plurality of time positions from the frame-by-frame features through a time adaptive selection network, connecting the features corresponding to the selected time positions in series, and mapping the serial features to a prediction score of video quality; and outputting a visual quality evaluation result of the ultra-high-definition video. According to the method, the dimensionality disaster is avoided by adaptively fusing all the block-by-block features, and the computing resource demand is reduced by adaptively selecting a small amount of time and position.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video quality evaluation, and particularly to a method for evaluating the visual quality of ultra-high-definition videos based on adaptive spatial fusion and temporal selection. Background Art

[0002] With the rapid development of photography and display devices, ultra-high-definition (UHD) videos have become popular in the mass media. Compared with standard-definition videos, UHD videos have higher resolutions and generally higher frame rates, providing viewers with a vivid and smooth experience. However, various spatio-temporal distortions and artifacts may occur during the acquisition, transmission, and storage of UHD videos. Therefore, there is an urgent need to develop objective models to evaluate the visual quality of UHD videos.

[0003] Video quality assessment (VQA) aims to simulate the human visual system (HVS) by calculation, fitting the human visual quality perception process, enabling machines to evaluate video quality in a manner related to human subjective opinions. Although VQA research has achieved great success in standard-definition (SD) videos, most VQA models are not specifically designed for UHD videos. They still face the following two major challenges when evaluating UHD videos: The first challenge is called the spatial challenge, which is caused by the high-resolution characteristics of UHD videos. On the one hand, given a spatial scene, due to the global-priority visual attribute, the overall perception of the entire scene is crucial. This means that visual stimuli from all spatial positions should be considered. On the other hand, due to the attention mechanism, human vision is only sensitive within a limited viewing range. This means that stimuli from different spatial positions should not contribute equally.

[0004] The second challenge is called the temporal challenge, which comes from the high frame rate characteristics of UHD videos. On the one hand, the visual signals in UHD videos are highly redundant because similar patterns repeat multiple times over time. This redundancy should be greatly reduced to prevent some quality perception information from being masked by repeated signals. On the other hand, the temporal changes in visual signals indicate some important factors of video quality, such as temporal artifacts. Therefore, given a video to be evaluated, it is essential to describe its temporal dimension.

[0005] Both of the above two challenges will lead to the problems of the curse of dimensionality and high demand for computing resources.

[0006] Therefore, the existing technologies need to be improved. Summary of the Invention

[0007] The technical problem to be solved by the present invention is that, aiming at the defects of the existing technologies, the present invention provides a method for evaluating the visual quality of ultra-high-definition videos based on adaptive spatial fusion and temporal selection, so as to solve the problems of the curse of dimensionality and high demand for computing resources in the existing video quality assessment methods.

[0008] The technical solution adopted by the present invention to solve the technical problem is as follows: In a first aspect, the present invention provides a method for evaluating the visual quality of ultra-high-definition videos based on adaptive spatial fusion and temporal selection, including: Obtain the ultra-high-definition video to be evaluated; Perform intra-frame processing on each frame of the ultra-high-definition video through a spatial adaptive fusion network to fuse block-by-block features and obtain frame-by-frame features; Select several temporal positions from the frame-by-frame features through a temporal adaptive selection network, concatenate the features corresponding to the selected temporal positions, and map the concatenated features to a predicted score of video quality; Output the visual quality evaluation result of the ultra-high-definition video.

[0009] In one implementation, the performing intra-frame processing on each frame of the ultra-high-definition video through a spatial adaptive fusion network to fuse block-by-block features and obtain frame-by-frame features includes: Adaptively aggregate the block-by-block features within each frame of the ultra-high-definition video in a weighted summation manner to obtain local aggregation features; Capture the global information of cross-block features through cascaded vision transformer blocks and encode the adjusted weights in the global information to obtain global interaction features; Associate the local aggregation features and the global interaction features to obtain the frame-by-frame features.

[0010] In one implementation, the local aggregation features are calculated by the following formula: ; where f tl represents the local aggregation feature of the t-th frame, w m represents the m-th element in w, SAT represents the self-attention module; b m represents the block-by-block feature of the t-th frame, and the dimension of the local aggregation feature is the same as that of the block-by-block feature.

[0011] In one implementation, the capturing the global information of cross-block features through cascaded vision transformer blocks and encoding the adjusted weights in the global information to obtain global interaction features includes: Take each of the block-by-block features as an input token of the cascaded vision transformer block, generate M feature tokens for all tiles, and combine the class token as a learnable embedding to obtain the global information; Introduce a common weighted bias matrix W in the multi-head self-attention module of the cascaded vision transformer block, and encode the adjusted weights in the global information based on the common weighted bias matrix W to obtain the global interaction features.

[0012] In one implementation, the method of the time-adaptive selection network selecting several time positions from the frame-by-frame features, concatenating the features corresponding to the selected time positions, and mapping the concatenated features to the predicted score of the video quality includes: Perform grouping processing on the frame-by-frame features, calculate the feature mean and feature difference of each group of features, and connect the average features of all groups to obtain the finally connected feature map; Map the connected features to an importance vector d to predict the predicted score of the video quality.

[0013] In one implementation, the feature mean is calculated by the following formula: ; where represents the i-th frame-by-frame feature in the h-th group; The feature difference is calculated by the following formula: ; where and represent two different features in the group.

[0014] In one implementation, the method of mapping the connected features to an importance vector d to predict the predicted score of the video quality includes: Map the connected features to an importance vector d through a convolutional layer, a multi-layer perceptron, and a Sigmoid activation function, and gate the feature groups in a binary quantization manner, and use the gating mechanism to retain the features with greater importance and exclude other features at the same time; Predict the predicted score of the video quality based on the retained features.

[0015] In a second aspect, the present invention provides a super-high-definition video visual quality evaluation system based on adaptive spatial fusion and time selection, including: An acquisition module for acquiring a super-high-definition video to be evaluated; An intra-frame processing module for performing intra-frame processing on each frame of the super-high-definition video through a spatial adaptive fusion network to fuse block-by-block features to obtain frame-by-frame features; An inter-frame processing module for selecting several time positions from the frame-by-frame features through a time-adaptive selection network, concatenating the features corresponding to the selected time positions, and mapping the concatenated features to the predicted score of the video quality; An output module for outputting the visual quality assessment result of the ultra-high definition video.

[0016] In a third aspect, the present invention provides a terminal, including: a processor and a memory, where the memory stores an ultra-high definition video visual quality assessment program based on adaptive spatial fusion and temporal selection. When the ultra-high definition video visual quality assessment program based on adaptive spatial fusion and temporal selection is executed by the processor, it is used to implement the operations of the ultra-high definition video visual quality assessment method based on adaptive spatial fusion and temporal selection as described in the first aspect.

[0017] In a fourth aspect, the present invention further provides a medium, which is a computer-readable storage medium. The medium stores an ultra-high definition video visual quality assessment program based on adaptive spatial fusion and temporal selection. When the ultra-high definition video visual quality assessment program based on adaptive spatial fusion and temporal selection is executed by a processor, it is used to implement the operations of the ultra-high definition video visual quality assessment method based on adaptive spatial fusion and temporal selection as described in the first aspect.

[0018] The present invention adopts the above technical solutions and has the following effects: The present invention proposes a spatial adaptive fusion network, which fuses block-by-block features with adaptive spatial weights in a way of local aggregation and global interaction, avoiding the dimensionality disaster caused by information loss; and, a temporal adaptive fusion network is proposed, which adaptively selects positions according to the temporal characteristics of the video, and at the same time combines frame-level differential features and frame-level mean features to calculate the importance degree of different sampling positions, can minimize the number of selected positions, effectively eliminate redundant data, and reduce the computational resource requirements. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.

[0020] Figure 1 is a flowchart of the ultra-high definition video visual quality assessment method based on adaptive spatial fusion and temporal selection in the present invention.

[0021] Figure 2 is a schematic diagram of the overall framework of the ultra-high definition video visual quality assessment algorithm based on adaptive spatial fusion and temporal selection in the present invention.

[0022] Figure 3 is a schematic diagram of the spatial adaptive fusion network in the present invention.

[0023] Figure 4 It is a schematic diagram of the time - adaptive fusion network in the present invention.

[0024] Figure 5 It is a schematic diagram of an application scenario of the spatial weight in the present invention.

[0025] Figure 6 It is the functional schematic diagram of the terminal in one implementation manner of the present invention.

[0026] The realization of the purpose of the present invention, functional features and advantages will be further described in conjunction with embodiments with reference to the accompanying drawings. Detailed Embodiments

[0027] To make the purpose, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0028] Exemplary Method Existing VQA (Video Quality Assessment) models are not specifically designed for ultra - high - definition videos. They still face the following two major challenges when evaluating ultra - high - definition videos: The first challenge is called the spatial challenge, which is caused by the high - resolution characteristics of ultra - high - definition videos. On the one hand, given a spatial scene, due to the globally - preferred visual attributes, the overall perception of the entire scene is crucial. This means that visual stimuli from all spatial positions should be considered. On the other hand, due to the attention mechanism, human vision is only sharp within a limited viewing angle range. This means that stimuli from different spatial positions should not contribute equally.

[0029] The second challenge is called the temporal challenge, which comes from the high - frame - rate characteristics of ultra - high - definition videos. On the one hand, the visual signals in ultra - high - definition videos are highly redundant because similar patterns repeat many times over time. This redundancy should be greatly reduced to prevent some quality - perception information from being masked by repeated signals. On the other hand, the temporal changes in visual signals indicate some important factors of video quality, such as temporal artifacts. Therefore, given a video to be evaluated, it is essential to describe its temporal dimension.

[0030] Therefore, when existing VQA models face these two challenges, they will both lead to the problem of the curse of dimensionality and the problem of high demand for computing resources.

[0031] In view of the above technical problems, an ultra-high-definition video visual quality evaluation method based on adaptive spatial fusion and temporal selection is provided in an embodiment of the present invention. The method includes: obtaining an ultra-high-definition video to be evaluated; performing intra-frame processing on each frame of the ultra-high-definition video through a spatial adaptive fusion network to fuse block-by-block features and obtain frame-by-frame features; selecting several temporal positions from the frame-by-frame features through a temporal adaptive selection network, concatenating the features corresponding to the selected temporal positions, and mapping the concatenated features to a predicted score of video quality; outputting the visual quality evaluation result of the ultra-high-definition video. By adaptively fusing all block-by-block features, the present invention avoids the curse of dimensionality, and by adaptively selecting a small number of temporal positions, reduces the computational resource requirements.

[0032] As Figure 1 shown, an embodiment of the present invention provides an ultra-high-definition video visual quality evaluation method based on adaptive spatial fusion and temporal selection, including the following steps: Step S100, obtaining an ultra-high-definition video to be evaluated; In VQA (Video Quality Assessment) research, both the spatial and temporal aspects of the input video must be considered simultaneously. Given a UHD (Ultra-High-Definition) video to be evaluated, the dimensionality of its spatio-temporal data is very high, which makes the VQA task challenging.

[0033] From a spatial perspective, existing common strategies of randomly cropping or downsampling the original video frames can effectively reduce the data dimensionality, but inevitably result in information loss. Even though the dimensionality of video frames is very high, in this embodiment, it is desired to include and utilize all its spatial blocks to obtain a frame-by-frame representation. Therefore, the design idea of this embodiment is to extract a quality-aware feature for each block and then adaptively fuse all these block-by-block features. When fusing the block-by-block features into a frame-by-frame representation, there are at least two considerations in this embodiment. First, the contributions of the block-by-block features should not be equal. Human vision is only sharp within a limited viewing angle range. Therefore, when humans watch a high-resolution video, blocks at different spatial positions will play different roles in quality perception. Second, the frame-by-frame representation should include not only the local description within the block but also the global description across the blocks. This is because both local and global appearances have a great impact on video quality. Therefore, fusing features in a way of local aggregation and global interaction can avoid information loss.

[0034] From a temporal perspective, video data has a high degree of redundancy because similar visual patterns repeat over time. A commonly used strategy is to select video frames at predetermined or random time positions. However, in a VQA environment, these selected positions in the temporal dimension are hardly optimal. The solution idea in this embodiment is to adaptively select a small number of time positions from the complete video according to the frame-by-frame features of the video. Therefore, the selected positions should include those where the frame-by-frame features can well represent the spatial distortion of the entire video. Secondly, the change of inter-frame features within the neighborhood is also crucial. Considering the above two points, the number of selected positions should be minimized to effectively eliminate redundancy.

[0035] Based on the above theory, a VQA model for UHD videos is proposed in this embodiment, and its framework is as Figure 2 shown. This framework consists of two cascaded parts: intra-frame processing and inter-frame processing.

[0036] In this embodiment, based on the above VQA model, first obtain the ultra-high-definition video to be evaluated and use it as the input of the VQA model; assume that the total number of frames of the ultra-high-definition video to be evaluated is T. During the intra-frame processing, for the t-th frame of the video, the goal of intra-frame processing is to generate its spatial representation f t (1 ≤ t ≤ T). This goal is mainly achieved through a spatial adaptive fusion network (i.e., the SAFNet network), where block-by-block features are adaptively weighted and fused. In inter-frame processing, L time positions are selected from the frame-by-frame features {f t , …, f T} through a temporal adaptive selection network (i.e., the TASNet network). Finally, the concatenated feature F of the selected L features is mapped to the predicted score of the video quality.

[0037] As Figure 1 shown, an embodiment of the present invention provides a method for visual quality assessment of ultra-high-definition videos based on adaptive spatial fusion and temporal selection, including the following steps: Step S200, perform intra-frame processing on each frame of the ultra-high-definition video through a spatial adaptive fusion network to fuse block-by-block features and obtain frame-by-frame features.

[0038] In this embodiment, in order to implement intra-frame processing, a spatial adaptive feature fusion network (SAFNet) is designed, and the architecture of SAFNet is as Figure 3 shown. It fuses all block-by-block features in two ways: local aggregation and global interaction. In the former, block features are adaptively aggregated in a weighted summation manner. In the latter, the global information of cross-block features is captured by a cascaded vision transformer (ViT) block and the adjusted weights are encoded therein.

[0039] Specifically, in one implementation of this embodiment, step S200 includes the following steps: Step S201, adaptively aggregate the per-block features within each frame of the ultra-high-definition video in a weighted summation manner to obtain local aggregation features.

[0040] In this embodiment, the local aggregation features are calculated by the following formula: ; where f tl represents the local aggregation feature of the t-th frame, w m represents the m-th element in w, SAT represents the self-attention module; b m represents the per-block feature of the t-th frame, and the dimension of the local aggregation feature is the same as that of the per-block feature b m .

[0041] Step S202, capture the global information of the cross-block features through cascaded vision transformer blocks and encode the adjusted weights in the global information to obtain global interaction features; Specifically, the capturing the global information of the cross-block features through cascaded vision transformer blocks and encoding the adjusted weights in the global information to obtain global interaction features includes: taking each per-block feature as the input token of the cascaded vision transformer block, generating M feature tokens for all patches, and combining the class token as a learnable embedding to obtain the global information; introducing a common weighted bias matrix W in the multi-head self-attention module of the cascaded vision transformer block, and encoding the adjusted weights in the global information based on the common weighted bias matrix W to obtain the global interaction features.

[0042] Step S203, associate the per-frame features based on the local aggregation features and the global interaction features.

[0043] In this embodiment, for the global aggregation operation, in order to capture the interaction between different patches (i.e., per-block features), a ViT (cascaded vision transformer) architecture is adopted, which can well describe the long-range feature relationship. Specifically, each patch feature is regarded as the input token of ViT, so as to generate M feature tokens for all patches. In addition, the ViT architecture in this embodiment also includes a class token cls-token as a learnable embedding. This embodiment introduces a common weighted bias matrix W in the multi-head self-attention (MHSA) module of the ViT block. The common weighted bias matrix W can effectively consider the importance between different tokens, making the global aggregation features more representative.

[0044] As an example, such as Figure 2As shown in the figure, in this embodiment, taking the t-th frame of the input ultra-high-definition video as an example, it is divided into multiple tiles through Cropping processing. For each tile, after the encoder assigns shared weights, the per-tile feature is obtained; then, all the per-tile features are input into the spatial adaptive fusion network for fusion with the initial weights, and the per-frame feature is output.

[0045] Specifically, as Figure 3 shown, the processing process of the spatial adaptive fusion network is as follows: 1) Spatial weight adjustment: For the input per-tile features (b 1 ,..., b m ,..., b M ), each per-tile feature is passed through a fully connected layer FC, a conditional modulation layer CML, and a ReLU function layer for feature transformation processing; then, through a Concatenation operation, multiple tensors of the same dimension are concatenated together to obtain the feature x; furthermore, the feature x is mapped to the adjustment vector z, and the initial weights (s 1 ,..., s m ,..., s M ) are used for weight adjustment to determine the weights (w 1 ,..., w m ,..., w M ) after spatial weight adjustment.

[0046] 2) Local feature aggregation: Using the weights after spatial weight adjustment, perform summation in a weighted summation manner, and combine the mechanism of the self-attention module to adaptively aggregate the per-tile features (b 1 ,..., b m ,..., b M ) to obtain the local aggregated feature .

[0047] 3) Global interaction fusion: For the input per-tile features (b 1 ,..., b m ,..., b M ), combine them with the class token cls-token through a combination function (Combination) to obtain the combined feature B; then, in the multi-head self-attention (MHSA) module of the ViT block, use the common weighted bias matrix W to assign weights to each feature to obtain the global interaction feature; among them, the common weighted bias matrix W is calculated based on the weights after spatial weight adjustment.

[0048] 4) Concatenation operation: Concatenate the local aggregated feature and the global interaction feature together to obtain the per-frame feature 。

[0049] As Figure 1 shown, an embodiment of the present invention provides a method for evaluating the visual quality of ultra - high - definition videos based on adaptive spatial fusion and temporal selection, including the following steps: Step S300: Select a number of temporal positions from the per - frame features through a temporal adaptive selection network, concatenate the features corresponding to the selected temporal positions, and map the concatenated features to a predicted score of the video quality; Step S400: Output the visual quality evaluation result of the ultra - high - definition video.

[0050] In this embodiment, as an example, as Figure 2 shown, for all per - frame features (f 1 ,..., f t ,..., f T ) output by the spatial adaptive fusion network, input them into the temporal adaptive selection network for inter - frame processing to obtain a concatenated feature F; and input the concatenated feature F into the quality score prediction module, use the gated recurrent unit (GRU) of the quality score prediction module for gating, and use a multi - layer perceptron (MLP) to output quality scores (p 1 ,..., p l ,..., p L ), and after merging, obtain the predicted score of the final video quality.

[0051] Specifically, in an implementation manner of this embodiment, step S300 includes the following steps: Step S301: Group - process the per - frame features, calculate the feature mean and feature difference of each group, and connect the average features of all groups to obtain a finally - connected feature map; Step S302: Map the connected features to an importance vector d, and predict to obtain the predicted score of the video quality.

[0052] In this embodiment, in order to implement inter - frame processing, a temporal adaptive selection network (TASNet) is proposed. The architecture of TASNet is as Figure 4 shown. It selects a small number of per - frame features along the time dimension according to two main factors. In one factor, in this embodiment, it is considered whether the selected features are representative. In another factor, in this embodiment, it is considered whether the temporal changes of the selected features appear in consecutive frames. Based on the above two factors, the temporal importance of the per - frame features can be obtained.

[0053] Specifically, the per-frame features are grouped, with each group consisting of 7 frames (7 is an empirical parameter). Then, the mean and variance of each group of features are calculated. The feature mean can be solved by formula (2), and the feature variance can be obtained from formula (3).

[0054] (2); (3); For formula (2), where represents the i-th per-frame feature in the h-th group; the average features of all H groups are concatenated to obtain the final feature map. As Figure 4 shown, the feature map is then transformed by cascaded ViT blocks, where the average feature of each group is regarded as a token. Due to the interaction ability of ViT, the final output features of the ViT blocks contain information about the token relationships. For formula (3), where and represent two different features in the group.

[0055] In one implementation of this embodiment, mapping the concatenated features to the importance vector d and predicting the predicted score of the video quality includes: mapping the concatenated features to the importance vector d through a convolutional layer, a multi-layer perceptron, and a Sigmoid activation function, and gating the feature groups in a binary quantization manner, using the gating mechanism to retain the features with greater importance and exclude other features at the same time; predicting the predicted score of the video quality based on the retained features.

[0056] For the variance features of formula (3) above, different from the average feature map, this feature map is then transformed by residual blocks (RB), as Figure 4 shown. In these residual blocks (RB), only one-dimensional (1D) kernels are included, and convolutional operations are performed within each group. In this way, the variance features of a group can be enhanced without being interfered by other groups. The final output features of the residual blocks (RB) are expected to contain information about the feature variations within each group. To obtain the temporal importance of each group, in this embodiment, by concatenating F A and F D along the channel dimension to combine the above two factors.

[0057] As Figure 4 shown, the concatenated features are mapped to the importance vector d through a convolutional layer, a multi-layer perceptron (MLP), and a Sigmoid activation function. Then, in this embodiment, the feature groups are gated in a binary quantization manner, using the gating mechanism to retain the features with greater importance and exclude other features at the same time. It should be noted that the number of features selected for different videos may be different, and finally, the video quality score is predicted through these selected features.

[0058] To prove the effectiveness of the proposed model in this embodiment, it is compared with some traditional and state-of-the-art NR VQA models, including: V-BLIINDS model, LVQM model, VSFA model, VIDEVAL model, GSTVQA model, BVQA model, FAST-VQA model, TB-VQA model, STFEE model, STI-VQA model, and HR-VQA model. In this embodiment, tests are conducted on four publicly available datasets, including: Waterloo IVC dataset, BVI-SR dataset, YouTube-UGC dataset, and ITM-HDR-VQA dataset. The performance of different models can be reflected by calculating the Spearman rank correlation coefficient (SROCC) and Pearson linear correlation coefficient (PLCC). The larger the value of the metric, the better the performance.

[0059] Table I Performance of Different VQA Models on 4K Videos

[0060] As shown in Table 1, compared with the current state-of-the-art no-reference video quality assessment methods, the method proposed in this embodiment has achieved the best level in algorithm performance metrics. Bold indicates the best performance, and underlined indicates the second-best performance. In contrast, the performance of competitors is poor. For example, the STFEE model achieved the second-best performance on 4K videos in the Waterloo IVC dataset. This is attributed to the use of a saliency map during patch cropping. However, the model did not optimize the contribution of each patch in the VQA task and ignored the global interaction of local features. On the BVI-SR dataset, the second-best model is the STI-VQA model, which fuses quality-aware features from multiple scales. However, it did not consider the influence of patch location. The HR-VQA model ranked second on the YouTube-UGC dataset and the ITM-HDR-VQA dataset. The main reason is that the HR-VQA model can alleviate spatial challenges through the GMS strategy. However, it treats all patches equally, which is unreasonable for UHD videos. All these results indicate that the proposed model is effective in evaluating the visual quality of UHD videos.

[0061] In addition, the effectiveness of the innovation is also verified in this embodiment. As Figure 5 shown, Figure 5 in (a) represents a video frame with two image patches, highlighted in red and cyan respectively. Figure 5 in (b) represents the corresponding saliency map of (a). Figure 5 in (c) represents the football player image patch and its spatial weight. Figure 5 in (d) represents the lawn image patch and its spatial weight.

[0062] Figure 5 An example of weight adjustment is shown, which includes a video frame with two highlighted patches and the corresponding saliency map. The basic weights of these two patches are 0.4577 and 0.0639 respectively. Their weight ratio can be calculated to be approximately 7.16. Obviously, the football player area is more attractive than the grass area. However, after adaptive weight adjustment, their weight ratio is reduced to about 1.64. This means that the contribution of the grass area in the VQA task has increased. This result is quite reasonable because serious blocking artifacts can be observed in the grass area.

[0063] This embodiment achieves the following technical effects through the above technical solutions: This embodiment proposes a spatial adaptive fusion network, which fuses patch-level features with adaptive spatial weights in a way of local aggregation and global interaction, avoiding the dimensionality disaster caused by information loss; and, a temporal adaptive fusion network is proposed, which adaptively selects positions according to the video temporal characteristics, and combines frame-level differential features and frame-level mean features to calculate the importance of different sampling positions, which can minimize the number of selected positions, effectively eliminate redundant data, and reduce the computational resource requirements.

[0064] Exemplary device Based on the above embodiments, the present invention further provides a super high-definition video visual quality evaluation system based on adaptive spatial fusion and temporal selection, including: An acquisition module, configured to acquire the super high-definition video to be evaluated; An intra-frame processing module, configured to perform intra-frame processing on each frame of the super high-definition video through a spatial adaptive fusion network to fuse patch-level features and obtain frame-level features; An inter-frame processing module, configured to select several temporal positions from the frame-level features through a temporal adaptive selection network, concatenate the features corresponding to the selected temporal positions, and map the concatenated features to a predicted score of the video quality; An output module, configured to output the visual quality evaluation result of the super high-definition video.

[0065] This embodiment achieves the following technical effects through the above technical solutions: This embodiment proposes a spatial adaptive fusion network, which fuses patch-level features with adaptive spatial weights in a way of local aggregation and global interaction, avoiding the dimensionality disaster caused by information loss; and, a temporal adaptive fusion network is proposed, which adaptively selects positions according to the video temporal characteristics, and combines frame-level differential features and frame-level mean features to calculate the importance of different sampling positions, which can minimize the number of selected positions, effectively eliminate redundant data, and reduce the computational resource requirements.

[0066] Based on the above embodiments, the present invention further provides a terminal, and its principle block diagram can be as Figure 6 shown.

[0067] The terminal includes: a processor, a memory, an interface, a display screen, and a communication module connected through a system bus; wherein, the processor of the terminal is used to provide computing and control capabilities; the memory of the terminal includes a storage medium and an internal memory; the storage medium stores an operating system and a computer program; the internal memory provides an environment for the operation of the operating system and the computer program in the storage medium; the interface is used to connect external devices; the display screen is used to display corresponding information; the communication module is used to communicate with a cloud server or other devices.

[0068] When the computer program is executed by the processor, it is used to implement the operations of the ultra-high-definition video visual quality assessment method based on adaptive spatial fusion and temporal selection.

[0069] Those skilled in the art can understand that Figure 6 the principle block diagram shown in

[0070] merely shows the block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0071] In one embodiment, a terminal is provided, which includes: a processor and a memory, and the memory stores an ultra-high-definition video visual quality assessment program based on adaptive spatial fusion and temporal selection. When the ultra-high-definition video visual quality assessment program based on adaptive spatial fusion and temporal selection is executed by the processor, it is used to implement the operations of the ultra-high-definition video visual quality assessment method based on adaptive spatial fusion and temporal selection as described above.

[0072] Those of ordinary skill in the art can understand that all or part of the processes in the above embodiments of the method can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and volatile memories.

[0073] In summary, the present invention provides a method for evaluating the visual quality of ultra-high-definition videos based on adaptive spatial fusion and temporal selection, including: obtaining an ultra-high-definition video to be evaluated; performing intra-frame processing on each frame of the ultra-high-definition video through a spatial adaptive fusion network to fuse block-by-block features and obtain frame-by-frame features; selecting several temporal positions from the frame-by-frame features through a temporal adaptive selection network, concatenating the features corresponding to the selected temporal positions, and mapping the concatenated features to a predicted score of video quality; and outputting the visual quality evaluation result of the ultra-high-definition video. The present invention avoids the curse of dimensionality by adaptively fusing all block-by-block features, and reduces the computational resource requirements by adaptively selecting a small number of temporal positions.

[0074] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for ultra-high-definition video visual quality assessment based on adaptive spatial fusion and temporal selection, characterized in that: include: Obtain the ultra-high-definition video to be evaluated; Performing intra-frame processing on each frame of the ultra-high-definition video through a spatial adaptive fusion network to fuse block-by-block features to obtain frame-by-frame features; Selecting a plurality of time positions from the frame-by-frame features through a time adaptive selection network, connecting the features corresponding to the selected time positions in series, and mapping the connected features to a prediction score of the video quality; Output a visual quality assessment result of the ultra-high definition video.

2. The method for evaluating ultra-high definition video visual quality based on adaptive spatial fusion and temporal selection according to claim 1, characterized in that: The intra-frame processing of each frame of the ultra-high-definition video is performed by the spatial adaptive fusion network to fuse block-by-block features to obtain frame-by-frame features, including: Adaptively aggregating block-by-block features in each frame of the ultra-high-definition video in a weighted summation manner to obtain local aggregated features; Capturing the global information of cross-block features by cascading visual transformer blocks and encoding the adjusted weights in the global information to obtain global interaction features; The frame-by-frame features are obtained based on the association between the local aggregate features and the global interaction features.

3. The method for evaluating ultra-high definition video visual quality based on adaptive spatial fusion and temporal selection according to claim 2, characterized in that: The local aggregation feature is calculated by the following formula: ; Among them, f tl represents the local aggregated features of the tth frame, w m represents the mth element in w, SAT represents the self-attention module; b m represents the block-by-block features of the t-th frame, and the dimension of the local aggregated features is the same as the dimension of the block-by-block features.

4. The method for evaluating ultra-high definition video visual quality based on adaptive spatial fusion and temporal selection according to claim 2, characterized in that: The method captures the global information of cross-block features by cascading visual transformer blocks and encodes the adjusted weights in the global information to obtain global interaction features, including: Using each of the block-wise features as an input tag of the cascaded visual transformer block, generating M feature tags for all blocks, and combining the class tags as learnable embeddings to obtain the global information; A common weighted bias matrix W is introduced into the multi-head self-attention module of the cascaded visual transformer block, and based on the common weighted bias matrix W, the adjusted weights are encoded in the global information to obtain the global interaction feature.

5. The method for evaluating ultra-high definition video visual quality based on adaptive spatial fusion and temporal selection according to claim 1, characterized in that: The method of selecting a plurality of time positions from the frame-by-frame features through a time adaptive selection network, connecting the features corresponding to the selected time positions in series, and mapping the connected features to a prediction score of the video quality includes: The frame-by-frame features are grouped, and the feature mean and feature difference of each group of features are calculated, and the average features of all groups are connected to obtain a final connected feature map; The connected features are mapped to the importance vector d, and the prediction score of the video quality is predicted.

6. The method for evaluating ultra-high definition video visual quality based on adaptive spatial fusion and temporal selection according to claim 5, characterized in that: The characteristic mean is calculated by the following formula: ; in, represents the i-th frame-by-frame feature in the h-th group; The characteristic difference is calculated by the following formula: ; in and Represents two different features in a group.

7. The method for evaluating ultra-high definition video visual quality based on adaptive spatial fusion and temporal selection according to claim 5, characterized in that: Mapping the connected features to the importance vector d and predicting the predicted score of the video quality includes: The connected features are mapped to an importance vector d through a convolutional layer, a multi-layer perceptron, and a Sigmoid activation function, and the feature group is gated by a binary quantization method, and the gating mechanism is used to retain features with greater importance and exclude other features at the same time; A prediction score of the video quality is obtained based on the retained feature prediction.

8. An ultra-high-definition video visual quality assessment system based on adaptive spatial fusion and temporal selection, characterized in that: include: An acquisition module, used for acquiring an ultra-high-definition video to be evaluated; An intra-frame processing module, used for performing intra-frame processing on each frame of the ultra-high-definition video through a spatial adaptive fusion network to fuse block-by-block features to obtain frame-by-frame features; An inter-frame processing module, configured to select a plurality of time positions from the frame-by-frame features through a time adaptive selection network, concatenate the features corresponding to the selected time positions, and map the concatenated features to a prediction score of the video quality; An output module is used to output the visual quality assessment result of the ultra-high-definition video.

9. A terminal, characterized in that: include: A processor and a memory, wherein the memory stores an ultra-high definition video visual quality assessment program based on adaptive spatial fusion and temporal selection, and when the ultra-high definition video visual quality assessment program based on adaptive spatial fusion and temporal selection is executed by the processor, it is used to implement the operation of the ultra-high definition video visual quality assessment method based on adaptive spatial fusion and temporal selection as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an ultra-high-definition video visual quality assessment program based on adaptive spatial fusion and temporal selection, and the ultra-high-definition video visual quality assessment program based on adaptive spatial fusion and temporal selection, when executed by a processor, is used to implement the operation of the ultra-high-definition video visual quality assessment method based on adaptive spatial fusion and temporal selection as described in any one of claims 1-7.