Video quality evaluation method with rich information representation
By extracting video frame features in the RGB and YIQ color spaces, combining CA attention mechanism and Euclidean distance calculation, a rich feature vector is constructed and temporal modeling is performed using TCN. This solves the problems of insufficient accuracy and robustness in existing video quality assessment methods and achieves more efficient video quality assessment.
Patent Information
- Application Number
- CN202510708988.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-10-21
AI Technical Summary
Existing video quality assessment methods rely on subjective human evaluation, which is time-consuming, labor-intensive, and easily influenced by personal emotions. Furthermore, existing machine learning models lack sufficient representation of the relationships between video frames, resulting in insufficient accuracy and robustness in the evaluation.
This paper proposes a video quality assessment method consisting of four parts: spatial feature extraction, motion feature extraction, temporal modeling, and quality prediction. By extracting video frame features in the RGB and YIQ color spaces, combining CA attention mechanism and Euclidean distance calculation, rich feature vectors are constructed. Temporal modeling is performed using TCN, and finally, video quality is predicted through a fully connected layer.
This improves the accuracy and robustness of video quality assessment, enhances the model's adaptability to different video types and viewing conditions, and improves the overall performance of video quality assessment.
Smart Images

Figure CN120823148A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video quality evaluation methods, and more specifically, to a video quality evaluation method with rich information representation. Background Art
[0002] Traditional video quality evaluation mostly relies on manual subjective evaluation, which is not only time-consuming and labor-intensive, but also easily influenced by the evaluator's personal emotions and preferences. With the rapid development of technologies such as deep learning, computer vision, and audio processing, video quality evaluation models based on machine learning and artificial intelligence have gradually emerged. Such models can automatically and accurately evaluate the quality of videos through multi-dimensional analysis of video content, greatly improving the efficiency and accuracy of video quality evaluation.
[0003] Although this type of model has made significant progress, it still faces many challenges, such as how to accurately extract and characterize key quality features in complex and changing video scenes, how to effectively integrate multiple features to improve the accuracy and robustness of evaluation, and how to make the model better adapt to different video types, encoding formats and viewing conditions. In addition, the current mainstream method of video quality evaluation first focuses on extracting features from the spatial domain of video frames. These features are usually static attributes of the image, such as color, texture, brightness distribution, etc. Subsequently, these static spatial features are input into the time series model to capture and model the changes of these features over time, thereby realizing the prediction of the overall quality of the video. This method lacks rich information representation. This method only treats each frame as an independent image for processing, which breaks the relationship between video frames. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a video quality evaluation method with rich information representation. By setting four parts: spatial feature extraction, motion feature extraction, temporal modeling and quality prediction, the spatial feature extraction part of the model takes into account the different characteristics of the video frames in RGB and YIQ color spaces to fully extract the features of the video frames in the spatial domain. In addition, in order to fully understand the connection between adjacent frames, the motion information in the temporal sequence is supplemented by extracting the motion features of the video, so that the model has better performance and improves the accuracy of the model prediction. The overall model has better performance and accuracy in video quality evaluation, and the overall performance of the model is improved through each part to solve the problems raised in the above background technology.
[0005] To achieve the above objectives, the present invention provides the following technical solutions: a video quality assessment method with rich information representation, including spatial feature extraction, motion feature extraction, time series modeling and quality prediction;
[0006] Spatial feature extraction: First, traverse each frame of the video to obtain all RGB image frames, then convert all RGB image frames to YIQ color space, thereby obtaining all image frames in YIQ color space. After that, use the modified ResNet-50 network pre-trained on ImageNet to extract spatial features on RGB and YIQ of all video frames at the same time. The modified network model removes the last two layers of the ResNet-50 network, namely the average pooling layer and the fully connected layer, and directly outputs 2048-dimensional features. At the same time, the CA attention mechanism is introduced. The CA attention mechanism dynamically adjusts the feature weights of different positions by analyzing the horizontal and vertical coordinate information of each position in the feature map. Subsequently, the features after the CA attention mechanism are adaptively average pooled and flattened. Finally, the feature vectors extracted from the two color spaces are spliced in the column direction to obtain a spatial feature vector with 4096 dimensions.
[0007] Motion feature extraction: First, each frame of the video sequence is converted into a grayscale image. By simplifying the image information, the interference of color on motion feature extraction is removed. Then, the Euclidean distance between corresponding pixels of adjacent grayscale images is calculated to capture the motion information in the video. To obtain a more robust and representative global motion feature, the calculated Euclidean distance of all pixels is average pooled to obtain the average Euclidean distance between adjacent grayscale images, which reflects the overall motion intensity and degree of change between video frames. The extracted motion feature vector is then adjusted to be spliced with the spatial domain feature on the column level for subsequent time series modeling.
[0008] Temporal modeling: By extracting spatial and motion features and fusing them, we obtain a 4097-dimensional feature vector for each frame. This feature vector is input into the TCN for temporal modeling. The TCN output channel is set to 128, and the dimensionality is reduced. That is, the output feature dimension after TCN is 128.
[0009] Quality prediction: To improve the computational efficiency of the model, a fully connected layer is first constructed after the TCN to directly map the reduced feature vector to the video quality score. Specifically, the 128-dimensional feature vector output by the TCN after dimensionality reduction is input into a fully connected layer. The output dimension of the fully connected layer is 1, which represents the video quality score. The calculation process is as follows: , where Score is the video score of the evaluation, FC is the fully connected layer, and F represents the fused spatial features and motion feature vectors.
[0010] In a preferred embodiment, the spatial feature extraction part uses a pre-trained ResNet-50 network to perform feature extraction in the RGB and YIQ color spaces of the video frame, and introduces the CA attention mechanism to enhance the spatial features.
[0011] In a preferred embodiment, the motion feature extraction part calculates the Euclidean distance between pixels of adjacent frames to represent the degree of motion change of the video, and fuses the extracted spatial features and the calculated motion features to construct a feature vector with information representation.
[0012] In a preferred embodiment, the fused feature vector is input into a time-domain convolutional network for temporal modeling, and the feature dimension is reduced to 128. The 128-dimensional feature is directly mapped to the quality score of the video through a fully connected layer.
[0013] In a preferred embodiment, the spatial feature extraction part uses the different characteristics of the video frame in the RGB and YIQ color spaces to extract the features of the video frame in the spatial domain, and extracts the motion features of the video to supplement the temporal motion information.
[0014] In a preferred embodiment, the formula for the spatial feature extraction process is as follows: , where ResNet-50 is the training network, is the spatial feature, is the input video frame sequence, i=1, 2, ..., T, T is the total number of video frames, and after passing through the ResNet-50 network, the video frame has 2048-dimensional spatial domain features;
[0015] The calculation process of the CA attention mechanism is as follows: for an input video frame feature map x, the CA attention mechanism selects the input feature map and performs pooling operations in the horizontal and vertical directions respectively. The calculation formula is as follows: 、 ,in represents the Cth channel of the input feature map, is the relevant output after the C-th channel pooling, Z c h Represents the horizontal output after the C-th channel pooling, Z c w Represents the vertical output after the C-th channel pooling, x c (h, i) represents the i-th pixel with height h in the C-th channel, x c (j, w) represents the jth pixel of the Cth channel with a width of w. The feature maps after pooling in the horizontal and vertical directions are then spliced together, and a 1×1 convolution operation and activation operation are performed. The calculation formula is as follows: ,in, represents a 1×1 convolution dimension reduction function, Indicates activation operation, Z h , Z w Represents the horizontal and vertical outputs of all channels after pooling, thereby obtaining a feature map f of size C / r×(H+W)×1. Then, f is separated into feature maps F in the horizontal and vertical directions in the spatial dimension. h and F w , and then perform a 1×1 convolution dimension increase operation. The calculation formula is as follows: 、 ,in, and are the convolutions in the horizontal and vertical directions respectively, is the Sigmoid activation function, so we get the attention vector and ,Finally, the output formula of the CA attention mechanism is as follows: , where x c (i, j) is the pixel of channel C, g c h (i) and g c w (j) are the horizontal vector and vertical vector of the attention vector of the C-th channel, respectively, so that the spatial domain features of the entire video The size is T×4096.
[0016] In a preferred embodiment, the conversion formula for converting each frame of the video sequence into a grayscale image is: , where R, G, and B represent the three color channels of the image;
[0017] The calculation formula that reflects the overall motion intensity and degree of change between video frames is as follows: ,in represents the grayscale image of the t-th frame, Represents the grayscale image of the t+1th frame, D represents the global average Euclidean distance, and the total number of frames of the input video is T. Therefore, the Euclidean distance values of T-1 adjacent frames can be obtained. Since the value range of the grayscale image is [0, 255], the calculated Euclidean distance value is also between 0 and 255. Then the T-1 Euclidean distances are normalized. This process will use Min-Max normalization. At this time, the Euclidean distance value is mapped to [0, 1]. The conversion formula is as follows: ,in, (i=1, 2, ..., T-1) is the Euclidean distance between the i-th frame and the i+1-th frame, min(D) and max(D) represent their maximum and minimum values respectively. Therefore, the motion characteristics of the entire video are The size is T-1×1. So far, the spatial features of the entire video size T×4096 and the motion features of T-1×1 are extracted;
[0018] The process of adjusting the extracted motion feature vector is as follows: first, the first row of the motion feature vector is filled with 0 to make its size T×1, thereby matching the T×4096 dimension of the spatial feature. Then, the adjusted motion feature vector and the spatial feature are spliced on the column to construct a feature matrix of size T×4097. The fused feature matrix contains both the spatial information of the video and the motion information between frames.
[0019] In a preferred embodiment, the calculation process of inputting the feature vector into TCN for time series modeling is as follows: , where F k and F d They are spatial features and motion feature vectors respectively.
[0020] The technical effects and advantages of the present invention are as follows:
[0021] The present invention sets four parts: main spatial feature extraction, motion feature extraction, time series modeling and quality prediction. The spatial feature extraction part of the model takes into account the different characteristics of video frames in RGB and YIQ color spaces to fully extract the features of video frames in the spatial domain. In addition, in order to fully understand the connection between adjacent frames, the motion features of the video are extracted to supplement the motion information in the time series, so that the model has better performance and improves the accuracy of model prediction. The overall model has better performance and accuracy in video quality evaluation, and the overall performance of the model is improved through each part, further improving the overall use effect of the video quality evaluation system. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a schematic diagram of the overall framework flow of the model of the present invention;
[0023] Figure 2 Schematic diagram of the feature extraction backbone network of the present invention. DETAILED DESCRIPTION
[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0025] As attached Figure 1 To the attached Figure 2The video quality evaluation method shown in FIG has rich information representation, including spatial feature extraction, motion feature extraction, time series modeling and quality prediction.
[0026] Spatial feature extraction: First, traverse each frame of the video to obtain all RGB image frames. Then, all RGB image frames are converted to YIQ color space to obtain all image frames in YIQ color space. After that, the modified ResNet-50 network pre-trained on ImageNet is used to extract spatial features on RGB and YIQ of all video frames at the same time. The modified network model removes the last two layers of the ResNet-50 network, namely the average pooling layer and the fully connected layer, and directly outputs 2048-dimensional features. The formula for the spatial feature extraction process is as follows: , where ResNet-50 is the training network, is the spatial feature, is the input video frame sequence, i=1, 2, ..., T, T is the total number of video frames. After passing through the ResNet-50 network, the video frame has a 2048-dimensional spatial domain feature. At the same time, the CA attention mechanism is introduced. The CA attention mechanism analyzes the horizontal and vertical coordinate information of each position in the feature map, thereby dynamically adjusting the feature weights of different positions. Its calculation process is as follows: For an input video frame feature map x, the CA attention mechanism selects the input feature map for pooling operations in the horizontal and vertical directions respectively. The calculation formula is as follows: 、 ,in represents the Cth channel of the input feature map, is the relevant output after the C-th channel pooling, Z c h Represents the horizontal output after the C-th channel pooling, Z c w Represents the vertical output after the C-th channel pooling, x c (h, i) represents the i-th pixel with height h in the C-th channel, x c (j, w) represents the jth pixel of the Cth channel with a width of w. The feature maps after pooling in the horizontal and vertical directions are then spliced together, and a 1×1 convolution operation and activation operation are performed. The calculation formula is as follows: ,in, represents a 1×1 convolution dimension reduction function, Indicates activation operation, Z h , Z w Represents the horizontal and vertical outputs of all channels after pooling, thereby obtaining a feature map f of size C / r×(H+W)×1. Then, f is separated into feature maps F in the horizontal and vertical directions in the spatial dimension. h and Fw , and then perform a 1×1 convolution dimension increase operation. The calculation formula is as follows: 、 ,in, and are the convolutions in the horizontal and vertical directions respectively, is the Sigmoid activation function, so we get the attention vector and ,Finally, the output formula of the CA attention mechanism is as follows: , where x c (i, j) is the pixel of channel C, g c h (i) and g c w (j) are the horizontal vector and vertical vector of the attention vector of the Cth channel respectively;
[0027] Then, the features after the CA attention mechanism are adaptively averaged and flattened, and finally the feature vectors extracted from the two color spaces are spliced in the column direction to obtain a spatial feature vector with 4096 dimensions, making the spatial feature of the entire video The size is T×4096;
[0028] Motion feature extraction: First, each frame of the video sequence is converted into a grayscale image. By simplifying the image information, the interference of color on motion feature extraction is removed. The conversion formula is: , where R, G, and B represent the three color channels of the image. Then, the Euclidean distances between corresponding pixels in adjacent grayscale images are calculated to capture the motion information in the video. To obtain a more robust and representative global motion feature, the calculated Euclidean distances of all pixels are average-pooled to obtain the average Euclidean distance between adjacent grayscale images, reflecting the overall motion intensity and degree of change between video frames. The calculation formula is as follows: ,in represents the grayscale image of the t-th frame, Represents the grayscale image of the t+1th frame, D represents the global average Euclidean distance, and the total number of frames of the input video is T. Therefore, the Euclidean distance values of T-1 adjacent frames can be obtained. Since the value range of the grayscale image is [0, 255], the calculated Euclidean distance value is also between 0 and 255. Then the T-1 Euclidean distances are normalized. This process will use Min-Max normalization. At this time, the Euclidean distance value is mapped to [0, 1]. The conversion formula is as follows: ,in, (i=1, 2, ..., T-1) is the Euclidean distance between the i-th frame and the i+1-th frame, min(D) and max(D) represent their maximum and minimum values, respectively;
[0029] Motion characteristics of the entire video The size is T-1×1. So far, the spatial features of the entire video T×4096 and the motion features of T-1×1 are extracted. Then the extracted motion feature vector is adjusted so that it can be spliced with the spatial features on the column for subsequent time series modeling. First, the first row of the motion feature vector is filled with 0 to make its size T×1, thereby matching the T×4096 dimension of the spatial features. Then, the adjusted motion feature vector is spliced with the spatial features on the column to construct a feature matrix of size T×4097. The fused feature matrix contains both the spatial information of the video and the motion information between frames. Temporal modeling: By extracting the spatial features and motion features above and fusing them together, we obtain a 4097-dimensional feature vector for each frame. This feature vector is then fed into the TCN for temporal modeling. The calculation process is as follows: , where F k and F d They are spatial features and motion feature vectors respectively. The TCN output channel is set to 128 and the dimensionality reduction output is performed, that is, the output feature dimension after TCN is 128;
[0030] Quality prediction: To improve the computational efficiency of the model, a fully connected layer is first constructed after the TCN to directly map the feature vector after dimensionality reduction to the video quality score. Specifically, the 128-dimensional feature vector output by the TCN is input into a fully connected layer. The output dimension of the fully connected layer is 1, which represents the video quality score. The calculation process is as follows: , where Score is the video score of the evaluation, FC is the fully connected layer, and F represents the fused spatial features and motion feature vectors.
[0031] The spatial feature extraction part uses a pre-trained ResNet-50 network to extract features in the RGB and YIQ color spaces of the video frames. The CA attention mechanism is then introduced to enhance the spatial features to improve the feature representation ability. The motion feature extraction part creatively calculates the Euclidean distance between pixels in adjacent frames to represent the degree of motion change in the video. Secondly, the extracted spatial features and the calculated motion features are fused to construct a feature vector with rich information representation. The fused feature vector is input into the temporal convolutional network for temporal modeling and the feature dimension is reduced to 128. Finally, the 128-dimensional features can be directly mapped to the video quality score through a fully connected layer. The spatial feature extraction part uses the different characteristics of video frames in RGB and YIQ color spaces to fully extract the spatial features of the video frames. In addition, in order to improve the connection between adjacent frames, the motion features of the video are extracted to supplement the temporal motion information and improve the accuracy of its model prediction. The spatial feature extraction and motion feature extraction together constitute the feature extraction module.
[0032] To verify the effectiveness of the proposed model, 10 experiments were conducted on the CVD2014, LIVE-VQC, and KoNViD-1k datasets. The datasets are shown in Table 1. The experimental results are the average of the 10 experiments and compared with 12 state-of-the-art algorithm models, such as VIDEVAL, PVQ, ChipQA, RAPIQUE, BVQA-2022, 2BiVQA, HVS-5M, FAST-VQA, DisCoVQA, DOVER, CONVIQT, and ReLaX-VQA. The comparison indicators are PLCC and SROCC to ensure the scientificity and objectivity of the evaluation results. The experimental results are shown in Table 2, which shows the performance of each model on the three datasets. In order to better identify the model performance, the best result is marked in red bold and the second best result is marked in black bold. Through experimental analysis and comparison with the 12 mainstream algorithms, it is verified that the proposed model has better performance and accuracy in video quality evaluation, which fully proves that each part has a certain contribution to the overall performance of the model.
[0033]
[0034]
[0035] Through the observation and analysis of Table 2, the proposed model has better performance than other mainstream methods, and its average performance on the three datasets ranks first. Specifically, on the CVD2014 dataset, the proposed model ranks first in PLCC and second in SROCC. FAST-VQA considers local quality through grid patch sampling, and covers global quality through contextual relationships by sampling mini patches in a uniform grid, further building Fragment Attention. The network is designed to achieve efficient end-to-end deep VQA and learn effective video quality-related representations. ReLaX-VQA uses fragments of residual frames and optical flow, as well as different expressions of spatial features of sampled frames, to enhance motion and spatial perception. In addition, the model enhances features by adopting layer stacking technology in deep neural networks. On LIVE-VQC and KoNViD-1k, DOVER ranks second overall. DOVER analyzes quality from both technical and aesthetic perspectives, breaking the tradition of only analyzing perceptual quality at the technical level and introducing aesthetic analysis to compensate for the aesthetic perspective. The aesthetic branch retains the semantics and composition of the original video through spatial downsampling and temporal sparse frame sampling, obtaining an aesthetic view. The video clips are analyzed in the technical branch, and the two branches are combined to achieve better performance. Overall, the PLCC and SROCC proposed by this model reached an average of 0.904 and 0.890, which are approximately 1.573% and 0.674% higher than the second-ranked model DOVER's PLCC and SROCC, respectively, indicating that this model has better performance and generalization ability.
[0036] To further validate the effectiveness of each component within the overall model, three ablation experiments were conducted on various datasets, focusing on the two feature extraction components and the temporal modeling component. Specifically, to verify the importance of spatial feature extraction, in the first ablation experiment, only motion information was extracted and applied to the subsequent temporal modeling process for perceptual quality prediction. In contrast to the first ablation experiment, in the second ablation experiment, only spatial features were extracted for temporal modeling before quality prediction, discarding temporal motion information to verify that motion features contribute to the model. The specific data results of the two ablation experiments are shown in Table 3, where SF represents the spatial feature extraction component and MF represents the motion feature extraction component. In the previous discussion, various time series prediction methods have been thoroughly explored, including but not limited to long short-term memory networks, gated recurrent units, and temporal convolutional networks. Therefore, to further verify that TCN can improve the performance of this chapter's model, in the third ablation experiment, the impact of three temporal modeling methods, LSTM, GRU, and TCN, on the model's prediction results are systematically analyzed. The experimental data are shown in Table 4, with the best-performing combination indicated in bold black.
[0037]
[0038]
[0039] Through the observation and analysis of Table 3 and Table 4, we can see that in the first ablation experiment, when the model structure only has the spatial feature extraction part, it performs best on the CVD2014 dataset, and the evaluation indicators PLCC and SROCC reach 0.903 and 0.901 respectively. Compared with the same dataset CVD2014 with only the motion feature extraction part, the two indicators are 0.291 and 0.298 higher. In the second ablation experiment, when the model structure only has the motion feature extraction part, the PLCC and SROCC that show the best performance on the KoNViD-1k dataset are only 0.701 and 0.690. It can be seen that in video quality evaluation, spatial features are more important than motion features. Dynamic features are more important. When spatial features and motion features are combined for time series modeling, the overall effect of the model reaches the best, which shows that spatial features and motion features each have certain positive contributions to the model performance, and once again verifies the effectiveness and comprehensiveness of this model. In order to verify that the TCN time series modeling method is better than LSTM and GRU in the model of this chapter, in the third ablation experiment, we compared the model prediction index effects when using different time series prediction methods on three data sets. The experimental results show that TCN achieved better performance in the model of this chapter. This is due to the fact that TCN uses convolutional layers with different expansion rates, which can capture patterns and long-term dependencies of different scales in the sequence.
[0040] Working principle of the present invention:
[0041] First of all, this model mainly consists of four parts: spatial feature extraction, motion feature extraction, time series modeling and quality prediction. Specifically, in order to obtain rich feature representation, all frames of the input video are first obtained and all video frame images are converted into RGB and YIQ color spaces. Secondly, the feature extraction part of the pre-trained ResNet-50 network is used to extract spatial feature information of each frame in these two color spaces. Then, in order to more accurately capture the key details and position information in each video frame, the CA attention mechanism is applied to the feature map obtained after feature extraction to dynamically adjust the different positions in the feature map. In order to further enhance the expressive power of features, in addition, considering the characteristics of video as time series data, the motion information in the time series is supplemented by calculating the Euclidean distance between the pixels of adjacent frames in the video sequence. This not only captures the changes between frames, but also reflects the dynamic characteristics of the video content, providing a more comprehensive perspective for video quality assessment. Finally, the extracted spatial features and the calculated motion features are fused to construct a feature vector containing rich information. TCN is used to perform time series modeling on the fused feature vector, and a fully connected layer is used to reduce the dimension to achieve video quality score prediction.
[0042] Finally, a few points should be explained: First, in the description of this application, it should be noted that, unless otherwise specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense, and may refer to mechanical or electrical connections, internal communication between two components, or direct connection. "Up," "down," "left," and "right" are only used to indicate relative positional relationships. When the absolute positions of the objects being described change, the relative positional relationships may also change.
[0043] Secondly: The drawings of the embodiments disclosed in the present invention only involve structures related to the embodiments disclosed in the present invention. Other structures may refer to conventional designs. The same embodiment and different embodiments of the present invention may be combined with each other without conflict.
[0044] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A video quality assessment method with rich information representation, characterized by: Including spatial feature extraction, motion feature extraction, time series modeling and quality prediction; Spatial feature extraction: Traverse each frame of the video to obtain all RGB image frames, then convert all RGB image frames to YIQ color space to obtain all image frames in YIQ color space. After that, use the modified ResNet-50 network pre-trained on ImageNet to extract spatial features on RGB and YIQ of all frames of the video at the same time. The modified network model removes the last two layers of the ResNet-50 network, namely the average pooling layer and the fully connected layer, outputs the features, and introduces the CA attention mechanism. The CA attention mechanism dynamically adjusts the feature weights of different positions by analyzing the horizontal and vertical coordinate information of each position in the feature map. Then, the features after the CA attention mechanism are adaptively average pooled and flattened. Finally, the feature vectors extracted from the two color spaces are spliced in the column direction to obtain the spatial feature vector. Motion feature extraction: Each frame of the video sequence is converted into a grayscale image. By simplifying the image information, the interference of color on motion feature extraction is removed. Then, the Euclidean distance between corresponding pixels of adjacent grayscale images is calculated to capture the motion information in the video. The calculated Euclidean distance of all pixels is averaged and pooled to obtain the average Euclidean distance between adjacent grayscale images, which reflects the overall motion intensity and degree of change between video frames. The extracted motion feature vector is then adjusted to be spliced with the spatial domain feature on the column level for subsequent time series modeling. Time series modeling: By extracting the spatial features and motion features, and fusing the two features, a new feature vector is obtained. This new feature vector is input into the TCN for time series modeling and dimensionality reduction output. Quality prediction: A fully connected layer is constructed after TCN to directly map the feature vector after dimensionality reduction to the video quality score. Specifically, the feature vector output by TCN is input into a fully connected layer. The output dimension of the fully connected layer is 1, which represents the quality score of the video. The calculation process is as follows: , where Score is the video score of the evaluation, FC is the fully connected layer, and F represents the fused spatial features and motion feature vectors.
2. The video quality assessment method with rich information representation according to claim 1, characterized in that: The spatial feature extraction part uses a pre-trained ResNet-50 network to extract features in the RGB and YIQ color spaces of the video frame, and introduces the CA attention mechanism to enhance the spatial features.
3. The video quality assessment method with rich information representation according to claim 1, characterized in that: The motion feature extraction part calculates the Euclidean distance between pixels of adjacent frames to represent the degree of motion change in the video. The extracted spatial features and the calculated motion features are fused to construct a feature vector with information representation.
4. The video quality assessment method with rich information representation according to claim 3, characterized in that: The fused feature vector is input into the time domain convolutional network for temporal modeling, and the feature dimension is reduced to 128. The 128-dimensional feature is directly mapped to the video quality score through a fully connected layer.
5. The video quality assessment method with rich information representation according to claim 1, characterized in that: The spatial feature extraction part uses the different characteristics of video frames in RGB and YIQ color spaces to extract the features of video frames in the spatial domain, and extracts the motion features of the video to supplement the temporal motion information.
6. The video quality assessment method with rich information representation according to claim 1, characterized in that: The formula for the spatial feature extraction process is as follows: , where ResNet-50 is the training network, is the spatial feature, is the input video frame sequence, i=1, 2, ..., T, T is the total number of video frames, and after passing through the ResNet-50 network, the video frame has 2048-dimensional spatial domain features; The calculation process of the CA attention mechanism is as follows: for an input video frame feature map x, the CA attention mechanism selects the input feature map and performs pooling operations in the horizontal and vertical directions respectively. The calculation formula is as follows: 、 ,in represents the Cth channel of the input feature map, is the relevant output after the C-th channel pooling, Z c h Represents the horizontal output after the C-th channel pooling, Z c w Represents the vertical output after the C-th channel pooling, x c (h, i) represents the i-th pixel with height h in the C-th channel, x c (j, w) represents the jth pixel of the Cth channel with a width of w. The feature maps after pooling in the horizontal and vertical directions are then spliced together, and a 1×1 convolution operation and activation operation are performed. The calculation formula is as follows: ,in, represents a 1×1 convolution dimension reduction function, Indicates activation operation, Z h , Z w Represents the horizontal and vertical outputs of all channels after pooling, thereby obtaining a feature map f of size C / r×(H+W)×1. Then, f is separated into feature maps F in the horizontal and vertical directions in the spatial dimension. h and F w , and then perform a 1×1 convolution dimension increase operation. The calculation formula is as follows: 、 ,in, and are the convolutions in the horizontal and vertical directions respectively, is the Sigmoid activation function, so we get the attention vector and ,Finally, the output formula of the CA attention mechanism is as follows: , where x c (i, j) is the pixel of channel C, g c h (i) and g c w (j) are the horizontal vector and vertical vector of the attention vector of the C-th channel, respectively, so that the spatial domain features of the entire video The size is T×4096.
7. The video quality assessment method with rich information representation according to claim 1, characterized in that: The conversion formula for converting each frame of the video sequence into a grayscale image is: , where R, G, and B represent the three color channels of the image; The calculation formula that reflects the overall motion intensity and degree of change between video frames is as follows: ,in represents the grayscale image of the t-th frame, Represents the grayscale image of the t+1th frame, D represents the global average Euclidean distance, and the total number of frames of the input video is T. Therefore, the Euclidean distance values of T-1 adjacent frames can be obtained. Since the value range of the grayscale image is [0, 255], the calculated Euclidean distance value is also between 0 and 255. Then the T-1 Euclidean distances are normalized. This process will use Min-Max normalization. At this time, the Euclidean distance value is mapped to [0, 1]. The conversion formula is as follows: ,in, (i=1, 2, ..., T-1) is the Euclidean distance between the i-th frame and the i+1-th frame, min(D) and max(D) represent their maximum and minimum values respectively. Therefore, the motion characteristics of the entire video are The size is T-1×1. So far, the spatial features of the entire video size T×4096 and the motion features of T-1×1 are extracted; The process of adjusting the extracted motion feature vector is as follows: first, the first row of the motion feature vector is filled with 0 to make its size T×1, thereby matching the T×4096 dimension of the spatial feature. Then, the adjusted motion feature vector and the spatial feature are spliced on the column to construct a feature matrix of size T×4097. The fused feature matrix contains both the spatial information of the video and the motion information between frames.
8. The video quality assessment method with rich information representation according to claim 1, characterized in that: The calculation process of inputting the feature vector into TCN for time series modeling is as follows: , where F k and F d They are spatial features and motion feature vectors respectively.
Citation Information
Cited By
Park abnormal behavior few-sample real-time monitoring method and system
CN121259512A
A park abnormal behavior few-sample real-time monitoring method and system
CN121259512B