No-Reference Screen Video Quality Assessment Method and Device Based on Multi-Scale Features and Channel Attention
Through the referenceless screen video quality evaluation method with multi-scale features and channel attention, the accuracy and stability of screen video quality evaluation in the prior art is solved, and efficient screen video quality evaluation is achieved, which is in line with human visual characteristics.
Patent Information
- Application Number
- CN202311112440.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-17
- Filing Date
- 2023-08-31
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-08-31
AI Technical Summary
The existing video quality evaluation algorithms are mainly concentrated in the field of natural videos, and it is difficult to effectively evaluate the quality of screen content videos. The evaluation algorithm without reference is poor in screen videos and cannot meet user experience needs.
The video quality evaluation method of referenceless screen video based on multi-scale features and channel attention is adopted, and video frames are obtained through random sampling, and multi-scale features are extracted using the pre-trained VGG16 model, and the video quality score is calculated in combination with the channel attention module and the video timing feature extraction module.
It improves the accuracy and stability of screen video quality evaluation, conforms to human visual characteristics, can effectively evaluate the quality of screen video, and reduces the memory space requirement.
Smart Images

Figure CN117173609B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a no-reference screen video quality evaluation method and device based on multi-scale features and channel attention. Background Art
[0002] With the rapid development of mobile Internet and portable communication devices, a large number of new media forms have emerged. A new type of video data type represented by screen video content is widely used in scenarios such as game live broadcast, online conferencing, and online education. Research on screen video quality evaluation has always been a hot issue in the field of computer vision. Different from traditional natural videos, screen content videos mainly refer to videos generated by computers, usually including computer graphics texts, mixed scenes of natural scenes and graphics texts, and computer-generated animations, etc., containing complex textures and sharp edge features.
[0003] During the processes of acquisition, transmission, and display of screen videos, various distortions usually occur, affecting the video quality. These distortions will affect the user experience and reduce the visually perceived subjective effect. Therefore, it is very important to propose an algorithm that conforms to the human visual characteristics and can accurately and quickly evaluate the quality of screen videos.
[0004] At present, most video quality evaluation algorithms mainly focus on the field of natural videos and are mainly full-reference video quality evaluations. However, due to the different spatio-temporal characteristics of screen content videos and natural videos, directly migrating natural video-related quality evaluation algorithms to screen content videos has relatively poor effects, and no-reference video quality evaluation is more practically significant. Therefore, designing a quality evaluation algorithm that conforms to human visual characteristics and the characteristics of screen videos has important theoretical research significance and practical application value. Summary of the Invention
[0005] Aiming at the above-mentioned technical problems, the purpose of the embodiments of the present application is to propose a no-reference screen video quality evaluation method and device based on multi-scale features and channel attention to solve the technical problems mentioned in the above background art section.
[0006] In a first aspect, the present invention provides a no-reference screen video quality evaluation method based on multi-scale features and channel attention, including the following steps:
[0007] Obtain video frames randomly sampled from the video;
[0008] Build a video quality evaluation model and train it to obtain a trained video quality evaluation model. The video quality evaluation model includes a feature extraction module, a channel attention module, a video temporal feature extraction module, and an average pooling layer connected in sequence. The feature extraction module is used to extract multi-scale features in video frames. The channel attention module is used to weight the multi-scale features. The video temporal feature extraction module is used to extract features to obtain spatio-temporal dimensional features, and calculate the quality score corresponding to the video through the average pooling layer. The channel attention module includes an adaptive average pooling layer, an adaptive max pooling layer, two three-dimensional convolutional neural network layers, and a Sigmoid activation function layer;
[0009] Input the video frames into the trained video quality evaluation model to obtain the quality score of the video.
[0010] Preferably, the feature extraction module uses a pre-trained VGG16 model. Input the video frames into the pre-trained VGG16 model, and extract the first feature, the second feature, and the third feature from the second convolutional layer, the seventh convolutional layer, and the thirteenth convolutional layer of the pre-trained VGG16 model. The formulas are as follows:
[0011]
[0012]
[0013]
[0014] Where frame represents the video frame extracted from the video, Conv2, Conv7, Conv13 represent the corresponding second convolutional layer, seventh convolutional layer, and thirteenth convolutional layer in the pre-trained VGG16 model, i represents the number of frames of the video frames obtained at different sampling rates, represent the first feature, the second feature, and the third feature respectively.
[0015] Preferably, in the channel attention module, the multi-scale features are respectively input into the adaptive average pooling layer and the adaptive max pooling layer to obtain the multi-scale average feature and the multi-scale max feature. The multi-scale average feature and the multi-scale max feature are respectively input into two three-dimensional convolutional neural network layers to obtain the fourth feature and the fifth feature. The fourth feature and the fifth feature are combined by addition and passed through the Sigmoid activation function layer to obtain the purified feature.
[0016] Preferably, the multi-scale features are respectively input into the adaptive average pooling layer and the adaptive max pooling layer to obtain the multi-scale average feature and the multi-scale max feature. The specific operations are as follows:
[0017]
[0018]
[0019]
[0020]
[0021] Among them, AAP2d() and AMP2d() represent the adaptive average pooling operation and the adaptive maximum pooling operation respectively, represents concatenating channels, and stack represents stacking frame-level features into video features, represents the multi-scale average feature of the video, represents the multi-scale maximum feature of the video, and n represents the number of video frames obtained at different sampling rates.
[0022] Preferably, the multi-scale average feature and the multi-scale maximum feature are respectively input into two three-dimensional convolutional neural network layers to obtain a fourth feature and a fifth feature. The fourth feature and the fifth feature are combined by addition and passed through a Sigmoid activation function layer to obtain a purified feature. The specific operations are as follows:
[0023]
[0024]
[0025]
[0026]
[0027] Among them, 3D CNN represents a three-dimensional convolutional neural network, represents the multi-scale channel average feature, represents the multi-scale channel maximum feature, w represents the weight assigned to the key area, f vid represents the purified feature.
[0028] Preferably, in the video quality evaluation model, the purified feature is input into a video temporal feature extraction module to extract spatio-temporal dimensional features, and the spatio-temporal dimensional features are input into an average pooling layer to obtain the quality score of the video. The specific operations are as follows:
[0029] Q = AvgPooling(VFMNet(f vid ));
[0030] Among them, VFMNNet() represents the video temporal feature extraction module, and AvgPooling() represents the average pooling layer.
[0031] Preferably, the video temporal feature extraction module includes four three-dimensional convolution modules, an adaptive average pooling layer, and two fully connected layers connected in sequence. The three-dimensional convolution module includes a three-dimensional convolutional neural network layer, a batch normalization layer, and a ReLU activation function layer connected in sequence.
[0032] In a second aspect, the present invention provides a no-reference screen video quality evaluation device based on multi-scale features and channel attention, including:
[0033] A video frame acquisition module configured to acquire video frames randomly sampled from a video;
[0034] A model construction module configured to construct a video quality evaluation model and perform training to obtain a trained video quality evaluation model. The video quality evaluation model includes a feature extraction module, a channel attention module, a video temporal feature extraction module, and an average pooling layer connected in sequence. The feature extraction module is used to extract multi-scale features in the video frames. The channel attention module is used to perform feature weighting on the multi-scale features. The video temporal feature extraction module is used to perform feature extraction to obtain spatio-temporal dimensional features, and calculate the quality score corresponding to the video through the average pooling layer. The channel attention module includes an adaptive average pooling layer, an adaptive max pooling layer, two three-dimensional convolutional neural network layers, and a Sigmoid activation function layer;
[0035] An evaluation module configured to input the video frames into the trained video quality evaluation model to obtain the quality score of the video.
[0036] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.
[0037] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] (1) The no-reference screen video quality evaluation method based on multi-scale features and channel attention proposed by the present invention explores the influence of different sampling rates on video quality evaluation related tasks and solves the problem of high repetition degree of consecutive frames in video sequences.
[0040] (2) The no-reference screen video quality evaluation method based on multi-scale features and channel attention proposed by the present invention focuses on considering the characteristics of the human visual system and the characteristics of screen videos. Considering that the subjective perception of visual information by the human system has a hierarchical nature, a pre-trained VGG16 model is used for feature extraction, and a channel attention module is adopted to focus on some regions, which conforms to the unique visual attention mechanism of humans.
[0041] (3) The no-reference screen video quality evaluation method based on multi-scale features and channel attention proposed by the present invention fully considers the characteristics of the human visual system and the characteristics of screen videos in the spatial dimension, channel dimension, and time dimension. The algorithm has relatively high stability and robustness, and has a good screen video quality evaluation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0043] Figure 1 is an exemplary device architecture diagram to which an embodiment of the present application can be applied;
[0044] Figure 2 is a schematic flowchart of the no-reference screen video quality evaluation method based on multi-scale features and channel attention according to the embodiment of the present application;
[0045] Figure 3 is a schematic diagram of the video quality evaluation model of the no-reference screen video quality evaluation method based on multi-scale features and channel attention according to the embodiment of the present application;
[0046] Figure 4 is a schematic diagram of the feature extraction module of the no-reference screen video quality evaluation method based on multi-scale features and channel attention according to the embodiment of the present application;
[0047] Figure 5 is a schematic diagram of the channel attention module of the no-reference screen video quality evaluation method based on multi-scale features and channel attention according to the embodiment of the present application;
[0048] Figure 6 is a schematic diagram of the video temporal feature extraction module of the no-reference screen video quality evaluation method based on multi-scale features and channel attention according to the embodiment of the present application;
[0049] Figure 7Schematic diagram of a no-reference screen video quality evaluation device based on multi-scale features and channel attention according to an embodiment of the present application;
[0050] Figure 8 It is a schematic structural diagram of a computer device of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0051] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] Figure 1 An exemplary device architecture 100 is shown that can apply the no-reference screen video quality evaluation method based on multi-scale features and channel attention or the no-reference screen video quality evaluation device based on multi-scale features and channel attention according to the embodiments of the present application.
[0053] As Figure 1 shown, the device architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0054] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various applications, such as data processing applications, file processing applications, etc., may be installed on the terminal devices 101, 102, 103.
[0055] The terminal devices 101, 102, 103 may be hardware or software. When the terminal devices 101, 102, 103 are hardware, they may be various electronic devices, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they may be installed in the above-listed electronic devices. It may be implemented as multiple software or software modules (for example, software or software modules for providing distributed services), or may be implemented as a single software or software module. No specific limitation is made here.
[0056] Server 105 may be a server that provides various services, such as a background data processing server that processes files or data uploaded by terminal devices 101, 102, and 103. The background data processing server may process the acquired files or data to generate a processing result.
[0057] It should be noted that the no-reference screen video quality evaluation method based on multi-scale features and channel attention provided in the embodiments of the present application can be executed by server 105, or can be executed by terminal devices 101, 102, and 103. Correspondingly, the no-reference screen video quality evaluation device based on multi-scale features and channel attention can be set in server 105, or can be set in terminal devices 101, 102, and 103.
[0058] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0059] Figure 2 FIG. shows a no-reference screen video quality evaluation method provided by an embodiment of the present application, including the following steps:
[0060] S1, obtaining video frames randomly sampled from the video.
[0061] Specifically, video frames are screened from the video in a random sampling manner to solve the problem of redundant information in the video. Embodiments of the present application use four different sampling rates for comparative experiments to screen video frames and explore their impact on video quality evaluation-related tasks. First, the input screen video sequence is subjected to comparative experiments at four different sampling rates of every 5 frames, every 10 frames, every 15 frames, and every 20 frames to screen video frames, and the corresponding number of frames of the collected video frames is 60, 30, 20, and 15.
[0062] S2, constructing and training a video quality evaluation model to obtain a trained video quality evaluation model. The video quality evaluation model includes a feature extraction module, a channel attention module, a video temporal feature extraction module, and an average pooling layer connected in sequence. The feature extraction module is used to extract multi-scale features in the video frames. The channel attention module is used to perform feature weighting on the multi-scale features. The video temporal feature extraction module is used to perform feature extraction to obtain spatio-temporal dimensional features, and the quality score corresponding to the video is calculated through the average pooling layer. The channel attention module includes an adaptive average pooling layer, an adaptive max pooling layer, two three-dimensional convolutional neural network layers, and a Sigmoid activation function layer.
[0063] In a specific embodiment, the feature extraction module uses a pre-trained VGG16 model. The video frames are input into the pre-trained VGG16 model, and the first feature, the second feature, and the third feature are extracted from the second convolutional layer, the seventh convolutional layer, and the thirteenth convolutional layer of the pre-trained VGG16 model. The formula is as follows:
[0064]
[0065]
[0066]
[0067] Where frame represents the video frames extracted from the video, Conv2, Conv7, and Conv13 represent the corresponding second convolutional layer, seventh convolutional layer, and thirteenth convolutional layer in the pre-trained VGG16 model, and i represents the number of frames of the video frames obtained at different sampling rates. respectively represent the first feature, the second feature, and the third feature.
[0068] Specifically, referring to Figure 3 and Figure 4 , considering that the subjective perception process of the human visual system for visual information has a hierarchical nature, a pre-trained VGG16 model is used to extract the texture features, contour features, and local key features in the video from the second, seventh, and thirteenth convolutional layers respectively.
[0069] In a specific embodiment, in the channel attention module, the multi-scale features are respectively input into the adaptive average pooling layer and the adaptive max pooling layer to obtain the multi-scale average feature and the multi-scale max feature. The multi-scale average feature and the multi-scale max feature are respectively input into two three-dimensional convolutional neural network layers to obtain the fourth feature and the fifth feature. The fourth feature and the fifth feature are combined by addition and passed through the Sigmoid activation function layer to obtain the purified feature.
[0070] In a specific embodiment, the multi-scale features are respectively input into the adaptive average pooling layer and the adaptive max pooling layer to obtain the multi-scale average feature and the multi-scale max feature. The specific operations are as follows:
[0071]
[0072]
[0073]
[0074]
[0075] Among them, AAP2d() and AMP2d() represent the adaptive average pooling operation and the adaptive max pooling operation respectively, represents concatenating channels, and stack represents stacking frame-level features into video features, represents the multi-scale average feature of the video, represents the multi-scale max feature of the video, and n represents the number of video frames obtained at different sampling rates.
[0076] In a specific embodiment, the multi-scale average feature and the multi-scale max feature are respectively input into two three-dimensional convolutional neural network layers to obtain a fourth feature and a fifth feature. The fourth feature and the fifth feature are combined by addition and passed through a Sigmoid activation function layer to obtain a purified feature. The specific operations are as follows:
[0077]
[0078]
[0079]
[0080]
[0081] Among them, 3D CNN represents a three-dimensional convolutional neural network, represents the multi-scale channel average feature, represents the multi-scale channel max feature, w represents the weight assigned to the key region, f vid represents the purified feature.
[0082] Specifically, considering that the visual attention mechanism is a signal processing mechanism unique to the human brain, a channel attention module is used to adaptively assign weights to key regions. Referring to Figure 5 , the first feature, the second feature, and the third feature extracted by the feature extraction module are respectively input into an adaptive average pooling layer and an adaptive max pooling layer, and two pooling methods, namely the adaptive average pooling function and the adaptive max pooling function, are used for pooling processing to ensure dimension alignment during concatenation, obtaining a multi-scale average feature and a multi-scale max feature. The multi-scale average feature and the multi-scale max feature are input into two 3D CNNs. First, the channel dimension of the input video tensor is reduced, then the video channel dimension is enlarged, and then they are combined by addition and passed through a Sigmoid activation function layer to introduce non-linearity. Finally, a purified feature f vid .
[0083] In a specific embodiment, in the video quality evaluation model, the purified features are input into the video temporal feature extraction module to extract spatio-temporal dimensional features. The spatio-temporal dimensional features are input into the average pooling layer to obtain the quality score of the video. The specific operations are as follows:
[0084] Q = AvgPooling(VFMNet(f vid ));
[0085] where VFMNet() represents the video temporal feature extraction module, and AvgPooling represents the average pooling layer.
[0086] In a specific embodiment, the video temporal feature extraction module includes four three-dimensional convolutional modules, an adaptive average pooling layer, and two fully connected layers connected in sequence. The three-dimensional convolutional module includes a three-dimensional convolutional neural network layer, a batch normalization layer, and a ReLU activation function layer connected in sequence.
[0087] Specifically, referring to Figure 6 , considering the importance of the time dimension for video tasks, a video temporal feature extraction module mainly based on 3DCNN with time dimension convolution is used to extract video temporal features. Finally, the quality score of the video is obtained through an average pooling operation, realizing the modeling of the time dimension features of the video sequence and the mapping of the video quality score, and having a good screen video quality evaluation effect.
[0088] Collect a distorted screen content dataset, divide the distorted screen content dataset into a training set, a test set, and a validation set in a ratio of 6:2:2, and use the distorted screen content dataset to train the video quality evaluation model. The loss function used during the training process is L1 Loss.
[0089] S3. Input the video frames into the trained video quality evaluation model to obtain the quality score of the video.
[0090] Specifically, input the extracted video frames into the trained video quality evaluation model to evaluate the quality of the video and obtain the quality score of the video.
[0091] Embodiments of this application verify the effectiveness of the channel attention mechanism in the proposed video quality evaluation model and the effect of the random sampling strategy on saving memory space. Here, the experimental results under four sampling rates are compared, and the comparison results are shown in Table 1. The best experimental results are shown in bold black. CA represents the channel attention module, and PLCC, SROCC, and RMSE are performance indicators for objective video quality evaluation. PLCC is mainly used to measure the prediction accuracy of objective algorithms, and its value range is [-1, 1]. The closer it is to 1, the better the algorithm performance. SROCC is mainly used to measure the consistency between the change trend of the objective quality evaluation score and the change trend of the MOS value, that is, the monotonicity of the objective algorithm. The value range is [-1, 1]. The closer it is to 1, the better the algorithm performance. RMSE represents the absolute error between the score obtained by the objective quality evaluation algorithm and the MOS value, and is used to measure the accuracy of the objective algorithm. The smaller its value, the better the algorithm performance.
[0092] Table 1 Experimental results of the proposed algorithm on the SCVD database
[0093]
[0094] As can be seen from Table 1, when the sampling rate is one frame per 10 frames and the channel attention is added, the experimental effect reaches the best. This method explores the influence of different sampling rates on video quality evaluation-related tasks respectively, then uses the pre-trained VGG16 model to extract multi-scale frame-level features, and uses the channel attention module to perform weighted purification on video features in the channel dimension. Finally, by designing a video temporal feature extraction module mainly based on temporal convolution, the modeling of the temporal dimension features of the video sequence and the mapping of the video quality score are realized, and it has a good screen video quality evaluation effect.
[0095] The above steps S1 - S3 do not represent the order between steps, but are only step symbol representations.
[0096] Further referring to Figure 7 , as an implementation of the methods shown in the above figures, this application provides an embodiment of a no-reference screen video quality evaluation device based on multi-scale features and channel attention. This device embodiment corresponds to the method embodiment shown in Figure 2 , and this device can be specifically applied to various electronic devices.
[0097] This application embodiment provides a no-reference screen video quality evaluation device based on multi-scale features and channel attention, including:
[0098] A video frame acquisition module 1, configured to acquire video frames randomly sampled from the video;
[0099] The model construction module 2 is configured to construct a video quality evaluation model and perform training to obtain a trained video quality evaluation model. The video quality evaluation model includes a feature extraction module, a channel attention module, a video temporal feature extraction module, and an average pooling layer connected in sequence. The feature extraction module is used to extract multi-scale features in video frames. The channel attention module is used to perform feature weighting on the multi-scale features. The video temporal feature extraction module is used to perform feature extraction to obtain spatio-temporal dimensional features, and calculate the quality score corresponding to the video through the average pooling layer. The channel attention module includes an adaptive average pooling layer, an adaptive maximum pooling layer, two three-dimensional convolutional neural network layers, and a Sigmoid activation function layer;
[0100] The evaluation module 3 is configured to input video frames into the trained video quality evaluation model to obtain the quality score of the video.
[0101] The following refers to Figure 8 , which shows a schematic structural diagram of a computer device 800 suitable for implementing the embodiments of the present application (such as Figure 1 the server or terminal device shown). Figure 8 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0102] As Figure 8 shown, the computer device 800 includes a central processing unit (CPU) 801 and a graphics processing unit (GPU) 802, which can execute various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 803 or the programs loaded from the storage section 809 into the random access memory (RAM) 804. In the RAM 804, various programs and data required for the operation of the device 800 are also stored. The CPU 801, GPU 802, ROM 803, and RAM 804 are connected to each other through a bus 805. The input / output (I / O) interface 806 is also connected to the bus 805.
[0103] The following components are connected to the I / O interface 806: an input section 807 including a keyboard, a mouse, etc.; an output section 808 including, for example, a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 809 including a hard disk, etc.; and a communication section 810 including a network interface card such as a LAN card, a modem, etc. The communication section 810 performs communication processing via a network such as the Internet. The drive 811 can also be connected to the I / O interface 806 as needed. A removable medium 812, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 811 as needed so that the computer program read from it can be installed into the storage section 809 as needed.
[0104] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 810, and / or installed from the removable medium 812. When the computer program is executed by the central processing unit (CPU) 801 and the graphics processing unit (GPU) 802, the above functions defined in the method of the present application are executed.
[0105] It should be noted that the computer-readable medium described in the present application can be a computer-readable signal medium, a computer-readable medium, or any combination of the two. The computer-readable medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor device, apparatus, or component, or any combination of the above. More specific examples of the computer-readable medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution device, apparatus, or component. And in the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution device, apparatus, or component. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0106] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can also be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0107] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of apparatuses, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based device that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0108] The modules described in the embodiments of this application can be implemented in software or in hardware. The described modules can also be provided in a processor.
[0109] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: obtain video frames randomly sampled from the video; construct a video quality evaluation model and train it to obtain a trained video quality evaluation model. The video quality evaluation model includes a feature extraction module, a channel attention module, a video temporal feature extraction module, and an average pooling layer connected in sequence. The feature extraction module is used to extract multi-scale features from the video frames. The channel attention module is used to perform feature weighting on the multi-scale features. The video temporal feature extraction module is used to perform feature extraction to obtain spatio-temporal dimensional features, and calculate the quality score corresponding to the video through the average pooling layer. The channel attention module includes an adaptive average pooling layer, an adaptive max pooling layer, two three-dimensional convolutional neural network layers, and a Sigmoid activation function layer; input the video frames into the trained video quality evaluation model to obtain the quality score of the video.
[0110] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solution formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present application.
Claims
1. A no-reference screen video quality assessment method based on multi-scale features and channel attention, characterized in that Including the following steps: Obtain video frames randomly sampled from the video; Construct a video quality evaluation model and train it to obtain a trained video quality evaluation model. The video quality evaluation model includes a feature extraction module, a channel attention module, a video temporal feature extraction module, and an average pooling layer connected in sequence. The feature extraction module is used to extract multi-scale features from the video frames. The channel attention module is used to perform feature weighting on the multi-scale features. The video temporal feature extraction module is used to perform feature extraction to obtain spatio-temporal dimensional features, and the average pooling layer is used to calculate the quality score corresponding to the video. The channel attention module includes an adaptive average pooling layer, an adaptive max pooling layer, two 3D convolutional neural network layers, and a Sigmoid activation function layer. The multi-scale features are respectively input into the adaptive average pooling layer and the adaptive max pooling layer to obtain multi-scale average features and multi-scale max features. The multi-scale average features and multi-scale max features are respectively input into two 3D convolutional neural network layers to obtain a fourth feature and a fifth feature. The fourth feature and the fifth feature are combined by addition and passed through the Sigmoid activation function layer to obtain purified features. The purified features are input into the video temporal feature extraction module to extract spatio-temporal dimensional features. The spatio-temporal dimensional features are input into the average pooling layer to obtain the quality score of the video. The specific operations are as follows: Q = AvgPooling(VFMNet(f vid )); Wherein, VFMNet( ) represents the video temporal feature extraction module, and AvgPooling( ) represents the average pooling layer; Input the video frames into the trained video quality evaluation model to obtain the quality score of the video.
2. The no-reference screen video quality evaluation method based on multi-scale features and channel attention according to claim 1, wherein The feature extraction module uses a pre-trained VGG16 model. Input the video frames into the pre-trained VGG16 model, and extract a first feature, a second feature, and a third feature from the second convolutional layer, the seventh convolutional layer, and the thirteenth convolutional layer of the pre-trained VGG16 model. The formula is as follows: f i 2 = VGG16(ReLU((Conv2(frame)))); f i 7 = VGG16(ReLU((Conv7(frame)))) f i 13 = VGG16(ReLU((Conv13(frame)))); Among them, frame represents the video frame extracted from the video, Conv2, Conv7, and Conv13 represent the corresponding second convolutional layer, seventh convolutional layer, and thirteenth convolutional layer in the pre-trained VGG16 model, i represents the number of video frames obtained at different sampling rates, and f i 2 , f i 7 , f i 13 represent the first feature, the second feature, and the third feature respectively.
3. The no-reference screen video quality evaluation method based on multi-scale features and channel attention according to claim 1, characterized in that The multi-scale features are respectively input into the adaptive average pooling layer and the adaptive max pooling layer to obtain multi-scale average features and multi-scale max features. The specific operations are as follows: Among them, AAP2d() and AMP2d() represent the adaptive average pooling operation and the adaptive maximum pooling operation respectively, represents concatenating channels, and stack represents stacking frame-level features into video features, represents the multi-scale average feature of the video, represents the multi-scale maximum feature of the video, and n represents the number of video frames obtained at different sampling rates.
4. The no-reference screen video quality evaluation method based on multi-scale features and channel attention according to claim 3, characterized in that The multi-scale average features and multi-scale max features are respectively input into two 3D convolutional neural network layers to obtain a fourth feature and a fifth feature. The fourth feature and the fifth feature are combined by addition and passed through the Sigmoid activation function layer to obtain purified features. The specific operations are as follows: Among them, 3D CNN represents a three-dimensional convolutional neural network, represents the multi-scale channel average feature, represents the multi-scale channel maximum feature, w represents the weight assigned to the key area, and f vid represents the purified feature.
5. The no-reference screen video quality evaluation method based on multi-scale features and channel attention according to claim 1, characterized in that The video temporal feature extraction module includes four 3D convolutional modules, an adaptive average pooling layer, and two fully connected layers connected in sequence. The 3D convolutional module includes a 3D convolutional neural network layer, a batch normalization layer, and a ReLU activation function layer connected in sequence.
6. A no-reference screen video quality evaluation device based on multi-scale features and channel attention, characterized in that Including: A video frame acquisition module configured to obtain video frames randomly sampled from the video; A model construction module, configured to construct and train a video quality evaluation model to obtain a trained video quality evaluation model. The video quality evaluation model includes a feature extraction module, a channel attention module, a video temporal feature extraction module, and an average pooling layer connected in sequence. The feature extraction module is used to extract multi-scale features in the video frame. The channel attention module is used to perform feature weighting on the multi-scale features. The video temporal feature extraction module is used to extract features to obtain spatio-temporal dimensional features, and calculate the quality score corresponding to the video through the average pooling layer. The channel attention module includes an adaptive average pooling layer, an adaptive max pooling layer, two three-dimensional convolutional neural network layers, and a Sigmoid activation function layer. The multi-scale features are respectively input into the adaptive average pooling layer and the adaptive max pooling layer to obtain multi-scale average features and multi-scale max features. The multi-scale average features and multi-scale max features are respectively input into the two three-dimensional convolutional neural network layers to obtain a fourth feature and a fifth feature. The fourth feature and the fifth feature are combined by addition and passed through the Sigmoid activation function layer to obtain a purified feature. The purified feature is input into the video temporal feature extraction module to extract spatio-temporal dimensional features. The spatio-temporal dimensional features are input into the average pooling layer to obtain the quality score of the video. The specific operations are as follows: Q = AvgPooling(VFMNet(f vid )); Wherein, VFMNet( ) represents the video temporal feature extraction module, and AvgPooling( ) represents the average pooling layer; An evaluation module, configured to input the video frame into the trained video quality evaluation model to obtain the quality score of the video.
7. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-5.
Citation Information
Patent Citations
No-reference video quality evaluation method based on spatio-temporal multi-scale analysis
CN113313682A
No-reference video quality evaluation method and device, equipment and storage medium
CN113888502A