Non-reference rendering video quality evaluation method and system based on deep learning

Through the reference-free rendered video quality evaluation method based on deep learning, using video clip feature sequences and motion estimation, the problem of difficult distortion types such as jagged, molar and flicker in the rendered video is solved, and more accurate rendered video quality evaluation is achieved, assisting users in optimizing rendering settings.

CN120047384AActive Publication Date: 2025-05-27GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411962735.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-27
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

Existing video quality evaluation methods are difficult to effectively evaluate distortion types such as jagged, molar and flicker in rendered videos, resulting in difficulty in direct application in the rendering field.

Method used

The reference-free rendered video quality evaluation method is adopted based on deep learning. By segmenting the video data into video clips, randomly selecting image frames to construct feature sequences, motion estimation and image differential calculations are performed, and a comprehensive evaluation is carried out by combining pre-trained image difference detectors and multi-layer perceptrons.

Benefits of technology

In the absence of reference video, the quality of the rendered video can be more accurately evaluated, especially for the common distortion types such as jagged, molar and flicker in rendered videos, providing more accurate references, helping users to trade off between picture quality and rendering overhead and optimize rendering settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047384A_ABST
    Figure CN120047384A_ABST
Patent Text Reader

Abstract

The invention provides a non-reference rendering video quality evaluation method and system based on deep learning, and the method comprises the steps: 1, segmenting video data, randomly extracting image frames, and constructing a feature sequence to obtain an image quality score; 2, segmenting the video data and performing motion estimation to obtain a second video subset; 3, performing image difference calculation on the second video subset, and inputting the second video subset into a pre-trained image difference detector and a multi-layer perceptron to obtain a quality score reflecting the time stability of the video; and 4, comprehensively evaluating the quality score to obtain a final evaluation score. According to the method, the video quality is comprehensively evaluated from the time domain and the space domain, additional optimization is carried out on the distortion types which are likely to occur in the rendered video under the condition that no reference video exists, the advantages and disadvantages of the rendered picture can be judged, a user is assisted in balancing between the picture quality and the rendering cost, and therefore setting of the rendered pictures on different platforms is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a method and system for evaluating quality of a reference-free rendered video based on deep learning. Background Art

[0002] With the development of rendering technology and the popularity of real-time rendering applications, rendered videos are becoming an important part of people's daily media. Quality assessment of these videos helps to optimize rendering settings and improve rendering quality. Subjective video quality assessment based on human eyes can obtain accurate evaluation results, but this method is expensive and inefficient.

[0003] In many applications, automatic objective video quality assessment indicators are more practical. These objective video quality assessment indicators can be divided into reference video quality assessment and non-reference video quality assessment. Among them, reference video quality assessment requires a high-quality, lossless video as a reference object. The target video and the reference video are compared to get the relative video quality score. However, in rendering applications, it is usually difficult to obtain a reference video suitable for comparison because it is difficult to reproduce exactly the same camera motion and object motion. In this case, the non-reference video quality assessment indicator is more suitable as a quality assessment standard because it does not require the video as a reference object input.

[0004] Traditional no-reference video quality assessment methods use local contrast, brightness, chromaticity and other image features to measure video quality. However, these manually designed features are difficult to perform sufficiently accurate quality assessment on complex video content. After deep learning was introduced into the field of video quality assessment, new video quality assessment methods have delegated the task of video feature extraction to deep neural networks, which improves the accuracy of the assessment. At the same time, some methods also evaluate the quality influencing factors of videos in the time domain.

[0005] For example, reference 1 (Quality assessment of in-the-wild videos [C], Proceedings of the 27th ACM international conference on multimedia, 2019: 2351-2359) provides a video quality assessment algorithm VSFA, which introduces a gated recurrent unit to simulate the human eye's perception of video quality changes, extending the dimension of video quality assessment from the spatial domain to the temporal domain. However, this method extracts information from a single image frame and then further evaluates the video quality.

[0006] For another example, the invention application with publication number CN118741086A discloses an image quality inspection method, device, vehicle and storage medium for vehicle-mounted video transmission. The method is based on a convolutional neural network and obtains original video stream data from a transmitting end in real time, and extracts each original video frame image of the original video stream data; obtains multiple second video frame images from a receiving end in real time; identifies and recognizes the multiple second video frame images to obtain second identification information corresponding to each second video frame image and multiple de-identified third video frame images; determines the first image transmission quality based on the difference between each first identification information and the corresponding second identification information; inputs each original video frame image and each third video frame image into a preset image quality assessment model, respectively, to obtain the second image transmission quality; and determines the target image transmission quality based on the first image transmission quality and the second image transmission quality.

[0007] The above video quality assessment indicators based on deep neural networks have achieved high accuracy on the video dataset captured by the camera. However, there are great differences between rendered videos and real content videos. Past video quality assessment methods focus more on distortion types such as blur, overexposure, underexposure, and noise caused by camera shooting. For rendered videos, distortions such as aliasing, moiré, and flicker caused by insufficient sampling rate in rendering require extra attention, which is ignored by past video quality assessment methods, making it difficult for these methods to be directly applied to the rendering field. Summary of the invention

[0008] The purpose of the present invention is to provide a reference-free rendering video quality assessment method and system based on deep learning, which comprehensively evaluates the video quality from the time domain and spatial domain, pays attention to the local loss in the video quality assessment, and is used to solve the distortion problems such as aliasing, moiré, and flicker caused by insufficient sampling rate in rendering, so as to provide a more accurate reference for subsequent rendering images.

[0009] To achieve the above-mentioned object of the invention, an embodiment provides a method for evaluating quality of a non-reference rendered video based on deep learning, comprising the following steps:

[0010] Step 1: Segment the video data into a number of video segments with a preset step length, randomly select a frame from each video segment to form a first data set, construct a feature sequence based on the first data set, and obtain an image quality score of the first data set;

[0011] Step 2: Evenly divide the video data into a number of first video subsets, randomly extract continuous image frames from the first video subsets and perform motion estimation to obtain motion vectors of the first video subsets, and obtain the second video subsets by reverse deformation and occlusion removal processing of the first video subsets;

[0012] Step 3: performing image difference calculation on the motion vectors of the second video subset, and inputting the difference calculation results into the pre-trained image difference detector and the first multi-layer perceptron to obtain a quality score of the video temporal stability;

[0013] Step 4: Input the image quality score of the first data set and the quality score of the video temporal stability into a second multilayer perceptron to obtain a final evaluation score of the video data, wherein the number of neurons of the second multilayer perceptron is less than the number of neurons of the first multilayer perceptron.

[0014] The present invention comprehensively evaluates the video quality from the time domain and spatial domain, obtains the video quality score in the absence of a reference video, and performs additional optimization for the types of distortion that are prone to occur in rendered videos. The final video quality score helps to judge the quality of the rendered picture, assists users in balancing image quality and rendering overhead, and thus optimizes the rendering picture settings on different platforms.

[0015] In one embodiment, dividing the video data into a plurality of video segments at a preset step length includes: dividing the total number of image frames of the video data by the frame rate of the video data to set a corresponding division step length.

[0016] In one embodiment, the method of constructing a feature sequence based on the first data set to obtain the image quality score of the first data set includes: using a pre-trained Swin Transformer to extract features from each data in the first data set, and splicing the extracted features to construct a corresponding feature sequence, and inputting the feature sequence into a nonlinear layer to output the image quality score of the first data set.

[0017] In one embodiment, the method of randomly extracting continuous image frames from the first video subset and performing motion estimation to obtain the motion vector of the first video subset includes: randomly selecting N continuous frames of images from each first video subset and cropping them to a uniform size, performing motion estimation on the first N-1 frames of images in each first video subset using a dense optical flow tracking algorithm to obtain a motion vector relative to the Nth frame, and masking pixels to which the motion vector is not mapped.

[0018] In one embodiment, the reverse deformation includes performing reverse deformation and occlusion removal processing on the first N-1 frames of images after mask annotation:

[0019] F′ k→(N) =Warping(F k , M k )·D

[0020] Among them, F′ k→(N) represents the image frame after reverse deformation, Warping(·) is the reverse deformation operation, Fk represents the kth image frame, M k It represents the motion vector estimation of the first N-1 frames relative to the Nth frame, and D represents the overlapping image in the occlusion area mark of the N-1 frame.

[0021] In one embodiment, the pre-trained image difference detector and the first multi-layer perceptron are optimized in parameters through a loss function, and the loss function is a weighted combination of a Pearson linear correlation coefficient and a ranking loss function.

[0022] In one embodiment, the Pearson linear correlation coefficient L PLCC :

[0023]

[0024] Where s represents the number of videos. represents the final evaluation score obtained by the prediction of the i-th video, qi represents the true score of the i-th video manually annotated, a represents the mean of the final evaluation scores obtained by all video predictions, and b represents the mean of the true scores of all video manually annotated;

[0025] In one embodiment, the ranking loss function L ranking :

[0026]

[0027] Where s represents the number of videos. represents the final evaluation score obtained by the j-th video prediction, q j represents the true score of the jth video manually annotated, and sgn(·) represents the sign function.

[0028] In order to clearly demonstrate the no-reference rendering video quality assessment method based on deep learning, the present invention also provides a no-reference rendering video quality assessment system based on deep learning, the no-reference rendering video quality assessment system includes an image quality assessment unit, a temporal stability assessment unit and a video quality comprehensive assessment unit;

[0029] The image quality assessment unit is used to randomly select image frames from the input video data to output corresponding image quality scores;

[0030] The temporal stability evaluation unit is used to perform inverse deformation and occlusion removal processing on the input video data, and perform image difference calculation and feature regression on the processing results to output a corresponding quality score of the video temporal stability;

[0031] The video quality comprehensive evaluation unit performs a comprehensive evaluation based on the image quality score and the quality score of the video temporal stability to output a corresponding final evaluation score.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] (1) By segmenting the video data, randomly extracting image frames from the video data and constructing a feature sequence to obtain the image quality score, we can better consider the changes of each pixel in the time domain and focus on the local quality loss of the rendered video;

[0034] (2) Perform motion estimation and image difference calculation on the video. These difference images can reflect the temporal continuity of video clips at different temporal frequencies and help detect distortion in rendered videos.

[0035] (3) The final evaluation score is obtained by comprehensively evaluating the video quality in the time domain and spatial domain. In the absence of a reference video, additional optimization is performed on the types of distortion that are prone to occur in the rendered video. This helps to judge the quality of the rendered image and assists users in balancing image quality and rendering overhead, thereby optimizing the rendering image settings on different platforms. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0037] Figure 1 A flowchart of a method for evaluating quality of a non-reference rendered video based on deep learning provided in this embodiment;

[0038] Figure 2 A schematic diagram of the process of obtaining the image quality score and the quality score of the video time stability;

[0039] Figure 3 A schematic diagram of the structure of the image difference detector and the multi-layer perceptron provided in this embodiment;

[0040] Figure 4 A schematic diagram of the structure of a deep learning-based no-reference rendering video quality assessment system provided in this embodiment. DETAILED DESCRIPTION

[0041] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the following is a

[0042] The present invention is further described in detail in the following examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.

[0043] To solve the distortion problems such as aliasing, moiré, and flickering caused by insufficient sampling rate in rendering, the embodiment provides a reference-free rendering video quality assessment method based on deep learning. Figure 1 and Figure 2 The present invention specifically introduces a method for evaluating quality of a non-reference rendered video based on deep learning, which includes the following steps:

[0044] S1. Divide the video data into a number of video segments with a preset step size, randomly select a frame from each video segment to form a first data set, construct a feature sequence based on the first data set, and obtain an image quality score of the first data set.

[0045] In the embodiment, the input video data V includes z frames and the frame rate is r. 1 , F 2 , ..., F z}, divide the video data V evenly into Fragments {C 1 , C 2 , ..., C N}, each segment lasts for 1 second, where F represents an image frame and C represents a set of evenly divided image frames;

[0046] Next, one frame is randomly selected from each segment to form the first data set {C′ 1 , C′ 2 ,…,C′ N}, these image frames are then input into the pre-trained Swin Transformer to extract features;

[0047] Finally, the extracted features are concatenated to construct a corresponding feature sequence, which is input into a nonlinear layer to output the image quality score q of the first data set. a .

[0048] S2, dividing the video data into a plurality of first video subsets, randomly extracting continuous image frames from the first video subsets and performing motion estimation to obtain motion vectors of the first video subsets, and obtaining second video subsets by reverse deformation and occlusion removal processing of the first video subsets.

[0049] In the embodiment, the video data V={F 1 , F 2 , ..., F z} is divided into 10 first video subsets containing 5 consecutive frames to construct the first video subset set {K1 , K 2 , ..., K 10}, for each first video subset K i A cropping coordinate is randomly generated, and based on this coordinate, each image frame in the first video subset is cropped into a segment {F′ j , F′ j+1 , F′ j+2 , F′ j+3 , F′ j+4}, F′ represents the cropped image frame, and j is a random number.

[0050] It should be noted that the five frames in the same first video subset still use the same cropping coordinates to maintain the temporal continuity of the five frames. i The first four frames in the image are used for motion estimation, and the motion vector estimation M of the first four frames relative to the fifth frame is obtained. k Afterwards, by marking M k The pixels without motion vectors mapped in can obtain image D with marked occluded areas.

[0051] Next, use M k Perform reverse deformation on the first 4 frames:

[0052] F′ k→(j+4) =Warping(F k , M k )·D (1)

[0053] Where k∈{j, j+1, j+2, j+3}, Warping(·) is the reverse deformation operation, and D is the four occlusion region marking images {D j , D j+1 , D j+2 , D j+3} in the overlapping part.

[0054] After the deformation is completed, 5 aligned frames can be obtained, which are recorded as the second video subset K′ i ={F′ j→(j+4) , F′ (j+1)→(j+4) , F′ (j+2)→(j+4) , F′ (j+3)→(j+4) , F′ j+4}, construct the second video subset set {K′ 1 , K′ 2 , ..., K′ 10}.

[0055] S3. Perform image difference calculation on the motion vector of the second video subset, input the difference calculation result into a pre-trained image difference detector and a first multi-layer perceptron, and obtain a quality score of the video temporal stability.

[0056] In the embodiment, Figure 3 As shown, the second video subset set {K′ 1 , K′ 2 , ..., K′ 10}, perform image difference calculation on each second video subset, perform image difference operation on each pair of adjacent frame images, and each pair of frame images with an interval of 1 frame, 2 frames, and 3 frames, thereby obtaining 10 image difference results. These difference images reflect the time continuity of video clips at different time frequencies, which are helpful to detect distortions such as flickering and moving aliasing that are easy to occur in rendering.

[0057] Afterwards, the difference results of these 10 images are input into a neural network based on depthwise separable convolution to obtain a feature vector with a length of 84.

[0058] Finally, the features of all subsets are regressed through average pooling and the first multi-layer perceptron (MLP) to obtain the quality score q reflecting the temporal stability of the video. b .

[0059] S4. Input the image quality score of the first data set and the quality score of the video temporal stability into a second multilayer perceptron to obtain a final evaluation score of the video data, wherein the number of neurons of the second multilayer perceptron is less than the number of neurons of the first multilayer perceptron.

[0060] In the embodiment, the image quality score q obtained based on S1 a and the time stability quality score q obtained by S3 b , the contribution to obtaining the final evaluation score of the video data presents a nonlinear relationship, so S4 uses a second multilayer perceptron (MLP) to map this nonlinear state. The structure of the second multilayer perceptron (MLP) in S4 is roughly the same as the first multilayer perceptron (MLP) in S3. The only difference is that the second multilayer perceptron (MLP) only uses a quarter of the number of neurons in the first multilayer perceptron (MLP) to prevent overfitting.

[0061] During the training process, the pre-trained image difference detector and the first multi-layer perceptron (MLP) are optimized through the loss function, which uses the Pearson linear correlation coefficient L PLCC And the ranking loss function L ranking ; The final evaluation score obtained by predicting multiple videos in S4 is expressed as The true scores of multiple videos manually annotated are expressed as Q = {q 1 ,q 2 , ..., q s},in, represents the final evaluation score obtained by the s-th video prediction, q s represents the true score of the manual annotation of the s-th video.

[0062] The Pearson linear correlation coefficient L was used PLCC as follows:

[0063]

[0064] Where s represents the number of videos, represents the final evaluation score obtained by the prediction of the i-th video, gi represents the true score of the i-th video manually annotated, a represents the mean of the final evaluation scores obtained by all video predictions, and b represents the mean of the true scores of all video manually annotated;

[0065] The Pearson linear correlation coefficient L was used ranking as follows:

[0066]

[0067] Where s represents the number of videos. represents the final evaluation score obtained by the prediction of the th video, represents the final evaluation score obtained by the i-th video prediction, q i represents the true score of the manual annotation of the i-th video, qj represents the true score of the manual annotation of the j-th video, and sgn(·) represents the sign function;

[0068] The final loss function LOSS is composed of the Pearson linear correlation coefficient L PLCC And the ranking loss function L ranking Weighted combination:

[0069] LOSS=L PLCC +α·L ranking (4)

[0070] Among them, α is the weight and is set to 0.3.

[0071] In order to test the actual effect of the technical solution provided by this embodiment, this embodiment is then used to evaluate on the ReVQ-2k rendering video dataset and other types of video datasets. When the dataset provides a temporal stability quality score as a label, training can be performed based on the temporal stability quality score to obtain better results. The method based on the temporal stability quality score is marked as "Ours-" in Table 1. When training is performed only based on the overall quality score, this method is marked as "Ours" in Table 1. The effect of quality assessment is measured using the two indicators SRCC and PLCC. The effect of the present invention on the rendering video dataset ReVQ-2k is shown in Table 1 (the larger the value, the better).

[0072] Table 1

[0073]

[0074] In addition, the present invention also achieved good results on the non-rendered video dataset, as shown in Table 2 (the larger the value, the better), which shows the robustness of the deep learning-based no-reference rendered video quality assessment method provided by the present invention.

[0075] Table 2

[0076]

[0077]

[0078] In order to clearly demonstrate the no-reference rendering video quality assessment method based on deep learning, this embodiment also provides a no-reference rendering video quality assessment system based on deep learning, which is used to complete the above-mentioned no-reference rendering video quality assessment method based on deep learning, such as Figure 4 As shown, the reference-free rendering video quality assessment system includes an image quality assessment unit, a temporal stability assessment unit and a video quality comprehensive assessment unit;

[0079] The image quality assessment unit is used to randomly select image frames from the input video data to output corresponding image quality scores;

[0080] The temporal stability evaluation unit is used to perform inverse deformation and occlusion removal processing on the input video data, and perform image difference calculation and feature regression on the processing results to output a corresponding quality score of the video temporal stability;

[0081] The video quality comprehensive evaluation unit performs a comprehensive evaluation based on the image quality score and the quality score of the video temporal stability to output a corresponding final evaluation score.

[0082] By using the above system, the video quality is comprehensively evaluated in the time domain and spatial domain. In the absence of reference video, additional optimization is performed on the types of distortion that are prone to occur in rendered videos. This helps to judge the quality of the rendered images and assist users in balancing image quality and rendering overhead, thereby optimizing the rendering image settings on different platforms and better implementing a reference-free rendered video quality assessment method based on deep learning.

[0083] In addition, it should be understood that after reading the above description of the present invention, those skilled in the art may make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the claims attached to this application.

Claims

1. A method for evaluating quality of no-reference rendered videos based on deep learning, characterized in that: The following steps are involved: Step 1: Segment the video data into a number of video segments with a preset step length, randomly select a frame from each video segment to form a first data set, construct a feature sequence based on the first data set, and obtain an image quality score of the first data set; Step 2: Divide the video data into a number of first video subsets, randomly extract continuous image frames from the first video subsets and perform motion estimation to obtain motion vectors of the first video subsets, and obtain the second video subsets by performing inverse deformation and occlusion removal processing on the first video subsets; Step 3: performing image difference calculation on the motion vectors of the second video subset, and inputting the difference calculation results into the pre-trained image difference detector and the first multi-layer perceptron to obtain a quality score of the video temporal stability; Step 4: Input the image quality score of the first data set and the quality score of the video temporal stability into a second multilayer perceptron to obtain a final evaluation score of the video data, wherein the number of neurons of the second multilayer perceptron is less than the number of neurons of the first multilayer perceptron.

2. The method for assessing video quality without reference rendering according to claim 1, characterized in that: The step of dividing the video data into a plurality of video segments at a preset step length includes: dividing the total number of image frames of the video data by the frame rate of the video data to set a corresponding dividing step length.

3. The method for assessing video quality without reference rendering according to claim 1, characterized in that: The method of constructing a feature sequence based on the first data set to obtain the image quality score of the first data set includes: using a pre-trained SwinTransformer to extract features from each data in the first data set, and splicing the extracted features to construct a corresponding feature sequence, and inputting the feature sequence into a nonlinear layer to output the image quality score of the first data set.

4. The method for assessing video quality without reference rendering according to claim 1, characterized in that: The method of randomly extracting continuous image frames from the first video subset and performing motion estimation to obtain the motion vector of the first video subset includes: randomly selecting N continuous frames of images from each first video subset and cropping them to a uniform size, using a dense optical flow tracking algorithm to perform motion estimation on the first N-1 frames of images in each first video subset to obtain a motion vector relative to the Nth frame, and masking and marking pixels to which the motion vector is not mapped.

5. The method for assessing video quality without reference rendering according to claim 1, characterized in that: The reverse deformation and occlusion removal process includes performing reverse deformation and occlusion removal processes on the first N-1 frames of images after mask annotation: F′ k→(N) =Warping(F k ,M k )·D Among them, F′ k→(N) represents the image frame after reverse deformation, Warping(·) is the reverse deformation operation, F k represents the kth image frame, M k It represents the motion vector estimation of the first N-1 frames relative to the Nth frame, and D represents the overlapping image in the occlusion area mark of the N-1 frame.

6. The method for assessing video quality without reference rendering according to claim 1, characterized in that: The pre-trained image difference detector and the first multi-layer perceptron are optimized in parameters through a loss function, and the loss function is formed by a weighted combination of a Pearson linear correlation coefficient and a ranking loss function.

7. The method for assessing video quality without reference rendering according to claim 6, characterized in that: The Pearson linear correlation coefficient L PLCC : Where s represents the number of videos, represents the final evaluation score obtained by the i-th video prediction, q i represents the true score of the manual annotation of the i-th video, a represents the mean of the final evaluation scores obtained by all video predictions, and b represents the mean of the true scores of all video manual annotations.

8. The method for assessing video quality without reference rendering according to claim 6, characterized in that: The ranking loss function L ranking : Where s represents the number of videos, represents the final evaluation score obtained by the j-th video prediction, represents the final evaluation score obtained by the i-th video prediction, q i represents the true score of the manual annotation of the i-th video, q j represents the true score of the jth video manually annotated, and sgn(·) represents the sign function.

9. A deep learning-based no-reference rendering video quality assessment system, characterized in that: Used to complete the deep learning-based no-reference rendering video quality assessment method according to any one of claims 1 to 8, the no-reference rendering video quality assessment system comprises an image quality assessment unit, a temporal stability assessment unit and a video quality comprehensive assessment unit; The image quality assessment unit is used to randomly select image frames from the input video data to output corresponding image quality scores; The temporal stability evaluation unit is used to perform inverse deformation and occlusion removal processing on the input video data, and perform image difference calculation and feature regression on the processing results to output a corresponding quality score of the video temporal stability; The video quality comprehensive evaluation unit performs a comprehensive evaluation based on the image quality score and the quality score of the video temporal stability to output a corresponding final evaluation score.

Citation Information

Patent Citations

  • Image quality inspection method and device for vehicle-mounted video transmission, vehicle and storage medium

    CN118741086A

  • No-reference video quality evaluation method fusing spatio-temporal features

    CN112954312A

  • No-reference video quality evaluation method and system based on space-time perception feature fusion

    CN117876851A

  • Image quality evaluation method and device, storage medium and electronic equipment

    CN118279230A