Video quality evaluation method, device, equipment, medium and program product

By adjusting the frame rate and extracting feature information from high frame rate videos, the problem of inaccurate evaluation of high frame rate videos in existing technologies is solved, achieving more accurate video quality evaluation and improving the accuracy and efficiency of evaluation results.

CN121746995APending Publication Date: 2026-03-27BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies, when assessing the quality of high frame rate videos, suffer from inaccurate temporal feature extraction, leading to discrepancies between the assessment results and the actual situation, thus reducing the accuracy of video quality assessment.

Method used

By acquiring high frame rate videos and their corresponding high frame rate original videos, the frame rate is adjusted based on a frame rate less than or equal to a preset frame rate to obtain a low frame rate original video. Temporal feature information is extracted based on the low frame rate video, and temporal feature upsampling is performed. Combined with spatial domain feature information, the results are input into a video quality assessment model for evaluation.

Benefits of technology

It improves the accuracy of high frame rate video quality assessment, ensures consistency between assessment results and subjective perception, reduces the number of assessments, and improves efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746995A_ABST
    Figure CN121746995A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video quality evaluation method and device, equipment, a medium and a program product. The method comprises the following steps: acquiring a to-be-evaluated first video and a first original video corresponding to the first video; obtaining a second original video, wherein the second original video is obtained by performing frame rate adjustment on the first original video based on a second frame rate; obtaining first time domain feature information, the first time domain feature information being obtained by performing time domain feature up-sampling based on second time domain feature information and a first frame rate, and the second time domain feature information being obtained by extracting time domain features based on a second original video; obtaining spatial domain feature information, wherein the spatial domain feature information is obtained by extracting spatial domain features based on the first video and the first original video; and inputting the first time domain feature information and the spatial domain feature information into a video quality evaluation model for quality evaluation to obtain a quality evaluation result of the first video. Through the technical scheme of the embodiment of the invention, the accuracy of quality evaluation of the high-frame-rate video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to computer technology, and more particularly to a video quality assessment method, apparatus, device, medium, and program product. Background Technology

[0002] With the rapid development of computer technology, video quality assessment is often required. Existing video quality assessment methods directly extract features from the video and then assess the quality based on the extracted feature information. However, in implementing this disclosure, it was found that the existing technology has at least the following problems:

[0003] Existing methods for quality assessment of high frame rate videos may result in inaccurate temporal feature extraction, leading to discrepancies between the assessment results and the actual situation, thus reducing the accuracy of video quality assessment. Summary of the Invention

[0004] This disclosure provides a video quality assessment method, apparatus, device, medium, and program product to improve the accuracy of quality assessment for high frame rate video.

[0005] In a first aspect, embodiments of this disclosure provide a video quality assessment method, including:

[0006] Obtain the first video to be evaluated and the first original video corresponding to the first video, wherein the original frame rate corresponding to the first original video and the first frame rate corresponding to the first video are both greater than the preset frame rate.

[0007] A second original video is obtained, which is obtained by adjusting the frame rate of the first original video based on a second frame rate, wherein the second frame rate is less than or equal to the preset frame rate.

[0008] First temporal feature information corresponding to the first video is obtained. The first temporal feature information is obtained by temporal feature upsampling based on the second temporal feature information and the first frame rate. The second temporal feature information is obtained by extracting temporal features based on the second original video.

[0009] The spatial domain feature information corresponding to the first video is obtained, and the spatial domain feature information is obtained by extracting spatial domain features based on the first video and the first original video.

[0010] The first temporal feature information and the spatial feature information are input into the video quality assessment model for quality assessment to obtain the quality assessment result of the first video.

[0011] Secondly, embodiments of this disclosure also provide a video quality assessment device, comprising:

[0012] The video acquisition module is used to acquire the first video to be evaluated and the first original video corresponding to the first video, wherein the original frame rate corresponding to the first original video and the first frame rate corresponding to the first video are both greater than the preset frame rate.

[0013] The second original video acquisition module is used to acquire a second original video, which is obtained by adjusting the frame rate of the first original video based on a second frame rate, wherein the second frame rate is less than or equal to the preset frame rate.

[0014] The first temporal feature information acquisition module is used to obtain the first temporal feature information corresponding to the first video. The first temporal feature information is obtained by temporal feature upsampling based on the second temporal feature information and the first frame rate. The second temporal feature information is obtained by extracting temporal features based on the second original video.

[0015] A spatial domain feature information acquisition module is used to obtain spatial domain feature information corresponding to the first video, wherein the spatial domain feature information is obtained by extracting spatial domain features based on the first video and the first original video.

[0016] The video quality assessment module is used to input the first temporal feature information and the spatial feature information into the video quality assessment model to perform quality assessment, so as to obtain the quality assessment result of the first video.

[0017] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:

[0018] One or more processors;

[0019] Storage device for storing one or more programs.

[0020] When the one or more programs are executed by the one or more processors, the one or more processors implement the video quality assessment method as described in any of the embodiments of this disclosure.

[0021] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video quality assessment method as described in any of the embodiments of this disclosure.

[0022] Fifthly, embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the video quality assessment method as described in any of the embodiments of the present disclosure.

[0023] In this embodiment, by acquiring a high frame rate first video to be evaluated and a corresponding high frame rate first original video, the frame rate of the first original video is adjusted based on a second frame rate less than or equal to a preset frame rate to obtain a low frame rate second original video. Temporal features are extracted based on the low frame rate second original video to obtain accurate second temporal feature information. Temporal features are then upsampled based on the second temporal feature information and the first frame rate to achieve adaptation from low frame rate temporal features to high frame rate temporal features, thereby obtaining accurate first temporal feature information corresponding to the first video. Spatial features are extracted based on the first video and the first original video to obtain accurate spatial feature information corresponding to the first video. This enables the video quality assessment model to accurately assess the video quality of the high frame rate first video based on the first temporal feature information and spatial feature information, thereby improving the accuracy of high frame rate video quality assessment. Attached Figure Description

[0024] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0025] Figure 1 This is a schematic flowchart of a video quality assessment method provided in an embodiment of this disclosure;

[0026] Figure 2 This is a flowchart illustrating a video quality assessment process according to an embodiment of this disclosure;

[0027] Figure 3 This is a flowchart illustrating another video quality assessment method provided in this embodiment of the disclosure;

[0028] Figure 4 This is a flowchart illustrating yet another video quality assessment method provided in this disclosure embodiment;

[0029] Figure 5 This is a schematic diagram of the structure of a video quality assessment device provided in an embodiment of this disclosure;

[0030] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0031] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0032] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0033] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0034] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0035] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0036] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0037] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0038] Before describing the technical solutions provided in the embodiments of this disclosure, VMAF (Video Multi-method Assessment Fusion), an existing video quality assessment method, will be introduced. VMAF is a machine learning-based video quality assessment method. VMAF directly extracts multiple spatiotemporal features from the video and performs quality assessment based on the extracted temporal and spatial feature information. However, when extracting temporal features, VMAF typically measures the difference between adjacent frames; that is, the greater the difference between frames, the richer the temporal information obtained, thus tending to give a higher quality score. For videos with "consistent subjective image quality across frames," if only the video frame rate is reduced (e.g., from 60fps to 30fps) without changing the image quality of individual frames (i.e., ensuring consistency of video spatial feature information), existing VMAF may actually give a higher quality score, which clearly deviates from subjective perception. Therefore, existing VMAF systematically underestimates the quality of high frame rate videos, reducing the accuracy of high frame rate video quality assessment.

[0039] Figure 1 This is a flowchart illustrating a video quality assessment method provided in an embodiment of the present disclosure. This embodiment is applicable to the situation of quality assessment of high frame rate videos. The method can be executed by a video quality assessment device, which can be implemented in the form of software and / or hardware, or optionally by an electronic device, such as a mobile terminal, a PC, or a server.

[0040] like Figure 1 As shown, the video quality assessment method specifically includes the following steps:

[0041] S110. Obtain the first video to be evaluated and the first original video corresponding to the first video. The original frame rate corresponding to the first original video and the first frame rate corresponding to the first video are both greater than the preset frame rate.

[0042] The first video refers to the high frame rate video whose image quality needs to be evaluated. The first frame rate refers to the frame rate of the first video. The first frame rate is greater than the preset frame rate. The preset frame rate is a pre-set maximum frame rate for videos that can accurately extract temporal features. For example, the preset frame rate could be 30fps. The first raw video refers to the video directly obtained from a video data source, such as video shot by a camera. The first video is obtained by transcoding the first raw video. The raw frame rate refers to the frame rate of the first raw video. The raw frame rate is also greater than the preset frame rate. The first raw video is also a high frame rate video. The first frame rate and the raw frame rate can be the same or different. It should be noted that when the first frame rate and the raw frame rate are the same, the quality evaluation result of the first video obtained by performing the following steps is more accurate.

[0043] In this embodiment of the disclosure, the high frame rate first original video is transcoded, that is, the first original video is decoded and then re-encoded to obtain the high frame rate first video to be evaluated.

[0044] In some optional implementations, there are at least two ways to obtain the first video. The first method is to obtain the first video by pre-transcoding the first original video. If the first original video has not yet been formally transcoded, pre-transcoding can be performed to assess the quality of the pre-transcoded first video, thus saving transcoding resources and predicting the video quality encoded by the encoding parameters used in the pre-transcoding. The second method is to obtain the first video by formally transcoding the first original video. If the first original video has already been formally transcoded, the formally transcoded first video can be directly obtained, and its quality can be directly assessed by performing the following steps. If the first original video has not yet been formally transcoded, it can also be formally transcoded, and its quality can be assessed by performing the following steps, thus improving the accuracy of the quality assessment through formal transcoding.

[0045] S120. Obtain the second original video. The second original video is obtained by adjusting the frame rate of the first original video based on the second frame rate. The second frame rate is less than or equal to the preset frame rate.

[0046] The second frame rate is a frame rate that is less than or equal to the preset frame rate. The frame rate of the second original video is the second frame rate. The second original video is the low frame rate video.

[0047] In this embodiment of the disclosure, the frame rate of the first original video is reduced based on the second frame rate to obtain a second original video with the second frame rate, i.e., a second original video with a low frame rate, such as... Figure 2 As shown.

[0048] S130. Obtain the first temporal feature information corresponding to the first video. The first temporal feature information is obtained by temporal feature upsampling based on the second temporal feature information and the first frame rate. The second temporal feature information is obtained by extracting temporal features based on the second original video.

[0049] The temporal feature information can be used to characterize the degree of motion change in the video sequence, reflecting the video features in the time domain. The second temporal feature information can be temporal feature information extracted from a second original video with a low frame rate. The second temporal feature information can include at least one temporal feature information. Each temporal feature information in the second temporal feature information can be frame-level temporal feature information. The first temporal feature information refers to temporal feature information with the same temporal resolution as the first video. The first temporal feature information can also be frame-level temporal feature information.

[0050] In the embodiments disclosed herein, such as Figure 2 As shown, each type of second temporal feature information at the frame level in the second original video is determined based on every two adjacent frames in the second original video with a low frame rate. Because the frame rate of the second original video is low, accurate second temporal feature information can be extracted. Based on the first frame rate, each type of second temporal feature information is upsampled to a temporal resolution consistent with the first video, thus achieving adaptation of the temporal features of the low frame rate video to the temporal features of the high frame rate video. For example, if the frame rate of the second original video is 30fps, meaning there are 30 video frames per second, the second temporal feature information extracted from the second original video per second can be a feature vector composed of 30 elements, where each element represents the temporal feature value corresponding to a video frame. If the first frame rate is 60fps, meaning there are 60 video frames per second, the extracted frame-level second temporal feature information needs to be upsampled by a factor of 2, resulting in a first temporal feature information composed of a feature vector composed of 60 elements, thereby obtaining temporal feature information matching the first frame rate.

[0051] It should be noted that extracting second temporal feature information from the second original video, compared to extracting it from the second video itself, can better simulate the temporal masking effect of the human visual system, thus ensuring the accuracy of video quality assessment. The second video is obtained by transcoding the second original video. Extracting second temporal feature information from the second video would potentially reduce motion complexity due to video compression, thereby affecting the accuracy of video quality assessment.

[0052] In some optional implementations, step S130, "upsampling temporal features based on the second temporal feature information and the first frame rate to obtain the first temporal feature information corresponding to the first video", may include: interpolating and aligning the second temporal feature information with the time axis based on the first frame rate and a preset interpolation method to obtain the first temporal feature information corresponding to the first video.

[0053] The preset interpolation method is a pre-defined method used to interpolate adjacent feature values ​​in the second temporal feature information at the frame level. For example, the preset interpolation method can be a repeating interpolation aligned by timestamps, linear interpolation, or spline interpolation, or it can be an alignment method based on motion estimation or optical flow.

[0054] Specifically, by using a preset interpolation method based on the first frame rate corresponding to the first video, the second temporal feature information at the frame level is interpolated and aligned with the time axis, thereby achieving upsampling of temporal features and obtaining first temporal feature information consistent with the time resolution corresponding to the first frame rate, thus obtaining the first temporal feature information of the high frame rate video and realizing the accurate extraction of temporal features of the high frame rate video.

[0055] S140. Obtain the spatial domain feature information corresponding to the first video. The spatial domain feature information is obtained by extracting spatial domain features from the first video and the first original video.

[0056] The spatial domain feature information may include at least one of Visual Quality Fidelity (VIF) and Detail Loss Measure (DLM). Visual Quality Fidelity measures the degree to which image information is preserved. The Detail Loss Measure is used to detect the loss of detail caused by compression or distortion. In this embodiment, the spatial domain feature information may be frame-level spatial domain feature information.

[0057] In the embodiments disclosed herein, such as Figure 2 As shown, the first original video is used as a reference video. Spatial domain features are extracted from the first video to obtain all extracted spatial domain feature information, such as the visual information fidelity and detail loss index of the first video. For example, the process of determining visual information fidelity is as follows: based on information theory and natural scene statistical models, the visual information fidelity is determined by comparing the information difference between the first original video and the first video after being filtered by the human visual system. The process of determining the detail loss index is as follows: by calculating the local gradient of the image (such as the Sobel operator), comparing the gradient difference between the first original video and the first video, and quantifying the high-frequency information loss, the detail loss is obtained. Non-natural noise (such as blockiness, blurring, etc.) is analyzed, and the interference intensity is quantified by frequency domain or spatial domain methods to obtain additional impairment. The detail loss index is obtained by weighted summation of the detail loss and additional impairment.

[0058] It should be noted that since the video frame rate does not affect the accuracy of spatial domain feature extraction, spatial domain features can be extracted directly based on the first frame rate with a high frame rate and the first original video to obtain the spatial domain feature information corresponding to the first video.

[0059] S150. Input the first temporal domain feature information and spatial domain feature information into the video quality assessment model for quality assessment to obtain the quality assessment result of the first video.

[0060] The video quality assessment model can be a pre-trained machine learning model used for video quality assessment based on spatiotemporal features. For example, the video quality assessment model can be, but is not limited to, a Support Vector Machine (SVM). In this embodiment, the video quality assessment model can be the one used in VMAF. The quality assessment result can be represented using a quality assessment score from 0 to 100; a higher score indicates higher video quality.

[0061] In this embodiment of the disclosure, the quality assessment result of the first video can be obtained in at least two ways. As a first way, since the first temporal feature information and spatial feature information are both frame-level feature information, the first temporal feature information and spatial feature information corresponding to each video frame in the first video can be input into the video quality assessment model to perform frame-level quality assessment, so as to obtain the quality assessment score corresponding to each video frame in the first video, and the quality assessment scores of all video frames are averaged to obtain the average score as the quality assessment result of the first video.

[0062] In the second implementation, the first video is the video segment to be evaluated. In this case, the first temporal feature information at the frame level corresponding to the first video is averaged to obtain the average temporal feature information of the first video, and the spatial feature information at the frame level corresponding to the first video is averaged to obtain the average spatial feature information of the first video. The average temporal feature information and the average spatial feature information of the first video are input into the video quality assessment model to perform overall quality assessment of the video segment, and the quality assessment score output by the video quality assessment model is directly used as the quality assessment result of the first video. This can reduce the number of quality assessments and improve the efficiency of video quality assessment.

[0063] It should be noted that, since the video quality assessment model is input with accurate first temporal and spatial feature information corresponding to high frame rate videos, the video quality assessment model can accurately assess the quality of high frame rate videos, thereby obtaining accurate quality assessment results and improving the consistency between objective evaluation and subjective perception of high frame rate videos.

[0064] The technical solution of this disclosure involves acquiring a high frame rate first video to be evaluated and a corresponding high frame rate first original video. Based on a second frame rate less than or equal to a preset frame rate, the frame rate of the first original video is adjusted to obtain a low frame rate second original video. Temporal features are extracted from the low frame rate second original video to obtain accurate second temporal feature information. Temporal features are then upsampled based on the second temporal feature information and the first frame rate to adapt the low frame rate temporal features to high frame rate temporal features, thereby obtaining accurate first temporal feature information corresponding to the first video. Spatial features are extracted based on the first video and the first original video to obtain accurate spatial feature information corresponding to the first video. This allows the video quality assessment model to accurately evaluate the quality of the high frame rate first video based on the first temporal and spatial feature information, thus improving the accuracy of high frame rate video quality assessment.

[0065] In some alternative implementations, temporal and spatial features can be extracted based on the existing spatiotemporal feature extraction module in VMAF. However, this module extracts temporal and spatial features from the input original video and the transcoded video (i.e., the video after transcoding the original video). Furthermore, the temporal feature information extracted by this module for high frame rate videos is inaccurate, leading to an underestimation of the quality of high frame rate videos by the video quality assessment model in VMAF.

[0066] Based on this, the process of obtaining the first temporal feature information can be as follows: transcode the second original video to obtain a second video with a second frame rate; input the second original video and the second video into the spatiotemporal feature extraction module for spatiotemporal feature extraction to obtain the extracted second temporal feature information. At this time, the extracted spatial domain feature information can be deleted and ignored, and temporal feature upsampling is performed based on the second temporal feature information and the first frame rate to obtain the first temporal feature information corresponding to the first video. It should be noted that since the spatiotemporal feature extraction module can accurately extract the temporal feature information of low frame rate videos, the second original video and the second video with low frame rates can be input into the spatiotemporal feature extraction module for spatiotemporal feature extraction, and only the extracted temporal feature information can be retained, while the extracted spatial domain feature information can be ignored and deleted.

[0067] The process of obtaining spatial domain feature information can be as follows: The first video and the first original video are input into the spatiotemporal feature extraction module for spatiotemporal feature extraction to obtain the extracted spatial domain feature information. It should be noted that since the video frame rate does not affect the accuracy of the spatiotemporal feature extraction module in extracting spatial domain features, high-frame-rate first videos and the first original video can be directly input into the spatiotemporal feature extraction module for spatiotemporal feature extraction. Furthermore, only the extracted spatial domain feature information can be retained, while the extracted temporal domain feature information, being inaccurate, can be ignored and deleted.

[0068] Step S150 may include: inputting the extracted first temporal feature information and spatial feature information into the existing video quality assessment model in VMAF for quality assessment, so as to obtain the quality assessment result of the first video.

[0069] The above method allows for more accurate quality assessment of high frame rate videos without altering the structural parameters of the existing spatiotemporal feature extraction module and video quality assessment model in VMAF. This achieves an intrusive modification to the model structure and training process, enabling online deployment. Experiments demonstrate that, under the same encoding parameter information (e.g., the same CRF (Conditional Random Field), implementing the technical solution of this disclosure can reduce the mean VMAF difference between high and low frame rate videos from 2.5 to 0.8, thereby improving the accuracy of high frame rate video assessment.

[0070] Figure 3 This is a flowchart illustrating another video quality assessment method provided in this disclosure. Based on the above-described embodiments, this disclosure provides a detailed description of the extraction process of the second temporal feature information. Explanations of terms that are the same as or corresponding to those in the above-described embodiments are not repeated here.

[0071] like Figure 3 As shown, the video quality assessment method specifically includes the following steps:

[0072] S310. Obtain the first video to be evaluated and the first original video corresponding to the first video. The original frame rate corresponding to the first original video and the first frame rate corresponding to the first video are both greater than the preset frame rate.

[0073] S320. Obtain the second original video. The second original video is obtained by adjusting the frame rate of the first original video based on the second frame rate. The second frame rate is less than or equal to the preset frame rate.

[0074] S330. For each original video frame in the second original video, determine the first temporal deviation information corresponding to the original video frame based on the original video frame and the previous video frame of the original video frame.

[0075] In this embodiment of the disclosure, for each original video frame in the second original video, such as the Nth original video frame, the motion change between the Nth original video frame and the (N-1)th original video frame is statistically analyzed to obtain the first temporal deviation information corresponding to the Nth original video frame.

[0076] In some optional implementations, step S330 may include: performing filtering processing on the original video frame and the previous video frame respectively to obtain the processed original video frame and the previous video frame; determining the first absolute error sum between the original video frame and the previous video frame based on the processed original video frame and the previous video frame; and determining the first temporal deviation information corresponding to the original video frame based on the first absolute error sum.

[0077] Specifically, filtering processes, such as convolutional filtering, are applied to both the original video frame (i.e., the Nth original video frame) and the previous video frame (i.e., the (N-1)th original video frame) to smooth the transitions. Based on the pixel values ​​of each pixel in the processed original video frame and the previous video frame, a first absolute error sum between the original and previous video frames is determined. This first absolute error sum is then used to measure the difference between the original and previous video frames. The first absolute error sum can be directly determined as the first temporal deviation information corresponding to the original video frame, or it can be normalized and used as the first temporal deviation information corresponding to the original video frame.

[0078] It should be noted that for the first original video frame in the second original video, since it has no preceding video frame, the first temporal deviation information corresponding to the first original video frame can be directly determined to be 0. Using the above method, the first temporal deviation information corresponding to each original video frame in the second original video can be accurately determined.

[0079] S340. Based on the original video frame and the next video frame after the original video frame, determine the second temporal deviation information corresponding to the original video frame.

[0080] In this embodiment of the disclosure, for each original video frame in the second original video, such as the Nth original video frame, the motion change between the Nth original video frame and the (N+1)th original video frame is statistically analyzed to obtain the second temporal deviation information corresponding to the Nth original video frame.

[0081] In some optional implementations, step S340 may include: performing filtering processing on the original video frame and the subsequent video frame respectively to obtain the processed original video frame and the subsequent video frame; determining the second absolute error sum between the original video frame and the subsequent video frame based on the processed original video frame and the subsequent video frame; and determining the second temporal deviation information corresponding to the original video frame based on the second absolute error sum.

[0082] Specifically, filtering processes, such as convolutional filtering, are applied to both the original video frame (i.e., the Nth original video frame) and the subsequent video frame (i.e., the N+1th original video frame) to smooth them out. Based on the pixel values ​​of each pixel in the processed original video frame and the pixel values ​​of each pixel in the subsequent video frame, a second absolute error sum between the original and subsequent video frames is determined. This second absolute error sum is then used to measure the difference between the original and subsequent video frames. The second absolute error sum can be directly determined as the second temporal deviation information corresponding to the original video frame, or it can be normalized and used as the second temporal deviation information corresponding to the original video frame.

[0083] It should be noted that for the last original video frame in the second original video, since there is no subsequent video frame, the second temporal offset information corresponding to the last original video frame can be directly determined to be 0. Using the above method, the second temporal offset information corresponding to each original video frame in the second original video can be accurately determined.

[0084] S350. Based on the first time-domain deviation information and the second time-domain deviation information, determine the second time-domain feature information corresponding to the original video frame.

[0085] In this embodiment of the disclosure, the second temporal feature information corresponding to the original video frame may include two different types of second temporal feature information, namely, third temporal feature information and fourth temporal feature information.

[0086] In some optional implementations, step S350 may include: determining the first temporal deviation information as the third temporal feature information corresponding to the original video frame; and determining the smaller value between the first temporal deviation information and the second temporal deviation information as the fourth temporal feature information corresponding to the original video frame.

[0087] Specifically, for each original video frame in the second original video, the first temporal deviation information corresponding to that original video frame can be directly determined as the third temporal feature information corresponding to that original video frame. If the first temporal deviation information corresponding to that original video frame is greater than the second temporal deviation information corresponding to that original video frame, then the second temporal deviation information corresponding to that original video frame is determined as the fourth temporal feature information corresponding to that original video frame. If the first temporal deviation information corresponding to that original video frame is less than the second temporal deviation information corresponding to that original video frame, then the first temporal deviation information corresponding to that original video frame is determined as the fourth temporal feature information corresponding to that original video frame, thereby determining the frame-level third temporal feature information and the frame-level fourth temporal feature information. By utilizing the third and fourth temporal feature information corresponding to each original video frame, the temporal deviation information in the second original video can be more comprehensively characterized, thereby more accurately characterizing the motion complexity of the second original video, so as to perform quality scoring based on the motion complexity, thereby simulating the temporal masking effect of the human visual system.

[0088] S360. Based on the second temporal feature information and the first frame rate, temporal feature upsampling is performed to obtain the first temporal feature information corresponding to the first video.

[0089] S370. Obtain the spatial domain feature information corresponding to the first video. The spatial domain feature information is obtained by extracting spatial domain features from the first video and the first original video.

[0090] S380. Input the first temporal domain feature information and spatial domain feature information into the video quality assessment model for quality assessment to obtain the quality assessment result of the first video.

[0091] The technical solution of this disclosure, for each original video frame in the second original video, determines the first temporal deviation information corresponding to the original video frame based on the original video frame and the preceding video frame, and determines the second temporal deviation information corresponding to the original video frame based on the original video frame and the following video frame. Based on the first and second temporal deviation information, the second temporal feature information corresponding to the original video frame can be determined more comprehensively and accurately, thereby obtaining accurate second temporal feature information in advance, and thus ensuring the accuracy of high frame rate video quality assessment.

[0092] Figure 4 This is a flowchart illustrating another video quality assessment method provided in this disclosure. Based on the above-described embodiments, this disclosure provides a detailed description of the quality assessment process when the first video is obtained by pre-transcoding a first original video. Explanations of terms that are the same as or corresponding to those in the above-described embodiments are not repeated here.

[0093] like Figure 4 As shown, the video quality assessment method specifically includes the following steps:

[0094] S410. Obtain the first video to be evaluated and the first original video corresponding to the first video. The original frame rate corresponding to the first original video and the first frame rate corresponding to the first video are both greater than the preset frame rate.

[0095] S420. Obtain the second original video. The second original video is obtained by adjusting the frame rate of the first original video based on the second frame rate. The second frame rate is less than or equal to the preset frame rate.

[0096] S430. Obtain the first temporal feature information corresponding to the first video. The first temporal feature information is obtained by temporal feature upsampling based on the second temporal feature information and the first frame rate. The second temporal feature information is obtained by extracting temporal features based on the second original video.

[0097] S440. Obtain the spatial domain feature information corresponding to the first video. The spatial domain feature information is obtained by extracting spatial domain features from the first video and the first original video.

[0098] S450. Based on the first video, determine the target feature information corresponding to the first video, wherein the target feature information includes at least one of pre-coded feature information, peak signal-to-noise ratio, and structural similarity index.

[0099] The target feature information refers to the additional feature information needed when the first video is obtained by pre-transcoding the first original video, in order to ensure the accuracy of quality assessment based on the pre-transcoded video. Precoding feature information refers to the precoding feature information used in the pre-transcoding process of the first video. For example, precoding feature information may include at least one of prediction mode feature information, bit type feature information, and distortion type feature information. Prediction mode feature information may include, but is not limited to, the proportion of intra-frame mode in P-frames, the proportion of merge mode in P-frames, and the proportion of skip mode in B-frames. Bit type feature information may include, but is not limited to, the percentage of P-frame bitrate used for residuals relative to the total bitrate of P-frames. Distortion type feature information may include, but is not limited to, the sum of the most-paired errors of the residuals in P-frames and the sum of the absolute errors of the residuals in B-frames. A P-frame is a predictive coded frame type in video coding, which achieves data compression by recording the difference between the current frame and a reference frame (I-frame or P-frame). A B-frame is a bidirectional predictive coded frame in video compression. Intra-frame mode, merge mode, and skip mode are three different prediction techniques used to improve compression efficiency. Intra-frame mode generates predicted values ​​using already encoded pixels within the current frame, removing spatial redundancy. Merge mode constructs a candidate list by reusing motion vectors in the spatial or temporal domains. Skip mode is a special form of Merge mode, suitable for static or regularly moving regions. Peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) are two objective quality evaluation metrics that measure the similarity between the first video and the first original video from different perspectives. SSIM is calculated by first acquiring two video frames at the same timestamp from both the first and first original videos, and then comprehensively evaluating the similarity between these two frames using three dimensions: brightness, contrast, and structure. The SSIM value ranges from -1 to 1, with values ​​closer to 1 indicating higher similarity, better aligning with human visual perception of image quality. By fusing the SSIM values ​​of each frame in the first video and the first original video (e.g., by taking the mean, median, or weighted average), the fused result serves as the final structural similarity index for the first video.

[0100] In this embodiment of the disclosure, when the first video is obtained by pre-transcoding the first original video, additional target feature information, namely pre-coded feature information, peak signal-to-noise ratio, and structural similarity index, is extracted from the first video. Since the extracted target feature information is all non-temporal domain feature information, it can be directly extracted from the high frame rate first video. It should be noted that the extracted target feature information can also be frame-level feature information in the first video.

[0101] S460. Input the target feature information, the first time domain feature information, and the spatial domain feature information into the video quality assessment model to perform quality assessment and obtain the quality assessment result of the first video.

[0102] In this embodiment of the disclosure, target feature information, first temporal feature information, and spatial feature information are input together into the video quality assessment model, so that the video quality assessment model can comprehensively use target feature information, first temporal feature information, and spatial feature information to perform a more accurate quality assessment, thereby improving the accuracy of quality assessment when the video to be evaluated is a pre-transcoded video.

[0103] The technical solution of this disclosure improves the accuracy of quality assessment when the video to be evaluated is a pre-transcoded video by determining the target feature information corresponding to the first video based on the first video when the first video is obtained by pre-transcoding the first original video, and simultaneously inputting the target feature information, the first temporal feature information and the spatial feature information into the video quality assessment model for quality assessment.

[0104] In some optional implementations, when the first video is obtained by pre-transcoding the first original video, the method may further include: obtaining the quality assessment result of each first video, wherein different first videos are obtained by pre-transcoding the first original video based on different candidate coding parameter information; determining target coding parameter information from multiple candidate coding parameter information based on the quality assessment result, and performing formal transcoding on the first original video based on the target coding parameter information to obtain the target video.

[0105] Specifically, the first original video can be pre-transcoded multiple times based on multiple candidate coding parameter information, such as multiple CRF values, to obtain a first video after each pre-transcoding. Each first video corresponds one-to-one with the candidate coding parameter information. For each pre-transcoded first video, a quality assessment can be performed according to the process described above to obtain a quality assessment result for each first video. The quality assessment results of all first videos are compared, and the candidate coding parameter information corresponding to the first video with the highest quality is determined as the target coding parameter information, thus obtaining the optimal coding parameter information. Then, the first original video is formally transcoded based on the target coding parameter information to obtain the target video with the best quality, thereby achieving the optimal selection of coding parameter information.

[0106] Figure 5 This is a schematic diagram of the structure of a video quality assessment device provided in an embodiment of the present disclosure, as shown below. Figure 5 As shown, the device specifically includes: a video acquisition module 510, a second original video acquisition module 520, a first temporal feature information acquisition module 530, a spatial domain feature information acquisition module 540, and a video quality evaluation module 550.

[0107] The system includes the following modules: a video acquisition module 510, which acquires a first video to be evaluated and a first original video corresponding to the first video, wherein both the original frame rate and the first frame rate of the first original video are greater than a preset frame rate; a second original video acquisition module 520, which acquires a second original video by adjusting the frame rate of the first original video based on a second frame rate, wherein the second frame rate is less than or equal to the preset frame rate; a first temporal feature information acquisition module 530, which acquires first temporal feature information corresponding to the first video, wherein the first temporal feature information is obtained by temporal feature upsampling based on the second temporal feature information and the first frame rate, and the second temporal feature information is obtained by extracting temporal features based on the second original video; a spatial domain feature information acquisition module 540, which acquires spatial domain feature information corresponding to the first video, wherein the spatial domain feature information is obtained by extracting spatial domain features based on the first video and the first original video; and a video quality assessment module 550, which inputs the first temporal feature information and the spatial domain feature information into a video quality assessment model for quality assessment to obtain the quality assessment result of the first video.

[0108] The technical solution provided in this disclosure acquires a high frame rate first video to be evaluated and a corresponding high frame rate first original video. Based on a second frame rate less than or equal to a preset frame rate, the frame rate of the first original video is adjusted to obtain a low frame rate second original video. Temporal features are extracted from the low frame rate second original video to obtain accurate second temporal feature information. Temporal features are upsampled based on the second temporal feature information and the first frame rate to achieve adaptation from low frame rate temporal features to high frame rate temporal features, thereby obtaining accurate first temporal feature information corresponding to the first video. Spatial features are extracted based on the first video and the first original video to obtain accurate spatial feature information corresponding to the first video. This enables the video quality assessment model to accurately assess the video quality of the high frame rate first video based on the first temporal feature information and spatial feature information, thereby improving the accuracy of high frame rate video quality assessment.

[0109] In some alternative implementations, the first video is obtained by pre-transcoding the first original video; or,

[0110] The first video is obtained by formally transcoding the first original video.

[0111] In some optional implementations, the first time-domain feature information acquisition module 530 includes:

[0112] The first temporal deviation information determination unit is used to determine the first temporal deviation information corresponding to each original video frame in the second original video based on the original video frame and the previous video frame of the original video frame.

[0113] The second temporal deviation information determination unit is used to determine the second temporal deviation information corresponding to the original video frame based on the original video frame and the video frame following the original video frame.

[0114] The second temporal feature information determination unit is used to determine the second temporal feature information corresponding to the original video frame based on the first temporal deviation information and the second temporal deviation information.

[0115] In some optional implementations, the first time-domain deviation information determination unit is specifically used for:

[0116] The original video frame and the preceding video frame are filtered separately to obtain the processed original video frame and the preceding video frame; based on the processed original video frame and the preceding video frame, the first absolute error between the original video frame and the preceding video frame is determined; based on the first absolute error, the first temporal deviation information corresponding to the original video frame is determined.

[0117] In some optional implementations, the second time-domain feature information includes: third time-domain feature information and fourth time-domain feature information;

[0118] The second time-domain feature information determination unit is specifically used for:

[0119] The first temporal deviation information is determined as the third temporal feature information corresponding to the original video frame; the smaller value between the first temporal deviation information and the second temporal deviation information is determined as the fourth temporal feature information corresponding to the original video frame.

[0120] In some optional implementations, the first time-domain feature information acquisition module 530 is specifically used for:

[0121] Based on the first frame rate and the preset interpolation method, the second temporal feature information is interpolated and aligned with the time axis to obtain the first temporal feature information corresponding to the first video.

[0122] In some alternative implementations, when the first video is obtained by pre-transcoding the first original video, the apparatus further includes:

[0123] The target feature information determination module is used to determine the target feature information corresponding to the first video based on the first video, wherein the target feature information includes at least one of pre-coded feature information, peak signal-to-noise ratio and structural similarity index;

[0124] The video quality assessment module 550 is specifically used to: input the target feature information, the first temporal feature information and the spatial feature information into the video quality assessment model for quality assessment, so as to obtain the quality assessment result of the first video.

[0125] In some optional implementations, the precoding feature information includes at least one of prediction mode feature information, bit type feature information, and distortion type feature information.

[0126] In some alternative implementations, when the first video is obtained by pre-transcoding the first original video, the apparatus further includes:

[0127] The quality assessment result acquisition module is used to acquire the quality assessment result of each first video. Different first videos are obtained by pre-transcoding the first original video based on different candidate encoding parameter information.

[0128] The target video determination module is used to determine target encoding parameter information from multiple candidate encoding parameter information based on the quality assessment results, and to perform formal transcoding on the first original video based on the target encoding parameter information to obtain the target video.

[0129] In some alternative implementations, the spatial domain feature information includes at least one of visual information fidelity and detail loss metrics.

[0130] The video quality assessment device provided in this disclosure can execute the video quality assessment method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0131] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0132] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 6 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 6The diagram below shows the structure of the terminal device or server 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0133] like Figure 6 As shown, electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.

[0134] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0135] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0136] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0137] The electronic device provided in this embodiment and the video quality assessment method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0138] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the video quality assessment method provided in the above embodiments.

[0139] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0140] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0141] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0142] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the aforementioned one or more programs, the electronic device causes the following actions: to acquire event description information to be identified; to process the event description information to generate multiple descriptive phrases corresponding to the event description information; to perform phrase matching based on the multiple descriptive phrases and multiple event phrases corresponding to each event in an event database to determine target events that may match the event description information from the event database; and to input the event description information and target event information from the event database into a target recognition model to identify whether the event description information matches the target event information, thereby obtaining a recognition result.

[0143] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0144] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the video quality assessment method provided in the above embodiments.

[0145] The computer program product includes a computer program carried on a non-transitory computer-readable medium, which contains program code for performing video quality assessment methods. The program code can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0146] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0147] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0148] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0149] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0150] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0151] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0152] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A video quality assessment method, comprising: Obtain the first video to be evaluated and the first original video corresponding to the first video, wherein the original frame rate corresponding to the first original video and the first frame rate corresponding to the first video are both greater than the preset frame rate. A second original video is obtained, which is obtained by adjusting the frame rate of the first original video based on a second frame rate, wherein the second frame rate is less than or equal to the preset frame rate. First temporal feature information corresponding to the first video is obtained. The first temporal feature information is obtained by temporal feature upsampling based on the second temporal feature information and the first frame rate. The second temporal feature information is obtained by extracting temporal features based on the second original video. The spatial domain feature information corresponding to the first video is obtained, and the spatial domain feature information is obtained by extracting spatial domain features based on the first video and the first original video. The first temporal feature information and the spatial feature information are input into the video quality assessment model for quality assessment to obtain the quality assessment result of the first video.

2. The video quality assessment method according to claim 1, wherein the first video is obtained by pre-transcoding the first original video; or, The first video is obtained by formally transcoding the first original video.

3. The video quality assessment method according to claim 1, wherein temporal features are extracted based on the second original video to obtain second temporal feature information, comprising: For each original video frame in the second original video, the first temporal deviation information corresponding to the original video frame is determined based on the original video frame and the previous video frame of the original video frame. Based on the original video frame and the next video frame after the original video frame, determine the second temporal deviation information corresponding to the original video frame. Based on the first temporal deviation information and the second temporal deviation information, the second temporal feature information corresponding to the original video frame is determined.

4. The video quality assessment method according to claim 3, wherein determining the first temporal deviation information corresponding to the original video frame based on the original video frame and the previous video frame of the original video frame includes: The original video frame and the previous video frame are filtered separately to obtain the processed original video frame and the previous video frame. Based on the processed original video frame and the previous video frame, determine the first absolute error between the original video frame and the previous video frame; Based on the first absolute error, the first temporal deviation information corresponding to the original video frame is determined.

5. The video quality assessment method according to claim 3, wherein the second temporal feature information includes: Third time-domain feature information and fourth time-domain feature information; The step of determining the second temporal feature information corresponding to the original video frame based on the first temporal deviation information and the second temporal deviation information includes: The first temporal deviation information is determined as the third temporal feature information corresponding to the original video frame; The smaller value between the first temporal deviation information and the second temporal deviation information is determined as the fourth temporal feature information corresponding to the original video frame.

6. The video quality assessment method according to claim 1, wherein the step of performing temporal feature upsampling based on the second temporal feature information and the first frame rate to obtain the first temporal feature information corresponding to the first video includes: Based on the first frame rate and the preset interpolation method, the second temporal feature information is interpolated and aligned with the time axis to obtain the first temporal feature information corresponding to the first video.

7. The video quality assessment method according to claim 2, wherein when the first video is obtained by pre-transcoding the first original video, the method further includes: Based on the first video, target feature information corresponding to the first video is determined, wherein the target feature information includes at least one of pre-coded feature information, peak signal-to-noise ratio, and structural similarity index; The step of inputting the first temporal feature information and the spatial feature information into a video quality assessment model for quality assessment to obtain the quality assessment result of the first video includes: The target feature information, the first temporal feature information, and the spatial feature information are input into the video quality assessment model for quality assessment to obtain the quality assessment result of the first video.

8. The video quality assessment method according to claim 7, wherein the precoding feature information includes: At least one of the following: prediction pattern feature information, bit type feature information, and distortion type feature information.

9. The video quality assessment method according to claim 2, wherein when the first video is obtained by pre-transcoding the first original video, the method further includes: Obtain the quality assessment result of each first video, wherein different first videos are obtained by pre-transcoding the first original video based on different candidate encoding parameter information; Based on the quality assessment results, target encoding parameters are determined from multiple candidate encoding parameter information, and the first original video is formally transcoded based on the target encoding parameter information to obtain the target video.

10. The video quality assessment method according to any one of claims 1-9, wherein the spatial domain feature information includes: At least one of the following metrics: visual information fidelity and detail loss.

11. A video quality assessment device, comprising: The video acquisition module is used to acquire the first video to be evaluated and the first original video corresponding to the first video, wherein the original frame rate corresponding to the first original video and the first frame rate corresponding to the first video are both greater than the preset frame rate. The second original video acquisition module is used to acquire a second original video, which is obtained by adjusting the frame rate of the first original video based on a second frame rate, wherein the second frame rate is less than or equal to the preset frame rate. The first temporal feature information acquisition module is used to obtain the first temporal feature information corresponding to the first video. The first temporal feature information is obtained by temporal feature upsampling based on the second temporal feature information and the first frame rate. The second temporal feature information is obtained by extracting temporal features based on the second original video. A spatial domain feature information acquisition module is used to obtain spatial domain feature information corresponding to the first video, wherein the spatial domain feature information is obtained by extracting spatial domain features based on the first video and the first original video. The video quality assessment module is used to input the first temporal feature information and the spatial feature information into the video quality assessment model to perform quality assessment, so as to obtain the quality assessment result of the first video.

12. An electronic device, the electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the video quality assessment method as described in any one of claims 1-10.

13. A storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video quality assessment method as described in any one of claims 1-10.

14. A computer program product comprising a computer program that, when executed by a processor, implements the video quality assessment method as described in any one of claims 1-10.