Video quality evaluation method, device and equipment
By using a pre-trained video quality assessment model that combines image, audio, and audio-visual synchronization assessment, the problems of high cost and low accuracy of manual assessment are solved, achieving efficient and accurate video quality assessment.
Patent Information
- Application Number
- CN202511530336.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2026-02-10
AI Technical Summary
Existing video quality assessment methods rely on manual scoring, which is costly, results in untimely and inaccurate feedback, leading to low assessment efficiency and accuracy.
A pre-trained video quality assessment model is used to comprehensively evaluate video quality through image quality, audio quality, and audio-visual synchronization assessment. The evaluation is performed using a multimodal model constructed with deep learning algorithms.
It enables fast and accurate video quality assessment, improves assessment efficiency and accuracy, and can comprehensively assess video quality from multiple dimensions.
Smart Images

Figure CN121504822A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and particularly relates to a video quality evaluation method, device and equipment. BACKGROUND
[0002] With the rapid development of intelligent devices and wireless networks, the types and volumes of video-related services such as video communication, high-definition video playing and short video socializing are also rapidly increasing. Therefore, how to evaluate the quality of videos to improve the processing quality of video-related services has become a focus of public attention.
[0003] For example, the quality of a video can be evaluated through artificial scoring and other feedback forms. However, artificial scoring and other feedback forms have problems such as high cost, untimely feedback and inaccurate feedback, which result in low evaluation efficiency and evaluation accuracy of video quality evaluation. Therefore, a technical solution is needed to improve the evaluation efficiency and evaluation accuracy of video quality evaluation to improve the processing quality of video-related services. SUMMARY
[0004] The embodiment of the present application aims to provide a technical solution capable of improving the evaluation efficiency and evaluation accuracy of video quality evaluation to improve the processing quality of video-related services.
[0005] To solve the above technical problems, the embodiment of the present application is implemented as follows: In a first aspect, the embodiment of the present application provides a video quality evaluation method, which comprises the following steps: receiving a quality evaluation request for a target video; in response to the quality evaluation request, acquiring a pre-trained video quality evaluation model; using the video quality evaluation model, performing image quality and audio quality evaluation processing on the target video respectively to obtain image quality evaluation results and audio quality evaluation results of the target video; using the video quality evaluation model, performing audio-visual synchronization evaluation processing on a first image and a first audio related to a detection object in the target video to obtain audio-visual synchronization evaluation results of the target video; based on the image quality evaluation results, the audio quality evaluation results and the audio-visual synchronization evaluation results of the target video, determining quality evaluation results of the target video.
[0006] In a second aspect, the embodiment of the present application provides a video quality evaluation device, which comprises: a request receiving module configured to receive a quality evaluation request for a target video; a model determining module configured to acquire a pre-trained video quality evaluation model in response to the quality evaluation request; The first evaluation module is configured to perform image quality and audio quality evaluation on the target video respectively by using the video quality evaluation model, and obtain image quality evaluation results and audio quality evaluation results of the target video. The second evaluation module is configured to perform audio-visual synchronization evaluation on the first image and the first audio related to the detection object in the target video by using the video quality evaluation model, and obtain audio-visual synchronization evaluation results of the target video. The quality evaluation module is configured to determine the quality evaluation results of the target video based on the image quality evaluation results, the audio quality evaluation results and the audio-visual synchronization evaluation results of the target video.
[0007] In a third aspect, an embodiment of the present application provides a video quality evaluation device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, and when the computer program is executed by the processor, the steps of the video quality evaluation method provided by the above-mentioned embodiments are implemented.
[0008] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the video quality evaluation method provided by the above-mentioned embodiments are implemented.
[0009] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the video quality evaluation method provided by the above-mentioned embodiments are implemented. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments described in the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0011] Figure 1 It is a flowchart of a video quality evaluation method of the present application; Figure 2 It is a flowchart of a model training process of the present application; Figure 3 It is a schematic diagram of a model training process of the present application; Figure 4 It is a flowchart of a determination process of audio-visual synchronization evaluation results of the present application; Figure 5A flowchart of a process for determining an audio quality evaluation result of the present application; Figure 6 A structural diagram of a video quality evaluation device of the present application; Figure 7 A structural diagram of a video quality evaluation device of the present application. DETAILED DESCRIPTION
[0012] Embodiments of the present application provide a video quality evaluation method, device and equipment.
[0013] In order for those skilled in the art to better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.
[0014] Embodiments of the present specification provide a video quality evaluation method, device and equipment. With the rapid development of intelligent devices and wireless networks, the types and volumes of video-related services such as video communication, high-definition video playback, and short video social networking have also rapidly increased. Therefore, how to evaluate video quality to improve the processing quality of video-related services has become the focus of public attention. For example, the quality of a video can be evaluated through artificial scoring and other feedback forms. However, artificial scoring and other feedback forms have problems such as high cost, untimely feedback, and inaccurate feedback, which result in low evaluation efficiency and accuracy of video quality evaluation. To improve the evaluation efficiency and accuracy of video quality evaluation and improve the processing quality of video-related services, in the present scheme, a quality evaluation request for a target video is received, a pre-trained video quality evaluation model is obtained in response to the quality evaluation request, the video quality evaluation model is used to perform image quality and audio quality evaluation processing on the target video respectively, and image quality evaluation results and audio quality evaluation results of the target video are obtained. The video quality evaluation model is used to perform audio-visual synchronization evaluation processing on a first image and a first audio related to a detection object in the target video, and an audio-visual synchronization evaluation result of the target video is obtained. Based on the image quality evaluation results, the audio quality evaluation results, and the audio-visual synchronization evaluation result of the target video, a quality evaluation result of the target video is determined. In this way, on the one hand, the video quality evaluation model can quickly and accurately evaluate the quality of the target video, and on the other hand, the multi-modal information of the target video can be used to comprehensively evaluate the video quality of the target video through multiple dimensions (i.e., image, audio, and audio-visual synchronization), so as to improve the accuracy of video quality evaluation while ensuring the efficiency of video quality evaluation. Specific processing can be referred to the specific content in the following embodiments.
[0015] As Figure 1 shown in the figure, the embodiment of the application provides a video quality evaluation method, the execution subject of the method can be a terminal device or a server, the terminal device can be a mobile terminal device such as a mobile phone, a tablet computer, a smart watch, etc., or a terminal device such as a computer, the server can be an independent server, or a server cluster composed of multiple servers. The method can specifically include the following steps: In step S102, a quality evaluation request for a target video is received.
[0016] Among them, the target video can be any video that needs to be evaluated for quality.
[0017] In implementation, the target video can be a video related to communication services, for example, the target video can be a real-time or recorded communication video stream generated during communication. Specifically, after receiving a video communication request from another user, the user can send a real-time communication video stream generated during the video communication process to the server to trigger a server request for quality evaluation of the real-time communication video, at which time the real-time communication video stream is the target video. Alternatively, the user can also record the video communication process and perform preprocessing such as desensitization on the recorded communication video locally, and then send the preprocessed communication video to the server to trigger a server request for quality evaluation of the preprocessed communication video, at which time the preprocessed communication video is the target video.
[0018] Alternatively, the target video can also be a video taken by a user using a certain camera device, and the user can send the taken video to the server to trigger a server request for quality evaluation of the taken video, at which time the taken video is the target video, and the user can determine the shooting quality evaluation result for the camera device according to the quality evaluation result of the target video.
[0019] Alternatively, the target video can also be a video to be played in a video playing platform (or a video social platform, etc.), and the server can determine these to-be-played videos as target videos and trigger a quality evaluation request for the target videos to optimize the video playing service according to the quality evaluation request for the target videos.
[0020] In addition, the above-mentioned triggering method of the quality evaluation request for the target video is an optional and implementable triggering method, and in actual application scenarios, there can be many different triggering methods, and different triggering methods can be selected according to different actual application scenarios, and the embodiments of the present application do not make specific limitations on this.
[0021] In step S104, in response to the quality evaluation request, a pre-trained video quality evaluation model is obtained.
[0022] The video quality assessment model can be a model built based on a preset deep learning algorithm, such as a preset multimodal large language model or a model built based on a convolutional neural network algorithm.
[0023] In step S106, the target video is evaluated using a video quality assessment model to obtain the image quality assessment results and audio quality assessment results of the target video.
[0024] In implementation, the video quality assessment model may include a first module for image quality assessment processing and a second module for audio quality assessment processing.
[0025] Taking the first module as an example, which is built based on the convolutional neural network algorithm, the server can use the convolutional neural network to perform image extraction processing on the target video based on the preset image extraction rules, and then perform image quality assessment processing on the extracted images to obtain the image quality assessment result of the target video.
[0026] The preset image extraction rules may include rules such as image extraction based on preset time intervals or image extraction based on preset detection objects.
[0027] For example, taking image extraction at a preset time interval as an example, the preset time interval can be determined according to the total duration of the target video. For example, the preset time interval can be 1% of the total duration of the target video. The server can use the first module to perform image extraction processing on the target video based on the preset time interval.
[0028] Alternatively, taking image extraction based on a preset detection object as an example, the server can use the first module to perform image extraction processing on images in the target video that contain the preset detection object.
[0029] The server can use the second module to extract the audio data corresponding to the target video, and based on the temporal and / or frequency domain features of the audio data, perform audio quality assessment processing on the target video to obtain the audio quality assessment result corresponding to the target video.
[0030] In step S108, the video quality assessment model is used to perform audio-visual synchronization assessment on the first image and the first audio related to the detection object in the target video, so as to obtain the audio-visual synchronization assessment result of the target video.
[0031] In implementation, the video quality assessment model may also include a third module for audio-visual synchronization assessment. The server can use the third module to extract multiple first images containing the detected object (such as a user, an object, etc.) in the target video based on the object features of the detected object.
[0032] The server can extract audio from the target video corresponding to the time range of multiple first images, and determine the extracted audio as the first audio corresponding to the first image.
[0033] The server can utilize a third module to perform audio-visual synchronization evaluation on multiple first images and first audio files to obtain the audio-visual synchronization evaluation result of the target video. For example, the server can perform text conversion on the first audio to obtain the corresponding first text data, and perform lip-sync recognition on the detected objects in multiple first images to determine the second text data based on the recognition results. Then, the server can determine the audio-visual synchronization evaluation result of the target video based on the matching of the first text data and the second text data at the same time point.
[0034] Furthermore, the above-mentioned audio-visual synchronization evaluation processing method is an optional and feasible processing method. In actual application scenarios, there can be a variety of different processing methods. Different processing methods can be selected according to different actual application scenarios. This specification does not specifically limit this in the embodiments.
[0035] In step S110, the quality assessment result of the target video is determined based on the image quality assessment result, audio quality assessment result, and audio-visual synchronization assessment result of the target video.
[0036] In implementation, the server can determine the weights of the image, audio, and audio-visual synchronization based on the application scenario of the target video. Then, based on these weights, the image quality assessment results, audio quality assessment results, and audio-visual synchronization assessment results of the target video are weighted and processed to obtain the quality assessment result of the target video.
[0037] The quality assessment results can be used to optimize video-related services, such as video playback devices, video generation devices, video shooting devices, and the video transmission process.
[0038] This invention provides a video quality assessment method. It receives a quality assessment request for a target video, and in response, acquires a pre-trained video quality assessment model. Using this model, it performs image and audio quality assessments on the target video, obtaining image and audio quality assessment results. Then, using the same model, it performs audio-visual synchronization assessment on a first image and first audio related to the detected object in the target video, obtaining an audio-visual synchronization assessment result. Based on these results, the overall quality assessment of the target video is determined. This approach offers two advantages: firstly, the video quality assessment model allows for rapid and accurate quality assessment of the target video; secondly, it leverages the multimodal information of the target video to comprehensively assess its quality across multiple dimensions (image, audio, and audio-visual synchronization), thereby improving both efficiency and accuracy in video quality assessment.
[0039] In practical applications, video quality assessment models can also be trained. There are various methods for model training; one optional method is provided below. Figure 2 As shown, the specific process may include the following steps S202 to S208.
[0040] In step S202, the first video used to train the video quality assessment model is obtained.
[0041] The first video can be any video that can be acquired within the video acquisition period. For example, the video processing service corresponding to the first video can be the same as the video processing service corresponding to the target video. Specifically, if the target video can be a video acquired in a video communication scenario, then the first video can be a communication video in a video communication scenario acquired within the video acquisition period (such as the last month, the last three months, etc.).
[0042] Alternatively, the first video can also be a video generated within the video acquisition period using a generative model based on a generative algorithm, which is related to the video processing business corresponding to the target video. Furthermore, there can be various methods for acquiring the first video, and different acquisition methods can be selected according to different application scenarios. This specification does not specifically limit these methods in the embodiments.
[0043] In step S204, the video configuration parameters and / or video playback parameters of the first video are adjusted to obtain the second video, and the third evaluation result and target quality evaluation result of the second video are obtained.
[0044] Among them, the third evaluation result can be used to characterize the audio-visual synchronization of the second video, the target quality evaluation result can be used to characterize the video quality of the second video, the video configuration parameters can include bitrate, resolution, etc., and the video playback parameters can include bandwidth capacity, packet loss rate, latency, jitter, etc.
[0045] In implementation, taking communication video as the target video as an example, the server can simulate a video call scenario and collect a video dataset containing human images as the first video, which can be denoted as G. For each segment g of the first video in the video dataset G, g∈G, a streaming conversion tool (FFmpeg, GStreamer, etc.) can be used to convert it into a video stream for playback. By adjusting the configuration parameters provided by the streaming conversion tool, the video configuration parameters of the first video can be adjusted, such as configuring different bitrates and resolutions to obtain the second video. Alternatively, during the playback of the first video, network traffic control tools (such as Clumsy on Windows, tc on Linux, etc.) can be used to dynamically adjust the bandwidth capacity, packet loss rate, latency, jitter, and other video playback parameters of the first video to obtain the second video.
[0046] In this way, various network conditions in a real video call can be simulated. For each original video segment g, g∈G, after adjustment using the above method, a second video can be obtained, which can be denoted as... That is, the first video and the second video can form a video pair. .
[0047] Then, the server can perform audio-visual synchronization evaluation on the second video to obtain a third evaluation result corresponding to the second video. The third evaluation result may include the audio-visual synchronization type of the second video, such as synchronized type, asynchronous type, etc.
[0048] For example, for each second video segment g', manual annotation can be used to determine whether the second video has audio-visual asynchrony. If it does, the third evaluation result can be marked as 1 (i.e., the audio-visual synchronization type of the second video is asynchronous); if it does not, the third evaluation result can be marked as 0 (i.e., the audio-visual synchronization type of the second video is synchronized). The audio-visual synchronization type can be marked as s. sync s sync ∈{0,1}.
[0049] For each second video segment The server can determine the target quality assessment result of the second video by using the Mean Opinion Score (MOS) of manually annotated videos. The annotation criteria for the MOS score are shown in Table 1 below.
[0050] Table 1
[0051] Thus, based on the annotation criteria in Table 1, the MOS score of the second video g' for each segment can be determined, and can be denoted as... Therefore, the target quality assessment result of the second video can be determined based on the MOS score, such as the MOS score being determined as the target quality assessment result of the second video.
[0052] In step S206, based on the first video, the second video is subjected to image quality and audio quality evaluation processing respectively to obtain the first evaluation result and the second evaluation result of the second video.
[0053] The first evaluation result can be used to characterize the image quality of the second video, and the second evaluation result can be used to characterize the audio quality of the second video.
[0054] In implementation, the server can determine the first evaluation result of the second video based on the difference between the pixel values of the image frames in the second video and the first video, and determine the second evaluation result of the second video based on the difference between the peak signal-to-noise ratio of the image frames in the second video and the first video, and the second evaluation result of the second video based on the difference between the peak signal-to-noise ratio of the image frames in the second video and the first video.
[0055] In practical applications, the specific processing method for performing image quality assessment processing on the second video based on the first video in step S206 above to obtain the first assessment result of the second video can be varied. The following provides an optional processing method, which may specifically include the processing steps A1 to A3.
[0056] In step A1, image extraction processing is performed on the first video and the second video respectively to obtain a first image and a second image with the same time position.
[0057] In step A2, the image quality score of the second video is determined based on the pixel values of the first image, the standard deviation of the pixel values of the first image, the pixel values of the second image, the standard deviation of the pixel values of the second image, and the covariance of the pixel values of the first image and the second image.
[0058] In step A3, the first evaluation result of the second video is determined based on the image quality score of the second video.
[0059] In implementation, for each pair of videos That is, the first video g and the second video g' can be randomly selected. For the first and second images with the same frame time position (e.g., n can be 10, 50, etc.), for two frames with the same time position... gi For the i-th first image, g i For the i-th second image, the SSIM index of structural similarity between them can be calculated, specifically:
[0060] in, and g i and g i 'Average pixel value, and Image g i and g i The standard deviation of the pixel values of ' For image g i and g i The covariance of the pixel values. C1 can be set to C2 can be set to The values of C1 and C2 can be adjusted according to the actual situation.
[0061] The image quality assessment score for the second video g' can be obtained using the average value of SSIM, i.e.:
[0062] in, The range of values can be Image quality score is recorded as The image quality score can be used as the first evaluation result for the second video.
[0063] In practical applications, the specific processing method for performing audio quality evaluation processing on the second video based on the first video in step S206 above to obtain the second evaluation result of the second video can be varied. The following provides an optional processing method, which may specifically include the processing steps B1 to B3.
[0064] In step B1, the first video is sampled, and the first peak signal-to-noise ratio corresponding to each sampling point in the first video and the second peak signal-to-noise ratio corresponding to each sampling point in the second video are determined.
[0065] In step B2, the audio quality score of the second video is determined based on the number of sampling points and the first peak signal-to-noise ratio and the second peak signal-to-noise ratio corresponding to the sampling points.
[0066] In step B3, a second evaluation result for the second video is determined based on the audio quality score of the second video.
[0067] In implementation, for each pair of videos That is, given the first video g and the second video g', the peak signal-to-noise ratio (PSNR) between them can be calculated to determine the PSNR of the second video. The audio quality assessment score, specifically, the PSNR calculation method can be:
[0068] in, This represents the number of sampling points. Furthermore, it can be determined that... The values are processed and normalized to... Between, record the current situation The value is The calculation method is as follows:
[0069] Where f(x) is the normalized audio quality score, which can be denoted as s. a ,s a The normalized audio quality score can be used as the second evaluation result for the second video if the score is ∈[0,1].
[0070] In step S208, the video quality assessment model is trained based on the first video, the first evaluation result, the second evaluation result, the third evaluation result, and the target quality assessment result to obtain the trained video quality assessment model.
[0071] In implementation, the server can use a video quality assessment model to perform image quality and audio quality assessment on the second video, respectively, to obtain the first image quality assessment result and the first audio quality assessment result of the second video. At the same time, the server can use the video quality assessment model to perform audio-visual synchronization assessment on the images and audio related to the detection object in the second video, to obtain the first audio-visual synchronization assessment result of the second video. Finally, the server can determine the predicted quality assessment result of the second video based on the first image quality assessment result, the first audio quality assessment result, and the first audio-visual synchronization assessment result.
[0072] Then, based on the first image quality assessment result, the first audio quality assessment result, the first audio-visual synchronization assessment result, the prediction quality assessment result, the first assessment result, the second assessment result, the third assessment result, and the target quality assessment result, the server determines the target loss value, so as to determine whether the video quality assessment model has converged according to the target loss value, and if the video quality assessment model has converged, the trained video quality assessment model is obtained.
[0073] Specifically, the server can determine a first loss value based on the first image quality assessment result and the first assessment result, a second loss value based on the first audio quality assessment result and the second assessment result, a third loss value based on the first audio-visual synchronization assessment result and the third assessment result, and a fourth loss value based on the predicted quality assessment result and the target quality assessment result. Finally, a target loss value can be determined based on one or more of the first, second, third, and fourth loss values, such as by weighting the first, second, third, and fourth loss values to obtain the target loss value.
[0074] Specifically, for each second video segment in the dataset The server can extract images from the second video at a preset frame interval, based on a preset extraction quantity D. For example, the preset frame interval can be 1 second, meaning one frame can be extracted every second (i.e., for a video with a frame rate of 30 frames per second, one frame can be extracted every 30 frames). Preset extraction quantity The frame rate can be 180, corresponding to a 3-minute video. To reduce computational load, each frame can be scaled to a fixed size, such as 521*521.
[0075] like Figure 3 As shown, the server can extract data using the 3D convolutional neural network in the first module of the video quality assessment model. Features of a frame image sequence are used to predict the image quality score of a video image. v .
[0076] Alternatively, the server can use a 3D convolutional neural network that includes the Inception module. For the Inception module, optional... , , These convolutional kernels of different sizes perform 3D convolutions, and the calculation results of these kernels are combined as input to the next layer of the neural network to extract features of the video image in different spaces and times. The last layer of the 3D convolutional neural network is a fully connected layer, and the feature vector calculated by the fully connected layer is denoted as F. image It can be used to characterize the second video. The image features are used to predict the image quality score of the second video based on this feature vector through the image quality prediction branch of the video quality assessment model. This score is denoted as s. v Simultaneously, the image quality score prediction loss L is calculated using MSE. image The calculation formula is: L image= .
[0077] For each second video segment The second video can be extracted. The corresponding audio data is denoted as g. audio For each g audio The file can be preprocessed to adjust it to a fixed sampling rate (optionally, use an audio sampling rate of 16K), a fixed duration (optionally, adjust to 3 minutes; for audio longer than 3 minutes, truncation can be performed, and for audio shorter than 3 minutes, zero padding can be performed), and a single channel (for audio that was originally multi-channel, a processing strategy of calculating the average of multiple channels can be adopted to process it into a single-channel audio file).
[0078] like Figure 3 As shown, the server can use the 1D convolutional neural network in the second module of the video quality assessment model to extract audio file g. audio Features over time. Using a 1D convolutional neural network (preferably, an Inception structure can be used in the 1D convolutional neural network to fuse features over different time lengths), a fully connected layer can be used in the last layer of the neural network to obtain features over the entire time domain data, denoted as F. audio_time。
[0079] For each audio file g audio The Mel spectrum can be used to convert audio files from time-domain data to spectral data, employing a Mel scale to make the transformed frequency-domain data more consistent with human auditory perception. The server can utilize a 2D convolutional neural network in the second module of the video quality assessment model to extract features from the Mel spectrum. In the final layer of the neural network, a fully connected layer is used to obtain features in the frequency-domain data, denoted as F. audio_frequency .
[0080] The features F in the time domain audio_time and the characteristic F in the frequency domain audio_frequency The vector obtained by connecting them is denoted as F. audio It can be used to characterize the second video. The audio features are used to predict the audio quality score of the second video, denoted as S, based on these audio features through the audio quality prediction branch of the video quality assessment model. a Meanwhile, the server can also use MSE to calculate the audio quality score prediction loss L. audio The calculation formula is: .
[0081] like Figure 3As shown, taking the detection object as a face as an example, the server can use a face detection model (such as the InsightFace model) to detect image frames containing faces in the second video g', and extract the image sequence containing faces based on a preset extraction quantity D (such as 100).
[0082] If the number of image frames containing faces in the second video is less than the preset extraction data D, the server can fill the gap with images whose pixel values are all 0. If the number of image frames containing faces in the second video is greater than the preset extraction data D, the server can truncate the extracted images containing faces.
[0083] The server can scale the image region containing the face in each extracted image to a fixed size (e.g., 120×120) to form a set of face region image sequences, which can be denoted as g. face_image' And record the time range t corresponding to the face region image sequence. range To extract the time range t range The audio data below can be denoted as g. face_audio' .
[0084] like Figure 3 As shown, the server can extract g using the 3D convolutional neural network in the third module of the video quality assessment model. face_image' The features of the image sequence are extracted using a 1D convolutional neural network in the third module of the video quality assessment model. face_audio' The audio features are extracted, and the extracted features are concatenated to obtain a vector, which can be denoted as F. sync .
[0085] Then, the audio-visual synchronization prediction branch in the third module can be used to predict the audio-visual synchronization category of the second video based on the concatenated vector, denoted as s. sync' Meanwhile, the server can also use cross-entropy to calculate the audio-visual synchronization branch prediction loss L. sync The calculation formula is: , Among them, ssync is the third evaluation result of the second video.
[0086] The server can use video image features F image Audio features F audio Audio-visual synchronization feature F sync Connect the vectors to obtain vector F. concat The dimension of a vector can be denoted as d. concat The server can use an attention mechanism to determine the MOS score of the second video. The three matrices Q, K, and V of the attention mechanism are calculated as follows:
[0087] Among them, W q W k W v These are the weight matrices Q, K, and V, respectively. The server can use A and F. concat Multiplying these features yields the fused attention features, which are then further processed using a fully connected layer to predict the MOS score of the second video, denoted as s. mos At the same time, the server can also use MES to calculate the MOS score prediction loss L. mos The calculation formula is: .
[0088] Thus, the overall loss (i.e., the target loss value) of the video quality prediction model can be calculated as follows:
[0089] Among them, L total For the target loss value, The weight values representing each loss (e.g., ) The server can train the video quality assessment model based on this target loss value until the model converges, resulting in the trained video quality assessment model, denoted as M. mos .
[0090] By collecting video datasets containing objects such as human faces for detection, and using streaming conversion tools to simulate different network conditions in real video calls, second videos of varying quality were obtained as training data. SSIM and PSNR were used to calculate the quality of video images and audio, and the presence of audio-visual desynchronization and MOS scores were manually labeled. A multi-branch, multi-task neural network prediction model was constructed, encompassing predictions of image quality, audio quality, audio-visual synchronization, and MOS scores. Then, an attention mechanism was used to fuse feature vectors from multiple models to predict the MOS score of the video.
[0091] Thus, in the main task of predicting video MOS scores, a multi-task learning framework that incorporates image quality prediction, audio quality prediction, and audio-visual synchronization prediction can improve the accuracy and robustness of the model's predictions. Furthermore, using the fusion features of face image sequence data and audio data to predict whether there is audio-visual desynchronization in the video can improve the accuracy and robustness of call video quality prediction.
[0092] Taking communication video as an example, the trained video quality assessment model M... mos Then, the model can be deployed on the test call terminal, and after the video call ends, the recorded video from the called end can be input into the trained video quality assessment model M.mos The recorded video is processed to perform a video quality assessment, resulting in a video quality evaluation result. This result is then used as the basis for evaluating the quality of video calls. If the video quality assessment results can be used to promptly identify areas and base stations with poor call quality, it can help improve the user's video call experience and increase user satisfaction.
[0093] In practical applications, step S208 above utilizes a video quality assessment model to perform audio-visual synchronization assessment processing on the first image and first audio related to the detection object in the target video. The specific processing methods for obtaining the audio-visual synchronization assessment results of the target video can vary. One optional processing method is provided below, such as... Figure 4 As shown, the specific process may include the following steps S2082 to S2088.
[0094] In step S2082, a first image related to the detected object is extracted from the target video using a video quality assessment model.
[0095] In step S2084, a sub-image containing the detected object is extracted from the first image, and the first audio corresponding to the time position of the sub-image in the target video is extracted.
[0096] In implementation, the server can extract the first image related to the detected object from the target video, and determine the image region where the detected object is located in each extracted first image as a sub-image to form a set of sub-image sequences of the detection region. Based on the time range corresponding to the set of sub-image sequences of the detection region, the audio data within that time range is determined as the first audio.
[0097] In step S2086, feature extraction processing is performed on the sub-image and the first audio respectively to obtain the first feature corresponding to the sub-image and the second feature corresponding to the first audio.
[0098] In implementation, such as Figure 3 As shown, the server can use the 3D convolutional neural network in the third module of the video quality assessment model to extract the first feature of the sub-image, and use the 1D convolutional neural network in the third module of the video quality assessment model to extract the second feature of the first audio.
[0099] In step S2088, the target video is subjected to audio-visual synchronization evaluation processing based on the first feature and the second feature to obtain the audio-visual synchronization evaluation result.
[0100] In implementation, the server can concatenate the first feature and the second feature to obtain the target feature vector. Then, it can use the audio-visual synchronization prediction branch in the third module to predict the audio-visual synchronization category corresponding to the target video based on the concatenated vector, so as to determine the audio-visual synchronization evaluation result of the target video according to the audio-visual synchronization category.
[0101] In practical applications, the video quality assessment model used in step S206 above to perform audio quality assessment processing on the target video can be processed in various ways to obtain the audio quality assessment results of the target video. One optional processing method is provided below, such as... Figure 5 As shown, the specific process may include the following steps S2062 to S2066.
[0102] In step S2062, the audio data corresponding to the target video is extracted using a video quality assessment model.
[0103] In step S2064, based on multiple preset time lengths, temporal feature extraction processing is performed on the audio data to obtain multiple temporal features.
[0104] In implementation, such as Figure 3 As shown, the server can use the 1D convolutional neural network in the second module of the video quality assessment model to extract features of audio data at multiple preset time lengths, thus obtaining multiple temporal features.
[0105] In step S2066, the frequency domain features of the audio data are extracted, and based on multiple time domain features and the frequency domain features of the audio data, the target video is subjected to audio quality assessment processing to obtain the audio quality assessment result of the target video.
[0106] In implementation, for the audio data corresponding to the target video, the server can use Mel spectrometry to convert the audio data from time-domain data to spectral data, and adopt Mel scale to make the transformed frequency domain data more consistent with human hearing. The server can use the 2D convolutional neural network in the second module of the video quality assessment model to extract the features of the Mel spectrometry, and use a fully connected layer in the last layer of the neural network to obtain the frequency domain features of the audio data.
[0107] Then, the server can concatenate multiple time-domain features and frequency-domain features of the audio data, and then predict the audio quality assessment result of the target video based on the concatenated features through the audio quality prediction branch in the video quality assessment model.
[0108] This specification provides a video quality assessment method. It receives a quality assessment request for a target video, and in response, acquires a pre-trained video quality assessment model. Using this model, it performs image and audio quality assessments on the target video, obtaining image and audio quality assessment results. Then, using the same model, it performs audio-visual synchronization assessment on a first image and first audio related to the detected object in the target video, obtaining an audio-visual synchronization assessment result. Based on these results, the overall quality assessment of the target video is determined. This approach offers two advantages: firstly, the video quality assessment model allows for rapid and accurate quality assessment of the target video; secondly, it leverages the multimodal information of the target video to comprehensively assess its quality across multiple dimensions (image, audio, and audio-visual synchronization), thereby improving both efficiency and accuracy in video quality assessment.
[0109] The above describes the video quality assessment method provided in the embodiments of this specification. Based on the same idea, the embodiments of this specification also provide a video quality assessment device, such as... Figure 6 As shown.
[0110] The video quality assessment device includes: a request receiving module 601, a model determination module 602, a first assessment module 603, a second assessment module 604, and a quality assessment module 605, wherein: The request receiving module 601 is used to receive a quality assessment request for the target video. The model determination module 602 is used to obtain a pre-trained video quality assessment model in response to the quality assessment request. The first evaluation module 603 is used to perform image quality and audio quality evaluation processing on the target video using the video quality evaluation model, and obtain the image quality evaluation result and audio quality evaluation result of the target video. The second evaluation module 604 is used to perform audio-visual synchronization evaluation processing on the first image and the first audio related to the detection object in the target video using the video quality evaluation model, so as to obtain the audio-visual synchronization evaluation result of the target video. A quality assessment module 605 is used to determine the quality assessment result of the target video based on the image quality assessment result, the audio quality assessment result, and the audio-visual synchronization assessment result. In this embodiment of the specification, the apparatus further includes: The video acquisition module is used to acquire a first video for training the video quality assessment model; The data enhancement module is used to adjust the video configuration parameters and / or video playback parameters of the first video to obtain the second video, and to obtain the third evaluation result and the target quality evaluation result of the second video. The third evaluation result is used to characterize the audio-visual synchronization of the second video, and the target quality evaluation result is used to characterize the video quality of the second video. The result acquisition module is used to perform image quality and audio quality evaluation processing on the second video based on the first video, and obtain a first evaluation result and a second evaluation result of the second video. The first evaluation result is used to characterize the image quality of the second video, and the second evaluation result is used to characterize the audio quality of the second video. The model training module is used to train the video quality assessment model based on the first video, the first evaluation result, the second evaluation result, the third evaluation result, and the target quality assessment result, so as to obtain the trained video quality assessment model.
[0111] In this embodiment of the specification, the second evaluation module 604 is used for: Using the video quality assessment model, a first image related to the detected object is extracted from the target video; Extract the sub-image containing the detected object from the first image, and extract the first audio from the target video corresponding to the time position of the sub-image; Feature extraction processing is performed on the sub-image and the first audio respectively to obtain a first feature corresponding to the sub-image and a second feature corresponding to the first audio. Based on the first feature and the second feature, the target video is subjected to audio-visual synchronization evaluation processing to obtain the audio-visual synchronization evaluation result.
[0112] In the embodiments of this specification, the result acquisition module is used for: Image extraction processing is performed on the first video and the second video respectively to obtain a first image and a second image with the same time position; The image quality score of the second video is determined based on the pixel values of the first image, the standard deviation of the pixel values of the first image, the pixel values of the second image, the standard deviation of the pixel values of the second image, and the covariance of the pixel values of the first image and the second image. Based on the image quality score of the second video, a first evaluation result for the second video is determined.
[0113] In the embodiments of this specification, the result acquisition module is used for: The first video is sampled, and the first peak signal-to-noise ratio corresponding to each sampling point in the first video and the second peak signal-to-noise ratio corresponding to each sampling point in the second video are determined. Based on the number of sampling points, the first peak signal-to-noise ratio and the second peak signal-to-noise ratio corresponding to the sampling points, the audio quality score of the second video is determined; Based on the audio quality score of the second video, a second evaluation result for the second video is determined.
[0114] In the embodiments described in this specification, the first evaluation module 603 is used for: Using the video quality assessment model, the audio data corresponding to the target video is extracted; Based on multiple preset time lengths, the audio data is subjected to temporal feature extraction processing to obtain multiple temporal features; The frequency domain features of the audio data are extracted, and based on the multiple time domain features and the frequency domain features of the audio data, the target video is subjected to audio quality assessment processing to obtain the audio quality assessment result of the target video.
[0115] This specification provides a video quality assessment device. It receives a quality assessment request for a target video, and in response, acquires a pre-trained video quality assessment model. Using this model, it performs image and audio quality assessments on the target video, obtaining image and audio quality assessment results. Then, using the same model, it performs audio-visual synchronization assessment on a first image and first audio related to a detection object in the target video, obtaining an audio-visual synchronization assessment result. Based on these results, the final quality assessment result for the target video is determined. Thus, on the one hand, the video quality assessment model allows for rapid and accurate quality assessment of the target video; on the other hand, it utilizes the multimodal information of the target video to comprehensively assess its quality across multiple dimensions (image, audio, and audio-visual synchronization), thereby improving both the efficiency and accuracy of video quality assessment.
[0116] The above are the video quality assessment devices provided in the embodiments of this specification. Based on the same idea, the embodiments of this specification also provide a video quality assessment device, such as... Figure 7 As shown.
[0117] The video quality assessment device can provide terminal equipment or servers, etc., for the above embodiments.
[0118] Video quality assessment devices can vary considerably depending on their configuration or performance. They may include one or more processors 701 and memory 702, with memory 702 storing one or more application programs or data. Memory 702 can be temporary or persistent storage. The application programs stored in memory 702 may include one or more modules (not shown in the figures), each module including a series of computer-executable instructions for the video quality assessment device. Furthermore, processor 701 may be configured to communicate with memory 702, executing the series of computer-executable instructions stored in memory 702 on the video quality assessment device. The video quality assessment device may also include one or more power supplies 703, one or more wired or wireless network interfaces 704, one or more input / output interfaces 705, and one or more keyboards 706.
[0119] Specifically, in this embodiment, the video quality assessment device includes a memory and one or more programs, wherein one or more programs are stored in the memory, and one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the video quality assessment device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: Receive a quality assessment request for the target video; In response to the quality assessment request, obtain a pre-trained video quality assessment model; Using the video quality assessment model, the target video is subjected to image quality and audio quality assessment processing respectively, to obtain the image quality assessment result and audio quality assessment result of the target video; Using the video quality assessment model, the first image and the first audio related to the detection object in the target video are subjected to audio-visual synchronization assessment processing to obtain the audio-visual synchronization assessment result of the target video; Based on the image quality assessment results, audio quality assessment results, and audio-visual synchronization assessment results of the target video, the quality assessment result of the target video is determined.
[0120] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the video quality assessment device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0121] This specification provides a video quality assessment device that receives a quality assessment request for a target video, and in response, acquires a pre-trained video quality assessment model. Using this model, it performs image and audio quality assessments on the target video, obtaining image and audio quality assessment results. Then, using the same model, it performs audio-visual synchronization assessment on a first image and first audio related to the detected object in the target video, obtaining an audio-visual synchronization assessment result. Based on these results, the device determines the overall quality assessment of the target video. Thus, on the one hand, the video quality assessment model allows for rapid and accurate quality assessment of the target video; on the other hand, it utilizes the multimodal information of the target video to comprehensively assess its quality across multiple dimensions (image, audio, and audio-visual synchronization), thereby improving both the efficiency and accuracy of video quality assessment.
[0122] Furthermore, based on the above Figures 1 to 5 The method shown in this specification, along with one or more embodiments, also provides a storage medium for storing computer-executable instruction information. In one specific embodiment, the storage medium can be a USB flash drive, optical disc, hard disk, etc. When the computer-executable instruction information stored in the storage medium is executed by a processor, it can achieve the following process: Receive a quality assessment request for the target video; In response to the quality assessment request, obtain a pre-trained video quality assessment model; Using the video quality assessment model, the target video is subjected to image quality and audio quality assessment processing respectively, to obtain the image quality assessment result and audio quality assessment result of the target video; Using the video quality assessment model, the first image and the first audio related to the detection object in the target video are subjected to audio-visual synchronization assessment processing to obtain the audio-visual synchronization assessment result of the target video; Based on the image quality assessment results, audio quality assessment results, and audio-visual synchronization assessment results of the target video, the quality assessment result of the target video is determined.
[0123] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the above-described storage medium embodiment is basically similar to the method embodiment, so the description is relatively simple; relevant parts can be referred to the description of the method embodiment.
[0124] This specification provides a storage medium that receives a quality assessment request for a target video, and in response to the request, acquires a pre-trained video quality assessment model. Using this model, it performs image and audio quality assessments on the target video, obtaining image and audio quality assessment results. Then, using the same model, it performs audio-visual synchronization assessment on a first image and first audio related to the detected object in the target video, obtaining an audio-visual synchronization assessment result. Based on these results, the final quality assessment result of the target video is determined. Thus, on the one hand, the video quality assessment model allows for fast and accurate quality assessment of the target video; on the other hand, it utilizes the multimodal information of the target video to comprehensively assess its quality across multiple dimensions (image, audio, and audio-visual synchronization), thereby improving both efficiency and accuracy in video quality assessment.
[0125] Furthermore, based on the above Figures 1 to 5 The method shown in this specification, along with one or more embodiments, also provides a computer program product including a computer program that, when executed by a processor, performs the following process: Receive a quality assessment request for the target video; In response to the quality assessment request, obtain a pre-trained video quality assessment model; Using the video quality assessment model, the target video is subjected to image quality and audio quality assessment processing respectively, to obtain the image quality assessment result and audio quality assessment result of the target video; Using the video quality assessment model, the first image and the first audio related to the detection object in the target video are subjected to audio-visual synchronization assessment processing to obtain the audio-visual synchronization assessment result of the target video; Based on the image quality assessment results, audio quality assessment results, and audio-visual synchronization assessment results of the target video, the quality assessment result of the target video is determined.
[0126] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the above-described embodiment of a computer program product is relatively simple in description because it is fundamentally similar to the method embodiment; relevant parts can be referred to the description of the method embodiment.
[0127] This specification provides a computer program product that receives a quality assessment request for a target video, and in response to the request, acquires a pre-trained video quality assessment model. Using this model, it performs image and audio quality assessments on the target video, obtaining image and audio quality assessment results. Then, using the same model, it performs audio-visual synchronization assessment on a first image and first audio related to the detected object in the target video, obtaining an audio-visual synchronization assessment result. Based on these results, the final quality assessment result for the target video is determined. Thus, on the one hand, the video quality assessment model allows for fast and accurate quality assessment of the target video; on the other hand, it utilizes the multimodal information of the target video to comprehensively assess its quality across multiple dimensions (image, audio, and audio-visual synchronization), thereby improving both efficiency and accuracy in video quality assessment.
[0128] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0129] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0130] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0131] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0132] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in one or more software and / or hardware.
[0133] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0134] Embodiments in this specification are described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable parallel device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable parallel device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0135] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable fraud device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0136] These computer program instructions can also be loaded onto a computer or other programmable device, causing a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0137] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0138] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0139] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0140] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0141] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0142] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0143] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0144] The above description is merely an embodiment of this specification and is not intended to limit this document. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. A video quality assessment method, characterized in that, The method includes: Receive a quality assessment request for the target video; In response to the quality assessment request, obtain a pre-trained video quality assessment model; Using the video quality assessment model, the target video is subjected to image quality and audio quality assessment processing respectively, to obtain the image quality assessment result and audio quality assessment result of the target video; Using the video quality assessment model, the first image and the first audio related to the detection object in the target video are subjected to audio-visual synchronization assessment processing to obtain the audio-visual synchronization assessment result of the target video; Based on the image quality assessment results, audio quality assessment results, and audio-visual synchronization assessment results of the target video, the quality assessment result of the target video is determined.
2. The method according to claim 1, characterized in that, Before obtaining the pre-trained video quality assessment model, the following is also included: Obtain the first video used to train the video quality assessment model; The video configuration parameters and / or video playback parameters of the first video are adjusted to obtain the second video, and the third evaluation result and the target quality evaluation result of the second video are obtained. The third evaluation result is used to characterize the audio-visual synchronization of the second video, and the target quality evaluation result is used to characterize the video quality of the second video. Based on the first video, the second video is subjected to image quality and audio quality evaluation processing respectively to obtain a first evaluation result and a second evaluation result of the second video. The first evaluation result is used to characterize the image quality of the second video, and the second evaluation result is used to characterize the audio quality of the second video. Based on the first video, the first evaluation result, the second evaluation result, the third evaluation result, and the target quality evaluation result, the video quality evaluation model is trained to obtain the trained video quality evaluation model.
3. The method according to claim 1, characterized in that, The step of using the video quality assessment model to perform audio-visual synchronization assessment processing on the first image and first audio related to the detection object in the target video to obtain the audio-visual synchronization assessment result of the target video includes: Using the video quality assessment model, a first image related to the detected object is extracted from the target video; Extract the sub-image containing the detected object from the first image, and extract the first audio from the target video corresponding to the time position of the sub-image; Feature extraction processing is performed on the sub-image and the first audio respectively to obtain a first feature corresponding to the sub-image and a second feature corresponding to the first audio. Based on the first feature and the second feature, the target video is subjected to audio-visual synchronization evaluation processing to obtain the audio-visual synchronization evaluation result.
4. The method according to claim 2, characterized in that, The step of performing image quality assessment processing on the second video based on the first video to obtain a first assessment result for the second video includes: Image extraction processing is performed on the first video and the second video respectively to obtain a first image and a second image with the same time position; The image quality score of the second video is determined based on the pixel values of the first image, the standard deviation of the pixel values of the first image, the pixel values of the second image, the standard deviation of the pixel values of the second image, and the covariance of the pixel values of the first image and the second image. Based on the image quality score of the second video, a first evaluation result for the second video is determined.
5. The method according to claim 2, characterized in that, The step of performing audio quality assessment processing on the second video based on the first video to obtain a second assessment result for the second video includes: The first video is sampled, and the first peak signal-to-noise ratio corresponding to each sampling point in the first video and the second peak signal-to-noise ratio corresponding to each sampling point in the second video are determined. Based on the number of sampling points, the first peak signal-to-noise ratio and the second peak signal-to-noise ratio corresponding to the sampling points, the audio quality score of the second video is determined; Based on the audio quality score of the second video, a second evaluation result for the second video is determined.
6. The method according to claim 1, characterized in that, The step of using the video quality assessment model to perform audio quality assessment processing on the target video to obtain the audio quality assessment result of the target video includes: Using the video quality assessment model, the audio data corresponding to the target video is extracted; Based on multiple preset time lengths, the audio data is subjected to temporal feature extraction processing to obtain multiple temporal features; The frequency domain features of the audio data are extracted, and based on the multiple time domain features and the frequency domain features of the audio data, the target video is subjected to audio quality assessment processing to obtain the audio quality assessment result of the target video.
7. A video quality assessment device, characterized in that, The device includes: The request receiving module is used to receive quality assessment requests for the target video. The model determination module is used to obtain a pre-trained video quality assessment model in response to the quality assessment request. The first evaluation module is used to perform image quality and audio quality evaluation processing on the target video using the video quality evaluation model, and obtain the image quality evaluation result and audio quality evaluation result of the target video. The second evaluation module is used to perform audio-visual synchronization evaluation processing on the first image and the first audio related to the detection object in the target video using the video quality evaluation model, so as to obtain the audio-visual synchronization evaluation result of the target video. The quality assessment module is used to determine the quality assessment result of the target video based on the image quality assessment result, the audio quality assessment result, and the audio-visual synchronization assessment result.
8. A video quality assessment device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the video quality assessment method as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the video quality assessment method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the steps of the video quality assessment method according to any one of claims 1 to 6.