A subjective quality evaluation method for real-time video communication
Through the acquisition and uniform sampling of user-generated content videos, combined with real-time communication testing platform and network simulation tools, a real-time video communication data set was constructed, which solved the problems of standardization, one-sidedness and lack of delayed evaluation of existing data sets, achieved more realistic and extensive subjective quality evaluation, and enhanced the effectiveness of the experiment.
Patent Information
- Application Number
- CN202111679772.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-11-24
- Filing Date
- 2021-12-31
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-31
AI Technical Summary
The existing real-time video subjective quality evaluation dataset has the problems of standardization of source video selection, distorted video one-sidedness and delayed evaluation, and it is difficult to truly reflect the user's subjective experience in real-time video communication.
By collecting user-generated content videos, uniform sampling is performed to cover different content characteristics and network conditions, combining real-time communication testing platform and network simulation tools, different network environments and damage situations are simulated, video images are recorded in real time and network damage are recorded, real-time video communication data sets are constructed, and videos are displayed on different terminals for users to score.
It improves the authenticity and breadth of subjective experimental samples, can more accurately reflect the visual quality deterioration caused by encoding and transmission damage in real-time video communication, enhances the effectiveness and accuracy of the experiment, and provides strong data support for future analysis of the influencing factors of real-time video quality.
Smart Images

Figure CN114401364B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video quality evaluation, and particularly to a subjective quality evaluation method for real-time video communication. Background Art
[0002] With the advent of the self-media era and the innovation of video live streaming technology, Real Time Communications (RTC) services are booming. Especially since the outbreak of the epidemic, live applications such as online meetings, telemedicine, and online education have received increasing attention and favor. Users are no longer restricted by the venue and can establish one-to-one or one-to-many connections anytime and anywhere. Through the audio and video devices on the computer or mobile phone, they can record audio and video streams in real time and send them to other user terminals to achieve end-to-end interaction. However, during the video transmission process, due to the dynamic changes in network conditions, the live video is often damaged in different types and degrees, resulting in a decline in visual quality, such as decreased clarity, stuttering, and latency. Currently, the live video system mainly uses the technology of adaptive bitrate streaming to transmit video, that is, according to the current network environment, without compromising the Quality of Experience (QoE) of users, dynamically adjust the video coding bitrate. However, how to measure the impact of coding damage and transmission damage on the human eye's visual quality is still an urgent problem to be solved. Currently, there are subjective quality evaluation and objective quality evaluation methods for video quality assessment. Among them, subjective quality evaluation usually requires a large number of users to score the video in the same experimental environment and establish a corresponding Mean Opinion Score (MOS) database for subjective quality, with relatively high accuracy and can be used as a standard for objective quality modeling.
[0003] However, the currently publicly available real-time video subjective quality evaluation datasets all have the following problems: (1) The source videos select standard high-definition quality videos as references. However, in actual real-time video communication sessions, the videos are usually generated by users and have great differences in content and quality, with characteristics such as high content complexity, diverse types, and rapid scene changes. (2) The distorted videos are obtained by adding different types and intensities of artificial coding damage and transmission damage to the high-definition source videos, which has certain one-sidedness and singularity and cannot reflect the real damage situation. (3) In addition, different from video-on-demand, real-time video communication has more stringent requirements for latency, and the currently publicly available datasets lack the evaluation dimension of interactive latency. Therefore, it is necessary to construct a real-time video dataset and a subjective quality evaluation method to obtain the subjective scores of users for real-time videos with different content characteristics and different damages, and provide data labels for objective quality modeling. Summary of the Invention
[0004] Aiming at the defects existing in the above-mentioned prior art, the main object of the present invention is to provide a method for subjective quality evaluation of real-time video communication, so as to further improve the authenticity and universality of subjective experiment samples.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] A method for subjective quality evaluation of real-time video communication includes the following steps:
[0007] (1) Collect user-generated content videos, and uniformly sample the videos according to perceptual information in different dimensions to ensure that the selected videos can be widely distributed in the feature space of spatial perceptual information - temporal perceptual information;
[0008] (2) Screen the network trajectories according to the average value and standard deviation of the bandwidth to ensure that the selected network trajectories can cover different network conditions;
[0009] (3) Based on a real-time communication test platform, transmit the source video in different network environments to simulate different coding damages and transmission damages;
[0010] (4) Record the video images at the sending end and receiving end in real time, and record the network damage situation of the video to construct a dataset for real-time video communication;
[0011] (5) Display the videos recorded in step (4) on different terminals to facilitate the subjects to score the visual quality of the videos.
[0012] Further, in the step (1), the user-generated content videos do not contain long-term stillness or black screen. After collecting the videos, first filter them according to the content, quality and resolution of the videos, and then classify them according to the generation method of the videos, and finally form three types of video sets: natural videos, unnatural videos and mixed videos.
[0013] Further, uniformly sample each type of video set according to spatial perceptual information and temporal perceptual information, and in the feature space composed of these two dimensions, the Euclidean distance between the same type of videos needs to be greater than a set threshold.
[0014] Further, in the step (2), the ratio between the standard deviation and the average value of the network bandwidth is used to characterize the volatility of the network, and the network trajectory data is uniformly sampled according to these two parameters of the ratio and the average value, where the time length of the network trajectory needs to be slightly longer than the playing length of the video.
[0015] Further, in the step (3), a source video is transmitted using an end-to-end real-time communication test platform, and different network environments are reproduced using network simulation tools and network traces, and the video images of the sending end and the receiving end are recorded in real-time and synchronously.
[0016] Further, in the step (5), a single excitation method is adopted to display the recorded video on different terminals, so that the subjects can score the visual quality from various evaluation indexes of the video.
[0017] Compared with the prior art, the method of the present invention uses real network traces to simulate the network transmission conditions of real-time video communication, and collects various visual quality deterioration conditions caused by coding and transmission damage of source videos with different content characteristics during the live broadcast process, avoiding the singularity and one-sidedness of artificial distorted videos, and improving the authenticity and universality of subjective experiment samples. In addition, according to the synchronously recorded sending video and receiving video images, the subjects can intuitively feel the impact of delay on the video viewing quality, improve the effectiveness and accuracy of the experiment, and provide strong data support for future analysis of the influencing factors of real-time video quality. Description of the Drawings
[0018] Figure 1 is a flowchart of the video subjective quality evaluation method provided by the embodiment of the present invention;
[0019] Figure 2 is an example diagram of the distribution of the source video data set in the SI-TI feature space before and after uniform sampling, (a) before uniform sampling, (b) after uniform sampling;
[0020] Figure 3 is an example of the distribution of the network trace data set in the bandwidth average - fluctuation parameter feature space before and after uniform sampling, (a)(c) before uniform sampling, (b)(d) after uniform sampling;
[0021] Figure 4 is an example diagram of the real-time video simulation transmission link provided by the embodiment of the present invention;
[0022] Figure 5 is an example diagram of the user scoring interface in the real-time video subjective experiment provided by the embodiment of the present invention;
[0023] Figure 6 is an example diagram of the user viewing interface in the real-time video subjective experiment provided by the embodiment of the present invention, (a) computer terminal, (b) mobile phone terminal. Detailed Embodiments
[0024] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0025] The overall system framework of the video subjective quality evaluation method according to the embodiments of the present invention is as follows Figure 1 as shown, and its specific working process is as follows:
[0026] (1) Collect existing user-generated content (UGC) videos on the network to establish an initial video dataset. Then, filter and classify the initial video dataset according to the content, quality, and resolution of the videos;
[0027] (2) Uniformly sample each category of video set in two dimensions: spatial perceptual information SI (Spatial perceptual Information) and temporal perceptual information TI (Temporal perceptual Information), to ensure that the selected videos are widely distributed in the SI-TI feature space, thereby ensuring the diversity of the final source video dataset;
[0028] (3) Collect publicly available network trajectory data, intercept the network bandwidth within a period of time, and divide it into three types of network environments: 3G, 4G, and Wi-Fi according to the source of the network trajectory;
[0029] (4) In each category of network trajectory set, sample according to two parameters: the average value and standard deviation of the bandwidth, to ensure that the selected network trajectories can cover different network conditions;
[0030] (5) Use an end-to-end real-time communication test platform based on the Web Real Time Communication (WebRTC) protocol framework to transmit the source video, and use network simulation tools and network trajectories to simulate different network environments, and record the video pictures of the sender and receiver in real time, and record the distribution of the sending bit rate and receiving bit rate of the video, as well as stuttering and interaction delay, etc.;
[0031] (6) The subjects watch the recorded videos on the computer side and the mobile phone side respectively, and give scores for the visual quality of the videos from the perspectives of video clarity, smoothness, interaction delay, etc.
[0032] l. Source video dataset
[0033] In a real-time video communication session, the video content is usually generated by users and varies greatly in terms of content and quality. In this embodiment, original videos with various content themes are collected from publicly available UGC video datasets on the Internet, such as speeches / lectures, games, concerts, sports, etc., which basically cover the daily application scenarios of real-time video communication. To avoid the interference of the video's own content on parameters such as stuttering and latency, the selected UGC videos do not contain long-term stillness or black screens. In addition, UGC videos with poor video quality and average per-frame image quality are filtered out and divided into natural videos, non-natural videos, and mixed videos according to their generation methods. Among them, natural videos are mainly from natural scenes captured by cameras, screen videos are composed of text, graphics, or animations generated by computers, and mixed videos have both.
[0034] The content characteristics of a video can be measured by SI and TI. SI represents the spatial complexity of the video:
[0035] SI = max time {std space [Sobel(F n )]} (1)
[0036] Among them, first, Sobel filtering is performed on each video frame, then the standard deviation of the filtered video frame is calculated, and finally, the maximum standard deviation is selected as SI. TI represents the temporal complexity of the video:
[0037] TI = max time {std space [M n (i, j)]} (2)
[0038] M n (i, j) = F n (i, j) - F n-1 (i, j) (3)
[0039] Among them, first, the pixel difference M n (i, j) at the same position between two adjacent frames is calculated, then the standard deviation of the frame difference image is calculated, and finally, the maximum standard deviation is selected as TI. In this experiment, to avoid the interference of abnormal peaks, the 90th percentile maximum values of SI and TI are respectively selected as the maximum values.
[0040] Calculate the SI and TI values for different categories of videos respectively and map them into the two-dimensional SI-TI feature space, as shown in Figure 2 (a) of, and the video set can be represented by a group of vectors:
[0041]
[0042] Among them, qi is a random variable representing the SI and TI values of the video. K represents the number of videos, and M represents the number of feature attributes of each video, that is, two dimensions of SI and TI. represents the probability mass function (PMF) of the selected video set in each interval. To ensure that the selected videos can cover different content characteristics, the goal of this experiment is to select a video set that can be evenly distributed in the SI-TI feature space:
[0043]
[0044] Among them, represents selecting N video sets, represents a uniform PMF distribution. Use a set of binary vectors to represent whether a video is selected, and use a set of binary matrices to represent whether the j-th video falls into the i-th quantization interval of the target distribution on the m-th feature attribute. 1 means falling into, and 0 means not falling into. H represents the number of quantization intervals. In this experiment, the number of quantization intervals for natural videos and non-natural videos is set to 6 and 3 respectively. The problem of uniform sampling can be modeled as an optimization problem, that is, to find a set of N video sets so that these points can be approximately evenly distributed on different feature attributes:
[0045]
[0046] To prevent the interference of the correlation between video SI and TI on the selection, it is also necessary to minimize the correlation between different feature attributes. Since the correlation is calculated by covariance, this constraint condition can be transformed into minimizing the covariance matrix between different characteristic attributes in the sub-video set:
[0047]
[0048] Among them, is the covariance matrix between sub-video sets . At the same time, it is also necessary to ensure that the Euclidean distance between selected similar videos is greater than the set threshold. In this experiment, 12 natural videos, 12 non-natural videos and 8 mixed videos are finally selected. As shown in Figure 2 (b), the video points are evenly distributed in the SI-TI feature space, characterizing the diversity of the source video dataset.
[0049] 2. Network trajectory dataset
[0050] Collect publicly available network trajectory datasets, and extract the network bandwidth within a certain period of time. The length of the extraction time needs to be slightly longer than the playback length of the video. The selected datasets are HSDPA, Belgium, and FCC, corresponding to three different types of networks: 3G, 4G, and Wi-Fi. Calculate the average value and standard deviation of the bandwidth for these three different types of network trajectories respectively, and characterize the network volatility through the ratio between the two. In the feature space of the bandwidth mean - volatility parameter, plot the distribution maps of different network trajectories, as shown in Figure 3 Figures (a) and (c) in Figure 3 Figures (b) and (d) in
[0051] 3. Synchronously record the transmitted video and the received video
[0052] Based on the WebRTC protocol framework, this embodiment builds an end-to-end real-time communication test platform to simulate the transmission process of real-time video, as shown in Figure 4 The local computer captures the source video through a virtual camera and sends it to the server in real time through the test platform. The server then forwards the real-time stream back to the local computer through the streaming media protocol, so that the local computer can synchronously display and record the transmitted video and the received video. In addition, the server controls the change of the network environment through network simulation tools and network trajectories, dynamically adjusts the network bandwidth, and can record real-time video data with different coding impairments and transmission impairments. During the communication process, record the distribution of the video transmission bitrate and reception bitrate, as well as stuttering and interaction delay, etc., for subsequent analysis of the control factors affecting the subjective quality of the video.
[0053] 4. Subjective quality evaluation experiment
[0054] After generating the real-time video dataset with different combinations of damage types and intensities, it is necessary to further conduct a subjective quality evaluation experiment to let users score the visual quality of the video, such as clarity, smoothness, response speed, etc. The rating criteria for scoring are as shown in Figure 5As shown in the figure. The videos are divided into 5 levels according to the visual quality, where the 5th level represents the highest video quality, with clear picture quality, smooth playback, and no communication delay. The 1st level represents the worst video quality, with blurred picture quality, stuttering, and intolerable delay. To ensure the stability and consistency of the experimental data, each subject should receive a certain amount of training before the experiment to understand the video degradation corresponding to different scoring levels. Following the recommendations of the International Telecommunication Union (ITU) standards, in this embodiment, a single-stimulus method is adopted during the subjective experiment, that is, only one recorded video is displayed on the computer or mobile phone each time, such as Figure 6 shown. To ensure the fairness of the experimental data, during the entire experimental process, the ambient light is normal, there is no noise interference, and the video display resolution is 1920×1080 on both the computer and mobile phone. Each video is scored by at least 15 subjects, and each subject takes a break every half hour to avoid visual fatigue.
[0055] The specific embodiments of the present invention have been described in detail above. It should be noted that the protection scope of the present invention is not limited thereto. Those skilled in the art can make changes and modifications to the present invention within the scope of the claims, but they all fall within the protection scope of the present invention.
Claims
1. A subjective quality evaluation method for real-time video communication, characterized in that, it includes the following steps: (1) Collect user-generated content videos, and uniformly sample the videos according to perceptual information in different dimensions to ensure that the selected videos can be widely distributed in the feature space of spatial perceptual information - temporal perceptual information; (2) Screen the network traces according to the average value and standard deviation of the bandwidth to ensure that the selected network traces can cover different network conditions; specifically: use the ratio between the standard deviation and the average value of the network bandwidth to characterize the volatility of the network, and uniformly sample the network trace data based on these two parameters of the ratio and the average value, where the time length of the network trace needs to be slightly longer than the playback length of the video; (3) Based on a real-time communication test platform, transmit the source video in different network environments to simulate different coding damages and transmission damages; among them, the transmission damage includes interaction delay; (4) Record the video images of the sending end and the receiving end in real time, and record the network damage situation of the video to construct a dataset for real-time video communication; (5) Display the videos recorded in step (4) on different terminals to facilitate the subjects to score the visual quality of the videos.
2. A subjective quality evaluation method for real-time video communication according to claim 1, characterized in that, in the step (1), the user-generated content video does not contain long-term stillness or black screen. After collecting the video, first filter it according to the content, quality and resolution of the video, and then classify it according to the generation method of the video, and finally form three types of video sets: natural videos, non-natural videos and mixed videos.
3. A subjective quality evaluation method for real-time video communication according to claim 2, characterized in that, uniformly sample each type of video set according to spatial perceptual information and temporal perceptual information, and in the feature space composed of these two dimensions, the Euclidean distance between the same type of videos needs to be greater than a set threshold.
4. A subjective quality evaluation method for real-time video communication according to claim 1, characterized in that, in the step (3), use an end-to-end real-time communication test platform to transmit the source video, and use network simulation tools and network traces to reproduce different network environments, and record the video images of the sending end and the receiving end in real time and synchronously.
5. A subjective quality evaluation method for real-time video communication according to claim 1, characterized in that, in the step (5), adopt a single-stimulus method to display the recorded videos on different terminals to facilitate the subjects to score the visual quality from each evaluation index of the videos.
Citation Information
Patent Citations
Method for subjective testing of video quality
CN101783971A
Real-time video communication quality evaluation method
CN109120924A