System and method for determining video codec performance in real-time communications over the Internet
By using network models and quality analyzer models in real-time Internet communication to evaluate the end-to-end delay and quality indicators of video codecs, the problem of inaccurate video codec performance evaluation in the existing technology is solved, and more accurate performance comparison is achieved.
Patent Information
- Application Number
- CN202310121472.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-05
- Filing Date
- 2023-02-15
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2043-02-15
AI Technical Summary
Existing general test conditions lack key measures when evaluating the performance of video codecs for Internet real-time communication (RTC) use cases. They are unable to effectively measure factors such as video freezes and delays, resulting in inaccurate evaluation results.
A network model and quality analyzer model are used to simulate real-world network conditions to evaluate the end-to-end delay and video quality indicators of video codecs, such as PSNR, SSIM, VMAF, etc., to recover packet loss through FEC and PR schemes, and to evaluate codec performance using lookup tables and calculation methods.
It enables reliable evaluation of video codec performance under different network conditions, provides more accurate video quality assessment results, and can compare the performance of different codecs.
Smart Images

Figure CN116614479B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] The present invention claims priority to U.S. patent application No. 63 / 311,205, filed on February 17, 2022, and entitled “SYSTEM AND METHOD FOR DETERMINING VIDEO CODEC PERFORMANCE IN REAL-TIME COMMUNCATION OVER INTERNET.” The present invention also claims priority to U.S. patent application No. 18 / 093788, filed on January 5, 2023, and entitled “SYSTEM AND METHOD FOR DETERMINING VIDEO CODEC PERFORMANCE IN REAL-TIME COMMUNCATION OVER INTERNET.” Technical Field
[0003] The present invention relates to real-time communication (RTC) on a network, and more particularly to a system and method for evaluating the quality of experience of RTC use cases on the Internet, and more particularly to a system and method for evaluating the performance of video codec solutions in RTC use cases on the Internet. Background Art
[0004] RTC on the Internet is widely used in many areas of our daily life and work. RTC video traffic is usually transmitted on the Internet in the form of data packets. Due to network congestion and changes in signal strength, data packet loss may occur in the transmission network. If one or more frame packets are lost, the receiving end may not be able to decode the corresponding frames. Strategies such as Forward Error Correction (FEC) or Packet Retransmission (PR) can be used to deal with the problem of packet loss in data transmission, but they usually also result in the transmission of more redundant data or additional delays. FEC inserts redundant data through channel coding and occupies part of the available bandwidth, while also causing a small delay. PR will repeatedly send lost data packets after receiving a retransmission request from the receiving end, resulting in greater delay.
[0005] Existing Common Test Conditions (CTC) lack key measures for evaluating the quality of experience for RTC use cases. Factors such as video freezes and latency need to be measured. Therefore, a new methodology and system architecture are needed to evaluate the performance of video codec solutions for RTC use cases and applications. When testing codecs within this system architecture, the measurement results for the test cases are reported and compared with those of a baseline video codec. Summary of the Invention
[0006] In general, the present invention provides a network model and a quality analyzer model based on various embodiments to evaluate the performance of video codecs in RTC applications. The present invention simulates some typical real-world network conditions. Under these conditions, the encoded video stream is transmitted from the transmitter to the receiver, and the end-to-end (E2E) delay and the smoothness of the received video are measured, as well as some existing video quality indicators, such as Peak Signal-To-Noise Ratio (PSNR), Structural Similarity Index (SSIM), Video Multimethod Assessment Fusion (VMAF), etc. They are used as performance indicators to measure RTC video quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] This patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the applicable fee.
[0008] The technical features of the present invention will be particularly pointed out in the claims. The present invention, its construction and use will also be better understood by reference to the following description and the accompanying drawings which form a part of the description. All drawings of the present invention also constitute a part of the present invention. In the drawings, like reference numerals represent like parts:
[0009] Figure 1 is a block diagram of an RTC system according to an embodiment of the present invention.
[0010] Figure 2 is a block diagram of an RTC device with an improved RTC application according to an embodiment of the present invention.
[0011] Figure 3 FIG. 4 is a block diagram of an RTC system according to an embodiment of the present invention.
[0012] Figure 4FIG. 4 is an example block diagram of an RTC system test architecture drawn according to an embodiment of the present invention.
[0013] Figure 5 FIG. 4 is a block diagram of a video codec performance evaluation system for RTC drawn according to an embodiment of the present invention.
[0014] Figure 6 FIG. 4 is a schematic diagram illustrating additional delay introduced by backward referencing according to an embodiment of the present invention.
[0015] Figure 7 The present invention is a flowchart of a process of determining a performance index of a video codec in real-time communication by a video codec performance evaluation system according to an embodiment of the present invention.
[0016] Figure 8 The present invention is a flowchart of a process of determining a frame decoding data item set as a quantization parameter value and a seed by a video coding performance evaluation system according to an embodiment of the present invention.
[0017] Figure 9 The present invention is a flowchart of a process of determining a frame decoding data item set as a quantization parameter value and a seed by a video coding performance evaluation system according to an embodiment of the present invention.
[0018] It will be appreciated by those skilled in the art that, in order to clearly and simply illustrate the above drawings, the various components in the drawings are not necessarily drawn to scale. The sizes of some components in the drawings may be exaggerated relative to other components to aid understanding of the present invention. In addition, the specific order of certain elements, parts, assemblies, modules, steps, operations, events and / or processes described or illustrated herein may not be necessary in actual applications. It will be appreciated by those skilled in the art that, for simplicity and clarity, those useful and / or necessary elements that are well-known and easily understood in existing feasible embodiments may not be described herein so that various embodiments of the present invention can be clearly presented. DETAILED DESCRIPTION
[0019] Please refer to the attached picture, Figure 1 is a block diagram of an RTC system 100. The RTC system 100 includes a set (i.e., one or more) of participating electronic devices, such as, Figure 1Electronic devices 102, 104, 106, and 108 are shown. The electronic devices 102, 104, 106, and 108 in the RTC system 100 communicate with each other via the Internet 110. When an electronic device (e.g., electronic device 102) sends video (or audio) data to other electronic devices (e.g., electronic devices 106 and 108), the electronic device (e.g., electronic device 102) is called a sender or a transmitter, and the other devices (e.g., electronic devices 106 and 108) are called receivers or receivers of the video (or audio) data. They are connected to the Internet 110 via a local area network (e.g., a Wi-Fi network, a public cellular phone network, an Ethernet network, etc.). Figure 2 Each of the electronic devices 102 , 104 , 106 , and 108 will be further described.
[0020] Please refer to Figure 2 , Figure 2 The present invention is a simplified block diagram of a real-time video communication device (e.g., electronic device 102). The electronic device 102 includes a processing unit 202 (e.g., a central processing unit (CPU)), a certain amount of memory 204 operably coupled to the processing unit 202, one or more user input interfaces 206 operably coupled to the processing unit 202 (e.g., a mouse interface, a keyboard interface, a touch screen interface, etc.), an audio output interface 210 operably coupled to the processing unit 202, a network interface 216 operably coupled to the processing unit 202, a video output interface 214 operably coupled to the processing unit 202 (e.g., a display screen), a video input interface 212 operably coupled to the processing unit 202 (e.g., a camera), and an audio input interface 208 operably coupled to the processing unit 202 (e.g., a microphone). The electronic device 102 also includes an operating system 220 and a dedicated real-time video communication software application 222 adapted for execution by the processing unit 202. The real-time video communication software application 222 is programmed using one or more computer programming languages (eg, C, C++, C#, Java, etc.) and includes components and modules for performing RTC communication over the Internet.
[0021] In the RTC use case, such as Figure 3 and Figure 4 As shown, in this embodiment, the transmitting end is electronic device 104 and the receiving end is electronic device 108. Video encoder 402 and packetizer 404 run on the transmitting end (i.e., electronic device 104), and depacketizer 406 and video decoder 408 run on the receiving end (i.e., electronic device 108). Figure 4An RTC system and its data stream for transmitting a video sequence 412 from a transmitting end to a receiving end are shown and are indicated by reference numeral 400. Electronic device 104 communicates with electronic device 108 via network 302 (e.g., the Internet). It should be noted that if electronic device 108 sends data (e.g., video data) to electronic device 104, electronic device 104 becomes the receiving end of the data, and electronic device 108 becomes the transmitting end of the data. Video encoder 402 encodes video sequence 412 and generates a coded bitstream 414, which is then sent to packetizer 404. Video sequence 412 consists of a number of pictures or video frames. A frame can be further divided into slices or slices. The video unit used for encoding can be a frame, a slice, or a slice.
[0022] Packetizer 404 packages the encoded bitstream 414 into a number of data packets 416 and transmits these data packets 416 (also referred to as transmission packets) via network 302 to a video decoder 408 at the receiving end (i.e., electronic device 108). In a simplified manner, network conditions can be represented by a packet loss rate r, a bandwidth limit b, and an end-to-end (E2E) network delay or latency d. To better cope with packet loss during transmission, schemes such as PR and FEC can be used to recover lost packets during transmission.
[0023] When the electronic device 108 acts as a receiving end, the depacketizer 406 receives a data packet (i.e., a received data packet 418) and parses the received data packet 418 according to a relevant scheme (e.g., FEC) to achieve depacketization. When the received data packets 418 of a video unit are fully received and recovered, they are then sent to the video decoder 408 as a recovered bitstream 420. The video decoder 408 then decodes and reconstructs the video unit. The video decoder 408 outputs the decoded video for playback or other purposes (e.g., storage and backup). The subsystem 450 of the RTC system 400 includes the packetizer 404, the network 302, and the depacketizer 406.
[0024] When a video unit depends on a reference frame, the reference frame must be successfully decoded before the video unit can be successfully decoded. If a video unit depends on a reference frame and the reference frame is not successfully decoded, the video unit may not be decoded. The reference frame cannot be successfully decoded because some of its data packets have been lost during transmission. To alleviate this problem, some solutions have the receiving end send feedback information 424 to the sending end indicating whether the video unit was successfully decoded, thereby helping the video encoder 402 select reliable reference frames or encode instantaneous decoder refresh (IDR) frames.
[0025] The video encoder 402, the packetizer 404, the depacketizer 406, and the video decoder 408 are all software components, which are a collection of computer programs written in a computer programming language (e.g., C, C++, C#, Java, etc.).
[0026] Because the network 302 is a time-varying system, when the video encoder 402 encodes the same video sequence 412 at two different times, it is almost impossible to reproduce exactly the same delay and smoothness results in a real-world network. Therefore, it is not feasible to evaluate and compare the performance of different video codecs and encoding tools in a real-world network. Video encoding tools include various technologies (e.g., inter-frame prediction, intra-frame prediction, etc.). Video codecs (such as H.264, H.265, AV1, and VP9) use video tools to encode and decode video data. As used herein, video codecs and video encoding tools are collectively referred to as video codecs. Therefore, it is necessary to find a method to reliably evaluate the performance of codecs in RTC use cases. The present invention discloses a network model that can be used to reproduce the same results under specific network conditions. As part of a new RTC video codec performance evaluation system, the network model can be used to evaluate the performance of new encoding tools relative to the original codec from the perspective of RTC quality.
[0027] In this paper, r represents the packet loss rate, b represents the network bandwidth limitation for transmitting media video streams on the network (including the effective video bit rate and redundant bit rate), and d represents the E2E network delay (also known as one-way delay (OWD)).
[0028] Because packet loss mitigation schemes like PR and FEC affect the effective values of r, b, and d, another parameter, s, is introduced to represent the effectiveness of these schemes. For simplicity, this article uses PR as an example. Parameter s is defined as the maximum number of times a packet can be retransmitted after its initial transmission. Typically, a packet is retransmitted because it was previously considered lost at the receiver and a notification (e.g., a negative acknowledgment (NACK)) is sent to the transmitter. When a lost packet is retransmitted multiple times, the end-to-end delay d of the video unit containing the lost packet increases, the actual packet loss rate r of the lost packet decreases, and the effective video bitrate b decreases. s = 0 indicates that PR is not used (i.e., a packet is sent only once). A positive value of s indicates a reduction in the actual packet loss rate. Retransmitting lost packets ensures that they are eventually received by the decoder (i.e., the receiver) in networks where r < 1. In practice, s typically has an upper limit, as a large value can result in excessive delays or waste of network bandwidth. In one embodiment of the present invention, this upper limit is set to 4, i.e., s = 4.
[0029] The present invention provides a new system and method for evaluating the performance of a video codec. Figure 5 The new test system is further described. Figure 5 , Figure 5 A schematic block diagram and data flow diagram of a novel system and method for evaluating the performance of a video codec in an RTC use case is shown, generally indicated by the reference numeral 500. In particular, the novel video codec performance evaluation system 500 includes a video codec performance evaluation network model 502 and a video codec quality analyzer 504. In one embodiment, both the network model 502 and the video codec quality analyzer 504 are computer software coded using one or more computer programming languages (e.g., C, C++, C#, Java, etc.).
[0030] Reference numerals 510 and 512 represent the video encoder and video decoder being tested, respectively. Reference numeral 522 represents an input video sequence (also referred to herein as test video data and test video sequence). Video sequence 522 comprises a set of video frames (or simply frames). Video encoder 510 encodes video sequence 522 into a coded bitstream 524. The coded bitstream 524 and its corresponding transmission timestamp are input to network model 502. Video codec performance evaluation network model 502 outputs a received bitstream 528 and its corresponding reception timestamp as input to video decoder 512. Video decoder 512 outputs decoded video 532. In a further embodiment, video decoder 512 outputs feedback information 534 and sends the feedback information 534 to video encoder 510.
[0031] The video codec performance evaluation network model 502 and the video codec quality analyzer 504 are used to evaluate the performance of a coding tool (e.g., a video codec, including a video encoder) relative to a reference video codec (e.g., a benchmark video codec) from the perspective of RTC quality. When the video codec performance evaluation network model 502 and the video codec quality analyzer 504 are run to test the performance of the video encoder 510, they are run on an electronic device (e.g., electronic devices 102, 104, 106, and 108).
[0032] The video codec performance evaluation network model 502 considers four parameters: s, r, d, and b. For codec evaluation purposes, the parameter b, which indicates network bandwidth, can typically be ignored because the video sequence 522 (also referred to herein as a test video sequence, video, or video sequence) is typically encoded at several different bit rates when evaluating a video codec. The video codec performance evaluation network model 502 selects encoded video sequences at specific bit rates (e.g., 800 kbps, 400 kbps, 200 kbps, and 100 kbps) that match the target network bandwidth. To further simplify testing, the network model 502 uses a limited set of quantization parameter (QP) values to encode the video sequence in a constant QP (CQP) mode. For example, when encoding a video sequence that complies with the H.264 video coding standard, the QP value ranges from 0 to 51. In one embodiment, the network model 502 sets the QP values to 22, 27, 32, and 37 when determining the RTC quality index of the video encoder 510. In another embodiment, when encoding a video sequence compliant with the AO Media Video 1 (AV1) video coding standard, the QP value is between 0 and 255. This method can roughly generate a bitstream with a relatively stable bit rate for the video sequence 522 containing similar content.
[0033] Significant variations in the size of the encoded frames can affect the latency and smoothness of the received video. Larger encoded frames are more likely to be packed into more packets, resulting in longer network transmission times and a higher probability of packet loss in lossy networks. Therefore, ideally, the video encoder in the RTC would be able to generate a bitstream with a constant bitrate. To mitigate the differences between CQP and constant bitrate (CBR) encoding modes, the disclosed systems and methods use video sequences whose video content does not vary significantly to evaluate codecs. For example, when video sequence 522 is captured by a camera moving steadily across a scene, the successive frames of the video change smoothly. In contrast, when the video sequence is a film clip containing dramatic changes in scenery, abrupt changes may occur between two consecutive video frames. In one embodiment, the size variation between two consecutive video frames in video sequence 522 is less than a predetermined percentage, such as 10% or 25%. In this case, the video sequence is considered to have a sudden change less than a predetermined video sudden change threshold.
[0034] When the network model 502 takes the parameters (r, d, s) into account, the network model 502 encodes the video unit 522 (the video sequence 522 includes multiple frames, each of which can be regarded as a video unit) (for example, a slice, a strip, or a frame) and packages it into N (N is a positive integer) data packets. The network model 502 creates a set of lookup tables whose index is N and whose values are f(r, s, N) and g(r, d, s, N). Alternatively, they can be pre-created by different computer software applications and referenced by the network model 502. In this case, it is also considered to be created by the network model 502 in this article. f(r, s, N) is the arrival probability of the video unit 522 encoded into N (N is a positive integer) received data packets, also referred to as the f value (or f value) in this article. g(r, d, N, s) is the expected arrival delay of the video unit 522, also referred to as the g value (or g value) in this article. To avoid precision fluctuations between different platforms (e.g., iPhone, iPad, Android smartphone, Android tablet, desktop computer running Windows operating system, laptop computer running Windows operating system), f(r, s, N) is scaled by a factor of 10,000 and rounded down to the nearest integer, and g(r, d, s, N) is also rounded down to the nearest integer. The scaling and rounding operations are performed by the network model 502. Alternatively, they can be created by other computer software applications. In this case, it is also considered to be created by the network model 502. The lookup table (LUT) is also referred to as the fg lookup table in this article. The parameters (r, d, s) are taken into account when the network model 502 evaluates the performance of the video codec, and the parameters (r, d, s) represent a network state.
[0035] The network model 502 further derives f(r, s, N) and g(r, d, s, N) by the following formulas:
[0036] f(r,s,N)=floor(10000(1-r s+1 ) N )
[0037]
[0038] Note that d does not affect f(r, s, N) and is proportional to g(r, d, s, N).
[0039] If the coding decision does not take feedback information 534 into account, d can be simply set to a fixed value, such as OWD. For example, when d = 100 ms is set, the network model 502 calculates g(r, 100, s, N) as follows:
[0040]
[0041] For different values of d, the network model 502 scales the above expression accordingly.
[0042] If the encoding decision takes feedback information 534 into account, the delay d will affect the estimated feedback time of the video encoder 510 (ie, the transmitter). In this case, different values of d need to be tested to evaluate the performance of the video encoder 510.
[0043] Therefore, for a given parameter triplet (r, d, s), the values of the functions f(r, s, N) and g(r, d, s, N) are calculated for different values of N, and the corresponding lookup tables are obtained by the network model 502. For example, for r = 0.25, d = 300 ms and s = 4, when N = 2, f(r, s, N) = 9980 and g(r, d, s, N) = 654.
[0044] The present invention defines a set of RTC test conditions for evaluating video codec quality. These test conditions are implemented using a set of lookup tables derived from the above formulas. In one embodiment, the network model 502 creates the lookup tables. Alternatively, the lookup tables can be constructed by different computer software applications or obtained from other programs. In this case, the lookup tables are also considered to be created or derived by the network model 502.
[0045] During the performance evaluation of the video encoder 510 and video decoder 512, a video frame may be encoded into one or more self-decoding units. A self-decoding unit can be a slice, slice, or even a frame from the video sequence input. Each coding unit is packaged into N data packets by the video codec performance evaluation network model 502. Under network conditions (r, d, s), the network model 502 finds the corresponding values of f(r, s, N) and g(r, d, s, N) in a lookup table. First, for each video unit, the network model 502 generates a pseudo-uniform random integer p in the range of 0 to 10,000. Alternatively, the network model 502 generates it. In this case, the pseudo-uniform random integer p is also considered to be generated by the network model 502. The network model 502 then compares it with f(r, s, N). If p ≤ f(r, s, N), all data packets for that video unit are considered to have been received at the decoder. The network model 502 then outputs the data packets as a received bitstream 528 to the video decoder 512. If p > f(r, s, N), then one or more packets of the video unit are considered lost in transmission. In this case, the network model 502 does not provide the packets of the video unit to the video decoder 512 as the received bitstream 528. In one embodiment, if any packet is lost or not received, the unit corresponding to the packet is considered lost.
[0046] To obtain consistent results across repeated experiments within the same test, the random number generator must use a fixed seed. To approximate the distribution of f(r, s, N), network model 502 performs several experiments, each using a different seed. Network model 502 then averages the results to produce the final result. To control overall test time, network model 502 selects a limited set of seeds, for example, three seeds (0, 1, 2), in experiments evaluating the performance of video encoder 510.
[0047] After the network model 502 determines that all units of a video frame have been received, the video decoder 512 checks whether its reference frame has also been received. When a reference frame is considered lost, it indicates that at least one of the data packets containing it has been lost. If any of the reference frames of the video frame is lost, the video decoder 512 marks the video frame as undecodable. Otherwise, the video frame is marked as decodable and its frame number and reception timestamp are stored or otherwise tracked. The timestamp is the most recently received timestamp across all packets for the video frame. After the test case is completed, all data collected at the encoder and decoder ends will be used to calculate the RTC quality metric to evaluate the performance of the video codec.
[0048] Figure 7 、 Figure 8 and Figure 9 The flowchart shown further illustrates the process of determining the RTC quality index of the video codec by the video codec performance evaluation system. Figure 7 , Figure 7 The flowchart shows a process of determining the RTC quality index of a video codec by a video codec performance evaluation system, which is represented by label 700. The process is formed for a network described by three parameters r, d and s. The video input to the process 700 is a video sequence 522. In step 702, the network model 502 creates a set of fg lookup tables based on r, d and s (for example, Table 1 and Table 2 shown below). In step 704, the network model 502 determines a finite set of quantization parameter values. In step 706, the network model 502 determines a finite set of seeds for generating random numbers. In step 708, the network model 502 determines a frame decoding data item set for each value in a set of quantization parameter values and each value in a set of seeds, thereby forming a first list including multiple frame decoding data item sets. For example, if there are 4 values in a set of quantization parameter values and 3 values in a set of seeds, the above list includes 12 frame decoding data item sets. In one embodiment, the frame decoding data item is a triplet Where n represents the index of the frame in the video sequence; Indicates the reception timestamp of frame n; 1 indicates that the frame has been received and can be decoded. In another example, A 0 in indicates that frame n is not decodable.
[0049] In step 710, the video codec quality analyzer 504 determines a decodable frame ratio based on each frame decoded data item set in the first list of multiple frame decoded data item sets. Thus, a second list of multiple decodable frame ratios is generated in step 710. In one embodiment, the decodable frame ratio is obtained using the following formula:
[0050]
[0051] N E represents the number of video frames encoded by the video encoder 510 from the video sequence 522, N D Represents the number of video frames successfully decoded by the video decoder 512 from the received frames of the video sequence 522 .
[0052] In step 712, the video codec quality analyzer 504 determines a plurality of decodable frame average delay values based on each frame decoded data item set in the first list including the plurality of frame decoded data item sets. Thus, a third list including a plurality of decodable frame average delay values is generated in step 712. In step 714, the video codec quality analyzer 504 determines a plurality of decodable frame maximum delay values based on each frame decoded data item set in the first list including the plurality of frame decoded data item sets. Thus, a fifth list including a plurality of decodable frame maximum delay values is generated in step 714. In one embodiment, the decodable frame average delay value and the decodable frame maximum delay value can be respectively obtained using the following formulas:
[0053]
[0054]
[0055] in, Is the decodable frame k i The receiving timestamp, Is the decodable frame k i The sending timestamp of the decodable frame is {k i |0≤k i ≤N E -1}, where i = 1, 2, ..., N D .
[0056] In step 716, the video codec quality analyzer 504 determines the video smoothness based on each frame decoded data item set in the first list of multiple frame decoded data item sets. As a result, a fourth list including multiple video smoothness is generated in step 716. In one embodiment, the video smoothness can be calculated using the following formula:
[0057]
[0058] In the above formula, Represents frame k i +1's expected time of receipt (also referred to herein as expected time of arrival). and The difference is the frame k i+1 The actual received reception timestamp and the expected reception timestamp Ideally, and The difference between is 0. However, due to frame loss or network delay, the difference may not be zero. Please note that there may also be This means that the receiver is in frame k i Frame k is received before the expected reception time is over i+1 This is also considered stuttering because it displays frames faster than expected. and It is used when the first or last frame of the video is lost.
[0059] In step 718, the video codec quality analyzer 504 determines a final decodable frame ratio based on the second list of decodable frame ratios. In one embodiment, the final decodable frame ratio is the average of the multiple decodable frame ratios in the second list of decodable frame ratios. In step 720, the video codec quality analyzer 504 determines a final decodable frame average delay value based on the third list of decodable frame average delay values. In one embodiment, the final decodable frame average delay value is the average of the multiple decodable frame average delay values in the third list of decodable frame average delay values. In step 722, the video codec quality analyzer 504 determines a final decodable frame maximum delay value based on the fifth list of decodable frame maximum delay values. In one embodiment, the final decodable frame maximum delay value is the average of the multiple decodable frame maximum delay values in the fifth list of decodable frame maximum delay values. In step 724, the video codec quality analyzer 504 determines a final video smoothness level based on the fourth list of multiple video smoothness levels. In one embodiment, the final video smoothness is an average of multiple video smoothness in a fourth list including multiple video smoothness. The real-time communication quality indicators of the video codec include a final decodable frame ratio, a final decodable frame average delay value, a final decodable frame maximum delay value, and a final video smoothness.
[0060] n represents the index of the current frame in the video sequence 522 , which is a positive integer starting from 0. fps represents the number of frames per second in the video sequence 522 . represents the reception timestamp (or time) of frame n, Represents the transmit timestamp of frame n. As used herein, for any particular frame, the difference between its receive timestamp and the timestamp of successful decoding is negligible and can be considered zero for purposes of valid performance evaluation of a video codec. Thus, the receive timestamp and the timestamp of successful decoding of that frame are considered the same. Similarly, as used herein, for any particular video frame, the difference between its transmit timestamp and the timestamp of successful encoding is also negligible and can be considered zero for purposes of valid performance evaluation of a video codec. Thus, the transmit timestamp and the timestamp of successful encoding of a video frame are considered the same.
[0061] Figure 8 The process of determining the frame decoding data item set in step 708 is further explained. Figure 8 The flowchart of FIG. 1 shows a process of determining a frame decoding data item set by a video codec performance evaluation system, wherein the data item set is a data item set corresponding to a specific QP value in a set of QP values and a specific seed in a set of seeds. The process is represented by reference numeral 800. In step 802, the network model 502 determines a transmission timestamp for each frame in the video sequence 522. In one embodiment, the network model 502 obtains the transmission timestamp of frame n from the video sequence 522. Where fps is the number of frames per second in the video sequence 522. In step 804, the network model 502 obtains (or in some way receives) the coded bit stream of the frame n. The video encoder 510 encodes the frame n into a coded bit stream. In step 806, the network model 502 divides the coded bit stream of the frame n into a plurality of (M) units. In step 808, the network model 502 packs each unit m of the M units into a certain number (N m ) data packets.
[0062] In step 810, the network model 502 determines the arrival probability value f(r, s, N m ) and the expected arrival delay g(r, d, s, N m ). In step 812, the network model 502 generates a random number p for each unit m. In one embodiment, the random number p is a pseudo-uniform random integer p in the range of 0 to 10,000 as described above. In step 814, the network model 502 compares the random number p with the arrival probability value f(r, s, N m) is compared to determine whether unit m has been received or lost. In one embodiment, if p≤f(r, s, N), then the video decoder 512 is deemed to have received all packets of unit m and the unit m. Otherwise, (p>f(r, s, N)), then one or more packets of the unit m are deemed to have been lost during transmission. Unit m is also deemed to have been lost. In step 816, the network model 502 determines the reception timestamp of the frame n. In one embodiment, the reception timestamp of frame n is derived using the following formula:
[0063]
[0064] In step 818, if all units of frame n have been received and frame n is decodable, the network model 502 records a frame decoding data item that indicates the reception timestamp of frame n and indicates that frame n is decodable. The video decoder 512 determines whether frame n is decodable. In one embodiment, the frame decoding data item is a triplet: In a further embodiment, the video decoder 512 determines some additional video quality indicators, such as PSNR, SSIM, and VMAF. PSNR can be calculated as SSIM can be calculated as In step 818, the network model 502 records a frame decoding data item, wherein the frame decoding data item includes the video quality measurement indicators PSNR, SSIM and VMAF. For example, in step 818, the frame decoding data item is recorded These frame decoding data items are used to determine the performance indicators of the video codec.
[0065] In step 820, if all elements of frame n are not received or frame n is not decodable, the network model 502 records a frame decoding data item that indicates the reception timestamp of frame n and indicates that frame n is not decodable. In one embodiment, the frame decoding data item is a triplet:
[0066] When the decoder side provides feedback information about the status of the received video frames to the encoder side, the network model 502 forwards the feedback information to the video encoder 510. In this case, Figure 9 The process of determining the frame decoding data item set in step 708 is shown, and the process is represented by reference numeral 900. In step 952, taking a specific frame k as an example, the network model 502 obtains or otherwise receives feedback information. The video frame decoding feedback information received in step 952 indicates the reception timestamp of the specific frame k. And whether the particular frame k is decodable. In step 954, the network model 502 determines whether to forward the video frame decoding feedback information to the video encoder 510. In one embodiment, if The video frame decoding feedback information should be forwarded to the video encoder 510 , otherwise it should not be forwarded. In step 956 , if it is determined in step 954 that forwarding is required, the network model 502 forwards the feedback information to the video encoder 510 .
[0067] It should be noted that the decoding process is performed on each unit of the frame. A frame of video consists of one or more units. The unit currently being decoded is also referred to as the current unit in this article. If a unit is decodable, it can continue to be used for encoding subsequent frames.
[0068] When the network model 502 evaluates the performance of the video codec, it determines the RTC quality index of the video transmission. When executing the test case, the maximum number of retransmissions is set to a finite number, such as 4 (s=4). In order to effectively evaluate the performance of the video encoder 510, the network model 502 sets the packet loss rate to predetermined values, such as 0%, 25% and 50%, which can represent typical network conditions. In addition, the network model 502 also sets the E2E delay value to a predetermined value, such as 50ms and 300ms, to represent the E2E delay value of the network transmission, for example, within a country or across the ocean. The specific numerical values cited in this paragraph are for illustrative purposes only. In different implementation plans, their values may vary.
[0069] When the video encoder 510 does not consider feedback information 534, d does not actually directly affect the final result. Furthermore, when r = 0%, the evaluation method of the present invention is the same as the traditional CTC case. Therefore, there are a total of six test cases in the no-feedback scenario, with packet loss rates of 25% and 50% and three different seeds. When the video encoder 510 considers feedback information 534, the network model 502 tests all combinations. In this case, six test cases are run. Because three different seeds are selected to evaluate the performance of the video codec, the network model 502 will execute 6 * 3 = 18 test cases.
[0070] In the test case where the video encoder 510 does not consider the feedback information 534, the confirmation message indicating whether the video decoder 512 successfully decoded all video frames is not sent to the video encoder 510. Alternatively, the confirmation message is sent but not considered. For the test case without feedback, the video encoder 510 periodically sends instantaneous decoding refresh frames (i.e., IDR frames) to the video decoder 512. The interval between two IDR frames is set according to the usage scenario, and is generally set to two seconds by default.
[0071] For the test case with feedback, when the video decoder 512 successfully decodes a video frame, the video decoder 512 sends a feedback message 534 to the video encoder 510, indicating that the video decoder 512 has successfully decoded the frame. In this case, the video encoder 512 can choose the encoding frame type more flexibly. If one or more frames are lost during transmission, it can send an IDR frame or a frame that refers to the successfully decoded frame as a reference frame. However, more reference frames will lead to stronger data dependency, which is not desirable for lossy network transmission. In order to limit this additional data dependency, only one reference frame is used in one embodiment of the present invention. In addition, in the case of no feedback mode, in one embodiment, the most recent reference frame is used to obtain the best compression performance. Because low latency is extremely important in RTC, the low latency encoding mode of the present invention is a more desirable approach. It should be noted that in the case of backward reference, more delays are introduced, such as Figure 6 As shown in the figure. If only one backward reference frame is used, the entire real-time communication system will have a delay of two frames. If more backward reference frames are used, a longer delay will be introduced. Therefore, the use of the backward reference frame type is not recommended in the RTC case.
[0072] The following section further describes how to determine the RTC quality indicator. The video codec quality analyzer 504 checks which frames are fully received and decoded at the decoding end based on the reference relationship. The video codec quality analyzer 504 further calculates the overall delay, thereby calculating the decodable video frame ratio and video freeze time. Video freeze time measures the smoothness of the received video.
[0073] A set of fg lookup tables can be generated by a script (e.g., a Python script) and referenced by the network model 502. As described herein, the fg lookup tables are considered to be created by the network model 502. For example, Tables 1 and 2 list six test cases for six different combinations of two network delay values (d = 50ms or d = 300ms) and three packet loss rate values (r = 50.0%, r = 25.0%, or r = 0.0%) for video units ranging in size from N = 1 to N = 256 packets, with a maximum number of retransmissions set to four (i.e., s = 4). In Tables 1-2, the retransmission parameter s is set to 4. Although the s parameter is not shown in Tables 1-2, it is actually incorporated into Tables 1-2. Therefore, Tables 1-2 can be considered to include the parameter s.
[0074] Table 1 – d = 300ms
[0075]
[0076]
[0077]
[0078]
[0079]
[0080]
[0081]
[0082] Table 2 – d = 50ms
[0083]
[0084]
[0085]
[0086]
[0087]
[0088]
[0089]
[0090] It is obvious from the above description that the present invention is susceptible to numerous other modifications and variations. Therefore, it is to be noted that within the scope of the appended claims, the present invention may be practiced otherwise than as specifically described.
[0091] The description of the present invention is intended to better illustrate and explain, and is not intended to be exclusive or to limit the present invention to the specific forms described above. The above description is intended to better explain the principles of the present invention and the practical application of these principles, so that relevant technical personnel in the field can better utilize the present invention to implement various embodiments and make various modifications in the intended appropriate uses. It should be recognized that the words "a" or "an" herein include both singular and plural forms. Conversely, where appropriate, references to multiple elements herein should also include their singular forms.
[0092] The scope of the present invention is not limited to the contents of the above description, but is determined by the scope of the claims. In addition, although the claims may be narrow, it should be recognized that the scope of the present invention is much broader than the scope of the claims. We will propose broader claims in one or more applications claiming priority from this application. If the contents disclosed in the above description and drawings are not within the scope of the claims, the contents of the invention are not disclosed to the public, and we reserve the right to file one or more patent applications in the future for the contents of the invention.
Claims
1. A method for determining performance indicators of a video codec in real-time communication, wherein: The method is performed by a video codec performance evaluation system, and the method includes: (1) Create a set of fg lookup tables, where each fg lookup table includes a packet loss rate value, a network delay value, a retransmission value, an arrival probability value, and an expected arrival delay value; (2) determining a set of quantization parameter values; (3) determining a set of seeds for generating random numbers; (4) determining a frame decoded data item set of a video sequence for each value in the set of quantization parameter values and each value in the set of seeds, thereby forming a first list including a plurality of frame decoded data item sets, the video sequence including a set of frames; (5) according to each frame decoded data item set in the first list including a plurality of the frame decoded data item sets, a. determining a decodable frame ratio, thereby forming a second list comprising a plurality of decodable frame ratios; b. Determine a decodable frame average delay value, thereby forming a third list comprising a plurality of decodable frame average delay values; and c. Determine a video smoothness, thereby forming a fourth list comprising a plurality of video smoothnesses; and (6) determining a final decodable frame ratio according to the second list including a plurality of decodable frame ratios; (7) determining a final decodable frame average delay value according to the third list including a plurality of decodable frame average delay values; and (8) Determine a final video fluency based on the fourth list including multiple video fluencies.
2. The method according to claim 1, wherein Determining a set of frame decoding data items includes, for each frame in the set of frames: (1) Determine a sending timestamp; (2) obtaining a coded bit stream of the frame; (3) dividing the coded bit stream into a group of units; (4) Packing each unit into a set of data packets; (5) for each unit in the set of units, a. Determine an arrival probability value according to the set of fg lookup tables; b. determining an expected arrival delay value according to the set of fg lookup tables; c. Generate a random number; and d. Comparing the random number and the arrival probability value to determine whether the unit has been received or lost; (6) Determine a receiving timestamp; (7) if all units of the frame have been received and the frame is decodable, recording a frame decoding data item, wherein the frame decoding data item indicates the reception timestamp of the frame and indicates that the frame is decodable; and (8) If all units of the frame are not received or the frame is undecodable, record a frame decoding data item, wherein the frame decoding data item indicates the reception timestamp of the frame and indicates that the frame is undecodable.
3. The method according to claim 2, wherein: The method further comprises: (1) determining a decodable frame maximum delay value according to each frame decoded data item set in the first list including a plurality of frame decoded data item sets, thereby forming a fifth list including a plurality of decodable frame maximum delay values; and (2) Determine a final decodable frame maximum delay value according to the fifth list including multiple decodable frame maximum delay values.
4. The method according to claim 2, wherein: The plurality of arrival probability values and the plurality of expected arrival delay values in the set of fg lookup tables are respectively calculated according to the following formulas: f(r,s,N)=floor(10000(1-r s+1 ) N ) Wherein r represents the preset packet loss rate, d represents the preset network delay value; s represents the preset number of retransmissions, and N represents the number of data packets in the unit.
5. The method according to claim 4, wherein The method further comprises: (1) Each frame decoding data item in the frame decoding data item set indicates the reception timestamp and indicates whether the frame is decodable; (2) the decodable frame ratio is a ratio of the number of frames in the set of frames to the number of decodable frames in the set of frames; and (3) The average decodable frame delay value is calculated by the following formula: in, Represents frame k i The receiving timestamp; Represents frame k i Sending timestamp; N D Indicates the number of decodable frames in a group of frames; (4) The video fluency is calculated by the following formula: Among them, fps means frames per second, N E Indicates the number of decodable frames in the set of frames; (5) The receiving timestamp is obtained by the following equation: Here, n represents the index of the frame in the set of frames.
6. The method according to claim 5, wherein The method further comprises: (1) Obtaining video frame decoding feedback information of a processed frame; (2) determining whether to forward the video frame decoding feedback information to the video encoder that encodes the frame; and (3) Forwarding the video frame decoding feedback information to the video encoder.
7. The method according to claim 6, wherein: when , forwarding the video frame decoding feedback information to the video encoder.
8. The method according to claim 2, wherein: The method further comprises: (1) Obtaining video frame decoding feedback information of a processed frame; (2) determining whether to forward the video frame decoding feedback information to the video encoder that encodes the frame; as well as (3) Forwarding the video frame decoding feedback information to the video encoder.
9. The method according to claim 5, wherein: The method further comprises: (1) determining a decodable frame maximum delay value according to each frame decoded data item set in the first list including a plurality of frame decoded data item sets, thereby forming a fifth list including a plurality of decodable frame maximum delay values; as well as (2) Determine a final decodable frame maximum delay value according to the fifth list including multiple decodable frame maximum delay values.
10. The method according to claim 9, wherein: The maximum delay value of the decodable frame is calculated by the following formula:
11. The method according to claim 5, wherein: The final decodable frame ratio is an average of a plurality of decodable frame ratios in the second list of the plurality of decodable frame ratios.
12. The method according to claim 5, wherein: The final decodable frame average delay value is an average value of a plurality of decodable frame average delay values in the third list comprising a plurality of decodable frame average delay values.
13. The method according to claim 5, wherein: The final video smoothness is an average of multiple video smoothnesses in the fourth list including multiple video smoothnesses.
14. The method according to claim 2, wherein: The final decodable frame ratio is an average of a plurality of decodable frame ratios in the second list of the plurality of decodable frame ratios.
15. The method according to claim 2, wherein: The final decodable frame average delay value is an average value of a plurality of decodable frame average delay values in the third list comprising a plurality of decodable frame average delay values.
16. The method according to claim 2, wherein: The final video smoothness is an average of multiple video smoothnesses in the fourth list including multiple video smoothnesses.
17. The method of claim 8, wherein The method further comprises: (1) determining a decodable frame maximum delay value according to each frame decoded data item set in the first list including a plurality of frame decoded data item sets, thereby forming a fifth list including a plurality of decodable frame maximum delay values; and (2) Determine a final decodable frame maximum delay value according to the fifth list including multiple decodable frame maximum delay values.
Citation Information
Patent Citations
Anti-error code method and system of multi-channel video conference system based on H264
CN105681342A
Bandwidth Adjustment For Real-time Video Transmission
CN107438187A