Method and apparatus for codec performance comparison
By improving the codec performance evaluation tool, fitting the codec's data points using nonlinear relationship fitting curves, solves the problem that existing tools cannot handle non-monotonic incremental data sets and require that input data must overlap, achieving a more flexible and accurate codec performance evaluation.
Patent Information
- Application Number
- CN202480004646.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-29
- Filing Date
- 2024-05-22
- Publication Date
- 2025-06-13
AI Technical Summary
Existing codec performance evaluation tools cannot handle non-monotonic incremental data sets and require that input data must be overlapped, limiting its versatility and flexibility.
The test codec performance is evaluated and BD-PSNR and BD-rate are calculated by encoding the reference video separately for each encoding parameter using an anchor and a test codec.
The evaluation of non-monotonic incremental data sets is achieved, supporting any given interval, even without overlap, significantly improving the flexibility and accuracy of the codec evaluation tool.
Smart Images

Figure CN120153398A_ABST
Abstract
Description
Incorporation by reference
[0001] This application claims the benefit of priority to U.S. Non - Provisional Application No. 18 / 621,713, filed on March 29, 2024, which claims the benefit of priority to U.S. Provisional Application No. 63 / 527,060, filed on July 16, 2023. The above - mentioned applications are hereby incorporated by reference in their entirety. Technical Field
[0002] The present disclosure generally relates to video coding techniques. More specifically, the disclosed techniques relate to enhancements to codec performance measurement and evaluation. Background Art
[0003] In recent decades, driven by the increasing interest in both live video content and video - on - demand content on various applications and platforms, video streaming applications have gained significant popularity. Thus, video streaming now represents a major source of Internet traffic and is expected to further surge due to the proliferation of video - centric applications and the advancement of video device capabilities. There is an urgent need to develop efficient video compression and delivery algorithms to effectively handle this expected growth. Summary of the Invention
[0004] The present disclosure describes various embodiments of methods, devices, and non - transitory computer - readable storage media for enhancing codec performance measurement and evaluation.
[0005] According to one aspect, embodiments of the present disclosure provide a method for processing video data and evaluating codec performance. The method includes: obtaining a first plurality of anchor data points generated based on an anchor video bitstream, where: the anchor video bitstream is encoded by an anchor video codec based on a reference video and corresponding coding parameters selected from a first plurality of coding parameters; each anchor data point represents the anchor codec performance using the corresponding coding parameter, and the anchor data point is formatted as a tuple that includes (i) the bitrate or a change in the bitrate and (ii) a quality measurement result; obtaining a second plurality of test data points generated based on a test video bitstream, where: the test video bitstream is encoded by a test video codec based on a reference video and corresponding coding parameters selected from a second plurality of coding parameters; each test data point represents the test codec performance using the corresponding coding parameter, and the test data point is formatted as a tuple; fitting an anchor curve to the first plurality of anchor data points, the anchor curve being based on an anchor polynomial, where the anchor polynomial is monotonic within a first axis range; fitting a test curve to the second plurality of test data points, the test curve being based on a test polynomial, where the test polynomial is monotonic within the first axis range; and evaluating the test codec performance based on the anchor curve and the test curve to obtain an evaluation result.
[0006] According to another aspect, an embodiment of the present disclosure provides an apparatus for evaluating a video codec. The apparatus includes: a memory storing instructions; and a processor communicatively coupled to the memory. When the processor executes the instructions, the processor is configured to cause the apparatus to perform the above method for processing video data and evaluating codec performance.
[0007] In another aspect, an embodiment of the present disclosure provides a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to perform the above method for processing video data and evaluating codec performance.
[0008] The above aspects and other aspects and their implementations are described in more detail in the drawings, the description, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0010] Figure 1 A schematic illustration showing a simplified block diagram of a computer / communication system (100) according to an example embodiment;
[0011] Figure 2 A schematic illustration showing a simplified block diagram of a computer / communication system (200) according to an example embodiment;
[0012] Figure 3 A schematic illustration showing a simplified block diagram of a video decoder according to an example embodiment;
[0013] Figure 4 A schematic illustration showing a simplified block diagram of a video encoder according to an example embodiment;
[0014] Figure 5 A block diagram of a video encoder according to another example embodiment;
[0015] Figure 6 A block diagram of a video decoder according to another example embodiment;
[0016] Figure 7A A schematic illustration showing an example data set used in Piecewise Cubic Hermit Interpolation (PCHIP);
[0017] Figure 7B A schematic illustration showing the average of rate-distortion (RD) curves considering two codecs Calculation of the Incremental Peak Signal-to-Noise Ratio ( Delta Peak Signal-to-Noise Ratio, BD-PSNR);
[0018] Figure 8A A dataset with data points following a monotonically increasing pattern is shown;
[0019] Figure 8B A dataset with data points not following a monotonically increasing pattern is shown;
[0020] Figure 8C Overlapping and non-overlapping datasets are shown.
[0021] Figure 8D Example datasets of two codecs and the coordinates of each data point in the datasets are shown.
[0022] Figure 9 An example codec evaluation method according to an example embodiment is shown.
[0023] Figure 10 An exemplary process for solving polynomial parameters using an iterative parameter optimization process is shown.
[0024] Figure 11 A schematic illustration of a computer system according to an example embodiment of the present disclosure is shown. Detailed Description of the Invention
[0025] The present invention will now be described in detail below with reference to the accompanying drawings, which form a part of the present invention and illustrate specific examples of the embodiments by way of illustration. However, note that the present invention can be implemented in various different forms, and thus, the subject matter covered or claimed is intended to be construed as not limited to any of the embodiments to be described below. Also note that the present invention can be implemented as a method, apparatus, component, or system. Therefore, the embodiments of the present invention can, for example, take the form of hardware, software, firmware, or any combination thereof.
[0026] Throughout the specification and claims, terms may have nuanced meanings that are presented or implied in contexts that go beyond the explicitly stated meaning. The phrases "in one embodiment" or "in some embodiments" as used herein do not necessarily refer to the same embodiment, and the phrases "in another embodiment" or "in other embodiments" as used herein do not necessarily refer to different embodiments. Similarly, the phrases "in one implementation" or "in some implementations" as used herein do not necessarily refer to the same implementation, and the phrases "in another implementation" or "in other implementations" as used herein do not necessarily refer to different implementations. For example, it is meant that the claimed subject matter includes combinations of the whole or parts of the exemplary embodiments / implementations.
[0027] Generally, terms can be understood at least in part according to their usage in context. For example, terms such as "and", "or", or "and / or" as used herein can have a variety of meanings, which depend at least in part on the context in which such terms are used. Generally, "or" (if used in connection with a list, e.g., A, B, or C) is intended to mean: A, B, and C, used herein in an inclusive sense; and A, B, or C, used herein in an exclusive sense. Additionally, depending at least in part on the context, the terms "one or more" or "at least one" as used herein can be used to describe any feature, structure, or property in a singular sense, or can be used to describe a combination of features, structures, and properties in a plural sense. Similarly, terms such as "a", "an", or "the" can likewise be understood to convey a singular usage or to convey a plural usage, depending at least in part on the context. Additionally, the terms "based on" or "determined by" can be understood to not necessarily be intended to convey an exclusive set of factors, and can alternatively allow for the existence of additional factors that are not necessarily explicitly described, again depending at least in part on the context.
[0028] Figure 1 FIG. shows a computer network or communication system 100. As Figure 1 shown, the terminal device can be implemented as a server, a personal computer, and a smart phone, but the applicability of the basic principles of the present disclosure is not limited thereto. Embodiments of the present disclosure can be implemented in a server or a desktop computer 120, a laptop computer 110, tablet computers 130 and 140, a media player, a wearable computer, dedicated video conferencing equipment, etc. The network (150) represents any number or type of network that transfers encoded video data between terminal devices, including, for example, wired (wired) and / or wireless communication networks. The communication network (150) can exchange data in circuit-switched channels, packet-switched channels, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0029] The codec performance measurement and evaluation method provided in the present disclosure can be applied to any electronic device, such as a server or a desktop computer 120, a laptop computer 110, a tablet computer 130, and 140, to compare the encoding performance of two encoders (also referred to as codecs), and further select the codec with better performance as the actual encoder for use. The corresponding decoder of the encoder can be used in any electronic device with decoding capabilities, such as a television terminal, a PC terminal, a mobile terminal, etc.
[0030] As an example of the application of the disclosed subject matter, Figure 2 illustrates the placement of a video encoder and a video decoder in a video streaming environment. The disclosed subject matter can be equally applicable to other video applications, including, for example, video conferencing, digital television broadcasting, gaming, virtual reality, storing compressed video on digital media including CDs, DVDs, memory sticks, etc. Specifically, the codec evaluation method described in the present disclosure can be utilized to measure, evaluate, and benchmark video encoders and / or decoders.
[0031] As Figure 2 shown, a video streaming system can include a video capture subsystem (213), and the video capture subsystem (213) can include a video source (201) such as a digital camera device for creating an uncompressed video picture or image stream (202). In the example, the video picture stream (202) includes samples recorded by the digital camera device of the video source (201). The video picture stream (202) is depicted as a thick line to emphasize the high data volume when compared with the encoded video data (204) (or encoded video bitstream), and the video picture stream (202) can be processed by an electronic device (220) including a video encoder (203) coupled to the video source (201). The video encoder (203) can include hardware, software, or a combination thereof to implement or enforce aspects of the disclosed subject matter described in more detail below. The encoded video data (204) (or encoded video bitstream (204)) is depicted as a thin line to emphasize the lower data volume when compared with the uncompressed video picture stream (202), and the encoded video data (204) can be stored on a streaming server (205) for future use or directly stored to a downstream video device (not shown). One or more streaming client subsystems such as Figure 2The client subsystems (206) and (208) therein can access the streaming server (205) to retrieve copies (207) and (209) of the encoded video data (204). The client subsystem (206) can include, for example, a video decoder (210) in an electronic device (230). The video decoder (210) decodes the incoming copy (207) of the encoded video data and creates an outgoing video picture stream (211) that is uncompressed and can be presented on a display (212) (e.g., a display screen) or other rendering device (not shown).
[0032] Figure 3 A block diagram of a video decoder (310) of an electronic device (330) according to any embodiment of the present disclosure below is shown. The electronic device (330) can include a receiver (331) (e.g., receiving circuitry). The video decoder (310) can be used instead of Figure 2 the video decoder (210) in the example of
[0033] As shown, in Figure 3 the receiver (331) can receive one or more encoded video sequences from a channel (301). To counter network jitter and / or handle playback timing, a buffer memory (315) can be provided between the receiver (331) and the entropy decoder / parser (320) (hereinafter referred to as "parser (320)"). The parser (320) can reconstruct symbols (321) from the encoded video sequences. The categories of these symbols include information for managing the operation of the video decoder (310) and potential information for controlling a rendering device such as a display (312) (e.g., a display screen). The parser (320) can parse / entropy decode the encoded video sequences. The parser (320) can extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the encoded video sequences. The subgroups can include Group of Picture (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. The parser (320) can also extract information such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc. from the encoded video sequences. The reconstruction of the symbols (321) can involve multiple different processing or functional units. The units involved and how these units are involved can be controlled by subgroup control information parsed by the parser (320) from the encoded video sequences.
[0034] The first unit may include a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) may receive quantized transform coefficients as (one or more) symbols (321) and control information from the parser (320), including information indicating which type of inverse transform to use, block size, quantization factor / parameter, quantization scaling matrix, etc. The scaler / inverse transform unit (351) may output a block including sample values that can be input into the aggregator (355).
[0035] In some cases, the output samples of the scaler / inverse transform (351) may belong to an intra-coded block, i.e., a block that does not use prediction information from a previously reconstructed picture but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) may use surrounding block information that has been reconstructed and stored in the current picture buffer (358) to generate a block having the same size and shape as the block being reconstructed. For example, the current picture buffer (358) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some implementations, the aggregator (355) may add, on a per-sample basis, the prediction information generated by the intra-picture prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351).
[0036] In other cases, the output samples of the scaler / inverse transform unit (351) may belong to an inter-coded and possibly motion-compensated block. In such a case, the motion-compensation prediction unit (353) may access the reference picture memory (357) based on the motion vectors to obtain samples for inter-picture prediction. After motion-compensating the obtained reference samples according to the symbols (321) belonging to the block, these samples may be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) (the output of unit 351 may be referred to as residual samples or a residual signal) to generate output sample information.
[0037] The output samples of the aggregator (355) may undergo various loop filtering techniques in the loop filter unit (356) including several types of loop filters. The output of the loop filter unit (356) may be a sample stream that may be output to the rendering device (312) and stored in the reference picture memory (357) for future inter-picture prediction.
[0038] Figure 4A block diagram of a video encoder (403) according to an example embodiment of the present disclosure is shown. The video encoder (403) may be included in an electronic device (420). The electronic device (420) may also include a transmitter (440) (e.g., transmission circuitry). The video encoder (403) may be used instead of Figure 4 the video encoder (403) in the example of
[0039] The video encoder (403) may receive video samples from a video source (401). According to some example embodiments, the video encoder (403) may encode and compress pictures of a source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by an application. Implementing an appropriate encoding speed constitutes a function of a controller (450). In some embodiments, the controller (450) may be functionally coupled to other functional units as described below and control the other functional units. Parameters set by the controller (450) may include parameters related to rate control (picture skipping, quantizer, λ value of rate-distortion optimization techniques...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc.
[0040] In some example embodiments, the video encoder (403) may be configured to operate in an encoding loop. The encoding loop may include a source encoder (430) and an (embedded) decoder (433) embedded in the video encoder (403). Even though the embedded decoder 433 processes the encoded video stream generated by the source encoder 430 without entropy encoding, the decoder (433) reconstructs the symbols in a manner similar to that which a (remote) decoder would create to create sample data (because in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols in the entropy encoding and the encoded video bitstream can be lossless). At this point, it can be observed that any decoder techniques that may exist only in the decoder other than parsing / entropy decoding may also necessarily exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter may sometimes focus on decoder operations related to the decoding part of the encoder. Since the encoder techniques are reciprocal to the decoder techniques described in detail, the description of the encoder techniques can be simplified. Only in certain areas or aspects, a more detailed description of the encoder is provided below.
[0041] During operation in some example implementations, the source encoder (430) may perform motion-compensated predictive coding that predictive-codes an input picture with reference to one or more previously encoded pictures designated as "reference pictures" from a video sequence.
[0042] The local video decoder (433) can decode the encoded video data of a picture that can be designated as a reference picture. The local video decoder (433) replicates the decoding process that can be performed by the video decoder on the reference picture, and can cause the reconstructed reference picture to be stored in the reference picture buffer (434). In this way, the video encoder (403) can locally store a copy of the reconstructed reference picture, which has the same content (no transmission errors) as the reconstructed reference picture to be obtained by the remote video decoder.
[0043] The predictor (435) can perform a prediction search for the encoding engine (432). That is, for a new picture to be encoded, the predictor (435) can search in the reference picture memory (434) for sample data (as a candidate reference pixel block) or some metadata that can be used as an appropriate prediction reference for the new picture, such as a reference picture motion vector, block shape, etc.
[0044] The controller (450) can manage the encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding the video data.
[0045] The outputs of all the above-mentioned functional units can undergo entropy encoding in the entropy encoder (445). The transmitter (440) can buffer the encoded video sequence(s) created by the entropy encoder (445) to prepare for transmission via the communication channel (460), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (440) can merge the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0046] The controller (450) can manage the operations of the video encoder (403). During encoding, the controller (450) can assign a certain encoded picture type to each encoded picture, which may affect the encoding technique that can be applied to the corresponding picture. For example, a picture can generally be assigned to one of the following picture types: an intra picture (I picture), a predictive picture (P picture), a bi-predictive picture (B picture), a multiple-predictive picture. As described in further detail below, the source picture can generally be spatially subdivided into a plurality of sample encoding blocks.
[0047] Figure 5A diagram showing a video encoder (503) according to another example embodiment of the present disclosure. The video encoder (503) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures, and encode the processing block into an encoded picture that is part of an encoded video sequence. The example video encoder (503) can be used instead of Figure 4 the video encoder (403) in the example.
[0048] For example, the video encoder (503) receives a matrix of sample values of the processing block. The video encoder (503) then uses, for example, rate-distortion optimization (RDO) to determine whether to best encode the processing block using an intra mode, an inter mode, or a bi-prediction mode.
[0049] In Figure 5 the example of, the video encoder (503) includes an inter encoder (530), an intra encoder (522), a residual calculator (523), a switch (526), a residual encoder (524), a general controller (521), and an entropy encoder (525) coupled together as shown in the example arrangement in Figure 5 .
[0050] The inter encoder (530) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a later picture in display order), generate inter prediction information (e.g., a description of redundancy information according to inter coding techniques, a motion vector, merge mode information), and calculate an inter prediction result (e.g., a prediction block) based on the inter prediction information using any suitable technique.
[0051] The intra encoder (522) is configured to receive samples of a current block (e.g., a processing block), compare the block with blocks that have already been encoded in the same picture, and generate transformed quantization coefficients, and in some cases also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques).
[0052] The general controller (521) can be configured to determine general control data, and control other components of the video encoder (503) based on the general control data to, for example, determine the prediction mode of a block, and provide a control signal to the switch (526) based on the prediction mode.
[0053] The residual calculator (523) can be configured to calculate the difference (residual data) between the received block and the prediction result of a block selected from the intra encoder (522) or the inter encoder (530). The residual encoder (524) can be configured to encode the residual data to generate transform coefficients. Then, the transform coefficients undergo quantization processing to obtain quantized transform coefficients. In various example embodiments, the video encoder (503) further includes a residual decoder (528). The residual decoder (528) is configured to perform inverse transformation and generate decoded residual data. The entropy encoder (525) can be configured to format the bitstream to include the encoded block and perform entropy encoding.
[0054] Figure 6 FIG. shows an example video decoder (610) according to another embodiment of the present disclosure. The video decoder (610) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In an example, the video decoder (610) can be used instead of Figure 4 the video decoder (410) in the example of
[0055] In Figure 6 the example of Figure 6 the video decoder (610) includes an entropy decoder (671), an inter decoder (680), a residual decoder (673), a reconstruction module (674), and an intra decoder (672) coupled together as shown in the example arrangement of
[0056] The entropy decoder (671) can be configured to reconstruct certain symbols representing the syntax elements that make up the encoded picture based on the encoded picture. The inter decoder (680) can be configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information. The intra decoder (672) can be configured to receive intra prediction information and generate a prediction result based on the intra prediction information. The residual decoder (673) can be configured to perform inverse quantization to extract the dequantized transform coefficients and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The reconstruction module (674) can be configured to combine the residual output by the residual decoder (673) with the prediction result (output by the inter prediction module or the intra prediction module, as appropriate) in the spatial domain to form a reconstructed block, and the reconstructed block forms part of the reconstructed picture as part of the reconstructed video.
[0057] Note that any suitable technology can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). In some example embodiments, one or more integrated circuits can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). In another embodiment, one or more processors executing software instructions can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610).
[0058] Online video streaming services (such as YouTube, Netflix, and Vimeo) constitute a significant portion of Internet traffic and their dominance will further increase. The surging popularity of video streaming services can be mainly attributed to the widespread availability of higher network bandwidth and improved compression efficiency. The proliferation of video playback devices such as smart phones, tablets, and smart TVs has facilitated this. As a result, users now expect to have seamless streaming capabilities on any device, anytime, anywhere. Video compression has become one of the most effective and crucial methods for reducing the size of media files, enabling faster and more efficient transmission and delivery over the network. Video encoding / decoding technology plays a vital role in optimizing bandwidth utilization while meeting the growing demand for high-quality video content. Considerable efforts are being made to develop more efficient video codecs.
[0059] One task in developing new video codecs or improving existing ones is to measure, evaluate, and compare codec performance. Video codec evaluation tools and algorithms evaluate the effectiveness and efficiency of codecs in encoding and decoding video and / or audio data. The evaluation tools / algorithms are crucial in determining the ability of a codec to deliver high-quality video output while minimizing the resulting file size and alleviating the demand for data transmission bandwidth. Well-designed and robust evaluation tools are essential for accurately quantifying the balance that a codec should achieve between maintaining video quality and optimizing the compression ratio. Obtaining a comprehensive and in-depth understanding of codec performance is crucial for ensuring the best multimedia experience across various platforms and applications. Factors such as streaming quality, file storage, and bandwidth requirements, as well as overall user satisfaction, are directly affected by the performance characteristics of the codec employed. Therefore, exhaustive codec performance testing plays a key role in delivering seamless and high-quality multimedia content to end users.
[0060] During the research and development process of multimedia codecs, Bjøntegaard Delta Peak Signal-to-Noise Ratio (BD-PSNR) and Bjøntegaard Delta Bitrate ( The delta rate (BD-rate) is usually used as a performance criterion for comparing the performance of different codecs. BD-PSNR and BD-BR are used to compare and measure the visual quality difference and bitrate difference respectively. The higher the BD-PSNR, the better the coding quality. On the other hand, the lower the BD-BR, the less bitrate is required for the coded bitstream with the same or similar quality.
[0061] Specifically, BD-PSNR is used to measure the average PSNR gain under the condition of the same bitrate, while BD-rate is used to measure the average bitrate saving under the condition of the same PSNR. Currently, the common algorithm is based on piecewise cubic Hermite interpolation (PCHIP), which uses cubic spline interpolation to construct a cubic function for every two adjacent points. The data between these two points is obtained by interpolation. Here, refer to Figure 7A to briefly describe this process. In this example, each group of data has four data points as input, and the four data points are represented as (x1, y1), (x2, y2), (x3, y3), (x4, y4). Take two points in sequence from left to right. The first step is to interpolate between the first point (x1, y1) and the second point (x2, y2). The interpolation polynomial is solved by the input values and the first derivative values of the two points, and then the points between x1 and x2 are solved by the interpolation polynomial. Repeat the above solving process in sequence to obtain the piecewise interpolation polynomial in the interval from x1 to x4 (i.e., from x1 to x2, from x2 to x3, and from x3 to x4 in sequence).
[0062] Refer to Figure 7B , Figure 7 is an example implementation of a codec evaluation tool based on PCHIP. Based on the data points (or operating points, shown as solid dots in Figure 7B ) corresponding to the reference codec (or called the anchor codec), a reference interpolation curve 702 is generated, and based on the data points corresponding to the test codec (marked as "X" in Figure 7B ), a test interpolation curve 704 is generated. The interpolation curve can also be called a fitting curve because it is designed to fit the data points. The codec evaluation tool then calculates the area of the region (i.e., the region 706 formed by the two fitting curves and two vertical dotted lines) through the interpolation curves, and then calculates BD-PSNR and BD-rate based on the difference in the region between the reference fitting curve and the test fitting curve.
[0063] Existing codec performance evaluation algorithms such as PCHIP, even though commonly used, have some limitations and defects.
[0064] First, existing codec evaluation tools require the input data (data points) to follow a monotonically increasing pattern. When the input data is non-monotonic, the codec evaluation tools will fail, resulting in unstable and unreliable evaluation results. Therefore, existing codec evaluation tools require the input data to strictly follow a monotonically increasing pattern, which makes them unable to handle input data that does not meet this requirement. This limitation affects the generality of codec evaluation tools because for some codecs, the input data may not always follow a monotonic pattern (such as a monotonically increasing pattern).
[0065] Figure 8A shows data points that meet the monotonicity requirement, while Figure 8B shows data points that do not meet the monotonicity requirement. Note that in Figure 8A and Figure 8B the x-axis represents the bit rate, for example, the bit rate in logarithmic format, and the y-axis represents the quality metric, for example, the mean average precision (mAP).
[0066] Second, for existing codec evaluation tools used to compare the performance of two codecs, the input data for each codec must have an overlap, and the performance evaluation interval is limited to the overlapping part of the two data point sets. As Figure 8C shown, the solid dot data set represents the performance measurement results of codec A, while the "X" data set represents the performance measurement results of codec B. Existing codec evaluation tools may only produce valid evaluation results within the indicated overlapping range (i.e., the overlapping bit rate range). In addition, for data sets generated from various codecs, the overlap condition may not always be met. For example, as Figure 8C shown, the "+" data set represents the performance measurement results of codec C, and the "+" data set does not overlap with the "X" data set. In this scenario, existing codec evaluation tools will not be able to evaluate and compare the performance of codec B and codec C.
[0067] In the present disclosure, various embodiments are described that are designed to address the above-mentioned deficiencies and limitations. These embodiments can support non-monotonically increasing data sets. They also support any given interval, even in the absence of an overlap. Therefore, codec evaluation tools according to these embodiments offer significant flexibility. These embodiments provide solutions for calculating BD-PSNR and BD-rate regardless of whether the input data follows a monotonic pattern. Therefore, codec evaluation tools according to these embodiments have considerable practical value and guiding significance.
[0068] Figure 9 shows an example codec evaluation method based on steps.
[0069] Input:
[0070] The input to the codec evaluation tool can include: a first video codec and a second video codec. Exemplarily, the first video codec can be a reference video codec, which is referred to as an anchor or anchor codec; the second video codec can be a codec to be evaluated, which is referred to as a test or test codec. The input can also include a set of encoding parameters [c1, c2, ……, c n , where n is an integer. These parameters can include, for example, spatial and temporal parameters (e.g., the resolution of the video), quantization parameters, encoding bitrate, color bit depth, chroma subsampling scheme, etc. The input can also include one or more reference videos S and a performance evaluation interval. The interval can include a range on the x-axis (first axis), which is represented as [xa, xb] (ranging from xa to xb, including both ends), or a range on the y-axis (second axis), which is represented as [ya, yb]. The x-axis can represent the bitrate or a transformation of the bitrate (e.g., the bitrate in logarithmic format). The y-axis can represent a quality metric, which can include but is not limited to peak signal-to-noise ratio (PSNR), mAP, multiple object tracking accuracy (MOTA), incremental MOTA ( Delta MOTA, BD-MOTA), structural similarity index (SSIM), or video quality metric (VQM). The quality metric can be used to evaluate the fidelity or perceptual quality of the video content.
[0071] In some example implementations, the same set of encoding parameters can be applied to both the first video codec and the second video codec.
[0072] In some example implementations, the first video codec and the second video codec can use different sets of encoding parameters.
[0073] In some example implementations, for each encoding parameter, the encoded video bitstream can be encoded by the codec. And the corresponding data points (whether anchor data points or test data points) can be calculated.
[0074] In some example implementations, multiple data points can be calculated from the same encoded video bitstream.
[0075] Output:
[0076] The expected output may include a quantitative measurement of the test performance relative to the anchor. Commonly used metrics may include the average target quality improvement at the same bitrate (e.g., BD-PSNR), and the average bitrate savings at the same target quality (e.g., BD-rate).
[0077] The following steps may be applicable to any existing codec performance comparison method or tool.
[0078] Step 1:
[0079] The first step involves separately encoding a reference video source (denoted as S) using the anchor and the test for each coding parameter (or i.e., operating point). This results in two sets of encoded bitstreams to be used as source signals for codec performance evaluation, one set denoted as [s a1 , s a2 , ……, s an , and the other set denoted as [s t1 , s t2 , ……, s tn . Here, "s ak " represents the source signal (or encoded bitstream) encoded from the reference video S by the anchor codec using the ck configuration. Similarly, "s tk " represents the source signal (or encoded bitstream) encoded from the reference video S by the test codec using the ck configuration. Where k is an integer from 1 to n.
[0080] Step 2:
[0081] For each of the operating points, the second step is to calculate the bitrate value r and the quality score q of the encoded source signals. This results in 4 datasets: 1. [ra1, ra2, ……, r an : This set represents the bitrate measurements of the source signals encoded by the anchor codec. For example, ra1 is the bitrate measurement of the source signal sa1 (encoded by the anchor codec). 2. [qa1, qa2, ……, q an : This set represents the target quality measurements of the videos encoded by the anchor codec. For example, q a1 is the target quality measurement of the source signal s a1 . 3. [rt1, rt2, ……, r tn : This set represents the bitrate measurements of the source signals encoded by the test codec. For example, rt1 is the bitrate measurement of the source signal st1 (encoded by the test codec). 4. [qt1, qt2, ……, q tn:This set represents the target quality measurement results of the source signal encoded by the test codec. For example, qt1 is the target quality measurement result of the source signal st1.
[0082] By taking the bitrate measurement result r as the x-axis and the target quality q as the y-axis, these four data sets together form an anchor data point set for the anchor codec performance and a test data point set for the test codec performance. Referring to the example Figure 8D , the four solid dots form the anchor data point set for the anchor codec, while the four "x"s form the test data point set for the test codec. Note that these data points are discrete data points. If the two data sets (represented by "x" and dots) contain points with the same x-axis value, it is possible to determine which codec (anchor or test) performs better by comparing their respective y-axis values. However, in practice, finding such points may be challenging or impossible. Therefore, it is proposed to fit curves with a non-linear relationship to fit each data point set (i.e., the anchor data point set and the test data point set). This method enables a quantitative comparison between the anchor codec and the test codec, thus facilitating a more meaningful evaluation of their performance.
[0083] In some example implementations, the anchor data point set for the anchor codec and the test data point set for the test codec have an overlap (e.g., on the x-axis).
[0084] In some example implementations, the anchor data point set for the anchor codec and the test data point set for the test codec do not have an overlap (e.g., on the x-axis). In this case, it is possible to extrapolate at least one data set.
[0085] In some example implementations, when obtaining the fitted curve, a specific range of x-axis values, such as [xa, xb], can be applied. This range can be determined by actual requirements, such as the bitrate range suitable for a specific video streaming environment. It is worth noting that this range can be adjusted to adapt to different use case scenarios.
[0086] In some example implementations, when obtaining the fitted curve, a specific range of y-axis values, such as [ya, yb], can be applied. This range can be determined by actual requirements such as the quality requirements of a specific video streaming environment. It is worth noting that this range can be adjusted to adapt to different use case scenarios.
[0088] Step 3:
[0089] Numerically fit the anchor data points, that is, having [ra1, ra2, ……, r an as the x-axis values and [qa1, qa2, ……, q anData points as y-axis values (e.g., (ra1, qa1), (ra2, qa2), ……, (ran, q an )) . Assume that the relationship between r (bit rate, x-axis value) and q (target quality, y-axis value) can be represented by the following polynomial function: fa(x) = b 0 x 3 + b 1 x 2 + b 2 x + b 3 (1)
[0090] Where b0, b1, b2, and b3 are coefficients that control the shape of the curve, and x is the input parameter (in this example, it corresponds to r), and y (f(x)) is the output value (in this example, it corresponds to q).
[0091] In some example implementations, when numerically fitting the anchor data points, special constraints for this polynomial can be added. The special constraint requires that its first derivative be positive (or non - negative). Additional constraints can be added. The additional constraint requires that its second derivative be negative (or non - positive). These constraints can be imposed within a given x range (e.g., bit rate range). The first constraint can ensure that the fitted curve is monotonically increasing, and the second constraint can ensure that the fitted curve is convex. Another special constraint for this polynomial is that its first derivative is positive within a given x range, which ensures that the curve is monotonically increasing. Another special constraint for this polynomial is that its first derivative is negative within a given x range, which ensures that the curve is monotonically decreasing. The first derivative is: f'(x) = 3b 0 x 2 + 2b 1 x + b 2 (2) And the second derivative is: f'(x) = 3b 0 x 2 + 2b 1 x + b 2 (3)
[0092] In some example implementations, the first derivative of the polynomial needs to be negative.
[0093] In some example implementations, additional constraints are imposed such that f(min(x)) > 0 and f(max(x)) < 100, where min(x) and max(x) represent the minimum and maximum values of x respectively.
[0094] The fitting curve can be obtained by solving, for example, a typical non-linear curve fitting problem. For example, the parameters b0, b1, b2, b3 can be obtained by minimizing the least square error between the observed values and the fitted values using any suitable optimization method.
[0095] As an example, let (x_i, y_i) for i in [1, 2, ……, N] represent the bit rate and metric value of a data point (e.g., the data points as Figure 8D shown). N = 6 is the number of encoding parameters (e.g., 6 quantization parameters (QuantizationParameter, QP)). Using any suitable optimization method, the least square error between the observed values and the fitted values can be solved using the following equation:
[0096] As another example, Figure 10 shows an exemplary process for solving the parameters b0, b1, b2, b3 using an iterative parameter optimization process. In Figure 10 it, the parameters a, b, c, and d to be optimized respectively correspond to the parameters b0, b1, b2, and b3 in the above polynomial function.
[0097] After applying the above optimization process, we obtain a non-linear curve (i.e., fa(x)) between the anchor bit rate and the quality measurement results (i.e., [ra1, ra2, ……, r an and [qa1, qa2, ……, q an ). Figure 7B shows an example non-linear curve 704 of the fitting curve as anchor codec data points (represented by "X").
[0098] Step 4:
[0099] For the test data (i.e., [rt1, rt2, ……, r tn and [qt1, qt2, ……, q tn ), repeat the non-linear relationship curve solving process as described in step 3 for the anchor codec. We obtain a non-linear relationship curve ft(x) between [rt1, rt2, ……, r tn and [qt1, qt2, ……, q tn . Figure 7B shows an example non-linear curve 702 of the fitting curve as test codec data points (represented by solid dots).
[0100] The fitting curve obtained using the aforementioned constraints can better capture the non-linear relationship between the bit rate and the performance metric (quality metric). Specifically, it has the following characteristics: 1) At the intermediate bit rate points, there is an almost linear relationship between the performance metric and the bit rate, and the performance metric rises rapidly as the bit rate increases. 2) When the bit rate increases to a certain extent, further increasing the bit rate may not bring additional gain in the performance metric, that is, the performance metric enters the saturation region. 3) When the bit rate decreases to a certain level, the performance metric will not decrease as the bit rate decreases.
[0101] In addition, the proposed constrained cubic curve fitting method has the following advantages: 1) As long as there is an overlapping region between the two curves (e.g., overlapping x-axis values), this method gives valid values for all test cases. 2) Compared with other methods (such as the BDExtend(pareto) method), this method is more transparent and consistent because all data points can be used for evaluation. 3) The cubic curve fitting method achieves a smaller fitting error than other methods such as the BDExtend(pareto) method. Therefore, this method can provide a more accurate quality metric value, such as the BD-rate value.
[0102] Step 5:
[0103] For a given performance evaluation interval [xa, xb], calculate the difference in the integral areas between fa(x) (i.e., the curve 704 in Figure 7B and ft(x) (i.e., the curve 702 in Figure 7B within the evaluation interval. This result will be denoted as BD-PSNR, which represents the average target PSNR gain at the same bit rate. BD-PSNR can be calculated using the following formula:
[0104] For the example interval [xa, xb], refer to Figure 7B .
[0105] Step 6:
[0106] For a given performance evaluation interval [ya, yb], calculate the difference in the integral areas between fa(x) (i.e., the curve 704 in Figure 7B and ft(x) (i.e., the curve 702 in Figure 7B within the evaluation interval [ya, yb]. This result can be further converted into BD-rate, which represents the average bit rate gain at the same target PSNR. BD-rate can be calculated using the following formula:
[0107] For the example interval [ya, yb], refer to Figure 7B 。
[0108] In some embodiments, the codec performance evaluation method described above can be applied to video streaming applications. In such an application, a video server can serve multiple end-users simultaneously (e.g., transmit video streams to them). Depending on the transmission channel / link, the bitrate of video streaming can be different. For example, User A may be covered by excellent WiFi signal and the link speed exceeds 200 Mbps (megabits per second), User B is under cellular coverage and the link speed may be around 50 Mbps, and User C is under poor coverage and the link speed is only around 10 Mbps. For the same video source, there can be multiple versions of encoded bitstreams, and these bitstreams can be encoded by different codecs. The codec performance evaluation tool according to the present disclosure can be used to compare 3 codecs, with 3 bitrate ranges (in Mbps): [0, 20], [20, 100], and [100, 300]. After the comparison, it turns out that: · Codec A performs excellently in the bitrate range [0, 20]; · Codec B performs excellently in the bitrate range [20, 100]; and · Codec C performs excellently in the bitrate range [100, 300].
[0109] Therefore, based on the coverage condition of each user, the video stream server can serve: · User A (link speed 200 Mbps), where the bitstream is encoded by Codec C; · User B (link speed 50 Mbps), where the bitstream is encoded by Codec B; and · User C (link speed 10 Mbps), where the bitstream is encoded by Codec A.
[0110] In some embodiments, a similar concept can be applied to a single user with network coverage conditions that change, for example, due to user movement. When the user moves, the available bandwidth may change due to factors such as changing cellular signal strength or network coverage conditions. In such scenarios, the video streaming platform dynamically switches the video bitstream based on the codec used for encoding. Specifically, the video streaming platform can employ a dynamic approach to deliver the best viewing experience based on the available bandwidth / link speed of the user. By leveraging multiple encoded bitstreams (for the same video source, e.g., the same movie), the platform can intelligently select the bitstream encoded by a codec that outperforms other codecs within a specific bitrate range corresponding to the user's current link speed. This adaptive strategy ensures that for a given bandwidth constraint, the most efficient codec is used to compress the selected bitstream, thereby delivering maximized video quality adapted to the current link speed. In other words, for the same video block, multiple video bitstreams can be encoded by different codecs. Using the encoding evaluation tools / algorithms provided in the present disclosure, the best bitstream encoded by a codec that outperforms other codecs within a specific bitrate range can be selected, and the specific bitrate range is determined by the bandwidth of the video streaming session.
[0111] Exemplary methods that follow the principles for measuring and evaluating codec performance underlying the above implementations may include some or all of the following steps: Step 1: Obtain m anchor data points, each anchor data point generated based on a corresponding anchor-encoded video bitstream, where m is an integer, and where: the corresponding anchor-encoded video bitstream is a bitstream encoded by an anchor video codec based on a reference video and a corresponding encoding parameter selected from m encoding parameters; each of the m anchor data points represents the anchor codec performance using the corresponding encoding parameter and is a pair formed by an x-axis value representing the bitrate or a change in the bitrate and a y-axis value representing the quality measurement result; Step 2: Obtain n test data points, each test data point generated based on a corresponding encoded test video bitstream, where n is an integer, and where: the corresponding encoded test video bitstream is a bitstream encoded by a test video codec based on a reference video and a corresponding encoding parameter selected from n encoding parameters; each of the n test data points represents the test codec performance using the corresponding encoding parameter and is a pair formed by an x-axis value representing the bitrate or a change in the bitrate and a y-axis value representing the quality measurement result; Step 3: Fit the m anchor data points with an anchor curve, the anchor curve being based on an anchor polynomial, where the anchor polynomial is monotonic within the x-axis range; Step 4: Fit the n test data points with a test curve, the test curve being based on a test polynomial, where the test polynomial is monotonic within the x-axis range; and Step 5: Evaluate the test codec performance based on the anchor curve and the test curve to obtain an evaluation result.
[0112] In any part or combination of the above implementation manners, at least one of the following conditions is satisfied: the m anchor data points are monotonically increasing; the m anchor data points are not monotonically increasing; the n test data points are monotonically increasing; or the n test data points are not monotonically increasing.
[0113] In any part or combination of the above implementation manners, the m encoding parameters are the same as the n encoding parameters, and m is equal to n.
[0114] In any part or combination of the above implementation manners, at least one of the following conditions is satisfied: the x-axis ranges of the m anchor data points and the x-axis ranges of the n test data points overlap; or the x-axis ranges of the m anchor data points and the x-axis ranges of the n test data points do not overlap.
[0115] In any part or combination of the above implementation manners, each of the anchor polynomial and the test polynomial is monotonically increasing.
[0116] The embodiments in the present disclosure can be used to evaluate and compare two codecs, for example, a first codec and a second codec. In some example implementations, the first codec can be an anchor codec (benchmark codec), and the second codec can be a test codec. Specifically, the test codec can be a codec under development and / or a codec in the standardization process to be adopted by a video encoding / decoding standard.
[0117] The above operations can be combined or arranged in any number or order as needed. Two or more of the steps and / or operations can be executed in parallel. The embodiments and implementation manners in the present disclosure can be used alone or in any order combination. The steps in an embodiment / method can be divided to form multiple sub-methods, and each of the sub-methods can be independent of the other steps in the embodiment and can form an independent solution. In addition, each of the methods (or embodiments) can be executed by a device, and the device can be implemented by a processing circuit system (for example, one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0118] The above technology can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, FIG. 13 shows a computer system (1800) suitable for implementing certain embodiments of the disclosed subject matter.
[0119] Computer software can be encoded using any suitable machine code or computer language, and any such suitable machine code or computer language can be subject to mechanisms such as assembly, compilation, linking, etc. to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), Graphics Processing Units (GPUs), etc. or executed through interpretation, microcode execution, etc.
[0120] Instructions can be executed on various types of computers or their components, and such various types of computers or their components include, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0121] The components for the computer system (1800) shown in FIG. 13 are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the present disclosure. The configuration of the components should also not be construed as having any dependency or requirement related to any one or combination of the components shown in the exemplary embodiment of the computer system (1800).
[0122] The computer system (1800) may include certain human - machine interface input devices. The input human - machine interface devices may include one or more of the following (only one of each depicted): keyboard (1801), mouse (1802), touchpad (1803), touch screen (1810), data glove (not shown), joystick (1805), microphone (1806), scanner (1807), camera device (1808).
[0123] The computer system (1800) may also include certain human - machine interface output devices. Such human - machine interface output devices can stimulate the senses of one or more human users through, for example, haptic output, sound, light, and smell / taste. Such human - machine interface output devices may include: haptic output devices (e.g., haptic feedback through the touch screen (1810), data glove (not shown), or joystick (1805), but there may also be haptic feedback devices that do not serve as input devices); audio output devices (e.g., speakers (1809), headphones (not depicted)); visual output devices (e.g., screen (1810), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch - screen input capability, each with or without haptic feedback capability - some of which may be able to output two - dimensional visual output or more than three - dimensional output through means such as stereoscopic output; virtual reality glasses (not depicted); holographic displays and fog machines (not depicted)); and printers (not depicted).
[0124] The computer system (1800) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1820) with media (1821) such as CD / DVD, thumb drives (1822), removable hard disk drives or solid state drives (1823), traditional magnetic media such as magnetic tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), etc.
[0125] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transient signals.
[0126] The computer system (1800) may also include an interface (1854) to one or more communication networks (1855). The network may be, for example, a wireless network, a wired network, an optical network. The network may also be a local area network, a wide area network, a metropolitan area network, a vehicular and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include: local area networks (such as Ethernet, wireless LAN); cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television; vehicular and industrial networks including CAN bus, etc.
[0127] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the core (1840) of the computer system (1800).
[0128] The core (1840) may include one or more central processing units (CPUs) (1841), a graphics processing unit (GPU) (1842), a dedicated programmable processing unit in the form of a field programmable gate area (FPGA) (1843), a hardware accelerator for certain tasks (1844), a graphics adapter (1850), etc. These devices, together with a read-only memory (ROM) (1845), a random access memory (1846), and an internal mass storage device (e.g., an internal hard disk drive, SSD, etc. that is not user-accessible) (1847), may be connected via a system bus (1848). In some computer systems, the system bus (1848) may be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus (1849) to the system bus (1848) of the core. In an example, a screen (1810) may be connected to the graphics adapter (1850). The architecture of the peripheral bus includes PCI, USB, etc.
[0129] A computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be those specially designed and constructed for the purposes of this disclosure, or they may be of the type well-known and available to those having skill in the art of computer software.
[0130] Although this disclosure has described several exemplary embodiments, there are variations, permutations, and various alternative equivalents that fall within the scope of this disclosure. Accordingly, it should be recognized that those skilled in the art will be able to conceive of many systems and methods that, although not explicitly shown or described herein, embody the principles of this disclosure and are thus within the spirit and scope of this disclosure.
Claims
1. A method for processing video data, comprising: A first plurality of anchor data points generated based on an anchor video bitstream is obtained, wherein: The anchor video bitstream is encoded by an anchor video codec based on a reference video and corresponding encoding parameters selected from a first plurality of encoding parameters; Each anchor data point represents the performance of the anchor codec using the corresponding encoding parameters, the anchor data point being formatted as a two-tuple comprising (i) a bit rate or a change in the bit rate and (ii) a quality measurement; A second plurality of test data points generated based on a test video bitstream is obtained, wherein: The test video bitstream is encoded by a test video codec based on the reference video and corresponding encoding parameters selected from a second plurality of encoding parameters; Each test data point represents a test codec performance using the corresponding encoding parameters, the test data point being formatted as the two-tuple; fitting the first plurality of anchor data points with an anchor curve, the anchor curve being based on an anchor polynomial, wherein the anchor polynomial is monotonic over a first axis; fitting the second plurality of test data points with a test curve, the test curve being based on a test polynomial, wherein the test polynomial is monotonic over the first axis; and The test codec performance is evaluated based on the anchor curve and the test curve to obtain an evaluation result.
2. The method according to claim 1, wherein: At least one of the following conditions is met: The first plurality of anchor data points are monotonically increasing; The first plurality of anchor data points are not monotonically increasing; The second plurality of test data points are monotonically increasing; or The second plurality of test data points is not monotonically increasing.
3. The method according to claim 1, wherein: The first plurality of encoding parameters are the same as the second plurality of encoding parameters.
4. The method according to any one of claims 1 to 3, wherein: At least one of the following conditions is met: The first axis range of the first plurality of anchor data points and the first axis range of the second plurality of test data points have an overlap; or The first axis range of the first plurality of anchor data points and the first axis range of the second plurality of test data points have no overlap.
5. The method according to any one of claims 1 to 3, wherein: Each of the anchor polynomial and the test polynomial is monotonically increasing.
6. The method according to any one of claims 1 to 3, wherein: The anchor polynomial is represented by the following formula: fa(x)=b0x 3 +b1x 2 +b2x+b3 wherein: b0, b1, b2 and b3 are coefficients, the x-axis of the equation represents the bit rate or a change in the bit rate, and the y-axis of the equation represents the quality measurement result; and The test polynomial is represented by the following formula: ft(x)=c0x 3 +c1x 2 +c2x+c3 Where: c0, c1, c2 and c3 are coefficients, the x-axis of the equation represents the bit rate or the change in the bit rate, and the y-axis of the equation represents the quality measurement result.
7. The method according to claim 6, further comprising: The anchor polynomial is obtained by using a constraint such that within the first axis range, the first-order derivative of the anchor polynomial is positive, wherein the first-order derivative of the anchor polynomial is represented by the following formula: f′(x)=3b0x 2 +2b1x+b2; and The anchor polynomial is derived using a constraint such that within the first axis range, the first-order derivative of the test polynomial is positive, wherein the first-order derivative of the test polynomial is represented by the following formula: f′(x)=3c0x 2 +2c1x+c2。 8. The method according to claim 7, further comprising: deriving the anchor polynomial with an additional constraint such that within the first axis, the second derivative of the anchor polynomial is negative; as well as The test polynomial is derived with an additional constraint such that the second derivative of the test polynomial is negative.
9. The method according to claim 7 further comprises using the following formula to derive Incremental Peak Signal-to-Noise Ratio BD-PSNR: in, xa and xb are the start and end of the first axis range respectively.
10. The method according to claim 7 further comprises using the following formula to derive BD-rate: in, ya and yb are the beginning and end of the quality measurement interval on the y-axis.
11. The method according to any one of claims 1 to 3, further comprising: In response to the evaluation result indicating that the anchor video codec outperforms the test video codec within the first axis range and a link speed of a video streaming session falls within the first axis range, selecting an encoded video bitstream encoded by the anchor video codec for the video streaming session; as well as In response to the evaluation result indicating that the test video codec outperforms the anchor video codec within the first axis range and a link speed of a video streaming session falls within the first axis range, selecting an encoded video bitstream encoded by the test video codec for the video streaming session.
12. The method according to claim 11, further comprising: The selected encoded video bitstream is transmitted in the video streaming session.
13. An apparatus for processing video data, the apparatus comprising a memory for storing computer instructions and a processor in communication with the memory, wherein: When the processor executes the computer instructions, the processor is configured to cause the device to: A first plurality of anchor data points generated based on an anchor video bitstream is obtained, wherein: The anchor video bitstream is encoded by an anchor video codec based on a reference video and corresponding encoding parameters selected from a first plurality of encoding parameters; Each anchor data point represents the performance of the anchor codec using the corresponding encoding parameters, the anchor data point being formatted as a two-tuple comprising (i) a bit rate or a change in the bit rate and (ii) a quality measurement; A second plurality of test data points generated based on a test video bitstream is obtained, wherein: The test video bitstream is encoded by a test video codec based on the reference video and corresponding encoding parameters selected from a second plurality of encoding parameters; Each test data point represents a test codec performance using the corresponding encoding parameters, the test data point being formatted as the two-tuple; fitting the first plurality of anchor data points with an anchor curve, the anchor curve being based on an anchor polynomial, wherein the anchor polynomial is monotonic over a first axis; fitting the second plurality of test data points with a test curve, the test curve being based on a test polynomial, wherein the test polynomial is monotonic over the first axis; and The test codec performance is evaluated based on the anchor curve and the test curve to obtain an evaluation result.
14. The device according to claim 13, wherein: At least one of the following conditions is met: The first plurality of anchor data points are monotonically increasing; The first plurality of anchor data points are not monotonically increasing; The second plurality of test data points are monotonically increasing; or The second plurality of test data points is not monotonically increasing.
15. The device according to any one of claims 13 to 14, wherein: At least one of the following conditions is met: The first axis range of the first plurality of anchor data points and the first axis range of the second plurality of test data points have an overlap; or The first axis range of the first plurality of anchor data points and the first axis range of the second plurality of test data points have no overlap.
16. The device according to any one of claims 13 to 14, wherein: Each of the anchor polynomial and the test polynomial is monotonically increasing.
17. The device according to any one of claims 13 to 14, wherein: The anchor polynomial is represented by the following formula: fa(x)=b0x 3 +b1x 2 +b2x+b3 wherein: b0, b1, b2 and b3 are coefficients, the x-axis of the equation represents the bit rate or a change in the bit rate, and the y-axis of the equation represents the quality measurement result; and The test polynomial is represented by the following formula: ft(x)=c0x 3 +c1x 2 +c2x+c3 Where: c0, c1, c2 and c3 are coefficients, the x-axis of the equation represents the bit rate or the change in the bit rate, and the y-axis of the equation represents the quality measurement result.
18. The device according to claim 17, wherein: When the processor executes the computer instructions, the processor is configured to further cause the device to: The anchor polynomial is obtained by using a constraint such that within the first axis range, the first-order derivative of the anchor polynomial is positive, wherein the first-order derivative of the anchor polynomial is represented by the following formula: f′(x)=3b0x 2 +2b1x+b2; and The anchor polynomial is derived using a constraint such that within the first axis range, the first-order derivative of the test polynomial is positive, wherein the first-order derivative of the test polynomial is represented by the following formula: f′(x)=3c0x 2 +2c1x+c2。 19. The apparatus according to claim 18, when the processor executes the computer instructions, the processor is configured to further cause the apparatus to: deriving the anchor polynomial with an additional constraint such that within the first axis, the second derivative of the anchor polynomial is negative; and The test polynomial is derived with an additional constraint such that the second derivative of the test polynomial is negative.
20. A non-transitory storage medium for storing computer-readable instructions that, when executed by a processor, cause the processor to: A first plurality of anchor data points generated based on an anchor video bitstream is obtained, wherein: The anchor video bitstream is encoded by an anchor video codec based on a reference video and corresponding encoding parameters selected from a first plurality of encoding parameters; Each anchor data point represents the performance of the anchor codec using the corresponding encoding parameters, the anchor data point being formatted as a two-tuple comprising (i) a bit rate or a change in the bit rate and (ii) a quality measurement; A second plurality of test data points generated based on a test video bitstream is obtained, wherein: The test video bitstream is encoded by a test video codec based on the reference video and corresponding encoding parameters selected from a second plurality of encoding parameters; Each test data point represents a test codec performance using the corresponding encoding parameters, the test data point being formatted as the two-tuple; fitting the first plurality of anchor data points with an anchor curve, the anchor curve being based on an anchor polynomial, wherein the anchor polynomial is monotonic over a first axis; fitting the second plurality of test data points with a test curve, the test curve being based on a test polynomial, wherein the test polynomial is monotonic over the first axis; and The test codec performance is evaluated based on the anchor curve and the test curve to obtain an evaluation result.