Methods and devices for producing a bit rate ladder for video streaming
Patent Information
- Application Number
- EP2023755403
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-17
- Filing Date
- 2023-08-14
- Publication Date
- 2025-06-25
Smart Images

Figure 1.1
Abstract
Description
[0001] Methods and devices for generating a bit rate ladder for video streaming
[0002] The present invention relates to methods and apparatus for generating a bit rate ladder for encoding representations of a video segment.
[0003] When transmitting video data, the quality of the video depends on the bitrate. The amount of data required can be so large that difficulties arise during data transmission over networks with limited bandwidth. Examples include the broadcast of a digital television program and image / video transmission over the internet or mobile networks.
[0004] Despite the common compression of image or video data before it is stored or transmitted over a network, the data volume of a video quality often cannot be reduced sufficiently for networks with limited bandwidth.
[0005] Streaming services therefore typically provide multiple versions of the same video, each with different quality levels. These different versions of the same video are also called representations of a video. They have different bitrates. The different bitrates are achieved by setting different coding parameters in the encoder. For example, the quantization step width can be set differently for different representations. This set of representations is called the bitrate ladder.
[0006] Since the desired image quality should be as high as possible, it is desirable to adapt the bit rate selection to the user's available bandwidth without having to accept significant loss of image quality. The object of the present invention is therefore to efficiently generate a bit rate ladder that meets predetermined quality specifications.
[0007] This object is achieved by the independent claims. The dependent claims define advantageous embodiments.
[0008] Some embodiments of the present invention allow a set of representations of video sections to be created such that the maximum quality difference in a quality measure is minimized while taking into account the costs of coding and storage. According to a first aspect, the present invention relates to a method for generating a bitrate ladder for coding representations of a video section. The method comprises determining a first set of sampling points, wherein a sampling point indicates a quality of a representation based on bitrate and resolution, and the quality is based on a comparison with an original representation. The method further comprises generating a second set of sampling points based on the first set of sampling points, wherein the second set contains more sampling points than the first set.Furthermore, the method comprises selecting a subset of support points of the second set, taking into account quality specifications for generating the bit rate ladder based on the subset of support points.
[0009] According to an embodiment of the present invention, determining the first set of support points may comprise selecting a first grid of value pairs in a bit rate resolution space, and determining qualities of representations at the value pairs of the first grid to obtain a first set of support points.
[0010] In one embodiment, the first grid may contain at least the predetermined value pairs maximum bit rate, maximum resolution, and minimum bit rate, minimum resolution, wherein the minimum bit rate for the minimum resolution is determined taking quality specifications into account, the maximum bit rate for the maximum resolution is determined taking quality specifications into account, the maximum resolution corresponds to a resolution of the original representation, and the minimum resolution corresponds to a predetermined resolution that is smaller than the resolution of the original representation.
[0011] For example, the quality specifications may contain at least two target quality levels, corresponding to a minimum target quality and a maximum target quality. Furthermore, a quality of a representation generated based on the minimum bit rate and the minimum resolution may fall below the minimum target quality, and a quality of a representation generated based on the maximum bit rate and the maximum resolution may exceed the maximum target quality.
[0012] In one embodiment, generating the second set of nodes may further comprise generating a second grid of value pairs in a bitrate-resolution space that contains value pairs of the first set, and generating qualities for the value pairs of the second set based on the nodes of the first set. According to one embodiment, generating qualities for the value pairs of the second set may comprise at least one of the following:
[0013] Interpolation of the support points, and / or
[0014] Processing by a neural network, and / or a combination thereof.
[0015] For example, processing by a neural network may comprise obtaining nodes of the first set or an interpolation of nodes of the first set as input data, and generating output data comprising processing the input data by one or more layers of the neural network.
[0016] In one embodiment, output data of the neural network may be processed by filtering the output data to satisfy monotonicity conditions and / or limiting the range of values of the predicted qualities.
[0017] For example, the quality specifications may include at least two target quality levels corresponding to a minimum target quality and a maximum target quality. Furthermore, selecting the subset of sampling points for each target quality level from the quality specifications may include determining a bit rate for a bit rate specification of an encoder. Furthermore, determining a bit rate for a bit rate specification may include determining a bit rate for each resolution whose associated predicted quality meets the quality specifications for the respective target quality level, and selecting the minimum bit rate from the determined bit rates as the bit rate specification.
[0018] In one embodiment, the determination of the bit rate for the bit rate specification may comprise an interpolation based on the support points of the second set.
[0019] For example, selecting the subset of nodes may comprise generating a representation comprising encoding the video portion with the respective bit rate specification for each target quality level from the quality specifications.
[0020] In one embodiment, the method may further comprise determining a quality of the generated representation and comparing the determined quality with the quality specifications. If the determined quality meets the quality specifications, the method may further comprise including the representation in the bitrate ladder. If the determined quality does not meet the quality specifications, the method may further comprise determining a new representation based on a new bitrate specification.
[0021] The present invention further relates, according to a second aspect, to a method for encoding representations of a video section. The method comprises the above-mentioned generation of a bitrate ladder, wherein the bitrate ladder contains two or more quality levels. The method further comprises generating a representation for each of the quality levels of the bitrate ladder, wherein generating the representation comprises encoding the video section according to the respective quality level.
[0022] According to an advantageous embodiment, a computer program is provided which comprises program instructions stored on a non-transferable, computer-readable medium and which, when executed on one or more processors, cause the one or more processors to perform the steps of one of the above-mentioned methods.
[0023] According to a third aspect, the present invention further relates to a device for generating a bitrate ladder for encoding representations of a video section. The device comprises a unit for determining a first set of support points, wherein a support point indicates a quality of a representation based on bitrate and resolution, and the quality is based on a comparison with an original representation. The device further comprises a unit for generating a second set of support points based on the first set of support points, wherein the second set contains more support points than the first set. Furthermore, the device comprises a unit for selecting a subset of support points of the second set, taking into account quality specifications for generating the bitrate ladder based on the subset of support points.
[0024] The present invention further relates, according to a fourth aspect, to a device for encoding representations of a video segment. The device comprises the above-mentioned device for generating a bit rate ladder. The device further comprises a unit for generating a representation for each of the quality levels of the bit rate ladder, wherein generating the representation comprises encoding the video segment according to the respective quality level.
[0025] Additional advantages and benefits of the present invention will become apparent from the detailed description of a preferred embodiment and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Fig. 1 shows a block diagram of an exemplary device for determining a bit rate ladder.
[0027] Fig. 2 shows exemplary relationships between bitrate and quality.
[0028] Fig. 3 shows examples of quality loss and maximum quality loss in a
[0029] Quality-bitrate diagram.
[0030] Fig. 4 shows an example classification into quality levels.
[0031] Fig. 5 shows an example determination of the maximum quality level.
[0032] Fig. 6 shows examples of the acceptance rate and the VMAF rating for
[0033] Video sections longer than 30 seconds.
[0034] Fig. 7 shows examples of the acceptance rate and the VMAF score for video segments shorter than 30 seconds.
[0035] Fig. 8 shows the determined dependence between MOS and VMAF rating.
[0036] Fig. 9 shows an exemplary subdivision into quality levels based on the VMAF rating.
[0037] Fig. 10 shows an exemplary block diagram of a scaler and encoder.
[0038] Fig. 11 shows an exemplary flowchart for generating a bit rate ladder.
[0039] Fig. 12 shows an exemplary flowchart for determining a first set
[0040] Support points.
[0041] Fig. 13 shows an exemplary flow chart for determining a second set of support points.
[0042] Fig. 14 shows an example flowchart for selecting representations based on the second set.
[0043] Fig. 15 shows an example of a first set of support points.
[0044] Fig. 16 shows an example of a second set of support points. Fig. 17 schematically shows the structure of a neural network for generating estimated quality values.
[0045] Fig. 18a-d schematically show the monotonicity filtering of estimated qualities.
[0046] Fig. 19 shows an example of a linear interpolation of support points of the second
[0047] Theorem for a constant spatial resolution.
[0048] Fig. 20 shows an example of a bit rate specification selection in a target range.
[0049] Fig. 21 shows an example of a generated bitrate ladder in the bitrate resolution space.
[0050] Fig. 22 shows an exemplary device that can execute program instructions.
[0051] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENT
[0052] In the following, a preferred embodiment of the present invention will be described in detail with reference to the drawings.
[0053] Fig. 1 shows a device 100 for determining a bitrate ladder, which can be used to encode a video sequence 140 into a plurality of representations 109 of different quality levels. Such a bitrate ladder can be determined individually for each video or for each video section. However, it can also be determined once and then used as a template for encoding a plurality of video sequences. The encoding is performed by the encoder 150. The encoder 150 can be a standardized encoder, such as H.264 / AVC ("Advanced Video Coding"), H.265 / HEVC ("High-Efficiency Video Coding"), H.266 / VVC ("Versatile Video Coding"), or AV1 ("AOMedia Video 1"). The present invention can be used with any encoders as long as they are parameterizable so that a desired bit rate and / or quality of the coded video sequence can be set by one or more coding parameters.
[0054] A video sequence 140 is a sequence of a plurality (two or more) images, which can also be abbreviated to "video" or "video signal". The term "video section" is also used below to emphasize that a video sequence to be coded, for example a film, does not necessarily have to be coded in its entirety, but in one or more sections. A video section can, on the one hand, be a temporal section, i.e., a subset of the total number of images in a video sequence. However, a video section can instead or additionally be a spatial section, e.g., a subpicture of an overall image.
[0055] The device 100 for determining a bit rate ladder may, for example, contain a device 110 for determining the quality levels.
[0056] Alternatively, the device 100 for determining a bit rate ladder can, for example, receive predetermined quality specifications as input parameters.
[0057] Quality is measured using a predefined quality metric. Preferably, the quality metric correlates with the quality perceived by viewers. Determining the quality levels involves defining a quality range in which the majority of representations should be located, as well as the levels themselves (number and / or distribution of levels within the quality range).
[0058] After the quality levels have been determined, the bitrate ladder can be determined in a device 120 based on the specific quality specifications containing the quality levels. This can be done, for example, for a specific codec. In general, however, it is also possible to use different codecs for specific quality levels. It can be advantageous to encode the representations of a bitrate ladder associated with high quality with a more efficient codec in terms of coding efficiency than the representations associated with low quality. If the less efficient codec is associated with a shorter encoding time, the encoding time can be reduced.
[0059] A bitrate ladder is a set of representations, each associated with a bitrate and a spatial resolution corresponding to the respective quality levels predetermined (in device 110). For example, a representation in the bitrate ladder is determined such that it leads to one of the quality levels. A bitrate here refers to the bitrate of an encoded video sequence (or video section). The spatial resolution refers to the number of samples, or pixels, in the horizontal and vertical directions that the video sequence (or video section) has. A specific codec or encoder 150 typically allows for adjustment of the bitrate. The bitrate ladder can therefore be determined by testing different bitrate settings. The video is encoded with each of the bitrate settings, and the quality is determined.Then, those bit rates whose qualities are closest to the predetermined quality levels are selected. However, such a procedure can require numerous encodings until a suitable bit rate ladder is found. It would be desirable to reduce the number of encodings. In Fig. 1, this search for bit rates is represented by loop 121 – device 120 configures the bit rate settings and the video sections to be encoded for encoder 150, and encoder 150 outputs a coded bit stream, which is decoded by decoder 155. The quality of the decoded video section is determined. The quality determination can be performed in decoder 155 or in device 120 based on the decoded video section.
[0060] It should be noted that different video sequences (e.g., with different content) can result in different qualities after encoding and decoding (also referred to as reconstruction), even with the same bitrate setting. Therefore, the bitrate ladder can be determined based on a plurality of encoded video sections 101 (provided as input 140 of the encoder 150). Furthermore, an encoder 150 does not have to directly support a bitrate input. The bitrate setting can be done indirectly, e.g., by setting the spatial (or temporal) resolution of the video, the quantization step (i.e., the quantization step width), the bit depth, or through other coding parameters. The above-mentioned facilities are functional and can all be implemented in any software and / or hardware.Streaming services use adaptive bitrates (ABR) to offer different video signal quality levels to end users with different bandwidths. With ABR streaming, the video signal is streamed at different bitrates (R). 1; ..., R k , .... RK encoded. These different bit rates R ± , ..., R k , ... , R K correspond to different quality levels Q , Q k , Q K . An encoded video signal of a specific bit rate and associated quality level is a representation (R k , Qk) and the set of all K representations (R 1; Qi), ... , (R K , Q K is a bitrate ladder.
[0061] The quality Q of a digital video signal increases with the bit rate R, as shown in Fig. 2. Furthermore, the quality associated with a bit rate can depend on the content of the video segment. More complex content in a video signal with low redundancy typically has lower quality than content with higher redundancy at the same bit rate.
[0062] The video signal can be scaled before encoding to achieve a different (spatial) resolution. For example, a video signal has an original resolution of 1920x1080 pixels, which is scaled to achieve a smaller spatial resolution, e.g., 640x360. Scaling can, for example, involve dropping pixels (“downsampling”), often in combination with prior low-pass filtering, and / or interpolating the pixels. In such a case, the quality Q can also depend on this resolution S. In other words, the quality can be specified as a function of bit rate and resolution: Q(R,s
[0063] Consequently, a bitrate ladder can be extended by the resolution as an additional parameter: a representation k can be determined depending on the bitrate R k and resolution S k specify as (R k ,S k , Q k (R k >Sk)
[0064] If predefined bitrates are used for all video content to create a bitrate ladder, this results in wasted data rate or storage for less complex content. Similarly, it can happen that insufficient data rate is provided for more complex content, leading to a reduction in subjective quality (as perceived by viewers (users).
[0065] Content-dependent bitrate ladders can be optimized for complete video content, such as an entire movie ("per-title encoding"), or for more detailed subdivisions, such as individual scenes within a movie ("per-scene encoding"). By taking the resulting quality into account, both data rate and storage space can be saved.
[0066] For example, the K bitrates in the bitrate ladder are sorted as follows: R ± < ... < R k < ... < R K. Accordingly, the corresponding quality levels < ... < Q k < ... < Q K . The K local resolutions are preferably sorted so that < ... < S k < ... < S K If the representations are not present in this sorting, they can be reordered into this sorting. The present invention therefore also applies to all sortings.
[0067] Each end-user device can request and stream content from a Content Delivery Network (CDN) at a bitrate suitable for the user's individual internet connection speed T. There are several possible bitrate selection strategies. For example, the highest possible bitrate that is lower than the individual internet connection speed T can be selected, i.e.
[0068] Furthermore, it is possible, for example, to distinguish between different representations, e.g. R P > Qp) and (Rp + i, Q p+1 ), after certain time periods in order to efficiently utilize the available transmission rate. However, the present invention is not limited to these examples. When using a set of representations with discrete bit rates R lt ... , R k , ... , R K the streamed video has a lower quality if the individual transmission rate T is not in the set R ± , ... , R k , ..., R K This difference defines the loss of quality
[0069] AQ(T) = Q(T) - Q(R P (T)), where Q(T) denotes the quality level that the user could receive based on his individual transmission rate, and QR P (T) denotes the maximum quality level, that the user can receive due to the discrete set of representations. This loss of quality is illustrated in Fig. 3.
[0070] In addition, a maximum quality loss Q max This maximum quality loss describes the difference in quality between two consecutive bit rates R p and R p+1 with the corresponding quality levels Q p and Q p+1 ,
[0071] Fig. 3 shows an example of such a maximum quality loss. If there are significant differences between the quality levels of consecutive representations, a quality loss Q(T) can lead to a significant subjective quality loss.
[0072] A large number of representations is required to minimize the maximum quality loss for all users, taking into account both low-bandwidth users, e.g., in mobile networks, and high-bandwidth users, e.g., in fiber optic connections. However, this leads to high coding and storage costs for operators. Consequently, the maximum quality loss in a quality measure should be minimized while taking into account the coding and storage costs.
[0073] To automate the generation of the set of representations, the subjective user perception is estimated using an objective quality measure. Such an objective quality measure can be an estimate of a subjective quality. Examples of objective quality measures that estimate subjective quality are VMAF, ITU-T P.1203, or the structural similarity index (SSIM). However, the present invention is not limited to the use of the aforementioned examples, and other, non-standardized quality measures can be used.
[0074] The quality measure can be a VMAF (Video Multi-Method Assessment Fusion, VMAF) metric. The VMAF metric is an objective metric for algorithmically evaluating image quality in videos. It evaluates a modified video (e.g., through transcoding and / or scaling) based on a comparison with an undisturbed reference (original). The undisturbed reference (original representation) corresponds to the original video signal to be encoded at its original resolution.
[0075] The VMAF metric assigns a rating between 0 and 100 to a video signal. A rating of 0 corresponds to a low subjective quality estimate, while a rating of 100 corresponds to a high subjective quality estimate. The mean of the VMAF ratings of all frames of a video signal is defined below as the VMAF rating of the video signal. The quality Q corresponds to the VMAF rating VMAF. This results in the quality difference.
[0076] Determination of quality levels
[0077] An example of how to determine quality levels is described below. Using a quality measure such as the one described above, a bitrate ladder consisting of a set of representations can be generated such that a predefined maximum quality loss is maintained between adjacent quality levels.
[0078] Using a quality measure, a minimum quality level Q min and a maximum quality level Q max can be determined. Starting from the minimum or maximum quality level, a set of quality levels can be generated. This set of quality levels consists of K quality levels, where K > 2.
[0079] The lowest quality level is below the minimum quality level Q min or is equal to the minimum quality level Q min . It applies < Qmin . The highest quality level Q K is above the maximum quality level Q max or is equal to the maximum quality level Q max - Q applies K > Q max -
[0080] In other words, the range of values between the lowest and the highest quality level < Q min < Q k < Q max In this exemplary determination of quality specifications, QK is divided into sections that do not exceed the maximum quality difference. The maximum quality difference between each pair of directly consecutive representations (R k , Q k ) and R k+1 , Q k+1 ) is less than or equal to AQ max for all transmission rates T in the range R < T < R K The classification based on such a maximum quality difference is shown in Fig. 4 as an example for a VMAF assessment.
[0081] The maximum quality level can be determined by a quality at which a predetermined number of viewers cannot distinguish the representation corresponding to the quality from an original representation.
[0082] The predetermined number of viewers can be derived from standardized test methods. One example is the well-known and standardized "Double Stimulus Impairment" test method according to ITU-R BT.500 (ITU-R., "Rec. BT.500-14: Methodologies for the subjective assessment of the quality of television images" (2019)). However, the present invention is not limited to the use of the aforementioned example. Another methodology can be determined and applied.
[0083] An example determination of the maximum quality level is shown in Fig. 5. To obtain a possible set of ratings on the VMAF scale, tests were conducted with test subjects. Based on the determined mean opinion score (MOS) for various VMAF ratings of test video sequences, an example of a smallest possible maximum quality level is achieved with a VMAF rating of 95. In this example, the MOS is specified on a scale between 0 (very disturbing impairments) and 10 (unnoticeable impairments).
[0084] The minimum quality level Q min can be determined using an acceptance measure. This acceptance measure can specify a minimum quality at which a predetermined number of viewers find the representation corresponding to the minimum quality acceptable.
[0085] An example determination of the maximum quality level is shown in Figs. 6 and 7. Fig. 6 shows an example acceptance rate for video sequences longer than 30 seconds. A distinction is made between free and paid streaming services. Fig. 7 shows an example acceptance rate for video sequences shorter than 30 seconds. Acceptance is indicated with 0 (unacceptable subjective quality) or 1 (acceptable subjective quality). The acceptance rate is calculated from the mean value for all test subjects.
[0086] If an acceptance rate of 0.5 is required, a possible minimum quality level of 55 on the VMAF scale was determined. This lower limit can be changed by further criteria. For example, the minimum rating on the VMAF scale for a first streaming service should be 10 to 15 higher than for a second streaming service. For video sequences longer than 30 seconds, the minimum VMAF quality level should be 70 for second-type streaming services or 85 for first-type streaming services. The first and second streaming services can be paid or free, respectively, but they do not have to be paid or free.
[0087] A minimum number of quality levels can be determined by the maximum quality level, the minimum quality level, and the maximum quality gap. The minimum number can result from the generation of the quality levels.
[0088] A determined relationship between the VMAF rating and the MOS is approximately linear, justifying a constant, maximum quality gap for all neighboring pairs in the set of representations. This approximately linear relationship is illustrated as an example in Fig. 8 in a MOS-VMAF diagram.
[0089] The maximum quality difference can be chosen so that a subjective quality of the video signal for R k and R k+1 is the same for a predetermined number of viewers. To determine the maximum quality difference AQ maxUsing the VMAF metric as an example, all pairs of VMAF ratings and corresponding opinion scores (OS) can be evaluated. The lower VMAF rating is denoted by VMAF and the higher VMAF rating is denoted by VMAF h This results in values for maximum quality differences
[0090] AVMAF max = VMAF h - VMAF .
[0091] AVMAF max can be determined, for example, using Fig. 8. Non-overlapping confidence intervals of the measured MOS values can indicate a distinguishability of the quality. Accordingly, AVMAF max be chosen so large that overlapping confidence intervals of the measured MOS values result in order to achieve identical subjective quality. For the maximum quality difference, a distance of AVMAF max = 2.
[0092] For a maximum quality level of 95 on the VMAF scale, a minimum quality level of 55, and a maximum quality gap of 2, at least 21 quality levels result. However, the present invention is not limited to the use of the exemplary values mentioned. Different values may result when using a different quality measure.
[0093] Fig. 9 shows an example classification in a VMAF bitrate diagram. For clarity, a maximum quality difference of 5 on the VMAF scale was chosen in this example. This example results in nine quality levels in the value range that includes the minimum and maximum quality levels. However, this example representation does not correspond to an ideal set of quality levels, as quality differences may be noticeable between the levels. However, it is possible that the representation corresponds to a practical set.
[0094] Creating a bitrate ladder
[0095] As mentioned above, a specific codec or encoder 150 typically allows setting the desired bitrate, but not a direct setting of a target quality. Testing different bitrate settings typically requires encoding a video at each of these bitrate settings and determining the corresponding quality. It is desirable to minimize the number of such test encodings to reduce processing overhead. At the same time, it is desirable to maintain predetermined quality specifications and also minimize storage requirements.
[0096] The quality specifications may contain quality levels. These quality levels can, for example, be defined by a maximum target quality Q max , a minimum target quality Q min and a maximum quality gap AQ maxSuch quality levels can be generated, for example, according to the procedure in section Determination of Quality Levels. In the case of the above-mentioned VMAF metric, target quality levels can be obtained, for example, according to the following values: VMAF max = 95, VMAF min = 79 and AVMAF max = 2. The present invention is not limited to these exemplary values, in particular AVMAF max also depend on the absolute VMAF value of the respective quality level.
[0097] Furthermore, the quality specifications may contain one or more parameters 6;, the values of which indicate permissible deviations from target quality levels.
[0098] For example, the lowest quality level could in the bitrate ladder by a maximum of a predetermined value of a parameter of the minimum target quality Q min differ: Qmin - < Qi < Q minSimilarly, the highest quality level Q K in the bitrate ladder deviate from the maximum target quality Qmax by a maximum of a predetermined value of a parameter e2: Q max QC Qmax + e 2- The distance between two adjacent quality levels Q k and Q k+1 could, for example, be located in an area defined by the maximum quality distance AQ max and the value of a third parameter e3 is determined: AQ max - e3< Qk+i ~ Qk ^ AQmax- The values of the parameters 6; can be different or the values of the parameters 6; can be the same, ie e1= e2= e3= e. The values are preferably all positive. The values e3 can be different for each k, ie AQ max - e 3 k < Q k+ - Q k < AQ maxFor example, possible values of one or more parameters for the VMAF metric lie in an interval of [0.05; 0.5]. Smaller values may allow the generation of representations closer to the desired target quality, but a larger number of trial encodings may be required to achieve this accuracy. On the other hand, larger allowable deviations may allow a reduction in the number of trial encodings required to generate a representation.
[0099] To generate a bitrate ladder considering quality specifications, an initial set of nodes is determined. A node specifies the quality of a representation based on bitrate and resolution. In other words, a node is a tuple consisting of bitrate R, resolution S, and the associated quality (Q, R, S, QR, sy).
[0100] A method for generating a bit rate ladder is shown as an example in the flow chart in Fig. 11.
[0101] The first set of sampling points can be determined S1110 by selecting a group of value pairs in a bit rate-resolution space. The determination of a first set of sampling points is illustrated, for example, in the flowchart in Fig. 12.
[0102] The group of value pairs can be arranged as a grid. For example, three bit rates and three resolutions can be selected to generate a 3x3 grid of value pairs (Ri. Sj). The grid can also be generated with any other number of bit rates and / or any other number of resolutions. Such a grid can have any dimension N R x N s have, where N R and N s Integers greater than or equal to 1.
[0103] For example, an N R x / V s-grid, e.g. a 5x5 grid, at pairs of values Rt.Sj) can be generated S1220. From this, a subset of S support points can be selected S1230. Such a subset can, for example, be a 3x3 grid, the main diagonal of the N R x N s - grid, a checkerboard pattern, or something similar.
[0104] The quality, which depends on the bit rate R and the resolution S of the (scaled and) encoded video section (representation), can be determined using the selected value pairs S1240.
[0105] The determination includes, for example, scaling and encoding the original video section (original representation) in the original resolution. An exemplary embodiment of such an encoder is shown in Fig. 10. The original video section 1010 is scaled to obtain a resolution different from the original resolution. This target resolution S Coded1020 is an input parameter for the scaler and encoder 1040. Furthermore, the exemplary scaler and encoder 1040 receives the corresponding bit rate of the value pair as the target bit rate R RC 1130. The encoder 1040 generates an encoded representation 1050 with resolution S Coded and bit rate R Co ded- The bitrate R Co ded of the encoded video signal 1050 may differ from the target bit rate R RC It is also possible that the encoder does not work deterministically and that repeated encoding with the same parameters results in slightly different bit rates R CoThe scaler and encoder do not necessarily have to be combined into a single unit, as shown in Fig. 10. The scaler and encoder can be separate units. For example, a scaler can output a video signal with a modified resolution, and an encoder can receive a video section with the target resolution to encode it.
[0106] In this example, the encoded representation is decoded and, if necessary, scaled to the original resolution. Such scaling ("upsampling") can be achieved by interpolating the decoded video signal, e.g., by bicubic filtering. The decoded (and scaled) video section is compared with the original representation to determine the (objective) quality of the encoded representation. This objective quality can be specified, for example, with the VM AF metric or any other objective video metric.
[0107] In general, specifying a sampling point does not necessarily require generating an encoded video segment to determine quality based on a comparison with an original representation. For example, a sampling point can also be generated by specifying an estimated quality Q. An estimated quality can be obtained, for example, through interpolation, extrapolation, processing by a neural network, or similar.
[0108] The first grid, which contains the bitrate-resolution value pairs for the first set of support points, contains at least the (predetermined) value pairs maximum bitrate R NR , maximum resolution S Ns , and minimum bitrate R l t minimum resolution
[0109] The maximum resolution typically corresponds to the original resolution. The minimum resolution, for example, corresponds to a predetermined resolution that is smaller than the resolution of the original representation.
[0110] In an exemplary embodiment, the minimum bit rate R ± for the minimum resolution taking into account quality specifications S1210. The minimum bit rate can be selected so that a predetermined minimum quality is not reached. Likewise, the maximum bit rate R NR for maximum resolution S Ns determined taking quality specifications into account. The maximum bitrate can be selected so that a predetermined maximum quality is exceeded.
[0111] As described above, the quality specifications can contain at least two target quality levels corresponding to a minimum target quality Q min and a maximum target quality Q maxare equivalent to.
[0112] The minimum bit rate can be chosen so that a corresponding representation for the smallest resolution a quality Q(R , S ) is achieved which is less than or equal to the minimum target quality and the permissible deviation e1: Q(R , S ) < Q min - e1. The maximum bit rate can be chosen so that a corresponding representation for the highest resolution S Ns a quality (?(R WR , SI) which is greater than or equal to the maximum target quality and the permissible deviation e2: Q(R NR ,S NS ) > Q max + e2.
[0113] The other (N R - 2) Bit rates R2, ... , R NR -I can be calculated, for example, as follows: where f (bitrate) indicates that a function is applied to the bitrate. For example, the logarithm to base 2 can be used, ie
[0114] In addition to the maximum and minimum resolution, further N s - 2 local resolutions S2, ...,S Ws -i. It should be assumed that each spatial resolution S n should be larger than For example, typical local resolutions W x H, e.g. 1920x1080, 1280x720, 640x360, 320x180, 160x140 can be used.
[0115] An exemplary 5x5 grid may contain resolutions with a width W of {512; 768; 1024; 1280; 1920} pixels. The bit rates of such a 5x5 grid may be arranged logarithmically between a predetermined minimum bit rate and a predetermined maximum bit rate. A predetermined minimum bit rate and a predetermined maximum bit rate may be determined, for example, based on quality specifications. In an exemplary implementation, the minimum bit rate R1 = 250 kbit / s and the maximum bit rate R ±= 2000 kbit / s is used. The logarithmically arranged additional bit rates in this example are R2 = 420 kbit / s, R3 = 707 kbit / s, and R4 = 1189 kbit / s. A subset of this 5x5 grid can be selected as the first set of sampling points, as described above. Alternatively, all value pairs of the example 5x5 grid can be used to determine the first set of sampling points.
[0116] As described above, for each selected pair of values (Ri. Sj), a coded representation is generated and the quality is determined to generate a support point for the first set of support points. A first set of support points is shown as an example in Fig. 15. The corners or intersection points of the displayed grid represent the determined support points R n ,S m , QR n ,S m )) in the VMAF metric.
[0117] Based on the first set of support points, a second set of support points is generated S1120, wherein the second set contains more support points than the first set. The second set may contain one or more support points from the first set. Preferably, all support points from the first set are included in the second set.
[0118] The generation of the second set of sampling points comprises, for example, predicting (generating) qualities Q for value pairs Rt,Sj) in the bitrate-resolution space that are not contained in the first set of sampling points. The generation comprises, for example, interpolation, extrapolation, processing by a neural network, a combination thereof, or other methods for generating additional sampling points based on the first set of sampling points. A constraint for a prediction comprises, for example, that for each resolution S m As the bitrate increases, the quality also increases:
[0119] To generate the second set of support points, a second grid of value pairs can be created. The second grid has, for example, an arbitrary dimension M R x M s have, where M R > N R and M s > N s An exemplary second grid contains the value pairs of the first set. For example, such a second grid includes 45 resolutions and 129 bit rates. The present invention is not limited to these exemplary numerical values. As described above, the second grid may include any number of value pairs greater than the number of value pairs in the first set.
[0120] The qualities Q at the value pairs of the second set are generated (predicated) based on the support points of the first set. The support points of the first set contain, as described above, the determined qualities Q at the value pairs of the first set.
[0121] As already indicated, generating qualities Q for the value pairs of the second set may comprise at least one of the following: interpolation of the support points and / or processing by a neural network, and / or a combination thereof.
[0122] In the exemplary second grid with 45 resolutions and 129 bit rates mentioned above, 5805 estimated qualities are generated.
[0123] In a first exemplary embodiment, the sampling points of the first set are interpolated to obtain (estimated) qualities Q for the value pairs of the second set. Interpolation for the resolution can be performed, for example, using a cubic interpolation polynomial. Interpolation for the bit rate can be performed, for example, using a power series model with one or more terms.
[0124] In a second exemplary embodiment, the sampling points of the first set are processed by a neural network to obtain estimated qualities Q for the value pairs of the second set. For example, the neural network receives the sampling points of the first set as input data, wherein the sampling points each contain a bit rate, a resolution, and the associated determined (measured) quality Q. Additionally, the neural network can receive the value pairs of the second set as input parameters. For example, the neural network is trained to output estimated qualities Q for the value pairs of the second set. The neural network processes the input data through one or more layers to generate output data. The output data contains estimated qualities Q for the value pairs of the second set.
[0125] In a third exemplary embodiment, illustrated in the flow chart in Fig. 13, the sampling points of the first set are interpolated S1310, analogously to the first exemplary embodiment, to obtain estimated qualities Q for the value pairs of the second set. This estimate can be refined by using the sampling points of the first set as input data for a neural network. Such a neural network is trained, for example, to improve (refine) the qualities Q estimated by the interpolation. The neural network processes the input data S1320 through one or more layers to generate output data. The output data contains estimated qualities Q for the value pairs of the second set.
[0126] An example structure of a neural network is shown in Fig. 17. An input layer 1710 receives the input data as a two-dimensional matrix. A neural network can, for example, contain one or more convolutional layers that can operate with different convolution matrices (“kernels”) and strides of different sizes. Normalizing the output of a convolutional layer can increase its efficiency. Typically, a (normalized) output of such a convolution is processed by a non-linear activation function. Multiple blocks 1720, 1730, 1740 consisting of a convolution, a normalization, and a non-linear activation can be applied both in parallel and in series. A possible application in series is indicated in Fig. 17 by “1x,” “2x,” etc.For example, normalization can be applied to a small number of data sets from the previous layer (“batch normalization”). A nonlinear activation function can be, for example, a sigmoid function, a hyperbolic tangent, or a rectifier (“rectified linear unit,” ReLU). The neural network can also contain fully connected layers 1750 and further filters 1751, which can, for example, reduce the dimension of the weights and / or deactivate individual neurons from a previous layer (“drop-out layer”). Further layers 1760 can generate the desired dimension of the output data (“depth to space”). Any data generated in parallel can be combined by element-wise addition 1770. An output layer 1780 generates the output data described above.
[0127] However, the present invention is not limited to a network of this exemplary structure. In general, the neural network can include any combination of (different) layers that generate the desired output data from the input data described above. While convolutional networks can be advantageous due to their ability to effectively compress two-dimensional correlated data, the present invention is not limited to the application of convolutional networks.
[0128] The output data of a neural network can be further processed by filtering and / or limiting the value range of the predicted (estimated) qualities.
[0129] In an exemplary implementation, monotonicity conditions can be met by filtering S1330. Such a monotonicity condition includes, for example, that for each resolution S mAs the bit rate increases, the quality also increases: VMAF(ß n ,S m ) > VMAF(R n- , S m ). This can be achieved, for example, by local (low-pass) filtering.
[0130] Fig. 18 shows an example of filtering using the VMAF metric. If, as in Fig. 18a, a minimum 1810 can be found, e.g., by changing the sign of the gradient, a new quality value VMAF neu 1840 from the old value VMAF ait 1810 and the two neighboring values VMAF N1 1820 and VMAF N2 1830, e.g. VMAF neu = 0.50 VMAF ait + 0.25 VMAF N1 + 0.25 VMAF N2The adjustment of a value is shown in Fig. 18b. If the new value 1840 again represents a minimum, the step is repeated to obtain another new value 1850, as shown in Fig. 18c. The quality as a function of the bit rate can now have a minimum at another point 1860. The determination of a new quality value is repeated until the function no longer has a minimum, as shown in Fig. 18d. The quality as a function of the bit rate is thus monotonically increasing.
[0131] For example, the range of predicted qualities can be restricted. It is possible that the neural network estimates quality values that lie outside the range of the quality measure used. For example, the VMAF metric allows values between 0 and 100 and can be restricted as follows: f 100 ; > 100
[0132] VMAF (R n ,S m ~) = 0 ; ) < 0
[0133] (yMAF(Rn ,S m ) ;
[0134] A second set of support points is shown as an example in Fig. 16. The corners or intersection points of the displayed grid represent the generated support points (R n ,S m , Q(R n , S m )) with the estimated qualities Q in the VMAF metric.
[0135] The support points can, for example, be additionally weighted S1350 by predetermined criteria, such as bit rate, expected coding time or spatial resolution.
[0136] A subset of the support points of the second set is selected S1130, taking quality specifications into account, to generate the bit rate ladder S1140. Figure 14 shows an exemplary flow chart for generating the bit rate ladder from the support points of the second set.
[0137] As described above, the quality specifications can contain at least two target quality levels corresponding to a minimum target quality Q min and a maximum target quality Q max In addition, the quality specifications can contain further target quality levels, which are defined, for example, by a maximum quality difference between two adjacent quality levels, as described above. Alternatively, further target quality levels can also be explicitly specified in the quality measure used. For each of the K target quality levels from the quality specifications, a representation can be determined and included in the bitrate ladder. The determination of the representation is described below using the example of a current k-th target quality level. Representations for further quality levels can be generated analogously. The bitrate ladder can be based on the minimum target quality Q min, also starting from the maximum target quality Q max Fig. 14 shows an example flowchart for generating the bit rate ladder starting from the maximum target quality Q max , i.e. starting with the K-th level of the bitrate ladder.
[0138] For a current level from the target quality levels, initial coding parameters, e.g. resolution and bit rate specification, are calculated S1410. For this purpose, a bit rate is determined for a bit rate specification of an encoder. For each resolution in the second set, a bit rate is determined whose predicted quality meets the quality specifications. There may be local resolutions whose associated qualities do not meet the quality specifications. These are not taken into account when selecting a bit rate specification. Furthermore, further boundary conditions with regard to the local resolution can also be specified. For example, only local resolutions that have a minimum size can be taken into account, e.g. all local resolutions above 1280x720 samples. However, only local resolutions that are less than or equal to a specified size can be taken into account, e.g. all local resolutions less than or equal to 1280x720 samples.
[0139] The quality specifications for certain (measured) qualities Q as well as for predicted qualities Q include, for example, the conditions described above:
[0140] Qmin ^1 — Ql — Qmin>
[0141] Qma X — QK — Qma X + ^2> Qma X 3 — Qk+1 Qk — ^ Qma X -
[0142] From the bitrates determined in this way, the smallest bitrate is selected as the bitrate specification. The resolution associated with the selected bitrate is used as the target resolution. When selecting the bitrate for the kth level of K quality levels, it can also be taken into account that the spatial resolution S k for increasing quality < ... < Q k < ... < Q K should not become smaller: Si < ... < S k < ... < S K .
[0143] The determination of the bit rate for the bit rate specification includes, for example, an interpolation based on the support points of the second set. For example, as shown in Fig. 19 for the estimated values of the VMAF metric, an interpolation of the (estimated) quality as a function of the bit rate at a constant resolution S m Such an interpolation can, for example, be a linear interpolation.
[0144] Fig. 20 shows an example of determining a bit rate control R RC for a target quality in the range between VMAF ziei and VMAF ziei + e. The bit rate specification R RC is chosen so that the value of the quality VMAF ziei + e / 2, which was determined by interpolating the predicted qualities. This increases the probability that the actual quality is within the desired range.
[0145] A representation can be generated for the current target quality level from the quality specifications. This involves encoding S1420 the video section with the respective selected bitrate specification. If necessary, the video section can be scaled to the corresponding spatial resolution before encoding.
[0146] The quality Q of the generated representation can be determined (measured). This can be done, as described above, by an objective comparison with the original representation. The determined quality Q can be compared with the quality specifications S1430.
[0147] If the determined quality meets the quality specifications ("Yes" in S1430), the representation can be included in the bitrate ladder S1440. After the representation for the current k-th target quality level has been included in the bitrate ladder, a representation for the (k-1)-th target quality level can be determined analogously S1460, provided that the lowest target level of the bitrate ladder has not yet been reached ("No" in S1450), i.e., k > 1. If k = 1 ("Yes" in S1450), the bitrate ladder has been completely generated and the exemplary process in Fig. 14 is terminated.
[0148] The further representations to be generated are determined based on the respective previous representation included in the bitrate ladder. This can be determined by the quality specification for neighboring quality levels AQ max - e3< Q k+1 - Q k < AQ max , which shows the maximum quality difference AQ maxand the corresponding permissible deviation e3. Furthermore, the spatial resolution S k for increasing quality not become smaller.
[0149] If the determined quality of the generated representation does not meet the quality specifications ("No" in S1430), a new representation can be determined based on a new bit rate specification. For a redetermination of the coding parameters S1470, for example, the (estimated) quality as a function of the bit rate, which is shown by way of example in Fig. 19, can be supplemented by the new, measured quality of the generated representation. For example, the interpolation described in detail with reference to Fig. 20 can be repeated with the added (determined) quality value. For example, with the added (determined) quality value, a bilinear interpolation can be performed between the already encoded and predicted support points in order to reduce possible deviations from estimated and determined quality values. With the newly determined coding parameters, another representation is generated in the loop S1480.
[0150] This process of creating representations and comparing the respective determined qualities with the quality specifications can be repeated until a representation is created that meets the quality specifications.
[0151] Generating (estimated) qualities Q enables improved determination of coding parameters and can thus reduce the number of test codes required to generate a representation of the bitrate ladder. The number of test codes required may depend on the permissible deviations 6; from target quality levels.
[0152] An example bitrate ladder is shown in Fig. 21. The dots mark the quality levels of the bitrate ladder.
[0153] The described exemplary embodiments for generating a bit rate ladder can be combined as desired, unless explicitly stated otherwise.
[0154] A bitrate ladder generated as described above can be used to encode representations of another video section. In other words, the generated bitrate ladder can be used to encode one or more other video sections. The generated bitrate ladder contains two or more quality levels. For each of the quality levels of the bitrate ladder, a representation of the other video section can be generated. The generation includes encoding the other video section according to the respective quality level. The quality level includes an associated target bitrate and a target resolution.
[0155] Although the embodiments of the invention have been described based on coding video data, the invention is not limited thereto but can also be used for coding still images.
[0156] Embodiments of the present invention and their functions may be implemented in hardware, software, firmware, or a combination thereof, as shown by way of example in Fig. 22. When embodiments are implemented in software, the functions may be stored on a computer-readable storage medium 2230 or transmitted over a communication channel 2240 (e.g., a bus) as instructions or code executed by a hardware-based processor unit 2220. For example, a computer-readable storage medium 1130 may be a RAM, ROM, EEPROM, CD-ROM or other optical storage medium, a magnetic storage medium, flash memory, or other storage medium that can be used to store program code in the form of instructions so that it can be read by a computer.
[0157] Instructions can be executed by one or more processors, such as digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits, programmable logic gates (Field Programmable Gate Arrays, FPGAs), or other integrated or discrete logic circuits. Accordingly, the term "processor" can refer to one of the structures mentioned or other structures suitable for implementing the methods described above. Furthermore, the described functionalities can be implemented in dedicated hardware and / or software modules configured to encode and / or decode image data, including within the framework of a combined codec. The methods can also be implemented in one or more circuits or logic elements.
[0158] The processor 2220 can thus implement the device 110 or 120, or the apparatus 100 for determining a bit rate ladder.
[0159] An apparatus for determining the quality specifications for encoding representations of a video section comprises a unit that determines the maximum and minimum quality levels as described above, and a unit that determines the set of quality levels with predefined maximum quality distance between adjacent quality levels as described above.
[0160] A device for generating a bit rate ladder for encoding representations of a video section, comprising a unit for determining a first set of support points, wherein a support point indicates a quality of a representation based on bit rate and resolution and the quality is based on a comparison with an original representation, a unit for generating a second set of support points based on the first set of support points, wherein the second set contains more support points than the first set, and a unit for selecting a subset of support points of the second set taking into account quality specifications for generating the bit rate ladder based on the subset of support points.
[0161] An apparatus for encoding representations of a video portion, comprising a unit for generating the bit rate ladder as described above, and a unit for generating a representation for each of the quality levels of the bit rate ladder, comprising encoding the video portion according to the respective quality level.
[0162] In summary, the present invention relates to methods and devices for generating a bitrate ladder for encoding representations of a video segment. The generation involves generating a set of nodes, each node indicating a quality of a representation based on bitrate and resolution. A subset of nodes is selected for generating the bitrate ladder, taking quality specifications into account.
Claims
CLAIMS 1 . A method for generating a bitrate ladder for encoding representations of a video segment, comprising: Determining a first set of nodes, wherein a node indicates a quality of a representation based on bit rate and resolution, and the quality is based on a comparison with an original representation; Creating a second set of support points based on the first set of support points, the second set containing more support points than the first set, Selecting a subset of sampling points of the second set taking into account quality specifications for generating the bitrate ladder based on the subset of sampling points.
2. The method according to claim 1, wherein determining the first set of support points comprises: Selecting a first grid of value pairs in a bitrate-resolution space, and Determining qualities of representations on the value pairs of the first grid to obtain a first set of support points.
3. Method according to claim 2, wherein the first grid contains at least the predetermined value pairs maximum bit rate, maximum resolution, and minimum bit rate, minimum resolution, wherein the minimum bit rate for the minimum resolution is determined taking into account quality specifications, the maximum bit rate for the maximum resolution is determined taking into account Quality specifications are determined, the maximum resolution corresponds to a resolution of the original representation, and the minimum resolution corresponds to a predetermined resolution that is smaller than the resolution of the original representation.
4. The method according to claim 3, wherein the quality specifications include at least two target quality levels corresponding to a minimum target quality and a maximum target quality, a quality of a representation generated on the basis of the minimum bit rate and the minimum resolution falls below the minimum target quality, and a quality of a representation generated on the basis of the maximum bit rate and the maximum resolution exceeds the maximum target quality.
5. The method according to any one of claims 1 to 4, wherein generating the second set of support points comprises: Creating a second grid of value pairs in a bitrate-resolution space containing value pairs of the first set, and Generate qualities for the value pairs of the second set based on the support points of the first set.
6. The method of claim 5, wherein generating qualities for the value pairs of the second set comprises at least one of the following: Interpolation of the support points, and / or Processing by a neural network, and / or a combination thereof.
7. The method according to claim 6, wherein the processing by a neural network comprises: Obtaining support points of the first set or an interpolation of support points of the first set as input data, Generation of output data comprising processing the input data by one or more layers of the neural network.
8. Method according to one of claims 6 or 7, wherein output data of the neural network are Filtering the output data to meet monotonicity conditions, and / or Limiting the range of values of the predicated qualities.
9. The method according to any one of claims 1 to 8, wherein the quality specifications contain at least two target quality levels corresponding to a minimum target quality and a maximum target quality, and selecting the subset of support points comprises: for each target quality level from the quality specifications, determining a bit rate for a bit rate specification of an encoder comprising Determining a bit rate for each resolution whose associated predicted quality meets the quality specifications for the respective target quality level, and Select the minimum bitrate from the specified bitrates as the bitrate default.
10. The method according to claim 9, wherein the determination of the bit rate for the bit rate specification comprises an interpolation based on the support points of the second set.
11. Method according to one of claims 9 or 10, wherein selecting the subset of support points further comprises Creating a representation comprising encoding the video portion with the respective bitrate specification for each target quality level from the quality specifications.
12. The method according to claim 11, further comprising Determination of the quality of the generated representation, Comparing the determined quality with the quality specifications; if the determined quality meets the quality specifications: adding the representation to the bitrate ladder; if the determined quality does not meet the quality specifications: determining a new representation based on a new bitrate specification.
13. A method for encoding representations of a video segment, comprising: Generating a bitrate ladder according to any one of claims 1 to 12, wherein the bitrate ladder includes two or more quality levels; for each of the quality levels of the bitrate ladder: generating a representation comprising encoding the video portion according to the respective quality level.
14. A computer program comprising: program instructions stored on a non-transferable, computer-readable medium that, when executed on one or more processors, cause the one or more processors to perform steps of any of methods 1 to 13. Apparatus for generating a bitrate ladder for encoding representations of a video section, comprising: a unit for determining a first set of support points, wherein a support point indicates a quality of a representation based on bitrate and resolution, and the quality is based on a comparison with an original representation; a unit for generating a second set of support points based on the first set of support points, wherein the second set contains more support points than the first set, a unit for selecting a subset of support points of the second set, taking quality specifications into account, for generating the bitrate ladder based on the subset of support points.Apparatus for encoding representations of a video portion, comprising: means for generating a bitrate ladder according to claim 15; a unit for generating a representation for each of the quality levels of the bitrate ladder, comprising encoding the video portion according to the respective quality level.