Prediction of rate distortion curves for video encoding
The dynamic selection of candidate average bitrates based on video characteristics addresses the issue of inconsistent quality and resource inefficiency in existing methods, improving video encoding by optimizing segment selection and maintaining consistent quality.
Patent Information
- Application Number
- JP2025134028
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-03
- Filing Date
- 2025-08-12
- Publication Date
- 2025-12-16
AI Technical Summary
Existing video encoding methods using static lists of candidate average bitrates fail to optimize video quality and resource utilization due to varying video characteristics, leading to suboptimal transcoding and inconsistent quality gaps.
A pre-analysis optimization process dynamically selects candidate average bitrates based on video characteristics, using rate-distortion curves and machine learning to generate an optimized list for each video portion, ensuring consistent quality and resource efficiency.
This approach enhances video quality by optimizing segment selection for adaptive bitrate streaming, minimizing storage and distribution footprint while maintaining consistent quality levels.
Smart Images

Figure 2025183209000001_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001]
[0001] This application is a continuation-in-part application, pursuant to 35 U.S.C. § 120, of "DYNAMIC SELECTION OF CANDIDATES" filed March 6, 2023. This application is entitled to and claims the benefit of an earlier filed application, U.S. Application No. 18 / 179,281, entitled "DYNAMIC SELECTION OF CANDIDATE BIT RATES FOR VIDEO ENCODING," the contents of which are incorporated herein by reference in their entirety for all purposes. [Background technology]
[0002] One method of delivering video to client devices uses adaptive bitrate streaming (ABR). Adaptive bitrate streaming is based on providing multiple streams (often called variants or profiles) encoded with different levels of video attributes, such as different bitrates and / or quality. A profile ladder lists different profiles available for a client to use when streaming segments of video. The client can dynamically select a profile based on network conditions and other factors. The video is segmented (e.g., divided into individual segments, typically several seconds in length), and the client can switch from one profile to another at segment boundaries when network conditions change. For example, when a network condition with higher available bandwidth is experienced, a video delivery system may want to provide the client with a profile with a higher bitrate that improves the quality of the video being streamed. When a network condition with lower available bandwidth is experienced, the video delivery system may want to provide the client with a profile with a lower bitrate so that the client can play the video without any playback issues, such as rebuffering or download failures. [Brief explanation of the drawings]
[0003] The included drawings are for illustrative purposes and merely serve to provide examples of possible structures and operations of the disclosed inventive systems, apparatus, methods, and computer program products. These drawings in no way limit any changes in form and detail that may be made by those skilled in the art without departing from the spirit and scope of the disclosed implementations. [Figure 1]
[0004] FIG. 1 illustrates a system for dynamically selecting a list of candidate average bit rates according to some embodiments. [Figure 2]
[0005] FIG. 1 illustrates an example of a portion of a video according to some embodiments. [Figure 3A]
[0006] FIG. 10 illustrates an example of generating a list of candidate average bit rates according to some embodiments. [Figure 3B]
[0007] FIG. 10 illustrates an example of generating encoded segments according to some embodiments. [Figure 4]
[0008] FIG. 10 illustrates an example of clustering encoded segments into multiple pools according to some embodiments. [Figure 5]
[0009] FIG. 10 illustrates an example of a selection process according to some embodiments. [Figure 6]
[0010] FIG. 10 illustrates an example of a graph of rate-distortion curves that may be used to select coded segments for pooling according to some embodiments. [Figure 7]
[0011] FIG. 10 illustrates an example of selected encoded segments for a per-segment profile according to some embodiments. [Figure 8]
[0012] 4A and 4B illustrate examples of different rate-distortion curves for video content according to some embodiments. [Figure 9]
[0013] 10A-10C illustrate different characteristics using different encoding configurations according to some embodiments. [Figure 10]
[0014] FIG. 10 illustrates an example of using static candidate average bit rates for different rate-distortion curves according to some embodiments. [Figure 11]
[0015] FIG. 10 illustrates an optimized candidate average bitrate list according to some embodiments. [Figure 12]
[0016] FIG. 1 illustrates a more detailed example of a segment quality-driven adaptive (SQA) system and pre-analysis optimization process according to some embodiments. [Figure 13]
[0017] FIG. 1 illustrates a more detailed example of a rate-distortion (RD) prediction system according to some embodiments. [Figure 14]
[0018] FIG. 10 illustrates the output of a predictive network according to some embodiments. [Figure 15]
[0019] FIG. 1 illustrates a simplified flowchart of a method for performing an optimization process for selecting a list of candidate average bit rates according to some embodiments. [Figure 16]
[0020] 4A and 4B illustrate an example of determining boundaries for a list of candidate average bit rates according to some embodiments. [Figure 17]
[0021] 10 illustrates an example of eliminating candidate average bit rates based on quality according to some embodiments. [Figure 18]
[0022] FIG. 10 illustrates an example in which a minimum gap is used to eliminate candidate average bit rates according to some embodiments. [Figure 19]
[0023] 6 is a graph illustrating when adding candidate average bit rates may be advantageous according to some embodiments. [Figure 20]
[0024] 10A and 10B illustrate the decision to add candidate average bit rates according to some embodiments. [Figure 21]
[0025] FIG. 1 illustrates an example of an RD prediction system according to some embodiments. [Figure 22]
[0026] FIG. 1 illustrates an example of features that may be extracted according to some embodiments. [Figure 23]
[0027] 1 illustrates an example of a frame of video according to some embodiments. [Figure 24]
[0028] FIG. 1 illustrates a simplified flowchart of a prediction method according to some embodiments. [Figure 25]
[0029] 10 is a graph listing bitrate quality values for one target resolution according to some embodiments. [Figure 26]
[0030] 10A and 10B are diagrams illustrating examples of direct prediction modes according to some embodiments. [Figure 27]
[0031] FIG. 1 illustrates an example of an indirect prediction mode system according to some embodiments. [Figure 28]
[0032] 1A and 1B illustrate examples of proxy and target encoding results according to some embodiments. [Figure 29]
[0033] FIG. 10 illustrates an example of using direct and indirect quality values, according to some embodiments. [Figure 30A]
[0034] 10 is a graph of a rate-distortion curve for a single resolution according to some embodiments. [Figure 30B]
[0035] Graph showing an RD map according to some embodiments. [Figure 31]
[0036] FIG. 1 illustrates a video streaming system that communicates with multiple client devices over one or more communication networks according to one embodiment. [Figure 32]
[0037] 1 is a schematic diagram of a device for viewing video content and advertisements. DETAILED DESCRIPTION OF THE INVENTION
[0004]
[0038]
[0013] Techniques for video distribution systems are described herein. In the following description, for purposes of explanation, numerous examples and specific details are set forth in order to provide a thorough understanding of some embodiments. Some embodiments, as defined by the claims, may include some or all of the features in these examples, alone or in combination with other features described below, and may further include modifications and equivalents of the features and concepts described herein.
[0005]
[0039] The system can adaptively generate a list of bitrates to be used to encode a video. The list of bitrates may be referred to as candidate average bitrates (CABs). The encoder transcodes segments of the video using each bitrate in the list of candidate average bitrates. In some embodiments, the system can dynamically select a list of candidate average bitrates for different portions of the video, such as for different chunks of the video. A chunk may be an independent coding unit that the encoder encodes using the same settings. A video may include one or more chunks, and each chunk may include multiple segments. In some embodiments, the list of candidate average bitrates may be set at the chunk level. Although the list of candidate average bitrates is described as being set at the chunk level, the list of candidate average bitrates may be set for different portions of the video.
[0006]
[0040] The encoder can encode segments of the video using bit rates in a list of candidate average bit rates to generate multiple candidate segments. A segment quality-driven adaptation (SQA) process can select a segment from the candidate segments to use for a profile in a profile ladder. The goal of the process is to optimize (e.g., minimize) the storage or distribution footprint of a portion of the video while maintaining similar quality.
[0007]
[0041] Each video may have different characteristics. Similarly, different portions of the same video may have different characteristics. Using a static list of candidate average bit rates for all portions of a video or for multiple videos may not provide optimal results. For example, a static list of candidate average bit rates may encode a video with simple video content with a bit rate higher than required. Also, a video with complex video content may be encoded at low quality due to insufficient bit rate. Furthermore, a static list of candidate average bit rates may generate segments with irregular quality gaps from an encoding perspective. For example, adjacent profiles may have similar video quality that is redundant with each other or may have an unacceptably large quality gap. Having similar video quality for adjacent profiles may be unnecessary and may not provide much benefit in viewing quality. For example, if two bit rates in the list of candidate average bit rates result in an encoded segment with similar quality, transcoding the segment using those two bit rates may be redundant and waste resources. Also, having a large quality gap may result in a poor viewing experience during playback, as the quality may change abruptly when playback switches from one profile to another.
[0008]
[0042] To overcome the above drawbacks, the pre-analysis optimization process can dynamically select bit rates within a list of candidate average bit rates for a video. To select the list of candidate average bit rates, the pre-analysis optimization process can analyze portions of the video and output an optimized list of candidate average bit rates for the portions. For example, the pre-analysis optimization process can analyze characteristics of each portion and output a list of candidate average bit rates for each portion. In some embodiments, the pre-analysis optimization process can predict characteristics of the portions, such as a rate-distortion curve that describes quality versus bit rate for the portion. The pre-analysis optimization process uses the respective rate-distortion curve to determine the optimal list of bit rates for each portion.
[0009]
[0043] The optimization process provides many advantages. For example, it provides an optimal selection of transcoded segments to select when selecting segments for a profile in the profile ladder. If the list of candidate average bit rates is set at a static value for the entire video and / or is the same for multiple different videos, suboptimal transcoding may occur. Different videos, and even different portions of the same video, may have diverse characteristics. Thus, a static list of candidate average bit rates may be suboptimal for some videos or portions of videos. The use of a dynamic list of candidate average bit rates based on the characteristics of portions of videos may result in higher quality video and viewing experiences, since the segment quality-driven adaptation process may have a better selection of encoded segments to select to form profiles for the profile ladder.
[0010]
[0044] system
[0045] FIG. 1 illustrates a system 100 for dynamically selecting a list of candidate average bit rates, according to some embodiments. The system 100 includes a content delivery network 102, a client 104, and a video delivery system 106. The source files may contain different types of content, such as video, audio, or other types of content information. Video may be used for illustrative purposes, but other types of content may be understood. In some embodiments, the source files may be received in a format that requires encoding into another format, as described below. For example, the source file may be a mezzanine file containing compressed video. The mezzanine file may be encoded to generate other files, such as different profiles of the video.
[0011]
[0046] A content provider may operate the video delivery system 106 to provide a content distribution service that allows entities to request and receive media content. The content provider may use the video delivery system 106 to coordinate the distribution of media content to the clients 104. While a single client 104 is described, multiple clients 104 may be using the service. The media content may be different types of content, such as on-demand video from a library of videos and live video. In some embodiments, live video may be where video is available based on a linear schedule. Video may also be provided on-demand. On-demand video may be content that can be requested at any time and is not limited to viewing on a linear schedule. The video may be a program, such as a movie, a show, or an advertisement.
[0012]
[0047] The clients 104 may include different computing devices such as smartphones, living room devices, televisions, set-top boxes, tablet devices, etc. The clients 104 include a media player 112 capable of playing content such as videos. In some embodiments, the media player 112 may receive segments of videos and play these segments. The clients 104 may send a request for a segment to one of the content delivery networks 102 and then receive the requested segment for playback on the media player 112. A segment may be a portion of a video, such as 6 seconds of a video.
[0013]
[0048] A video may be encoded into a profile ladder including multiple profiles. Each profile may correspond to a different configuration, which may be a different level of bitrate and / or quality, but may also include other characteristics such as codec type, computing resource type (e.g., computer processing unit), etc. Each video may have an associated profile with a different configuration. The profiles may be categorized into different levels, and each level may be associated with a different configuration. For example, a level may be a combination of bitrate, resolution, codec, etc., and each level may be associated with a different bitrate, such as 400 kilobytes per second (kbps), 650 kbps, 1000 kbps, 1500 kbps, ...12000 kbps. Each level may also be associated with another characteristic, such as a quality characteristic (e.g., resolution). Profile levels may be referred to as higher or lower, such that a profile with a higher bitrate or quality may be rated higher than a profile with a lower bitrate or quality. An encoder may use the characteristics to encode the source video. For example, an encoder may encode a source video at a target bitrate of 1500 kbps.
[0014]
[0049] The content delivery network 102 includes a server that can deliver video to the client 104. The content delivery network 102 receives requests for segments of video from the client 104 and delivers the segments of video to the client 104. The client 104 may request a segment of video from one of the profile levels based on the current playback state. The playback state may be any state experienced based on the playback of the video, such as available bandwidth, buffer length, etc. For example, the client 104 may use an adaptive bitrate algorithm to select a profile for the video based on the current available bandwidth, buffer length, or other playback state. The client 104 may continuously evaluate the current playback state and switch between profiles during playback of a segment of video. For example, during playback, the media player 112 may request a different profile for the video asset. For example, if low-bandwidth playback conditions are being experienced, the media player 112 may request a lower profile associated with a lower bitrate for the upcoming segment of the video. However, if playback conditions of higher available bandwidth are being experienced, the media player 112 may request a higher level profile associated with higher bandwidth for the upcoming segment of video.
[0015]
[0050] A segment quality-driven adaptive processing system (SQA system) 108 can encode segments using a list of candidate average bit rates. The SQA system 108 then selects segments for each profile using an optimization process. For example, the SQA system 108 can adaptively select segments with optimal bit rates for each profile in a profile ladder while maintaining similar quality levels. The SQA system 108 enables the system to maintain quality similar to or matching the target bit rate while minimizing the number of bits required to store or distribute the content.
[0016]
[0051] The pre-analysis optimization process 110 can dynamically generate a list of candidate average bit rates for the portions of the video. In some embodiments, the pre-analysis optimization process 110 can predict characteristics of each of the portions of the video, such as a rate-distortion curve. The pre-analysis optimization process 110 then selects candidate average bit rates for the portions of the video based on analyzing the characteristics of each of the portions of the video.
[0017]
[0052] The following first describes the segment quality-driven adaptive processing process, and then describes the dynamic selection of the list of candidate average bit rates in more detail.
[0018]
[0053] Segment Quality Driven Adaptation Process
[0054] As described above, the optimization process 110 can dynamically select a list of candidate average bit rates for portions of a video. The portions of a video may be different sizes. FIG. 2 shows an example of a portion of a video according to some embodiments. A video 200 may be divided into different portions at the segment level and the chunk level. In some embodiments, at 202, a video 200 may be divided into chunk-level portions. For example, multiple chunks, chunk_0, chunk_1, ..., chunk_m, may be included in the video 200. Each respective chunk may be divided into smaller portions, which may be called segments. For example, at 204, chunk_0 is divided into segments, segment_0, segment_1, ..., segment_n. Similarly, although not shown, chunk_1 may be divided into its own respective segments, segment_0, segment_1, ..., segment_n. A segment may be shorter in length than a chunk. For example, a chunk may be a two-minute video, and a segment may be a five-second video.
[0019]
[0055] In a segment quality-driven adaptation process, the SQA system 108 can process each segment of the video 200 to generate multiple encodings of each respective segment based on a list of candidate average bit rates. For illustrative purposes, the optimization process 110 selects a list of candidate average bit rates for each chunk, but lists of candidate average bit rates may be selected for different portion sizes, such as per segment, for multiple chunks, etc. The bit rates included in each respective list of candidate average bit rates can be optimized based on characteristics associated with each portion (e.g., chunk and / or segment) of the video using the list of candidate average bit rates. Given different characteristics for different chunks, the respective lists of candidate average bit rates may be different. However, it may be possible for the bit rates for multiple chunks in each respective list of candidate average bit rates to be the same.
[0020]
[0056] FIG. 3A illustrates an example of generating a list of candidate average bit rates according to some embodiments. Chunks chunk_0, chunk_1, chunk_2, ..., chunk_n are shown at 202. At 302, optimization system 110 has a list of candidate average bit rates selected for each chunk based on characteristics of each respective chunk. For example, for chunk_0, list #0 of candidate average bit rates is based on characteristics for chunk_0. Also, list #1 of candidate average bit rates is based on characteristics for chunk_1, and so on. In some examples, for chunk_0, list #0 of candidate average bit rates may include bit rates of 8500, 7750, 7000, 6250, 5500, 4750, 4000, and 3250 kilobytes per second (Kbps). For chunk_1, list #1 of candidate average bit rates may include bit rates of 7000, 6250, 4750, 4000, 3250, 2000, and 1250 Kbps.
[0021]
[0057] The list of candidate average bit rates may include bit rates to be used by the encoder to encode each segment. Traditionally, the candidate average bit rates may have statically included the same bit rates. Sometimes, two types of bit rates are used for all chunks. The first type may be a target average bit rate, and the second type may be an intermediate average bit rate. The target average bit rate may be a base bit rate associated with a profile in a profile ladder for adaptive bit rate encoding. The intermediate average bit rate may be supplemental to the target average bit rate. For example, additional bit rates between the target average bit rates may be added. The use of the intermediate average bit rate may provide additional bit rates for encoding additional encoding segments that may have different characteristics, such as quality, from the segments encoded from the target average bit rate. In some cases, the optimization process 110 may include bit rates from the target average bit rate and / or the intermediate average bit rate in the list of candidate average bit rates. For example, the optimization process 110 may include the target average bit rate in the list of candidate average bit rates, but may dynamically select other bit rates. In another example, the optimization process may dynamically select a bitrate in the list of candidate average bitrates based solely on the characteristics of the chunk.
[0022]
[0058] As described above, the encoder generates encoded segments for the chunks. Figure 3B shows an example of generating encoded segments according to some embodiments. At 204, the segments for chunks in chunk_0 are denoted segment_0, segment_1, segment_2, ..., segment_n. At 304, a list of candidate average bitrates (CAB list) for chunk_0 is used. In some embodiments, the same list of candidate average bitrates for chunk_0 is used for all segments of the chunk. However, multiple different lists of candidate average bitrates may be used for different segments of the chunk. The encoder then encodes the segments of chunk_0 using the list of candidate average bitrates.
[0023]
[0059] At 306, the encoded segments for each segment are listed. For each segment, the encoder encodes the segment using an average bit rate in the list of candidate average bit rates. The encoder can target each average bit rate when encoding the segment. This results in a set of encoded segments for each segment of the chunk, such as ENC_S0_CAB_0, ENC_S0_CAB_1, ENC_S0_CAB_2, ..., ENC_S0_CAB_n for segment_0. In the notation, ENC_S0 represents the encoded segment for segment_0, and CAB_0, CAB_1, CAB_2, etc. represent candidate average bit rates. For example, CAB_0 may be 8500 Kbps, CAB_1 may be 7750 Kbps, and CAB_2 may be 7000 Kbps. Each encoded segment may be encoded at the same quality level, such as 1080p. The process may be repeated for another quality level using the list of candidate average bit rates.
[0024]
[0060] For each segment, the optimization process 110 clusters the encoded segments into multiple pools. Each pool may correspond to one profile. FIG. 4 illustrates an example of clustering encoded segments into multiple pools according to some embodiments. At 402, multiple encoded segments are shown for candidate average bit rates. Each segment may have an associated value for a quality indicator. For example, encoded segment ENC_S0_CAB_0 may have a quality of quality_S0_c0, encoded segment ENC_S0_CAB_1 may have a quality of quality_S0_c1, etc. In the notation, quality_S0 represents the encoded segment of segment_0, and c0, c1, c2, etc. represent the quality for this encoded segment.
[0025]
[0061] Different methods may be used to include encoded segments in pools 401-1, 404-2, and 404-p. For example, each pool may have or be associated with a profile. Each profile may be associated with a target bitrate, which may be the maximum bitrate that can be used to encode segments for the associated profile. The SQA system 108 may include encoded segments starting with the highest average bitrate that can be used for the associated profile for the pool. The SQA system 108 may then add other encoded segments at other bitrates less than the maximum bitrate. This may result in different encoded segments being included in each pool. For example, pool S0_Pool_0 may include segments ENC_S0_CAB_0, ENC_S0_CAB_1, ENC_S0_CAB_2, etc. Pool S0_Pool_1 may also include encoded segments ENC_S0_CAB_2, ENC_S0_CAB_3, ENC_S0_CAB_4, etc. Thus, pool S0_Pool_1 may contain coded segments that start at a bit rate lower than the maximum bit rate in pool S0_Pool_0. If the coded segments are coded at bit rates of 8500, 7750, 7000, 6250, 5500, 4750, 4000, 3250 Kbps, pool S0_pool_0 may start with coded segments with average bit rates of 8500, 7750, 7000, etc., and pool S0_pool_1 may start with coded segments with average bit rates of 7000, 6250, 5500, etc. In some examples, exemplary bit rates for pools may be pool_0: 8500, 7700, 7000, 6250, 5500, 4750, pool_1: 7000, 6250, 5500, 4750, 4000, and pool_p: 5500, 4750, 4000, 3250.
[0026]
[0062] From each pool, the SQA system 108 may select one encoded segment based on using a selection process. FIG. 5 shows an example of a selection process according to some embodiments. The following process may be performed for each pool. At 404-1, pool S0_pool_0 from FIG. 4 is shown with its respective encoded segments. The SQA system 108 may use one or more rules to select an encoded segment for each pool. At 502, the SQA system 108 selects an encoded segment ENC_S0_CAB_1 for pool S0_pool_0. In some embodiments, the SQA system 108 may attempt to select an encoded segment with the minimum bitrate that has a quality value that meets a criterion. In some examples, the SQA system 108 may start with the first encoded segment in the pool, such as the segment with the highest bitrate. Then, the SQA system 108 selects a neighboring encoded segment in the pool, such as the encoded segment with the next highest bitrate. If the first and second encoded segments have similar quality (e.g., within a threshold), the SQA system 108 selects the encoded segment with the lowest bitrate. The SQA system 108 may continue the comparison using adjacent encoded segments in the pool, such as the second and third encoded segments. If the adjacent encoded segments do not have similar quality, the process may terminate. Other methods may be used, such as starting with the encoded segment with the lowest bitrate. The process may also select the segment with the lowest bitrate that has a quality within the threshold of another segment, such as the segment with the highest bitrate. The following describes an example of a process using a rate-distortion curve.
[0027]
[0063] 6 shows an example of a rate-distortion curve graph 600 that may be used to select coding segments for pooling according to some embodiments. In graph 600, the Y-axis is quality and the X-axis is bitrate. Curve 602 defines the relationship between quality and bitrate. For example, the curve may plot the rate and distortion of a segment or chunk, although the curve may also plot other characteristics of quality and bitrate.
[0028]
[0064] The coded segments may be listed as A, B, C, D, E, and F on the curve 602 based on their respective rates and distortions. At 604, an example of coded segments having similar quality is shown. In this case, coded segment C and coded segment D have similar bit rates and similar quality. For example, the quality difference between coded segment C and coded segment D may meet a threshold value min_gap (e.g., equal and / or smaller). In this case, the SQA system 108 may select coded segment D because the quality difference is minimal, since this coded segment has a lower bit rate compared to coded segment C, but segment D provides similar quality compared to segment C.
[0029]
[0065] The SQA system 108 can also collapse coded segments whose quality exceeds an upper boundary. For example, the upper boundary at 606 may be the boundary used to determine coded segments as candidates for collapse. In this case, the SQA system 108 may select one or more of the segments above the upper threshold, such as selecting only one segment (e.g., segment B), or selecting fewer segments found to exceed the upper threshold (e.g., selecting two of four segments). In another example, coded segments A and B may be removed. The SQA system 108 may also remove coded segments whose quality falls below a lower boundary. For example, a lower threshold is shown at 608. The SQA system 108 may select one or more of the segments below the lower threshold, such as selecting only one segment (e.g., segment F), or selecting fewer segments found to fall below the lower threshold. In another example, coded segments E and F may be removed. The upper and lower thresholds may be used to restrict segments for a profile that exceed or fall below a desired bit rate or quality. One reason for using an upper limit is to restrict the bit rate used to encode a segment, and one reason for using a lower limit is to restrict a bit rate from being too low. After processing the encoded segments to remove unencoded segments, the SQA system 108 may select a segment for the profile. For example, the SQA system 108 may select the encoded segment with the lowest bit rate that has a quality level that meets the threshold, such as within the gap with the highest-quality segment. In this case, the SQA system 108 may select encoded segment D.
[0030]
[0066] While the above rules may be used to select segments, other processes may also be used. For example, the selection of an encoded segment may be based on which encoded segments have been selected for other profiles. In some examples, the selected segment may be based on reducing the storage of encoded segments where a profile may reuse segments from other profiles. Thus, the SQA system 108 may optimize quality while minimizing the bitrate used for encoded segments that fall between the lower bound and the ceiling.
[0031]
[0067] FIG. 7 shows examples of selected encoded segments for each segment's profile, according to some embodiments. At 702, 704, 706, and 708, encoded segments are shown for Profile_0, Profile_1, Profile_2, and Profile_P, respectively. Within a profile, the SQA system 108 can select different encoded segments having different candidate average bit rates for different segments. For example, for Profile_0, Segment_0 was encoded using candidate average bit rate CAB_1, Segment_1 was encoded using candidate average bit rate CAB_0, Segment_2 was encoded using candidate average bit rate CAB_0, etc. In some examples, in Profile_0, Segment_0 was encoded using a bit rate of 7750 Kbps, Segment_1 was encoded using a bit rate of 8500 Kbps, and Segment_2 was encoded using a bit rate of 8500 Kbps. For profile_1, segment_0 was coded using CAB_4, segment_1 was coded using CAB_2, and segment_2 was coded using CAB_3. For example, for profile 1, segment_0 was coded using a bitrate of 5500 Kbps, segment_1 was coded using a bitrate of 7000 Kbps, and segment_2 was coded using a bitrate of 6250 Kbps.
[0032]
[0068] The following then describes an optimization process for dynamically generating a list of candidate average bit rates.
[0033]
[0069] Optimization Process
[0070] As mentioned above, video content may have diverse characteristics, such that content in different videos may have different characteristics, and content within the same video may also have different characteristics. For example, some content, such as cartoons or news, may be easy to encode. However, some content, such as live-action movies or sports, may be difficult to encode. The encoding characteristics may vary. The following describes different characteristics for content.
[0034]
[0071] 8 shows an example of different rate-distortion curves for video content according to some embodiments. The rate-distortion curve is used to indicate the relationship between quality and bitrate, although other metrics may be used to indicate the relationship between quality and bitrate for video content. Different rate-distortion curves may be shown for different chunks of video, although the rate-distortion curves may be different for different portions of the video, such as segments, chunks, multiple chunks, or different videos.
[0035]
[0072] Three chunks, chunk_A, chunk_B, and chunk_C, are shown along with graphs 802, 804, and 806 of rate-distortion curves for the chunks, respectively. In graph 802, quality varies steeply at lower bit rates, but does not vary much at higher bit rates. In graph 804, quality varies as the bit rate increases with a constant correlation. In graph 806, quality at lower bit rates may vary only minimally, while quality increases steeply at higher bit rates.
[0036]
[0073] In addition to different content generating different rate-distortion curves, different encoding configurations may also generate different encoding results. Different encoding configurations may include using different encoders (e.g., x264, x265, etc.) or different encoding parameters (e.g., rate-distortion optimization (RDO) level, B-frames, number of references, etc.). Figure 9 illustrates different characteristics using different encoding configurations according to some embodiments. For the same segment or chunk, a first encoding configuration in 902 results in different characteristics compared to a second encoding configuration shown in 904. Encoding configuration A results in a rate-distortion curve similar to chunk_A above, and encoding configuration B results in a rate-distortion curve similar to chunk_B above, even though these rate-distortion curves are for the same content.
[0037]
[0074] Considering that the rate-distortion curves may be different, using a static list of candidate average bit rates may not be optimal. For example, using the same list of candidate average bit rates for different rate-distortion curves may not provide optimal results. FIG. 10 shows an example of using static candidate average bit rates for different rate-distortion curves according to some embodiments. Graphs 802, 804, and 806 show different rate-distortion curves for the different chunks shown in FIG. 8. The dotted lines in each graph indicate different bit rates in the list of candidate average bit rates. Some problems may arise when using a fixed list of candidate average bit rates. For example, in graph 802, the two highest candidate average bit rates at 1008 may be redundant because they have similar quality to the third candidate average bit rate at 1010. That is, to provide an encoded segment with similar quality, only one bit rate, such as the bit rates listed at 1010, may need to be encoded.
[0038]
[0075] In graph 804, at 1012, two candidate average bit rates may be redundant because these two coded segments have similar quality compared to the coded segment with the next lowest bit rate shown at 1014. As above, only one bit rate may need to be coded, such as the lowest bit rate at 1014, to give coded segments with similar quality.
[0039]
[0076] In graph 806, the three lowest candidate average bit rates may produce encoded segments with similar quality at 1016. Also, the candidate average bit rates may be too far apart, so that the difference in quality between the encoded segments may be too large at 1018. That is, to minimize the difference in quality between the candidate average bit rates, it may be more desirable to have more candidate average bit rates with smaller quality differences.
[0040]
[0077] 11 illustrates an optimized candidate average bitrate list according to some embodiments. In graph 802, the SQA system 108 can dynamically select candidate average bitrates to optimize the quality observed in the encoded segment. For example, in 1102, the SQA system 108 can increase the number of candidate average bitrates at bitrates where the curve is steep. Also, in 1103, the SQA system 108 can decrease the number of candidate average bitrates where the curve does not change the quality much.
[0041]
[0078] In graph 804, at 1104, the SQA system 108 may remove candidate average bitrates from the lowest bitrate where quality may be redundant, and at 1106, the SQA system 108 may add additional bitrates to obtain varying quality at higher bitrates.
[0042]
[0079] In graph 806, at 1108, the SQA system 108 may eliminate bit rates at the lower end of the curve. Also, at 1110, the SQA system 108 may space the candidate average bit rates more evenly to capture different quality levels in more even increments.
[0043]
[0080] Pre-analysis optimization process design
[0081] 12 shows a more detailed example of the SQA system 108 and pre-analysis optimization process 110 according to some embodiments. A chunk to be encoded is received. An encoding configuration may also be received that defines settings for encoding the chunk. The encoding configuration may include an encoder type, a quality level, etc.
[0044]
[0082] The pre-analysis optimization process 110 can receive chunks and encoding configurations and output an optimized list of candidate average bit rates. The RD prediction system 1202 can predict rate-distortion curves for segments and / or chunks within the chunks. While predicting a rate-distortion curve for a segment or chunk may be described, rate-distortion curves may be generated for different portions of the video, such as for multiple chunks and / or multiple segments. As described in more detail below, the RD prediction system 1202 can use machine learning logic to generate a prediction of the rate-distortion curve for a segment.
[0045]
[0083] The predicted rate-distortion curve is output to the CAB list optimization system 1204. The CAB list optimization system 1204 can optimize a list of candidate average bit rates for the chunk based on, for example, the predicted rate-distortion curves for the segments within the chunk. The optimized list of candidate average bit rates may be based on the characteristics of each chunk and may be different for chunks with content having different characteristics. This process is described in more detail below.
[0046]
[0084] The CAB list optimization system 1204 outputs the optimized list of candidate average bit rates to the SQA system 108. The SQA system 108 includes an encoding system 1206 that receives the optimized list of encoding configurations, chunks, and candidate average bit rates. The encoding system 1206 then uses each candidate average bit rate in the list to encode each segment of the chunk. After encoding each segment using the list of candidate average bit rates, the selection system 1208 selects an encoded segment for each profile in the profile ladder using the selection process described above. The selection system 1208 outputs the selected encoded segments for the profiles in the profile ladder.
[0047]
[0085] The following describes the prediction of segment characteristics and then optimization to select a list of candidate average bit rates.
[0048]
[0086] RD prediction system
[0087] 13 shows a more detailed example of the RD prediction system 1202 according to some embodiments. A feature extraction system 1302 receives chunks of video. The feature extraction system 1302 may then extract values for features that may convey information related to video transcoding. Some examples of features may be about the video content, encoding settings, etc. The extracted features may provide a better prediction of the characteristics of the segments of the chunks. The values for the features are output to a prediction network 1304.
[0049]
[0088] The prediction network 1304 can use the trained model to generate characteristics for segments of chunks, such as predicted rate-distortion curves. The prediction network 1304 may use different machine learning algorithms, such as support vector machine (SVM) regression, convolutional neural networks (CNN), boosting, etc. The trained model can be trained based on a particular machine learning algorithm.
[0050]
[0089] The prediction network 1304 can receive values for the features in addition to other inputs such as a segment location, an encoding configuration, and a target bit rate. The segment location may be the segment location (e.g., that segment within the video) for which to generate a rate-distortion curve, the encoding configuration may include the configuration used to encode the segment, and the target bit rate may include an output bit rate range for the segment. The prediction network 1304 can output a rate-distortion curve for the segment between the output bit rate ranges based on the features.
[0051]
[0090] FIG. 14 illustrates the output of the prediction network 1304 according to some embodiments. At 204, the segments for the chunk include segment_0, segment_1, segment_2, ..., segment_n. A rate-distortion curve may be generated for each segment within each chunk of video. For example, at 1402, a rate-distortion curve is output for each segment. The rate-distortion curve for segment_0, the rate-distortion curve for segment_1, etc. are shown. Each rate-distortion curve is based on characteristics for the respective segment. A list of candidate average bitrates for the chunk may be generated based on the rate-distortion curves. Chunk-level rate-distortion curves may also be output.
[0052]
[0091] List of candidate average bitrate optimizations
[0092] 15 shows a simplified flowchart 1500 of a method for performing an optimization process for selecting a list of candidate average bit rates according to some embodiments. At 1502, the CAB list optimization system 1204 determines boundaries for the list of candidate average bit rates. For example, the boundaries may be maximum and minimum bit rates that may be used for the list of candidate average bit rates. Different methods may be used to determine the boundaries, and are described in more detail in FIG. 16.
[0053]
[0093] At 1504, the CAB list optimization system 1204 generates a list of potential candidate average bitrates with optimal bitrate allocations. In some embodiments, one list of potential candidate average bitrates is generated for the chunk based on the maximum and minimum bitrates determined in 502. The list of potential candidate average bitrates may be generated using different methods. One method may be to use a predetermined list that falls between the minimum and maximum bitrates. For example, the predetermined list may include bitrates from the target average bitrate and the median average bitrate. For example, a bitrate from a predetermined list within a minimum and maximum range may be used. Another method may determine the total number of potential candidate average bitrates and divide the bitrate range between the minimum and maximum bitrates into intervals. Different examples may be used, such as:
[0054]
number
[0055]
number
[0056]
number
[0057]
number
[0058] where interval_i is the interval value of i, interval_(i+1) is the interval value + 1, interval_(i+2) is the interval value + 2, and delta is a predetermined value.
[0059]
[0094] The total number of intervals may be set to a number, such as 10. The intervals for interval_i may be set based on the above method by dividing the range into total numbers. The CAB list optimization system 1204 then selects bit rates based on the interval values to divide the range of bit rates between the minimum and maximum bit rates into a list of bit rates. For example, a minimum bit rate of 2000 and a maximum bit rate of 10,000 with an interval of 1500 and a total number of bit rates of 5 may result in a list of bit rates of 10,000, 7500, 5000, 3500, and 2000 when using equal division.
[0060]
[0095] At 1506, the CAB list optimization system 1204 refines the list of potential candidate average bit rates using an optimal quality allocation to generate an optimized list of candidate average bit rates. The quality allocation may examine the quality for each segment and determine whether the quality satisfies one or more rules. For example, redundant candidate average bit rates, such as candidate average bit rates with similar quality, may be eliminated. Additionally, additional candidate average bit rates may be added as needed, such as when neighboring candidate average bit rates have a quality gap that exceeds a threshold, such as a too large difference. The process is described in more detail in Figures 17, 18, and 19.
[0061]
[0096] As illustrated in 1502 of FIG. 15, the CAB list optimization system 1204 determines boundaries for a list of candidate average bit rates. FIG. 16 shows an example of determining boundaries for a list of candidate average bit rates according to some embodiments. The following process is described, but other processes may be understood. For example, settings may be used to determine minimum and maximum bit rates. In this example, the CAB list optimization system 1204 may analyze the minimum and maximum bit rates for the rate-distortion curve for each segment in the chunk and determine what the minimum and maximum bit rates should be at the chunk level.
[0062]
[0097] At 1602, the rate-distortion curves for each segment are received and analyzed. The CAB list optimization system 1204 may then select a minimum bitrate and a maximum bitrate for each segment based on the respective rate-distortion curve for the segment. For example, for segment_0, the minimum and maximum bitrates may be selected based on characteristics of the rate-distortion curve for segment_0. For example, the CAB list optimization system 1204 may set a maximum quality threshold and a minimum quality threshold and use the rate-distortion curve to determine a minimum bitrate corresponding to the minimum quality threshold and a maximum bitrate corresponding to the maximum quality threshold. For segment_1, the CAB list optimization system 1204 selects a minimum and maximum bitrate based on characteristics of the rate-distortion curve for segment_1, and so on.
[0063]
[0098] The above analysis was performed at the segment level. The CAB list optimization system 1204 then analyzes the segment-level results to determine minimum and maximum values at the chunk level. At 1606, the CAB list optimization system 1204 determines a maximum value from among the values for the maximum bitrate for the segment, from max_bitrate_0, max_bitrate_1, max_bitrate_2, ..., max_bitrate_n, etc. The CAB list optimization system 1204 also determines a minimum value from among the values for the minimum bitrate for the segment, from among min_bitrate_0, min_bitrate_1, min_bitrate_2, ..., min_bitrate_n, etc.
[0064]
[0099] At 1608, the CAB list optimization system 1204 outputs minimum and maximum bit rates for the chunk. In this case, the lowest minimum bit rate is selected from the minimum bit rates for the segment, and the highest maximum bit rate is selected from the maximum bit rates for the segment. The selection process may take into account the individual characteristics of the rate-distortion curve for the segment and select minimum and maximum bit rates that may include all of the minimum and maximum bit rates determined at the segment level. For example, if the minimum bit rates are 2000, 3000, and 3500, the minimum bit rate selected will be 2000. Similarly, if the maximum bit rates are 10000, 9000, and 8500, the maximum bit rate selected will be 10000. While the above process may be used, other methods of selecting minimum and maximum bit rates may be understood, such as taking an average of the values.
[0065]
[0100] As illustrated in 1506 of FIG. 15, the CAB list optimization system 1204 defines a list of candidate average bit rates with optimal quality allocations. Part of the allocation includes eliminating candidate average bit rates based on similar quality. Quality similarity may be defined in different ways. For example, the CAB list optimization system 1204 determines the distance between quality values to determine whether some candidate average bit rates should be eliminated. FIG. 17 shows an example of eliminating candidate average bit rates based on quality according to some embodiments. In this example, the CAB list optimization system 1204 may determine whether the quality levels of two adjacent candidate average bit rates meet a threshold, such as being within the threshold. The candidate with the higher bit rate may then be eliminated.
[0066]
[0101] At 1702, each segment may have an associated potential removal list of encoded segments that may potentially be removed. As shown, for segment_0, the CAB list optimization system 1204 has determined that the candidate average bit rates of S0_CAB_0, S0_CAB_3, and S0_CAB_4 may be removed. These candidate average bit rates may be removed because the encoded segments may have similar quality levels that meet the threshold with respect to adjacent encoded segments. Similarly, for segment_1, the CAB list optimization system 1204 has determined that the candidate average bit rates of S1_CAB_0 and S1_CAB_2 may be removed, and for segment_n, the CAB list optimization system 1204 has determined that the candidate average bit rates for Sn_CAB_0 and Sn_CAB_3 may be removed. The segment is not removed for segment_2 because the segments are not determined to have similar quality levels within the threshold.
[0067]
[0102] The above analysis was at the segment level. Then, at 1704, the CAB list optimization system 1204 can use the segment-level candidate average bit rates to determine candidate average bit rates to be removed at the chunk level. For example, the CAB list optimization system 1204 may select a candidate average bit rate for the chunk level based on the occurrence of the candidate average bit rate in different segments of the potential removal list. In some embodiments, the CAB list optimization system 1204 can select a candidate average bit rate and calculate the total number of occurrences in the potential removed candidate pool. If the total number of this candidate average bit rate meets a threshold, such as at or above the threshold, the CAB list optimization system 1204 places this candidate average bit rate in the removal candidate list at the chunk level. For example, the candidate average bit rate CAB_0 is found in three of the above-mentioned segments (e.g., segment_0, segment_1, and segment_n), meeting the threshold of "3." The CAB list optimization system 1204 then places the candidate average bit rate of CAB_0 in the removal candidate list. Candidate average bit rates CAB_2, CAB_3, and CAB_4 may not meet the threshold because the bit rates occur in two or fewer segments in the potential removal list. Therefore, the CAB list optimization system 1204 does not include these candidate average bit rates in the removal candidate list. Other methods of selecting which candidate average bit rates to remove may be understood.
[0068]
[0103] The above analysis was performed at the segment level and merged to the chunk level. However, the process may be performed at a different level. For example, the analysis may be used to merge candidate average bitrates from multiple chunks to a portion of the video covering multiple chunk levels, or from multiple chunks to the video level.
[0069]
[0104] The following describes an example of removing candidate average bit rates. Figure 18 shows an example of using a minimum gap to remove candidate average bit rates according to some embodiments. For example, in 1802, candidate average bit rates C and D have similar quality levels that satisfy a threshold, such as threshold min_gap. In this case, this candidate average bit rate is adjacent to candidate average bit rate D, and candidate average bit rate C has a larger bit rate than candidate average bit rate D but has a minimum quality advantage, so the CAB list optimization system 1204 determines that one of the candidate average bit rates, such as candidate average bit rate C, should be removed. In this case, the CAB list optimization system 1204 may compare the difference between the quality values for the adjacent candidate average bit rates with threshold min_gap and remove one of the candidate average bit rates when the threshold is met.
[0070]
[0105] Another part of quality allocation involves adding candidate average bit rates based on gaps in quality. Figure 19 shows a graph 1900 illustrating when it may be advantageous to add candidate average bit rates according to some embodiments. The CAB list optimization system 1204 may use a threshold, such as the maximum gap max_gap in 1902, to determine when to add a candidate average bit rate. For example, if there is a gap in quality values between adjacent candidate average bit rates that is greater than the threshold max_gap, such as between candidate average bit rates C and D in graph 1900, the CAB list optimization system 1204 may add a candidate average bit rate between candidate average bit rates C and D on the rate-distortion curve.
[0071]
[0106] Different methods may be used to determine how many new candidate average bit rates should be added. The CAB list optimization system 1204 may add "i" new candidates when two candidates are separated by a ratio-based threshold. For example, different ratios may constitute the gap between the added candidates, such as 1:1, meaning each gap is equal, 1:1.5, meaning each gap is 1.5, or other ratios.
[0072]
[0107] In one possible process, the variable i is set to i=1, and the CAB list optimization system 1204 adds i new candidate average bit rates based on the ratio. For example, one candidate average bit rate, labeled "F," may be added between points C and D. Then, if all gaps between the new adjacent candidate average bit rates are smaller than the threshold max_gap, the process ends. However, if this is not the case, the value of the variable i is incremented, such as by "2," and two new candidates are added between the candidate average bit rates based on the ratio. For example, two or more candidate average bit rates may be added between points C and F and between points F and D. Then, the process continues as described above. Once the candidate average bit rates have been added such that there is no gap between the candidate average bit rates D and C that is larger than the threshold, the candidate average bit rate is output.
[0073]
[0108] The above process is determined for each segment. The CAB list optimization system 1204 may then take the potential added candidate average bit rates at the segment level and merge the candidate average bit rates at the chunk level. Figure 20 illustrates the determination of adding candidate average bit rates according to some embodiments. In 2002, each segment may have candidate average bit rates that can potentially be added. For segment_0, two candidate average bit rates may be added midway between the candidate average bit rates CAB_0 and CAB_1. Also, one candidate average bit rate may be added between CAB_4 and CAB_5. For segment_1, two candidate average bit rates may be added between the candidate average bit rates CAB_0 and CAB_1. For segment_n, two candidate average bit rates may be added between the candidate average bit rates CAB_0 and CAB_1. Thus, three segments add two candidate average bit rates between candidate average bit rates CAB_0 and CAB_1, and one segment adds one candidate average bit rate between CAB_4 and CAB_5.
[0074]
[0109] The CAB list optimization system 1204 may use the segment-level candidates to determine an additional candidate list for the chunk. For example, to be added to the additional candidate list at the chunk level, the CAB list optimization system 1204 may determine whether a potential added candidate average bit rate at the segment level is found within a threshold, such as the number of segments. If the threshold is 70% of the segments, the CAB list optimization system 1204 adds two candidates between candidate average bit rates CAB_0 and CAB_1 because these candidates are found in more than 70% of the segments (three out of four segments). The CAB list optimization system 1204 does not add a candidate between CAB_4 and CAB_5 because this addition is found only in segment_0, which is less than 70% of the segments. In this case, adding an additional candidate average bit rate may not be needed because only one segment requires addition, and adding candidate average bit rates for all other segments of the chunk when only one segment is affected may not be useful. However, adding two candidate average bit rates between candidate average bit rates CAB_0 and CAB_1 may be beneficial since over 70% of the segments had potential additions.
[0075]
[0110] The output of the CAB list optimization system 1204 is a list of candidate average bit rates per chunk. For example, a list of candidate average bit rates per chunk, such as that depicted in FIG. 3A, is output by the CAB list optimization system 1204 according to some embodiments.
[0076]
[0111] conclusion
[0112] Therefore, the list of candidate average bit rates can be optimized based on the characteristics found in each segment. This produces an improved list of candidate average bit rates for each chunk that is optimized for the characteristics of each chunk. The candidate average bit rates can improve the quality of the selection of encoded segments available for each chunk profile selection. This can improve the quality of the video in addition to improving the playback experience.
[0077]
[0113] Rate-Distortion Curve Prediction
[0114] The RD prediction system 1202 can generate a prediction of the relationship between bitrate and quality for a video. As described above, the relationship may be referred to as a rate-distortion curve (RD curve). The rate-distortion curve may be predicted for different target configurations. For example, for adaptive bitrate video transcoding, a high-resolution source video may be converted and encoded into multiple resolutions (e.g., 4K (3840x2160 pixels), 1080p (1920x1080), 720p (1280x720), 360p (480x360), etc.). For each resolution, the RD prediction system 1202 can predict a rate-distortion curve. For example, if the resolutions include 1080p, 720p, and 360p, the RD prediction system 1202 may generate three rate-distortion curves at different bitrates for each of the three resolutions. The multiple rate-distortion curves at different resolutions may be referred to as a rate-distortion map. Although three rate-distortion curves are described, more rate-distortion curves may be needed for the video. In addition to more resolutions, there may be multiple target configurations that require new rate-distortion maps, such that new rate-distortion maps are needed for different combinations of settings for encoding the video. If there are two encoder types, encoder #1 and encoder #2, the RD prediction system 1202 may generate different target configurations for each encoder. Then, for each encoder, the RD prediction system 1202 may generate three rate-distortion curves for a total of six rate-distortion curves.
[0078]
[0115] Conventionally, rate-distortion maps can be used in different ways, such as to design an optimal transcoding system for generating encoded bitstreams for adaptive bitrate systems. Conventionally, rate-distortion maps can only be obtained through multiple actual encodings of a given video. For example, one encoding job can generate one quality value for one bitrate and resolution pair. Multiple encoding jobs must be performed at each bitrate to generate a rate-distortion curve for a target configuration. When there are multiple target resolutions and target configurations, multiple encoding jobs must be performed to generate a rate-distortion curve. Therefore, the cost of processing time and computational resources to generate a rate-distortion curve is very high. When a service is transcoding multiple videos, the processing time and computational resources required to generate a rate-distortion curve may be impractical.
[0079]
[0116] In some embodiments, the RD prediction system 1202 predicts one or more rate-distortion maps for the video. This improves the use of computational resources in that an actual encoding job may not need to be performed to generate the rate-distortion curves for each target configuration and resolution. Furthermore, the prediction may be generated faster than performing the actual encoding job. The prediction may also be improved, such as by using proxy encoding information. The proxy encoding information may use encoding results from the actual encoding to generate the prediction. However, the proxy encoding information may be collected from an encoding configuration that is different from the target configuration used in the prediction, such as by using a fast preset setting when encoding the video. The fast preset setting may include simplifications to allow the encoding to be performed faster, such as a smaller resolution of the input video, a lower target bit rate, thinned video frames, etc. Proxy encoding information is described in more detail below.
[0080]
[0117] System Overview
[0118] FIG. 21 illustrates an example of an RD prediction system 1202 according to some embodiments. The RD prediction system 1202 may generate a rate-distortion curve for the rate-distortion map. As mentioned above, the rate-distortion map may be important because video content may have different characteristics. For example, some content may be easy to encode, such as cartoons or news, while some content may be difficult to encode, such as movies or sports. The rate-distortion curves for different content may differ, as discussed above in FIG. 8, which illustrates different rate-distortion curves for different chunks of video. Also, FIG. 9 illustrates different rate-distortion curves for different encoding configurations for the same segment. FIGS. 10 and 11 illustrate the benefits of using the rate-distortion curve to select a bitrate for an adaptive bitrate algorithm. As mentioned above, the SQA system 108 can dynamically select candidate average bitrates to optimize the quality found in the encoded segment. For example, the SQA system 108 may increase the number of candidate average bitrates at bitrates where the curve is steep, decrease the number of candidate average bitrates where the curve does not change quality much, eliminate candidate average bitrates from the lowest bitrates where quality may be redundant, add additional bitrates to capture changing quality at higher bitrates, eliminate bitrates at the lower end of the curve, and space candidate average bitrates more evenly to achieve different levels of quality in more equal increments. While rate-distortion curves may be used in adaptive bitrate algorithms, prediction of quality values may also be used for other purposes. In some embodiments, rate-distortion curve prediction may be used in different coding optimization systems. For example, in per-title and per-segment coding, the system can predict rate-distortion curves for the title or segment level of a video. The system can determine a dynamic target bitrate for higher quality at the same bitrate, or a lower bitrate at the same quality, for the title or segment.For encoding parameter optimization, the system may use different encoding parameters to predict different rate-distortion curves. The system may then select appropriate encoding parameters for several videos. When generating adaptive profile ladders for different profiles (e.g., bit rates and resolutions), the system may predict rate-distortion curves for different resolutions and bit rates. The system may then select different resolution and bit rate groups based on the rate-distortion curves to optimize the bit rates and resolutions within the groups to set adaptive profile ladders for different videos. The RD prediction system 1202 may receive video, such as frames of video, and generate rate-distortion curves for a rate-distortion map. The feature extraction system 2102 may receive frames of video. The feature extraction system 2102 may then extract values for a list of features based on characteristics of each frame of the video. The feature integration system 2104 may integrate frame-level features into portions of the video. For example, as described above, a segment may be a portion of a video, such as 6 seconds or multiple frames of video. The feature integration system 2104 may integrate frame-level features into segment-level features that describe features at the segment level. The start and end frames of a segment may be determined in different ways. For example, settings defining the segment may be received, or the segment may be dynamically determined by analyzing characteristics of the video to generate the segment. Although segment-level characteristics are described, other parts of the video, such as frame-level, chunk-level, etc., may also be used.
[0081]
[0119] The prediction network 2106 may receive segment-level features and a target configuration. The target configuration may include different combinations of parameters. For example, the parameters of the target configuration may include a target start frame / end frame, a target resolution, a target bit rate, a target quality indicator, and a target encoder. For example, the parameters may include a target resolution (640x360, 1280x720, 1920x1080, etc.), a target bit rate (500kbps, 1Mbps, 3Mbps, 6Mbps, etc.), a target image quality indicator (PSNR, VMAF, EPS, etc.), a target encoding configuration (e.g., a target encoder (AVC, HEVC, AV1, etc.), and a target encoding setting (fast preset, slow preset, etc.)). Different combinations of parameters may be generated. For example, the first target configuration may be a target resolution (640x360, 1280x720, 1920x1080, etc.), a target bitrate (500kbps, 1Mbps, 3Mbps, 6Mbps, etc.), a target quality index (PSNR), a target encoding configuration (e.g., a target encoder (AVC)), and a target encoding setting (high-speed preset). The second target configuration may be a target resolution (640x360, 1280x720, 1920x1080, etc.), a target bitrate (500kbps, 1Mbps, 3Mbps, 6Mbps, etc.), a target quality index (VMAF), a target encoding configuration (e.g., a target encoder (HEVC), a target encoding setting (low-speed preset). Although each combination may be listed as a target configuration, a list of possible parameter settings may be received, and then the RD prediction system 1202 generates different combinations. Using the first target configuration, the prediction network 2106 may predict a list of quality values (PSNR) for multiple bitrates (500 kbps, 1 Mbps, 3 Mbps, 6 Mbps, etc.) at each resolution (640x360, 1280x720, 1920x1080, etc.) for the AVC encoder using faster presets.Using the second target configuration, the prediction network 2106 may predict a list of quality values (VMAF) for multiple bitrates (500 kbps, 1 Mbps, 3 Mbps, 6 Mbps, etc.) at each resolution (640 x 360, 1280 x 720, 1920 x 1080, etc.) for the HEVC encoder using a slower preset. The output result may be multiple quality values for the bitrates. The RD map generator 2108 may generate rate-distortion curves from the quality values and then generate a rate-distortion map from the rate-distortion curves. The rate-distortion map may be for one target configuration. If multiple target configurations are being processed, the RD prediction system 1202 may generate a rate-distortion map for each target configuration.
[0082]
[0120] The following now describes different parts of the RD prediction system 1202 in more detail.
[0083]
[0121] Feature Extraction
[0122] The feature extraction system 2102 can extract different types of features for frames of video. In some embodiments, the feature extraction system 2102 can extract features associated with computer vision features, spatial domain features, time domain features, frequency domain features, and proxy coding features, although other features may also be used. FIG. 22 shows an example of features that may be extracted according to some embodiments. At 2202, the computer vision features may include different features related to the visual characteristics of the frame. The computer vision features may describe some detailed information about the frame. For example, the Sobel gradient means that if the value is higher, the frame is more complex, and the encoder may use more bitrate to encode the frame. Therefore, if the value is lower, the encoding bitrate will be higher and the quality will be lower, and vice versa. For example, features may be used based on the mean, variance, and histogram of pixel values, the gradient of Sobel and Laplace operations, blur intensity, noise intensity, etc. The values may be organized at the pixel level, block level, frame level, etc.
[0084]
[0123] At 2204, spatial domain features can analyze content differences, such as similarity and / or redundancy, of frame content. The spatial domain features can describe frame similarity and redundancy, such that if the content within a frame has a lot of redundancy and similarity, the coding bit rate can be lower and the quality can be higher, and vice versa. The features can be based on intra-prediction to calculate sum of absolute differences (SAD) to determine the similarity or redundancy of frame content. The features can be organized differently, such as by different block sizes, such as 4x4, 8x8, 16x16, etc.
[0085]
[0124] In 2206, the time-domain features may be based on features associated with multiple frames. The time-domain features may describe the speed and complexity of motion of adjacent frames; thus, if the content in these frames moves slowly and predictably, the coding bit rate may be lower and the quality may be higher, and vice versa. For example, the time-domain features may be based on the similarity, speed, and complexity of motion of adjacent frames. The feature extraction system 2102 may use inter-prediction to calculate motion vectors (MVs) and sums of absolute differences between objects within frames to convey the similarity, motion, speed, or complexity of motion of adjacent frames. Features may also be organized by different block sizes.
[0086]
[0125] In 2208, frequency domain features may be based on frequency domain information of the frame. The frequency domain features may describe frequency domain information of the frame, which is another view of the frame. If a frame is complex in the frequency domain, the encoder may use more bit rate to encode the frame, and therefore the encoding bit rate may be higher and the quality may be lower, and vice versa. The feature extraction system 2102 may use different frequency domain information, such as transform coefficients of the discrete cosine transform (DCT) / discrete sine transform (DST). The features may also be organized by block size.
[0087]
[0126] At 2210, the proxy coding features can be based on the actual coding of the video. The proxy coding features have an association with the coding result and a positive correlation with the prediction of the video coding result. The configuration used to perform the actual coding can differ from the target configuration being processed for prediction. In some embodiments, the proxy coding features can be determined based on lower computing resource consumption coding compared to the target configuration. In other embodiments, a faster preset of the encoder can be used to generate the proxy coding. The encoder can have different presets such as fast, slow, medium, etc. that use different amounts of computing resources (e.g., slow may use more computing resources but produce higher quality coding). The proxy coding can also have other simplifications, such as a smaller resolution of the input video, a lower target bitrate, thinned video frames (e.g., fewer video frames), a different encoder, etc. Thus, the proxy coding configuration may not be an exact copy of the target configuration, but may be designed to run faster. Proxy coding functions can also be used for multiple target configurations. In some embodiments, one proxy coding is performed and used in predicting multiple target configurations. The proxy encoding results may be used to generate proxy encoding characteristics. Some examples of proxy encoding characteristics include frame type quantization parameters, quality, bit rate, etc. resulting from the actual proxy encoding.
[0088]
[0127] Feature Integration
[0128] The feature integration system 2104 can integrate frame-level features into segment-level features. As mentioned above, segment-level features can be processed, but this step may not be necessary if frame-level features are used for prediction. Figure 23 shows an example of frames of a video according to some embodiments. At 2302-i to 2302-i+1, 2302-i+j, a segment X including frame_i, frame_i+1, and frame_i+j is shown. Each frame can be associated with multiple features in the segment at 2304, such as feature_0_i, feature_1_i, feature_2_i, and feature_M_i for frame_i of the frames.
[0089]
[0129] The feature integration system 2104 may integrate each feature to generate segment-level features at feature_0_output, feature_1_output, feature_2_output, and feature_M_output 2306. Different methods may be used to determine the segment-level features. In some embodiments, an average value of each feature may be calculated from each frame of the segment. For example, the feature integration system 2104 may generate an average value of the feature value of feature_0 for each frame. Average values are similarly calculated for the other features.
[0090]
[0130] In another example, before the average is calculated, the feature integration system 2104 may remove some values from some of the frames that meet a threshold. In some embodiments, some outliers from the frames may be removed. For example, one frame may have a feature value that may be significantly different (e.g., above a threshold) from the features from other frames, skewing the segment-level value so that it does not represent most of the frames of the segment. Different methods of calculating outliers may be used. In some embodiments, the average of the feature for the frames of the segment may be calculated, and the standard error of the feature is calculated. A score comparing the average to the standard error may be used to determine whether the score of the feature is an outlier. If the score meets the threshold, the feature may be removed. If the threshold is not met, the score may not be removed. The feature integration system 2104 may then calculate the average of the feature based on the list of unremoved features.
[0091]
[0131] The result of integrating the frame-level features to the segment level may be a per-feature average. Although an average is described, other methods of integrating or combining features may be used, such as using the median value from the frame. After integrating the features at the segment level, a rate-distortion map prediction may be generated.
[0092]
[0132] prediction
[0133] The prediction network 2106 can predict a list of quality values for multiple bitrates based on the target configuration. The prediction network 2106 can use segment-level features from the feature integration system 2104 and the target configuration to output a list of predicted quality values. The prediction network 2106 can use one or more models that can be trained to perform the prediction. In some embodiments, the prediction network 2106 may predict quality values for each bitrate of the target resolution. For example, for a target resolution of 640x360, the prediction network 2106 predicts quality values for target bitrates of 500 Kbps, 1 Mbps, 3 Mbps, 6 Mbps, etc. Then, for a target resolution of 1280x720, the prediction network 2106 predicts quality values for target bitrates of 500 Kbps, 1 Mbps, 3 Mbps, 6 Mbps, etc. If multiple target configurations are used, the following process can be performed for each target configuration.
[0093]
[0134] FIG. 24 shows a simplified flowchart 2400 of a prediction method according to some embodiments. The following process may be performed for each segment of a video. That is, a quality value for a rate-distortion map is generated for each segment. At 2402, the RD prediction system 1202 constructs a model based on the target configuration. For example, each target configuration may be associated with a model capable of predicting a quality value. In other embodiments, a single model may predict quality values for multiple target configurations. Different methods, such as supervised training and unsupervised training, may be used to construct the model.
[0094]
[0135] After constructing the model, a quality value may be predicted. Different methods for generating quality values are described at least in FIG. 26 and FIG. 27, although other methods may be used. At 2404, the RD prediction system 1202 determines whether the last target resolution has been processed. For example, target resolutions may include 640×360, 1280×720, 1920×1080, etc. If the last target resolution has been processed, the process may end. If the last target resolution has not been processed, at 2406, the RD prediction system 1202 determines whether the last bitrate has been processed. As described above, the prediction network 2106 may predict quality values for multiple bitrates for each target resolution.
[0095]
[0136] If the last bitrate has not been processed, at 2408, the RD prediction system 1202 predicts a quality value for the current target resolution and current bitrate using the model of the prediction network 2106. The prediction can receive features and target configuration for the segment and output a quality value for the bitrate. For example, the prediction can be a quality value for a target resolution of 640x360 and a target bitrate of 500 kbps.
[0096]
[0137] At 2410, the RD prediction system 1202 moves to the next target bitrate for the resolution. For example, after a bitrate of 500 kbps, the next target bitrate may be 1 Mbps. The process then returns to 2406. For each bitrate, the prediction network 2106 predicts a quality value. For example, other bitrates may be 3 Mbps, 6 Mbps, etc. When the last bitrate has been predicted, the process moves to 2412, where the next target resolution is processed. For example, another target resolution may be 1280x720. The process then returns to 2404, where the RD prediction system 1202 determines whether the last target resolution has been processed. If not, the process continues to process bitrates for the new target resolution. The same bitrates as described above may be used. When all target resolutions have been processed, such as after the target resolution of 1920x1080 has been processed, the process ends.
[0097]
[0138] The prediction may predict a quality value for each resolution and bitrate pair instead of a rate-distortion curve. Predicting all the details to generate a rate-distortion curve may require predicting many detailed information, such as some parts of the curve increasing faster and some parts increasing slower, the slope of parts of the curve being different, or the shape of the curve being completely different. Therefore, predicting a rate-distortion curve directly may be difficult. However, the RD prediction system 1202 may use more detailed information to generate points on the rate-distortion curve, which may then be used to estimate the curve. However, the prediction may predict a quality measurement for a rate-distortion curve based on a single input.
[0098]
[0139] 25 shows a graph 2500 listing quality values of bitrate for one target resolution, according to some embodiments. Graph 2500 provides a list of quality values for a single target configuration at only one resolution. The X-axis may be bitrate and the Y-axis may be quality. Each point in graph 2500 may represent a quality value output by prediction network 2106. For example, point 2502 is a quality value prediction for a first bitrate, and point 2504 is a quality value for a second bitrate. Multiple resolutions may have their associated points as shown in FIG. 25.
[0099]
[0140] The prediction of the quality value for the current target resolution and bitrate as described in 2408 may be performed in different ways. The following describes two methods: direct prediction mode and indirect prediction mode.
[0100]
[0141] Direct Prediction Mode
[0142] FIG. 26 shows an example of a direct prediction mode according to some embodiments. At 2602, features of segment x are input to a prediction network 2106. Additionally, a target configuration is input to the prediction network 2106. The features used may include those described in FIG. 22: computer vision features, spatial domain features, time domain features, frequency domain features, and / or proxy coding features. The prediction network 2106 may be trained to generate predictions based on the feature input. For example, based on values for the features, the prediction network 2106 may generate a quality value. The proxy coding features are used as input along with other features, and the prediction network 2106 is trained to output a quality value based on the values of the proxy coding features and the other features. The output quality value is for the target configuration. As described in more detail below, the proxy coding features may be used differently in the indirect prediction mode by predicting a quality value for the proxy coding configuration.
[0101]
[0143] The prediction network 2106 outputs a predicted quality value for each bit rate. In some embodiments, prediction may be performed for each bit rate and resolution pair. Using proxy coding features to determine the prediction may improve the quality value. For example, using several actual coding results may provide better information than simply using features not based on the actual coding. Proxy coding results may have a strong correlation with accurate prediction of the rate-distortion curve. Having several points from the actual coding may provide a generated shape for the rate-distortion curve, and predictions based on those actual points may be improved when given guidance from the actual coding results. In Figure 25, there were eight quality values. The prediction network 2106 may be implemented to generate eight different quality values using eight different bit rates for one resolution.
[0102]
[0144] Indirect Prediction Mode
[0145] The indirect prediction mode can predict a quality value using a quality offset prediction. The quality offset can be based on the difference between the proxy coding prediction and the target coding prediction. The proxy coding prediction can be a quality value based on the proxy coding configuration. The target coding prediction can be a target quality value based on the difference between the proxy coding configuration and the target configuration. For example, the proxy coding configuration can include a fast setting, while the target configuration can include a normal setting. The process can determine an offset for adjusting the proxy coding prediction based on the difference between the proxy coding configuration and the target configuration. The target coding prediction can be a desired quality value for the target configuration similar to the quality value generated in FIG. 26.
[0103]
[0146] 27 shows an example of an indirect prediction mode system according to some embodiments. In indirect prediction mode, two sub-modules may be used: a prediction network 2106 and a proxy encoding quality calculation engine 2702.
[0104]
[0147] The features of segment x 2602 are received by the prediction network 2106 and the proxy encoding quality calculation engine 2702. The proxy encoding quality calculation engine 2702 may use the proxy encoding configuration to calculate a proxy quality value based on the current target resolution and the current target bitrate. In some embodiments, the actual encoding of the video may be used to generate the proxy encoding features. An encoder may receive frames of segment x, encode segment x, and output an encoded bitstream. The proxy encoding features may be determined by characteristics of the encoded bitstream. The proxy encoding quality calculation engine 2702 may generate a proxy quality value at the proxy encoding point.
[0105]
[0148] FIG. 28 shows an example graph 2800 of proxy quality values and target quality values according to some embodiments. Line 2808 represents the proxy quality value, and line 2810 represents the target quality value. The proxy encoding values are at points where quality values are predicted from the proxy encoding features and are shown at 2806-1, 2806-2, and 2806-3. The encoder may use settings such as fast encoding presets to generate proxy encoding results at the proxy encoding points. The prediction network 2106 can then use the features from the proxy encoding to generate the corresponding proxy quality value. However, the proxy encoding points may not be enough points to generate the required points for the target quality value. Therefore, the proxy encoding quality calculation engine 2702 may generate additional proxy encoding values to generate the target encoding value.
[0106]
[0149] To generate additional proxy encoding values, the proxy encoding quality calculation engine 2702 may generate estimated proxy encoding values at other bit rates. For example, the estimated proxy quality value 2802 is generated based on the actual proxy quality value. In some examples, an interpolation or fitting algorithm may estimate the estimated proxy quality value in 2802 for another bit rate. As shown, the estimated proxy quality value 2802 may use the values of the proxy quality values 2806-2 and 2806-3 to determine the value of the estimated proxy quality value 2802, such as by estimating the value of the estimated proxy quality value 2802 at a bit rate between the proxy quality values 2806-2 and 2806-3. The calculation engine 2702 may generate other proxy quality values in a similar manner. The proxy encoding quality calculation engine 2702 outputs the proxy encoding point quality value to the combiner 2704.
[0107]
[0150] In addition to the proxy encoding quality calculation, the prediction network 2106 is also configured to generate a quality offset for the bitrate. For example, the prediction network 2106 may be trained to determine a quality offset based on input features for the segment. The quality offset may be a prediction of the difference between the target configuration and the proxy encoding configuration. For example, the quality offset may estimate the difference between a proxy quality value and a target quality value, which is estimated based on the difference between the proxy encoding configuration and the target configuration. The prediction network may be trained differently to predict an offset instead of a target quality value.
[0108]
[0151] 28, the prediction network 2106 can predict an offset, which is the difference between point 2802 and the target coding point. For example, the prediction network 2106 can predict the offset of the quality value for the target bit rate associated with the proxy quality value at 2802. The prediction network 2106 then outputs the offset for each bit rate associated with the proxy quality value to the combiner 2704.
[0109]
[0152] The combiner 2704 may combine the proxy quality values and the quality offsets. For example, for each proxy quality value, a respective quality offset may be received. The combiner 2704 may then combine the proxy quality values with the associated quality offset values to generate a target quality value, as shown at 2804 in FIG. 28 . In some embodiments, the combiner 2704 may add the quality offsets to each proxy quality value to generate the target quality value. Other methods of combining the offsets, such as using the offset as a multiplier or subtracting the offset, may also be understood. The combiner 2704 then outputs a target quality value for the target configuration. The above process may be performed for each target configuration to generate a target quality value.
[0110]
[0153] The direct prediction mode and the indirect prediction mode may be performed alternatively, such as only one of the modes being performed to generate a quality value. For example, the indirect mode may be determined to be more accurate when certain characteristics of the video are encountered, and the direct mode may be more accurate when certain characteristics are encountered. In some examples, the direct mode may be more accurate when simple content such as a cartoon is being encoded. However, when a movie is being encoded with more complex content, the indirect mode may be more accurate. The indirect mode may be more accurate because the proxy encoding may use the actual encoding to generate a quality value. Then, the prediction of the target quality value may be based on several quality values from the actual encoding, and a straightforward prediction of the quality value may be difficult due to the complexity of the content.
[0111]
[0154] In other embodiments, the direct predicted quality value and the indirect predicted quality value may be used in combination. Figure 29 shows an example of using a direct quality value and an indirect quality value according to some embodiments. In some embodiments, the cross-validation engine 2902 may use the direct quality value and the indirect quality value to output a verified quality value. The verified quality value may be based on a combination of the direct quality value and the indirect quality value for the bit rate. For example, an average of the two values may be used. In other embodiments, the cross-validation engine 2902 may select one of the direct quality value or the indirect quality value to output. For example, a direct predicted measurement may be selected, or an indirect predicted measurement may be selected.
[0112]
[0155] In another example, the cross-validation engine 2902 can validate the measurement value, such as by comparing the direct quality value and the indirect quality value for the bit rate. If the difference between the direct quality value and the indirect quality value meets a threshold, such as being within a threshold of each other, the cross-validation engine 2902 can validate the quality value. Otherwise, the cross-validation engine 2902 cannot validate the quality value and can output an error. The cross-validation engine 2902 can also merge direct mode results and indirect mode results based on the bit rate range. For example, the cross-validation engine 2902 can use direct mode to predict quality values in lower bit rate ranges and indirect mode to predict quality values in higher bit rate ranges, because different prediction modes may have different advantages in different bit rate ranges. Furthermore, the cross-validation engine 2902 can use maximum or minimum values when merging values together, such as when averaging values.
[0113]
[0156] RD Map
[0157] After generating the target quality values for multiple resolutions and bit rates, the RD map generator 2108 generates an RD map. The RD map generator 2108 can use fitting or interpolation methods to link the quality values being generated. The RD map generator 2108 can generate rate-distortion curves for different resolutions.
[0114]
[0158] 30A shows a graph 3000 of a rate-distortion curve for a single resolution, according to some embodiments. Quality value points are shown, and the RD map generator 2108 uses a method to link the points as a curve 3004. Different fitting or interpolation methods can be used to draw the curve based on the quality values.
[0115]
[0159] 30B shows a graph 3002 illustrating an RD map according to some embodiments. The RD map may include rate-distortion curves for three resolutions: 1920x1080, 1280x720, and 640x360 at 306-1, 306-2, and 306-3, respectively.
[0116]
[0160] conclusion
[0161] Therefore, the rate-distortion curves and rate-distortion maps can be predicted using the target configuration without actually encoding each rate-distortion curve, which saves computational resources and time. Also, the rate-distortion curves can be predicted more accurately using the direct and indirect modes.
[0117]
[0162] system
[0163] Features and aspects disclosed herein may be implemented in conjunction with a video streaming system 3100 that communicates with multiple client devices over one or more communications networks, as shown in Figure 31. Aspects of the video streaming system 3100 are described merely to provide one example of an application for enabling distribution and delivery of content prepared in accordance with the present disclosure. It should be understood that the present technology is not limited to streaming video applications and may be adapted to other applications and delivery mechanisms.
[0118]
[0164] In one embodiment, a media program provider may include a library of media programs. For example, the media programs may be aggregated and provided through a site (e.g., a website), an application, or a browser. A user may access the media program provider's site or application and request a media program. The user may be limited to requesting only media programs provided by the media program provider.
[0119]
[0165] In system 3100, video data may be obtained from one or more sources, e.g., video source 3110, for use as input to video content server 3102. The input video data may comprise raw or edited frame-based video data in any suitable digital format, e.g., Moving Picture Experts Group (MPEG)-1, MPEG-2, MPEG-4, VC-1, H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), or other formats. Alternatively, the video may be provided in a non-digital format and converted to a digital format using a scanner or transcoder. The input video data may include various types of video clips or programs, e.g., television episodes, movies, and other content generated as primary content of interest to consumers. The video data may also include audio, or audio alone may be used.
[0120]
[0166] The video streaming system 3100 may include one or more computer servers or modules 3102, 3104, and 3107 distributed across one or more computers. Each server 3102, 3104, 3107 may include or be operatively coupled to one or more data stores 3109, such as databases, indexes, files, or other data structures. The video content server 3102 may access a data store (not shown) of various video segments. The video content server 3102 may provide video segments as directed by a user interface controller that communicates with client devices. As used herein, a video segment refers to a distinct portion of frame-based video data, such as a television episode, a motion picture, a recorded live performance, or other video content that may be used in a streaming video session to view the video content.
[0121]
[0167] In some embodiments, the video ad server 3104 may access a data store of relatively short videos (e.g., 10-second, 30-second, or 60-second video ads) configured as advertisements for particular advertisers or messages. The advertisements may be provided to advertisers in exchange for some type of payment, or may include promotional messages, public service messages, or some other information for the system 3100. The video ad server 3104 may serve video ad segments as directed by a user interface controller (not shown).
[0122]
[0168] The video streaming system 3100 may also include a pre-analysis optimization process 110 .
[0123]
[0169] The video streaming system 3100 may further include an integration and streaming component 3107 that integrates video content and video advertisements into streaming video segments. For example, the streaming component 3107 may be a content server or a streaming media server. A controller (not shown) can determine the selection or configuration of advertisements within the streaming video based on any suitable algorithm or process. The video streaming system 3100 may include other modules or units not shown in FIG. 31 , such as a management server, a commerce server, a network infrastructure, an advertisement selection engine, etc.
[0124]
[0170] The video streaming system 3100 may be connected to a data communications network 3112. The data communications network 3112 may include a local area network (LAN), a wide area network (WAN), such as the Internet, a telephone network, a wireless network 3114 (e.g., a wireless cellular telecommunications network (WCS)), or some combination of these or similar networks.
[0125]
[0171] One or more client devices 3120 can communicate with the video streaming system 3100 via the data communications network 3112, the wireless network 3114, or another network. Such client devices may include, for example, one or more laptop computers 3120-1, desktop computers 3120-2, “smart” mobile phones 3120-3, tablet devices 3120-4, network-enabled televisions 3120-5, or combinations thereof, via a router 3118 for a LAN, a base station 3117 for the wireless network 3114, or some other connection. In operation, such client devices 3120 may send and receive data or instructions to the system 3100 in response to user input or other input received from a user input device. In response, the system 3100 can provide video segments and metadata from the data store 3109 to the client device 3120 in response to a selection of a media program. The client device 3120 can use a display screen, projector, or other video output device to output video content from streaming video segments in a media player and receive user input to interact with the video content.
[0126]
[0172] Delivery of audio-video data from the streaming component 3107 to remote client devices via computer networks, telecommunications networks, and combinations of such networks can be implemented using various methods, such as streaming. In streaming, a content server continuously streams audio-video data to a media player component running at least partially on the client device, and the media player component can play the audio-video data simultaneously as it receives the streaming data from the server. Although streaming is described, other delivery methods can be used. The media player component can begin playing the video data immediately after receiving the first portion of data from the content provider. Traditional streaming technologies use a single provider that delivers a stream of data to a set of end users. Delivering a single stream to a large audience can require high bandwidth and processing power, and the provider's required bandwidth can increase as the number of end users increases.
[0127]
[0173] Streaming media can be delivered on-demand or live. Streaming allows for instant playback at any point within a file. End users can skip through a media file, start playback, or change playback to any point within the media file. Thus, end users do not have to wait for a file to download incrementally. Streaming media is typically delivered from a small number of dedicated servers with high bandwidth capabilities through dedicated devices that accept requests for video files and use information about the format, bandwidth, and structure of those files to deliver only the amount of data needed to play the video, at the speed required to play the video. Streaming media servers can also take into account the transmission bandwidth and capabilities of the media player on the destination client. The streaming component 3107 can communicate with the client device 3120 using control and data messages to adapt to changing network conditions as the video is played. These control messages can include commands to enable control functions such as fast-forwarding, fast-rewinding, pausing, or seeking to specific parts of the file at the client.
[0128]
[0174] Because the streaming component 3107 transmits video data only when needed and at the required rate, precise control over the number of streams served can be maintained. Viewers cannot watch high-data-rate video over a lower-data-rate transmission medium. However, a streaming media server (1) provides users with random access to video files, (2) allows monitoring of who is watching which video programs and for how long, (3) uses transmission bandwidth more efficiently because only the amount of data needed to support the viewing experience is transmitted, and (4) video files are not stored on the viewer's computer but are discarded by the media player, thus allowing for more control over the content.
[0129]
[0175] The streaming component 3107 may use TCP-based protocols such as Hypertext Transfer Protocol (HTTP) and Real-Time Messaging Protocol (RTMP). The streaming component 3107 can also deliver live webcasts and can multicast, allowing two or more clients to tune into a single stream, thus conserving bandwidth. Streaming media players may not rely on buffering the entire video to provide random access to any point in a media program. Instead, this is achieved using control messages sent from the media player to the streaming media server. Other protocols used for streaming are HTTP Live Streaming (HLS) or Dynamic Adaptive Streaming over HTTP (DASH). The HLS and DASH protocols deliver video over HTTP via a playlist of small segments, typically made available at various bitrates from one or more content delivery networks (CDNs). This allows the media player to switch both bitrate and content source for each segment. This switching helps compensate for network bandwidth fluctuations and infrastructure failures that may occur during video playback.
[0130]
[0176] Delivery of video content via streaming can be accomplished under a variety of models. In one model, users pay to view video programs, for example, by paying a fee for access to a library of media programs or a limited portion of a media program, or by using a pay-per-view service. In another model, widely adopted by broadcast television shortly after its inception, sponsors pay for the presentation of media programs in exchange for the right to present advertisements during or adjacent to the presentation of the program. In some models, advertisements are inserted at predetermined times, sometimes called "ad slots" or "ad breaks," within the video program. In streaming video, media players can be configured to prevent client devices from playing video without playing predetermined advertisements during designated ad slots.
[0131]
[0177] Referring to Figure 32, a schematic diagram of an apparatus 3200 for viewing video content and advertisements is shown. In selected embodiments, the apparatus 3200 may include a processor (CPU) 3202 operably coupled to a processor memory 3204, the processor memory holding binary-coded functional modules for execution by the processor 3202. Such functional modules may include an operating system 3206 for handling system functions such as input / output and memory access, a browser 3208 for displaying web pages, and a media player 3210 for playing videos. The memory 3204 may hold additional modules not shown in Figure 32, for example, modules for performing other operations described elsewhere herein.
[0132]
[0178] The bus 3214 or other communication components may support communication of information within the device 3200. The processor 3202 may be a dedicated or special-purpose microprocessor configured or operable to perform particular tasks in accordance with the features and aspects disclosed herein by executing machine-readable software code that defines those tasks. The processor memory 3204 (e.g., random access memory (RAM) or other dynamic storage device) may be coupled to the bus 3214 or directly to the processor 3202 and may store information and instructions executed by the processor 3202. The memory 3204 may also store temporary variables or other intermediate information during the execution of such instructions.
[0133]
[0179] A computer-readable medium in storage device 3224 is connected to bus 3214 and can store static information and instructions for processor 3202; for example, storage device (CRM) 3224 can store modules for operating system 3206, browser 3208, and media player 3210 when device 3200 is powered off, and from which the modules can be loaded into processor memory 3204 when device 3200 is powered on. Storage device 3224 can include a non-transitory computer-readable storage medium that holds information, instructions, or some combination thereof, for example, instructions that, when executed by processor 3202, configure or enable device 3200 to perform one or more operations of the methods described herein.
[0134]
[0180] A network communications (comm.) interface 3216 may also be connected to the bus 3214. The network communications interface 3216 may provide or support bidirectional data communications between the device 3200 and one or more external devices, such as the streaming system 3100, optionally via a router / modem 3226 and a wired or wireless connection 3225. Alternatively, or additionally, the device 3200 may include a transceiver 3218 connected to an antenna 3229, through which the device 3200 may communicate wirelessly with a base station for a wireless communications system or with the router / modem 3226. Alternatively, the device 3200 may communicate with the video streaming system 3100 via a local area network, a virtual private network, or other network. In another alternative, the device 3200 may be incorporated as a module or component of the system 3100 and communicate with other components via the bus 3214 or by some other modality.
[0135]
[0181] Device 3200 may be connected (e.g., via bus 3214 and graphics processing unit 3220) to display unit 3228. Display 3228 may include any suitable configuration for displaying information to an operator of device 3200. For example, display 3228 may include or utilize a liquid crystal display (LCD), a touchscreen LCD (e.g., a capacitive display), a light emitting diode (LED) display, a projector, or other display device to present information to a user of device 3200 in a visual display.
[0136]
[0182] One or more input devices 3230 (e.g., an alphanumeric keyboard, microphone, keypad, remote control, game controller, camera, or camera array) may be connected to bus 3214 via user input port 3232 for communicating information and commands to device 3200. In selected embodiments, input device 3230 may provide or support control over cursor positioning. Such cursor control devices, also referred to as pointing devices, may be configured as mice, trackballs, trackpads, touchscreens, cursor direction keys, or other devices for receiving or tracking physical movements and converting the movements into electrical signals indicative of cursor movement. A cursor control device may be integrated into display unit 3228, for example, with a touch-sensitive screen. The cursor control device may communicate directional information and command selections to processor 3202 and control cursor movement on display 3228. A cursor control device may have two or more degrees of freedom, for example, allowing the device to specify a cursor position in a plane or three-dimensional space.
[0137]
[0183] Some embodiments may be embodied in a non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, system, or machine. The computer-readable storage medium includes instructions for controlling a computer system to perform methods described by some embodiments. The computer system may include one or more computing devices. The instructions, when executed by one or more computer processors, may be configured or operable to perform those described in some embodiments.
[0138]
[0184] As used in this description and throughout the claims that follow, "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Also, as used in this description and throughout the claims that follow, the meaning of "in" includes "in" and "on" unless the context clearly dictates otherwise.
[0139]
[0185] The above description shows various implementations, along with examples of how aspects of some embodiments may be implemented. The above examples and embodiments should not be considered the only embodiments, but are presented to illustrate the flexibility and advantages of some embodiments as defined by the following claims. Based on the above disclosure and the following claims, other configurations, embodiments, implementations, and equivalents may be employed without departing from the scope of the invention as defined by the claims.
Claims
1. determining, by a computing device, feature values for a portion of the video and a target configuration, wherein the target configuration is associated with encoder parameters and includes a set of bit rates and a resolution; generating, by the computing device, a plurality of quality values for the set of bit rates and the resolution based on the feature values; generating, by the computing device, a representation of an association between bit rate and the plurality of quality values for the resolution; analyzing, by the computing device, the representation to determine a list of bit rates for the portion of the video; outputting, by the computing device, the list of bit rates to use for encoding the portion of the video using the resolution; and A method for providing the above.
2. The method of claim 1 , wherein the parameters of the encoder comprise settings for encoding the portion of video.
3. generating the plurality of quality values The method of claim 1 , comprising generating a quality value for the resolution for each bit rate in the set of bit rates.
4. The resolution comprises a plurality of resolutions, and wherein generating the plurality of quality values comprises: selecting a resolution from the plurality of resolutions; generating a quality value for the resolution for each bit rate in the set of bit rates; 2. The method of claim 1, further comprising: continuing to select new resolutions from among the plurality of resolutions until all resolutions from the plurality of resolutions have been processed; and generating the quality values for each bit rate from the set of bit rates for the new resolutions.
5. The method of claim 1 , wherein the representation of the relationship between bit rate and the plurality of quality values comprises a rate-distortion curve.
6. generating a plurality of quality values for said set of bit rates for a plurality of resolutions; The method of claim 1 , further comprising generating a plurality of representations of the association between bit rate and the plurality of quality values for the plurality of resolutions.
7. generating a map comprising the plurality of representations at the plurality of resolutions; Using the map to determine the list of bit rates. The method of claim 6 further comprising:
8. generating the plurality of quality values The method of claim 1 , comprising inputting the feature values into a predictive network to generate the plurality of quality values.
9. The method of claim 8 , wherein the feature values comprise one or more of computer vision features, spatial domain features, time domain features, frequency domain features, and proxy coding features.
10. The method of claim 9 , wherein the proxy encoding characteristics are based on an actual encoding of the video, wherein settings in the actual encoding are different from the parameters in the target configuration.
11. generating the plurality of quality values determining proxy encoding feature values from an actual encoding of the video; and inputting the proxy encoding feature values into a prediction network to generate a plurality of proxy quality values.
12. generating the plurality of quality values The method of claim 11 , comprising generating the plurality of quality values based on the plurality of proxy quality values.
13. generating the plurality of quality values based on the plurality of proxy quality values, generating a plurality of offset values using a prediction network, wherein the plurality of offset values are based on differences between settings in the actual encoding from the parameters in the target configuration; 13. The method of claim 12, comprising using the plurality of offset values and the plurality of proxy quality values to generate the plurality of quality values.
14. generating the plurality of quality values determining a first plurality of quality values using a first prediction method; determining a second plurality of quality values using a second prediction method, wherein the second prediction method predicts an offset for the proxy quality values determined based on a difference between settings in the actual encoding from the parameters in the target configuration; and comparing the first plurality of quality values to the second plurality of quality values to determine the plurality of quality values.
15. generating the plurality of quality values analyzing characteristics of the portion of the video; selecting one of a first prediction method and a second prediction method based on the characteristic, wherein the second prediction method predicts an offset to a proxy quality value determined based on a difference between settings in actual encoding from the parameters in the target configuration; and using the selected one of the first prediction method and the second prediction method to determine the plurality of quality values.
16. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by a computing device, determining feature values for a portion of the video and a target configuration, wherein the target configuration is associated with encoder parameters and includes a set of bit rates and a resolution; generating a plurality of quality values for the set of bit rates and the resolution based on the feature values; generating a representation of the association between bit rate and the plurality of quality values for the resolution; analyzing the representation to determine a list of bit rates for the portion of the video; and outputting the list of bit rates to use for encoding the portion of the video using the resolution.
17. determining, by the computing device, proxy encoding feature values from an actual encoding of a portion of the video using a proxy encoding configuration associated with a first parameter of the encoder, the proxy encoding configuration including a set of bit rates and a resolution; generating, by the computing device, a plurality of proxy quality values for a set of bit rates and resolutions based on the proxy coding feature values using a first prediction method; determining, by the computing device, a target feature value for the portion of the video, wherein the target feature value is based on a target configuration associated with a second parameter of the encoder; generating, by the computing device, a plurality of offset quality values for the set of bit rates and the resolution based on the target feature values using a second prediction method, wherein the second prediction method predicts an offset relative to a proxy quality value determined based on a difference between the first parameter and the second parameter; generating, by the computing device, a plurality of quality values using the plurality of proxy quality values and the plurality of quality offset values; A method comprising:
18. the plurality of quality values comprises a first plurality of quality values; determining a second plurality of quality values using a third prediction method that predicts quality values based on the target configuration; 18. The method of claim 17, further comprising comparing the first plurality of quality values to the second plurality of quality values to determine a quality value to output for the set of bit rates and the resolution.
19. Comparing the first plurality of quality values to the second plurality of quality values to determine a quality value to output includes: analyzing characteristics of said portion of video; and selecting one of the first plurality of quality values and the second plurality of quality values based on the characteristic.
20. Comparing the first plurality of quality values to the second plurality of quality values to determine a quality value to output includes:
20. The method of claim 18, comprising combining the first plurality of quality values and the second plurality of quality values based on a characteristic.